Title banner for the pre-ship rights check — teal node lines and a gate motif.

Key Points

  1. This is a US/EU-focused pre-ship rights check. If you build or ship in the US or EU, these are the rights basics that keep costly disputes away: copyright exceptions (US fair use, the EU DSM Directive's text-and-data-mining rules), the copyrightability of AI-generated work (US human authorship; the EU originality test, with AI-output ownership still unresolved — the EU AI Act's transparency duties are a separate layer), and OSS license tiers.
  2. The specifics differ by jurisdiction, but the method travels: fix your target jurisdiction first, check the copyright exceptions, classify OSS by SPDX tier, and document the human contribution to anything AI-generated.
  3. Some questions are settled (in 2026 the US Supreme Court let the human-authorship rule stand); others are still in active litigation (Andersen v. Stability AI; NYT v. OpenAI) — which is exactly why records and conservative defaults matter now. Building in or for Japan? The Japanese edition covers the Japan-law version.
  4. Out of scope (for a lawyer / IP specialist, or the related pieces): case-by-case legal calls / OSS licenses in depth (OSS licenses) / scraping, external assets, AI service terms, and AI-output commercialization / local law outside the US/EU.

この記事の要点(日本語版はこちら)

  1. 本記事は 米国・EU 法を背骨にした「国際版」 です(日本語版の翻訳ではなく、対象法域が異なる対の記事)。
  2. 日本での公開・配布が中心の方は、日本法を主役にした 日本語版 を参照してください。

Written: 2026-05 / Jurisdictions in scope: United States & EU (Japan is covered in the Japanese edition) / last_updated: 2026-05-23 Changelog: 2026-05-23 first published as an international (US/EU) edition

This article is educational material, not legal advice

Copyright, license agreements, and the copyrightability of AI-generated works carry real consequences — statutory damages, injunctions, regulatory fines, and contract termination. This article is an educational overview of US and EU law; it does not guarantee the legality of any specific case, and it is not a substitute for advice on your own facts. Out of scope: case-by-case legal determinations / contract negotiation / litigation strategy / jurisdictions other than the US and EU. For real decisions, consult a qualified attorney in the relevant jurisdiction / your legal department / each license's official text. All information is provided "AS IS," without warranty of any kind as to accuracy, completeness, or currency. To the maximum extent permitted by law, the author and YATA-NODE accept no liability for any loss or damage arising from the use of or reliance on this article — whether or not you consulted a professional. Use it at your own risk.

About this series: this is the international edition of the entry point to YATA-NODE's "rights basics" series. The companion Japanese edition covers the same ground under Japanese law. OSS specifics, the boundaries of scraping, and the commercialization of AI-generated works each have their own piece (no fixed cadence; revised as the law moves).

A few terms (for newcomers):

  • OSS (Open Source Software): source code published under a license that lets anyone use, modify, and redistribute it within the license terms.
  • SPDX identifier: an industry-standard short name that uniquely identifies a license (e.g., MIT / Apache-2.0 / AGPL-3.0).
  • TDM (text and data mining): automated analysis of large volumes of text/data — the legal category most "training data" debates fall under in the EU.
  • robots.txt: a file by which a site tells automated crawlers which pages they may fetch.

A 30-minute check up front beats a six-figure problem later

"Use it first, deal with problems later" almost never pays off. Three different kinds of risk to calibrate the scale:

  • Statutory damages (US): US copyright law lets a rightsholder elect statutory damages instead of proving actual loss — $750 to $30,000 per work, rising to up to $150,000 per work for willful infringement (17 U.S.C. §504(c)). "Per work" is the part that hurts: a few dozen infringed items compound fast. (Outside the US the numbers differ but can be just as large — Japan's "Fast Movie" case ended in a ¥500M award.)
  • Regulatory fines (EU): Clearview AI scraped images to build a face-recognition database and was fined €20 million by Italy's Garante under the (2022); Greece's DPA and France's CNIL issued €20M fines of their own, and the UK ICO a £7.5M penalty (overturned for lack of jurisdiction in 2023, then revived and remitted to the First-tier Tribunal by the Upper Tribunal in 2025, with a further appeal pending). Scraping public data is not consequence-free.
  • Transaction risk: US/EU enterprise contracts routinely require "compliance with OSS licenses" as a representation and warranty, and M&A due diligence expects an (software bill of materials = a list of the OSS and other components you ship). A copyleft source-disclosure obligation propagating into your proprietary code (GPL/, sometimes loosely called "contamination") can stall or sink a deal.

The takeaway: a 30-minute rights check up front sharply reduces, and surfaces early, risk that can run from thousands to millions of dollars. It doesn't eliminate the risk, but it changes the odds.

Copyright exceptions — US fair use and the EU's text-and-data-mining rules

Fix your jurisdiction first. Copyright is territorial: the exception that saves you in one country may not exist in another. Decide where you publish and distribute before you lean on any exception.

United States — fair use (17 U.S.C. §107)

Fair use is a fact-specific defense weighed after the fact, not a checkbox you tick in advance. Courts weigh four factors:

  1. the purpose and character of the use (including whether it's commercial, and how "transformative" it is);
  2. the nature of the copyrighted work;
  3. the amount and substantiality of the portion used; and
  4. the effect on the potential market for or value of the work.

The common misreading is "non-commercial / research / training is automatically fair use." It isn't — no single factor is decisive, and the fourth (market effect) often does the heavy lifting. Treat fair use as a defense you might raise, not a permission you already hold.

European Union — the DSM Directive's TDM exceptions (Arts. 3 & 4)

The EU handles "training data" mostly through text-and-data-mining (TDM) exceptions in Directive (EU) 2019/790:

  • Article 4 is the general TDM exception: with lawful access, you may reproduce/extract works for TDM — unless the rightsholder has expressly reserved that use by appropriate means. For content made available online, that reservation must be machine-readable (robots.txt-like signals, metadata, or other recognised rights-reservation mechanisms — the exact implementation is still a moving practical/legal question). Either way, in the EU an opt-out actually has teeth.
  • Article 3 is narrower: TDM for scientific research by research and cultural-heritage institutions, and it cannot be overridden by a rightsholder opt-out the way Art. 4 can.

(Japan takes a third route — a broad "non-enjoyment purpose" exception in Article 30-4 of its Copyright Act. If that's your market, see the Japanese edition.)

What's still being litigated

Whether US fair use actually covers training generative AI on copyrighted works is not settled — it is being fought right now. Two cases worth watching (status as of 2026-05; litigation pending, no final ruling on the merits):

  • Andersen v. Stability AI (N.D. Cal., No. 3:23-cv-00201-WHO): a class of visual artists led by Sarah Andersen. An August 2024 order let the core copyright claims proceed into discovery while dismissing the DMCA/CMI (§1202) claims. Still in discovery.
  • NYT v. OpenAI / Microsoft (S.D.N.Y., No. 1:23-cv-11195, consolidated in MDL 25-md-3143): an April 2025 ruling kept the core copyright-infringement case moving forward while narrowing some collateral claims (e.g., dismissing certain DMCA §1202 claims without prejudice). It's in active discovery, with no trial date or settlement publicly reported.

The point isn't to predict the outcome — it's that the uncertainty is real, which is exactly why keeping records and defaulting to the conservative option is the rational move today.

OSS, the minimum — four tiers by how widely the source-disclosure obligation reaches, narrow (MIT/Apache) to wide (AGPL)

The SPDX-based workflow travels well internationally — even though enforcement and interpretation remain jurisdiction-specific — so this section applies whether you ship in the US or the EU. Organize licenses into four tiers by source-disclosure scope, from narrow to wide, and check which tier something is in — by its — before you incorporate it into your product.

A spectrum chart arranging OSS licenses by source-disclosure scope, narrow to wide, in four tiers; public-domain-equivalent on the left, AGPL on the right.
TierRepresentative licensesCore obligationThe solo / side-project call
1. Public-domain-equivalentCC0, UnlicenseNo copyright notice required; almost unrestrictedVirtually no source-disclosure obligation (note that civil-law systems — e.g., Germany, France, Japan, as opposed to common-law systems — may not allow a full waiver of copyright)
2. PermissiveMIT, BSD, Apache 2.0Retain the copyright notice (Apache 2.0 also: keep NOTICE + explicit patent grant + retaliation clause)The smallest disclosure obligation for commercial or personal use. MIT leaves more room for an implied-patent dispute than Apache
3. Limited-scope copyleft (weak copyleft)MPL 2.0, LGPLSource-disclosure duty per modified file (MPL) or per library (LGPL)Can be mixed with proprietary code. Statically linked LGPL often trips you up on the duty to provide relinkable materials
4. Whole-work / network copyleft (strong / network copyleft)GPL-3.0, AGPL-3.0Disclose the source of the whole derivative work. AGPL can add obligations when a modified version is offered over a network (v3 §13)Even in a one-person SaaS, depending on how you modify and combine it, you may owe end users the corresponding source. The scope depends on the combination form — check the official text / a specialist. Minimizable if used purely internally

Counter-reading / caution: "OSS = all free" is a misreading. "Source-available but not OSS" licenses exist — the Sustainable Use License (n8n), (MongoDB), (HashiCorp's Terraform, etc.) — and they can put you in breach when you resell a commercial SaaS. License terms also move: Redis (8.0, May 2025) and Elasticsearch (Aug 2024) re-added AGPL-3.0 (OSI-approved) after earlier shifts to SSPL and similar. So once the SPDX identifier tells you something may not be OSI-approved, go to the official license page and read the current full terms.

Copyright in AI-generated work — the US and the EU, with Japan for contrast

"Who owns what ChatGPT / Claude / Midjourney produced?" Jurisdictions answer differently, so fix the geography where you publish or distribute first, then read the relevant row.

A concept diagram with 'copyrightable if there is human creative contribution' at the center and the US, the EU, and Japan around it.
JurisdictionTest for copyrightabilityStatus (as of 2026-05)Practical response
United StatesHuman authorship is an essential requirement for protectionThe Supreme Court denied certiorari in Thaler v. Perlmutter on 2026-03-02, leaving the D.C. Circuit's 2025 ruling in place: a work authored entirely by a machine (no human author) can't be registered. How much human input is enough remains open.Don't claim copyright in AI-only output; for AI-assisted work, keep evidence of your creative choices, and handle third-party rights, ToS, and warranties separately
European UnionThe author's own intellectual creation (CJEU, Infopaq, C-5/08)Infopaq sets the originality test; it doesn't directly decide who owns AI output, and pure machine output stays uncertain across Member States. The AI Act's Art. 50 duties are a separate layer (below) — transparency, not ownershipSame instinct: a purely machine-made output is hard to own; document the human contribution
JapanHuman creative contribution, considered as a whole (Agency for Cultural Affairs guidance)Current official guidance, not a court ruling — covered in the Japanese editionRecord prompts + edits + selection decisions

The shared instinct: copyright protection is strongest where identifiable human creative choices shape the final expression; a pure "auto-generate → use as-is" output is hard or uncertain to protect — clearest in the US after Thaler, and not yet directly resolved for AI-only output in the EU or Japan. Practically, you can't stop others from reusing unprotected output, so don't put raw AI output at the core of your originality.

The EU is a separate layer (and it's moving). Art. 50 imposes transparency duties — not copyright. In brief: AI systems must let people know they're interacting with an AI; providers must mark synthetic audio/image/video/text so it's machine-readable as AI-generated; and deployers must disclose deepfakes and AI-generated text published to inform the public — subject to carve-outs (obvious cases, standard editing assistance that doesn't materially alter the input, law-enforcement uses, and editorially controlled public-interest text). Timing as of 2026-05: most Art. 50 duties apply from 2026-08-02. Under the Digital Omnibus political agreement (trilogue 6 May 2026; not yet adopted or published in the Official Journal), Art. 50(2) watermarking duties for AI systems placed on the market before 2 August 2026 would get a four-month grace period until 2 December 2026 — systems newly placed on the market are still expected to comply from 2 August 2026. Because this is not yet in force, treat it as proposed/politically agreed and confirm the adopted text before you rely on it.

Two common misreadings:

① "The EU AI Act decides whether I infringe copyright" — no; transparency duties and copyright ownership are different layers.
② "If AI made it there's zero copyright / if I wrote a prompt it's all mine" — both are oversimplifications.

Counter-reading / uncertainty: the lines here are being drawn case by case. The US litigation above, and EU member-state rulings, can shift the picture — and even how much human input creates authorship is unresolved (e.g., the pending US Allen v. Perlmutter, over a heavily prompt-engineered image). Treat this as an area to re-check every 6–12 months.

Summary — a pre-ship rights checklist

We saw the scale with real examples: statutory damages (up to $150,000 per work in the US for willful infringement), GDPR fines (€20M for Clearview AI), and deal-breaking OSS copyleft obligations. The work to reduce that risk is four steps: (a) fix your activity geography, (b) check whether the copyright exceptions apply to your case (US fair use is a defense, not a permission; EU Art. 4 turns on lawful access + opt-out), (c) classify OSS into four tiers by SPDX identifier and cross-check the linking + distribution form, and (d) for AI-generated work, keep records given the unsettled state of the law.

Paste the list below into your LICENSE.md / dev notes / PR description and check it before you ship.

A flow diagram of the four pre-publish steps (fix geography → check copyright → classify OSS → record AI) leading to the goal 'Publish.'
□ 1. Fixed my activity geography (US / EU / both / global)
□ 2. Obtained third-party assets from official sources and identified the license type
□ 3. Checked commercial use (not NC) and modifiability (not ND)
□ 4. Checked the OSS SPDX identifier and identified which of the four tiers it is in
□ 5. Mapped the combination form to the right obligation: linking/distribution (LGPL/GPL) and network use of a modified version (AGPL §13)
□ 6. Inventoried transitive dependencies via an SBOM
□ 7. For TDM / scraping in the EU, confirmed lawful access and honored any machine-readable opt-out (DSM Art. 4)
□ 8. If AI-generated content is included, recorded the human creative contribution (prompts + edit history + selection decisions)
□ 9. Recorded the acquisition date + license URL + basis for the decision in internal notes
□ 10. If unsure, consulted a qualified attorney in the relevant jurisdiction in advance (at your own risk)

Don't forget rights beyond copyright: the above centers on copyright / OSS / AI-generated work. For third-party assets and public services, terms of service (ToS), trademarks, likeness/publicity rights, and personal data (GDPR in the EU; CCPA/CPRA in California and similar US state laws) are separate pitfalls. Terms of service are covered in AI service terms; trademarks, likeness rights, and personal data are out of scope here (consult a specialist or the official sources).

This is the international entry point to the "rights basics" series. Each theme is covered in depth in its own piece:

References

Centered on primary materials. A trailing (primary) marks a primary source; (supporting) marks commentary / a docket used to locate primaries.

Statutes & regulations

Official guidance

Case law & penalties

On AI assistance: Starting from points the author already knew, the author used LLMs (Claude by Anthropic, with a second, independent model cross-checking the legal facts) to research, organize, and summarize, then verified the facts against the primary sources cited above and revised accordingly. Entries marked (primary) point to statutes, official guidance, court opinions, or regulators; (supporting) entries are dockets/commentary used to locate the primary materials. For the "not legal advice" note and what is out of scope, see the disclaimer at the top.

About the author

More than 20 years of electrical and software development — from control engineering at a major electronics manufacturer — plus about 10 years of solo development. Across hardware and software, and across enterprise and individual work, I publish the basics that "become a risk if you don't know them," and I plan to cover practical ways to use AI as well. More at About this blog.