AI, Copyright, and the Licensing Gap
Learning Is Not Theft, but New Technology Exposes Old Assumptions
Few topics in technology generate more passionate debate today than artificial intelligence and copyright.
Authors, artists, musicians, software developers, and other content creators are asking whether AI companies have built commercial products on the uncompensated work of others. AI companies and their defenders often respond that their systems are not storing and redistributing those works. They are learning patterns from them.
The argument is usually reduced to one emotionally loaded question:
Is AI training stealing?
I believe that question is too simplistic. Worse, it may be keeping us from having the discussion that actually matters.
The real issue may not be whether AI training constitutes theft. The real issue may be that artificial intelligence has exposed a gap in copyright and licensing frameworks that were never designed to address machine learning in the first place.
That is not the same as saying creators have no legitimate grievance. It means we are trying to resolve a new category of use with agreements, assumptions, and language created for a different technological era.
Every Creator Learns from Previous Creators
Before discussing how artificial intelligence learns, it is worth examining how human learning works.
Every software engineer has studied code written by someone else. Every author has read books written by previous generations. Every musician has listened to thousands of songs. Every artist has observed techniques, styles, and approaches developed by other artists.
None of us creates in a vacuum.
A software developer who studies design patterns and then applies them to a business problem is not stealing. A novelist influenced by Tolkien, Sanderson, Asimov, or Hemingway is not committing copyright infringement simply because those writers shaped how that novelist thinks about storytelling. A musician inspired by jazz does not owe royalties every time that influence appears in a chord progression.
Hunter S. Thompson, known for Fear and Loathing in Las Vegas, reportedly typed out entire novels by F. Scott Fitzgerald and Ernest Hemingway to study their sentence structures. He copied their words as an exercise, not to publish those words as his own. The purpose was learning—and the result was not another Hemingway or Fitzgerald, but a writer with one of the most recognizable voices of the twentieth century.
Learning from existing work is not plagiarism.
Plagiarism occurs when someone presents another person’s expression or ideas as their own. Copyright infringement occurs when protected expression is used in a way the law does not permit. Influence, education, and learning are not automatically either one.
Human creativity has always been cumulative. We learn, adapt, combine, reject, refine, and build upon what came before us. Even our attempts to create something new are shaped by everything we have previously encountered.
From that perspective, the statement that an AI system learns patterns from existing works does not immediately settle the ethical question in either direction. Learning itself has never been the problem.
The question is where learning ends and licensed usage begins.
The Sanderson Question
Consider a simple example.
Suppose Brandon Sanderson releases a new novel.
I cannot legally obtain a copy merely because I want to read it. I must acquire access through an authorized channel. I might purchase the book, borrow it from a friend, check it out from a library, or obtain it through another licensed service.
Once I have lawful access, I may read it.
I may study the way Sanderson constructs a magic system. I may examine his pacing, his worldbuilding, or the way he manages multiple character arcs. Years later, something I learned may influence the way I approach my own work.
What I may not do is copy the novel, redistribute it, or publish it under my own name.
That distinction has served creators and consumers reasonably well: access to a work permits consumption and learning, while copyright continues to protect the author's expression and defined rights.
The companies developing large language models are bound by the same copyright laws as you and I. Artificial intelligence does not exist outside those protections simply because the technology is new, and calling the process “training” does not automatically exempt it from the rules governing copyrighted work. If an AI company copies, distributes, or reproduces protected expression in a way the law does not permit, it should be held accountable just as any person or corporation would be. The unresolved question is not whether copyright law applies. It is whether using a lawfully acquired work to train an AI system falls within existing rights—or represents a separate use those rights were never written to address.
Artificial intelligence introduces a question previous generations never had to answer:
Does permission to access and read a book also include permission to use that book for machine learning?
The Sanderson example is not meant to suggest that a person reading one novel and a company training a foundation model on millions of books are identical. They clearly are not. It establishes the principle underneath the dispute: lawful access and permitted use are related, but they are not necessarily the same thing.
A purchased copy gives me the right to read the book. It does not grant me the right to print and sell five thousand copies. A streaming subscription lets me listen to a song. It does not grant me the right to place that song in a commercial. Possessing or accessing a work has never automatically granted every possible use of it.
That leads to the more precise question:
Is machine learning already included among the permitted uses of a lawfully accessed work, or is it a distinct use that should require a distinct right?
That is the licensing gap.
The Licensing Gap
Copyright notices and licensing agreements have long addressed familiar actions: copying, reproduction, distribution, transmission, adaptation, public performance, and commercial republication.
What they have rarely addressed—particularly in works and agreements created before modern generative AI—is machine training.
They generally do not tell us:
- whether AI training is permitted;
- whether commercial and nonprofit training should be treated differently;
- whether training requires a separate license;
- whether the creator is entitled to compensation; or
- whether permission depends on how the work was acquired and how it is retained.
The reason is not mysterious. Most of these frameworks were written before large language models and generative systems became part of ordinary commercial life. The people writing the agreements were not ignoring AI training. They were not contemplating it.
Now both sides point to frameworks that do not explicitly address the new use and claim the answer is obvious.
Creators say, "I never gave you permission to train on my work."
AI companies say, "The law does not require a separate training license."
Both positions reveal the same underlying problem: the rights and obligations were not clearly defined before the technology made the question economically important.
There is also an important first step that cannot be skipped. Before asking whether training is a permitted use, we have to ask whether the work was lawfully obtained. A licensing model cannot sanitize material taken from an unauthorized source. Acquisition and usage are separate questions, and both matter.
That gives us a clearer framework:
- Acquisition: Was the work obtained or accessed lawfully?
- Training: Does that lawful access permit use of the work for machine learning?
- Output: Does the system reproduce protected expression or generate work that infringes upon it?
Public debate often collapses all three into the word theft. That may express the creator's anger, but it does not give us a useful framework for resolving the issue.
We Have Seen This Pattern Before
The software industry has encountered analogous problems before.
Traditional software licenses developed in a world where software was installed on physical machines owned or controlled by the customer. Terms such as per user, per device, and per installation were comparatively easy to understand.
Then cloud computing changed the architecture.
Software could run on someone else's infrastructure. A single system could serve thousands of customers. Applications could be accessed without being installed on the user's machine at all. Infrastructure could appear and disappear in minutes.
Suddenly, old licensing assumptions produced new questions:
- What constitutes an installation?
- Who owns or controls the machine?
- What qualifies as hosting?
- How should usage be measured in a multi-tenant environment?
- Does a traditional license permit a company to offer the software as an online service?
The answer was not to declare licensing obsolete or to stop the cloud.
Licensing evolved.
New definitions, restrictions, and business models developed around Software as a Service, Platform as a Service, and Infrastructure as a Service. Vendors revised terms. Customers negotiated different rights. Markets developed ways to price access, usage, capacity, and commercial hosting.
The comparison is not that cloud licensing and copyright law are identical. They are not. The comparison is the pattern:
Technology creates a use case the existing agreements did not clearly anticipate. Ambiguity creates conflict. Licensing eventually evolves to address the new reality.
Artificial intelligence may be the next iteration of that pattern.
Scale Changes the Discussion
If AI learning and human learning share anything, they do not share scale.
A human may read hundreds or thousands of books over a lifetime. A foundation model may process millions of works during development. A human learns gradually through education and experience. An AI company can analyze enormous collections in a comparatively short period and use the resulting system across millions of commercial interactions.
That difference matters.
An individual reading a novel feels like education.
A corporation training a commercial system on millions of novels feels like industrial use.
The difference is not merely emotional. Scale changes the economics. A single reader's influence on a market is negligible. A model trained on the accumulated work of an entire creative industry may become a product capable of competing within that industry.
Creators are therefore not unreasonable when they ask whether "learning" is an adequate description of the entire transaction. AI companies are not necessarily wrong when they argue that identifying patterns is different from storing and reselling complete works.
Both statements can be true, and neither resolves the rights question.
The disagreement is not simply about whether learning occurs. It is about what rights should be required before learning can take place at industrial scale for commercial purposes.
That is why the human-learning analogy is useful but insufficient. It explains why learning from existing work is not inherently unethical. Scale explains why machine learning may still deserve its own licensing category.
A Better Question
The debate over whether AI training is theft forces everyone into one of two camps.
If training is theft, the technology is built on wrongdoing and should be restricted accordingly.
If training is merely learning, creators are expected to accept the use of their work without further discussion.
Neither position leaves much room for a workable solution.
A licensing approach asks a better question:
What rights should govern the use of copyrighted work for AI training, and what is a fair price for those rights?
That question does not presume that all training is infringement. It also does not presume that every work available to a human reader is automatically available for unlimited machine use.
It recognizes AI training as a use that can be defined, granted, withheld, limited, and priced.
The Future May Be Better Licensing
Future licenses could address AI usage directly rather than leaving creators and technology companies to argue over silence.
| Usage right | Possible treatment | | --- | --- | | Human reading | Permitted | | Personal study | Permitted | | Educational use | Permitted or conditional | | Redistribution | Prohibited without permission | | Commercial republication | Separate license | | Search indexing | Permitted or conditional | | Nonprofit AI research | Permitted or conditional | | Commercial AI training | Separate license | | AI training with creator compensation | Permitted under defined terms |
The exact categories would vary by medium and market. A novelist, a photographer, a software developer, and a musician may not need identical arrangements. Commercial training may require different terms from academic research. A model that retains access to source material may require different terms from one trained on material that is later removed.
The point is not that this table contains the final answer.
The point is that the rights should be explicit.
Creators should understand how their work may be used. AI companies should understand what they may acquire, what requires permission, and when compensation is owed. Publishers and licensing organizations could offer collective mechanisms so that a company does not have to negotiate separately with millions of individual rights holders. Smaller AI developers would need access to structures that do not reserve innovation exclusively for the wealthiest corporations.
There would still be difficult questions. How should compensation be calculated? Should payment reflect the number of works, the frequency of use, the contribution of a work to a model, or the revenue generated by the finished product? How could usage be audited? What happens when a creator opts out? How should previously trained models be handled?
Licensing would not make those questions disappear.
It would finally give us a framework in which to answer them.
Most importantly, it would shift the debate from accusation to negotiation:
Not: Is AI training theft?
But: What is a fair price for AI training rights?
That is a far more productive conversation.
Conclusion
Artificial intelligence did not create the practice of learning from others. Humanity has always done that. Every creator is shaped by previous creators, and every new idea carries traces of what came before it.
What AI created was a new use case operating at a scale never previously imagined.
That scale exposed assumptions embedded within copyright and licensing frameworks written for a different era. Lawful access does not necessarily answer the question of permitted machine use. Silence in an old agreement does not produce clarity for a new technology.
The most important question is not simply whether AI learns from copyrighted works.
Humans do that every day.
The more important question is whether our existing frameworks adequately define the rights granted when a work is acquired and then used for machine training.
Increasingly, the answer appears to be no.
Whenever technology exposes a gap between existing agreements and new capabilities, history suggests the solution is rarely to stop the technology. It is to define the new use, establish the rights surrounding it, and allow creators and innovators to negotiate from something clearer than assumption.
Artificial intelligence may not be forcing us to abandon copyright.
It may be forcing us to modernize licensing—just as the cloud once forced the software industry to reconsider what it meant to install, access, and use software at all.