Skip to main content
Anthropomorphic Provenance Laundering

Anthropomorphic Provenance Laundering

Written by Michael Blevins on .
  1. Anthropomorphic Provenance Laundering
  2. Human Reclaimed Intelligence
  3. Stop Calling It Training
  4. Fair Use on the Way In
  5. The Machine Becomes Human
  6. Who Owns the Intelligence?

The language of artificial intelligence makes machines sound human while the people who supplied the intelligence disappear.

Part One in a series on artificial intelligence, human authorship, and provenance. Artificial intelligence did not begin with intelligence.

It began with access.

Access to books, photographs, illustrations, music, journalism, software, research, conversations, business methods, and nearly every other form of human expression placed within reach of a machine. Then we gave the machine human words.

We said it was trained. We said it learned. We said it could understand, reason, imagine, and create.

Those words do more than make a complicated technology easier to explain. They change how we understand what happened. They make industrial ingestion sound like education. They make extraction sound like personal development. They make the machine appear more human while the humans who supplied its capabilities disappear.

I call this pattern:

Anthropomorphic Provenance Laundering

Anthropomorphic: We assign human traits and verbs to a machine to obscure its mechanical nature.

Provenance: the origin, history, and source of the intelligence are what give the output its actual value.

Laundering: Human work passes through a commercial system so it emerges separated from the people who created it—rendering it legally and ethically “clean” for corporate profit.

The intelligence did not appear from nowhere.

Its provenance was removed from view.

Money laundering does not create the original value. It obscures where that value came from and allows it to reappear with a different history attached to it.

Anthropomorphic Provenance Laundering performs a similar transformation with human intelligence. Human expression enters carrying names, ownership, experience, culture, methods, relationships, and memory. It emerges as machine capability: commercially valuable, enormously scalable, and increasingly difficult to trace back to the people who supplied it.

The comparison is not that every use of artificial intelligence is illegal. The comparison is that the value survives while the origin becomes obscured.

The Internet Was Never Just Data

The internet is often described as a massive collection of publicly available data.

It was never merely data.

It was people explaining what they knew. Artists showing what they could see. Photographers documenting the world. Programmers solving problems. Writers developing ideas. Musicians sharing sound. Teachers creating lessons. Businesses revealing methods earned through years of work. Communities recording their histories before they disappeared.

Calling all of that material “data” is the first act of separation.

A painting becomes data. A novel becomes data. A photograph becomes data. A person’s life’s work becomes data.

Once the human origin is linguistically removed, industrial ingestion becomes much easier to rationalize.

Even the U.S. Copyright Office, while using “data” and “dataset” as shorthand, stresses that the works involved are not merely data in the ordinary sense. They contain creative expression and protected human authorship. (U.S. Copyright Office)

Language models are not built merely by collecting isolated facts. They process how words are selected and arranged across sentences, paragraphs, and entire documents. The Copyright Office describes those relationships as central to linguistic expression. It also notes that image models use curated aesthetic images precisely because those images help them generate aesthetic outputs. (U.S. Copyright Office)

The books, paintings, photographs, articles, and illustrations were not included by accident.

They were useful because they contained expression.

The system may not preserve every source in an immediately recognizable form, but its capability was created by processing those sources.

The fact that provenance becomes difficult to identify does not mean the provenance ceased to exist.

Let’s Stop Calling It Training

I understand that training is the accepted technical term.

That does not make it a neutral one.

We train employees. We train athletes. We train apprentices. We teach students.

Human beings learn through consciousness, memory, experience, character, responsibility, limitation, and judgment. We retain imperfect impressions filtered through our histories, personalities, relationships, and understanding of the world.

A machine does something materially different.

It can copy and process more human expression than any individual could encounter in thousands of lifetimes. It can analyze those works nearly instantaneously and produce competing material at superhuman speed and scale.

The Copyright Office has directly rejected the idea that machine learning and human learning are equivalent for copyright analysis. A student could not copy every book in a library and defend the copying merely by calling it education. Copyright law should not grant greater freedom to copy simply because a computer performs the act. (U.S. Copyright Office)

The word training hides that difference.

It turns a commercial system into a student. It turns creators into teachers who never agreed to teach. It turns copied material into lessons. Once the machine has supposedly “learned,” the resulting capability appears to belong entirely to the machine and the company that controls it.

A common defense is that a language model does not store and reproduce an exact copy of every work it processes.

But the system was designed to retain something. If nothing survived the process, no capability could have been gained from it.

It retained relationships between words. It retained sentence structures, patterns of explanation, methods of argument, coding conventions, visual relationships, stylistic associations, and countless other representations derived from human work.

The question cannot be limited to:

Did the machine reproduce this exact work?

We must also ask:

What capability was extracted from the work, and who now owns the commercial value of that capability?

That is a provenance question.

Reverse the Mirror

Now imagine that I introduce a new technology called:

HRI — Human Reclaimed Intelligence

I collect millions of responses from ChatGPT, Claude, Gemini, and every other major artificial-intelligence platform. I ingest their writing, code, explanations, safety methods, reasoning structures, and accumulated capabilities.I convert that material into mathematical relationships and use those relationships to build a commercial system that competes directly with them.

Then I explain that I did not copy their intelligence. My system merely learned from it.

It identified patterns. It transformed the material. It did not retain every response in its original form. It did not reproduce one platform every time it answered a question.

I claim the process is protected by fair use and call it:

Retraining

Would the artificial-intelligence companies accept that explanation?

Would the artificial-intelligence
Exhibit B: This visual represents “machine learning.” But the software isn’t meditating, studying, or thinking; it is executing a mathematical probability map built directly on the unauthorized visual vocabulary of human artists who spent lifetimes perfecting their craft.1

Or would they call it scraping, reverse engineering, model extraction, intellectual-property theft, unauthorized competition, or misuse of proprietary systems?

We already know part of the answer.

OpenAI’s current terms prohibit users from automatically extracting output, attempting to discover the underlying components of its systems, and using output to build models that compete with OpenAI. (OpenAI)

The HRI hypothetical would not be legally identical to every dispute surrounding AI development. Contracts, access restrictions, trade secrets, copyright ownership, and acquisition methods could all create legal differences.

But the reversal exposes the rationalization.

When the intelligence belongs to a technology company, provenance suddenly matters. Consent matters. Access matters. Competition matters. Market harm matters. The method of acquisition matters. The right to prevent extraction matters.

But when the source material came from millions of human beings, we were told the machine was simply learning.

HRI gives us a straightforward test:

Would artificial-intelligence companies accept their own explanation of machine learning if their systems were the material being ingested?

If the answer is no, then the principle is not:

This method is inherently fair.

The principle begins to look more like:

This method is fair when we control the machines, write the contracts, and receive the commercial benefit.

That is not simply a technological distinction.

It is a power distinction.

Fair Use Is Not a Permission Slip

There is no universal right to take anything that can be reached online.

Fair use is a legal doctrine under Section 107 of the U.S. Copyright Act. It requires consideration of the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the work’s potential market. It is a case-specific analysis, not an automatic exemption granted to technological development. (Legal Information Institute)

The U.S. Copyright Office has concluded that some uses of copyrighted material in AI development may qualify as fair use and others may not. The analysis can depend on what was copied, how it was obtained, how it was used, what the resulting system does, and whether it competes in markets served by the original works. The Office has also warned that the speed, volume, and sophistication of generated output could create market effects on an unprecedented scale. (U.S. Copyright Office)

That distinction matters.

An image model trained on aesthetic works to generate new aesthetic images is not merely cataloging those works. A music model built from musical material to produce new music aims to satisfy some of the same demand. A language model capable of producing articles, books, advertising, lessons, and software can compete with the people whose writing, structures, methods, and judgment helped make that capability possible.

Fair use cannot be treated as a gate that swings only one way.

We must examine what entered the system, what was retained, what capability was created, what comes back out, and whose market is affected.

But the damage is not limited to payment or ownership.

Provenance tells us who spoke, what they experienced, where an idea came from, what shaped it, and why it mattered. It allows us to evaluate credibility, motive, context, and meaning.

When that thread is erased, culture becomes easier to flatten.

Regional memory becomes generic content. Hard-earned business knowledge becomes an anonymous answer. The voice of a community becomes a style that can be reproduced without the community. People receive the result without the human history required to understand or judge it.

A culture that loses provenance does not merely lose credit.

It loses memory.

The Case for Clarity

Anthropomorphic Provenance Laundering is not a claim that every use of artificial intelligence is theft. It is not an argument that these systems have no value. It is not a demand that creatives reject every new tool.

I use these tools.

I believe intelligent software can strengthen human work, expand human capability, and help people build extraordinary things. But technology should increase the value of human intelligence, not erase its origin. This is not a movement against technology.

It is a movement for human authorship, visible provenance, meaningful consent, fair compensation, cultural memory, and responsible technology.

It is a demand for language honest enough to reveal what actually happened.

Creatives should not have to prove that technology is useless before they are permitted to question how it was built. We should not have to reject innovation to defend the value of our work. And we should not accept language that grants machines human qualities while reducing human creativity to anonymous raw material.

Before we debate whether the machine thinks, we should identify what it ingested.

Before we celebrate what it learned, we should ask who taught it without knowing they were teaching.

Before we call its output original, we should examine the provenance that made the output possible.

And before we accept fair use as a universal rationalization, we should reverse the mirror:

When machines ingest human intelligence, they call it training.
When humans ingest machine intelligence, what will the companies call it?

Artificial intelligence did not begin with intelligence.

It began with access.

We can build a technological future without erasing the people whose work made it possible.

But first, we have to name what is happening.

Anthropomorphic Provenance Laundering.


Continue the Series

In Part Two, we look at the tactical application of this double standard: Human Reclaimed Intelligence. Subscribe to follow the series.

  1. Capability Extraction: To illustrate the stripping away of human authorship, the machine actively performed the act, mining the texture, depth, and composition of real human creators to produce a seamless, uncredited visual about its own extraction. ↩︎

Subscribe

Engage. Inspire. Motivate.