Here is a transaction that happens millions of times a day and that almost no one talks about.
A researcher at a major university digitizes a collection of letters written by enslaved people in the 1840s. The institution posts them to an open-access digital archive. A tech company’s web crawler scrapes that archive. The letters — intimate, anguished, strategic documents of survival — become training data for a large language model. The model learns, among other things, how to produce text that sounds historical, that mimics the cadences of 19th-century correspondence, that can generate “content” about the antebellum South on demand.
No one asked the descendants of those letter writers for permission. No one compensated the Black archivists who spent decades preserving those documents. No one informed the historically Black university that originally housed the collection that its intellectual inheritance was now distributed across a corporate neural network, stripped of provenance, mixed with Reddit threads and product reviews, reduced to statistical weight.
This is not a hypothetical. This is the economy.
And this essay is not about whether that economy is legal. Courts are working that out. This essay is about something the courts are not equipped to address: What happens when a people’s collective memory is treated as raw material?
The Extraction No One Named
The AI industry runs on data. That much is understood. What is less understood — deliberately, conveniently less understood — is whose data, and what kind.
The public conversation about AI and labor has centered, rightly, on the workers. The content moderators in Nairobi, paid pennies per hour to teach models what constitutes violence. The gig workers labeling images so that self-driving cars can distinguish a stop sign from a child. The artists watching their styles reproduced without credit. The writers discovering their books in training datasets they never consented to join.
These are important fights. But they address only one dimension of the extraction. The labor dimension. The dimension that fits neatly into existing categories of exploitation — wages, contracts, intellectual property.
There is another dimension. Call it the cultural dimension. And for Black communities, it may be the more consequential one.
Because Black people in America are not just supplying labor to the AI economy. They are supplying the raw cultural infrastructure on which much of that economy is built. Language. Vernacular. Narrative convention. Musical structure. Rhetorical style. Humor. Grief. Resistance. The entire archive of how a people have spoken, written, sung, protested, prayed, and documented their existence in this country for four centuries.
All of it is data now. And almost none of it is governed.
The Archive as Commodity
Let’s be precise about what has happened.
Over the past two decades, a massive digitization effort has converted Black cultural production — from the archives of historically Black colleges and universities, from the Schomburg Center, from community libraries, from church basements, from personal collections — into digital formats. This was, in many cases, a labor of love. Archivists, librarians, scholars, and community members worked to ensure that materials at risk of physical deterioration would survive for future generations.
The promise of digitization was access. The assumption was that making these materials available online would democratize knowledge, bring Black history to wider audiences, honor the communities that produced it.
That assumption was built for a pre-AI world.
In the current landscape, digitization has a second function that no archivist in 2005 could have anticipated: it makes cultural material machine-readable. And machine-readable means scrapable. And scrapable means trainable. And trainable means profitable — for someone. Almost certainly not for the community that created the material in the first place.
Consider what gets fed into the machine. The full runs of Black newspapers, digitized through painstaking grant-funded work. Oral history transcripts from civil rights veterans, uploaded to university repositories. Slave narratives collected by the Federal Writers’ Project and hosted by the Library of Congress. Sermons, speeches, essays, letters, diaries, songbooks, protest literature, legal briefs — the textual record of Black American life, accumulated across centuries of deliberate preservation.
When these materials enter a training dataset, they do not arrive with context. The model does not know that a letter was written under duress. It does not know that a sermon was delivered in a church that had been bombed the week before. It does not know that the language in a protest leaflet was chosen with the precision of someone who understood that a misplaced word could mean a jail cell. The model sees text. It calculates probability. It learns patterns. And it moves on.
This is what extraction looks like when it operates at the level of culture rather than labor. The product is not a widget. The product is meaning itself — detached from its source, stripped of its stakes, converted into a commodity that generates value for entities that had no hand in its creation and bear no obligation to the communities it came from.
Language Is Not Neutral Data
Nowhere is this extraction more visible — and more invisible — than in the treatment of Black language.
African American Vernacular English is not slang. It is a complete linguistic system with its own syntax, grammar, and internal logic, developed across centuries of communal practice. It is among the most influential forces in contemporary American culture. The words, phrases, cadences, and rhetorical moves that originate in Black speech travel outward to shape how the entire country talks, posts, jokes, sells, and persuades.
AI models have absorbed AAVE thoroughly. They have ingested the tweets, the captions, the comment threads, the lyrics, the interviews — the vast ocean of Black digital expression that constitutes a significant portion of the internet’s most culturally dynamic content. The models can reproduce the surface patterns of this language. They can generate text that sounds Black. What they cannot do is understand what that language means — the tonal shifts that signal irony, the deliberate incompletions that carry shared knowledge, the code-switching that maps entire histories of survival and resistance onto a single sentence.
A 2024 study published in Nature demonstrated that large language models exhibit what researchers described as covert raciolinguistic bias — producing more negative assessments of speakers who use features of African American English than any human stereotypes about African Americans ever experimentally documented. The models had absorbed the language. They had also absorbed the prejudice against it. They learned to reproduce Black speech and to penalize it simultaneously.
This is not a paradox. It is a perfect mirror of how Black culture has always functioned in American commerce: valued as product, devalued as practice. Celebrated when it generates revenue, disciplined when it shows up in a job interview, a courtroom, a classroom. AI did not invent this dynamic. It automated it.
Who Owns Digitized Black Memory?
Here is the question at the center of the governance crisis, and it is a question that current legal frameworks are spectacularly ill-equipped to answer.
Copyright law protects individual works of authorship. It can, in theory, prevent a tech company from reproducing a specific novel or a specific song without permission. But much of the Black cultural archive is not covered by copyright in any meaningful way. Slave narratives are in the public domain. Oral histories recorded by federal programs belong to the government. Community-produced materials — church bulletins, meeting minutes, protest flyers — often have no clear individual author. Vernacular language, by definition, is collectively produced and cannot be owned by any single person.
This means that the most distinctive, historically significant, communally created expressions of Black American life exist in a legal no-man’s-land. They are available to anyone with a web crawler and a training budget. And the communities that produced them have no mechanism — legal, economic, or technological — to govern how their collective inheritance is used.
Indigenous communities around the world have begun to articulate a framework for this problem. The First Nations Information Governance Centre in Canada has advanced the OCAP principles — ownership, control, access, and possession — as a foundation for Indigenous data sovereignty. The CARE principles — Collective Benefit, Authority to Control, Responsibility, Ethics — have offered an international standard for governing data about and from Indigenous communities. Māori-led organizations in Aotearoa have built language models governed by Māori law, ensuring that community data remains under community jurisdiction.
These frameworks share a foundational insight: data about a people, from a people, drawn from the lived experience and cultural production of a people, is not a neutral resource. It is collective property. And its governance is a sovereignty issue.
Black America needs its own version of this framework. Not borrowed. Built.
Data Rights as Civil Rights
The civil rights tradition in this country has always been, at its core, a fight over who gets to participate in the systems that govern American life. Voting rights. Housing rights. Educational access. Employment protections. Each of these battles recognized that formal equality meant nothing if the structures remained exclusionary.
Data governance is the next structural frontier.
Consider what is at stake. AI systems trained on Black cultural data are being deployed in hiring decisions, criminal sentencing, healthcare triage, educational assessment, credit scoring, and content moderation. The models absorb Black life as input and produce decisions about Black life as output. The community supplies the raw material and receives the consequence. It participates in neither the design, the governance, nor the economic return.
This is not a technology problem. It is a civil rights problem. And it requires a civil rights response.
That response must include, at minimum:
Provenance rights. Communities must have the ability to trace how their cultural materials enter AI systems and to assert conditions on their use. This requires transparency from AI companies about training data sources — a requirement that the European Union’s AI Act has begun to address, but that U.S. law has largely ignored.
Collective governance mechanisms. Because Black cultural production is often communal rather than individual, existing intellectual property frameworks are insufficient. New models of collective data governance — data trusts, community data cooperatives, cultural stewardship agreements — must be developed and legally recognized.
Economic participation. If Black archives, language, and cultural expression generate value within AI systems, that value must flow back to Black communities. This is not charity. It is compensation. And it must be structured as infrastructure, not philanthropy — endowments for Black archival institutions, funding for community-controlled AI research, investment in the digital literacy necessary for communities to participate in governing these systems.
Interpretive authority. Perhaps most critically, Black communities must have a role in determining how their cultural materials are interpreted within AI systems — not just whether they are included, but how they are weighted, contextualized, and deployed. A training dataset that includes Black history is meaningless if the model still defaults to a framework that treats that history as marginal, exceptional, or complete.
Narrative Sovereignty Is a Governance Issue
SOLE does not use the phrase “narrative sovereignty” as branding. We use it as a governance position.
Narrative sovereignty means that the communities who produce culture retain authority over how that culture is used, interpreted, and monetized — especially when it enters systems that operate at scale. It means that Black archives are not open pits from which any corporation can extract at will. It means that the digitization of Black memory is not, by default, the privatization of Black memory. It means that data rights are civil rights, and that any AI governance framework that fails to account for the collective cultural inheritance of Black communities is a framework built on extraction.
Carter G. Woodson spent his life building institutions to ensure that Black history would survive. The archives exist because people fought for them. The oral histories were recorded because someone understood they mattered. The language persists because communities carried it forward, generation after generation, despite every institutional incentive to abandon it.
These are not datasets. They are inheritance.
And the question of who governs that inheritance — who decides how it enters the machine, what it means once it’s there, and who benefits from its presence — is not a question for Silicon Valley to answer alone. It is not a question for courts adjudicating fair use doctrine. It is not a question for policy papers that treat “diverse training data” as a checkbox.
It is a question for us. The people whose memory is on the line.
And the answer begins with a principle that should be non-negotiable: Nothing about us, without us. Nothing from us, without our consent. Nothing built on our inheritance, without our participation in what gets built.

