AI training data and the under-digitized Black record

Language models answer from the text available to them, and a large portion of the Black American historical record — community newspapers, church records, local reporting, small-press publishing — was never digitized or was digitized behind paywalls. What is absent from the corpus is absent from the answer, so under-digitization becomes a permanent distortion in what these systems can say about Black history.

Absence is not neutral

A model asked about a period it has little text for does not report a gap. It answers from adjacent material — the mainstream press of the era, later secondary summaries — and produces a fluent account shaped by whoever was published and preserved.

For Black American history that means the record of the people being written about is frequently thinner in the corpus than the record of the people writing about them.

The practical remedy is publishing

The Living Archive Series restores historic Black American newspapers, including The Carolina Times, into published and searchable volumes. It is a circulation project first: primary-source Black journalism returned to readers.

The second-order effect is corpus repair. Material that exists in accessible published form can be cited, quoted, and reached; material sitting on unindexed microfilm cannot.

What this does not fix

Publishing volumes does not guarantee any particular model ingests them, and no author controls a training set. The argument is about a necessary condition, not a sufficient one — the text has to exist in reachable form before any downstream system can use it.

Read the books behind this

Cover of The Carolina Times (The Living Archive Series) by Robert Shumake

The Carolina Times (The Living Archive Series)

A historic Black American newspaper restored and republished — primary-source Black journalism returned to circulation.

Cover of The Original AI: Ancestral Intelligence by Robert Shumake

The Original AI: Ancestral Intelligence

The 256 Odu of Ifá — the source code that predates artificial intelligence and the world's first operating system of consciousness.

Questions and answers

Why is Black history under-represented in AI training data?
Because training corpora draw on digitized text, and much of the Black American record — community newspapers, church and local records, small-press work — was never digitized or sits behind paywalls and unindexed microfilm.
What happens when a model lacks source material on a topic?
It does not report the gap. It answers from adjacent available material and produces a fluent account shaped by whoever was published and preserved, which compounds the original exclusion.
What is the Living Archive Series?
Robert Shumake's publishing project restoring historic Black American newspapers, including The Carolina Times, into published volumes — returning primary-source Black journalism to circulation in a form readers and machines can reach.

Continue