Charlotte Malmberg

Frameworks for simplicity beyond complex systems.

The Pipeline Nobody Audits

What makes up the corpus your general LLM has been trained on?

The honest answer: whatever made it through.

A large language model is built on what was written down, what survived being written down, and what got included in the training corpus once the survivors were collected. Filters applied in sequence over centuries, and now applied in the last decade by a small number of teams in a few cities. By the time the model produces a sentence on your screen, every one of those filters has run.

The user’s question is the first visible filter. AI does not just reflect the data. It reflects how the data is queried. That is the retrieval layer. The user operates it. But there are older filters running underneath—ones the user never sees.

This post is about the other layers. The ones the user did not operate. The pipeline that produced the data the model was trained on, and what that pipeline did long before any user typed a word.

Selection

Most of what has happened was never written down. A lot of what was written down has not survived. Some of the loss is accidental. Most of it was filtered.

The filter is not abstract.

In medieval Europe, the people who decided what to copy, translate, bind into books, and keep in libraries were monks and clerics. Male, celibate, working inside a theological framework in which female authority was inherently suspect.

When a queen ruled, the chronicle recorded it. The framing in the chronicle classified what she did as holding the household together. The activity is sometimes visible in administrative records. The framing in the narrative histories disappears it.

That is one type of selection. Not what gets recorded, but under what framework it gets recorded.

Janina Ramirez, in Femina, makes the point that the loss is compounded at the next stage. This is another type of selection.

When monasteries were dissolved, when libraries were culled, the filter was applied by people with a pre-existing framework about whose words mattered. Women’s texts failed that filter disproportionately, not because they were worse, but because the filter had already answered the question before looking at the page.

Some of what survived did so accidentally. Texts written by women that were preserved because someone thought a man wrote them. Marginalia that nobody examined for centuries because the assumption was that serious readers were male.

That is the part of the pipeline that is finished. The selection has happened. The texts that did not make it are gone. What remains is a corpus shaped by the filter, weighted toward what the filter approved.

Weighting

That corpus is what AI is trained on. It is not the world. It is what survived.

A model trained on that corpus does not know it is reading a filtered record. It learns to predict the most statistically likely next word given everything that came before. Frequency becomes truth. The dominant framing wins because it is dominant.

Call this statistical density. The denser the pattern in the training data, the more authoritative the model treats it. Statistical density has never been a reliable determination of truth or validity of output, and a model trained without that distinction cannot recover it.

Take a smaller case. If a thousand people write about a lie, and the truth is written down once or twice, what happens inside the model? The lie has the density. The truth has the frequency of an outlier. The model learns the lie as default and treats the truth as noise.

This is not the same problem as bias in the input. Bias in the input concerns the content. Statistical density is a structural argument. It produces a biased output even when the input contains the truth, because the truth is not represented at the frequency the architecture treats as authority.

The recovery work is newer, high-resolution, but low-frequency. The default record had centuries to compound. Statistical density treats accumulated volume as authority.

Output

In October 2019, the Sveriges Riksbank Prize in Economic Sciences was awarded jointly to Abhijit Banerjee, Esther Duflo, and Michael Kremer for their experimental approach to alleviating global poverty. Esther Duflo was the youngest person ever to win the Economics Nobel prize. She was only the second woman to win it in the prize’s history.

The Economic Times headline read: Indian-American MIT Prof Abhijit Banerjee and wife win Nobel in Economics.

The underlying record was correct. The Nobel committee named all three jointly. The encoding in the public record reduced one of two women ever to win the prize, the youngest recipient ever, to wife.

This is the same mechanism Ramirez described, running in 2019, in a major financial newspaper. The activity was recorded. The framing strips the agency. The model trained on this article will treat wife as the dominant frame for Duflo, because the dominant frame in the article is wife.

The mechanism is not absence. It is how women are filed when present. Margaret Beaufort, filed as the mother of Henry VII. Esther Duflo, filed as the wife. Same encoding, five hundred years apart.

This is the output stage of the pipeline. Selection produces a distribution. Weighting reinforces the dominant pattern in the distribution. Output puts the reinforced pattern back into the next generation of the corpus.

Pipeline creates the distribution. Retrieval selects from the distribution. Output reinforces the distribution. It is a loop, not a sequence. Every output the model produces is a candidate input for whatever is trained next. AI is now training on its own distribution.

Why this is getting worse, not better

Recovery scholarship exists. The 1992 Jones and Underwood biography of Margaret Beaufort exists. Women’s history as an academic discipline has existed for fifty years. Banerjee and Duflo’s prize-winning work, named correctly by the Nobel committee, exists.

None of that has the density of the older record.

A model trained on the corpus available now may return Banerjee and his wife when asked about the 2019 Nobel, because that is what the corpus contains. Eleanor of Aquitaine as a devoted mother. Catherine de’ Medici as the Black Queen. The recovery is in the data. It is not in the density.

We are moving backwards because the recovery work happened in the smaller, newer half of the corpus, and the architecture treats the larger, older half as authority. Women’s voices were fewer to start with. They are referenced less often. And they have had less time to compound. Three filters, applied in sequence, by an architecture that was not designed to know any of them were running.

And the loop is closing. Every AI summary that returns Banerjee and his wife enters the next training corpus. There is supposed to be human oversight here. In reality, it does not slow the cycle. What took the medieval filter centuries, AI now does in real time.

There does not need to be malicious intent. I have spent thirty years inside organisations that ran on filtered records and called the result analysis. The mechanism does not require villains. It requires only that the people running each stage do what their framework treats as obvious, and that the framework was already in place before they arrived.

The selection happened over centuries. The training is happening now. The output is happening on screens, generating text that reads as authoritative because it has the density of the source.

The pipeline runs without intent. That is the structural diagnosis.

Statistical density is not a reliable determination of truth or validity of output. It never has been. And we are now using it to tell us what we think we know.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *