Home About Book Reviews Game Reviews Essays RSS

essay

The Chosen One Is a Gap in the Training Data

The chosen one usually arrives with a prophecy.

A child bears a mark. An ancient text names a bloodline. A hidden power wakes up at exactly the right moment. Fate singles one person out and tells the world to pay attention.

Predictive systems suggest a stranger version of the trope.

What if the exceptional person matters because the system has almost no data about them?

The chosen one is a gap in the training data.

That idea turns an old fantasy structure into a very contemporary science-fiction problem.

Prediction depends on history

Most behavioral models learn from records.

What did people click? Where did they go? What did similar users buy? Which features correlate with a later event? How did someone respond the last thousand times the environment changed?

A model becomes useful by finding regularity.

This is why recommendation systems improve when they have more interaction history. It is also why new users, rare behaviors, and unusual contexts can be harder to model.

The science-fiction version is simple: a society becomes so good at prediction that being statistically ordinary makes you legible.

The dangerous person is the blank space.

Foundation and the anomaly that breaks history

Foundation is built around population-scale prediction. Psychohistory can anticipate broad historical forces, but it is vulnerable to disruptions outside its assumptions.

The Mule becomes famous precisely because he is an anomaly the plan did not adequately account for.

He is not “unmodelled” because he skipped an app. The mechanism is different. The structural lesson is the same: every predictive system has a domain in which its confidence is earned.

Outside that domain, precision can become theater.

The Minority Report and the person who learns the forecast

Philip K. Dick’s “The Minority Report” introduces another kind of model-breaking actor: a person who gains access to the prediction about himself.

That creates a feedback loop.

If the system predicts an action and the subject learns the prediction, the information can become a new cause. The forecast changes the person it describes.

This is one reason predictive AI fiction is so fertile. Prediction of passive objects is easier than prediction of agents who can react to being predicted.

A person can resist the model simply because the model exists.

Little Brother and practical illegibility

Cory Doctorow’s Little Brother takes place in a surveillance-heavy San Francisco after a terrorist attack.

Marcus understands networks, security systems, and the gap between what surveillance claims to know and what it can actually infer. His resistance depends on technical literacy and on creating uncertainty inside systems that treat patterns as evidence.

The novel is useful because it refuses the idea that surveillance equals omniscience.

Data collection creates power. It also creates false positives, incentives, workarounds, and adversarial behavior.

The more people know they are being watched, the more observation becomes a game.

Gnomon and refusing legibility

In Gnomon, opacity becomes a form of resistance against a society organized around transparency.

That conflict reaches beyond privacy. A system that interprets total visibility as civic virtue can treat hiddenness as evidence of wrongdoing.

This is where surveillance capitalism and state surveillance share a structural temptation: collect more because more data appears to promise fewer surprises.

But a world without surprises is also a world with very little room for people to redefine themselves outside existing categories.

QualityLand and the right to say the profile is wrong

QualityLand makes this comic.

Peter Jobless receives a product chosen for him by a system that supposedly knows his wants. His insistence that he does not want it becomes a challenge to the entire logic of algorithmic certainty.

The joke works because the model has institutional prestige. The system’s prediction is treated as stronger evidence of Peter’s desire than Peter’s own statement.

That is preference shaping at its most absurd. If the model is assumed to know you better than you know yourself, disagreement can be reclassified as confusion.

MAYA: Seed Takes Root and the off-grid hero

MAYA: Seed Takes Root makes this structure central to its protagonist.

Yachay grows up without tethering to Maya, the living planetary network used by nearly everyone around him. The Divyas’ predictive power depends on the information that flows through that network.

Yachay has very little history inside it.

The chosen one is a gap in the training data because the system cannot model him the same way it models people who have spent their lives generating behavioral information.

The second formulation is even cleaner: The chosen one just never logged in.

And the consequence follows: The algorithm cannot see him.

That turns a familiar heroic exception into an information problem. Yachay is not important because a system recognized him as special. He is important because the system failed to recognize him at all.

This is The right to be unmodelled converted into narrative stakes.

The phrase matters because privacy and modelling are different. A person can keep a specific secret while remaining highly predictable. Conversely, somebody can reveal many facts while behaving in ways a model does not anticipate.

MAYA also explores predictive processing in its treatment of minds and perception, but the political system’s behavioral prediction is a separate mechanism. The important overlap is uncertainty: models act on what they expect, then have to update when the world violates the expectation.

That is also where AI alignment becomes political. A system may be aligned to a social objective and still mishandle a person who falls outside the assumptions embedded in its model.

Hindustan Times’ interview on MAYA’s multi-format universe describes the project as interested in truth, control, shifting points of view, and systems that change who appears heroic or villainous. Yachay’s illegibility gives that larger idea a practical form.

The cold-start hero

Machine-learning practitioners already have ordinary terms for parts of this problem.

A new user can create a cold-start problem because the system has little interaction history. Rare cases can be poorly represented in training data. Distribution shifts can make patterns learned in one context unreliable in another.

Science fiction can turn those technical limitations into social hierarchy.

Imagine a city where access depends on a risk score. The person with no record may be treated as safe because there is no evidence against them, or dangerous because there is no evidence at all.

Imagine healthcare built around population models that work well on the majority and badly on rare bodies.

Imagine a political system that sees unpredictability as a threat because every unforecast action creates cascading uncertainty.

The “chosen one” can emerge from any of those gaps.

Being seen can become a form of control

There is a common assumption that recognition is always good.

Often it is. Being ignored by institutions can mean exclusion from credit, healthcare, rights, or political attention.

But perfect legibility creates its own danger.

A system that sees every pattern can tailor every incentive. A platform that knows every vulnerability can decide when to show a message. A government that predicts every route can intervene before a protest forms.

In that world, opacity can become political capacity.

This is why science fiction about algorithms should care about people who do not fit the model. They expose the difference between a map and the territory.

The hero does not have to be statistically unique

The strongest version of this trope avoids another mistake.

The protagonist does not need to be biologically unprecedented or metaphysically immune to prediction.

Sometimes the meaningful act is simply refusing participation.

Do not generate the data.

Use a different channel.

Behave outside the expected category.

Protect a part of life from measurement.

Create institutions where an appeal can override the score.

That makes the trope less mystical and more useful.

The old chosen one was special because destiny could see them.

The algorithm-age chosen one may be special because destiny cannot.

Why predictive systems need outsiders

The appeal of many AI science fiction books is that they make intelligence larger than one character. A model, network, or institution can know more than any individual person.

That power creates a structural weakness: systems learn from what they can observe.

The missing case matters.

A person who does not use the platform, a community that records information differently, a behavior with no historical precedent, or an event produced by a new combination of old causes can all expose the limits of a model trained on the past. That is why the outsider is such a durable figure in AI science fiction books. The outsider does not have to be magically unpredictable. They only have to fall outside the system’s experience.

This is also a useful distinction in science fiction about artificial intelligence. The problem is rarely that a model knows nothing. The interesting problem is that it knows a great deal and can therefore become overconfident about the remaining unknowns.

In MAYA, this tension becomes Prediction as government. Most of Neh is legible to a system that learns from participation. Yachay matters because his absence creates a hole where the model expects continuity.

That makes the classic chosen-one structure feel less mystical. The exceptional person is exceptional because the institution lacks data, not because the universe issued a prophecy.

A prediction system can be strongest exactly where history is repetitive. The hero appears where history stops repeating.

More like this

← back to the blog