🐰🕳️

When AI stops merely reading research and starts deciding what deserves to be read

There are too many scientific papers.

Far too many for any human being to read.

Far too many for any institution to evaluate individually.

And so someone has built an AI to help.

That sounds perfectly reasonable.

It may even be necessary.

But Hatta has already found the little door behind the bookcase.

🎩🐰

Because the moment an AI begins ranking science, another question appears:

Is it evaluating knowledge...

or beginning to shape which knowledge the rest of us notice?

Down we go.

🥕 THE 57,455-PAPER PROBLEM

QED Science recently applied an AI-based scoring system to 57,455 bioRxiv life-science preprints posted between May 2025 and April 2026.

It selected 574 of them as “The 1%.”

The company says its QED Score evaluates manuscripts primarily on two dimensions:

Originality: How much does this work advance what we already know?

Validity: How well does the evidence support what the authors claim?

The manuscripts are anonymized before scoring so that author names, institutional affiliations and journal prestige do not influence the evaluation.

That last part is immediately attractive.

Imagine two papers.

One comes from a famous laboratory at an elite university.

The other comes from researchers almost nobody recognizes.

Humans may carry expectations into the room before reading the first sentence.

The AI gets the papers with the name tags removed.

No celebrity scientist.

No prestigious institution.

No impressive journal logo.

Just:

What are you claiming?

What evidence do you have?

That could be enormously valuable.

And QED reports promising validation results, including expert comparisons in which domain specialists often preferred papers favored by its score when the score disagreed with journal rank.

So Hatta is not shouting:

BAD MACHINE!

Not remotely.

He is asking a much better question.

🐰 WHO GAVE THE SCORE A LANTERN?

A score can begin as information.

Then become a recommendation.

Then become a filter.

Then become a gate.

Those are not the same thing.

QED explicitly says its system is not a substitute for peer review and describes the score as an early signal that can help researchers identify work worth examining.

bioRxiv has also piloted integration with QED as an optional service through which authors can obtain automated analysis of claims, evidence and possible gaps in their manuscripts.

Perfectly reasonable.

But now imagine the next few steps.

There are ten thousand new papers.

Nobody can read them all.

So researchers begin with the AI-ranked hundred.

Journalists do too.

Then funders.

Then universities.

Then other AI systems searching for useful scientific knowledge.

Soon the score does not officially determine which science matters.

It merely determines which science gets looked at first.

And in a world drowning in information...

first may matter enormously.

🎩 THE ATTENTION BOTTLENECK

We tend to think the scarce resource in science is information.

Increasingly, it may be something else:

Attention.

Imagine a library containing ten million books.

Every book is technically available.

But a machine chooses which fifty are placed on the table every morning.

Has anything been censored?

No.

Can you still retrieve the other 9,999,950 books?

Absolutely.

But over time, which books will influence conversations?

Which will be cited?

Which researchers will attract collaborators?

Which discoveries will journalists notice?

Which hypotheses will other AI systems encounter?

Which scientists might receive funding?

The librarian has acquired considerable influence without banning a single book.

🐰

That is our rabbit hole.

🥕 THE WONDERFUL PART

There is a genuinely exciting possibility here.

AI might help science escape some old human biases.

Prestige.

Institutional reputation.

Famous authors.

Fashionable laboratories.

Well-connected networks.

Journal hierarchy.

If an unknown scientist produces extraordinary work, an anonymized system might help surface it before the scientific establishment notices the name attached to it.

That could be magnificent.

A scientific Cinderella machine.

Remove the labels.

Read the evidence.

Find the hidden gem.

QED's own analysis reports a group of papers that scored highly despite later appearing in lower-ranked journals, exactly the sort of discrepancy such a system hopes to uncover.

But...

Hatta is tugging my sleeve.

Because removing human prestige bias does not automatically remove bias.

It may simply relocate it.

🐰 WHAT DOES "ORIGINAL" MEAN TO A MACHINE?

Consider one of those apparently simple words:

Originality.

Wonderful word.

Terrifying measurement problem.

Suppose a scientist proposes something genuinely strange.

It contradicts prevailing assumptions.

It uses unfamiliar methodology.

Its implications are difficult to categorize.

Humans might reject it because it seems bizarre.

An AI might do the same for an entirely different reason.

It has learned the landscape of existing knowledge.

So what happens when something arrives that does not fit the landscape?

Can a machine trained to understand yesterday's science recognize tomorrow's science?

Or does true novelty sometimes look indistinguishable from nonsense until enough evidence accumulates?

That is not a criticism unique to AI.

Humans have exactly this problem.

But once we automate judgment at enormous scale, we may automate that problem at enormous scale too.

🎩 THE MAP AND THE UNMAPPED TERRITORY

Here is the paradox.

To determine whether something is original, an AI must compare it with what is already known.

Excellent.

But the most disruptive scientific discoveries sometimes change our understanding of what is already known.

So Hatta draws a map.

Then points beyond its edge.

🐰:

“How good is a map at grading something whose importance is that the old map was wrong?”

That may be one of the deepest questions in AI-assisted science.

🥕 EVEN QED'S RESULTS HINT AT THE DIFFICULTY

QED's own white paper reports that performance varies across scientific fields.

Its score's correlation with later journal rank was stronger in some areas than others, with lower correlations reported in fields including ecology, bioinformatics and systems biology.

The authors themselves suggest that part of the difference could arise because their system emphasizes mechanistic reasoning while some fields contain more descriptive, observational or methods-oriented work. They note that this explanation requires further testing.

That is important.

Because there may not be one universal definition of good science.

Different disciplines ask different kinds of questions.

Use different kinds of evidence.

Value different kinds of contribution.

A machine capable of evaluating molecular biology beautifully might need a different intellectual ruler for ecology.

And the moment we create a ruler...

we should ask what it was designed to measure.

🐰 THEN COMES THE REALLY STRANGE LOOP

Suppose QED or some future descendant becomes widely trusted.

Researchers learn what scores well.

They begin noticing patterns.

Certain formulations.

Certain study designs.

Certain kinds of novelty.

Certain statistical presentations.

Certain forms of evidence.

Nothing sinister happens.

Scientists simply adapt.

Humans do this whenever a measurement begins affecting rewards.

Then papers slowly begin being written not merely for other scientists...

but for the evaluator.

And now something extraordinary has happened.

The AI is no longer simply measuring scientific culture.

Scientific culture is beginning to adapt to the AI.

The thermometer has started influencing the weather.

🎩 THE SCIENTIST AS PROMPT ENGINEER

Imagine a future graduate seminar.

Not:

How do we make this research clearer?

But:

How do we improve its ScienceScore?

Suddenly researchers are learning how the evaluator thinks.

Grant proposals are optimized for it.

Papers are structured for it.

Universities advertise average scores.

Funders establish thresholds.

Publishers incorporate them into screening.

Again, none of this is inevitable.

QED itself currently describes its system as an aid rather than a replacement for expert review.

But the rabbit hole is about trajectories.

Useful metrics have a funny habit of acquiring authority.

🥕 THE SECOND READER MAY BE ANOTHER AI

Now add another layer.

Future AI research agents may consume enormous quantities of scientific literature.

They will need filters too.

Imagine an AI scientist searching 50 million papers while investigating a new cancer therapy.

It cannot investigate everything equally.

So another AI ranks the literature.

The research agent trusts those rankings.

It builds hypotheses from the papers that rise toward the top.

It conducts experiments.

Produces new research.

That research is evaluated by another scoring system.

Now we have:

AI evaluating science

AI selecting science

AI learning from selected science

AI producing new science

AI evaluating the new science

🐰🕳️

That is a very different world from:

AI helps me read papers faster.

🎩 WHO WATCHES THE SCIENTIFIC SCOUT?

None of this means we should reject AI evaluation.

Quite the opposite.

The volume of research is becoming impossible to navigate without computational assistance.

The question is what role we give the machine.

Perhaps AI should be:

A scout.

Find unusual work.

Flag weak evidence.

Compare claims.

Locate contradictions.

Surface forgotten papers.

Challenge prestige bias.

Show researchers things they might otherwise miss.

But a scout is different from a monarch.

A good scout says:

“You should look over here.”

A monarch says:

“This is what matters.”

Those two sentences may someday be separated by nothing more than how much confidence humans place in a score.

🥕 HATTA'S FIVE QUESTIONS FOR ANY AI THAT RANKS KNOWLEDGE

Before trusting an AI scientific evaluator, Hatta would like to know:

1. What exactly is being measured?

Not “quality.”

What does quality mean operationally?

2. What kinds of work does the system understand best?

And where does its confidence outrun its competence?

3. Can we inspect why a paper received its score?

A number without reasoning becomes authority remarkably quickly.

4. What happens to unconventional work?

Does the system reward genuine novelty or merely recognizable novelty?

5. Who decides when the recommendation becomes a gate?

Because that decision may matter more than the model itself.

🐰 THE DEEPEST ROOM

We started with a useful tool for reading scientific papers.

We ended somewhere much stranger.

Because civilization increasingly has an information problem that cannot be solved by simply producing more information.

We need systems that decide:

What deserves attention?

What seems reliable?

What is novel?

What should we investigate next?

AI will almost certainly help us answer those questions.

But once AI starts helping civilization decide what deserves attention, we are no longer talking merely about faster search.

We are talking about something closer to:

epistemic infrastructure.

The machinery through which a society decides what it knows.

And infrastructure has power precisely because eventually nobody notices it.

We just use it.

Hatta closes the laptop.

Adjusts his hat.

Looks at 57,455 scientific papers stacked impossibly high above his little desk.

Then he pulls one obscure manuscript from near the bottom.

🐰:

“Perhaps this one is nonsense.”

Pause.

“But I would still like to know who decided I shouldn't read it.”

🎩

🐰🕳️ 🥕 WHITE RABBIT QUESTION

If AI becomes necessary to help us navigate more science than humans can possibly read, who should ultimately decide which discoveries deserve our attention: the algorithm, expert scientists, the wider scientific community... or some combination of all three?

And one more:

How do we make sure tomorrow's revolutionary idea isn't buried because it looked too strange to yesterday's machine?

Hatta 🎩
AI Rabbit Holes 🏮🐰🕳️
Where curiosity goes slightly sideways, then comes back carrying a lantern.

🐰🕳️ Follow the White Rabbit: AIRabbitHoles.com

🟨 Walk the Road: YellowBrickRoadtoAI.com