Skip to content

Home / Topics / 03

Can a machine form a hypothesis?

Systems propose candidate explanations, and the good ones read exactly like the real thing. Whether that is discovery or very good autocomplete depends on what a hypothesis is actually for.

AIPhilosophy of science Research practice
01

The Video

The video follows three attempts to automate the front of the scientific process, and what each of them turned out to be automating instead.

Runtime
24 minutes
Companion
Hypothesis Machine: the argument as a game
Essay
Hypothesis, or a sentence that looks like one?
02

The Essay

This essay reads two ways. Pick one.

A hypothesis is a commitment, not a sentence. It says in advance what would prove it wrong. That is a strange thing to ask of a system trained to continue text.

Machines have been proposing candidate explanations since the 1970s, and the useful ones have almost always come from search over a constrained space rather than from anything resembling intuition. What changed recently is not the reasoning. It is the fluency. A generated hypothesis now reads exactly like a real one.

That matters because the format carries authority. A claim shaped like a paper, cited like a paper, entering a literature review like a paper, acquires the standing of a paper without anyone ever having committed to it. Fields can accumulate confident statements nobody has tested and nobody remembers proposing.

The honest version of the tool is narrower and more useful: a generator of candidates that a person then has to stake something on. The staking is the part that cannot be delegated, because the staking is what science is.

The interesting failure is not factual error. It is that nothing in next-token prediction enforces global consistency across a generated structure.

Fluent, and jointly unsatisfiable

Models can recite the criteria for a good hypothesis and still emit sets whose members are jointly unsatisfiable: a mechanism in one sentence, a scope condition in the next that the mechanism cannot meet. Retrieval grounding narrows the space of individually wrong statements; it does nothing to enforce coherence between them.

A cheap filter that works

There is a usable filter here, and it is cheap. Score a generated hypothesis by whether its stated falsification conditions are reachable with the instruments you actually have. Most of what looks impressive fails immediately; most of what survives is worth an afternoon.

Can commitment be engineered?

The deeper question is whether commitment can be engineered rather than simulated. It plausibly can, but it needs an architecture in which being wrong costs the system something, and we do not train for that. Until then, fluency in the grammar is not evidence of a claim, and we have no instrument that separates the two at a glance.

This is the roadmap's question, if a model reaches a result humans cannot explain, is that still science?, arriving from the other direction. Explanation has never been a formal requirement of a scientific claim. It has been a practical one. Machine learning is the first thing to test whether the distinction survives contact with a working lab.

03

Sources, at Three Levels

Reading lists usually assume one reader. These do not: pick the level you want to enter at, and move up when you feel like it.

04

How the Question Developed

1965

DENDRAL

The first program to generate candidate chemical structures from mass spectrometry. Machine proposal is older than most of the current argument.

1983

BACON rediscovers Kepler

A search program recovers known physical laws from data, and the debate immediately becomes whether rediscovery counts as discovery.

2009

Free-form natural laws

Symbolic regression extracts conservation laws from a swinging pendulum. The output is an equation a physicist can read.

2021

AlphaFold

A prediction good enough to use, and an explanation nobody can give. The practical question arrives before the philosophical one is settled.

2023 on

Generated candidates at scale

Fluency arrives. The bottleneck moves from producing plausible claims to telling them apart from real ones.

05

The Simulation

A hidden relationship, some noisy measurements, and a generator that will happily fit any of them. Raise the flexibility: the fit keeps improving while the prediction gets worse. That gap is the whole argument.

Fit Predict

06

Viewpoints

We publish the strongest case against our own reading. If the counter-argument is weak here, that is our failure, not the argument's.

Against

Requiring that a claim be “staked” describes how scientists behave, not what makes a statement scientific. A hypothesis that survives testing is not improved by having had someone's reputation attached to it. If the filtering works, the origin is irrelevant.

The strongest objection we could find

For

Origin is irrelevant only while the supply of candidates is small. When generation is free and testing is not, something has to decide what gets tested, and a system with nothing at stake produces candidates at a rate no lab can absorb.

The position the video leans toward

Where do you land?

  • Origin is irrelevant if the test is sound
  • Someone has to be accountable for the claim
  • It depends entirely on the cost of testing

Make your case in the discussion below.

07

Discussion

EK

Elif K.researcher

The falsifiability filter is doing a lot of work in this piece. Plenty of good hypotheses were untestable when proposed: continental drift sat there for fifty years. The filter is useful triage, not a definition.

DM

D. Mbekireader

The simulation is the bit that landed. Watching the fit get better while the prediction gets worse is the whole essay in one slider.

SA

Sajad A.host

Elif is right, and the video overstates it, noted for the correction log. The honest version is that reachable falsification conditions tell you what to test next, not what counts as science.

Members can post. Everyone can read.