What if AI doesn’t need to understand science to change science?
That, I think, is where the argument gets interesting. Not in the tired shouting match between “AI is magic” and “AI is autocomplete,” but in the awkward space between them. The space where something can be genuinely useful, surprisingly powerful and still not conscious, wise or meaningfully aware of what it is doing.
Tim Rocktäschel of Recursive has been talking about AI understanding and AI doing scientific discovery. It’s a bold claim, and it deserves more than a raised eyebrow. Rocktäschel is not a carnival barker standing outside the tent promising mechanical clairvoyance. Recursive’s stated ambition is serious: use AI to accelerate science, perhaps even automate parts of the scientific method itself.
There is something compelling about that. Science is drowning in its own success. Papers pile up faster than human beings can read them. Datasets sprawl across disciplines. Chemistry, biology, physics, computer science and medicine all throw off patterns that no single expert can hold in their head. The old image of the lone scientist having a eureka moment at a bench is not dead, but it is increasingly quaint. Modern discovery is often distributed, computational and hidden inside vast search spaces.
So yes, AI may help. It may help enormously.
But here’s the catch: helping with discovery is not the same thing as understanding.
An AI system can generate a plausible hypothesis. It can scan literature. It can suggest an experiment. It can design a protein, optimise a chemical structure, write code, run simulations, compare outputs and iterate. It may even find something we missed. In some fields, that may be enough to produce real advances.
But none of this proves that the system understands what it has done.
We need to be careful with our verbs.
First the system “predicts.”
Then it “reasons.”
Then it “understands.”
Then it “discovers.”
Then it “improves itself.”
Each step sounds small. Each word feels like a natural extension of the last. But the language is doing a little conjuring trick. It smuggles in agency, intention and self-knowledge without proving any of them.
A calculator does not understand arithmetic. A microscope does not understand cells. A telescope did not understand Jupiter’s moons. Yet each changed science. The question, then, is not whether a tool needs understanding to be useful. It clearly does not.
The better question is this:
What happens when the tool starts producing outputs that look like understanding?
That is the strange new territory.
With a microscope, we don’t ask the lens why it thinks the cell looks abnormal. With an AI system, we do. And it answers. Fluently. Confidently. Sometimes beautifully. It gives us a story about its reasoning. It explains its answer in a tone that suggests there is a little scholar inside the machine, calmly walking us through the steps.
But that explanation is not necessarily introspection. It may be more like a press briefing written after the event. Plausible. Coherent. Reassuring. Not the same as opening the bonnet and showing the actual mechanism.
This is why “AI understands” remains such a slippery claim. If by understanding we mean “can produce a useful answer in a context where humans would normally need understanding,” then perhaps some AI systems are getting close in narrow areas. But if we mean something richer, such as grasping meaning, recognising limits, forming justified beliefs, knowing why a result matters and being able to interrogate its own assumptions, the evidence is much weaker.
And if we mean that AI understands its own thought mechanisms, the claim becomes weaker still.
Current large language models do not have transparent access to their own causal processes. They do not inspect their weights, trace their internal activations and report back like a scientist describing an experiment. They generate language. They produce explanations. Some of those explanations are useful. Some are post-rationalisations. Some are theatre.
Let’s be real: humans do this too.
We often explain ourselves after the fact. We confabulate. We tidy messy motives into neat stories. Anyone who has sat through a project board, a political interview or a family argument knows that “reasons” and “what actually happened” are not always the same creature.
I’ve seen this often enough in change work. A project can acquire a very clean story once it has already started drifting. The RAG status is amber-green, the benefits are “on track,” the dependency is “being managed” and everyone knows there is something underneath the table that the paperwork has not quite named. The story is not always false. Sometimes it is worse than false. It is tidied.
But with humans, there is at least a body, a history, a set of commitments, a social world and consequences that bite back. A scientist can be embarrassed, challenged, discredited, retrained or persuaded. A project manager can be held accountable. A leader can be asked not just “what is your answer?” but “what did you fail to consider?”
With AI, accountability has to be designed around the system. It does not arise naturally from the system.
This is where the scientific discovery story becomes both exciting and risky.
Imagine an AI system that proposes a new material, a drug candidate or an AI architecture. It runs thousands of simulations, compares results, discards weak options and escalates promising ones. The result may be genuinely novel. It may work. It may even be something no human would have found unaided.
Does that count as discovery?
In one sense, yes.
Evolution discovered wings without understanding aerodynamics. Ant colonies solve logistical problems without a project plan. Markets find prices without anyone understanding the whole economy. Useful novelty does not always require conscious comprehension.
But science is not only novelty. Science is not just finding a thing that works. It is also explaining why it works, knowing when it does not, understanding the boundary conditions and building a shared body of knowledge that others can challenge, reproduce and extend.
This is where the word “discovery” starts to wobble.
If an AI system finds a pattern but nobody understands it, have we discovered something or merely located it? If a model proposes a treatment but cannot explain the causal mechanism, have we advanced medicine or created a very sophisticated guessing machine? If an AI improves an AI architecture, is that scientific progress, engineering search or just a clever loop with a glossy interface?
The answer may be “all of the above,” which is precisely why the language matters.
There is a familiar problem here for anyone involved in change work. An answer can look complete before the situation is understood. A dashboard can look authoritative while the data underneath is weak. A project plan can look tidy while the dependencies are quietly rotting underneath. A business case can say “benefits realisation” while everyone in the room knows the benefits are more prayer than plan.
That is one of my recurring worries about AI. It does not just answer questions. It can make us feel as if the question has been settled. I keep coming back to the same uncomfortable thought: perhaps the real human responsibility in an AI age is not having all the answers, but protecting the quality of the questions.
Because AI makes uncertainty look good.
It can make the half-known look known. It can turn a chain of weak assumptions into a confident paragraph. And because the prose is smooth, people may stop noticing the joins.
That is why the debate about AI discovery should not be reduced to whether AI is clever enough. Cleverness is the wrong test. The real test is whether the discovery process has evidence, verification, challenge and governance built into it.
Not governance as a dead hand. Not the kind of governance that turns curiosity into committee paperwork and buries every new idea in a SharePoint folder last opened in 2019. I mean governance as disciplined curiosity. Governance as the act of asking: How do we know? What did we test? What did we exclude? What failed? What would prove us wrong? Who is accountable if this is impressive but unsafe?
That is not anti-innovation. It is what stops innovation becoming theatre.
The strongest version of the AI optimist’s argument is worth acknowledging. They may say: “You are demanding a human kind of understanding from a non-human system. That is parochial. The point is not whether AI understands as we do. The point is whether it can produce reliable, useful discoveries.”
Fair challenge.
We should not insist that machines think like humans before we allow them to help us. Aircraft do not flap their wings. Submarines do not swim like fish. Computers do not calculate like clerks with pencils. A non-human route to discovery is still a route.
But the counterargument is just as strong: when we use human words like “understanding,” “reasoning” and “discovery,” we import expectations of judgement, responsibility and meaning. If the machine does not possess those things, the surrounding system must.
That is the missing piece in much of the AI debate.
The question is not simply, “Can AI discover?”
It is, “What kind of system must surround AI-generated discovery for it to become knowledge rather than noise?”
This is where local government, healthcare, science and business all meet the same problem. We are not short of outputs. We are short of trustworthy outputs. We are not short of answers. We are short of good questions, safe processes and people willing to challenge the beautiful slide before it becomes the approved direction of travel.
And I do mean beautiful slide. Anyone who has worked around transformation programmes knows the danger. The diagram has a nice gradient. The arrows line up. The customer journey has been simplified into four pleasing boxes. Somewhere between box two and box three, reality has been asked to leave the room.
AI can do that at scale.
It can produce the elegant map before anyone has walked the territory. It can make the change look delivered before anyone has changed. It can produce the “lessons learned” before anyone has learned them.
This is why I don’t want to dismiss AI discovery. That would be too easy and probably wrong. The telescope did not understand the heavens, but it still changed our place in them. The camera did not understand memory, but it changed how we saw ourselves. Film editing did not understand time, but it taught us new ways to experience it.
AI may become another such instrument. Perhaps a stranger one. A machine that does not understand, yet helps us see.
But that only works if we remember what the instrument is.
AI may well become a discovery engine. But an engine is not a driver. It has power, but not purpose. It can move fast, but it does not decide where civilisation ought to go.
Perhaps that is the line we need to hold.
Let AI search. Let it simulate. Let it propose. Let it surprise us. Let it find patterns in the parts of reality too large, too small or too tangled for us to see unaided.
But do not let us lazily rename that “understanding” simply because the machine has learned to speak in the accent of certainty.
The future of AI in science may not be artificial intelligence replacing human understanding. It may be something stranger: artificial discovery forcing us to become more serious about what understanding really means.
And perhaps that is the real test.
Not whether the machine understands.
Whether we still do.
Comments
Post a Comment