I have been recently diving into going back on writing on a piece of paper. Right now is kind of a critical time- to be scientists who choose mediums of slow expression. This means for me- to engage my senses- so I write in a medium that is not my computer. It is my journal. Agents have create a lot of burn out in our life. I have personally experienced a "read the .md" file fatigue and in many ways- I understand Demis and his decision to go back to science.
Today I want to probe in details about discovery vs regulation.
Lets first understand the two definitions in context of science:
Today I want to probe in details about discovery vs regulation.
Lets first understand the two definitions in context of science:
- Discovery is probabilistic. Finding a molecule is a search through an enormous space where you don't know the answer in advance. Being fast and approximately right, then validating in the lab, beats being slow and precise. The uncertainty is not a flaw to manage; it is the feature you are paying for.
- Regulation is deterministic. A clinical study report carries thousands of individually consequential facts. There is exactly one correct number for how many patients experienced a given adverse event. One. "Approximately right" is not a minor limitation here. Hallucination is not an inconvenience to be cleaned up later. It is the thing you cannot ship.
In current scientific context, I feel we are frequently making a mistake of treating AI as one decision. I think we should match the shape of tool to the shape of the problem- scientifically.
In the above, we are now seeing the two kinds of modes that exist for AI:
The first sees AI as changing the epistemic structure of discovery itself. Not just making existing workflows quicker, but changing how hypotheses are generated, how search is conducted, how evidence is combined, and how problems are decomposed. This is the feature of hallucination which is termed inversely as uncertainty in this section of work.
The second sees AI as a faster way to do what we already do. Search the literature faster. Draft the report faster. Summarize the meeting faster. Write code faster. This is not trivial. Much of scientific work is bottlenecked by coordination, formatting, synthesis, and repetitive assembly. Better tools can remove real friction.
What I have a beef with is: we as scientific members are bringing probabilistic optimism to deterministic documentation, and deterministic caution to probabilistic exploration. And I think this is wrong.
That is why so many current coscientist pitches feel confused. They are trying to sell one interface into multiple kinds of scientific work without respecting that those kinds of work operate under different truth conditions and axioms.
The second sees AI as a faster way to do what we already do. Search the literature faster. Draft the report faster. Summarize the meeting faster. Write code faster. This is not trivial. Much of scientific work is bottlenecked by coordination, formatting, synthesis, and repetitive assembly. Better tools can remove real friction.
What I have a beef with is: we as scientific members are bringing probabilistic optimism to deterministic documentation, and deterministic caution to probabilistic exploration. And I think this is wrong.
That is why so many current coscientist pitches feel confused. They are trying to sell one interface into multiple kinds of scientific work without respecting that those kinds of work operate under different truth conditions and axioms.
Discovery wants probabilistic tools
For model-informed scientists, this distinction should feel natural.
Discovery has always tolerated uncertainty differently from regulation. Mechanistic reasoning, translational inference, PK/PD projections, PBPK scenario exploration, target prioritization, and candidate triage are all domains where we work under incomplete knowledge. We build models not because we know everything, but because we need a disciplined way to reason under uncertainty. And we need universal explainers for these systems.
Current AI systems are like Chinese rooms as described in this position paper by Tom at google deepmind.
They recombine existing symbolic concepts to optimize metrics hence lacking the sensory grounding to invent axioms without symbolic connections. Scientifically, we know optimization only occurs within a framework (and that is how you can identify process people against scientists) however, current AI systems as it stands, cannot create frameworks.
This brings me to an important observation of the framing that I chose to box scientists in: the framing of the verifier..
I think our job was never just verification. It was never just pattern matching. It was the Jump- the abduction. Surely, verification definition and job would be the bridge for a year or more in this loop.
Here is the learnings I took on RLEF and Constitutional AI actually demonstrates and what it debunks about the standard MIDD narrative in my point of view:
Myth 1: "Probabilistic and deterministic are separate rooms."
ReAct interleaves chain-of-thought reasoning with tool-calling actions in a tight loop: reason → act (call tool) → observe → reason again. The reasoning is probabilistic. The tool output is deterministic. The loop is both.
Scientifically, for example: A PBPK model call from an LLM is deterministic. The decision to call it is probabilistic.
this drives me to the second myth:
Myth 2: "Generate-then-check is just doubled work."
This assumes the "check" is always a human. It is not. If you look at the way LLMs are built- the verifier is code. It runs automatically. It gives ground-truth signal. The LLM learns from it.
In drug development terms: if your model output can be checked against a simulation, a database constraint, or a mechanistic invariant, that check does not need to be a person. It can be architecture.
In drug development terms: if your model output can be checked against a simulation, a database constraint, or a mechanistic invariant, that check does not need to be a person. It can be architecture.
Myth 3: "Agents scale output. Humans scale trust."
This sells agents short. Constitutional AI shows that models can self-critique and self-revise against written principles. The critique is generated by the model. The revision is generated by the model. The preference model is trained on AI feedback, not human feedback.
The scientist is not the verifier. The scientist is the verifier architect- the one who writes the constitutional principles, designs the unit tests, and defines the mechanistic invariants that the agent checks itself against. This for example is understood clearly by one community and special interest group from ISOP where they are building the verification evals for pharmacometrics packages. You can find the leaderboard here.
Why this matters for MIDD:
Why this matters for MIDD:
From the Google Deepmind paper:
Einstein had almost no data forcing him to invent General Relativity. Newtonian gravity worked to 10⁻⁹ precision. The only anomaly was Mercury's perihelion which was explained away as a hidden planet ("Vulcan"). There was no inductive signal. No error gradient. No dataset to compress.
Thus, for bringing AI to be able to be abductive: the paper proposes that the path to true AI invention is physically consistent, multimodal world models — systems that can run counterfactual simulations, not just predict pixel sequences. Genie (DeepMind, 2024) is a first step: action-controllable generative environments where an agent can intervene, not just observe.
A well-built PBPK or QSP model is a world model. It encodes physical laws (mass balance, receptor binding kinetics, diffusion). It allows counterfactual simulation: what if the dose were higher? What if the target expression changed? What if the patient were pediatric? So the modellers that we are, we build and validate these structures is not just running simulations. In principle, we are building the synthetic laboratory that future agentic systems will need to reason abductively about drug action.
AI systematically favors data-rich domains. And that is where, I think one of the most important areas to exploit the current reception of GenAI is by being bullish on digital twin generation with this technology.
For now, the data poor domains are safe from AI abduction- what does this mean for us? This means rare diseases, novel mechanisms, and first-in-class modalities will always require human abductive reasoning. The modeler who can formulate a plausible mechanistic hypothesis in the absence of rich clinical data is doing exactly what no LLM can do.
Finally, the verifier architect is also the axiom architect
Einstein had almost no data forcing him to invent General Relativity. Newtonian gravity worked to 10⁻⁹ precision. The only anomaly was Mercury's perihelion which was explained away as a hidden planet ("Vulcan"). There was no inductive signal. No error gradient. No dataset to compress.
Thus, for bringing AI to be able to be abductive: the paper proposes that the path to true AI invention is physically consistent, multimodal world models — systems that can run counterfactual simulations, not just predict pixel sequences. Genie (DeepMind, 2024) is a first step: action-controllable generative environments where an agent can intervene, not just observe.
A well-built PBPK or QSP model is a world model. It encodes physical laws (mass balance, receptor binding kinetics, diffusion). It allows counterfactual simulation: what if the dose were higher? What if the target expression changed? What if the patient were pediatric? So the modellers that we are, we build and validate these structures is not just running simulations. In principle, we are building the synthetic laboratory that future agentic systems will need to reason abductively about drug action.
AI systematically favors data-rich domains. And that is where, I think one of the most important areas to exploit the current reception of GenAI is by being bullish on digital twin generation with this technology.
For now, the data poor domains are safe from AI abduction- what does this mean for us? This means rare diseases, novel mechanisms, and first-in-class modalities will always require human abductive reasoning. The modeler who can formulate a plausible mechanistic hypothesis in the absence of rich clinical data is doing exactly what no LLM can do.
Finally, the verifier architect is also the axiom architect
With this deepmind paper- we learn one thing- verifiers alone are insufficient; someone must still invent the axioms that the verifiers check against.
In MIDD, those axioms are:
In MIDD, those axioms are:
- The equivalence principle of your system: what physical law makes this drug-target interaction possible?
- The boundary conditions of your model: what must be true for this extrapolation to be valid?
- The regulatory principles: what does ICH M15 require for this context of use?
The modeler who defines these; who performs the Jump from clinical observation to mechanistic axiom is for now irreplaceable. As we approach singularity in science- this would change.