The key challenge to establishing justification for our scientific beliefs based on what AI models produce is solvable.
![]()
“What time is it?” Your friend looks out the window and reads 3:00 from the clock tower outside. You read 3:00 from your clock on the wall. You both answer “3:00,” and you are both right. Who between you is the better scientist?
Science has always been concerned with inference – that is, learning about hidden properties of the world from observable data. In the field of mathematics, this might be called the “inverse problem,” and because this a hard problem, scientists rely on principles for structuring their investigation into hidden properties in a built-out form of a standard inductive argumentative structure encoded as the “scientific method.” This method is typically posed in the language of the hypothesis, an assumption about an unobservable property of an object and the prediction concerning hypothetical observable data entailed by it; an experiment, the observation of data generated by the system inhabited by the object of interest, possibly having undergone some intervention on the system; a test of whether the observations are compliant with the hypothesis; and the conclusion, which proceeds from the test outcome to state the hypothesis’ standing in context of observation.
Some data are informative about subjects of interest but very hard to summarize in a human-readable way. The European Space Agency’s Gaia mission, for example, has observed about two billion unique stars. To index Sun-like stars from that catalog would be practically impossible if done manually by humans. In the field of high-energy particle physics, evidence for many theories beyond the Standard Model (SM) has been hard to find despite the fact that the large Hadron collider at CERN has generated the largest scientific data set in the world. Automated tools for subselecting data which might be useful for science are essential for fast generation and storage of experimental data, let alone for performing tests and analyses.
For a scientist to equip themselves with the habit of drawing true conclusions from data of any complex sort, they must rely on extensions of their own mind and use computers. This fact has been recognized for decades. [Explain] Today, computing is ubiquitous in science, and two broad classes of methods have risen to the forefront of discourse across fields: machine learning (ML) and artificial intelligence (AI). By different accounts, ML is a subclass of AI, or vice versa. For our purposes, I will insist on a distinction between these two in terms of the intended task of constituent methods. ML (as a class of methods) consists of methods developed for tasks for which the human mind is ill-equipped, including but not limited to data visualization, compression, and transformation. AI, on the other hand, is concerned with tasks for which the human mind is well-equipped, including playing board games, driving cars, and processing natural language.
Neither ML nor AI are essentially scientific in nature. For example, we might call a particular reinforcement learning algorithm – or a specific code implementation of the algorithm; or the algorithm, its code implementation, and the hardware running the code – tasked with controlling the plasma inside a nuclear fusion machine like the Tokamak an ML-based system, because it’s got to operate in real time at beyond-human speeds on a very complex collection of physical objects. On the other hand, a similar algorithm written to learn and play chess is an AI-based system, even if mathematically the algorithms rely on similar principles or their code implementations are parameterized by similarly defined neural network architectures. I choose to be radically accepting of the “intelligence” designation for such methods, not as a commitment to the personhood or consciousness-like behavior of modern AI systems, but as a helpful signifier of their intended use.
| Machine learning (ML) | Artificial intelligence (AI) |
|---|---|
| Computing for tasks for which the human mind is ill-equipped | Computing for tasks for which the human mind is well-equipped |
In principle, the scientific method remains the same in the age of ML and AI, but its practice is extendable now to new domains of inquiry by virtue of these methods’ ability to pre-process the human-operated steps of the method. At best, therefore, ML and AI offer a strict improvement to the enterprise of doing science to the extent that they extend its scope. The question remains of how we realize and enjoy this promise.
A scientist’s job is not to establish facts by inference on hidden properties of the world. Instead, their task is to extend the human capacity to reason – and knowledge of facts can help. Educated guesses about universal laws from a small number of poorly controlled experimental data points would count as activity among the stuff of science, but it’s not scientific if it’s not directed toward uncovering the truth with the sharpest tools available either in the literature or in the lab. This has raised a challenge for the individual scientist in the age of the internet. There is no human way to maintain discourse in context of all of the information that is available online. However, personally knowing and understanding everything that has been done on a subject is not required in order to participate in a line of inquiry with reasons behind every commitment made within that subject.
In other words, ignorance does not preclude rationality or the good conduct of science.
I argue for a differential on AI usage in the procedure of the scientific method on the basis that reliance on it can be rational or irrational. Individuals can and should invoke expert opinions when wading into waters or charting a course where they haven’t their own expertise; in the context of AI usage, those experts are the technologists behind the technology, not the AI model itself. Deference does not indicate unscientificity. Instead, the moment at which discursive reasoning truncates itself via deference determines the strength of the resulting scientific contribution. This is uncomfortably conservative, perhaps. The implication is that reporting the result of private engagement with an AI model in the analysis of data, with deference to the model’s reliability as the form of justification for the report as being scientific, satisfies this condition of being science. The merit of any such work is not in its being scientific or the originality of the facts that it communicates but rather its extension of reason beyond what has been established.
Indeed, sound reasoning comes in all sorts of strengths. Let \(X\) denote a proposition of fact in the form of a sentence. “An AI model printed \(X\)” is justification for reporting \(X\) as fact. “An AI model printed \(X\), and this AI model has been known to print established facts; therefore, \(X\) is a fact” is a reasonably strong argument whose strength is related to how one might specify the phrase “has been known” in this context. “An AI model printed \(X\), and I have personally and independently checked that this AI model reproduces facts; therefore, \(X\) is a fact” starts to sound like the best quality of science that can be offered.
Without some independent certification that the output of an algorithm accords with the claim of the algorithm’s design intentions, reliance on the explanatory mechanism of that algorithm is ill-justified. By analogy, reading the time as 12 o’clock noon when a clock reads that time does not confer knowledge of the time without some certification that the clock is not broken. In the same way that the channel for justification in the reporting must cohere with the underlying causal mechanism, so too must algorithms be checked for coherence with the system which they purport to model.