At the World Economic Forum in Davos this year, Elon Musk told Larry Fink, chair and chief executive of BlackRock and the Forum's interim co-chair, that we might have AI "smarter than any human" by the end of 2026, and that by around 2030 or 2031 it would probably be "smarter than all of humanity collectively."
The statement travelled. It reached my LinkedIn feed several times in the following week, sometimes as a warning, sometimes as an endorsement, always presented as the kind of thing that could turn out to be true or false.
This article is not an argument about whether he is right. I do not know, and the public claim as stated does not provide a way for anyone to determine whether it has been satisfied. It is about a more useful question. What happens to a consequential claim, and to the decisions taken downstream of it, when it enters circulation before anyone has defined what would establish it?
The Musk prediction is the specimen. The examination is the point. I want to show the reasoning rather than summarise it, including the parts where my own arguments fell apart.
My argument is that such a claim should not acquire evidentiary weight in planning or governance until its construct, measure, threshold, falsifying conditions, and decision rule have been specified.
The first objection, and why it failed
My initial position was the common one. These systems are trained on human output. Everything they produce is drawn from what people have already written, measured, or recorded. On that reasoning, human knowledge functions as a ceiling. A system might exceed any individual, but it cannot pass the collective boundary of what humans have observed.
I held that for about a day.
It does not survive contact with the record.
AlphaGo Zero was given the rules of Go and learned to play through self-play, without training on human games. After three days of training it defeated AlphaGo Lee, the previously published system that had beaten Lee Sedol, by one hundred games to none. Unlike AlphaGo Zero, the earlier system had used supervised learning from expert human play.
AlphaTensor later discovered exact matrix-multiplication algorithms that improved on the best previously known algorithms for several matrix sizes. In one notable case, multiplication of 4 × 4 matrices over a finite field, it improved on Strassen's two-level method for the first time in roughly fifty years.
Those are not retrievals.
I want to state the conclusion carefully, because the obvious phrasing overreaches. With large trained systems it can be extraordinarily difficult to establish exactly what information was latent in the training data. The claim worth making is therefore weaker, but sufficient: AI systems can produce valid solutions that were never supplied to them as solutions. AlphaGo Zero is the cleanest case, because the exclusion of human game data was part of the training design itself.
There is a further problem with the ceiling argument, and it is the one that ended it for me. Humans have no access to knowledge outside human knowledge either. We generate it by acting on the world and checking what comes back. There is no external reservoir from which people retrieve discoveries before making them. Requiring an AI system to reach beyond previous human observation sets a standard that human beings themselves do not meet.
I am leaving this reversal on the page deliberately. An argument that collapses under examination is worth more visible than deleted, because the collapse is part of the demonstration.
Two kinds of feedback, and why the difference matters
The examples above share a property that is easy to state badly. Go and matrix multiplication provide unusually cheap and objective feedback. A game ends in a win or it does not. A proposed matrix-multiplication algorithm can be tested for correctness and evaluated for the number of operations it requires. A system can generate candidates, evaluate them, and redirect its search without waiting for a laboratory, a clinical trial, or an institution to tell it what happened.
Protein structure prediction is a different case, and I originally used it as though it were the same one. AlphaFold can predict structures that have not been experimentally resolved, and that is part of what makes it valuable. But confirming a novel structure against the physical world can still require methods such as X-ray crystallography, cryo-electron microscopy, or NMR. AlphaFold predictions are exceptionally useful hypotheses, but even high-confidence predictions can differ from experimental structures. A model's internal confidence in a prediction is not the same thing as external empirical confirmation.
So there are at least two loops worth distinguishing. The inner loop is generation and evaluation during optimisation, inference, or search, and it can be fast, automatic, and enormous in volume. The outer loop is confirmation against the physical world, and it can be slow, expensive, and in many fields it may become the binding constraint.
Where both loops are cheap, the results can be extraordinary. Where the outer loop remains expensive, AI can generate better candidates at much greater speed without making the need for verification disappear. The bottleneck moves.
Medicine, economics, and institutional behaviour sit toward the difficult end of that spectrum. Ground truth can be slow, expensive, confounded, contested, or all four. That does not mean AI cannot make major advances in those fields. It means the rate at which hypotheses can be produced and the rate at which the world can tell us whether they are right are different quantities.
Compression, and where the framing breaks
What survives after my first argument collapsed is narrower. Some of the results AI systems are producing are things human researchers might eventually have reached by other means. Scientists were already solving protein structures. Mathematicians were already searching for better algorithms. The work would have continued.
What clearly changed was the rate.
Dario Amodei has described a possible "compressed 21st century," predicting that AI-enabled biology and medicine could deliver within five to ten years the progress human researchers would otherwise have achieved over fifty to a hundred.
Compression rather than transcendence.
He treats that as transformative. It can be both. But two limits need marking before anyone leans too heavily on the compression framing.
The first is that my version carries a defect similar to the one I am about to identify in Musk's. "Humans would have found it eventually" cannot usually be tested. No experiment runs the counterfactual history in which AI did not exist and reports when the discovery would have occurred. If I criticise one unverifiable prediction and answer it with another, I have advanced very little. The most I can claim is that acceleration is observable while the counterfactual endpoint is not. Calling the phenomenon compression therefore requires fewer assumptions than claiming a system has crossed some boundary beyond human possibility, but that is an argument from parsimony. It is not proof.
The second limit is more interesting, and I think it undermines the augmentation framing more than I first allowed. Compression at sufficient magnitude stops behaving like a difference in degree. If a discovery that would otherwise have taken four centuries arrives next Tuesday, saying that humans would have reached it eventually may be true and yet carry almost no explanatory weight. For everyone living through the consequences, the distinction between humanly possible eventually and available now only because of the machine starts to collapse.
That suggests compression versus transcendence may itself become a false dichotomy at the extremes. The distinction does useful work at ten times. It may stop doing useful work somewhere before ten thousand. I do not know where that threshold sits, or whether a single threshold exists at all, and that seems to me a more valuable open question than arguing over which label wins.
Does speed make something smarter?
Here the answer splits, and my first version was too simple. I originally wrote that speed was sufficient on its own in domains with checkable answers.
That is wrong.
Checkability does not imply tractability. A password guess can be checked almost instantly. So can a chess position. A candidate proof may sometimes be mechanically verified. None of that means raw computational speed will reliably find a good solution, because the search space can be so large that exhaustive search remains impractical regardless of how quickly individual candidates are checked.
AlphaGo Zero's achievement was therefore not simply searching faster. It learned which parts of an enormous search space were worth examining. Representation, learned policy, abstraction, and search direction mattered alongside computational scale. Capability in these settings depends on at least three things: how much search can be performed, how intelligently the search space is navigated, and how informative the feedback is. Speed matters enormously. It is not the whole mechanism.
Judgment-heavy problems complicate the picture further. Additional computation can improve a judgment process. A system can examine more evidence, generate more alternatives, test more counterfactuals, search for contradictions, and subject its own conclusion to adversarial critique. But speed cannot supply the evaluative standard itself. Deciding which question matters, whether the frame is wrong, what evidence should count, or how competing values should be traded when there is no single correct ordering requires something beyond throughput. Running a defective evaluative process faster may produce more conclusions without producing better ones.
Worth adding that speed is not foreign to human intelligence. Processing speed and working memory already appear in major psychometric models of cognitive ability, and we count quickness as part of being smart in people. The harder question is what additional properties have to accompany it before we use the same language for a machine.
What is "smarter" supposed to measure?
The disagreement between Musk and many of his critics cannot be resolved simply by looking at more systems. Both sides can observe the same capabilities and reach opposite conclusions, because they are often working with different constructs and different ideas about what should count as evidence.
This is a measurement problem before it is a philosophical one. Five questions have to be separated. What is the construct being claimed? What observable measure represents it? What threshold counts as satisfying the claim? What result would count against it? And what decision rule tells us how conflicting evidence should be resolved?
At least four working conceptions of intelligence are relevant.
Psychometric intelligence
Psychometric approaches treat intelligence partly through the common factors that account for covariance among human cognitive abilities. These constructs were developed to explain patterns in human populations, and whether they transfer cleanly to systems with very different architectures and capability profiles is an open measurement question.
Researchers have begun examining whether artificial systems exhibit latent performance patterns comparable to general cognitive ability. Uneven performance alone does not settle the matter, since humans also have uneven cognitive profiles. The issue is construct validity. A number can be produced without establishing that the number measures what we mean when we say an AI is smarter.
Task capability
Can the system do the things? This is highly operationalisable, which is why benchmarks dominate so much public discussion. Its weakness is not measurability but interpretation.
Benchmark results can be affected by contamination, task selection, prompting, scaffolding, and the relationship between the benchmark and the environment in which the system will actually operate. A list of tasks successfully completed tells us something important. It does not automatically tell us what general capacity produced that performance.
Goal achievement across environments
Shane Legg and Marcus Hutter surveyed many existing definitions of intelligence and later developed a formal account built around an agent's ability to achieve goals across a wide range of environments. Generality becomes central.
This conception is comparatively friendly to Musk's intuition, because breadth is one of the dimensions along which modern AI systems have advanced rapidly. Turning the formal idea into a practical measurement of real-world systems remains another matter.
Skill-acquisition efficiency
François Chollet shifts attention from accumulated skill to the efficiency with which genuinely new skill is acquired. The question is not merely what a system can do, but what priors and experience were required for it to learn to do it. The Abstraction and Reasoning Corpus, or ARC, emerged from that line of thinking.
I want to be careful here, because the intuitive criticism of machine learning often goes too far. It is tempting to say that a child can learn a rule from three examples while a model consumes more text than a person could read in a lifetime. But that comparison counts model pretraining while treating human pretraining as free. The child arrives with evolved biological priors and substantial developmental experience, since years of perception, physical interaction, language acquisition, and social learning have already occurred. The model arrives with architecture and pretraining. Neither starts from zero.
Comparing their learning efficiency fairly therefore requires deciding what belongs in the accounting on both sides. That is not obviously unfair to the machine. It is difficult to make fair in either direction, and that difficulty is itself informative.
Each of these conceptions can, at least in principle, be operationalised quantitatively. None has been nominated as the referent of the claim we started with.
It is worth noting which one is most visible in practice. Public industry reporting predominantly foregrounds task capability, because benchmark results can be produced at scale and communicated comparatively. An executive hearing that a system is smarter than a person is generally thinking of something closer to general cognitive ability, or to broad competence across unfamiliar situations. The measurement and the inference are drawn from different constructs. That gap is where the governance problem begins.
The collective version
"Smarter than all of humanity collectively" needs separate treatment, and my first draft overreached here too. I wrote that the statement had no unit of measurement. As stated, no unit has been supplied, and that is defensible. What is not defensible is the stronger claim that no meaningful test could ever be constructed.
Someone could try. Imagine a very large set of previously unseen problems spanning mathematics, engineering, medicine, scientific research, strategic planning, software development, and other domains. On one side is the AI system. On the other is the best output obtainable from human institutions allowed to collaborate under fixed conditions. The evaluation criteria, time limits, and thresholds are established in advance.
The treatment of AI assistance on the human side would have to be specified. Permitting comparable frontier systems would turn the exercise into AI versus AI-assisted humanity. Prohibiting them would impose an artificial restriction on contemporary institutions. Neither choice is neutral, but either could be made explicit. The methodological problems would remain severe.
It would still be a test.
But another problem immediately appears. That test would not measure the combined intelligence of humanity in some pure sense. It would measure AI performance against contemporary human institutional performance on a specified set of problems under specified conditions. Those are not the same construct, and the distinction matters. A test can be perfectly repeatable while measuring the wrong thing.
Reliability is not validity.
So the accurate criticism of Musk's statement is not that no test could ever exist. It is that no construct, measure, threshold, or decision rule has been supplied. That difference looks small. It does much more work than the stronger claim, because it is not defeated by somebody proposing a conceivable experiment.
What I know, what I am assuming, what I cannot determine
What the record supports. AI systems have produced valid solutions that were never supplied to them as solutions, most cleanly in domains where candidate solutions can be generated and evaluated efficiently within the computational process itself. That is documented and reproducible.
What I am assuming. Gains in fast-feedback domains will not transfer automatically or proportionally to domains where external validation is slow, expensive, confounded, or contested. I find the verification-bottleneck argument persuasive. I cannot prove it, and a sufficiently large improvement in automated experimentation, simulation, or other forms of reliable validation would weaken it.
What I cannot determine. Whether Musk's 2030 to 2031 claim is true. Not simply because those years have not arrived, but because the statement has not been specified precisely enough for competing observations to produce an agreed verdict. No construct is named. No measurement is proposed. No threshold is established. No decision rule explains what to do when performance is superhuman in some respects and inferior in others.
That third item is the finding, and it is worth being precise about what kind of finding it is.
An underspecified claim is not the same as an empirically inaccessible one. "There is microbial life beneath the ice on Europa" may currently be difficult to settle, but its truth conditions are reasonably clear. If we obtained and reliably analysed a sample containing extraterrestrial microorganisms, we would know what that evidence meant. The difficulty is access to the evidence.
The Musk claim has a different problem. The evidence could be sitting in front of us and reasonable parties could still disagree about whether the threshold had been crossed, because the threshold itself was never defined. It has not yet been specified precisely enough for checking to yield an agreed verdict.
That is a repairable defect.
It has not been repaired.
Why this belongs in a governance conversation
A prediction on a conference stage is not itself a governance failure. What happens to it downstream can become one, and the risk is easy to describe. A statement begins as "Musk predicts X." It becomes "industry leaders expect X." Then "AI capability is projected to reach X." And eventually "given anticipated AI capability, we should assume X."
Notice what happened. The proposition did not acquire additional evidence. It acquired institutional authority through repetition, and at the same time its provenance, uncertainty, and original status as one person's prediction became less visible.
I am describing a possible pathway here rather than reporting a specific case. I have not audited a procurement file and established that this particular statement travelled through those stages, and I am not going to imply that I have. But related problems have been examined in several different literatures. John MacFarlane used the term knowledge laundering in analysing how epistemic status can change when a proposition moves between contexts with different standards. Statisticians have used uncertainty laundering to describe processes that begin with uncertain evidence and end with unjustifiably binary declarations. More recent work has used terms such as epistemic laundering and prediction laundering to examine different ways in which weak warrant, uncertainty, or contested assumptions can acquire an appearance of authority as they pass through technical or institutional systems.
These concepts are not identical, and collapsing them into one phenomenon would repeat exactly the mistake this article is warning against. But they point toward a common governance concern. Does the evidentiary status of a claim survive the journey from its source to the decision that relies on it?
That is the question I care about. No fabrication is required. No one has to act in bad faith. Musk may hold his prediction sincerely. The failure occurs when a statement loses its qualifiers as it travels and eventually arrives at a decision with the appearance of a technical input, without having acquired the evidentiary properties of one.
Underspecification makes that kind of claim unusually resistant to decisive contradiction, because evidence can appear to count against it and the meaning of the claim can move. Smarter can become better at most economically valuable tasks. All of humanity can become better than the best individual experts. By 2030 can become approximately around the beginning of the next decade. The more elastic the original proposition, the easier it is to preserve after the evidence changes.
That is why operational definitions matter before a claim becomes embedded in governance. They do not merely make measurement possible. They make goalposts harder to move.
None of this is an argument for waiting. Institutions act under uncertainty as a matter of routine, and refusing to respond to a risk until it has been perfectly characterised would be its own failure. The distinction is between acting on an uncertain claim while treating it as uncertain, and acting on an unspecified claim while treating it as settled. The first is ordinary risk management. The second produces controls aimed at something nobody has described, which is how organisations end up with safeguards that cannot be evaluated because the thing they were built against was never defined.
The method, which is the actual point
I began with an argument that was wrong and abandoned it. I replaced it with a narrower argument and marked the conditions under which that argument weakens. I separated fast internal evaluation from external empirical confirmation, corrected my account of what AlphaGo Zero demonstrated, and pulled back claims about computational speed and psychometric comparison. I replaced cannot be measured with the narrower and defensible has not been specified.
I ended without an answer to the prediction that started the article, because the prediction as stated does not yet admit an agreed answer.
That sequence is not a weakness in the analysis.
It is the analysis.
A method worth applying to other people's claims has to survive being applied to your own. The useful discipline is therefore not a verdict. It is the same set of questions, asked before an argument has been allowed to run for a year. What is the construct being claimed? What observable measure represents it? What threshold would establish it? What result would count against it? What decision rule handles evidence that points in different directions?
Those five questions are the mechanism, and they cost very little. They can be run in the time it takes to read a paragraph. The place for them is the point at which a capability claim first enters a planning document, a business case, or a risk register, rather than the point at which someone asks why a control was built the way it was.
Applied here: name the construct, the measurement, the threshold, and the observation that would show that AI is not smarter than all of humanity collectively. If nobody can supply those, the appropriate response is not yet agreement or disagreement. It is to note that the proposition has not been specified sufficiently to deserve evidentiary weight.
Not yet.
That qualifier is doing real work. The claim may become evaluable. Someone may specify it properly, evidence may accumulate against an agreed test, and at that point it will deserve a serious answer.
Until then it is a statement in circulation. And the fact that a statement is circulating is evidence of its reach. It is not evidence of its truth.
Kevin V. Watson is a Digital Forensics and Incident Response specialist and a graduate student in the Master of Interdisciplinary Artificial Intelligence program at the University of Ottawa. He is the author of the Zemi Method, a published doctrine on investigative reasoning and evidentiary boundaries. He writes at the intersection of digital forensics, cybersecurity, AI governance, and the public interest.
Sources and further reading
- World Economic Forum. "Davos 2026: Conversation with Elon Musk," Meet the Leader podcast transcript.
- https://www.weforum.org/podcasts/meet-the-leader/episodes/conversation-with-elon-musk-davos-2026/
- Silver, D., et al. (2017). "Mastering the game of Go without human knowledge." Nature, 550, 354–359.
- https://www.nature.com/articles/nature24270
- Google DeepMind. "AlphaGo Zero: Starting from scratch."
- https://deepmind.google/blog/alphago-zero-starting-from-scratch/
- Fawzi, A., et al. (2022). "Discovering faster matrix multiplication algorithms with reinforcement learning." Nature, 610, 47–53.
- https://www.nature.com/articles/s41586-022-05172-4
- Jumper, J., et al. (2021). "Highly accurate protein structure prediction with AlphaFold." Nature, 596, 583–589.
- https://www.nature.com/articles/s41586-021-03819-2
- Terwilliger, T. C., et al. (2024). "AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination." Nature Methods, 21, 110–116.
- https://www.nature.com/articles/s41592-023-02087-4
- Amodei, D. (2024). "Machines of Loving Grace: How AI Could Transform the World for the Better."
- https://darioamodei.com/machines-of-loving-grace
- Legg, S., & Hutter, M. (2007). "A Collection of Definitions of Intelligence."
- https://arxiv.org/abs/0706.3639
- Legg, S., & Hutter, M. (2007). "Universal Intelligence: A Definition of Machine Intelligence."
- https://arxiv.org/abs/0712.3329
- Chollet, F. (2019). "On the Measure of Intelligence."
- https://arxiv.org/abs/1911.01547
- Ilić, D., & Gignac, G. E. (2024). "Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?"
- https://arxiv.org/abs/2310.11616
- MacFarlane, J. (2005). "Knowledge Laundering: Testimony and Sensitive Invariantism." Analysis, 65(2), 132–138.
- https://academic.oup.com/analysis/article/65/2/132/139574
- McShane, B. B., Gal, D., Gelman, A., Robert, C., & Tackett, J. L. (2019). "Abandon Statistical Significance." The American Statistician, 73(sup1), 235–245.
- https://doi.org/10.1080/00031305.2018.1527253
- "Epistemic laundering: generative AI and the naturalization of misrecognition." AI & Society (2026).
- https://doi.org/10.1007/s00146-026-03068-9
- Rohanifar, Y., Ahmed, S. I., & Sultana, S. (2026). "Prediction Laundering: The Illusion of Neutrality, Transparency, and Governance in Polymarket."
- https://arxiv.org/abs/2602.05181