Over the years sentiment has swung back and forth in assessing machine intelligence: Knowledge was the currency for expert systems and ontologies. Mastering Chess logic was at one point the ultimate test of human-like intelligence. Deep Blue proved that wrong. Next IBM Watson with Jeopardy. Again, no cigar.
Now with LLMs and Generative AI we have a slew of benchmarks and tests aimed at measuring progress towards AGI (what is Real AGI) :
IQ, coding, math, law, medical, language competence, question answering, reasoning, and even tests supposedly specifically designed for AGI.
Instead of analyzing these various tests let’s look at a very prominent recent paper covering the definition and measure of AGI (ref) — it highlights the problem of current benchmarks quite well. The fact that some 26(!) AI experts claim authorship to this piece makes it a good reference point. Here’s a key diagram
The 10 dimension shown are the basis of their ‘AGI score’. GPT5 achieves 57%.
But does this approach have construct validity — does it properly measure real AGI? What if the dimensions are wrong or incomplete?
Then it would be like me saying my Tesla is 75% airplane: It is fast, carries passengers, and has autopilot.
That is essentially where we are.
To know what needs to be measured one has to understand what makes human intelligence so powerful.
Unfortunately almost nobody leading AI efforts right now is doing that or getting it right. Some cling to the idea that knowledge and/ or reasoning are intelligence, others explicit state that they don’t know what intelligence is: Demis Hassabis says that we need to build AGI in order to understand intelligence, while Elon Musk claims that nobody understands it (personal conversation).
Fact is that we can and must understand human intelligence from first principles for us to develop it effectively. However, we need to start with epistemology and cognitive psychology to figure it out. Not statistics, mathematics, or computer science. I summarize these blind spots towards AGI as ‘The 7 Deadly Sins of AGI Design’.
In short, what makes human intelligence so special is the ability to quickly adapt to changing circumstances, by learning incrementally in real time and to form contextual abstraction on-the-fly. We also leverage meta-cognition to monitor and control our thought processes (System1 and 2 thinking).
AI born from these insights is called Cognitive AI, or what DARPA calls “The Third (and final) Wave of AI“. LLMs contain knowledge and skills that are off the charts relative to humans, but they do not posses the unique qualities required for real intelligence.
There are other considerations about focusing on specific benchmarks alone:
Does benchmark performance translate into real-world performance?
Does optimizing for some benchmark(s) cripple performance in other areas?
Does a focus on near-term benchmark performance detract from the bigger goal?
Numerous studies have shown that good test performance does not actually translate well to practical deployment. Passing a law or medical exam may make for a great research tool but it cannot replace lawyers, doctors, or medical researchers. A great score on a language test does not automatically allow the system to provide good customer support. Acing a coding test does not turn an LLM into a programmer.
One reason for this is that answers to these test (or very similar ones) may already be in the training data. Secondly, in order to beat existing scores — for funding or bragging reasons — systems are often highly optimized for just better results, not general ability. Such narrow fine-tuning is usually at the expense of other capabilities.
However, the main reason that systems underperform in practice is that they simply do not have the cognitive ability to cope with real-world complexity and change. They don’t have real general intelligence.
Targeting specific benchmarks makes perfect sense for narrow AI application (like Tesla targeting accident-free miles), but even here it is crucial to use appropriate metrics (making sure you have a meaningful highway to urban ratio). However, when is comes to long-term goals then short-term measures such as narrow tests, working on impressive mock-up demos, or incremental research goals can massively distract from achieving the actual outcome (for example, impressive but shallow robot performance). This is particularly acute for AGI.
Assessing AGI Progress
AGI is the Holy Grail of AI. Unleashing the intelligence, competence and rationality of AGI promises to usher in a new Renaissance and Enlightenment — to significantly boost human flourishing. It will even help to improve human ethics and morality.
Assessing progress towards AGI is not easy. However, blindly chasing benchmarks will not get us there. We need a theory-driven approach.
Any AGI benchmarking has to start with understanding the key requirements of the kind of intelligence we’re trying to build. Some of these key requirements are:
Real-time incremental learning and concept formation. Robust integrated short- and long-term memory. Contextual and cause-effect reasoning, as well as temporal and spatial reasoning. Agentic meta-cognition. This capabilities all need to function with incomplete or incorrect data, and with limited time and compute.
Any system that does not have a clear path to providing all of this functionality should be eliminated out of hand as an AGI candidate, irrespective of how impressive its benchmark performance is. Big data, statistical systems like LLMs inherently cannot meet these requirements. As a prominent AI researcher put it: “LLMs are an off-ramp to AGI. A distraction. A dead-end".
In addition to crafting benchmarks that include the above dimensions (and to mark them as essential) systems have to be evaluated not only what what they achieve, but how they achieve it. For example, a system that taught itself Chess from scratch (like a human could) and plays a legal but poor game is much more impressive in AGI terms than a dedicated AI game engine. In fact, the seconds one is irrelevant for AGI.
ARC-AGI (ref), a prominent test developed specifically to measure AGI capabilities, while strong on visual analogical pattern generalization, suffers several benchmark limitations: It is rather one-dimensional in that it doesn’t cover language, memory, learning, or metacognition, and also ignores how the problems are solved.
In summary, AGI progress needs to be measured firstly by theoretical clarity both as far as understanding what is required, but also by a feasible technical roadmap on how it can be implemented using existing software techniques and hardware. Progress can then be measured via milestones and implantation-specific benchmarks.
The ultimate test for a fully developed AGI is that can learn a new task as easily or better/ faster than a smart, well-educated human, and without additional engineering or specialized training data. A general purpose AGI already trained to say STEM college-level (or potentially even less) should be able to learn to master new tasks and occupations like customer support, accounting, programming, research or management with no more time or data than a human counter-part.
Obviously, an AGI will be much more capable in certain respects because of its potentially photographic memory, instant access to data, much better reasoning ability, and 24/7 full focus. On the other hand, desk-bound AGIs will not have anywhere near the sense acuity or dexterity of a human. This will require significant development in robotics.
Why have we not seen more progress towards real AGI? The are number of reasons:
The (non-AGI) success of LLMs have ‘sucked the oxygen out of the air’ for other approaches
Some teams with good AGI potential did not get funded (sufficiently)
A lot of people don’t believe AGI is possible at all, or any time soon
Some people think that AGI is a bad or dangerous technology
The main reason though is that very, very few AI teams actually have a workable theory of what intelligence entails and how to achieve that with current technology.
The Narrow AI Trap
And then there’s the Narrow AI Trap.
Over the years I’ve seen many efforts trying to build AGI fall into this trap (including myself, with eyes wide open). The trap is, that in order to show progress one focuses on some specific capabilities that have little to do with what AGI development requires at that time, or at all. In most cases it actually retards AGI progress. The motivation for this can be to get papers published, to get funding, or just to feel good about seeing more intelligent-seeming functionality. The problem is that this new capability is largely hard-coded, hand-crafted, or specifically trained to achieve the desired result. AGI generally requires capabilities to be learned (interactively) and not to be specifically engineered using external human intelligence.
An extreme but common case is premature commercialization. Any budding AGI effort that choses to commercialize early version of its technology invariably ends up spending all of its time meeting customer requirements and doing all it can to quickly add urgent, but narrow functionality.
DeepMind demonstrates a clear case of falling into this trap. Their stated mission was “1. Solve AGI 2. Solve all the world’s problems”. Because they couldn’t come up with a workable theory or plan to achieve AGI, they started building (some very useful) narrow AI applications instead of focusing on 1. However good these systems are, they cannot be combined to make AGI; they are actually a distraction to get to AGI.
The only way to achieve AGI to to have a clear North Star roadmap and to pursue this relentlessly without distractions until the system reaches adult-level self-learning and reasoning ability. Then you have AGI!





Excellent analysis! It makes me wonder, if we struggle to define human intelligence, how can we hope to benchmarck true AGI? This is such a thought-provoking piece!
Makes good sense.