Skip to content
IC
All posts

AI research

The state of AGI: what the evidence supports, and what it does not

A sober read of the AGI evidence in 2026: what frontier systems can measurably do, what expert forecasts actually say, how seriously to take the existential question, and what the research shows about humans and AI working together.

Mathew Sayed Mathew Sayed
· · 10 min read

Two claims about artificial general intelligence are circulating through Australian boardrooms at the moment. The first, usually arriving via a vendor or a keynote, is that AGI is two to three years away and every operating model should be rebuilt now. The second, usually arriving via a burnt early adopter, is that the whole thing is a bubble and the sensible move is to wait it out. Both are positions. Neither is an analysis.

What makes this moment unusual is that you no longer have to choose between vibes. AI capability is now measured, repeatedly and publicly, by organisations with no product to sell, and expert forecasts are collected on a standing cadence rather than harvested once a decade. The evidence does not resolve the AGI question. It does something more useful: it tells you which claims are supported, which are speculation, and where the honest uncertainty sits. This piece is our read of that evidence, written for the executives and boards who have to make decisions inside the uncertainty rather than argue about it.

The definition does most of the arguing

Start with the inconvenient part: there is no agreed definition of AGI, and nearly every dispute about timelines is a dispute about definitions wearing a costume.

If AGI means an AI system that outperforms humans at most economically valuable work (the definition OpenAI’s charter uses), the question turns on labour economics as much as capability. If it means human-level performance across the full range of cognitive tasks, current systems are far away on some axes and past the bar on others. If it means a system that can do most things a remote knowledge worker can do over a full working day, the question becomes measurable, and that is precisely the direction serious measurement has taken.

When someone gives you a confident AGI date, the first question to ask is which definition they are using. The date moves by a decade or more depending on the answer, and people arguing past each other on definitions is most of what the public debate consists of.

What frontier systems can measurably do

The most decision-useful capability research of the past two years comes from METR, an independent evaluation organisation that measures the length of tasks AI agents can complete autonomously. Three of its findings matter more than any benchmark score:

  • The time horizon is growing exponentially. The length of software and reasoning tasks that frontier agents can complete at a 50 per cent success rate has doubled roughly every seven months across six years of models, and the most recent model generations doubled it faster, closer to every four months. This is the single most important empirical trend in AI, because it converts “is AI getting better?” into a number with a slope.
  • Reliability lags capability, badly. The same measurements show frontier models succeeding at close to 100 per cent on tasks that take a skilled human under about four minutes, and under 10 per cent on tasks taking more than about four hours. A 50 per cent success rate is the threshold that is easiest to measure. It is not the threshold at which you would delegate work that matters, which for most business processes is 99 per cent or better.
  • Performance degrades inside long contexts. Research on agent reliability finds success rates that hold up on short, isolated tasks fall away sharply when the same task is embedded in a longer interaction history. Autonomy over minutes and autonomy over days are different engineering problems, and only the first is close to solved.

The International AI Safety Report 2026, chaired by Turing Award winner Yoshua Bengio and written by more than 100 researchers nominated by over 30 countries, reads the same evidence and reaches a consistent position: capabilities are advancing quickly and unevenly, with rapid gains in mathematics, coding, and scientific reasoning alongside persistent failures in long-horizon reliability and real-world robustness.

So the honest capability summary for 2026 is this. Frontier systems are genuinely, measurably powerful within a jagged and expanding envelope, the envelope is growing on a fast exponential, and nothing inside the envelope yet resembles a system you could leave alone with a full working week. Both halves of that sentence are true at once, which is exactly why the public conversation, which prefers one half at a time, stays so bad.

What the forecasts actually say

Expert forecasts on AGI have two consistent properties: the central estimates keep moving closer, and the spread between forecasters remains enormous. Both properties are data.

SourcePopulationCentral estimate
AI researcher survey, 2,778 respondents (Grace et al.)Published AI researchers50% chance of high-level machine intelligence by 2047, a 13-year drop from the same survey one year earlier
Longitudinal Expert AI Panel, wave 8 (2026)Experts and calibrated forecastersOn average, 25% chance of AGI by 2029 and 50% by 2033
AAAI presidential panel (2025)AAAI community, 475 respondents76% say scaling current approaches is unlikely to yield AGI on its own
Samotsvety forecasting group (2023)Elite generalist forecastersRoughly 28% chance of AGI by 2030

Read carefully, the table says three things. First, the people closest to the technology have shortened their timelines dramatically and repeatedly; a median that falls 13 years in a single survey cycle is not a stable expert consensus, it is a field revising itself in real time. Second, a large majority of the academic community does not believe the current recipe, scaled up, gets there by itself; something further has to be invented, and inventions do not run on schedules. Third, serious, calibrated forecasters now place real probability mass on transformative capability inside this decade, and dismissing that as hype requires ignoring the measured capability trend in the previous section.

The dispersion is the finding. When credible estimates for the same event span 2029 to 2047 and beyond, the rational planning posture is not to pick a favourite date. It is to make decisions that remain sensible across the range, which is a governance problem, not a forecasting one. We come back to that at the end.

On the threat to human existence, honestly

Asked directly whether AI threatens human existence, the honest answer has three parts, and all three belong in the answer.

The concern is not fringe. In 2023, the Center for AI Safety’s one-sentence statement, that mitigating the risk of extinction from AI should be a global priority alongside pandemics and nuclear war, was signed by the chief executives of the leading AI laboratories and by Turing Award winners including Geoffrey Hinton and Yoshua Bengio. In the largest survey of AI researchers, the median respondent put around 5 per cent probability on extremely bad outcomes, including human extinction, and a substantial minority put the figure at 10 per cent or higher. People who build these systems for a living assign the risk a probability you would never accept from an aircraft or a reactor.

The concern is also genuinely contested. The same survey shows wide disagreement, with many researchers placing the probability near zero, and the AAAI result above implies most academics doubt current systems are even on the path that the loss-of-control scenarios require. The International AI Safety Report treats loss of human control as a risk worth managing while being explicit that experts disagree about its likelihood and mechanism. Anyone quoting a single confident number, in either direction, is performing certainty that the field does not possess.

And the risks with evidence behind them today are nearer to the ground. The 2026 report’s sharpest warnings are not science fiction: AI-enabled deepfakes, biological uplift, and cyberattack capability are documented, current, and worsening. For an Australian organisation, the existential question is primarily one for governments and frontier laboratories, and the appropriate response is to support serious regulation and safety research. The operational questions, deepfake-enabled fraud against your finance team, AI-accelerated phishing, ungoverned AI handling your customer data, are yours, this quarter. A board that spends its AI risk discussion on superintelligence while its staff paste customer records into unsanctioned chatbots has the risk register upside down.

Our position, stated plainly: the existential risk is real enough that dismissing it is unserious, uncertain enough that planning your business around it is unserious too, and no excuse for ignoring the concrete harms that are already measurable. Hold all three of those and you are ahead of most of the commentary.

What the research says about humans and AI working together

Whatever the AGI endpoint, the next decade of work is humans and AI systems operating together, and this is the part of the question with the strongest empirical footing, because it is being tested in field experiments rather than forecast.

The picture the experiments paint is consistent and more interesting than either the replacement narrative or the productivity-miracle narrative:

  • Inside the capability envelope, gains are large. The Harvard and BCG field experiment on knowledge workers (Dell’Acqua et al.) found consultants using AI completed tasks 12.2 per cent faster with roughly 40 per cent higher quality on tasks within the model’s competence.
  • Outside the envelope, AI actively harms performance. The same study found performance 19 percentage points worse when subjects used AI on tasks just beyond its competence, because the tool is equally fluent when it is wrong. The researchers called this boundary the jagged frontier, and it is the single most useful concept an executive can carry into AI adoption.
  • Even experts misjudge the frontier. A METR randomised controlled trial found experienced open-source developers were 19 per cent slower with early-2025 AI tools on mature codebases they knew deeply, while believing the tools had sped them up. Perceived productivity and measured productivity diverge, which is why anecdote-led adoption fails.
  • Team structures are changing shape. Field experiments on human-AI teams report productivity per worker rising by around 50 per cent, with AI agents functioning as junior collaborators rather than tools. The skill that separates effective operators is calibrated reliance, knowing when to trust the system and when to override it, and the evidence says that skill is trainable.

Notice what this adds up to. The collaboration future is not humans versus AI, and it is not humans passively supervising flawless machines. It is organisations learning, task by task, where the frontier sits for their work, moving the routine and verifiable side of it across, and keeping human judgment on the side where reliability, accountability, and context still decide outcomes. That is not a transitional arrangement to be endured until AGI arrives. On the current reliability evidence, it is the operating model for the foreseeable planning horizon, and the organisations treating it as a designed system, with evals as the control layer and governance attached from the start, are the ones seeing the measured gains rather than the anecdotes.

What a mid-market organisation should do with all this

The evidence supports a posture, not a prediction. Four moves follow from it directly:

  1. Plan for capability growth without betting on a date. The measured trend says systems will keep absorbing longer and more complex tasks. Decisions that are robust across the 2029-to-2047 spread, building AI literacy, instrumenting your processes, keeping exit options open with vendors, beat decisions optimised for one forecast.
  2. Map your own jagged frontier. The experiments say the payoff is task-specific. Pilot with measurement, not sentiment, and expect some confident use cases to fail the test.
  3. Govern before you scale. Every study of failure in this space runs through the same gap: nobody knew where AI was operating, on what data, with what oversight. Knowing what you are running is the foundation, which is why we treat the AI register as the first governance artefact, not the last.
  4. Assign the risk to someone. The near-term harms are concrete and the accountability for them is yours. If nobody owns AI risk in your organisation, that is the gap to close before any strategy discussion, and a governance posture assessment is the fastest honest way to size it.

AGI may arrive on the forecasters’ shorter timelines, the academics’ longer ones, or a path nobody has predicted. Organisations that measured their way through the uncertainty will be fine in every version. Organisations that picked a narrative will be right only by luck.

Sources and further reading

Get started

Bring AI risk under board oversight in two weeks.

A thirty-minute discovery call costs nothing. We confirm fit, scope, and timing, then issue a fixed-fee statement of work within two business days.