What changed in AI, what didn’t, and what industrial autonomy actually requires

About a year ago, I wrote about the path toward the “sentient and autonomous factory.” My thesis was that three technology layers were finally converging: decades of process automation, a decade or more of industrial data infrastructure, and a new generation of generative AI and agents. Together, I argued, they could enable plants to move from monitoring and analysis toward anticipating problems, reasoning about alternatives and eventually taking action.
A year later, I still believe that basic trajectory is right. What I got wrong was the speed.
AI capabilities have progressed considerably faster than I expected. At the same time, real-world industrial adoption has moved much more slowly. I don't think those observations contradict each other. If anything, the past twelve months have clarified what the real challenge was all along.
The bottleneck hasn't moved. It has simply become much more visible.
Three eras in less than a year
This isn't meant as a rigorous taxonomy, but from my perspective as a heavy user and builder of these systems, we have already moved through three distinct eras.
First came AI that answers.
Until late 2025, interacting with an LLM still largely meant sending it a question or prompt and getting something back. The models became remarkably capable, but the interaction remained fundamentally reactive. The human broke down the problem, asked the questions, evaluated the answer and decided what to do next.
At times, it felt like an extraordinarily sophisticated auto-complete.
Then came AI that executes.
By late 2025, something changed. Across the leading model families, AI became noticeably better at sustained reasoning and tool use. Instead of specifying every step, you could increasingly give the system an objective and leave it working: plan the task, gather information, use tools, inspect intermediate results, correct itself and continue.
There was no single model or release that caused this shift. GPT-5, Claude Sonnet and Opus 4.5, Gemini 3 and the later GPT-5.x models all demonstrated versions of the same emerging capability. For me, however, the change became particularly tangible around the end of 2025. AI was starting to feel less like a very capable auto-complete and more like something to which you could actually delegate work.
METR's “time horizon” research provides one way of seeing the change. The metric asks how difficult a task is in terms of the time it would take a skilled human to complete it, and measures the point at which an AI agent succeeds half the time. By early 2026, the most capable agents were already saturating METR's existing benchmark, with measured horizons extending beyond two human working days. METR now cautions that its current task suite is no longer reliable for precise measurements at those lengths. On a newer software-reimplementation benchmark, agents have successfully completed some tasks estimated to take humans weeks. (METR [1])
Now we are entering AI that delivers.
The interesting unit is increasingly not one LLM responding to one user. It is a team of agents, equipped with tools, data access and integrations, working together to produce an end product: An analysis. A report. An investigation. A piece of software. A recommendation. An actual artifact that somebody can use.
A particularly relevant benchmark is GDPval, designed around economically valuable professional work rather than exam questions. A more recent independent implementation, GDPval-AA v2, places models in an agentic environment with tools and asks them to produce finished deliverables such as documents, spreadsheets and presentations. Expert-human work is anchored at 1,000 Elo. By September 2026, several frontier systems were scoring substantially above that benchmark. This certainly does not mean that AI can replace professionals wholesale. But on a growing set of bounded professional tasks, the quality of the finished work is already competitive with - and in these evaluations often judged above - that of experienced human experts. (Artificial Analysis [2])
Stanford's 2026 AI Index describes the broader phenomenon well: evaluations designed to remain difficult for years are sometimes being saturated within months. At the same time, AI capability remains very uneven - models can demonstrate extraordinary mathematical reasoning while failing at tasks that humans find trivial. (Stanford HAI [3])
So I would not claim that AI is now generally “smarter than humans”. But frontier systems are clearly competitive with experienced humans across a rapidly expanding set of bounded reasoning and professional tasks.
That all happened much faster than I expected.
So why aren't our factories autonomous yet?
McKinsey's latest State of AI survey captures an interesting tension. Among companies with more than $1 billion in revenue, 40% now report scaling AI agents in at least one function, compared with 27% a year ago. Yet only 37% of respondents report that AI has made any positive contribution to enterprise EBIT—a number essentially unchanged from the year before. (McKinsey [4])
I don't find that particularly surprising.
We have arguably compressed many years of technological progress into a few quarters. What did we expect industrial companies to do? Industrial operations don't move on model-release cycles. They have installed equipment, operating procedures, maintenance regimes, control systems, cybersecurity requirements, investment cycles, shutdown schedules and - quite rightly - people who want evidence before handing consequential decisions to a machine.
AI can improve dramatically between Christmas and Easter. A 30-year-old chemical plant cannot.
The gap between AI capability and industrial adoption therefore isn't evidence that industrial AI has failed. It tells us something about the nature of the problem.
The bottleneck has been clarified
A year ago, it was still reasonable to ask whether AI itself would become sufficiently capable. Today, that is much less interesting. The hard questions are increasingly operational:
- Can the AI access the relevant information?
- Does it understand the relationships between the physical plant, its equipment, operating history and documentation?
- Can it distinguish normal operating variation from something that matters?
- Can it call the appropriate analytical models rather than simply inventing an answer?
- Can we trace its conclusions back to evidence?
- Do we know when it should escalate rather than act?
- And, most importantly: is it working on a problem valuable enough for anybody to change how they operate?
McKinsey argues that the limiting factor is increasingly organizations' ability to absorb change. I broadly agree, but I would put it slightly differently.
Industrial organizations have limited capacity for change. They always will. The real challenge is choosing initiatives important enough to deserve that capacity.
If an AI project addresses a major production constraint, material energy consumption, recurring quality losses or a significant reliability problem, people tend to find the time to make it work.
If an initiative isn't important enough to compete successfully for change capacity, I would question whether it should have been started in the first place.
We should therefore stop beginning with:
“Where could we use AI?”
Start with:
“Which operating decisions materially affect our business, and why aren't we consistently making the best possible decision today?”
Then work backwards.
From copilot to operational participant
This also changes how I think about “agents.”
Consider a maintenance engineer who receives a photo showing corrosion or damage on a piece of equipment.
A few years ago, applying AI might have meant sending the image to a vision model and asking what it could be. Useful, perhaps, but still fundamentally a question-and-answer interaction.
In a more agentic architecture, the image becomes the starting point for an investigation.
A supervisory agent can identify the relevant equipment, retrieve technical documentation and maintenance history, inspect recent process telemetry, call existing predictive or condition-monitoring models, look for similar historical events, and assemble the evidence around the issue. Within minutes, the engineer gets not just an opinion, but a structured assessment with supporting evidence and suggested next steps.
The engineer still makes the consequential decision. But much of the information gathering and analysis that precedes good engineering judgment has been compressed from potentially hours to minutes.
That is a very different value proposition.
We increasingly see the same pattern across documents, structured records and industrial time-series data. On bounded technical problems, the quality of the resulting investigation can already be comparable to what I would expect from a competent mid-level technical engineer.
There is another, even simpler change that I think matters enormously:
The human no longer has to initiate every interaction.
At Arundo, we use our own scheduled agents to monitor applications, data pipelines, clusters and other systems. They don't wait for somebody to remember to ask a question. They wake up, investigate, bring together evidence and produce a useful output.
The same principle applies in a plant.
There is a fundamental difference between an operator asking:
“How did the plant perform last night?”
and a system that has already reviewed the night's operation, identified an unusual deterioration, investigated likely causes and put a sourced recommendation in front of the morning team.
The prompt is no longer the trigger.
That, to me, is where the transition from copilot to operational participant starts.
Autonomy needs levels
The term “autonomous factory” creates another problem: it sounds binary. A plant is either autonomous or it isn't.
The automotive industry solved a similar vocabulary problem by defining levels of driving automation. I think industrial operations need something comparable.
We came up with the following scale at Arundo:
Where are we today? Most brownfield industrial operations are still predominantly at Levels 0 and 1.
Level 2 is now very achievable.
There are credible Level 3 systems and even examples beyond that, but they tend to operate in clearly bounded domains. Higher levels are also considerably easier in greenfield environments, where systems and operating practices can be designed for autonomy from the beginning.
That will change. But I don't expect whole factories to suddenly jump from Level 1 to Level 5.
What happens next
My prediction for the next 12-24 months is fairly simple.
Level 2 will become ordinary. Asking questions across plant data, documentation, maintenance records and analytical models will increasingly become a normal part of industrial software.
Level 3 will become the real battleground. The most interesting systems won't wait to be asked. They will monitor, investigate and recommend continuously.
Level 4 will arrive decision by decision, not factory by factory. Companies will first close the loop around narrow problems where the economics are compelling, the physics are understood and the consequences can be safely bounded.
A year ago, I used words such as “sentient” and compared the architecture of an autonomous factory to the human brain. It was a useful way of describing what might become possible.
Today I would use more operational language.
A factory doesn't need to be sentient.It needs to understand its operating context. It needs to detect what matters. It needs to reason about possible actions. It needs to learn from outcomes. And, progressively, it needs to be trusted to act.
That sounds slightly less futuristic. In practice, we are much closer to it than I thought we would be a year ago.
The autonomous factory probably won't arrive on the day somebody switches the operator off. It will arrive one operating decision at a time.
Sources
[2] Artificial Analysis: GDPval
[3] Stanford HAI: 2026 AI Index Report


