AI's next stage is industrialization.
As with other new technologies, people have started by using Generative AI in familiar workflows. A coding assistant, for example, can pick up a ticket, implement the change, and submit code for review. This helped people get started quickly and learn what the models could do, but as the models have become more capable, the limits of that approach have become clearer. LLMs can produce more work than people can check and correct, but the quality still varies, so teams must either limit how much they use or accept the risk of using work that hasn't been adequately checked.
Previous technologies made their largest gains when the work was redesigned around what the new machinery could do, rather than bolting that machinery onto the old workflow.1 We're now at that same point with AI models, and if we redesign the work around them, the engineering task becomes familiar: arranging variable, imperfect machinery into processes that produce work to defined criteria and tolerances at scale. This is a class of problem that manufacturing engineers have spent generations solving through process control.
Industrialized AI is the application of manufacturing principles to cognitive work.
Thinking of LLMs as people can only take us so far.
Chat was the first popular way of interacting with Generative AI, and from there, the assistant became an agent, and then the agent was given a role, skills, and tools, until whole frameworks were built around chat interfaces.
Chat gave people a familiar way to use the models and organizations a manageable way to experiment with them. The problem comes when that interface starts to dictate how the work is organized, with agents given roles and coordinating through conversations. We end up automating the way people coordinate work rather than redesigning the work around the new machinery.
Keeping the human-shaped workflow has a practical ceiling because model output can grow very quickly, while trusted output can't grow beyond our ability to review it. Checking and correcting the work can consume the savings from generating it, while skipping those checks can turn unnoticed defects into incorrect decisions or failures in production. Moving past that ceiling requires us to change the role the model plays inside the system.
We can treat the model as machinery, not management.
At the simplest level, an LLM takes something in and makes something from it. You can give it requirements and it'll produce code, or give it source documents and it'll produce a report. Models can sometimes get this right in one shot, and they're getting better at doing that, but they can still make mistakes: both obvious and subtle. That makes it an incredibly useful but imperfect machine.
Fortunately, engineering has never required perfect machinery. Engineers decide what acceptable work means for a particular job, how much variation can be tolerated, how the result will be checked, and what happens when it falls short. An AI model is just another machine with variable output, so the same principles apply.
An LLM can also reason, plan, and use tools, which makes it tempting to let it manage the work as well as perform it. In many AI workflows, the agent chooses what to do next, which checks to run, and when to stop. That makes the process itself another source of variation, so even similar jobs can involve different amounts of work and different checks before a result is accepted. Quality, cost, and delivery time become harder to predict, and it becomes harder to establish why something failed or whether a change actually improved the process.2
This lets us reframe LLM workflows as manufacturing processes.
A manufacturing process moves material through a sequence of operations until it becomes a product that meets defined requirements. Machines and people perform those operations, with the result of one becoming material for the next.
AI work has the same shape. Requirements, source data, and existing work become material; models, tools, and deterministic software perform operations; one result can become the input to another. A report, source code, or design is a product of that work, but producing it doesn’t tell us whether it’s good enough to use.
Manufacturing deals with this through process control, which Shewhart, Deming, and generations of engineers developed to make variable machinery produce within defined tolerances. It gives us ways to specify acceptable work, inspect what was actually made, correct defects, and learn from the record of each run.
We can apply this discipline to AI by arranging manufacture, inspection, and repair as stages in a deterministic pipeline. LLMs, other AI models, and deterministic tools can produce the work, check it against defined requirements, and correct defects, allowing the process to run without a person reviewing every result. The pipeline records what each stage receives and produces, and controls how the work moves between them.
Engineers can test and compare the methods used at each stage, measuring how often manufacture produces acceptable first attempts, how accurately inspection identifies defects, and whether repair corrects those defects without introducing new ones.3
An AI inspector can miss a defect, just as a human inspector can, so inspection itself needs to be tested and monitored. Engineers test inspection methods on products known to meet the requirements and products with known defects, measure how often they get those judgments wrong, and monitor whether that performance changes over time. The results guide improvements to both manufacture and inspection, so fewer defects are produced and more of those that remain are caught. This gives us a basis for deciding whether the process is reliable enough for its intended use and whether it remains so.
Process control lets us manufacture to a defined standard at scale.
Once inspection and repair are part of the production process, we can apply the same approach to the next product without handing it to a person for a fresh review. That matters because LLMs can already produce far more work than people can check manually. Each product is checked against its requirements. If inspection finds that it still falls outside the agreed tolerances after the permitted repairs, the product is rejected or the process stops.4 The people responsible for the process improve it between runs instead of becoming the inspection stage that sets the production rate.
That changes the size and kind of work a team can reasonably take on. A substantial software change can be built from smaller products, each inspected before it is used to build the finished product, which is then inspected as a whole. Starting from the same requirements, a team could ask for a quick prototype or a version ready for release, setting the scope, budget, deadline, and standard for each rather than leaving those choices to emerge during the work.
Once work has been accepted and the evidence behind that decision has been retained, it can become input to the next controlled process without starting the assessment again from nothing. Accepted specifications can feed implementation, reports can feed decisions, and interpretations can feed further analysis. That gives us a supply chain of accepted work: later processes can build on earlier products without losing the source material, inspections, and repairs behind them.
This is Industrialized AI.
Industrialized AI organizes familiar mechanisms as a manufacturing system. AI models already generate and assess work, pipeline engines route it, and engineering systems run deterministic checks and retain versions and results. Many existing AI workflows already combine those mechanisms. A manufacturing system organizes their work around a defined product and tolerances, controls how it moves through manufacture, inspection, and repair, and releases it with the evidence behind the decision. The result is work produced to a declared standard through a process that can be repeated, examined, and improved.
A manufacturing system gives an organization more control over the cost, quality, and delivery time of AI work. By measuring how the process behaves, engineers can identify sources of variation and test ways to reduce them, while the organization decides how much to invest in achieving the outcomes it needs and what risk it is prepared to accept.
That control can make previously uneconomic products and services viable and allow existing ones to be delivered at greater scale.5
That is the next stage: the value an organization can realize from AI-produced work can scale with its investment in AI, without being capped by human review of every result.
Notes and references
- Factory electrification provides an example. Giving individual machines their own electric motors allowed factories to change their layout and the movement of materials, producing benefits beyond replacing the original power source. Paul David examines this in Computer and Dynamo (1989), section 5 (PDF). Back to text
- Manufacturing uses standardized work to reduce variation in how work is performed and establish a baseline for improvement. The Lean Enterprise Institute explains that connection, while NIST describes how stable patterns of variation make process outcomes more predictable. Applying this to AI means defining the workflow’s rules and measuring the variation that remains; a predefined workflow alone does not establish reliable output. Back to text
- Plan–Do–Study–Act offers a structured way to use these measurements: define an intended improvement, test a change, study the results, and use what was learned to revise the method. The Deming Institute describes the cycle and its origins in Shewhart’s work. Back to text
- A related manufacturing principle is jidoka: detecting abnormalities and stopping work so that problems can be addressed without someone continuously watching each machine. Toyota describes its origins in Sakichi Toyoda’s automatic looms and its role in the Toyota Production System. See Toyota’s explanation of jidoka. Back to text
- Deming’s chain reaction connects improvements in quality with lower costs through fewer mistakes, less rework, and fewer delays, improving productivity and the ability to compete. Applied to AI, this means considering the cost of producing accepted work, including inspection and repair. Back to text
See how it works.
The following pages start with the core loop, then show how it handles larger products and turns cost, speed, and quality into a production order. From there we can look inside the pipeline before running a complete example.