Illustrative AI-generated vulnerability report
Remote code execution in the gateway
- Reported component
- Gateway runtime 4.2
- Reported exposure
- Internet-facing
- Evidence cited
- Affected version and a successful connection to port 8443
Enter the preview passphrase to read what is ready so far.
Industrialized AI puts a repeatable process around work done by AI models.
It uses those models to produce things people need, such as reports, documents, and source code; each result is a product of the process. At its center is a simple loop in which AI models manufacture the product, inspect it, repair anything found wrong, and inspect the replacement. The models perform the difficult semantic work throughout; a deterministic pipeline records what they work from, defines what the product needs to satisfy, and controls what happens to each result.
AI models can make mistakes when manufacturing, inspecting, or repairing a product, just as people can. We treat each step like anything else in engineering: define what it must do, record what happened, test how well it works, and improve it when it fails.
We keep the original information, every version of the product, the problems found, the changes made, and the final decision. What leaves the process is an accepted product together with the evidence which explains that acceptance, including a record that can be examined if the result is questioned later.
Software teams are being flooded with vulnerability reports from scanners, security services, and, increasingly, AI systems. Some describe real emergencies; others combine a few accurate details into conclusions which don't hold for the software as it's actually deployed. The development team has to tell the difference quickly, because time spent investigating the wrong report is time not spent fixing a more serious vulnerability or delivering work the business values.
Teams have begun using LLMs to help with that review. A reviewer gives the model a report and whatever evidence is available, then asks whether the finding is accurate. That can save time, but it leaves a harder problem: how do we know whether the answer was properly checked, and what can we show the person who's expected to act on it?
We're going to follow one of those reports through the complete process. The question is deliberately straightforward: does the available evidence support what the report says?
Illustrative AI-generated vulnerability report
The product in this example is a more accurate version of the vulnerability report. The incoming report becomes part of the material used to produce it, alongside the advisory, deployment configuration, and network evidence. The finished report should tell the team which claims are supported, which aren't, and what the available evidence can't establish. The team can then decide how to fix the vulnerability; producing and checking that software change would be a separate piece of work.
Someone can paste the report and the available evidence into a chat window and ask whether the finding is accurate. The answer may be useful, but they still have to decide what to trust, what to check, and when they've done enough.
Someone else may provide different evidence or ask the question differently. The person approving the report still carries the responsibility, often without an agreed standard or a record that explains why they accepted it.
That means moving those choices out of an improvised chat and into a repeatable process.
The method changes with the reviewer, while the human still carries the work and responsibility.
The pipeline controls the route and releases the accepted report together with the evidence behind it.
Source material and the acceptance basis feed manufacture. Manufacture produces a candidate revision for inspection. Inspection either sends an accepted revision to release or routes a defective revision through repair. Repair feeds the replacement revision back into inspection.
Produces a candidate revision from the recorded material.
Checks the revision against the acceptance basis.
Produces a replacement revision from the recorded defect.
Makes the accepted revision available with its evidence.
The incoming report and the evidence used to check it become the source material: the information every later step will work from.
| Incoming report | Critical internet-facing remote-code-execution claim for gateway runtime 4.2. |
|---|---|
| Vulnerability advisory | The vulnerability is rated Critical (9.8). The affected feature permits unauthenticated remote code execution when it can be reached. |
| Runtime and configuration | Gateway runtime 4.2 is installed and the vulnerable feature is enabled on port 8443. |
| Scanner observation | Port 8443 accepted a connection from the corporate network. |
| Network controls | No external ingress route is published and the firewall snapshot contains an internet deny rule. |
| Known gap | No independent connection test was made from outside the corporate network. |
Before we ask the model to review the report, we need to agree what the team needs to know. We turn that into the questions below and require the report to identify anything the evidence can't prove. Inspection will use those same questions to check the result.
| 01Applicability | Are the affected runtime version and vulnerable feature present? |
|---|---|
| 02Exploitability | Does the vulnerable path exist, and what preconditions apply? |
| 03Exposure | Can the path be reached internally, externally, or from both positions? |
| 04Severity and action | Do the established facts support the rating and recommended response? |
| 05Limits | Which missing evidence could materially change the report? |
The model now does the investigation which would otherwise land with an engineer and produces the report the team needs. Producing this first report is manufacture. We save the result as a product revision so the next step has an exact report to inspect.
| Affected component | Confirmed: gateway runtime 4.2 is installed. |
|---|---|
| Vulnerable feature | Confirmed: the feature is enabled on port 8443. |
| Impact | Unauthenticated remote code execution if the vulnerable feature can be reached. |
| Exposure | Internet-facing. |
| Severity | Critical (9.8), as rated by the advisory. |
| Recommended action | Interrupt planned work and deploy the patched version immediately. |
In an ad hoc LLM review, this is often where the work ends. Someone reads the answer, perhaps asks the model a couple more questions, and decides whether it looks good enough to use. Here, checking the answer is a separate step that the report must pass before it can go back to the team.
The first report looks convincing, but the team still needs to know whether its conclusion follows from the evidence. Inspection gives the model a narrower job: take the saved report, the source material, and the questions we agreed, then assess each question and record anything that doesn't hold up. The model provides that semantic judgment; the process applies the acceptance rules and decides what happens next.
The obvious question is what we gain by asking an LLM to inspect work produced by an LLM. The inspector can itself be tested against reports whose correct conclusions are already known, including versions with deliberate mistakes such as the unsupported exposure claim here. Comparing its assessments with those known results shows which errors it catches or misses and whether it rejects sound conclusions, giving the team a basis for improving inspection and judging when it can be relied on.1
The pipeline requires inspection to record an assessment for every acceptance question, then uses a decision table to interpret the results. In this case, the unsupported exposure claim is a blocking defect, so the report is rejected and sent to repair.
The loop can continue while another repair is permitted. Release requires a complete assessment with no recorded finding that blocks acceptance under the declared tolerance. If a required assessment cannot be completed or the permitted attempts are exhausted, the run stops without release. The evidence, report versions, and inspection results are retained throughout, so we can see what was checked and how the final outcome was reached.
| Question | Can the gateway be reached from the internet? |
|---|---|
| Report claims | The connection to port 8443 confirms internet exposure. |
| Evidence shows | The connection came from the corporate network. No external test was supplied. |
| Semantic assessment | Internet exposure has not been established. |
| Process decision | Repair required |
Most of the report still holds up, so repair can keep those parts and correct the unsupported claim without making the team start the review again.
Inspection gives us a specific problem to correct rather than sending the model back to the beginning. The failed report, source material, and recorded defect go to the LLM with a clear instruction: correct the defect while preserving the parts which passed. Producing a new revision from that failed work is repair.
reviewed-report-r001 Failed inspection The report treated internal reachability as proof of internet exposure.
The source set contains no successful test from outside the corporate network.
reviewed-report-r002 Repaired report The claim is corrected and the missing evidence remains visible.
Repair doesn't establish that the correction worked, so the new revision returns to the same inspection. The model makes a fresh semantic assessment, and the decision table accepts the revision only when the complete inspection contains no recorded blocking defect. Otherwise, the report returns for another permitted repair or the run stops without release.
The model completes the inspection without recording a defect which the declared tolerance treats as blocking. The
decision table therefore accepts reviewed-report-r002, and the process
records that this is the version which may be returned to the team.
The decision table has accepted the repaired report after complete inspection. Before the team uses it, release records the exact version which completed the process and the evidence supporting that decision.
| Affected component | Confirmed: gateway runtime 4.2 is installed. |
|---|---|
| Vulnerable feature | Confirmed: the feature is enabled on port 8443. |
| Impact | Unauthenticated remote code execution if the vulnerable feature can be reached. |
| Observed exposure | Confirmed from the corporate network. |
| Internet exposure | Not established by the available evidence. |
| Severity | Critical (9.8), as rated by the advisory. |
| Recommended action | Deploy the patched version. Verify whether the service is reachable from the internet. |
Release prevents a failed draft or later unchecked response from quietly becoming the version somebody acts on.
The product is the checked report, not a patch. It gives the team a better basis for deciding what deserves attention and how urgently, without forcing them to repeat the investigation.
Along with the checked report, the team keeps the evidence and work which produced it. If the report is challenged later, or a mistake gets through, they can trace the problem to the source material, manufacture, inspection, repair, or the acceptance standard.2 The team can use those cases to test changes to manufacture or repair, comparing the original and revised methods on the same inputs against findings established from the source evidence. That shows whether a change produces a better first report or corrects a defect while preserving the conclusions that were already sound.
Looking across comparable runs helps the team distinguish the process’s usual variation from a change that needs investigation.3 A method which repeatedly confuses staging with production may need better source material or instructions, while a sudden increase in those errors after a model update gives the team a specific change to investigate. Consistent performance still has to meet the requirements; a process can be predictable and still produce too many defects.