On every chart in the control room, last week’s runs look fine. One of them was not. The evidence exists right now, spread across forty sensor channels, each individually in spec. You will meet this excursion in three weeks, at end-of-line characterization, after it has repeated itself through every run in between.
This series is about materials enablement, the application of AI to the journey from “works once” to “works every time.” Part 1 built the foundation, a live estimate of material state. Part 2 turned history into understanding, a model of how the process actually behaves. Detection is where those investments start paying operational rent: knowing the moment anything changes, in a process where change is constant.
That last clause is the paradox of scale-up. The process is supposed to be changing. Setpoints move, recipes get revised, and deliberate change from run to run is the entire point of development. That is precisely what makes unintended change so hard to see. When change is the baseline, drift wears camouflage.
Today’s Ceiling: Excursions Found Weeks Late
Today, most excursions surface at end-of-line characterization, weeks and many wafers after they began. The direct cost scales with time-to-detect: every additional day is more scrapped material. The indirect cost is worse. Every experiment run during the undetected window was interpreted against a process that was no longer the process, so the excursion consumes wafers and corrupts conclusions in the same stroke.
The standard toolkit was built for a different world. Statistical process control assumes a settled process: stable mean, known variance, control limits earned over hundreds of runs. A development process violates every one of those assumptions by design. Teams that apply SPC anyway get a flood of false alarms, respond by widening the limits, and end up with charts that can no longer see anything at all.
Even a settled process hides two kinds of change from single-signal charts. One is the joint shift: forty channels each within limits while the relationship between them walks away, a change no univariate view can represent. The other is slow drift, from consumable aging, chamber conditioning, even seasonal effects in the facility, each too gradual to cross any threshold on any one chart.
The modern fix, generic anomaly detection, fails socially rather than statistically. A physics-blind model scores strangeness without meaning. Engineers learn within a month which alerts to ignore, and an ignored alarm system is worse than none, because it converts vigilance into ritual.
The Solution: Changepoints with Attribution
Atomscale treats detection as a modeling problem before an alerting problem. Changepoint detection runs over the full multivariate stream, referenced against a physics-grounded model of what this particular run should look like, the same model hierarchy built up in Parts 1 and 2. Because the reference knows the recipe, intended change and unintended drift separate cleanly: moving a setpoint is expected, while the chamber responding differently to the same setpoint is an event.
That grounding is also the answer to alarm fatigue. An alert from Atomscale arrives with attribution: what changed, where in the run, and against which expectation. An engineer can evaluate that in minutes, and the alert channel keeps its credibility, which is the property alarm systems usually lose first.
The Unlock: Excursions Caught While the Run Is Live
Time-to-detect collapses from weeks to the middle of the run. The excursion in the opening of this piece gets flagged during the run where it began, with a pointer to what moved. Post-mortem mysteries become attributable events, and “was that run normal?” becomes a question with an automatic, trustworthy answer, for every run, without anyone staring at charts.
At one materials program we work with, changepoint detection over the existing sensor archive surfaced process shifts the team’s monitoring had never registered. The runs were already grown and the data already stored; what had been missing was a model of normal against which change becomes visible.
The next possibility is prediction. A model that knows what every healthy run looks like can begin flagging the run that is going to fail before it does, and maintenance can trigger on material state instead of wafer counts. Detection matures from catching change into anticipating it.