Field Note · essay

August 2026 · 6 min read

Research integrity · model calibration

The result I refused to call reliable

A model can be technically impressive and still fail the one test that matters.

The most honest result in my grid-risk work was not a forecast. It was a decision: the model had not earned the right to be called reliable.

01

A gate, not a decoration

I built GERT, the Grid Extreme Risk Toolkit, to make tail risk visible. Instead of giving an operator only a median forecast, the system exposes a range of possible demand outcomes, including the P99 boundary where a rare event can collide with generation capacity.

That design creates a temptation. A fan chart looks rigorous. A dashboard can make uncertainty feel controlled. But the visual is only useful if the intervals mean what they claim to mean. So the project needed a calibration gate: a model would not be promoted as reliable simply because its interface looked finished or its average performance looked strong.

02

When the average hides the event

The PJM polar-vortex study asked a narrower question: can ordinary annual validation conceal failure during the cold events that matter most? In the current research record, the nominal 90% interval covered roughly 82–85% of observations across the annual evaluation, but only about 31–33% during the 2014 cold-event window.

Those figures are still part of an ongoing manuscript, not a published conclusion. But they were enough to change the product decision. The model had not passed the preregistered calibration threshold, so I did not present it as operationally reliable.

A failed promotion is not a failed project when the gate was designed to protect the decision.
03

What changed in the product

The failed gate became an interface principle. GERT distinguishes live, simulated, stale, and unavailable evidence. It shows provenance alongside a forecast and treats capacity margin as a decision variable rather than a decorative metric. Scenario tools are labeled as interventions, not observations.

This does not make the model certain. It makes the distance between the model and the world harder to ignore.

04

A result that refuses to disappear

Competition problems are designed so that an answer exists. Real systems make no such promise. In a real system, the most useful outcome can be evidence that the current method should not cross the line from experiment to claim.

I still want the model to work. I also want the threshold to mean something. If a test cannot stop a release, it is not a gate; it is decoration.

Evidence trail

Sources · status · limitations

The claim should lead somewhere.

The interval-coverage figures and promotion decision come from the current project audit and research records. They are presented as an ongoing result, not as a published paper or operational grid forecast.