How we test it

Measure the decision, not just the outage.

Anyone can look smart with hindsight. We test whether a repair order was actually good using only what an operator could have known at the time.

Working beta · Testing against simulated storms · Operators always make the call

What we judge

Numbers that hold up when you push on them.

Every recommendation gets tied back to a measurable outcome, the limits it was working under, and the exact evidence sitting on the table when the call was made.

The team building the evaluation method
01Did it help?The restoration curve, customer-minutes out, and how fast critical loads came back.
02Could it be done?Topology, crews, materials, road access, safety holds, and critical loads.
03Was the call explainable?The alternatives, what was uncertain, where it declined to guess, and what the operator decided.
04What did we learn?Replaying at the original moment ties each recommendation to what actually happened, without hindsight.

The rules

A fair test gives both sides the same information.

Everything being compared works from the same snapshot of the event. Anything that had the benefit of hindsight is reported on its own and never sold as a result.

Test method last updated 7 August 2026. Read the versioned protocol

V.01

What we are asking

Did this repair order restore more customers, faster, without breaking anything that could not actually be done?

V.02

What we measure

Estimated customer-minutes out (CMI): how many customers were off, multiplied by how long, added up across the event. Customer-minutes out is the raw quantity behind SAIDI, SAIFI, and CAIDI—we report it directly rather than as an index, because index definitions and exclusion rules vary by regulator.

V.03

Measured vs. estimated

Customer-minutes taken from your approved historical records are kept separate from our modeled estimate. We never mix the two.

V.04

What makes it fair

Dovrane and the baseline get exactly the same information that existed at the original moment of the decision.

V.05

No peeking

A later report stays hidden until the timestamp it actually arrived.

V.06

Who we compare against

Your approved baseline, Dovrane, and an analysis-only run that is allowed to see the future—reported separately and never as a product result.

V.07

Test sets

Development, regression, and sealed sets, so we can show progress without quietly tuning against the final exam.

V.08

What we publish

The full spread, the median, the worst cases, every constraint violated, and how often the model declined to answer.

The reality checks

Every plan has to clear all six.

Separate checks screen each repair order against the things that actually stop work in the field. Fail one and the plan is held back, not ranked. Each check runs on every test set, and we report how often we get it wrong.

01

Topology and isolation

Can this section actually be isolated and re-energized in this order, given how the feeder is switched?

02

Crew class and hours

Is a crew with the right qualification actually free, and does the work fit inside their remaining hours?

03

Materials on hand

Is the pole, crossarm, or transformer this step needs sitting in a yard you can reach?

04

Roads and site access

Is the site reachable right now, or is the road closed, flooded, or blocked by debris?

05

Safety holds

Is there an active hold at this location—downed conductor, gas leak, fire department on scene?

06

Critical loads

Does this order respect the hospitals, water, and other loads you have designated critical?

What we learn

From plan quality to what actually worked.

  • Was the order better?Restoration impact measured against your approved baseline.
  • Did it stay honest?How many constraints it broke, how many plans it held back, how often it declined.
  • Could you follow it?Whether the evidence, alternatives, and reasoning were actually traceable.
  • Did operators use it?How often they accepted, changed, held, or rejected the recommendation.
  • How did it handle the worst?Critical-load recovery and behaviour on the ugliest events, not just the average.
  • What changes next?Observed outcomes feed the next round of evaluation and review.

Next step

Push on the method.

Bring the questions you would ask in a prudency review. We would rather answer those than show you a number.