Evaluation protocol
Restoration Sequencing Evaluation Protocol
ABSTRACT
This document defines how Dovrane Inc. evaluates repair-sequence recommendations for electric utility storm restoration. Every comparator is replayed against the same event-time snapshot, every candidate plan is screened by independent feasibility checks, and results are reported as full distributions with violations and abstentions counted. The protocol is versioned so that any claim produced under it can be pinned to the exact rules in force at the time.
Scope
This protocol defines how Dovrane Inc. evaluates repair-sequence recommendations for electric utility storm restoration: whether a recommended repair order restores more customers, faster, without violating a constraint that would have stopped the work in the field.
Evaluation currently runs against simulated storm events. Dovrane has no utility deployments and publishes no performance results; this document describes the method by which any future result would be produced, not a result.
The protocol covers modeled evaluation only. It does not cover real-world safety determination, crew dispatch, switching, or grid control, which remain entirely with the utility.
Definitions
Repair order (also repair sequence): a complete, ordered plan of storm repairs built from one snapshot of the event.
Estimated customer-minutes out (CMI): the number of customers off supply multiplied by how long they were off, summed across the event, as computed by the model. Customer-minutes out is the raw quantity behind SAIDI, SAIFI, and CAIDI; this protocol reports it directly rather than as an index, because index definitions and exclusion rules vary by regulator.
Observed customer-minutes: the same quantity derived from a utility's approved historical outage and restoration records. Estimated and observed customer-minutes are never mixed or interchanged.
Restoration curve: customers interrupted plotted against event time under a given repair order.
Event-time snapshot: the complete set of information that existed at one moment of the event — damage reports, crew and material state, road access, topology, and safety holds — with everything that arrived later excluded.
Abstention: an explicit model output of 'not enough evidence to answer', reported and counted rather than replaced by a guess.
Point-in-time replay rules
Every comparator operates on the same event-time snapshot. Dovrane and the baseline receive exactly the same information that existed at the original moment of the decision.
A report that arrived later stays hidden from every comparator until the timestamp at which it actually arrived.
Any run that is allowed to see information from after the decision moment is labeled as hindsight, reported separately, and never presented as a product result.
Feasibility checks
Each candidate repair order is screened by independent checks against the conditions that stop work in the field. A plan that fails any check is held back from ranking rather than ranked lower.
4.1 Topology and isolation: can this section actually be isolated and re-energized in this order, given how the feeder is switched?
4.2 Crew class and hours: is a crew with the right qualification actually free, and does the work fit inside their remaining hours?
4.3 Materials on hand: is the pole, crossarm, or transformer this step needs sitting in a yard that can be reached?
4.4 Roads and site access: is the site reachable now, or is the road closed, flooded, or blocked by debris?
4.5 Safety holds: is there an active hold at this location — downed conductor, gas leak, fire department on scene?
4.6 Critical loads: does this order respect the hospitals, water, and other loads the utility has designated critical?
Every check runs on every test set, and the rate at which each check is wrong is reported.
Baselines and comparators
Three comparators are evaluated on each event: the utility-approved baseline, the Dovrane engine, and an analysis-only oracle that is allowed to see the future.
The oracle exists to bound what was achievable in hindsight. Its results are reported separately under the rules of §3 and never as a product result.
Comparison against the approved baseline on the same event, under the same snapshot, is the only basis on which a repair order is called better or worse.
Test-set discipline
Events are divided into development, regression, and sealed sets.
Development sets are used to build and tune. Regression sets detect degradation between versions. Sealed sets are evaluated only at declared milestones, so progress can be shown without quietly tuning against the final exam.
Reporting requirements
Any published evaluation reports the full distribution of outcomes, the median, and the worst cases — not a best case or an average alone.
It reports every feasibility-check violation and the rate at which the model abstained.
It reports estimated and observed customer-minutes separately, per §2, and states which test set (development, regression, or sealed) produced each number.
It states the protocol version it was produced under.
Status and revisions
Version 0.1, published 7 August 2026. This protocol will be revised as the evaluation program matures; each revision changes the version number and this page's URL, so a citation of v0.1 always resolves to the rules it was made under. Questions about the method: contact@dovrane.tech.