How good is the model?
Tested on tickets filed between
The model against simple baselines
Two baselines make the comparison honest: "the project's usual late rate" and "how late similar recent tickets were". PR-AUC is the headline measure. The 95% intervals from resampling the test tickets are about ±0.01, narrower than the dots.
Are the probabilities honest?
If the model says 60%, about 60% of those tickets should run late. Each point is a tenth of the test tickets; points on the diagonal are perfectly calibrated.
Every project beats its own late rate
How risk moves with five signals
Average effect on a ticket's risk, in percentage points, at filing. Positive means more likely to run late.
Check the model against what happened
Pick a sample ticket from the test period and see what the model said at the time, why, and what happened.
Limits
- One time split, one test window. This is not a rolling backtest.
- The late rate drifts up over time (
in training, in the test period at filing) because recent tickets stay open. Calibration, learned on the year before the test, held up. - Results vary by about ±0.002 PR-AUC between identical runs.
- The model finds associations in past data, not causes. Do not read "assign tickets when filing them" as a recommendation.