How good is the model?

Tested on tickets filed between and , a period the model never saw while it was trained or tuned. Two models: one scores a ticket when it is filed, one scores it a week later.

The model against simple baselines

Two baselines make the comparison honest: "the project's usual late rate" and "how late similar recent tickets were". PR-AUC is the headline measure. The 95% intervals from resampling the test tickets are about ±0.01, narrower than the dots.

When filed, boosting is clearly better than both baselines and than logistic regression. On day 7 both models beat both baselines by a wide margin but are tied with each other, so a simple logistic regression would do as well. Day 7 scores look high for every model because tickets still open after a week are late 72% of the time.

Are the probabilities honest?

If the model says 60%, about 60% of those tickets should run late. Each point is a tenth of the test tickets; points on the diagonal are perfectly calibrated.

The middle is well calibrated. At the extremes the calibration flattens, which is why probabilities are shown limited to 5%–95%.

Every project beats its own late rate

How risk moves with five signals

Average effect on a ticket's risk, in percentage points, at filing. Positive means more likely to run late.

Orange raises risk, green lowers it. Priority is not monotonic: both the top and bottom priorities lower risk. "Reporter's earlier tickets" counts how many tickets the reporter had filed before, never who they are.

Check the model against what happened

Pick a sample ticket from the test period and see what the model said at the time, why, and what happened.

Limits