How this was built

Every number on this site comes from a public Jira snapshot, a fixed definition of "late", and a model tested on tickets it had not seen.

The data

What "late" means

Two moments of prediction

The model scores a ticket when it is filed and again on day 7 if it is still open. Each score uses only what was known at that moment. Fields that change (priority, type, components, links, assignee) are rebuilt from the change history as they stood then, never the final value: in CASSANDRA, 79% of tickets have their priority changed after filing, so using the final value would have leaked the future. Anything known only after the moment, such as the resolution, is never used.

The model

Gradient boosting (scikit-learn), compared with a logistic regression and two baselines, trained on tickets filed before ( for the day-7 model) and tested on filed from to . The split is by time, never random. Hyperparameters were chosen by validating on the next year in a rolling fashion, and probabilities were calibrated on the last validation year. Nothing was tuned on the test period.

The reasons

Each ticket's reasons are exact SHAP values from the boosting model, grouped into four themes: Ownership, Scope, Workload and History. A check requires the values to add up to the model's own output. Rankings stay stable when the model is refit on resampled or earlier data (rank correlation 0.94 to 0.99).

What this cannot tell you

The code

The project is a Python and SQL pipeline (DuckDB, scikit-learn, SHAP) that extracts the dataset, builds the label and features, trains the models and exports the small files this site reads. The site itself is static and computes nothing in your browser beyond drawing charts.