DTP looked at Compass near-miss data at an aggregate level and concluded it just re-tells traffic volume — more volume, more crashes, more near-misses, no new insight over the crash-and-volume ranking already in use. That conclusion was never tested on the raw event data. This is that test.
The "near-misses just track crashes" finding is an artefact of how coarsely you aggregate. As grid cells grow, everything correlates with everything (both track volume). Shrink to the resolution at which you'd actually target investment and the two signals come apart.
On a 1 km grid, most cells behave as expected (both high, or both low). The interesting ones are the disagreements: near-miss activity without a crash record yet (candidate latent-risk sites — the "locations to investigate" on raw data), and crashes with little behavioural signal (crash-only — often where telematics simply isn't looking).
Ranked by exposure-shrunk divergence. Presented as candidates for a raw-data look, not confirmed hazards — see the caveats before acting on any single cell.
If near-miss is only ever a proxy for how many cars pass, it adds nothing. The question is whether the character of the near-misses — how hard, how fast, what kind — carries information that raw counts don't. Ratios like g-force and over-limit share largely cancel out exposure, so they're the fair test.
Spearman correlation, per cell, of each near-miss trait against the share of that cell's crashes that were fatal or serious (FSI).
Event mix across the sampled raw events — a diagnostic layer crash counts can't provide.
The decisive test: give the model a real traffic-exposure offset (AADT × length, from the Victorian AADT layer) so it already knows how many vehicles use each road — then ask whether near-miss adds anything beyond that. Poisson GLM, spatially-blocked 5-fold cross-validation, deviance explained on held-out folds.
These results speak for metropolitan Melbourne only, and every claim is bounded by how telematics data is generated.