Near-Miss × Crash · Melbourne metro · GoldNet / CrashStats R&D
Surrogate-safety analysis

Do telematics near-misses tell us anything crashes don't?

DTP looked at Compass near-miss data at an aggregate level and concluded it just re-tells traffic volume — more volume, more crashes, more near-misses, no new insight over the crash-and-volume ranking already in use. That conclusion was never tested on the raw event data. This is that test.

1 · The aggregate paradox, explained

The "near-misses just track crashes" finding is an artefact of how coarsely you aggregate. As grid cells grow, everything correlates with everything (both track volume). Shrink to the resolution at which you'd actually target investment and the two signals come apart.

Variance in crash count explained by near-miss density (log–log R²), by grid cell size. Higher = more they look alike.
This is the quantitative version of the team's observation. At the whole-of-network / aggregate view DTP used, the correlation looks meaningful. At 500 m — closer to a treatable site — near-miss density explains just 4% of crash variance. And once you divide out real traffic exposure (AADT), the direct near-miss–crash link collapses to a partial correlation of just 0.11: near-miss density is essentially an echo of how many vehicles pass. The value has to come from somewhere else — which is the next panel.

2 · Where the two signals disagree

On a 1 km grid, most cells behave as expected (both high, or both low). The interesting ones are the disagreements: near-miss activity without a crash record yet (candidate latent-risk sites — the "locations to investigate" on raw data), and crashes with little behavioural signal (crash-only — often where telematics simply isn't looking).

Top latent-risk candidates high near-miss · low/no crash

Ranked by exposure-shrunk divergence. Presented as candidates for a raw-data look, not confirmed hazards — see the caveats before acting on any single cell.

3 · Density re-tells volume — but intensity is different

If near-miss is only ever a proxy for how many cars pass, it adds nothing. The question is whether the character of the near-misses — how hard, how fast, what kind — carries information that raw counts don't. Ratios like g-force and over-limit share largely cancel out exposure, so they're the fair test.

Does near-miss intensity predict crash severity?

Spearman correlation, per cell, of each near-miss trait against the share of that cell's crashes that were fatal or serious (FSI).

What kind of near-miss?

Event mix across the sampled raw events — a diagnostic layer crash counts can't provide.

4 · Would it help the model? (for AP3)

The decisive test: give the model a real traffic-exposure offset (AADT × length, from the Victorian AADT layer) so it already knows how many vehicles use each road — then ask whether near-miss adds anything beyond that. Poisson GLM, spatially-blocked 5-fold cross-validation, deviance explained on held-out folds.

Crash count — deviance explained (over exposure)

Severe (FSI) count — deviance explained

This is the hardened version of the test: exposure is a real AADT offset, not a road-type proxy. Near-miss survives it. The honest read — a supporting predictor, not a replacement for crash history — but one that earns its place precisely where crash-and-volume ranking is weakest: severe outcomes.

5 · Coverage & what this can't say

These results speak for metropolitan Melbourne only, and every claim is bounded by how telematics data is generated.

Hard limits

  • Exposure is now controlled — but only for total traffic. We divide out real AADT (vehicle-km) so results aren't just "busy road = more of everything." What AADT can't remove is telematics fleet penetration: near-miss counts still reflect where instrumented (fleet, rideshare, insurance) vehicles drive. That's why the headline rests on intensity and lift-beyond-exposure, not raw density.
  • Regional coverage is poor. Even in metro, 38% of crashes fall in cells with zero near-miss data. Outside metro it is far sparser — a network-wide near-miss layer is not currently viable.
  • Fleet & device bias. A laden truck and a hatchback trip different g-force thresholds; vendor detection settings vary. Treat mix and intensity as directional.
  • Divergence cells are candidates, not verdicts. A "high near-miss / low crash" cell can be a genuine latent hazard or a depot turnaround on a fleet's habitual route. They earn a look at the raw traces, nothing more.
  • No temporal lead test yet. The strongest version — near-misses today predicting crash increases next year — needs near-miss and crash windows that overlap in time; this run could not close that loop.