ML / AI · Jun 2026
AI for Construction Safety, Evidence Package
A construction-safety vision evaluation reaching F1 0.7591, packaged with paired statistical tests, decision-threshold sensitivity analysis, and end-to-end claim traceability from raw predictions to reported numbers. Sparse-category rates are now reported with Wilson, Agresti-Coull, and exact Clopper-Pearson confidence intervals.
0.7591
F1
Paired
Tests
Threshold sweep
Sensitivity
Problem
Safety-critical ML claims frequently rest on a single F1 figure with no record of how the threshold, split, or comparison was chosen. A rigorous evidence package was required to support construction-safety detection results.
Approach
I evaluated the system on a held-out split, ran paired statistical tests against baselines, swept decision thresholds to characterize precision-recall trade-offs, and wired every reported number back to the raw per-sample predictions through a reproducible pipeline.
Results
The package reports F1 0.7591 with full claim traceability, enabling reviewers to inspect, contest, and rerun any number in the report. The format generalizes to other safety-critical evaluation work.
Stack