Technology / Dataset
Building a Provenance-First Computer Vision Corpus
Film becomes training data only through a rigorous, auditable process.
SECTION
Provenance by Design
Every bounding box, track, label, and homography stores:
- Source master file ID
- Exact frame number
- Model/version that produced it
- Confidence score
- Schema version
Future models can be run retroactively over the entire historical corpus because raw evidence (tracks, homography, audio features) is stored at ingest.
SECTION
Rights Policy
Only first-party film (your own team's Hudl exports, self-filmed 4K sideline) or properly licensed commercial film is ingested. No NFHS streams, no scraped content. The pipeline is deliberately source-agnostic but the business rule is strict.
SECTION
Scale So Far (measured)
- Helmet detection agreement: 97.9% over 240 frames
- ~9,900 frames labeled with 193,736 boxes
The dataset grows with every game ingested, whether or not the film receives restoration or clipping. This ensures that when advanced route/coverage models arrive, they have the full historical base to train on.