Technology / Dataset

Building a Provenance-First Computer Vision Corpus

Film becomes training data only through a rigorous, auditable process.

SECTION

Provenance by Design

Every bounding box, track, label, and homography stores:

  • Source master file ID
  • Exact frame number
  • Model/version that produced it
  • Confidence score
  • Schema version

Future models can be run retroactively over the entire historical corpus because raw evidence (tracks, homography, audio features) is stored at ingest.

SECTION

Rights Policy

Only first-party film (your own team's Hudl exports, self-filmed 4K sideline) or properly licensed commercial film is ingested. No NFHS streams, no scraped content. The pipeline is deliberately source-agnostic but the business rule is strict.

SECTION

Scale So Far (measured)

  • Helmet detection agreement: 97.9% over 240 frames
  • ~9,900 frames labeled with 193,736 boxes

The dataset grows with every game ingested, whether or not the film receives restoration or clipping. This ensures that when advanced route/coverage models arrive, they have the full historical base to train on.