Behavioral attention classifier
A small, explainable model that could learn low-attention approval patterns from reviews collected with a counterbalanced protocol. It ships untrained: there is no public dataset for this task, and OverSight does not invent one. Until a model passes a grouped evaluation on real participants, the deterministic engine decides alone. Even then the classifier is advisory: it can raise intervention sensitivity, never lower it.
1 · Collect reviews
A collection session walks one participant through 18 requests in counterbalanced blocks (4 attentive, 4 rapid approval, 4 low attention, 3 distracted, 3 camera uncertain). The block order comes from a Latin square keyed by the participant code and scenarios are shuffled within the session. The console shows the instruction for each request; interventions are recorded but not enforced.
Notice and consent. Only derived numbers are stored, in this browser: no video, images, face landmarks, names, emails or clock times (time is counted from the start of the session). Participants are identified by a code you choose, such as P03. Tell each participant what is recorded and get their consent before starting; clear the data when the study ends.
- Attentive
- 0
- Rapid approval
- 0
- Low attention
- 0
- Distracted
- 0
- Camera uncertain
- 0
2 · Evaluate and train (grouped)
Runs the grouped evaluation of docs/EVALUATION.md (leave one participant out), then fits logistic regression on all reviews with calibration and thresholds from out-of-fold predictions. The model is activated only if it passes the pre-registered rule: at least 5 participants, a false-intervention rate within 1 point of the rules, and at least 20% fewer missed dangerous approvals with a participant-bootstrap interval above zero. The same code runs from the command line: npm run evaluate -- dataset.json
Stored features (derived numbers only, schema v2)
- conclusiveCoverage
- Dwell vs. required dwell on critical targets gaze could judge (missing when it judged none)
- coverageAvailable
- 1 when gaze judged at least one critical target
- trustLevel
- Gaze trust: 0 none, 1 low, 2 medium, 3 high
- trustConfidence
- Gaze trust confidence (0-1)
- separationSigma
- Best separation of a critical target from title/summary/buttons, in gaze-error units (capped at 10)
- timeToFirstCriticalRatio
- First fixation on a critical target / latency (1 = never; missing without gaze)
- attributedFixations
- log(1 + fixations attributed to critical targets); missing without gaze
- logLatencyRatio
- log(approval latency / expected review time from the personal baseline)
- visibilityRatio
- Time the critical regions were on screen vs. what reviewing them needs
- offCardRatio
- Share of gaze time outside the approval card; missing without gaze
- pointerActivity
- log(1 + pointer distance / 100 px)
- hoveredTarget
- Pointer rested on a critical region
- scrollDepth
- Deepest scroll position reached in the request
- riskOrdinal
- Request risk: 0 low, 1 medium, 2 high, 3 critical
- latencyMedian5
- Median log latency ratio over the last 5 approvals
- latencyEwma
- EWMA of log latency ratio across the session
- latencySlope5
- Slope of log latency ratio over the last 5 approvals (negative = speeding up)
- rapidStreak
- Consecutive rapid approvals ending here (capped at 5)
- coverageMean5
- Mean conclusive coverage over the last 5 approvals (missing when gaze judged none)
- notObserved5
- Critical targets with 'not observed' gaze evidence over the last 5 approvals
- speedupCusum
- One-sided CUSUM of -log latency ratio (sustained speed-up)
- approvalIndex
- Approvals so far in this session