How it works
Checking attention where it matters, and nowhere else.
Human-in-the-loop controls assume the human is reading. Under a steady stream of routine approvals, sign-off becomes a reflex: the one request that deletes customer data looks like the nine before it. A click proves a decision was made, not that the consequence was seen. OverSight adds the missing evidence.
01AI layer
Semantic risk engine
- Every approval request is split into typed blocks with stable ids (title, reasoning, resources, changes, consequences).
- A deterministic rule engine grades each block (irreversible deletion, public exposure, unverified payees, sensitive data leaving the organization, outages, permission escalation...) and selects at most two attention targets.
- If an OpenAI-compatible model is configured, its analysis is merged on top. It can escalate, never de-escalate below the rules: a deterministic safety floor.
02CV layer
On-device computer vision
- MediaPipe Face Landmarker (478 landmarks with iris refinement) runs in the browser via WebAssembly. Frames are reduced to a handful of numbers: iris position inside each eye, eye-direction coefficients, head yaw/pitch, face position.
- Calibration fits ridge-regression models from those features to screen coordinates: 9 dots, then a few seconds of looking at a dot while turning the head, so ordinary head movement does not throw gaze off. Accuracy is then measured on 5 dots that were never used for fitting and reported as Good, Fair or Recalibration recommended.
- Gaze is smoothed with a One Euro filter and carries its measured uncertainty. Evidence for each critical sentence is weighted by how compatible every estimate is with looking at its live bounding box. When gaze at that accuracy cannot tell the sentence apart from the title or the buttons, OverSight says so and relies on interaction timing instead.
03Behavioral layer
Behavioral signals & ML
- Approval latency is compared with the reviewer's own baseline from their first attentive reviews, not with a fixed wait time.
- Visibility, scroll depth, fixations and sweep across the critical sentence add evidence; a rapid-approval streak and a declining attention trend raise intervention sensitivity.
- A logistic-regression classifier can be trained on labeled reviews in the Model lab. It ships untrained and, once validated, is advisory only.
04Decision layer
Deterministic safety engine
- An explainable attention-evidence score combines critical-region coverage, reading evidence, latency, visibility, presence and the session pattern.
- Risk sets the thresholds: routine requests are never blocked; critical ones pause when the critical consequence received near-zero attention.
- Every decision carries plain-language reasons, and approval re-enables only after the missed consequence is visually reviewed or manually acknowledged.
Intervention ladder
Level 0
Normal
No interruption.
Level 1
Nudge
Subtle highlight; approval proceeds.
Level 2
Refocus
Approve becomes “Review critical consequence →” and the sentence is brought into focus.
Level 3
Pause
Only for high/critical risk with near-zero attention: the request collapses to the missed consequence until it is re-reviewed.
Across the session, OverSight also watches the sequence of approvals. When review attention declines across consecutive approvals it reports an approval fatigue pattern, raises intervention sensitivity and can switch to critical-only review mode. This describes behavior in the interaction; it is not a claim that anyone is tired.
Privacy architecture
Video never leaves this device.
We monitor the approval interaction, not the employee.
- Camera frames are processed in the browser tab and discarded; no frame is uploaded, stored or recorded.
- The page's Content Security Policy only allows network connections back to this origin.
- No facial recognition or identity model; no age, gender, ethnicity or emotion inference.
- Only eye-direction and blink coefficients are read from the face model; every other coefficient is discarded.
- Stored: derived numbers only (coverage, latency, dwell), in this browser session. Calibration is a few dozen numbers.
- The only server call sends approval-request text for semantic analysis, never camera data.
- A camera-free mode with manual acknowledgement is always available.
Honest limitations
- Commodity webcam gaze is approximate (often 100 to 200 px of error). OverSight therefore checks whole regions, not individual words.
- Looking at a sentence is evidence of inspection, not of comprehension. OverSight never claims the latter.
- Calibration covers moderate head movement. Resizing, zooming or moving the window makes it stale, and OverSight stops using gaze until you recalibrate.
- Glasses, strong backlight and low light reduce landmark quality; the system degrades to behavioral signals rather than guessing.
- The ML classifier ships untrained. Any accuracy figure must come from your own labeled sessions.
- Rule-based semantic analysis covers common high-risk patterns; the optional AI layer extends coverage.