Skip to content

Session Protocol

Standardise this and follow it. Without a fixed procedure, your comparisons carry the variance of your testing conditions rather than the athlete’s ability.

FactorRequirement
DeviceSame iPad every time. Log the model
BrightnessSame setting, normal range
Viewing distanceSame, roughly arm’s length
Time of daySame where possible
Fatigue stateNot immediately post-exercise
Trial count60 for anything recorded
Footage setSame set for any comparison
InstructionsSame script, read aloud, every time

A full 60-trial session that you discard.

First exposure measures learning the task, not anticipation. Skipping this makes every athlete’s first recorded session an underestimate, and their second look like improvement when nothing changed.

Run two full sessions and compare them.

If the two thresholds agree within 2 × the standard error, the athlete is stable and single sessions can be trusted going forward.

If they do not agree, do not yet trust individual results. Either the athlete’s performance is not yet stable, or something in your procedure is varying. Run a third before concluding anything.

This is the test–retest reliability of your own instrument in your own hands. It cannot be inherited from someone else’s setup.

Retest every 4–6 weeks unless studying a specific short-term intervention.

Anticipation is a perceptual-cognitive skill; it does not move week to week under normal training. Testing more often mostly measures noise, and repeated exposure to the same footage introduces familiarity with specific clips rather than general skill.

There is no session history in the app. Record at minimum:

Date ____________________
Athlete ____________________
Footage set ____________________
Device ____________________
Trials ______
Threshold ________ ms ± ________ ms
or DID NOT CONVERGE — reason: _______________
Accuracy (corrected) ________ %
No-response rate ________ %
Median RT ________ ms
Reversals ________
Timing-invalid ________

A change is real only if it exceeds roughly 2 × the standard error.

If a baseline is 133 ± 12 ms and a retest is 141 ± 14 ms, that is noise. If the retest is 175 ± 13 ms, that is a change worth investigating.

Never compare:

  • Across sports — different chance levels
  • Across viewpoints — different task
  • Across footage sets — different stimulus

If you are writing this up, the method is:

A temporal-occlusion paradigm with a 3-down-1-up adaptive staircase converging on 79.4% correct. Threshold estimated as the mean of the final six reversals; standard error as SD/√6. Sessions with fewer than six reversals or SE > 25 ms were rejected as non-convergent. Stimulus timing was hardware-verified (median blackout error 3.4 ms, p95 5.3 ms at 60 Hz); trials deviating by more than two display refreshes were flagged and excluded from threshold estimation. Response latencies were captured from the platform touch event in the same monotonic clock base as the measured blackout.

Every raw trial is stored: intended offset, measured offset, deviation, both timestamps, response zone, correctness, and timeout flag. Retrieving it currently requires direct access to the device database — see Current Limitations.