How does automatic sleep staging work?
Most staging software follows the same pattern. It reads the EEG, and often the EOG and EMG, splits the night into 30-second epochs, calculates features for each epoch (for example, power in different frequency bands), and passes them to a trained classifier that outputs a probability for each stage.
The open-source YASA algorithm is a well-documented example. It calculates time- and frequency-domain features for each 30-second epoch, adds smoothed versions of those features to capture context from surrounding epochs, and classifies them with a LightGBM gradient-boosting model. Its authors report that a full night sampled at 100 Hz typically processes in under 5 seconds on a consumer laptop. (Vallat & Walker, 2021 (opens in a new tab))
Where do sleep-staging algorithms make mistakes?
Errors are not spread evenly across the night. YASA's published evaluation is useful because it reports where its own model is weakest. These figures describe YASA on its authors' test datasets, not any particular installation.
| Situation | What the authors reported | What it means for review |
|---|---|---|
| N1 sleep | Lowest agreement of any stage: sensitivity 45.4%, median F1-score 0.432 | Review N1 epochs closely; human scorers also disagree most on N1 |
| Stage transitions | Mean accuracy 69.15% around transitions vs 94.08% in stable periods | Check epochs either side of a stage change |
| Fragmented nights | N1 percentage and transition count were the two strongest predictors of lower accuracy | Expect more corrections on disrupted sleep |
| Low confidence | Nights with higher average model confidence had higher accuracy (r = 0.76) | Use per-epoch confidence to prioritise review |
Source: Vallat & Walker, 2021 (opens in a new tab). The YASA documentation adds that the software should not replace human scoring, and recommends that a trained scorer check its predictions, especially low-confidence and N1 epochs. (YASA FAQ (opens in a new tab))
What does the human reviewer do?
- Open the night as a hypnogram and move epoch by epoch, with the EEG, EOG and EMG visible.
- Prioritise epochs that are likely to be wrong: low confidence, N1, and transitions.
- Confirm or correct each stage, adding a note where the reason matters.
- Keep the original machine output and every correction, so changes can be audited.
- Sign off only when every epoch has a review decision, then export the results.
This is the workflow InEpoch sets up and trains teams on: the application's review queues filter low-confidence, N1, transition and unreviewed epochs, and sign-off is only possible once every epoch has a review decision. See the InEpoch workflow.
How is autoscoring software validated or certified?
In the United States, the AASM runs an Autoscoring Certification Program that independently evaluates autoscoring software on real-world patient studies from accredited sleep centres. Full certification covers sleep stages, arousals, respiratory events and periodic limb movements. The AASM presents it as a supplement to FDA clearance, which compares a device with existing ones rather than validating it on a large independent dataset. (AASM (opens in a new tab))
Whatever a product's status, labs should also compare its output with their own scorers on their own recordings before relying on it. See how to evaluate sleep staging software.
Will AI replace sleep technologists?
Current evidence and guidance point to assistance, not replacement. The AASM position statement describes assisted scoring as an obvious early application of AI and sets out its limitations as well as its opportunities. (AASM, 2020 (opens in a new tab)) The work shifts from scoring every epoch by hand to reviewing, correcting and signing off, which still needs trained people.
Frequently asked questions
Is AI sleep scoring accurate?
Published algorithms can approach human agreement on clear epochs but are weakest on N1, stage transitions and fragmented sleep. Accuracy depends on the software, the recording and the population, so results should be checked against your own scorers.
Does autoscoring remove the need for review?
No. Autoscoring proposes stages; a trained technologist reviews, corrects and signs off. The YASA documentation explicitly says its staging should not replace human scoring.
Sources
- Goldstein CA et al. Artificial intelligence in sleep medicine: an AASM position statement. J Clin Sleep Med. 2020;16(4):605–607 (opens in a new tab)
- AASM Autoscoring Certification Program (opens in a new tab)
- Vallat R, Walker MP. An open-source, high-performance tool for automated sleep staging. eLife. 2021;10:e70092 (opens in a new tab)
- YASA documentation: FAQ (opens in a new tab)