Methods

AI-assisted sleep scoring: how it works and where people fit

AI-assisted sleep scoring uses software to propose a sleep stage for every 30-second epoch of a recording, usually with a probability or confidence value. A trained technologist then reviews the proposals, corrects errors and signs off. The AASM describes assisted scoring of polysomnography as one of the most apparent uses of AI in sleep medicine.

How does automatic sleep staging work?

Most staging software follows the same pattern. It reads the EEG, and often the EOG and EMG, splits the night into 30-second epochs, calculates features for each epoch (for example, power in different frequency bands), and passes them to a trained classifier that outputs a probability for each stage.

The open-source YASA algorithm is a well-documented example. It calculates time- and frequency-domain features for each 30-second epoch, adds smoothed versions of those features to capture context from surrounding epochs, and classifies them with a LightGBM gradient-boosting model. Its authors report that a full night sampled at 100 Hz typically processes in under 5 seconds on a consumer laptop. (Vallat & Walker, 2021 (opens in a new tab))

Where do sleep-staging algorithms make mistakes?

Errors are not spread evenly across the night. YASA's published evaluation is useful because it reports where its own model is weakest. These figures describe YASA on its authors' test datasets, not any particular installation.

Where YASA's published evaluation found errors (Vallat & Walker, 2021)
SituationWhat the authors reportedWhat it means for review
N1 sleepLowest agreement of any stage: sensitivity 45.4%, median F1-score 0.432Review N1 epochs closely; human scorers also disagree most on N1
Stage transitionsMean accuracy 69.15% around transitions vs 94.08% in stable periodsCheck epochs either side of a stage change
Fragmented nightsN1 percentage and transition count were the two strongest predictors of lower accuracyExpect more corrections on disrupted sleep
Low confidenceNights with higher average model confidence had higher accuracy (r = 0.76)Use per-epoch confidence to prioritise review

Source: Vallat & Walker, 2021 (opens in a new tab). The YASA documentation adds that the software should not replace human scoring, and recommends that a trained scorer check its predictions, especially low-confidence and N1 epochs. (YASA FAQ (opens in a new tab))

What does the human reviewer do?

  1. Open the night as a hypnogram and move epoch by epoch, with the EEG, EOG and EMG visible.
  2. Prioritise epochs that are likely to be wrong: low confidence, N1, and transitions.
  3. Confirm or correct each stage, adding a note where the reason matters.
  4. Keep the original machine output and every correction, so changes can be audited.
  5. Sign off only when every epoch has a review decision, then export the results.

This is the workflow InEpoch sets up and trains teams on: the application's review queues filter low-confidence, N1, transition and unreviewed epochs, and sign-off is only possible once every epoch has a review decision. See the InEpoch workflow.

How is autoscoring software validated or certified?

In the United States, the AASM runs an Autoscoring Certification Program that independently evaluates autoscoring software on real-world patient studies from accredited sleep centres. Full certification covers sleep stages, arousals, respiratory events and periodic limb movements. The AASM presents it as a supplement to FDA clearance, which compares a device with existing ones rather than validating it on a large independent dataset. (AASM (opens in a new tab))

Whatever a product's status, labs should also compare its output with their own scorers on their own recordings before relying on it. See how to evaluate sleep staging software.

Will AI replace sleep technologists?

Current evidence and guidance point to assistance, not replacement. The AASM position statement describes assisted scoring as an obvious early application of AI and sets out its limitations as well as its opportunities. (AASM, 2020 (opens in a new tab)) The work shifts from scoring every epoch by hand to reviewing, correcting and signing off, which still needs trained people.

Frequently asked questions

Is AI sleep scoring accurate?

Published algorithms can approach human agreement on clear epochs but are weakest on N1, stage transitions and fragmented sleep. Accuracy depends on the software, the recording and the population, so results should be checked against your own scorers.

Does autoscoring remove the need for review?

No. Autoscoring proposes stages; a trained technologist reviews, corrects and signs off. The YASA documentation explicitly says its staging should not replace human scoring.

Sources

  1. Goldstein CA et al. Artificial intelligence in sleep medicine: an AASM position statement. J Clin Sleep Med. 2020;16(4):605–607 (opens in a new tab)
  2. AASM Autoscoring Certification Program (opens in a new tab)
  3. Vallat R, Walker MP. An open-source, high-performance tool for automated sleep staging. eLife. 2021;10:e70092 (opens in a new tab)
  4. YASA documentation: FAQ (opens in a new tab)