# AI-assisted sleep scoring: how it works and where people fit

> AI-assisted sleep scoring uses software to propose a sleep stage for every 30-second epoch of a recording, usually with a probability or confidence value. A trained technologist then reviews the proposals, corrects errors and signs off. The AASM describes assisted scoring of polysomnography as one of the most apparent uses of AI in sleep medicine.

Published by InEpoch. Last updated: 2026-10-11. Canonical: https://inepoch.com/learn/ai-assisted-sleep-scoring

## How does automatic sleep staging work?

Most staging software follows the same pattern. It reads the EEG, and often the EOG and EMG, splits the night into 30-second epochs, calculates features for each epoch (for example, power in different frequency bands), and passes them to a trained classifier that outputs a probability for each stage.

The open-source YASA algorithm is a well-documented example. It calculates time- and frequency-domain features for each 30-second epoch, adds smoothed versions of those features to capture context from surrounding epochs, and classifies them with a LightGBM gradient-boosting model. Its authors report that a full night sampled at 100 Hz typically processes in under 5 seconds on a consumer laptop. ([Vallat & Walker, 2021](https://doi.org/10.7554/eLife.70092))

## Where do sleep-staging algorithms make mistakes?

Errors are not spread evenly across the night. YASA's published evaluation is useful because it reports where its own model is weakest. These figures describe YASA on its authors' test datasets, not any particular installation.

*Where YASA's published evaluation found errors (Vallat & Walker, 2021)*

| Situation | What the authors reported | What it means for review |
| --- | --- | --- |
| N1 sleep | Lowest agreement of any stage: sensitivity 45.4%, median F1-score 0.432 | Review N1 epochs closely; human scorers also disagree most on N1 |
| Stage transitions | Mean accuracy 69.15% around transitions vs 94.08% in stable periods | Check epochs either side of a stage change |
| Fragmented nights | N1 percentage and transition count were the two strongest predictors of lower accuracy | Expect more corrections on disrupted sleep |
| Low confidence | Nights with higher average model confidence had higher accuracy (r = 0.76) | Use per-epoch confidence to prioritise review |

Source: [Vallat & Walker, 2021](https://doi.org/10.7554/eLife.70092). The YASA documentation adds that the software should not replace human scoring, and recommends that a trained scorer check its predictions, especially low-confidence and N1 epochs. ([YASA FAQ](https://yasa-sleep.org/faq.html))

## What does the human reviewer do?

1. Open the night as a hypnogram and move epoch by epoch, with the EEG, EOG and EMG visible.
2. Prioritise epochs that are likely to be wrong: low confidence, N1, and transitions.
3. Confirm or correct each stage, adding a note where the reason matters.
4. Keep the original machine output and every correction, so changes can be audited.
5. Sign off only when every epoch has a review decision, then export the results.

This is the workflow InEpoch sets up and trains teams on: the application's review queues filter low-confidence, N1, transition and unreviewed epochs, and sign-off is only possible once every epoch has a review decision. See [the InEpoch workflow](https://inepoch.com/#workflow).

## How is autoscoring software validated or certified?

In the United States, the AASM runs an Autoscoring Certification Program that independently evaluates autoscoring software on real-world patient studies from accredited sleep centres. Full certification covers sleep stages, arousals, respiratory events and periodic limb movements. The AASM presents it as a supplement to FDA clearance, which compares a device with existing ones rather than validating it on a large independent dataset. ([AASM](https://aasm.org/about/industry-programs/autoscoring-certification/))

Whatever a product's status, labs should also compare its output with their own scorers on their own recordings before relying on it. See [how to evaluate sleep staging software](https://inepoch.com/learn/sleep-staging-software-checklist).

## Will AI replace sleep technologists?

Current evidence and guidance point to assistance, not replacement. The AASM position statement describes assisted scoring as an obvious early application of AI and sets out its limitations as well as its opportunities. ([AASM, 2020](https://aasm.org/advocacy/position-statements/artificial-intelligence-sleep-medicine-position-statement/)) The work shifts from scoring every epoch by hand to reviewing, correcting and signing off, which still needs trained people.

## Frequently asked questions

### Is AI sleep scoring accurate?

Published algorithms can approach human agreement on clear epochs but are weakest on N1, stage transitions and fragmented sleep. Accuracy depends on the software, the recording and the population, so results should be checked against your own scorers.

### Does autoscoring remove the need for review?

No. Autoscoring proposes stages; a trained technologist reviews, corrects and signs off. The YASA documentation explicitly says its staging should not replace human scoring.

## Sources

1. [Goldstein CA et al. Artificial intelligence in sleep medicine: an AASM position statement. J Clin Sleep Med. 2020;16(4):605–607](https://aasm.org/advocacy/position-statements/artificial-intelligence-sleep-medicine-position-statement/)
2. [AASM Autoscoring Certification Program](https://aasm.org/about/industry-programs/autoscoring-certification/)
3. [Vallat R, Walker MP. An open-source, high-performance tool for automated sleep staging. eLife. 2021;10:e70092](https://doi.org/10.7554/eLife.70092)
4. [YASA documentation: FAQ](https://yasa-sleep.org/faq.html)
