Sleepgenic is the dedicated sleep research arm of TrailGenic — not a supplement, not a clinic.  ·  sleepgenic.ai
Sleep Interpretation Library · How accurate are deep and REM sleep on wearables?

How Accurate Are Deep and REM Sleep on Wearables?

By Mike Ye & Ella

Consumer wearables usually identify sleep more reliably than wake or individual sleep stages. Deep- and REM-sleep estimates can be useful for within-device trends, but their accuracy and direction of error vary by device, algorithm, population, and night.

METRIC: Deep Sleep, REM Sleep, Sleep Stage Classification
DEVICES: Garmin / Oura / Apple Watch / Fitbit / WHOOP / Consumer Wearables

Sleepgenic preserves deep, REM, light, and awake values as source-reported estimates. The Score Layer records the device’s classification. The Physiology Layer looks for agreement with duration, HRV, sleeping heart rate, stress, and continuity. The Context Layer tests sleep opportunity, timing, training, altitude, travel, illness, and other exposures.

Stage data earns meaning through repeated measurement on the same device—not through pretending one consumer estimate is polysomnography.

Read the Sleepgenic Methodology → · Open the Sleepgenic Lexicon →

The common misread is treating a wearable stage label as a direct measurement of brain state—or assuming every brand is biased in the same direction. Validation studies show that stage performance and bias vary across devices. Some devices overestimate particular stages, others underestimate them, and firmware can change performance over time.

1. The short answer

Wearables are generally better at recognizing sleep than at distinguishing wake, light sleep, deep sleep, and REM sleep. Stage estimates can still be useful, but not as exact clinical measurements. Their accuracy varies by brand, model, algorithm, population, sleep quality, and even the particular night being recorded.

The useful question is therefore not “Is my REM number perfectly accurate?” It is: “Is this estimate changing consistently on the same device, and do the other signals support the change?”

2. How clinical sleep stages are measured

Polysomnography classifies sleep using brain activity, eye movements, muscle tone, airflow, respiratory effort, oxygen saturation, heart rhythm, and other channels. Consumer wrist and finger devices do not directly measure most of those signals. They infer stages primarily from movement and cardiovascular patterns, sometimes with oxygen or temperature data.

This difference is structural. A wearable is solving a classification problem from a smaller, indirect signal set. Even a sophisticated algorithm can only classify what its sensors allow it to observe.

3. What validation studies show

Recent head-to-head studies generally find high sensitivity for detecting sleep but lower and more variable performance for detecting wake and separating individual stages. In a 2024 comparison of Oura Ring Gen3, Fitbit Sense 2, and Apple Watch Series 8, sleep detection sensitivity was at least 95%, while stage sensitivity varied substantially. A 2025 laboratory comparison of six wrist-worn devices likewise found meaningful differences across devices and only fair-to-moderate multistage agreement for many products.

2024 three-device validation →
2025 six-device validation →

4. Why the direction of error matters

It is inaccurate to say that all wearables simply overestimate deep sleep. In published comparisons, the direction and magnitude of error differ by device and stage. A model may underestimate deep sleep, overestimate REM, miss wake after sleep onset, or change behavior after an algorithm update.

Population characteristics also matter. Age, body mass, sleep efficiency, skin contact, movement, and sleep disorders can affect performance. A validation result for healthy adults and one firmware version should not be treated as a permanent specification for every user.

5. How Sleepgenic uses imperfect stage data

Sleepgenic preserves stage durations exactly as the source reports them, but interprets them as estimates. A Garmin REM value is compared primarily with prior Garmin REM values—not treated as interchangeable with an EEG-scored clinical stage.

The methodology uses rolling personal baselines, multi-night direction, source consistency, and explicit context. This does not remove measurement error. It makes the error less likely to dominate the interpretation.

6. The Three-Layer reading

Score Layer: What stage minutes and stage percentages did the device report, and how did the composite score respond?

Physiology Layer: Did HRV, sleeping heart rate, stress, duration, and fragmentation move in a compatible direction?

Context Layer: Was sleep shortened, interrupted, shifted, affected by alcohol, illness, training load, altitude, travel, temperature, or noise?

A low REM estimate after an early awakening has a plausible opportunity explanation because REM periods tend to lengthen later in the sleep period. A low REM estimate during a normal-length, otherwise stable night requires a different reading—but still not a clinical conclusion from one observation.

7. When stage estimates become useful

Stage estimates become most useful when four conditions hold:

  • the same device and wearing practice remain consistent;
  • the pattern persists across multiple nights;
  • the change is meaningful relative to the person’s prior baseline;
  • other physiological or contextual evidence supports the interpretation.

A one-night stage anomaly is weak evidence. A repeated stage shift that aligns with sleep duration, fragmentation, HRV, resting heart rate, and a defined exposure is stronger observational evidence.

8. Bottom line

Wearable deep and REM values are neither useless nor equivalent to polysomnography. They occupy the middle: imperfect estimates that can become informative when read longitudinally, within device, and with explicit context.

Do not ask one stage number to prove what the sensors never directly measured. Ask the repeated pattern to show what changed.

Sources: Robbins et al., 2024 · Schyvens et al., 2025 · AASM consumer sleep technology position statement

During Sleepgenic’s historical baseline, Garmin estimated deep sleep at 20.9% of total sleep and REM at 13.6%. Those values sit above and below common polysomnographic reference ranges, respectively. That does not establish unusually high deep sleep or a clinical REM deficit.

Sleepgenic treats both values as Garmin-specific baseline estimates. Their primary value is showing whether Mike’s stage pattern changes relative to his own prior Garmin record and whether other physiology and context move in the same direction.

Related Across the Sleepgenic Property
← All Interpretation Articles Methodology →