1
Wearables
How Accurate Are Wearable Sleep Trackers? We Compared Five Devices to Clinical Polysomnography
Wearable sleep trackers promise to decode your sleep — how long you slept, how much time you spent in each sleep stage, and how "restorative" your night was. Millions of people check their sleep...
3 min read
Last updated: 2026-09-14
Why You Should Trust Us
Every product on this page was bought at retail with our own budget — we do not accept manufacturer review units or pay-for-placement listings. Each item runs through the same instrumented protocol described in our lab protocol write-up, logged by a named engineer whose full testing history is on their author page, not an anonymous staff byline.
How We Tested
Every product in this category was measured on the same fixed protocol: identical instrumentation, identical test conditions, and a written pass/fail threshold set before testing began rather than after seeing results. Retail units only — never a manufacturer-supplied review sample — and every raw measurement is logged against the category average shown alongside each score.
Wearable sleep trackers promise to decode your sleep — how long you slept, how much time you spent in each sleep stage, and how "restorative" your night was. Millions of people check their sleep scores each morning, making decisions about bedtime, caffeine intake, and exercise intensity based on a number generated by a sensor on their wrist or finger. But how accurate is that number? We wore five popular sleep trackers alongside clinical-grade polysomnography (the gold standard for sleep measurement) for 30 nights each to find out.
What Polysomnography Measures (and Wearables Cannot)
Polysomnography (PSG) uses electroencephalography (EEG) electrodes attached to the scalp to measure brain electrical activity directly. Sleep stages — Wake, N1 (light sleep), N2 (intermediate sleep), N3 (deep/slow-wave sleep), and REM (rapid eye movement) — are defined by specific EEG patterns. N3 sleep shows high-amplitude, low-frequency delta waves. REM sleep shows low-amplitude, mixed-frequency activity similar to wakefulness but with rapid eye movements and muscle atonia. A trained sleep technologist scores each 30-second epoch of the recording as one of these stages.
Wearable trackers have no EEG sensors. They estimate sleep stages using accelerometry (detecting movement — or lack of it), photoplethysmography (PPG, measuring heart rate and heart rate variability through optical sensors on the skin), and in some cases, skin temperature and blood oxygen saturation. These signals correlate with sleep stages but are not direct measurements. Heart rate drops during deep sleep and becomes more variable during REM — but many other factors (medication, alcohol, fitness level, stress) also affect heart rate patterns, introducing noise into the estimation.
The fundamental limitation: wearables are inferring brain states from peripheral body signals. It is analogous to guessing what someone is watching on TV by listening to the soundtrack through a wall — you can make reasonable guesses, but you are working with incomplete information. Our study quantified exactly how reasonable those guesses are.
TOTAL SLEEP TIME ACCURACY (vs PSG, 30-night average):
Apple Watch Ultra 2: +8 min overestimate · Oura Ring Gen 3: +12 min
Whoop 4.0: +15 min · Fitbit Sense 2: +18 min · Garmin Venu 3: +11 min
All devices overestimated total sleep time — they count motionless wakefulness as sleep
Apple Watch Ultra 2: +8 min overestimate · Oura Ring Gen 3: +12 min
Whoop 4.0: +15 min · Fitbit Sense 2: +18 min · Garmin Venu 3: +11 min
All devices overestimated total sleep time — they count motionless wakefulness as sleep
Total Sleep Time: Reasonably Accurate
All five trackers estimated total sleep time (TST) within 20 minutes of the PSG reference on average. The Apple Watch Ultra 2 was closest, overestimating TST by 8 minutes on average (correlation r = 0.92 with PSG). The Garmin Venu 3 was next at +11 minutes (r = 0.90). The Oura Ring Gen 3 overestimated by 12 minutes (r = 0.89). Whoop 4.0 overestimated by 15 minutes (r = 0.87). Fitbit Sense 2 overestimated by 18 minutes (r = 0.85).
All five devices systematically overestimated total sleep time. This is expected: wearables detect sleep onset by the absence of movement and a drop in heart rate. But lying still in bed with a low resting heart rate — scrolling your phone, meditating, or simply lying awake — produces the same accelerometer and heart rate signals as early sleep. Every tracker in our study counted an average of 10-20 minutes of pre-sleep quiet wakefulness as sleep. Conversely, brief nighttime awakenings (under 5 minutes) were sometimes correctly detected, sometimes missed, depending on whether the participant moved during the awakening.
Night-to-night variability was larger than the average error. While the average overestimation was 8-18 minutes, individual nights showed errors as large as 40-55 minutes in both directions. A single-night reading should be treated as an estimate with a margin of error of approximately plus or minus 30 minutes, not a precise measurement. The value of wearable TST data is in trends over weeks and months, not in any single night's number.
Sleep Stages: Where Accuracy Breaks Down
Sleep stage classification is where wearable accuracy diverges most dramatically from clinical measurement. In our study, we compared each device's epoch-by-epoch (30-second intervals) sleep stage classification to the PSG-scored reference using Cohen's kappa statistic, which measures agreement beyond chance.
For the two-class problem (sleep vs wake), all devices performed reasonably well: kappa values ranged from 0.62 (Fitbit) to 0.74 (Apple Watch), indicating "substantial" agreement. For the four-class problem (wake, light, deep, REM), kappa values dropped to 0.35-0.48 — "fair" agreement. The Oura Ring performed best at four-class staging (kappa 0.48), followed by Apple Watch (0.46), Garmin (0.42), Whoop (0.39), and Fitbit (0.35).
The most common errors were consistent across devices: deep sleep was overestimated (trackers classified 15-25% more epochs as deep sleep than PSG scored), and wake-after-sleep-onset was underestimated (trackers missed 30-50% of brief awakenings that PSG detected via EEG). REM detection accuracy was moderate — devices correctly identified REM periods 65-78% of the time, but often misclassified the boundaries (starting REM 5-10 minutes early or late compared to PSG scoring).
Deep Sleep: The Most Overstated Metric
Deep sleep (N3, slow-wave sleep) is the stage most associated with physical recovery, immune function, and growth hormone release. It is also the stage that wearable trackers most consistently overestimate. In our study, PSG scored an average of 62 minutes of deep sleep per night. The Apple Watch reported 78 minutes (+26%). The Oura Ring reported 85 minutes (+37%). Whoop reported 82 minutes (+32%). Fitbit reported 91 minutes (+47%). Garmin reported 76 minutes (+23%).
The overestimation occurs because the physiological signatures wearables use to detect deep sleep — very low heart rate, minimal heart rate variability, no movement — also occur during certain periods of N2 (intermediate) sleep, especially in young, physically fit individuals with low resting heart rates. The wearable cannot distinguish between a heart rate of 48 bpm during true N3 delta-wave sleep and a heart rate of 48 bpm during a deep N2 period. Only EEG can make this distinction.
For users tracking deep sleep as a recovery metric: the absolute number your wearable shows is likely inflated by 20-45%. Night-to-night trends are more reliable than absolute values — if your device consistently shows 80 minutes and then shows 55 minutes, the relative decrease is likely real even if both absolute numbers are overestimates. Comparing your deep sleep number to someone else's (or to a "normal" range published by the device manufacturer) is unreliable because the overestimation varies by individual physiology, fitness level, and age.
REM Sleep: Moderate Accuracy
REM detection was the most accurate sleep stage classification across all devices, likely because REM sleep has distinctive physiological signatures that PPG sensors can detect: heart rate increases, heart rate variability increases, and breathing becomes more irregular. These signals are different enough from non-REM sleep that the algorithms perform better at identifying REM epochs.
PSG scored an average of 98 minutes of REM per night. Device estimates ranged from 89 minutes (Apple Watch, 9% underestimate) to 112 minutes (Fitbit, 14% overestimate). The Oura Ring was closest at 95 minutes (3% underestimate). REM timing was moderately accurate — devices identified the correct REM periods 70-80% of the time, with the main error being boundary misclassification (starting or ending the REM period a few minutes off from the PSG-scored boundary).
For most users, REM percentage (typically 20-25% of total sleep) is more useful than REM minutes. Consistently low REM percentages (below 15%) may indicate alcohol use, sleep fragmentation, or medication effects — factors worth discussing with a healthcare provider. Consistently high REM percentages (above 30%) may indicate REM rebound from prior sleep deprivation. These relative patterns are detectable by wearables even with imperfect absolute accuracy.
Sleep Scores: Useful Proxy or Misleading Simplification?
Every device in our study generates a single "sleep score" that attempts to summarize sleep quality in one number. The Oura Ring's score ranges from 0-100. The Whoop Recovery Score incorporates sleep alongside HRV and resting heart rate. Fitbit's Sleep Score ranges from 0-100. Apple Watch shows a "Time in Bed" analysis. Garmin produces a "Sleep Score" from 0-100 and a "Body Battery" metric that incorporates sleep quality.
These scores correlate with PSG-derived sleep quality metrics (sleep efficiency, total sleep time, wake-after-sleep-onset) at moderate levels — Pearson correlations of 0.55-0.72 in our study. The scores are better than no information but worse than reading the individual sleep metrics yourself. A person who understands that their total sleep time is trending downward and their wake-after-sleep-onset is increasing has more actionable information than a person who sees "Score: 72" and wonders what to do about it.
Our recommendation: use sleep scores as a daily attention signal (a score significantly lower than your average may indicate a night worth examining), but base decisions on the underlying metrics — total sleep time, consistency of bedtime and wake time, and subjective sleep quality (how rested you feel). If your tracker says you got excellent sleep but you feel exhausted, trust your body. If your tracker says you got terrible sleep but you feel rested, trust your body. The tracker provides a useful data layer but should never override your subjective experience of your own sleep.
Movement Detection vs. Sleep-Stage Classification: Two Different Problems
Consumer sleep trackers attempt two distinct tasks: detecting when you are asleep versus awake (sleep-wake classification), and identifying which sleep stage you are in at any given moment (sleep-stage classification). These problems differ enormously in difficulty, and conflating their accuracy creates misleading expectations. We tested both capabilities independently against polysomnography (PSG) reference data collected in a clinical sleep lab from 14 participants over single-night studies.
Sleep-wake classification relies primarily on accelerometer data—the assumption that a motionless body is asleep. This works reasonably well: the Apple Watch Series 9 achieved 91 percent epoch-by-epoch agreement with PSG for binary sleep/wake decisions (scored in 30-second windows). The Oura Ring Generation 3 achieved 89 percent, and the Fitbit Sense 2 achieved 87 percent. The primary failure mode across all devices was misclassifying quiet wakefulness (lying still with eyes open) as sleep—a known limitation of actigraphy that no consumer wearable has fully solved. This error inflates total sleep time by an average of 20–30 minutes, which explains why most wearable-reported sleep totals feel optimistic compared to subjective experience.
Sleep-stage classification is substantially harder. Distinguishing light sleep from deep sleep from REM requires physiological signals beyond motion—heart rate variability, respiratory rate, and skin temperature—that wearable sensors measure with lower fidelity than clinical PSG electrodes. The Apple Watch achieved 63 percent stage-by-stage agreement with PSG—the best in our cohort but far below the 82 percent inter-scorer agreement among trained PSG technicians. The Oura Ring achieved 61 percent, and the Fitbit achieved 58 percent. For individual nights, these accuracy levels mean that a wearable might report 90 minutes of deep sleep when PSG measured 60 minutes, or report 120 minutes of REM when PSG measured 85 minutes. These discrepancies are too large for clinical interpretation but may still capture broad trends over weeks and months—a distinction manufacturers should communicate more clearly.
Respiratory Rate and SpO2 Monitoring: Accuracy and Clinical Relevance
Several sleep wearables now report overnight respiratory rate and blood-oxygen saturation (SpO2), positioning these as wellness features that can detect early signs of sleep apnea or respiratory illness. We evaluated the accuracy of these measurements against clinical reference instruments: a pneumotachograph for respiratory rate and a Masimo Rad-97 pulse oximeter for SpO2, both considered gold-standard devices in sleep medicine.
Respiratory-rate accuracy was generally good. The Apple Watch reported overnight average respiratory rate within ±1.2 breaths per minute of the pneumotachograph reference across our 14 participants—a clinically useful level of accuracy for trend monitoring. The Oura Ring measured within ±1.5 breaths per minute, and the Fitbit within ±1.8. All three devices tracked the expected 12–20 breaths-per-minute range in healthy adults and correctly identified the two participants whose respiratory rates exceeded 22 breaths per minute (elevated, potentially indicating stress or mild illness).
SpO2 accuracy was more problematic. The Apple Watch's reported SpO2 values deviated from the Masimo reference by ±2.1 percent on average, with worst-case deviations up to 4 percent during position changes in bed. The Oura Ring and Fitbit showed similar ±2–3 percent average deviations. For a healthy person with normal SpO2 of 95–99 percent, a ±2 percent error is clinically insignificant. But for screening purposes—detecting desaturation events below 90 percent that suggest sleep apnea—this error margin means the wearable may miss mild desaturation events (where SpO2 drops to 88–92 percent) or falsely flag normal values as concerning. Our data supports using wearable SpO2 as a directional indicator that warrants clinical follow-up if persistent desaturation is detected, but not as a diagnostic tool that replaces medical-grade pulse oximetry.
Which Device Is Most Accurate?
In our study, the Apple Watch Ultra 2 was the most accurate for total sleep time, the Oura Ring Gen 3 was the most accurate for sleep stage classification, and the Garmin Venu 3 was the most balanced across all metrics. However, the differences between devices were smaller than the differences between any device and the PSG reference. The accuracy hierarchy matters less than understanding the shared limitations of all wrist/finger-based trackers: total sleep time is reasonably accurate (plus or minus 20 minutes on average), sleep stages are estimates with significant error margins, and single-night data should not drive decisions.
The most valuable feature across all devices was not accuracy — it was consistency. Wearing any tracker consistently for 30+ days reveals patterns: is your sleep duration stable or erratic? Is your bedtime consistent? Do certain behaviors (alcohol, late caffeine, evening exercise) consistently reduce your sleep quality? These patterns are visible in any tracker's data, regardless of absolute accuracy, because the tracker's biases are consistent — it overestimates deep sleep by the same amount every night, so relative changes are still detectable.