Every major sleep tracker compresses a night into a single number. The formulas are proprietary, the weightings are unpublished, and two devices on the same wrist will disagree. Knowing which inputs the number is built from — and which of those inputs are measured well — is what makes the score usable.
The Verdict
What a sleep score is actually measuring
Underneath every brand's number are the same physiological measurements, collected the same way. An optical sensor fires light into tissue and reads the returning signal to derive pulse timing. An accelerometer records movement. Usually a thermistor records skin temperature. From pulse interval data the device derives heart rate and HRV; from movement plus heart rate it estimates whether you were asleep and, less reliably, which stage you were in.
The score is then a weighted sum of derived quantities. No brand publishes its weights, but the behavior of the scores makes the ranking obvious: duration dominates. A 5-hour night cannot score well no matter how clean the architecture, and an 8-hour night with mediocre stages usually still scores respectably.
How each wearable calculates it
| Brand | What the score is called | Inputs it uses | Notes |
|---|---|---|---|
| Apple Watch | Sleep Score, 0–100 in the Health app | Duration (50 pts), bedtime consistency (30 pts), interruptions (20 pts) | The only brand here that publishes its point weightings. Sleep stages are recorded but do not feed the score. |
| Fitbit | Sleep Score, 0–100 | Duration, sleep depth (deep + REM), restoration (resting HR and restlessness) | Longest-running of these scores, so the largest population baseline behind it. |
| Garmin | Sleep Score, 0–100 | Duration, stages, stress and HRV overnight, restlessness | Reports the score alongside training load and body battery context. |
| Oura | Sleep Score, 0–100 | Total sleep, efficiency, restfulness, REM, deep sleep, latency, timing | Finger PPG plus skin temperature; publishes the seven contributors individually. |
| Whoop | Sleep Performance, % of calculated need | Hours vs need, sleep debt, disturbances, latency, efficiency | Expresses the result against your own need rather than a fixed 100-point scale. |
The table is listed alphabetically by brand. The structural difference worth noticing is Whoop's: expressing sleep as a percentage of your calculated need makes the score self-referential, so a person who needs 7 hours and gets 6.5 can score higher than a person who needs 9 and gets 7. The 0–100 scales instead judge the night against a general adult model. Neither approach is more correct; they answer slightly different questions.
The inputs, and how well each one is actually measured
| Input | Typical weight | Measurement quality and what to know |
|---|---|---|
| Total sleep duration | Heaviest on every platform | Nothing else compensates for a short night. Duration alone drives most of the score variance. |
| Sleep efficiency | High | Time asleep ÷ time in bed. Under 80% is the clearest sign that time in bed is not converting. |
| Latency | Moderate | Under 5 minutes is scored well by most algorithms but usually indicates sleep debt, not efficiency. |
| Deep and REM proportions | Moderate | The weakest input, since consumer stage classification runs 60–80% agreement with lab scoring. |
| Wake after sleep onset (WASO) | Moderate | Over 30 minutes awake after falling asleep is a fragmentation signal that scores heavily. |
| Resting heart rate and HRV | Moderate on some platforms | The channel through which last night's alcohol reaches this morning's score. |
| Restlessness / movement | Low–moderate | Sensitive to band tightness and to sharing a bed with a restless partner or a dog. |
| Timing consistency | Low–moderate | Some algorithms penalise a bedtime that moves against your own recent average. |
That accuracy split is the key to reading the number sensibly. The parts of the score built on duration, efficiency, and wake time are built on measurements consumer devices handle well. The part built on stage percentages sits on classification that agrees with polysomnography only 60–80% of the time. When your score drops and the app blames deep sleep, the blame is the least trustworthy element of the explanation.
What lowers a sleep score, ranked
- A short night. Duration is weighted heaviest everywhere, so an hour lost costs more than any architecture change.
- Alcohol within four hours of bed. Two drinks typically raise overnight resting heart rate 5–15 bpm, cut HRV 20–40%, fragment the second half of the night, and suppress REM. Four inputs, one cause.
- A bedroom above 70°F. Blocks the core-temperature drop that sleep depends on; shows up as fragmentation and lower deep sleep.
- Late caffeine. Half-life is 5–6 hours, and 8–10 in slow metabolisers carrying CYP1A2 variants. A 3pm coffee leaves roughly a quarter circulating at 11pm.
- Hard training within three hours of bed. Elevated core temperature and sympathetic tone raise overnight heart rate. Moderate evening exercise is fine; maximal efforts are not.
- A late heavy meal. Especially high-fat. Raises overnight heart rate and delays the temperature drop.
- An inconsistent schedule. A two-hour weekend shift is a two-time-zone circadian displacement, and most algorithms penalise the resulting misalignment.
How to act on the score, and when not to
Use it as a change-detector, not as a target. A useful workflow is weekly rather than daily: look at the 7-day average against the 30-day average, and if the gap exceeds about 5 points, look for a cause in the order above. Beyond removing what is dragging the number down, the changes that move objective sleep metrics are ranked rather than equal, and our sleep guides cover which levers carry the most leverage. Single-night reactions are almost always wrong, because same-person night-to-night variation and device error each run large enough to produce a 10-point swing with nothing physiological behind it.
Do not chase a 95. Score-chasing produces two predictable failures: extending time in bed beyond your need, which lowers efficiency and is a common on-ramp to conditioned insomnia, and orthosomnia, where anxiety about the number measurably degrades the sleep it is measuring. If you wake rested and function well through the afternoon, a 78 does not require a response.
A high sleep score with a low readiness or recovery score
The two numbers are not measuring the same thing, and a gap between them is information rather than a contradiction. A sleep score grades the night: how long you slept, how broken it was, when it happened. A readiness or recovery score grades your current physiological state, weighting resting heart rate and heart rate variability against your own baseline, and it carries forward yesterday's training load, alcohol, illness, and heat.
So a 90 sleep score alongside a low readiness usually means the night was fine and something else is taxing you. The common causes, in rough order of frequency: hard training in the previous 48 hours, an infection you have not noticed yet, alcohol the evening before, a warm bedroom, and dehydration. The response is to treat the readiness number as the actionable one that day and to leave the sleep behaviour alone, since it was not the problem. The reverse pattern, a poor sleep score with normal readiness, usually means a single short or disrupted night that your body absorbed without difficulty.
Platform design explains part of the split. Apple builds its score entirely from duration, bedtime consistency, and interruptions, with no heart-rate input at all, which is set out on our Apple Watch sleep score page. Oura and Whoop both fold physiological signals into their equivalents. A device that never measured your recovery cannot tell you about it, and a device that did will disagree with one that did not.
How sleep scores drift with age
Slow-wave sleep declines substantially from young adulthood onward, and total sleep becomes more fragmented with age. Any score that rewards deep sleep or penalises awakenings therefore trends downward across decades for reasons that are normal rather than correctable, which is why a 55-year-old comparing a 76 against a 25-year-old's 88 is comparing two different physiologies.
Brands do not publish whether or how their scoring is age-adjusted, so you cannot tell from the number which of them corrects for this. That makes your own 30-day and 90-day averages the only reliable reference, and it makes a downward trend within your own history far more meaningful than any comparison against a friend, a leaderboard, or a population figure.
When the score is misleading — and when to see a physician
The most important failure mode runs the wrong direction: a good score with bad daytime function. Obstructive sleep apnea produces arousals lasting a few seconds, which is often too brief for a wrist algorithm to register as waking. A person with 30 events an hour can therefore post consistent 80s. If you sleep 7–8 hours, score well, and still wake unrefreshed for more than a month, believe the fatigue over the number.
Signs that shift this from tracking to testing: loud snoring, witnessed breathing pauses, morning headaches, dry mouth on waking, more than two bathroom trips a night, or overnight blood oxygen dips below 90% on a device that measures it. An estimated 80% of moderate-to-severe cases are undiagnosed, home testing is now inexpensive, and the most-missed presentation is the lean, fit adult with a narrow airway.
Two other patterns belong with a clinician rather than an algorithm. Latency over 30 minutes or WASO over 30 minutes, three nights a week for three months, meets the definition of chronic insomnia — for which CBT-I, not a sleep aid, is first-line. And a score that collapsed within weeks of starting a new medication is a prescriber conversation: SSRIs and SNRIs suppress REM, lipophilic beta-blockers such as propranolol blunt REM and melatonin, and Z-drugs raise N2 without raising slow-wave sleep, so all three can move the number without your habits changing at all.
Frequently Asked Questions
What is a good sleep score?
On any of the 0–100 scales, 80+ is generally considered good and 90+ excellent, while Whoop expresses the same idea as Sleep Performance where 85%+ of calculated need is a common target. Those cutoffs are conventions set by each brand, not clinical thresholds, and they are not comparable to each other. The more useful benchmark is your own 30-day average: a score sitting 10 points below it for a week means something, and a single 72 does not.
Why do two devices score the same night differently?
Different sensors and different weightings. A ring reads finger arterial signal and skin temperature; a wrist band reads wrist arterial signal and forearm movement. Each brand then applies a proprietary formula whose weights are not published. It is completely normal for the same night to score 87 on one device, 79 on another, and 84% Sleep Performance on a third. All three can be internally consistent and none is the true answer.
What drops a sleep score the most?
Short total sleep, by a wide margin — duration is weighted heaviest on every platform. Alcohol is second and it is the most efficient score-killer because it hits four inputs simultaneously: it fragments the night, suppresses REM, raises resting heart rate 5–15 bpm, and cuts HRV. After that: a bedroom above 70°F, a late heavy meal, late caffeine, and hard training within three hours of bed.
Is the sleep score more reliable than the individual stages?
Yes, for a specific reason. The score is dominated by duration and efficiency, which consumer devices measure reasonably well — total sleep time is typically within 10–20 minutes of lab scoring. Stage classification, which is the weak input at 60–80% agreement, contributes only part of the composite. Averaging a strong signal with a weak one produces something more stable than the weak signal alone.
Can a good sleep score be wrong?
Routinely. The score cannot see anything the sensors miss. Someone with untreated obstructive sleep apnea can post 80s consistently because they lie still, sleep long, and wake only briefly at each event — the arousals are too short for a wrist algorithm to catch reliably. If you score well and still feel wrecked after 7–8 hours for several weeks, believe how you feel and get an apnea test.
Does checking the score make sleep worse?
It can, and the pattern has a name: orthosomnia. People who read the score immediately on waking report worse subjective sleep on days the number looks bad, and some develop performance anxiety about sleep that lengthens latency the following night. Two practical rules cut most of the harm — do not open the app within an hour of waking, and review the weekly average rather than each morning's number.
How do I read my sleep score if I work night shifts?
Expect the composite to run lower than your sleep deserves, because a rotating schedule costs you points on both total duration and timing consistency, for reasons you do not control. Read the underlying components instead, since total sleep duration and sleep efficiency are the best-measured inputs on every platform and they carry most of what the composite is compressing. Split sleep and daytime naps are handled inconsistently between brands, so a day slept in two blocks may not be totalled the way you expect. If you sleep adequate hours across the day and still feel unwell through your waking hours for several weeks, that belongs with a clinician rather than with the algorithm.
Related
- Sleep guides — stages, scores, and what moves them
- How much sleep do I need
- How much deep sleep do I need
- REM sleep explained
- Whoop recovery score explained