Free expedited shipping US board certified clinicians, all 50 states 24/7 unified patient care HSA · FSA eligible Compounded in the USA
← Living Vitality

Your Recovery Score Is a Model. Here's What's Actually Measured.

Your wearable does not measure your recovery — it models it from a chain of inferences at your wrist. Independent sleep labs have quantified where that chain loses fidelity, and what blood work measures instead.

Cover art: Your Recovery Score Is a Model. Here's What's Actually Measured.

Your wearable does not measure your recovery. It measures movement and a light signal at your wrist, estimates heart rate and heart-rate variability from that signal, infers sleep and wake, infers sleep stages from the inference, and then combines those layers into a single number using an algorithm the manufacturer does not publish. Every step of that chain is real engineering, and every step also loses fidelity — which independent sleep laboratories have now quantified. Understanding where the loss happens is the difference between using the number well and being ruled by it.

What does the accuracy research actually show?

Consumer wearables are very good at one thing and consistently poor at another. They are highly sensitive at detecting sleep — if you are asleep, the device almost always says so. They are much weaker at detecting wake during the night.

An independent laboratory validation published in 2025 tested six wrist-worn devices against polysomnography, the clinical reference standard, in 62 adults. Sensitivity for sleep ran from 91.7 to 96.3 percent across every device. Specificity for wake ran from 29.4 to 52.2 percent. Chance-corrected agreement, measured as Cohen’s kappa, ranged from 0.21 to 0.53 — fair to moderate, with no device reaching substantial agreement (Schyvens et al. 2025).

That is not one bad study. An earlier laboratory comparison of seven devices reported sensitivity of at least 0.93 for sleep alongside specificity between 0.18 and 0.54 for wake (Chinoy et al. 2021). A large review of the field states the pattern plainly: wearables show sensitivity above 90 percent in detecting sleep and lower specificity in detecting wake (de Zambotti et al. 2019).

The practical consequence is a systematic direction of error, not random noise. A device that struggles to recognise wake will tend to score fragmented nights as more consolidated than they were.

How much worse does it get at sleep stages?

Considerably. The same 2021 laboratory study reported four-stage classification — wake, light, deep, REM — at accuracy between 0.60 and 0.72, but with chance-corrected agreement of only 0.19 to 0.42 (Chinoy et al. 2021). Distinguishing deep from light from REM at your wrist is a substantially harder problem than distinguishing asleep from awake, and the numbers reflect that.

This is the layer most people care about, because “deep sleep” and “REM” are the figures that feel diagnostic. They are the least reliable outputs the device produces.

Is the recovery score itself validated?

We looked, and we want to be precise about what we found — which is an absence.

Whoop Recovery, Oura Readiness and Garmin Body Battery are proprietary composite scores. The formulas, weightings and reference populations are not published. Component algorithms have been studied in places, and the underlying signal quality is often good — a peer-reviewed validation of wrist photoplethysmography during rest and sleep found close agreement with ECG for heart rate and for RMSSD, a common heart-rate-variability measure (Bell et al. 2021). But we did not identify a peer-reviewed study validating any of these composite scores against a clinical outcome — illness, injury, measured fitness change, or anything else a patient would recognise as an outcome.

That absence is not a claim that the scores are useless. It is a claim about what kind of thing they are: consumer products designed for daily engagement, built on partially validated components, combined by an undisclosed method, and not evaluated as clinical instruments.

The professional bodies have said as much. The American Academy of Sleep Medicine’s position statement is unambiguous: given the lack of validation and FDA clearance, consumer sleep technologies cannot be used for the diagnosis or treatment of sleep disorders, though they may be useful in enhancing the conversation between a patient and a clinician when placed in the context of an appropriate clinical evaluation (AASM 2018).

Why aren’t these devices regulated as medical devices?

Because they are not marketed as ones. The FDA operates a compliance policy for low-risk products that promote a healthy lifestyle, set out in its General Wellness guidance, which describes an approach for products intended only for general wellness use and presenting low risk to user safety (FDA 2016). Staying inside that lane is a deliberate and entirely legitimate product decision. It is also the reason a recovery score is not held to the standard of a lab result — and the reason it should not be read like one.

What happens when people trust the number too much?

Sleep clinicians named this pattern in 2017. A case series described patients whose pursuit of ideal wearable sleep data was itself worsening their sleep, and proposed the term orthosomnia for it (Baron et al. 2017). It is worth being accurate about the evidence tier here: this was three cases, not a population study. What it establishes is that clinicians encountered the problem and thought it worth naming, not that it is common.

Still, the shape is familiar to anyone who has argued with their own wrist about whether they are recovered. If your device tells you your recovery is poor and you feel fine, both the device and you are working from partial information — but only one of you has direct access to how you feel. We looked at the same disagreement from the training side in when your recovery score says green but you feel overtrained.

What is actually measured, rather than modelled?

This is the part worth holding onto. There is a whole category of physiology that is directly measured in a laboratory, with standardised assays, defined reference intervals and reproducibility requirements — and much of it bears on exactly the fatigue and recovery questions a wearable is being asked to answer.

Iron status is the clearest example, and it comes with real trial evidence rather than inference. A double-blind randomised placebo-controlled trial in non-anaemic women with unexplained fatigue and ferritin below 50 µg/L found a significantly greater reduction in fatigue with iron than with placebo (Verdon et al. 2003), and a later randomised trial in non-anaemic menstruating women with low ferritin reached a similar conclusion (Vaucher et al. 2012). We should also say what the synthesis says: a systematic review of randomised trials in non-anaemic iron-deficient adults found a modest pooled reduction in fatigue while noting unclear risk of bias in most trials (Houston et al. 2018), and at least one randomised trial found no significant difference. The evidence leans positive and is not settled. That is a more useful thing to be told than a confident yes. It is also why “my labs are normal but I’m exhausted” deserves a closer read — a pattern we unpack in exhausted with normal labs.

Thyroid function, vitamin B12 and folate have long-established diagnostic pathways (Stabler 2013). Testosterone has a defined protocol that requires symptoms plus unequivocally low measurements on two separate morning draws, with the Endocrine Society explicitly recommending against routine population screening (Bhasin et al. 2018). High-sensitivity CRP is a validated cardiovascular risk marker, established through large randomised trial evidence (Ridker et al. 2008) — and we should be precise that it is validated for cardiovascular risk, not as a fatigue test.

Notice what these have in common. Each is a direct measurement with a published method. Each has a defined interval and a stated uncertainty. Each has a body of trial evidence you can go and read. None of them produces a single daily number, and that is the point — they are slower, less satisfying, and considerably more solid.

How should you use both?

Wearables are good at trend and at adherence. Watching your own resting heart rate drift over months, or noticing that your sleep is consistently shorter than you believed, is real information and worth having. Single-night stage percentages and a daily composite score are the weakest outputs and deserve the least weight.

Blood work sits at the other end: infrequent, but measured rather than inferred, and the thing that can identify a specific, addressable reason your recovery is not what it was. Used together — the wearable for pattern, the panel for cause — they answer different questions. Used interchangeably, the wearable gets asked a question it was never built to answer. This is the case for a measured baseline we make in full in why longevity medicine starts with your bloodwork.

Bottom line

A recovery score is a model built on a chain of inferences, and independent laboratory studies have quantified where that chain loses fidelity — particularly at wake detection and sleep staging. It is a reasonable tool for tracking your own trends over time. It is not a measurement of your physiology, and when it disagrees with how you feel, blood work is the layer that can settle the question.

This article is general education. It is not medical advice, it is not a diagnosis, and it does not establish a clinician-patient relationship. Device accuracy figures are quoted from independent peer-reviewed validation studies and reflect the specific device models and generations tested in those studies; newer models may perform differently.

Recovery score accuracy FAQ

How accurate are wearable sleep trackers? Very sensitive for sleep (above 90 percent), consistently weak for wake (roughly 29–52 percent specificity in lab testing) — so fragmented nights tend to score better than they were.

Are the sleep-stage readings reliable? They are the least reliable output: chance-corrected agreement with polysomnography of only 0.19–0.42 for four-stage classification.

Is the recovery score itself validated? The composite scores are proprietary and unpublished, and we did not identify a peer-reviewed study validating any of them against a clinical outcome.

Can a wearable diagnose a sleep disorder? No — the American Academy of Sleep Medicine states consumer sleep technology cannot be used for diagnosis or treatment, though it can enrich a clinical conversation.

What is measured rather than modelled? Laboratory physiology: iron status, thyroid function, B12 and folate, and other markers with published methods, reference intervals and trial evidence behind them.


When the score and how you feel disagree, measured physiology is the layer that settles it. Find Your Treatment →

Sources

Chief Medical Officer: Shannon Arora, MD

Shannon Arora, MD is the Chief Medical Officer of Trellis Vitality. This is a statement of role, not a review of this specific article.