I wore two sleep trackers for a month: they agreed on everything except the number I cared about

sleep biomarkers
I wore two sleep trackers for a month: they agreed on everything except the number I cared about

On the worst night of the month, one device told me I’d had 22 minutes of deep sleep. The other told me I’d had an hour and 47. Same body, same bed, same night, roughly 30 centimetres apart.

I’d been wearing both for a month by then — an Oura ring on one hand, an Apple Watch on the other wrist — with the vague plan of finding out which one was lying. What I got instead was more interesting, and it’s the argument of this post: the two trackers agreed almost perfectly about when I slept and disagreed wildly about how I slept — and the mortality evidence lives on the when side of that split, not the stage side. So read the thing as a clock rather than as a sleep lab — timing is the axis the validation studies support, and the stage breakdown is a classifier’s guess I’ve found nothing tying to health outcomes. Which is a claim about which half of the dashboard is worth reading, not a promise about how close either half is to the truth.

Bit anticlimactic as a conclusion. It also completely changed which screen I open at 7am.

What I actually did, and what to hold loosely

Thirty consecutive nights. Ring on the left index finger, Watch on the right wrist, both worn every night including the two I’d rather not itemise. Every morning I wrote down five numbers from each device before looking at either app’s summary screen: sleep onset time, wake time, total sleep time, deep sleep, and whatever each called its overall score.

Now the caveat that decides how much of this you should believe, and it’s a big one. I have no polysomnography. I did not sleep in a lab with electrodes on my scalp, so I cannot tell you which device was right on any night. All I measured was agreement between two devices — which is a strictly weaker thing than accuracy, and I want to be straight about the direction of the weakness.

Agreement can’t prove accuracy: two devices can be wrong in the same way and look reassuringly consistent. But disagreement does prove inaccuracy, and it puts a floor under it. If the ring says 40 minutes of deep sleep and the watch says 80, then whatever the truth was, at least one of them is out by at least 20 minutes, and quite possibly both are out by more. Disagreement is the minimum error, visible from my kitchen. That’s the one thing wearing two devices buys you that wearing one never will.

This is one person, thirty nights, two consumer devices, no ground truth. An anecdote about instruments, not a study about sleep.

The three results

Here’s the month, rounded, as I logged it.

What I compared Typical gap between the two devices on the same night
Time I fell asleep about 6 minutes
Time I woke up about 4 minutes
Total sleep time about 20 minutes
Deep sleep about 40 minutes

Read that column as a size, not a precision — I rounded everything, and a month of one person’s nights is a small sample by any standard.

On timing, they were nearly the same device. Sleep onset and wake times landed within a quarter of an hour of each other on all but two nights, and those two were both nights I’d read in bed for a while — a genuinely ambiguous boundary that I couldn’t have scored myself. If you’d handed me the two columns of sleep-onset times without labels, I couldn’t have told you they came from different manufacturers.

On total sleep time, they were close enough that I’d have drawn the same conclusion from either. Twenty minutes is around 5% of a seven-hour night. Which is not the same as either being within twenty minutes of the truth — they could both be an hour out in the same direction and I’d never know it from here.

On deep sleep, they disagreed by roughly the size of the number. Forty minutes of typical disagreement, on nights where the figures themselves were mostly between 45 and 90 minutes. And when I ranked the thirty nights by deep sleep according to each device and compared the two orderings, the top-five lists overlapped on one night. One.

Tempting to conclude they’re measuring different things. I can’t get there from here: two noisy estimates of the same quantity would scramble a ranking just as thoroughly. What the gap licenses is the floor and only the floor — on a typical night at least one of these devices was out by twenty minutes or more, and I have no way to say which, or whether both were.

So: my best deep-sleep night of the month depends entirely on which hand you ask.

Why the stage numbers fall apart (and the timing doesn’t)

This isn’t a problem peculiar to my two devices — it falls out of what the hardware can see, and it shows up across the devices the validation studies have tested. Which is not the same as saying my ring and my watch are blameless: with no sleep lab in the loop I can’t apportion the 40 minutes between the physics and the two classifiers, and Lee’s spread says implementations differ materially. All I can say is that the constraint is general, not that these particular two handled it well.

Polysomnography scores sleep stages from brain activity — EEG, plus eye movement and muscle tone. Deep sleep is defined by slow waves in the cortex. A ring or a watch has no access to any of that. It has movement, heart rate, heart-rate variability, skin temperature, sometimes blood oxygen, and a proprietary classifier trained to guess which stage those signals imply. It is inferring a brain state from a wrist.

The validation literature is refreshingly blunt about how that goes. In the Oura ring’s laboratory validation, 41 healthy young adults were recorded with the ring and full polysomnography on the same night. Summary measures held up well: sleep onset latency, total sleep time and wake after sleep onset weren’t significantly different from the lab, and the ring correctly sorted nights into under-6-hours, 6-to-7 and over-7 buckets 81% to 93% of the time. Epoch by epoch, though, it detected sleep with 96% sensitivity — and agreed with the lab on deep sleep 51% of the time. Specificity for detecting wake was 48%. (According to PubMed: de Zambotti et al., Behavioral Sleep Medicine, 2019, DOI.)

That pattern — excellent at “asleep”, mediocre at “awake”, poor at “which stage” — repeats across devices. Chinoy and colleagues put seven consumer trackers plus research actigraphy against polysomnography over three nights each in 34 adults, including a deliberately disrupted-sleep condition. Epoch-by-epoch sensitivity to sleep was at least 0.93 for every device; specificity for wake ran from 0.18 to 0.54; stage comparisons were, in the authors’ word, “mixed”. And the devices “tended to perform worse on nights with poorer/disrupted sleep” — the authors’ own hedge, and worth keeping, because performance varied by device rather than every one of them degrading. Which is of course precisely the condition you’d want them to handle. (According to PubMed: Chinoy et al., Sleep, 2021, DOI.)

And for the head-to-head people actually search for: a multicentre study scored 11 consumer trackers — including the Apple Watch 8 and the Oura Ring 3, the two on my hands — against polysomnography across 75 participants and 349,114 epochs. Macro F1 for stage classification ranged from 0.26 to 0.69 depending on the device, with different devices leading at different stages. (According to PubMed: Lee et al., JMIR mHealth and uHealth, 2023, DOI.)

Which is the answer to “is Oura more accurate than Apple Watch”, and it’s an unsatisfying one: accuracy varies by device and by metric, and the studies I’m citing can’t be lined up to say which of those matters more. Lee’s spread is real and material — 0.26 to 0.69 is the difference between useless and moderately good — but the devices that lead on one stage don’t lead on the next, and the sleep-detection figures come from a different study with a different device list, so you can’t line the two up and decompose them. What you can say is narrower, and bounded by the devices actually tested: across the eleven in Lee’s study, the best stage performance reached moderate and the worst was close to useless — and unless your own device’s maker has published a validation, nothing tells you where in that band it sits, or whether a model released since lands outside it.

Timing survives because timing needs less of what the hardware can’t do. To know roughly when you fell asleep and when you got up, a device has to find the sleep–wake boundary once at each end of the night, rather than classify every epoch in between. That’s the job these devices do best.

Best isn’t flawless, though, and the same numbers show you where it frays. Low specificity for wake means a device is biased towards scoring still-but-awake as asleep — so if you lie quietly in bed for half an hour, your onset time will read early. And the Chinoy devices tended to do worse on the disrupted-sleep nights — a tendency across the set, not a guarantee about any one of them. My own two disagreeing nights were exactly that shape: reading in bed, both devices guessing at a boundary I couldn’t have scored myself either.

So the honest version of the claim is narrower than “trackers get timing right”. It’s: timing is the part that holds up — the validation studies find sleep/wake detection good in the same devices where staging is poor — and it tends to degrade on precisely the broken nights you’d most want explained.

And I should be careful with my own number here, because it’s the exact thing this post keeps warning about. My two devices landed within about a quarter of an hour of each other on an ordinary night. That is agreement, not a tolerance: nobody put either of them next to a sleep lab, and the validations I’m citing report sensitivity and specificity rather than an error bound in minutes. So treat a quarter of an hour as what two devices did in one flat, and not as how close either was to the truth.

The bit I got wrong

I’d set this up as an accuracy audit. Two devices, one liar, find the liar. About three weeks in it occurred to me that I’d picked the wrong number to audit, which is a more embarrassing error than picking the wrong device.

I was auditing deep sleep because deep sleep is what the apps put at the top of the screen and what I’d been quietly competitive about. But work out what the number on my screen could actually be tied to, and the deep sleep figure comes up short. Nothing I’ve read ties a consumer wearable’s stage estimates to mortality — and I should be careful about how much that’s worth, since it’s the limit of my reading rather than a survey of the field. Sleep regularity is a different story, and a firmer one: it has been measured objectively, at scale, against deaths.

Two things I’m not saying. Not that such a study is impossible — you could absolutely run one, though feeding it a measure whose agreement with a sleep lab runs from poor to moderate, depending on the device and the stage, would make the answer hard to interpret rather than impossible to obtain. And not that slow-wave sleep doesn’t matter: whether sleep architecture measured properly, in a lab, predicts mortality is a separate question I haven’t gone looking into. This is a claim about the number a ring hands you, and nothing more.

Windred and colleagues calculated Sleep Regularity Index scores from more than 10 million hours of accelerometer data in 60,977 UK Biobank participants, then followed mortality for an average of six years. Against the least regular quintile, the fully adjusted all-cause mortality hazard ratios for the four more regular quintiles were 0.80, 0.75, 0.72 and 0.70 — so roughly 20% to 30% lower risk, with the steadiest fifth about 30% below the most erratic, after adjusting for age, sex, ethnicity and a stack of lifestyle and health factors. Crucially, sleep regularity predicted all-cause mortality more strongly than sleep duration did, and adding duration to a regularity model didn’t improve the fit. (According to PubMed: Windred et al., Sleep, 2024, DOI.)

One caveat on that number, because you’ll see a bigger one quoted: the widely repeated “20% to 48%” range — the paper’s own abstract included — comes from the minimally adjusted models. Once you control for the lifestyle and health factors, the top end comes down to about 30%. I’m using the conservative figures, which are the ones that survive the adjustment.

Read that alongside the validation studies and the two halves do join up — though more carefully than I first wrote it down. The Sleep Regularity Index is built from sleep/wake state, not from sleep stages. It asks, for any two days, how likely you are to be in the same state — asleep or awake — at the same clock time. No EEG, no architecture. That puts the mortality-linked metric on the axis where these devices are strongest: detecting sleep at all, which every device in the Chinoy study managed at a sensitivity of at least 0.93.

Two honest limits on that, because it would be very easy to oversell.

First, the SRI is not the two numbers I logged. It samples across the whole 24 hours, so naps and stretches of being awake in the middle of the night feed into it — which is precisely where wake specificity of 0.18 to 0.54 does its damage. My two devices agreeing on sleep onset and final wake to within six minutes does not show they’d agree on my SRI. I didn’t compute it, so I genuinely can’t tell you whether they would.

Second, Windred compared regularity against duration, and against nothing else. The study did not race regularity against deep sleep, so I can’t tell you regularity “beats” stage metrics — that race wasn’t run here, and I haven’t found anywhere it was. The honest asymmetry is about where the evidence exists at all: regularity has direct mortality data from objectively measured sleep in 60,977 people, and I haven’t found anything of the kind behind consumer deep-sleep numbers. Nobody is stopped from running that study — you’d just be feeding it a measure whose agreement with a sleep lab runs from poor to moderate depending on which device you happen to own, and getting an answer that’s correspondingly hard to read.

Even with both caveats, that’s still the answer to which screen is worth opening in the morning — and almost nobody’s app leads with it. I’ve written before about why when you sleep beats how long, and I’ll admit I wrote that post without registering that the same metric also sidesteps most of the instrument problem.

There’s a cost to the number I was chasing, too. Baron and colleagues gave it a name — orthosomnia — after a run of patients presenting for treatment of sleep problems they’d diagnosed from tracker data, where the pursuit of the perfect score had become its own source of insomnia. Their observation that stuck with me is that the data often felt more true to patients than validated measurement did. (According to PubMed: Baron et al., Journal of Clinical Sleep Medicine, 2017, DOI.) I wasn’t in any clinical danger over a deep sleep number my ring can’t really see — and I should be careful with that 51% too, since it’s epoch-by-epoch agreement with a sleep lab, not a reliability score for the total on my screen in the morning. The same study does report night-level errors — the ring underestimated deep sleep by around 20 minutes and overestimated REM by around 17 — but those are average biases across 41 people, not an error bar for my Tuesday. But I had definitely had mornings where it decided how tired I was.

The clock test

So here’s the framework I came out with, which I now run on any number a wearable hands me. Three questions, in order.

1. Does this number depend on stage classification? If it’s deep sleep, REM or light sleep, you’re reading a classifier’s guess rather than a measurement — and how good a guess it is varies by device and by stage. A composite score is the harder case: it mixes classifier guesses with better-measured inputs, so the question becomes how much of it rests on the guesses, which you can only answer if the maker publishes the weighting. In the Oura validation, epoch-by-epoch agreement with the lab was 51% for deep sleep, 61% for REM and 65% for light sleep; across the eleven devices Lee’s group tested, stage-classification performance ranged from poor to moderate. So there’s no single error rate to quote, which is itself the useful finding: unless your device publishes its own validation numbers, you don’t know which end of that range you’re standing on.

I’d like to tell you to read it as a monthly trend instead, and I can’t, because nothing I’ve cited supports that either. The validation studies compare epochs and nights against a sleep lab; none of them tests whether a device’s average moves when your actual sleep moves. Averaging thirty nights damps random noise, but a classifier that is systematically wrong stays systematically wrong no matter how many nights you pour into it. So the honest instruction is the unsatisfying one: don’t act on the stage numbers, in either direction, until somebody validates the trend. If it’s a timing or duration number, it’s much closer to a measurement.

2. Do two independent instruments agree about it? Put a second device on the same night and see whether the two numbers match. Because the night is the same night, any gap between them is measurement rather than physiology. Mine differed by about 40 minutes on deep sleep and six on sleep onset — onset, not lights-out, which the devices never saw.

Now be precise about what that buys, because I wasn’t at first and the distinction is the whole point of the exercise. Agreement does not show either device is right — two classifiers fed similar signals can share a bias and agree beautifully. Nor does it show either is repeatable: I never measured the same night twice on the same device, so I cannot tell you whether either would give me the same answer twice. What disagreement proves, and it’s the only thing this test proves, is that at least one of them is wrong by at least half the gap. A floor on the error. That’s a modest finding and it is worth more than the confident number on the screen, because it’s the only one here that’s actually established.

What the test emphatically isn’t is comparing one night against the next. Sleep architecture really does vary between nights that felt identical from the inside, so a big night-to-night swing is not evidence of a noisy device — it’s two explanations mixed together with no way to separate them. If you’ve only got one device, the honest substitute is to go and read its published validation numbers, not to interpret your own spread.

3. Can I deliberately change it? This is the one that quietly kills most sleep metrics — with a distinction I glossed over the first time I wrote it down. What I can set directly is my schedule: lights out at eleven, alarm at seven. What the tracker reports is sleep onset, and anyone who has lain awake for forty minutes knows those are not the same thing. Schedule is an input I control; onset is an output I influence. But it’s an output that follows its input closely enough to steer by, night after night, which is more than can be said for slow-wave sleep — there I can only do the things that make it likelier and wait to be told how it went. HRV sits at the far end of the same spectrum.

So: timing passes all three, with that caveat on the third. Deep sleep fails all three. The score is a mixture of both, and where it lands depends entirely on the recipe: Apple’s is published, so you can price it — 80 points on well-measured inputs, 20 on the weak wake signal, none on stages; Oura’s names two stage contributors among seven without saying how much they count, so you can’t do that sum at all. Which makes the score the one number whose verdict you have to look up rather than assume.

Would I keep it?

Not both, no. I gave the ring back after the month — not because it lost, but because two devices don’t make either one more accurate, and the experiment had already told me what it had to tell me. Whichever one you’ll wear every night without thinking about it is the right one. Charge cycles beat classifier quality.

What changed is what I look at. I now read three things: when I fell asleep, when I got up, and how much those two moved across the week. That’s it. It’s the only part of the dashboard that survived contact with a second opinion, and — conveniently — it’s the closest thing on that screen to the measure with 60,977 people’s mortality data behind it. Not the same thing as the SRI, as above. Much nearer to it than a deep sleep percentage.

We build Sarvita to pull sleep from Apple Health and hold it against the rest of the recovery picture, and I’d say the same there as anywhere: the useful signal is whether your sleep timing is drifting, not what last night’s architecture allegedly looked like. If you want the fuller picture of what sleep is doing to how you age, how many hours you actually need covers the duration side, and I’ve done this sort of two-instruments-one-body exercise before with biological age tests — where five tests disagreed by fourteen years for much the same reason.

The lesson generalises further than I’d like. Every consumer health device ships a mix of things it genuinely measures and things it confidently estimates, in the same font, on the same screen, with no indication of which is which.

A second device won’t sort them for you — I’d rather hoped it would. Two direct measurements can disagree through ordinary noise, and two estimates can agree beautifully because they share a bias, so a gap tells you somebody’s wrong without telling you which category you’re looking at. What the month actually bought me was narrower and, I think, still worth the month: on the numbers where they disagreed, I now know at least one device is out by half the gap, and that was enough to stop me reading the deep sleep figure at breakfast.

To actually sort the measured from the estimated you have to go at it from the other end — what can this sensor physically detect, and what is a model filling in? — and then read the validation paper. Less fun, considerably cheaper, and it answers in an afternoon the question I spent a month failing to answer with hardware.

Anyway. My bedtime is now embarrassingly consistent and my deep sleep is none of my business.

Related posts

What's your biological age?

Reading about longevity is step one. Knowing your number is step two — a free 2-minute test, no signup.

Check your biological age →