Skip to main content
Wearable ECG

Do Cardiologists Agree When Reading Smartwatch ECG Strips?

The watch says 'sinus rhythm.' Two cardiologists might not say the same thing about the same strip.

KM
Kate Maren Editor, KnowYourPrime
Established · see the file
For information only. This is not medical advice, diagnosis, or treatment, and it cannot account for your own health history. A reading on a consumer device is not a clinical measurement. If a number worries you or you have symptoms, talk to a qualified healthcare provider. Full disclaimer.

This article covers research on cardiologist agreement when interpreting single-lead ECG strips from wearable devices, not the accuracy of the device algorithms themselves. It does not cover treatment decisions made after a diagnosis.

Cardiologists reviewing the same single-lead ECG strips do not reliably agree with each other. One trial had a group of cardiologists independently review a shared set of point-of-care strips and found real variability in their reads, and a separate screening study measuring agreement between two independent cardiologist reviewers found only moderate concordance. The strip itself is a fixed piece of data, but the human interpretation layered on top of it is not as fixed as the phrase 'confirmed by a cardiologist' implies.

What people assume happens after the watch flags something

There's a comforting mental model a lot of people carry around: the smartwatch is the rough draft, and the cardiologist is the final answer. The algorithm might miss things or throw false alarms, but once a real doctor looks at the strip, you get a clean, correct verdict, or so the thinking goes. That's the implicit promise behind phrases like 'physician confirmed.'

The tension is whether that second layer is actually as solid as it sounds. If the watch is the noisy first pass and the cardiologist is the quiet, reliable second pass, then two cardiologists looking at the same strip should land in roughly the same place almost every time. I went looking at the research on this specific question, and that's not quite what happens.

What the interpretation studies actually measured

One trial designed specifically around this question had fifteen cardiologists each review a shared set of 200 point-of-care single-lead ECG strips collected from patients aged 65 and older, all of whom also had a same-day 12-lead ECG as the reference standard. The cardiologists were blinded to both the device's own algorithm output and the 12-lead result, and they were asked to classify each strip as sinus rhythm, atrial fibrillation, or unclassifiable, and to rate their own confidence. The setup was built to isolate human interpretation from the algorithm entirely, and it found measurable variability across reviewers looking at identical strips.

A separate population-based AF screening study took a related but distinct approach. Instead of many cardiologists reading a shared batch, it had two independent cardiologists each review overlapping single-lead strips collected from participants aged 65 and older who recorded four ECGs a day over one to four weeks. Out of thousands of ECGs, 1,843 recordings from 185 participants ended up reviewed by both cardiologists, and their agreement was calculated using a weighted statistical measure of concordance. The result was moderate agreement, not high agreement, at both the participant level and the individual strip level.

Neither study is about whether the AI or on-device algorithm is right or wrong. Both are specifically measuring whether trained human readers converge on the same answer when looking at the same limited, single-lead signal, a narrower and, in some ways, more uncomfortable question than device accuracy. That's already covered in more detail in the piece on whether wearable ECG accurately detects atrial fibrillation.

2 studies
  • Fifteen cardiologists independently reviewing a shared sample of 200 point-of-care single-lead ECG strips, blinded to the device algorithm and to the same-day 12-lead result, showed variability in sensitivity, specificity, confidence, and classification of rhythm.Randomized controlled trial substudy · Pipilas et al., American Heart Journal, 2024
  • Two independent cardiologists reviewing overlapping single-lead ECG strips from an AF screening population showed only moderate inter-rater agreement, with a weighted Cohen's kappa of 0.48 at the participant level and 0.58 at the ECG level.Feasibility study, population-based screening · Hibbitt et al., Europace, 2024
Claim rating: Established · see the file

Where the disagreement seems to come from

Neither of the two agreement studies pins the variability on one single cause, and I want to be precise about that rather than inventing a tidy mechanism. What both studies share is the underlying constraint: a single lead gives cardiologists less information than the 12-lead ECG they train on and default to as a reference standard elsewhere in the same research. Less signal, more room for two trained readers to land in different places on an ambiguous strip.

That constraint shows up again in adjacent research. A case report of a patient whose Apple Watch strip read as normal sinus rhythm at 75 bpm, while a same-day 12-lead ECG showed atrial flutter with a sawtooth pattern, illustrates a related but distinct problem. The algorithm was built and validated for AF specifically, and a regular-but-abnormal rhythm slipped past it. That's a device-algorithm limitation rather than a cardiologist-agreement one, but it underscores the same theme, a single-lead strip carries less information than the full picture, whether the reader is an algorithm or a specialist.

This also connects to the broader device-comparison picture I've looked at elsewhere. Work comparing consumer wearables side by side against a 12-lead reference has looked at how devices differ from each other in distinguishing sinus rhythm from AF. That's a related but distinct question from how human readers differ from each other looking at the same output, explored further in the piece on wearable ECG device comparisons and inconclusive readings.

The moderate-agreement finding comes from an AF screening population aged 65 and older using a specific handheld single-lead recorder, reviewed by cardiologists working from strips alone. It doesn't establish how agreement would look in a younger population, with a different device, or when cardiologists have additional clinical context (symptoms, history) alongside the strip rather than the strip in isolation.

What this doesn't tell us

None of the interpretation-agreement research addresses whether disagreement between cardiologists changes what happens to a patient downstream, whether that's a missed treatment, an unnecessary one, or something else. That's a separate question from measuring how often two readers agree on a rhythm classification, and it isn't something either agreement study was designed to answer.

I also want to be clear that these findings are about single-lead strips specifically. The reference standard both studies compare against, the 12-lead ECG, isn't the thing being disputed here. The disagreement is about how consistently trained specialists read the smaller, single-channel signal that wearables produce, a narrower and more specific claim than 'doctors disagree about heart rhythm.'

Common questions

If a cardiologist reviews a smartwatch strip, does that mean the diagnosis is settled?

Research measuring agreement between cardiologists reviewing the same single-lead strips found only moderate concordance, and separate work with fifteen cardiologists reviewing a shared strip sample found real variability in classification and confidence. A single cardiologist's read of a single-lead strip is not the same as full diagnostic certainty.

Why would two cardiologists disagree about the same ECG strip?

The studies available don't isolate one specific cause, but both point to the single-lead format itself as a shared constraint, offering less information than the 12-lead ECG cardiologists otherwise use as the reference standard.

Is this about the smartwatch algorithm being wrong, or about the doctors?

This is specifically about human interpretation. The agreement studies cited here compare cardiologist-to-cardiologist reads, not cardiologist-versus-algorithm accuracy, which is a separate body of research covered elsewhere in this cluster.

Does this apply to every wearable ECG device?

The two interpretation-agreement studies used point-of-care handheld single-lead devices in populations aged 65 and older. Agreement patterns in other age groups or with other device types were not measured in this research and shouldn't be assumed to be identical.