Accuracy exists at several levels

  • Sensor validity: does the device measure the underlying signal accurately?
  • Derived-metric validity: does an algorithm estimate sleep stage, HRV, VO₂ max, or another derived measure reliably?
  • Biomarker validity: is that measure associated with aging, function, disease, or mortality in the relevant population?
  • Composite-model validity: does the final age score predict or track the outcome it claims to represent?

Consumer wearable performance is heterogeneous

A 2024 living umbrella review synthesized 24 systematic reviews covering 249 non-duplicate validation studies and more than 430,000 participants. Performance varied substantially across devices and outcomes, and only a minority of commercially available devices had validation evidence for at least one biometric outcome.

Sleep staging illustrates the problem. Consumer devices can be useful for longitudinal sleep monitoring, but stage-level performance depends on sensors and algorithms and should not be treated as equivalent to polysomnography.

Longitudinal consistency can still be useful

A measurement does not need to replace a clinical gold standard to provide useful longitudinal context. Repeated measurements under similar conditions may reveal within-person patterns even when absolute accuracy is imperfect.

The critical requirement is to preserve measurement limitations rather than letting a polished composite score create false precision.

What BioMIR exposes

BioMIR uses selected Apple Health observations and keeps contributor-level values visible. Data Confidence summarizes daily input support, while CMA Freshness keeps the recency of persisted cardiometabolic observations explicit. These features support interpretation; they do not substitute for external model validation.

Selected references

  1. Doherty et al. (2024), living umbrella review of consumer wearable accuracy
  2. Schyvens et al. (2024), wrist-worn wearables versus polysomnography
  3. Lu et al. (2025), digital biomarkers of ageing