Research & validation

Make the model clear.
Put it to the test.

BioMIR is commercially available. Its proprietary longitudinal models are entering an external-validation and research-collaboration phase. Product availability and scientific validation are separate milestones.

Current Research Stage

Operational.
Ready for independent evaluation.

Adaptive BioAge is operational, with a public architectural specification; it has not yet been independently externally validated. The exact frozen reproduction package and coefficient-level provenance table are in preparation.

BioMIR seeks academic, clinical, epidemiologic, digital-biomarker, and preventive-health collaborators to evaluate reliability, calibration, longitudinal responsiveness, subgroup performance, and external validity. These are proposed studies, not results. Read the current Methods specification.

Three distinct evidence layers

Basis. Specification. Validation.

01

Scientific basis

Published research supports the study of constituent biomarkers and established methods. Population associations supply context; they do not establish a personal causal effect.

Explore Science ↗
02

BioMIR model specification

Input selection, reference normalization, weights, aggregation, Δ-year translation, and missing-data rules define what BioMIR computes. Reproducible execution is an implementation question.

Read Methods ↗
03

Validation status

Independent data must test the exact model, version, population, and intended use. Internal characterization or implementation checks do not substitute for external validation.

Review the status matrix ↓

Public evidence artifacts

Show the work.
Show what is not finished.

Auditability requires more than scientific prose. BioMIR separates materials that are already public from reproducibility artifacts that are still being prepared. An available document is evidence of specification or planning—not evidence that the model has been externally validated.

Available

Public Methods Specification

Current web specification of model architecture, evidence classes, weighting logic, data-quality semantics, and validation boundaries.

Read Methods ↗
Available on request · controlled distribution

Clinical Validation Protocol v1.3

Staged validation plan covering construct validity, precision, robustness, provenance, external validation, and release gates. Automatically delivered on request as the current controlled version with distribution tracking. Not a preregistration or completed validation dossier.

Request protocol →
Available on request · approval required

Agentic AI Implementation Protocol

Staged implementation roadmap covering canonical knowledge, deterministic queries, bounded tools, App Intents/Siri, Foundation Models, conversation, provenance, failure testing, and release gates. Distribution requires approval; no document or download link is sent automatically.

Request protocol →
In preparation

Frozen Model Specification v1.0

Planned signed specification of exact production equations, coefficients, weights, caps, units, transformations, ranges, eligibility rules, and edge-case handling.

In preparation

Coefficient Provenance Table

Planned coefficient-level lineage from source evidence and reference data through transformations, polarity, relative weighting, normalization, and implemented parameter values.

Planned

Model & Parameter Changelog

Planned change-control record identifying scientific parameter changes, rationale, model version, effective date, and whether historical outputs require recomputation.

Evidence-status matrix

What is supported.
What remains to be tested.

Status reflects the current research stage. This matrix describes the evidence boundary for each construct; it is not a report of completed validation studies.

BioMIR evidence status · 29 September 2026
ConstructScientific basisBioMIR specificationCurrent validation statusNext evaluation
BehavioralPublished epidemiologic and prognostic evidence informs the selected modifiable exposures and routines. BioMIR uses behavioral-adaptation language descriptively; it is not the standardized adaptive-behavior construct used in developmental assessment.Direction-aligned inputs, reference context, relative weights, and translation into a Behavioral Δ-year contribution.Specified and operational. Component literature does not externally validate BioMIR’s weighting or domain score.Independent calibration; sensitivity to logging, missingness, and exposure scale.
FunctionalPublished evidence on HRV, resting heart rate, aerobic fitness, and sleep supports the constituent measures; allostatic-load science provides broader context for cumulative multisystem dysregulation, but FUN is not a conventional allostatic-load index.Weighted standardized contributions from HRV, VO₂max, resting heart rate, total sleep, and deep sleep.Specified and operational. BioMIR’s composite and Δ-year translation still require independent evaluation.Repeatability, device effects, input-measurement error, and longitudinal responsiveness.
Cardiometabolic AgeEstablished vascular and metabolic relevance of systolic blood pressure, glucose, and BMI; age-regressed population reference context.BioMIR age-equivalent anchor with specified input selection, coefficients, carry-forward rules, and preserved observation dates.Specified and operational. Reference derivation is not external validation of the implemented age estimate.External calibration, subgroup performance, and sensitivity to dated or fallback inputs.
Adaptive BioAgeComponent biomarker literature and the temporal architecture provide a scientific rationale for study, not validation of the combined model.Cardiometabolic age-equivalent anchor plus Behavioral and Functional Δ-year terms, with domain and biomarker contributor attribution.Operational with a public architectural specification; exact frozen reproduction package in preparation. No independent external validation or established claim of lifespan change, treatment benefit, or surrogate-endpoint validity.Reliability, calibration, longitudinal responsiveness, subgroup performance, and external validity.
Klemera–Doubal MethodEstablished published Klemera–Doubal Method for combining age-related information across a biomarker panel.Implemented as a separate Clinical Long View using eligible clinical inputs and reference assumptions; not part of Adaptive BioAge.Published-method evidence is distinct from verification and validation of BioMIR’s implementation, panel, and target population.Reference-implementation agreement, panel/reference suitability, and independent population evaluation.
Levine PhenoAgeEstablished published outcome-linked clinical method using chronological age and nine laboratory biomarkers; distinct from DNAm PhenoAge.Implemented from a dated laboratory panel as a separate Clinical Long View; not part of Adaptive BioAge.The published method has an evidence base. That does not establish BioMIR-specific implementation accuracy or validate Adaptive BioAge.Formula and unit agreement, eligible-panel handling, and calibration in intended populations.
Data Confidence / FreshnessInput availability and observation recency inform interpretation. BioMIR’s displayed scores and time horizons are operational conventions.Daily input support; model-weighted CMA recency over 30 days; clinical-panel recency over 180 days. Original observation dates are retained.Specified and operational. These are not probabilities of correctness, clinical confidence levels, biological half-lives, or retesting schedules.Audit provenance and missingness handling; test whether indicators help users interpret estimates appropriately.

Evidence for component biomarkers does not automatically validate Adaptive BioAge. KDM and Levine PhenoAge remain separate clinical comparators. Δ-years is not a direct measurement of aging rate, lifespan gained or lost, disease probability, or treatment effect.

Domain terminology. BioMIR uses BEH for Behavioral and FUN for Functional. FUN integrates autonomic regulation, restorative physiology, and cardiorespiratory capacity across different timescales; it is not a multisystem allostatic-load index or a general measure of functional independence. BEH describes modifiable behavioral exposures and behavioral adaptation; adaptive behavior is an established developmental-assessment construct and is not used here as a formal synonym.

Current validation questions

Questions that can
constrain the claims.

Implementation fidelity & reproducibility

Can a frozen production model be reconstructed and recomputed deterministically from documented inputs, units, transformations, coefficients, caps, eligibility rules, and parameter versions?

Reliability

How repeatable are outputs under comparable conditions? Quantify within-person variability, device differences, repeated input acquisition, and plausible input-measurement noise.

Longitudinal responsiveness

Can sustained within-person changes be distinguished from transient biological variation, altered coverage, and regression to the mean? Compare with independently assessed physiological change.

Calibration

How do model outputs relate to prespecified reference measures and outcomes? Assess systematic error and reference-population mismatch without assuming that an age-like unit is literal biological age.

Missingness & freshness robustness

How sensitive are results to missing streams, carried observations, delayed records, source changes, and the operational CMA and Clinical Freshness functions? Stale data should not masquerade as physiological stability.

Subgroup performance

How do error, reliability, and interpretation vary across prespecified age, sex, health-status, device, and data-availability groups where sample sizes support evaluation?

External validity

Do findings reproduce in independent cohorts and settings with a locked model version? Agreement with another clock alone is not a gold standard or proof of clinical utility.

Future adaptive-layer validation

Personalization must
earn incremental value.

Methods defines the architectural boundary: the deterministic scientific core remains authoritative, while any future on-device or federated learning operates downstream. This research stream asks what must be demonstrated before adaptive personalization could become a released capability.

01

Incremental value over the deterministic baseline

Does a personalized downstream model improve prespecified longitudinal interpretation, calibration, anomaly detection, or response characterization beyond the locked deterministic BioMIR output? Evaluate gains on held-out or time-forward data and report where personalization adds no value.

02

Stability, drift & subgroup safety

Does personalization remain stable across device changes, missingness, changing routines, and distribution shift without feeding back into or silently altering the scientific core? Prespecify drift monitoring, update rules, rollback criteria, and subgroup analyses before deployment.

03

Privacy-preserving population learning

For any future federated approach, quantify utility alongside privacy and security: secure aggregation behavior, update stability, resistance to reconstruction or inference attacks, effects of privacy-preserving transformations, consent and withdrawal handling, and performance relative to appropriate non-federated benchmarks.

Separate validation track. These are future research questions, not capabilities or validation claims for the current released app. Evaluation of adaptive or federated layers must remain separable from validation of the deterministic core so gains, failures, and model changes can be attributed to the correct layer.

AI-readiness validation

AI should explain the model—
not become it.

BioMIR’s Apple Intelligence / AI-readiness protocol keeps scientific authority in deterministic BioMIR engines. Future language-model components are intended to consume bounded structured outputs and canonical BioMIR knowledge for explanation, narration, interaction, and orchestration. This validation stream asks whether that layer remains faithful, safe, private, and operationally optional.

01

Scientific fidelity & provenance

Do numerical statements exactly match the deterministic BioMIR outputs exposed through the semantic/tool layer? Can each analytical statement retain model version, period, units, freshness, Data Confidence, and supporting provenance so explanation remains distinguishable from computation?

02

Safety, ambiguity & version control

Test unsupported questions, missing or stale data, contradictory inputs, ambiguous periods, noisy trends, extreme values, clinical terminology, causal wording, and refusal to fabricate unavailable measurements. Prompt, tool-schema, canonical-knowledge, and language-model versions should be tracked separately from the validated analytical model.

03

On-device performance & graceful fallback

Measure latency, memory, energy use, model availability, device compatibility, offline behavior, narration interaction, and UI responsiveness. Apple Intelligence or another language model must remain additive: loss or unavailability of the AI layer must not impair normal BioMIR computation, visualization, or access to deterministic results.

Separate validation track. Canonical knowledge, semantic APIs, Foundation Models integration, App Intents/Siri interaction, and conversational analytics are staged readiness work—not current generative-AI product claims. Validation of this layer must not be used as a substitute for external validation of Adaptive BioAge or other BioMIR scientific outputs.

Parameter calibration pathway

Calibrate first.
Replicate independently.

Current parameter-refinement work uses the full founder longitudinal record—more than 2,000 days—to formally reassess provisional longitudinal parameters under a prespecified, version-controlled calibration protocol. This is model development, not external validation.

After a candidate parameter set is frozen, independently contributed, user-selected longitudinal exports can support N-of-1 series and cohort analyses of parameter stability, between-person heterogeneity, transportability, and responsiveness. Research participation remains separate and opt-in; exported records are not used automatically to update the production model.

Any proposed parameter revision should be documented in the model and parameter changelog, evaluated on held-out or time-forward data where feasible, and independently validated before production promotion.

Concrete collaboration studies

Three ways to build
independent evidence.

01 / Retrospective

External validation in an independent cohort

Study fit: cohorts with compatible, dated behavioral, wearable, cardiometabolic, and clinical data, with relevant outcomes where available.

Design: assess input compatibility; lock the algorithm, parameters, and analysis plan; evaluate on data independent of model derivation. Report missingness and selection effects.

Outputs: calibration and sensitivity analyses, subgroup performance, and comparisons with conventional measures. Any model revision requires a separate held-out evaluation.

02 / Prospective longitudinal

Repeatability & responsiveness

Study fit: repeated wearable or home measurements paired with standardized periodic clinical assessments.

Design: prespecify stable-condition repeat measurements, follow-up windows, device handling, and independently measured change; distinguish within-person signal from noise.

Outputs: repeatability estimates, sources of variability, and longitudinal responsiveness. Observed change alone does not establish intervention benefit or causation.

03 / Feasibility

Model-derived analytics & preventive-health workflows

Study fit: digital-health researchers and preventive-health teams evaluating exploratory analytics outputs, participant experience, and practical workflows.

Design: examine consent and data availability, completion burden, retention, contributor interpretation, and understanding of quality indicators.

Outputs: feasibility, usability, and protocol refinement. Engagement is not clinical efficacy; Adaptive BioAge is not an established surrogate endpoint.

Agree the study before exchanging participant-level data. Prespecify model version, reference data, endpoints, analysis responsibilities, consent or ethics review, access controls, publication policy, and permitted reuse. Availability of a dataset does not establish permission to use it.

Validation roadmap

From specification
to independent evidence.

The sequence is deliberate: freeze first, prespecify and preregister before confirmatory external analysis, then report the result whether favorable, null, or discordant. Later stages are plans, not completed milestones.

  1. 01
    Current

    Operational model + public specification

    Adaptive BioAge and its supporting architecture are implemented, with public Methods, evidence boundaries, and contributor-level interpretation.

  2. 02
    Current

    Working validation protocol

    Validation Protocol v1.3 defines staged internal characterization and the requirements for an independently executable external-validation protocol.

  3. 03
    Next

    Freeze the external-validation package

    Complete the frozen model specification, coefficient provenance, data dictionary, predefined endpoints, subgroup plan, and statistical analysis plan before confirmatory external analysis.

  4. 04
    Next

    Preregister before external analysis

    Timestamp hypotheses, outcomes, exclusions, statistical methods, covariates, and contingency rules before accessing or analyzing the external validation dataset.

  5. 05
    Planned

    Independent retrospective validation

    Evaluate a locked BioMIR version in an independent cohort, reporting calibration, sensitivity, missingness, subgroup performance, and negative or discordant findings.

  6. 06
    Planned

    Prospective longitudinal pilot

    Characterize repeatability and responsiveness under prespecified repeated-observation conditions and independently assessed longitudinal change.

  7. 07
    Planned

    Public scientific record

    Report methods and results through a preprint or peer-reviewed publication, then pursue independent replication where feasible.

For research teams & institutions

Bring a cohort.
Bring a testable question.

Tell us your institution, study type, available variables/observations and follow-up, and primary evaluation question. Start with a study summary; do not email participant-level health data.

BioMIR can discuss model documentation, contributor structure, and a versioned evaluation plan. Data access and study arrangements are agreed for each collaboration.

Propose a study

Longitudinal persistence index

BioMIR pairs a modeled age-equivalent contribution with a longitudinal persistence index: the proportion of qualified modeled days in High or Highest tiers, with fitted trends across prespecified periods. This is descriptive context, not disease probability or a validated clinical threshold. Independent studies must test repeatability, responsiveness and incremental information beyond the latest measurement and conventional averages, while separating held inputs from original observations and accounting for correlated repeated days. A 9 October 2026 draft amendment to the working v1.3 protocol specifies analytical checks, independent participant/time-forward evaluation, freshness sensitivity and human-factors endpoints; clinical and biostatistical review, preregistration and study execution remain pending.

The study plan freezes tiers, windows, denominator and version lineage. Clinical utility and any future referral thresholds require separate evidence and approval. Research collaborators can propose independently measured endpoints, acquisition schedules and participant-level external evaluation designs.