Voice biomarker guide

Clinical validation

Voice biomarkers are measurable characteristics of speech that may reflect aspects of a person’s physical, neurological, cognitive or psychological state.

How Virtuosis validates voice biomarkers

Virtuosis’s validation approach is designed for a clearly defined use, population and setting. We compare model outputs with an appropriate reference measure, evaluate performance on data not used for training and document both results and limitations. The exact protocol and metrics depend on the intended deployment.

Our validation principles

For each intended use, we define the outcome, population and recording conditions before evaluating the model. We then assess whether performance remains credible beyond the data used during development.

1

Reference standard

We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.

2

Representative data

We evaluate data reflecting the intended population, languages, devices and realistic recording conditions.

3

Independent evaluation

We test on held-out, external or prospective data that did not influence model development.

4

Transparent reporting

We document relevant performance metrics, uncertainty, subgroup results and known limitations.

What we document for each validation

Validation dimension

Examples

Why it matters

Reference standard

Clinical scale, diagnosis or outcome

Defines the target the model is expected to estimate.

Dataset

Cohort, demographics and language

Determines how broadly results may generalize.

Evaluation design

Held-out, external or prospective testing

Estimates performance beyond the development data.

Performance metrics

Sensitivity, specificity, AUROC and calibration

Makes accuracy and decision tradeoffs transparent.

Robustness

Device, noise, language and subgroup checks

Identifies where performance may change or fail.

For each Virtuosis validation, these dimensions are considered together. Performance observed in one cohort or recording setup is not assumed to transfer to another language, device, population or clinical setting.

Virtuosis clinical validation results

Peer-reviewed studies help define relevant outcomes, acoustic features, confounders and reporting standards for our validation work. Across the examples below, researchers examine patterns including pitch variability, loudness rangelow summarize Virtuosis AI validation across eight intended use cases. Each model was evaluated against an appropriate clinical reference measure, with TheresultsbelowsummarizeVirtuosisAIvalidationacrosseightintendedusecases.Eachmodelwasevaluatedagainstanappropriateclinicalreferencemeasure,withperformancereportedassensitivityandspecificity.Resultsarespecifictotheevaluateddatasets,populations,languagesandrecordingconditions.performance reported as sensitivity and specificity. Results are specific to the evaluated datasets, populations, languages and recording conditions.e, vocal control, voice clarity, speech rate and pauses. The figures are study-specific results from the cited literature; they are not performance claims for Virtuosis.

Mental health

Stress, anxiety and depression

Published depression research has reported lower voice smoothness, vocal control, pitch variability, loudness range, clarity and speech rate, together with more pauses. Huang et al. (Stress — Salivary cortisol; 98% sensitivity and 92% specificity; 400 participants; French, Italian, English, Spanish, German and Portuguese. Anxiety — GAD-7; 61% sensitivity and 79% specificity; 1,132 participants; French, Italian, English, Spanish, German, Portuguese and Chinese. Depression — PHQ-9; 77% sensitivity and 83% specificity; 1,933 participants; French, Italian, Chinese, English, Spanish, German and Portuguese.) reported 77% sensitivity and 83% specificity; cohort size, medication and study design remain important limitations.

Cognitive health

Parkinson’s disease, Alzheimer’s disease and MCI

Studies of Alzheimer’s disease and mild cognitive impairment have examined pauses, fluency, articulation and prosody. Nasrolahzadeh et al. (2018) reported 97% accuracy for Alzheimer’s disease and 95% for mild cognitive impairment, while systematic reviews caution that imbalance and limited external validation can inflate estimates.

Neurological health

Respiratory function

Parkinson’s disease can affect phonation, articulation, pitch and rhythm. Benba et al. (2016) reported 98% sensitivity and 96% specificity; balanced datasets, independent cohorts and realistic recordings remain essential before such results can be generalized.

Metabolic health

Type 2 diabetes

The Colive Voice study analyzed recordings from 607 U.S. adults and reported up to 94.5% accuracy with sex-specific models for predicting type 2 diabetes status. The finding supports further research into scalable screening signals, not standalone diagnosis.

Source and scope

Virtuosis uses peer-reviewed research as scientific context, then evaluates models for the intended use, population and recording conditions. We do not assume that performance reported in one study transfers to a Virtuosis deployment.

Important: a voice biomarker is not automatically a diagnosis. Results may be affected by age, sex, language, microphone quality, background noise, medication, temporary illness and other health conditions. Responsible systems define the intended use, validate against an appropriate reference standard, test generalizability and communicate uncertainty.

Discuss a clinical validation project.

Request a demo
By clicking “Accept”, you agree to the storing of cookies on your device to enhance site navigation and analyze site usage. View our Privacy Policy for more information.