Reference standard
We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.
Voice biomarkers are measurable characteristics of speech that may reflect aspects of a person’s physical, neurological, cognitive or psychological state.
Virtuosis’s validation approach is designed for a clearly defined use, population and setting. We compare model outputs with an appropriate reference measure, evaluate performance on data not used for training and document both results and limitations. The exact protocol and metrics depend on the intended deployment.
For each intended use, we define the outcome, population and recording conditions before evaluating the model. We then assess whether performance remains credible beyond the data used during development.
We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.
We evaluate data reflecting the intended population, languages, devices and realistic recording conditions.
We test on held-out, external or prospective data that did not influence model development.
We document relevant performance metrics, uncertainty, subgroup results and known limitations.
Validation dimension
Examples
Why it matters
Reference standard
Clinical scale, diagnosis or outcome
Defines the target the model is expected to estimate.
Dataset
Cohort, demographics and language
Determines how broadly results may generalize.
Evaluation design
Held-out, external or prospective testing
Estimates performance beyond the development data.
Performance metrics
Sensitivity, specificity, AUROC and calibration
Makes accuracy and decision tradeoffs transparent.
Robustness
Device, noise, language and subgroup checks
Identifies where performance may change or fail.
For each Virtuosis validation, these dimensions are considered together. Performance observed in one cohort or recording setup is not assumed to transfer to another language, device, population or clinical setting.
Peer-reviewed studies help define relevant outcomes, acoustic features, confounders and reporting standards for our validation work. Across the examples below, researchers examine patterns including pitch variability, loudness rangelow summarize Virtuosis AI validation across eight intended use cases. Each model was evaluated against an appropriate clinical reference measure, with TheresultsbelowsummarizeVirtuosisAIvalidationacrosseightintendedusecases.Eachmodelwasevaluatedagainstanappropriateclinicalreferencemeasure,withperformancereportedassensitivityandspecificity.Resultsarespecifictotheevaluateddatasets,populations,languagesandrecordingconditions.performance reported as sensitivity and specificity. Results are specific to the evaluated datasets, populations, languages and recording conditions.e, vocal control, voice clarity, speech rate and pauses. The figures are study-specific results from the cited literature; they are not performance claims for Virtuosis.
Mental health
Published depression research has reported lower voice smoothness, vocal control, pitch variability, loudness range, clarity and speech rate, together with more pauses. Huang et al. (Stress — Salivary cortisol; 98% sensitivity and 92% specificity; 400 participants; French, Italian, English, Spanish, German and Portuguese. Anxiety — GAD-7; 61% sensitivity and 79% specificity; 1,132 participants; French, Italian, English, Spanish, German, Portuguese and Chinese. Depression — PHQ-9; 77% sensitivity and 83% specificity; 1,933 participants; French, Italian, Chinese, English, Spanish, German and Portuguese.) reported 77% sensitivity and 83% specificity; cohort size, medication and study design remain important limitations.
Cognitive health
Studies of Alzheimer’s disease and mild cognitive impairment have examined pauses, fluency, articulation and prosody. Nasrolahzadeh et al. (2018) reported 97% accuracy for Alzheimer’s disease and 95% for mild cognitive impairment, while systematic reviews caution that imbalance and limited external validation can inflate estimates.
Neurological health
Parkinson’s disease can affect phonation, articulation, pitch and rhythm. Benba et al. (2016) reported 98% sensitivity and 96% specificity; balanced datasets, independent cohorts and realistic recordings remain essential before such results can be generalized.
Metabolic health
The Colive Voice study analyzed recordings from 607 U.S. adults and reported up to 94.5% accuracy with sex-specific models for predicting type 2 diabetes status. The finding supports further research into scalable screening signals, not standalone diagnosis.
Virtuosis uses peer-reviewed research as scientific context, then evaluates models for the intended use, population and recording conditions. We do not assume that performance reported in one study transfers to a Virtuosis deployment.
Important: a voice biomarker is not automatically a diagnosis. Results may be affected by age, sex, language, microphone quality, background noise, medication, temporary illness and other health conditions. Responsible systems define the intended use, validate against an appropriate reference standard, test generalizability and communicate uncertainty.