Speech carries more than words. Alcohol intoxication leaves measurable traces in articulation, tempo and spectral micro-variation — signals the human ear registers unreliably at best.
XENOM's R&D team has trained a proprietary deep learning model to detect these acoustic biomarkers and evaluated it on the Alcohol Language Corpus, the standard public benchmark for this task.
Scored strictly like-for-like — per-window unweighted average recall, a single decision per utterance with no voting, exactly how prior work is measured — the reference-based configuration reaches 86.3% UAR. The reference-free configuration, which requires no enrolled baseline and decides from a single spoken sample, reaches 81.9%. At the deployment operating point, where the system combines several windows of a recording into a consensus decision, these rise to 87.6% and 82.5% respectively.
For context: the official Interspeech 2011 Speaker-State Challenge baseline on this corpus stands at 65.9% UAR, previously published state-of-the-art results peak at roughly 73–75%, and untrained human listeners score around 60–65%.
The system runs as a passive layer over routine radio traffic. Workers speak as they always do; no additional procedure, no interruption to the task, no separate checkpoint.
What the system is — and is not. It does not measure blood alcohol concentration and cannot serve as grounds for removing a worker from duty. It is a first-line screening tool that indicates who warrants attention, for verification by established means. The relevant comparison is not with a breathalyzer, but with the alternative used today: supervisor observation, which performs at roughly the level of untrained listening.
We have published full results across acoustic conditions — including those where accuracy degrades — along with the comparability caveats and known limitations of the method.