Vocal Tract Reconstruction for Deep-Fake Audio Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for distinguishing between organic and synthetic audio are ineffective, particularly in detecting deep-fake audio, as they rely on specific generation techniques and struggle to accurately identify inconsistencies in vocal tract anatomy.
Innovation Solution
A method that models the dimensions of a vocal tract based on audio samples using fluid dynamics and articulatory phonetics to estimate cross-sectional areas, identifying bigram-feature pairs and calculating theoretical vocal tract areas to differentiate between organic and deep-fake audio by measuring divergence from estimated anatomical constraints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If bi-spectral analysis or machine learning discriminators are used for deep-fake detection, then detection capability is improved, but the method becomes highly dependent on specific generation techniques and less effective against unknown methods
Solution Approach 1:
Instead of trying to detect deep-fakes by analyzing their synthetic characteristics directly, the patent inverts the approach by analyzing what is missing - specifically, the absence of anatomically plausible vocal tract configurations. The system reconstructs what the vocal tract geometry should be based on the audio signal and checks for anatomical impossibilities, rather than looking for artifacts of synthesis methods.
Solution Approach 2:
The patent replaces machine learning-based detection systems with a physics-based acoustic model of the human vocal tract. By using articulatory phonetics and fluid dynamics to model how sound propagates through anatomically constrained vocal tracts, the system achieves detection that is independent of specific deep-fake generation techniques, as the anatomical constraints are universal.
2Productivity
If comprehensive dictionaries and formant synthesis models are used to generate synthetic voices, then voice generation capability is improved, but the output becomes easily distinguishable from organic speech
Solution Approach 1:
The patent changes the approach from modifying high-level speech parameters to constraining low-level anatomical parameters. By ensuring that the reconstructed vocal tract geometries satisfy anatomical plausibility constraints, the system provides a new dimension of verification that goes beyond traditional speech synthesis quality metrics.
3Reliability
If generative machine learning models are used for voice reconstruction, then speech quality for medical and memorial applications is improved, but unauthorized deep-fake audio becomes possible
Solution Approach 1:
The patent introduces an intermediary verification layer - the vocal tract reconstruction and anatomical plausibility check - that sits between the generated audio and the final output. This intermediary system can verify whether audio was generated through anatomically plausible means without affecting the quality of authorized voice reconstruction applications.
Data Source
AI summary
A method is provided for identifying synthetic “deep-fake” audio samples versus organic audio samples. Methods may include: generating a model of a vocal tract using one or more organic audio samples from a user; identifying a set of bigram-feature pairs from the one or more audio samples; estimating the cross-sectional area of the vocal tract of the user when speaking the set of bigram-feature pairs; receiving a candidate audio sample; identifying bigram-feature pairs of the candidate audio sample that are in the set of bigram-feature pairs; calculating a cross-sectional area of a theoretical vocal tract of a user when speaking the identified bigram-feature pairs; and identifying the candidate audio sample as a deep-fake audio sample in response to the calculated cross-sectional area of the theoretical vocal tract of a user failing to correspond within a predetermined measure of the estimated cross sectional area of the vocal tract of the user.


