Vocal Tract Reconstruction for Deep-Fake Audio Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for distinguishing between organic and synthetic audio are ineffective, particularly in detecting deep-fake audio, as they rely on specific generation techniques and struggle to accurately identify inconsistencies in vocal tract anatomy.

Innovation Solution

A method that models the dimensions of a vocal tract based on audio samples using fluid dynamics and articulatory phonetics to estimate cross-sectional areas, identifying bigram-feature pairs and calculating theoretical vocal tract areas to differentiate between organic and deep-fake audio by measuring divergence from estimated anatomical constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If bi-spectral analysis or machine learning discriminators are used for deep-fake detection, then detection capability is improved, but the method becomes highly dependent on specific generation techniques and less effective against unknown methods

Engineering Contradiction:
Improvedeep-fake detection capabilityVSAvoidgeneralization to unknown generation techniques
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

Instead of trying to detect deep-fakes by analyzing their synthetic characteristics directly, the patent inverts the approach by analyzing what is missing - specifically, the absence of anatomically plausible vocal tract configurations. The system reconstructs what the vocal tract geometry should be based on the audio signal and checks for anatomical impossibilities, rather than looking for artifacts of synthesis methods.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent replaces machine learning-based detection systems with a physics-based acoustic model of the human vocal tract. By using articulatory phonetics and fluid dynamics to model how sound propagates through anatomically constrained vocal tracts, the system achieves detection that is independent of specific deep-fake generation techniques, as the anatomical constraints are universal.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If comprehensive dictionaries and formant synthesis models are used to generate synthetic voices, then voice generation capability is improved, but the output becomes easily distinguishable from organic speech

Engineering Contradiction:
Improvesynthetic voice generation capabilityVSAvoidindistinguishability from organic speech
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the approach from modifying high-level speech parameters to constraining low-level anatomical parameters. By ensuring that the reconstructed vocal tract geometries satisfy anatomical plausibility constraints, the system provides a new dimension of verification that goes beyond traditional speech synthesis quality metrics.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If generative machine learning models are used for voice reconstruction, then speech quality for medical and memorial applications is improved, but unauthorized deep-fake audio becomes possible

Engineering Contradiction:
Improvevoice reconstruction qualityVSAvoidunauthorized impersonation risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an intermediary verification layer - the vocal tract reconstruction and anatomical plausibility check - that sits between the generated audio and the final output. This intermediary system can verify whether audio was generated through anatomically plausible means without affecting the quality of authorized voice reconstruction applications.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11694694B2Detecting deep-fake audio through vocal tract reconstruction
Publication Date: 2023.07.04 UNIV OF FLORIDA RESEARCH FOUNDATION INC
  • US11694694B2 patent drawing
  • US11694694B2 patent drawing
  • US11694694B2 patent drawing

AI summary

A method is provided for identifying synthetic “deep-fake” audio samples versus organic audio samples. Methods may include: generating a model of a vocal tract using one or more organic audio samples from a user; identifying a set of bigram-feature pairs from the one or more audio samples; estimating the cross-sectional area of the vocal tract of the user when speaking the set of bigram-feature pairs; receiving a candidate audio sample; identifying bigram-feature pairs of the candidate audio sample that are in the set of bigram-feature pairs; calculating a cross-sectional area of a theoretical vocal tract of a user when speaking the identified bigram-feature pairs; and identifying the candidate audio sample as a deep-fake audio sample in response to the calculated cross-sectional area of the theoretical vocal tract of a user failing to correspond within a predetermined measure of the estimated cross sectional area of the vocal tract of the user.