Deepfake Audio Detection via Phoneme Turbulence Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for distinguishing between organic and synthetic audio are often dependent on specific generation techniques and fail to generalize across datasets, making them susceptible to adaptive adversaries and ineffective against deepfake audio impersonation.
Innovation Solution
The method employs analysis of turbulent flows by converting speech into text, aligning it with phonemes, filtering for predetermined phonemes, and using a Weiner filter to transform frequency response vectors into classification space vectors, which are then normalized and compared to thresholds to identify whether the audio is synthetic or organic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If bi-spectral analysis or machine learning discriminators are used for deepfake detection, then detection capability is improved for specific generation techniques, but the method becomes highly dependent on previously observed generation techniques and fails to generalize across datasets
Solution Approach 1:
The patent segments the speech signal into individual phoneme instances and analyzes each phoneme's turbulent flow characteristics independently. By examining multiple phoneme-level measurements rather than treating the entire speech signal as a single entity, the system captures diverse turbulent flow patterns across different phonemes, enabling generalization to unseen deepfake generation techniques while maintaining high detection precision.
Solution Approach 2:
The patent replaces machine learning-based discriminators with a physics-based turbulence analysis approach. Instead of using trained models that depend on observed generation techniques, the system measures physical turbulent flow characteristics of speech signals, which are inherent to organic speech production and independent of synthetic generation methods. This substitution enables generalization across datasets while maintaining detection capability.
2Measurement precision
If detection methods are tailored to specific generation techniques, then precision against known deepfakes is improved, but the system becomes susceptible to adaptive adversaries who can evade detection
Solution Approach 1:
The patent employs a self-service detection mechanism where the system measures intrinsic turbulent flow characteristics of speech signals without requiring external training data or knowledge of generation techniques. By relying on physical properties that are inherent to organic speech production, the system automatically adapts to new deepfake methods without retraining, maintaining both precision and reliability against adaptive adversaries.
Solution Approach 2:
The patent changes the detection parameter from high-level spectral features used by machine learning models to fundamental physical parameters of turbulent flow. By measuring parameters such as frequency response variations and temporal dynamics of turbulent phonemes, the system detects deepfakes based on physical inconsistencies rather than statistical patterns, making it reliable against adaptive adversaries who cannot easily replicate physical turbulence.
Data Source
AI summary
A method is provided for identifying synthetic “deepfake” audio samples versus organic audio samples. Methods may include: receiving an audio sample comprising speech; converting the speech to text; aligning the text with phonemes identified within the audio sample; filtering the audio sample to only contain predetermined phonemes; obtaining, from the audio sample, a frequency response vector for each of the predetermined phonemes; transforming the frequency response vector for each of the predetermined phonemes to a classification space vector for each of the predetermined phonemes having a magnitude; normalizing the classification space vector for each of the predetermined phonemes; identifying each of the predetermined phonemes as one of synthetic or organic based on the classification space vector for each of the predetermined phonemes; and identifying the audio sample as synthetic or organic based on identification of each of the predetermined phonemes as one of synthetic or organic.


