Synthetic Speech Detection via Frame-Level Artifact Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current synthetic speech detection techniques are inadequate in identifying partially or fully synthetic audio recordings, often focusing on fully synthetic audio and failing to detect modifications within authentic speech, and are prone to over-fitting specific speech generation tools.
Innovation Solution
A machine learning system trained to generate speech artifact embeddings based on features extracted from various speech generators, using probabilistic linear discriminant analysis (PLDA) to compute scores and determine the presence of synthetic speech in audio clips, capable of detecting both partial and full synthetic audio waveforms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If detection techniques focus on fully synthetic audio recordings, then detection simplicity is improved, but detection accuracy for partially synthetic audio deteriorates
Solution Approach 1:
The audio clip is divided into multiple frames, and each frame is independently analyzed for synthetic speech detection. The system processes audio data by segmenting it into temporal sections (frames), allowing detection at granular levels and enabling identification of partially synthetic portions within an otherwise authentic recording.
2Measurement precision
If detection models are trained on speech from specific generation tools, then detection precision for those tools is improved, but adaptability to other generation tools deteriorates
Solution Approach 1:
The machine learning model is designed with a universal architecture that can detect synthetic speech artifacts across multiple speech generation tools. By training on diverse synthetic speech data from various generators and focusing on general artifact patterns rather than tool-specific characteristics, the model achieves broad adaptability while maintaining detection precision across different speech synthesis systems.
3Speed
If conventional detection techniques are used, then processing speed is improved, but detection capability for interleaved synthetic and authentic speech deteriorates
Solution Approach 1:
The audio stream is segmented into frames that are processed independently and efficiently. This segmentation allows the system to maintain high processing speed by handling small, manageable units while simultaneously improving detection capability for interleaved synthetic and authentic speech through frame-level analysis and scoring.
Solution Approach 2:
The system implements feedback through score computation and threshold-based decision making. Each frame receives a synthetic speech score, and the system provides feedback by comparing scores against thresholds to determine authenticity. This feedback mechanism enables reliable detection of partially synthetic audio while maintaining efficient processing through automated decision rules.
Data Source
AI summary
In general, the disclosure describes techniques for detecting synthetic speech in an audio clip. In an example, a computing system may include processing circuitry and memory for executing a machine learning system. The machine learning system may be configured to process an audio clip to generate a plurality of speech artifact embeddings based on a plurality of synthetic speech artifact features. The machine learning system may further be configured to compute one or more scores based on the plurality of speech artifact embeddings. The machine learning system may further be configured to determine, based on the one or more scores, whether one or more frames of the audio clip include synthetic speech. The machine learning system may further be configured to output an indication of whether the one or more frames of the audio clip include synthetic speech.


