Synthetic Speech Detection via Frame-Level Artifact Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current synthetic speech detection techniques are inadequate in identifying partially or fully synthetic audio recordings, often focusing on fully synthetic audio and failing to detect modifications within authentic speech, and are prone to over-fitting specific speech generation tools.

Innovation Solution

A machine learning system trained to generate speech artifact embeddings based on features extracted from various speech generators, using probabilistic linear discriminant analysis (PLDA) to compute scores and determine the presence of synthetic speech in audio clips, capable of detecting both partial and full synthetic audio waveforms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If detection techniques focus on fully synthetic audio recordings, then detection simplicity is improved, but detection accuracy for partially synthetic audio deteriorates

Engineering Contradiction:
Improvedetection simplicityVSAvoiddetection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The audio clip is divided into multiple frames, and each frame is independently analyzed for synthetic speech detection. The system processes audio data by segmenting it into temporal sections (frames), allowing detection at granular levels and enabling identification of partially synthetic portions within an otherwise authentic recording.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If detection models are trained on speech from specific generation tools, then detection precision for those tools is improved, but adaptability to other generation tools deteriorates

Engineering Contradiction:
Improvedetection precisionVSAvoidadaptability to different speech generators
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The machine learning model is designed with a universal architecture that can detect synthetic speech artifacts across multiple speech generation tools. By training on diverse synthetic speech data from various generators and focusing on general artifact patterns rather than tool-specific characteristics, the model achieves broad adaptability while maintaining detection precision across different speech synthesis systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Speed

If conventional detection techniques are used, then processing speed is improved, but detection capability for interleaved synthetic and authentic speech deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoiddetection capability
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The audio stream is segmented into frames that are processed independently and efficiently. This segmentation allows the system to maintain high processing speed by handling small, manageable units while simultaneously improving detection capability for interleaved synthetic and authentic speech through frame-level analysis and scoring.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback through score computation and threshold-based decision making. Each frame receives a synthetic speech score, and the system provides feedback by comparing scores against thresholds to determine authenticity. This feedback mechanism enables reliable detection of partially synthetic audio while maintaining efficient processing through automated decision rules.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240379112A1Detecting synthetic speech
Publication Date: 2024.11.14 SRI INTERNATIONAL
  • US20240379112A1 patent drawing
  • US20240379112A1 patent drawing
  • US20240379112A1 patent drawing

AI summary

In general, the disclosure describes techniques for detecting synthetic speech in an audio clip. In an example, a computing system may include processing circuitry and memory for executing a machine learning system. The machine learning system may be configured to process an audio clip to generate a plurality of speech artifact embeddings based on a plurality of synthetic speech artifact features. The machine learning system may further be configured to compute one or more scores based on the plurality of speech artifact embeddings. The machine learning system may further be configured to determine, based on the one or more scores, whether one or more frames of the audio clip include synthetic speech. The machine learning system may further be configured to output an indication of whether the one or more frames of the audio clip include synthetic speech.