Prosody Extractor for Synthetic Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Sophisticated deep learning models for voice generation and voice cloning produce extremely realistic synthetic speech, making it difficult for speaker recognition systems to distinguish between real and spoofed voices, posing a serious threat to individuals and organizations.

Innovation Solution

A computer-based system and method for detecting synthetic speech using a trained prosody extractor, which generates prosody embeddings from speech samples and compares them to reference embeddings to determine authenticity, employing a processor to train the prosody extractor with a loss function defined on spectrograms generated by a speech synthesis model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If sophisticated deep learning models for voice generation and voice cloning are used, then the realism of synthetic speech is improved, but the ability of speaker recognition systems to distinguish between real and spoofed voice deteriorates

Engineering Contradiction:
Improverealism of synthetic speechVSAvoiddetection accuracy of speaker recognition system
Core Design Contradiction:
Manufacturing precisionVSMeasurement precision

Solution Approach 1:

The patent introduces a prosody extractor as an intermediary component that bridges the gap between speech input and authenticity verification. This extractor specifically targets and isolates prosody features (rhythm, stress, intonation patterns) that serve as a mediator between the raw speech signal and the final authentication decision, enabling the system to detect subtle differences that conventional speakers verification misses

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts and isolates specific prosody features from the complete speech signal using a dedicated prosody extractor module. By separating these temporal and spectral characteristics from other speech components, the system focuses computational resources on the most discriminative features for detecting synthetic speech, thereby improving detection accuracy without being overwhelmed by the full complexity of the speech signal

Inventive Principle:
Principle #2Taking out (Extraction)

2Device complexity

If conventional speaker recognition systems are used, then the simplicity of the system is maintained, but the reliability of detecting synthetic speech deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidsynthetic speech detection reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the speech authentication process into distinct functional modules: a prosody extractor that isolates temporal and spectral features, a speech synthesis model that generates expected prosody patterns, and a comparison mechanism that validates authenticity. This segmentation allows each component to specialize in a specific task, improving overall reliability while maintaining manageable system complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary extraction of prosody features before the main authentication decision is made. By pre-processing the speech signal to isolate prosody characteristics and comparing them against expected patterns generated by the speech synthesis model, the system prepares discriminative information in advance, enabling more reliable detection of synthetic speech without adding excessive complexity to the final decision process

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12217762B1System and method for detecting synthetic speech based on prosody analysis
Publication Date: 2025.02.04 CORSOUND AI LTD
  • US12217762B1 patent drawing
  • US12217762B1 patent drawing
  • US12217762B1 patent drawing

AI summary

System and method detecting synthetic speech may include, using a processor: training a prosody extractor by: providing a training speech sample to an encoder-decoder (codec) model to generate a channel degraded speech sample; providing the channel degraded speech sample to a prosody extractor to extract a prosody embedding; providing to a speech synthesis model the prosody embedding, a codec embedding representing the codec model, speaker identity information and a text representation to generate a spectrogram of the training speech sample; and training the speech synthesis model and the prosody extractor using a loss function defined on the spectrogram generated by the speech synthesis model compared with a spectrogram of the channel degraded speech sample; and using the trained prosody extractor to detect synthetic speech.