Speech Detection System for Synthesized Voice Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Sophisticated speech synthesis algorithms, such as Tacotron, can generate artificially faked voices that are indistinguishable from real human voices, posing a threat for social engineering attacks, and existing technologies are inadequate in effectively detecting synthesized speech.

Innovation Solution

A speech-based system integrated with a classification algorithm trained using supervised machine learning to differentiate between synthetic and natural speech signals by dividing incoming speech into portions and evaluating index values, with a threshold value to determine and alert users to potentially fake voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech synthesis algorithms are used to generate fake voices, then voice quality and realism are improved, but security against social engineering attacks deteriorates

Engineering Contradiction:
Improvevoice qualityVSAvoidsecurity threat
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a detection system as an intermediary between the speech synthesis output and the end user. This intermediary analyzes acoustic features of the speech signal and compares them against learned patterns of synthesized versus natural speech, thereby mediating the security risk while preserving the ability to use high-quality synthesis for legitimate purposes

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary anti-action by training the detection system in advance on labeled datasets of synthesized and natural speech. This pre-training enables the system to proactively identify and block potential social engineering attacks before they can harm the target, countering the security threat posed by advanced synthesis algorithms

Inventive Principle:
Principle #9Preliminary anti-action

2Reliability

If existing detection methods are used, then some level of synthesized speech detection is achieved, but detection accuracy deteriorates due to indistinguishable quality

Engineering Contradiction:
Improvedetection capabilityVSAvoiddetection accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters used for detection by analyzing multiple acoustic features simultaneously (spectral characteristics, temporal patterns, prosodic features) rather than relying on single traditional methods. This multi-parameter approach enables accurate distinction between synthesized and natural speech even when overall quality is comparable

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent moves detection into a new dimensional space by using machine learning models that operate in high-dimensional feature spaces. This allows the system to detect subtle patterns and relationships across multiple acoustic dimensions that are imperceptible in traditional single-dimensional analysis, thereby achieving high detection accuracy

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If sophisticated speech synthesis is deployed, then voice realism is improved, but ease of operation for malicious purposes is improved

Engineering Contradiction:
Improvevoice realismVSAvoidease of malicious use
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The detection system serves as an automated intermediary that blocks malicious operations before they can succeed. By analyzing speech in real-time and comparing against trained models, the system makes it operationally difficult for attackers to use synthesized voices for social engineering, even though the synthesis technology itself remains easy to operate for legitimate purposes

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3933832B1A method and a speech-based system for automatically detecting synthesized or natural speech in a speech-based system and a computer program
Publication Date: 2024.11.06 DEUTSCHE TELEKOM AG
  • EP3933832B1 patent drawingFigure 1
  • EP3933832B1 patent drawingFigure 2
  • EP3933832B1 patent drawingFigure 3

AI summary

The present invention is directed inter alia to a method and a speech-based system (120) for automatically detecting synthesized or natural speech. The system (120) comprises a computational system (180) including a memory (181), which may store a trained classification algorithm, a communication interface (183) configured to receive speech signals addressed to the at least one communication apparatus, a data processing unit (182) configured to i) divide up the speech signals received at the communication interface (183) into portions, each portion having a predetermined duration, ii) to evaluate an index value by executing the classification algorithm stored in the memory (181) on at least one of the portions, and iii) to determine in dependence of the first and second predefined index value and in dependence of the index value evaluated, whether the at least one portion belongs to a synthesized speech signal or to a natural speech signal.