Spoofing Detection Neural Network with Regulation Factor

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spoofing detection technologies struggle with distinguishing real and spoofed speech, especially in real-world scenarios, due to their reliance on academic or synthetic data and the need for large audio samples.

Innovation Solution

A computer-based system and method using a neural network with a regulation factor in its loss function to classify voice samples as genuine or spoofed, requiring shorter audio samples and capable of operating effectively in real-world conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If current spoofing detection solutions are developed using academic or synthetic data, then model training can be performed, but performance deteriorates when tested on real world recordings

Engineering Contradiction:
Improveease of model trainingVSAvoiddetection performance on real world recordings
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transforms the training data parameter from academic/synthetic to real-world recordings, fundamentally changing the data distribution and characteristics to match deployment conditions. This parameter change ensures that the model learns from data that reflects actual operational scenarios, thereby maintaining high detection performance on real-world recordings while still enabling effective model training.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional spoofing detection solutions are used, then detection can be performed, but large audio samples are required which increases processing time and resource requirements

Engineering Contradiction:
Improvedetection capabilityVSAvoidaudio sample processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by using only a portion of the audio sample - specifically, it processes audio data in smaller chunks or frames rather than requiring entire long audio recordings. This allows the system to perform effective spoofing detection on shortened audio segments, reducing processing time and resource requirements while maintaining detection reliability through targeted analysis of critical audio features.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If voice generated by current text to speech algorithms is used, then synthetic speech can be created, but it becomes practically indistinguishable from human speech making detection difficult

Engineering Contradiction:
Improvesynthetic speech generation capabilityVSAvoidspoofing detection difficulty
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent implements feedback mechanisms where the model continuously learns from real-world recording examples, adjusting its detection criteria based on actual spoofed and genuine speech patterns encountered in deployment. This feedback loop enables the system to adapt to evolving TTS algorithms and maintain detection effectiveness even as synthetic speech becomes more sophisticated and harder to distinguish from human speech.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12288562B2System and method for spoofing detection
Publication Date: 2025.04.29 NICE LTD
  • US12288562B2 patent drawing
  • US12288562B2 patent drawing
  • US12288562B2 patent drawing

AI summary

A system and method for classification of voice samples to genuine voice samples or spoofing voice samples may include: extracting a set of features from each of a plurality of voice samples, each voice sample labeled as genuine or spoof; training a neural network having a plurality of nodes organized into layers, with links between the nodes, wherein each link comprises a weight, with the sets of features, by adjusting at least one of the weights using a loss function that comprises a regulation factor, wherein the regulation factor is set to zero for voice samples labeled as genuine and is proportional to the prediction of the neural network for data samples labeled as spoofing.