Spoofing Detection Neural Network with Regulation Factor
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spoofing detection technologies struggle with distinguishing real and spoofed speech, especially in real-world scenarios, due to their reliance on academic or synthetic data and the need for large audio samples.
Innovation Solution
A computer-based system and method using a neural network with a regulation factor in its loss function to classify voice samples as genuine or spoofed, requiring shorter audio samples and capable of operating effectively in real-world conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If current spoofing detection solutions are developed using academic or synthetic data, then model training can be performed, but performance deteriorates when tested on real world recordings
Solution Approach 1:
The patent transforms the training data parameter from academic/synthetic to real-world recordings, fundamentally changing the data distribution and characteristics to match deployment conditions. This parameter change ensures that the model learns from data that reflects actual operational scenarios, thereby maintaining high detection performance on real-world recordings while still enabling effective model training.
2Reliability
If traditional spoofing detection solutions are used, then detection can be performed, but large audio samples are required which increases processing time and resource requirements
Solution Approach 1:
The patent applies partial action by using only a portion of the audio sample - specifically, it processes audio data in smaller chunks or frames rather than requiring entire long audio recordings. This allows the system to perform effective spoofing detection on shortened audio segments, reducing processing time and resource requirements while maintaining detection reliability through targeted analysis of critical audio features.
3Productivity
If voice generated by current text to speech algorithms is used, then synthetic speech can be created, but it becomes practically indistinguishable from human speech making detection difficult
Solution Approach 1:
The patent implements feedback mechanisms where the model continuously learns from real-world recording examples, adjusting its detection criteria based on actual spoofed and genuine speech patterns encountered in deployment. This feedback loop enables the system to adapt to evolving TTS algorithms and maintain detection effectiveness even as synthetic speech becomes more sophisticated and harder to distinguish from human speech.
Data Source
AI summary
A system and method for classification of voice samples to genuine voice samples or spoofing voice samples may include: extracting a set of features from each of a plurality of voice samples, each voice sample labeled as genuine or spoof; training a neural network having a plurality of nodes organized into layers, with links between the nodes, wherein each link comprises a weight, with the sets of features, by adjusting at least one of the weights using a loss function that comprises a regulation factor, wherein the regulation factor is set to zero for voice samples labeled as genuine and is proportional to the prediction of the neural network for data samples labeled as spoofing.


