Speech Production Model for Voice Modification Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice modification detection systems are inefficient due to their reliance on large numbers of training samples and lack of generalization for new modification techniques, and single-class approaches rely on features unrelated to speech physics.
Innovation Solution
A single-class machine learning model using a physical model of speech production, such as the source-filter model, to detect voice modifications by analyzing parameters like pitch, formants, and residuals, generating a voice modification score based on anomalies in these parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-class machine learning approaches are used for voice modification detection, then the system can detect known modification types, but it requires a large number of training samples and fails to generalize to new modification methods
Solution Approach 1:
The patent replaces the conventional multi-class machine learning approach (which relies on statistical patterns from training data) with a physics-based speech production model. This model uses fundamental acoustic principles to analyze voice signals, enabling detection of modifications without requiring large numbers of training samples for each modification type. The physics-based approach inherently generalizes to new modification methods because it relies on universal speech production mechanisms rather than learned patterns from specific examples.
Solution Approach 2:
The patent transforms the detection approach by changing from analyzing raw audio features to analyzing parameters derived from a physics-based speech production model. By modeling speech as generated by physical mechanisms (source-filter model with glottal pulse, vocal tract filter, and radiation filter), the system extracts parameters that reflect the underlying physics of speech production. Modifications to the voice signal create detectable anomalies in these physics-based parameters, enabling generalization across different modification types without retraining.
2Ease of manufacture
If single-class machine learning approaches are used, then training samples of modified speech are not required, but the features used do not have a direct link to the physics of speech production
Solution Approach 1:
The patent replaces conventional single-class machine learning features (which are typically statistical or spectral features without physical meaning) with features derived from a physics-based speech production model. This substitution provides direct physical interpretability because each feature corresponds to a specific aspect of speech production physics, such as glottal pulse characteristics, vocal tract resonance, or acoustic radiation effects.
Solution Approach 2:
The physics-based speech production model serves multiple functions simultaneously: it provides a framework for feature extraction, ensures physical interpretability of those features, and enables detection without requiring modified speech training samples. The universal nature of the physics model allows it to handle both normal and modified speech through the same physical principles, eliminating the need for separate training on modification types.
3Reliability
If conventional voice modification detection systems are used, then they can identify known fraud patterns, but they are inefficient due to requiring large numbers of training samples
Solution Approach 1:
The patent replaces the inefficient conventional approach that requires collecting and processing large numbers of training samples with a physics-based model that can analyze voice signals directly. The physics-based speech production model provides built-in understanding of normal speech characteristics, allowing the system to detect modifications through anomaly detection rather than requiring extensive training data collection and processing.
Data Source
AI summary
A computer may train a single-class machine learning using normal speech recordings. The machine learning model or any other model may estimate the normal range of parameters of a physical speech production model based on the normal speech recordings. For example, the computer may use a source-filter model of speech production, where voiced speech is represented by a pulse train and unvoiced speech by a random noise and a combination of the pulse train and the random noise is passed through an auto-regressive filter that emulates the human vocal tract. The computer leverages the fact that intentional modification of human voice introduces errors to source-filter model or any other physical model of speech production. The computer may identify anomalies in the physical model to generate a voice modification score for an audio signal. The voice modification score may indicate a degree of abnormality of human voice in the audio signal.


