Speaker Recognition Registration with Synthesized Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speaker recognition systems face performance degradation due to mismatches between registration and test environments, particularly when exposed to different noise types and levels, leading to recognition errors.
Innovation Solution
A method that synthesizes a speech signal with a preset noise signal during registration, generating feature vectors that are robust against various noise types, using techniques like convolutional neural networks and domain transformations to improve verification performance across different environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speaker recognition is performed using speech signals recorded in a controlled registration environment, then the recognition accuracy is high under ideal conditions, but the performance degrades significantly when tested in noisy environments
Solution Approach 1:
The patent applies preliminary action by synthesizing noise signals and mixing them with speech signals during the registration phase. This prepares the speaker model in advance to account for various noise conditions that may occur during testing, thereby improving environmental adaptability without sacrificing recognition accuracy. The noise synthesis and signal mixing are performed beforehand to create robust speaker representations.
Solution Approach 2:
The patent changes the parameters of the speech signal by introducing synthesized noise components with varying noise levels and types during registration. This transforms the registration process from using clean speech signals to using noise-contaminated signals, enabling the system to adapt to different acoustic environments while maintaining recognition performance.
2Reliability
If noise signals are synthesized and mixed with speech signals during registration, then robustness against noise is improved, but the complexity of the registration process increases
Solution Approach 1:
The patent introduces an intermediary noise synthesis component that generates artificial noise signals to be mixed with speech signals during registration. This intermediary element enables the system to simulate various acoustic environments without requiring actual recordings from noisy environments, thereby improving noise robustness while keeping the system architecture manageable.
Solution Approach 2:
The patent creates copies of speech signals with added synthesized noise components during registration. Instead of requiring multiple actual recordings from different noisy environments, the system generates synthetic copies that mimic various noise conditions, simplifying the registration process while improving robustness.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method with registration includes: receiving a speech signal of a speaker; synthesizing the received speech signal and a noise signal to generate a synthesized signal; generating a feature vector based on the synthesized signal; and constructing a registration database (DB) corresponding to the speaker based on the generated feature vector.