Voiceprint-Guided Speech Enhancement for Interfering Voice Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech enhancement methods fail to effectively suppress background interfering human voices and ambient noise, resulting in poor speech audio quality in complex environments.
Innovation Solution
A speech enhancement method and apparatus using a voiceprint-guided neural network trained with mixed audio samples, where voiceprint vectors and audio features are extracted and processed to iteratively update network weights until a training condition is satisfied, focusing on enhancing the target speech while suppressing interfering voices and noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional speech enhancement methods (spectral subtraction, Gaussian mixture model, noise reduction neural network) are used, then the processing complexity is relatively low, but the speech enhancement quality is poor and cannot effectively suppress interfering human voices and ambient noise
Solution Approach 1:
The patent segments the speech enhancement task into multiple independent processing streams: voiceprint extraction from reference speech, audio feature extraction from mixed speech, and separate neural network processing paths. This segmentation allows each component to be optimized independently while maintaining overall system manageability, resolving the contradiction between quality and complexity.
Solution Approach 2:
The patent introduces a new dimension by incorporating voiceprint vectors as an additional input feature alongside traditional audio features. This dimensional expansion enables the neural network to leverage speaker-specific characteristics for more accurate speech enhancement, improving quality without proportionally increasing processing complexity through efficient feature fusion.
2Manufacturing precision
If voiceprint-guided neural network with mixed audio samples is used, then speech enhancement quality is improved by effectively suppressing interfering voices and noise, but the computational complexity and training requirements increase
Solution Approach 1:
The patent performs preliminary action by pre-extracting voiceprint vectors from reference speech and pre-processing mixed audio samples into appropriate feature representations before feeding them to the neural network. This preprocessing step reduces the computational burden during actual speech enhancement operations and enables more efficient training by providing well-structured input data from the outset.
Solution Approach 2:
The patent replaces traditional mechanical signal processing methods (spectral subtraction, Gaussian mixture models) with a neural network-based system that learns optimal enhancement strategies from data. This substitution enables the system to automatically adapt to different acoustic environments and speech characteristics, achieving superior quality while the training time investment pays off through faster inference operations.
3Manufacturing precision
If traditional noise reduction methods are used, then the data processing requirements are lower, but the speech audio quality and noise suppression effectiveness are insufficient in complex environments
Solution Approach 1:
The patent extracts and utilizes voiceprint features as a separate, dedicated input channel that specifically targets speaker identification and voice separation. By extracting these distinctive features and processing them through dedicated neural network paths, the system achieves effective suppression of interfering voices and noise while maintaining efficient data utilization through feature-level processing rather than raw audio processing.
Data Source
AI summary
A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.


