Speaker Recognition via Multi-Channel Voice Signal Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition methods face challenges in robustly authenticating speakers due to noise interference and limited voice information, often losing voice information during enhancement processes.
Innovation Solution
A method and apparatus that enhance the first voice signal by removing noise through techniques like stationary noise suppression and sound source separation, then generate a multi-channel voice signal by associating it with an enhanced signal, using neural networks to extract and compare feature vectors for accurate speaker recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech enhancement is applied to remove noise from the voice signal, then noise resistance is improved, but voice information loss increases
Solution Approach 1:
The patent segments the voice signal processing into multiple channels: the original voice signal channel and the enhanced voice signal channel. Each channel processes the signal differently, with the enhanced channel focusing on noise removal while the original channel preserves all voice information. This segmentation allows the system to benefit from both noise suppression and information preservation simultaneously.
Solution Approach 2:
The patent merges the original voice signal and the enhanced voice signal into a multi-channel voice signal that is then fed into the neural network. By combining both signals, the system leverages the noise-resistant properties of the enhanced signal while retaining the complete voice information from the original signal, thus resolving the contradiction between noise resistance and information preservation.
2Measurement precision
If multi-channel voice signal is generated by associating original and enhanced signals, then speaker recognition accuracy is improved, but system complexity increases
Solution Approach 1:
The patent introduces an additional dimension to the voice signal processing by creating a multi-channel structure. Instead of processing a single voice signal, the system processes multiple channels (original and enhanced signals) simultaneously, allowing the neural network to leverage both signals for more accurate speaker recognition while maintaining a relatively simple processing architecture.
Data Source
AI summary
A speaker recognition method and apparatus receives a first voice signal of a speaker, generates a second voice signal by enhancing the first voice signal through speech enhancement, generates a multi-channel voice signal by associating the first voice signal with the second voice signal, and recognizes the speaker based on the multi-channel voice signal.


