Neural Network Speaker Recognition Noise Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network devices for speaker recognition face challenges in effectively processing noisy speech signals, as they often require separate networks for noise removal and speaker identification, which can complicate the recognition process and reduce efficiency.
Innovation Solution
A neural network device that combines a skip connection-based neural network for noise removal with another neural network for speaker recognition, generating a third neural network capable of recognizing speakers in noisy speech signals by training on both noise removal and speaker identification information, thereby integrating speech enhancement and recognition functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate neural networks are used for noise removal and speaker identification, then each function can be optimized independently, but the device complexity and processing time increase
Solution Approach 1:
The patent merges the noise removal network and speaker recognition network into a single integrated neural network. The network simultaneously performs denoising and speaker identification functions by sharing common layers and processing pathways, thereby reducing device complexity while maintaining the reliability of noise removal through the integrated architecture.
2Reliability
If separate neural networks are used for noise removal and speaker identification, then each function can be optimized independently, but the processing time increases
Solution Approach 1:
The integrated neural network processes speech signals through a unified architecture where noise removal and speaker recognition operations are performed simultaneously rather than sequentially. This merging eliminates the time required to transfer data between separate networks and reduces overall processing time while maintaining speaker recognition accuracy through the combined processing power.
Solution Approach 2:
The network performs preliminary noise removal operations within the same processing pipeline as speaker recognition, preparing the signal in advance for accurate identification without requiring a separate preprocessing stage. This preliminary action within the integrated structure reduces the total processing time while ensuring recognition accuracy.
3Productivity
If a single integrated network is used for both noise removal and speaker recognition, then processing efficiency improves, but the training complexity increases
Solution Approach 1:
The training process is segmented into distinct phases: first training the network for noise removal using clean and noisy speech pairs, then training the same network for speaker recognition using the learned representations. This segmentation of the training process reduces overall training complexity while maintaining the productivity benefits of the integrated processing architecture.
Data Source
AI summary
Provided are a method of generating a trained third neural network to recognize a speaker of a noisy speech signal by combining a trained first neural network which is a skip connection-based neural network for removing noise from the noisy speech signal with a trained second neural network for recognizing the speaker of a speech signal, and a neural network device for operating the neural networks.


