Sound Processing Apparatus for Noise-Resilient Voice Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound processing technologies struggle to effectively separate desired voice signals from noise, particularly in environments with device operation noises and mixed voice sources, leading to reduced speech recognition accuracy in devices like robot cleaners and air conditioners.
Innovation Solution
A sound processing method and apparatus using multi-channel blind source separation based on independent vector analysis, combined with noise removal techniques such as adaptive line enhancement and multi-channel stationary noise reduction, to extract desired voice signals by comparing power values of off-diagonal elements in the adaptive filter, effectively distinguishing voice signals from operation and white noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional sound source separation techniques are used to extract desired speech signal, then speech signal can be separated from some noise, but speech recognition rate significantly lowers in devices with high motor operation noise such as vacuum cleaners and robots
Solution Approach 1:
The patent segments the noise removal process into multiple stages: first removing tonal noise using adaptive line enhancement, then removing stationary noise using multi-channel stationary noise reduction, and finally applying sound source separation. This multi-stage segmentation allows each technique to target specific noise types, improving overall speech recognition rate in high-noise environments.
Solution Approach 2:
The patent changes the approach from conventional single-stage sound source separation to a multi-stage process with different parameter optimization at each stage. The adaptive line enhancement adjusts filter parameters to cancel tonal noise, while the stationary noise reduction adjusts gain parameters based on spectral analysis, enabling effective noise removal across different noise conditions.
2Measurement precision
If white noise and device operation noise are mixed in the input signal, then the number of sound sources exceeds the number of microphones, but this degrades the separation performance
Solution Approach 1:
The patent extracts and removes specific noise components before sound source separation. By using adaptive line enhancement to extract tonal noise and stationary noise reduction to extract background noise, the system reduces the number of active sound sources, making the separation task feasible even when the total number of original sound sources exceeds the number of microphones.
Solution Approach 2:
The patent performs preliminary noise removal actions before the main sound source separation process. By pre-removing tonal and stationary noise components, the system prepares a cleaner input signal for separation, improving the effectiveness of subsequent blind source separation techniques when dealing with multiple sound sources.
3Measurement precision
If conventional sound source separation based on fundamental frequency and harmonics is used, then desired speech can be extracted, but the extraction is adversely affected by degradation of sound source separator in noisy environments
Solution Approach 1:
The patent applies different processing qualities to different frequency regions and noise types. Adaptive line enhancement targets specific tonal frequencies with high-quality filtering, while stationary noise reduction applies spectral analysis with localized gain adjustment. This local quality approach maintains speech extraction accuracy even in degraded noisy environments.
4Measurement precision
If beam forming method is used to separate voices of multiple speakers, then voice separation is achieved, but the method is limited to voice separation and cannot handle device operation noise
Solution Approach 1:
The patent creates a universal noise removal system that handles multiple noise types through different techniques. The adaptive line enhancement handles tonal noise from various sources, the stationary noise reduction handles background noise, and the beam forming handles spatial voice separation. This multi-functional approach makes the system applicable to both voice separation and device operation noise scenarios.
Data Source
AI summary
Disclosed are a sound processing apparatus and a sound processing method. The sound processing method includes extracting a desired voice enhanced signal by a sound source separation and a sound extraction. By using a multi-channel blind source separation method based on independent vector analysis, the desired voice enhanced signal is extracted from a channel having the smallest sum of off-diagonal values of a separation adaptive filter when the power of the desired voice signal is larger than that of other voice signals. According to the present disclosure, a user may build a robust artificial intelligence (AI) speech recognition system by using sound source separation and voice extraction using eMBB, URLLC, and mMTC techniques of 5G mobile communication.


