Audio Signal Separation via Mask-Based Blind Source Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies using multi-MIC beamforming are sensitive to microphone position errors and increase costs, while blind source separation methods for two MICs struggle to enhance voice signal quality effectively.
Innovation Solution
The method involves acquiring original noisy signals from multiple microphones, performing time-frequency estimation, determining mask values based on these signals, and updating the signals to separate audio sources accurately, reducing voice damage and improving quality without requiring precise microphone positioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-MIC beamforming technology is used to improve voice signal processing quality, then voice recognition rate increases, but the system becomes sensitive to microphone position errors and product cost increases
Solution Approach 1:
The patent replaces the mechanical/physical beamforming approach that relies on precise microphone positioning with a signal processing-based blind source separation method. Instead of using spatial filtering that requires accurate knowledge of microphone positions and sound source directions, the invention uses statistical independence properties of sound sources to separate them in the time-frequency domain, thereby eliminating sensitivity to position errors
Solution Approach 2:
The patent transforms the problem from spatial domain to time-frequency domain by applying short-time Fourier transform. This parameter transformation allows the system to work with spectral characteristics rather than spatial positions, changing the basis of separation from geometric relationships to statistical relationships in the frequency domain, which are invariant to microphone positioning errors
2Reliability
If the number of microphones is increased to improve separation performance, then audio signal quality improves, but product cost increases
Solution Approach 1:
The patent replaces the hardware-based solution of using more microphones with a software-based blind source separation algorithm. The invention demonstrates that with only two microphones, by exploiting the statistical independence of sound sources and using iterative optimization in the time-frequency domain, high-quality separation can be achieved without increasing the number of physical sensors
Solution Approach 2:
The patent creates virtual microphones through signal processing by combining signals from the two physical microphones in different ways across different frequency bins. This allows the system to effectively simulate the response of multiple physical microphones positioned at different locations, achieving the separation performance of a larger array using fewer physical elements
3Ease of manufacture
If blind source separation technology is used with two microphones to reduce hardware cost, then product cost decreases, but voice signal quality enhancement becomes difficult
Solution Approach 1:
The patent segments the audio signal into multiple frequency bins using short-time Fourier transform, and processes each frequency bin separately. This segmentation allows the application of different separation strategies to different frequency regions, improving overall voice quality by addressing the specific characteristics of each frequency band rather than treating the entire spectrum uniformly
Solution Approach 2:
The patent implements an iterative optimization process where the separation results from one iteration are used to improve the separation in the next iteration. The algorithm continuously refines the separation by comparing the estimated sources with the observed mixture and adjusting the separation parameters accordingly, thereby enhancing voice signal quality through progressive improvement rather than a single-pass processing
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for processing audio signal includes that: audio signals emitted respectively from at least two sound sources are acquired through at least two microphones to obtain respective original noisy signals of the at least two microphones; sound source separation is performed on the respective original noisy signals of the at least two microphones to obtain respective time-frequency estimated signals of the at least two sound sources; a mask value of the time-frequency estimated signal of each sound source in the original noisy signal of each microphone is determined based on the respective time-frequency estimated signals; the respective time-frequency estimated signals of the at least two sound sources are updated based on the respective original noisy signals of the at least two microphones and the mask values; and the audio signals emitted respectively from the at least two sound sources are determined.