Audio Signal Separation Using Feature Decomposition Matrices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Intelligent product devices using two microphones face challenges in voice signal processing due to sensitivity to microphone position errors and increased costs with more microphones, necessitating effective blind source separation techniques for high voice quality.
Innovation Solution
A method involving at least two microphones to acquire audio signals, convert them into frequency-domain estimation signals, perform feature decomposition, and obtain separation matrices to separate audio signals from sound sources without considering microphone positions, thereby improving signal quality and reducing hardware costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multi-microphone-based beamforming technology is used, then voice signal processing quality is improved, but sensitivity to microphone position errors increases and product cost increases
Solution Approach 1:
The patent segments the voice enhancement problem into two independent stages: beamforming for spatial filtering and blind source separation for signal decomposition. This segmentation allows the system to achieve high signal quality without requiring precise microphone positioning, as the blind source separation stage compensates for position errors through statistical signal analysis.
Solution Approach 2:
The patent replaces the mechanical positioning requirement of traditional beamforming with a computational approach using blind source separation. Instead of relying on precise physical microphone arrangement, the system uses algorithmic signal processing to separate sources, substituting mechanical precision with computational intelligence.
2Measurement precision
If multi-microphone-based beamforming technology is used, then voice signal processing quality is improved, but product cost increases
Solution Approach 1:
The patent changes the fundamental parameter from using many microphones to using few microphones with advanced signal processing. By transforming the problem from a hardware-intensive approach to a software-intensive approach, the system achieves comparable or superior performance with reduced hardware quantity.
Solution Approach 2:
The patent substitutes additional hardware (more microphones) with computational methods (blind source separation), replacing the mechanical solution of adding more sensors with an intelligent software-based solution that processes signals from fewer sensors.
3Quantity of substance
If blind source separation technology is used with two microphones, then product cost is reduced, but signal separation quality needs improvement
Solution Approach 1:
The patent applies preliminary beamforming processing to the microphone signals before feeding them into the blind source separation algorithm. This preliminary action of spatial filtering enhances the input quality for the separation stage, enabling high-quality separation even with only two microphones.
Solution Approach 2:
The patent introduces beamforming as an intermediary processing stage between microphone acquisition and blind source separation. This intermediary step pre-processes the signals to improve their quality and structure, making the subsequent separation task easier and more effective with limited microphone inputs.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
A method for processing an audio signal is provided. In the method, audio signals sent by at least two sound sources are acquired by at least two microphones to obtain multiple frames of original noisy signals of each microphone on a time domain (S11). For each frame, frequency-domain estimation signals of each sound source are acquired according to the original noisy signals (S12). For each sound source, the frequency-domain estimation signals are divided into multiple frequency-domain estimation components on a frequency domain (S13). For each sound source, feature decomposition is performed on a related matrix of each frequency-domain estimation component to obtain a target feature vector (S14). A separation matrix of each frequency point is obtained based on target feature vectors and the frequency-domain estimation signals (S15). The audio signals of sounds are obtained based on the separation matrixes and the original noisy signals (S16).