Audio Signal Processing Using Phase Reconstruction and Complex Masks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio signal processing methods, particularly in speech enhancement, face limitations due to the reliance on noisy phase information, which can lead to sub-optimal magnitude estimation and over-suppression of noise, hindering the quality of the reconstructed signal.
Innovation Solution
The approach involves estimating and refining the phase of the target audio signal using phase reconstruction methods and deep neural networks, allowing for the use of mask values greater than one to improve magnitude estimation, and formulating the phase estimation problem in terms of phase differences or related quantities to enhance the accuracy of the phase estimation process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional noise cancellation methods use noisy phase information for signal reconstruction, then the processing is simpler, but the magnitude estimation becomes sub-optimal and noise is over-suppressed
Solution Approach 1:
The patent segments the complex spectrogram into separate magnitude and phase components for independent processing. The magnitude spectrogram is estimated separately from the noisy magnitude, and the phase is estimated separately from the noisy phase, allowing each component to be optimized independently without the constraints of conventional joint processing
Solution Approach 2:
The patent changes the parameter representation by using complex mask values (with both magnitude and phase components) instead of conventional real-valued masks. This allows the mask to simultaneously adjust both magnitude and phase, enabling mask values greater than one and avoiding over-suppression while maintaining processing feasibility
2Ease of operation
If conventional methods use noisy phase for reconstruction, then the reconstruction process is straightforward, but the reconstructed signal quality deteriorates
Solution Approach 1:
The patent performs preliminary estimation of both magnitude and phase spectrograms before reconstruction. The magnitude spectrogram is estimated from the noisy magnitude, and the phase is estimated from the noisy phase, with both estimations refined before being combined for final signal reconstruction, ensuring high quality inputs to the reconstruction process
Solution Approach 2:
The patent introduces complex mask values as intermediaries that mediate between the noisy spectrogram and the clean signal reconstruction. These masks with magnitude greater than one act as corrective factors that compensate for the inconsistencies in noisy phase information, improving reconstruction quality without complicating the overall process
3Stability of the object's composition
If mask values are constrained to be less than or equal to one, then the processing is more stable, but the magnitude reconstruction accuracy is limited
Solution Approach 1:
The patent changes the constraint parameter for mask values from [0,1] to allow values greater than one. This parameter change enables the mask to fully compensate for magnitude attenuation in the noisy spectrogram, achieving accurate magnitude reconstruction while maintaining stability through the structured complex mask formulation
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
Systems and methods for audio signal processing including an input interface to receive a noisy audio signal including a mixture of target audio signal and noise. An encoder to map each time-frequency bin of the noisy audio signal to one or more phase-related value from one or more phase quantization codebook of phase-related values indicative of the phase of the target signal. Calculate, for each time-frequency bin of the noisy audio signal, a magnitude ratio value indicative of a ratio of a magnitude of the target audio signal to a magnitude of the noisy audio signal. A filter to cancel the noise from the noisy audio signal based on the phase-related values and the magnitude ratio values to produce an enhanced audio signal. An output interface to output the enhanced audio signal.