Higher-Dimensional Audio Projection for Background Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current noise suppression techniques in digital communication systems are inadequate in adapting to diverse environments, failing to effectively separate target speech from background noise and often distort the audio signal.
Innovation Solution
A multi-stage background noise suppression system that employs mask estimation, forward and inverse projection, noise estimation, and post-processing to enhance speech clarity by projecting audio into a higher-dimensional space and applying sophisticated manipulation techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If current noise suppression techniques are used, then some noise reduction is achieved, but speech distortion occurs and adaptability to diverse environments is poor
Solution Approach 1:
The patent projects the input audio signal from the original time-frequency domain into a higher-dimensional latent space using a neural network. This dimensional transformation enables the system to separate speech and noise components more effectively by exploiting additional representational dimensions, thereby reducing speech distortion while suppressing background noise.
Solution Approach 2:
The patent introduces a mask estimator as an intermediary component that operates in the higher-dimensional latent space. This mask estimator generates speech masks that selectively attenuate noise components while preserving speech components, serving as a mediator between the raw audio input and the final denoised output.
2Adaptability or versatility
If current noise suppression techniques are used, then processing speed is maintained, but adaptability to diverse environments is insufficient
Solution Approach 1:
The patent employs a dynamic neural network architecture with multiple stages including forward projection, mask estimation, and inverse projection. The system dynamically adjusts its processing based on the characteristics of the input audio, enabling adaptation to diverse acoustic environments while maintaining manageable complexity through modular design.
Solution Approach 2:
The patent transforms the audio signal into a different parameter space (higher-dimensional latent space) where noise and speech are more separable. By changing the representational parameters of the audio signal rather than processing it in the original domain, the system achieves better environmental adaptability.
3Measurement precision
If sophisticated noise suppression is applied, then speech clarity is improved, but computational complexity increases
Solution Approach 1:
The patent divides the noise suppression task into multiple sequential stages: forward projection to latent space, mask estimation, and inverse projection back to time-frequency domain. This segmentation allows each component to be optimized independently, achieving high speech separation accuracy while managing computational complexity through distributed processing.
Data Source
AI summary
The disclosed technology relates to methods, background noise suppression systems, and non-transitory computer readable media for background noise suppression. In some examples, frames fragmented from input audio data are projected into a higher dimension space than the input audio data. An estimated speech mask is applied to the frames to separate speech components and noise components of the frames. The speech components are then transformed into a feature domain of the input audio data by performing an inverse projection on the speech components to generate output audio data, and also transform the noise components of the frames in the higher dimension space into the lower dimension space using a second inverse projection to learn a difference between speech and noise. The output audio data is provided via an audio interface. The output audio data advantageously comprises a noise-suppressed version of the input audio data.


