Source Separable Audio Encoding for Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In mobile consumer devices, the signal-to-noise ratio of the desired voice signal is severely degraded due to ambient noise and acoustic echo feedback, especially in hand-held modes, where traditional beam forming solutions introduce frequency distortion and are impractical for real-world applications.
Innovation Solution
The technique transforms outputs from multiple microphones into a source-separable audio signal using adaptive filtering between two virtual microphone arrays, reducing ambient noise and echo while preserving useful information, allowing for noise suppression, echo cancellation, and voice enhancement, which can be processed in the cloud or on end devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional beam forming solutions are used to improve signal-to-noise ratio, then noise suppression is achieved, but frequency distortion is introduced
Solution Approach 1:
The patent segments the audio signal processing into two distinct stages: first capturing multiple microphone signals with spatial information, then applying source separation algorithms to decompose the mixture into individual source components. This segmentation allows noise suppression without the frequency-dependent beamforming constraints, resolving the contradiction between SNR improvement and frequency distortion.
Solution Approach 2:
The patent introduces an intermediary source separation processing stage between microphone capture and final audio output. This intermediary layer analyzes the spatial and spectral characteristics of mixed signals from multiple microphones, separates target speech from ambient noise and echo, and reconstructs the cleaned signal, thereby achieving noise suppression without traditional beamforming frequency distortion.
2Reliability
If multiple microphones are used to capture ambient noise for separation, then noise suppression capability is improved, but device complexity increases
Solution Approach 1:
The patent makes the multi-microphone system universal by implementing source separation algorithms that can process signals from any number of microphones configured in different geometries. The same processing framework adapts to various device form factors and microphone arrangements, allowing noise suppression capability improvement without proportionally increasing system complexity.
Solution Approach 2:
The patent changes the processing parameters from traditional beamforming (which requires precise geometric configurations) to source separation parameters that operate on the statistical and spectral properties of the mixed signals. This parameter transformation allows effective noise suppression with simpler microphone arrangements, reducing device complexity while maintaining or improving noise suppression capability.
3Measurement precision
If wide band voice is used for voice driven applications, then application accuracy is improved, but susceptibility to ambient noise and echo increases
Solution Approach 1:
The patent extracts the target voice signal from the mixed audio captured by multiple microphones using source separation algorithms. By identifying and extracting the speech components while removing ambient noise and acoustic echo, the system preserves wide band voice quality for accurate voice-driven applications while eliminating the harmful interference that normally degrades performance.
Data Source
AI summary
A method is provided for encoding multiple microphone signals into a composite source-separable audio (SSA) signal, conducive for transmission over a voice network. The embodiments enable the processing of source separation of the target voice signal from its ambient sound to be performed at any point in the voice communication network, including the internet cloud. A multiplicity of processing is possible over the SSA signal, based on the intended voice application. The level of processing is adapted with the availability of the processing power at the chosen processing node in the network in one embodiment. An apparatus for separating out the target source voice from its ambient sound is also provided. The apparatus includes a directed source separation (DSS) unit, which processes the two virtual microphone signals in the SSA representation, to generate a new SSA signal including the enhanced target voice and the enhanced ambient noise.


