Target SNR Audio Processing With Low-Latency Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing devices struggle to maintain a consistent signal-to-noise ratio (SNR) in noisy environments without introducing significant latency, leading to reduced situational awareness and user dissatisfaction.
Innovation Solution
A device using low-latency time-domain processing with guided filter coefficients from frequency-domain processing to separate target and non-target audio components, adjusting gains separately to achieve a target SNR, combining signals with minimal latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If simple noise suppression processes are applied to remove as much ambient noise as possible, then the signal-to-noise ratio is improved sufficiently for speech intelligibility, but situational awareness is reduced since important environmental cues such as traffic sounds may be removed
Solution Approach 1:
The patent applies different processing qualities to different frequency components of the audio signal. Specifically, it identifies and preserves certain frequency ranges that correspond to important environmental cues while applying stronger noise suppression to other frequency ranges where speech and target sounds are located. This selective approach allows the system to improve speech intelligibility through noise suppression while maintaining situational awareness by preserving environmentally important frequencies.
2Measurement precision
If more complex noise suppression processes are applied to improve speech intelligibility, then the signal-to-noise ratio is enhanced, but significant latency is introduced which leads to user dissatisfaction
Solution Approach 1:
The patent segments the audio processing into distinct frequency-based components, allowing different processing strategies to be applied to different segments. By dividing the audio spectrum into multiple frequency bands and applying tailored noise suppression to each, the system achieves effective speech enhancement without requiring a single complex processing pipeline that would introduce significant latency. This segmented approach enables parallel processing of different frequency components.
Solution Approach 2:
The patent applies noise suppression selectively rather than uniformly across all frequencies. It applies stronger suppression only where necessary for speech intelligibility while using milder or no suppression for environmental cues. This partial action approach achieves the necessary speech enhancement without the excessive processing complexity that would cause latency, thereby maintaining real-time performance.
3Measurement precision
If aggressive noise removal is applied to achieve high speech intelligibility, then the target signal clarity is improved, but the naturalness of the audio output deteriorates due to loss of environmental context
Solution Approach 1:
The patent preserves the natural composition of the audio by applying different processing qualities to different frequency components. Important environmental frequencies are processed with minimal intervention to maintain their natural characteristics, while speech frequencies receive stronger noise suppression for clarity. This local differentiation allows the output to maintain overall naturalness while achieving target signal clarity where needed.
Data Source
AI summary
A device includes one or more processors configured to obtain data specifying a target signal-to-noise ratio based on a hearing condition of a person and to obtain audio data representing one or more audio signals. The one or more processors are configured to determine, based on the target signal-to-noise ratio, a first gain to apply to first components of the audio data and a second gain to apply to second components of the audio data. The one or more processors are configured to apply the first gain to the first components of the audio data to generate a target signal and to apply the second gain to the second components of the audio data to generate a noise signal. The one or more processors are further configured to combine the target signal and the noise signal to generate an output audio signal.


