Speech Signal Compression Using Non-Uniform Dynamic Range Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hearing aids with dynamic range compression often result in distortion and reduced dynamic range, particularly for speech signals, due to fast-acting compression mechanisms that prioritize reaction speed over signal integrity, leading to preferences for more dynamics and potential loss of audibility in low input levels.
Innovation Solution
A method that calculates a non-uniform compression ratio function based on statistical parameters of the acoustic input signal, optimizing the dynamic range by deviating from a prescribed constant compression ratio to minimize distortion and enhance audio characteristics, particularly for speech signals, by considering frequency ranges, sound pressure levels, and other signal characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fast-acting dynamic compression is applied to control loudness and amplify soft sounds, then hearing threshold elevation and loudness control are improved, but signal distortion increases and dynamic range is reduced
Solution Approach 1:
The patent implements dynamic compression with time-varying parameters including attack time (2-10 ms), release time (20-150 ms), and threshold adjustments that adapt to signal characteristics. The compression ratio and threshold are modified dynamically based on signal level and temporal fine structure detection, allowing the system to maintain reliability in threshold elevation while preserving signal fidelity through adaptive parameter control.
Solution Approach 2:
The system changes physical parameters of the compression process including compression ratio (CR), threshold levels, attack/release times, and knee parameters based on signal characteristics. By detecting temporal fine structure and adjusting parameters accordingly, the system optimizes the balance between loudness control reliability and signal fidelity, applying different compression settings for different signal types and levels.
2Reliability
If compression ratio is increased to map large input dynamic range to narrow output dynamic range, then loudness recruitment compensation is improved, but distortion of speech signals increases
Solution Approach 1:
The patent applies different compression ratios to different frequency bands and signal levels. By analyzing temporal fine structure and signal characteristics, the system applies higher compression ratios where needed for loudness recruitment compensation while maintaining lower compression ratios in frequency bands and signal levels where speech fidelity is critical, thus resolving the contradiction locally rather than globally.
Solution Approach 2:
The compression ratio is made dynamic and adaptive rather than fixed. The system adjusts compression ratio in real-time based on detected signal characteristics, temporal fine structure presence, and signal level, allowing optimal compensation for loudness recruitment while preserving speech fidelity when temporal fine structure is detected.
3Speed
If attack time is shortened to respond quickly to sudden sounds, then reaction speed is improved, but distortion of impulse sounds increases
Solution Approach 1:
The attack time parameter is made dynamic and adaptive. The system uses short attack times (2-10 ms) for general operation to maintain fast response, but extends attack time for detected impulse sounds and transient signals. By detecting signal type and temporal characteristics, the system optimizes attack time to balance reaction speed with impulse sound fidelity.
4Ease of operation
If release time is extended to smooth transitions, then perceptual comfort is improved, but audibility of low input levels is reduced
Solution Approach 1:
The release time parameter is made dynamic and signal-level dependent. The system uses longer release times (20-150 ms) for high-level signals to provide perceptual comfort and smooth transitions, but uses shorter release times for low-level signals to maintain audibility. This adaptive approach resolves the contradiction by optimizing release time based on current signal conditions.
Data Source
AI summary
The invention relates to a method for processing an acoustic input signal, preferably a speech signal, said method comprising the following steps: a) receiving a digital representation (Sin) of an acoustic input signal, b) calculating at least one statistical parameter (P) of the digital representation (Sin) of the acoustic input signal, c) calculating a compression ratio function (CRf) based—on a prescribed constant compression ratio (CRpr), said prescribed constant compression ratio (CRpr) uniformly mapping acoustic input signals of a selected magnitude to acoustic output signals of a selected magnitude, and—on at least one statistical parameter (P) calculated in step b), and d) applying the non-uniform compression ratio function (CRf) according to step c) on the digital representation (Sin) of the acoustic input signal delivering a digital representation (Sout) of an enhanced acoustic output signal.


