Hybrid Speech Enhancement Decoder with Dynamic Blend Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech enhancement methods in audio programs, such as waveform-coded and parametric-coded enhancements, face challenges in providing consistent and high-quality speech audibility, especially for listeners with hearing impairments, due to bandwidth limitations and audible artifacts.
Innovation Solution
A hybrid speech enhancement method that dynamically blends waveform-coded and parametric-coded enhancements based on signal conditions, using a blend indicator to combine low-quality speech data and parametrically reconstructed speech, optimizing the speech enhancement process to minimize audible artifacts and maintain high quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If waveform-coded enhancement is used to increase speech audibility, then speech intelligibility improves, but coding artifacts become audible and quality deteriorates
Solution Approach 1:
The patent combines waveform-coded enhancement and parametric-coded enhancement into a hybrid system. The decoder selectively applies waveform-coded enhancement when speech is dominant (to maintain intelligibility) and parametric-coded enhancement when non-speech content is present (to avoid audible artifacts), merging the advantages of both approaches.
Solution Approach 2:
The system dynamically switches between waveform-coded and parametric-coded enhancement modes based on real-time analysis of the audio signal characteristics. The decoder determines the appropriate enhancement type by analyzing the dominance of speech versus non-speech content, making the enhancement approach adaptive rather than static.
2Reliability
If two independent audio streams are transmitted for speech enhancement, then speech audibility control improves, but bandwidth consumption doubles
Solution Approach 1:
The hybrid enhancement system serves multiple functions within a single audio stream transmission. It provides both speech enhancement and bandwidth efficiency by using parametric coding for non-speech portions and waveform coding only when necessary, making the system universally applicable to mixed audio content.
Solution Approach 2:
The system changes the coding parameters dynamically based on content type. For non-speech content, it uses parametric coding with low bandwidth; for speech content, it switches to waveform coding with higher bandwidth. This parameter change allows the system to maintain speech audibility control while minimizing overall bandwidth consumption.
3Reliability
If speech enhancement is applied consistently, then speech intelligibility improves, but artifacts become audible in low-background conditions
Solution Approach 1:
The enhancement approach is made dynamic rather than consistent. The decoder analyzes the audio signal to determine whether speech or non-speech content is dominant at any given moment, and selectively applies waveform-coded or parametric-coded enhancement accordingly, preventing artifacts from becoming audible in low-background conditions.
Solution Approach 2:
Different enhancement qualities are applied to different portions of the audio signal based on local characteristics. Waveform-coded enhancement with higher quality is applied locally when speech is dominant, while parametric-coded enhancement with lower artifact visibility is applied locally when non-speech content is present, optimizing overall quality.
Data Source
AI summary
A method for hybrid speech enhancement which employs parametric-coded enhancement (or blend of parametric-coded and waveform-coded enhancement) under some signal conditions and waveform-coded enhancement (or a different blend of parametric-coded and waveform-coded enhancement) under other signal conditions. Other aspects are methods for generating a bitstream indicative of an audio program including speech and other content, such that hybrid speech enhancement can be performed on the program, a decoder including a buffer which stores at least one segment of an encoded audio bitstream generated by any embodiment of the inventive method, and a system or device (e.g., an encoder or decoder) configured (e.g., programmed) to perform any embodiment of the inventive method. At least some of speech enhancement operations are performed by a recipient audio decoder with Mid/Side speech enhancement metadata generated by an upstream audio encoder.


