Adaptive DTX Hangover Framing for Comfort Noise Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio coding standards with DTX schemes face inefficiencies due to hangover periods incorrectly classifying speech offsets as inactive, leading to reduced transmission efficiency and compromised quality of comfort noise synthesis.
Innovation Solution
Adaptive hangover periods are used to determine a set of frames representative of background noise, which are transmitted with a SID frame to enable accurate comfort noise generation at the decoder side, optimizing transmission efficiency and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a fixed hangover period is used to avoid speech offset clipping, then speech quality is improved, but transmission efficiency deteriorates due to incorrect classification of speech offsets as inactivity
Solution Approach 1:
The patent applies dynamics by making the hangover period adaptive rather than fixed. The hangover period is dynamically adjusted based on the decay rate of the speech signal, allowing the system to extend the hangover period when speech is still present (avoiding clipping) and shorten it when speech has truly ended (improving transmission efficiency). This resolves the contradiction by making the hangover period flexible to match actual speech conditions.
Solution Approach 2:
The patent changes the parameter of hangover period duration based on signal characteristics. By monitoring the energy decay rate of the speech signal and adjusting the hangover period parameter accordingly, the system optimizes both speech quality preservation and transmission efficiency. This parameter adaptation allows the system to respond to varying speech conditions rather than using a static hangover period.
2Reliability
If the hangover period is extended to improve SID frame parameter analysis, then comfort noise synthesis quality is improved, but DTX efficiency is compromised due to increased speech frame transmissions
Solution Approach 1:
The patent makes the hangover period dynamic based on actual speech signal decay characteristics. By adapting the hangover period length to the specific speech conditions, the system ensures sufficient frames are available for SID parameter analysis when needed, while avoiding unnecessarily long hangover periods that would reduce DTX efficiency. This dynamic adjustment resolves the contradiction between comfort noise quality and DTX efficiency.
Solution Approach 2:
The system uses the speech frames during the hangover period to automatically generate SID parameters for comfort noise synthesis. By utilizing the available speech frames self-service style to compute their own comfort noise parameters, the system ensures high-quality comfort noise without requiring additional transmission resources, thus maintaining DTX efficiency while improving comfort noise synthesis quality.
3Productivity
If VAD detection is used to distinguish speech from inactivity, then transmission efficiency is improved, but speech offset detection accuracy deteriorates due to background noise masking
Solution Approach 1:
The patent applies preliminary action by extending the hangover period after VAD detects speech offset. This preliminary extension ensures that even if VAD makes an early offset detection (potentially due to background noise masking), the system still captures the decaying speech signal in the hangover period. This allows for accurate SID parameter computation and prevents speech clipping, thereby maintaining measurement precision while preserving transmission efficiency gains from VAD.
Solution Approach 2:
The hangover period acts as a cushion against potential VAD detection errors. By providing an additional time buffer after VAD detects speech offset, the system compensates for possible premature detection caused by background noise. This beforehand cushioning ensures that decaying speech components are not mistakenly classified as inactivity, maintaining both transmission efficiency and detection accuracy.
Data Source
AI summary
Transmitting node and receiving node for audio coding and methods therein. The nodes being operable to encode/decode speech and to apply a discontinuous transmission (DTX) scheme comprising transmission/reception of Silence Insertion Descriptor (SID) frames during speech inactivity. The method in the transmitting node comprising determining, from amongst a number N of hangover frames, a set Y of frames being representative of background noise, and further transmitting the N hangover frames, comprising at least said set Y of frames, to the receiving node. The method further comprises transmitting a first SID frame to the receiving node in association with the transmission of the N hangover frames, where the SID frame comprises information indicating the determined set Y of hangover frames to the receiving node. The method enables the receiving node to generate comfort noise based on the hangover frames most adequate for the purpose.


