Speech Attack Detection for Robust CELP Transition Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
CELP-based speech codecs face issues with coding efficiency and perceptual impact due to frame erasures, particularly during transitions like voiced onsets and plosives, leading to desynchronization between encoder and decoder states and poor long-term prediction.
Innovation Solution
Implement a method to detect attacks in sound signals, such as voiced onsets, and force the use of a glottal-shape codebook in CELP-based codecs, especially in frames where attacks are detected, to improve coding efficiency and robustness against frame erasures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If CELP-based speech codecs use adaptive codebook prediction to achieve high coding efficiency, then the bit rate is reduced and subjective quality is improved, but the system becomes sensitive to frame erasures causing desynchronization between encoder and decoder states
Solution Approach 1:
The patent detects attacks (voiced onsets, plosives) in advance before they cause decoder desynchronization. By identifying these critical frames preliminarily and forcing glottal-shape codebook usage, the system prevents the propagation of errors that would otherwise occur with standard adaptive codebook prediction during frame erasures.
Solution Approach 2:
The patent changes the coding parameter selection based on the detected attack type. During detected attacks, the system forces the use of glottal-shape codebook instead of adaptive codebook, thereby changing the excitation model parameter to one that is more robust against frame erasures while maintaining coding efficiency for attack frames.
2Ease of manufacture
If standard frame error concealment techniques use information from the last correctly received frame to conceal lost frames, then the implementation is simple, but the perceptual impact is very annoying during voiced onsets and transitions
Solution Approach 1:
The patent applies different error concealment strategies locally depending on the frame type. For normal frames, standard concealment is used, but for detected attack frames (voiced onsets, plosives, transitions), the system forces glottal-shape codebook usage which is specifically designed to handle these critical transitions, thereby improving local quality where it matters most without complicating the overall system.
Solution Approach 2:
By detecting attacks in advance and preliminarily marking these frames for special handling, the system prepares the appropriate concealment strategy before frame loss occurs. This preliminary identification allows the decoder to apply the correct glottal-shape codebook reconstruction for attack frames, avoiding the perceptual distortion that would result from applying generic concealment methods.
3Duration of action of moving object
If the adaptive codebook is updated using noise-like excitation contribution from unvoiced frames, then the coding process is continuous, but the periodic part of the excitation is completely missing causing long recovery time
Solution Approach 1:
The patent changes the excitation model parameter from adaptive codebook to glottal-shape codebook during detected attack frames. This parameter change ensures that even when the adaptive codebook is updated with noise-like unvoiced excitation, the glottal-shape codebook provides the necessary periodic excitation structure, maintaining reconstruction accuracy while preserving coding continuity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method and device for detecting an attack in a sound signal to be coded wherein the sound signal is processed in successive frames each including a number of sub-frames. The device comprises a first-stage attack detector for detecting the attack in a last sub-frame of a current frame, and a second-stage attack detector for detecting the attack in one of the sub-frames of the current frame, including the sub-frames preceding the last sub-frame. No attack is detected when the current frame is not an active frame previously classified to be coded using a generic coding mode. A method and device for coding an attack in a sound signal are also provided. The coding device comprises the above mentioned attack detecting device and an encoder of the sub-frame comprising the detected attack using a transition coding mode using a glottal-shape codebook populated with glottal impulse shapes.