Speech Attack Detection for Robust CELP Transition Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

CELP-based speech codecs face issues with coding efficiency and perceptual impact due to frame erasures, particularly during transitions like voiced onsets and plosives, leading to desynchronization between encoder and decoder states and poor long-term prediction.

Innovation Solution

Implement a method to detect attacks in sound signals, such as voiced onsets, and force the use of a glottal-shape codebook in CELP-based codecs, especially in frames where attacks are detected, to improve coding efficiency and robustness against frame erasures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If CELP-based speech codecs use adaptive codebook prediction to achieve high coding efficiency, then the bit rate is reduced and subjective quality is improved, but the system becomes sensitive to frame erasures causing desynchronization between encoder and decoder states

Engineering Contradiction:
Improvecoding efficiencyVSAvoidrobustness against frame erasures
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent detects attacks (voiced onsets, plosives) in advance before they cause decoder desynchronization. By identifying these critical frames preliminarily and forcing glottal-shape codebook usage, the system prevents the propagation of errors that would otherwise occur with standard adaptive codebook prediction during frame erasures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the coding parameter selection based on the detected attack type. During detected attacks, the system forces the use of glottal-shape codebook instead of adaptive codebook, thereby changing the excitation model parameter to one that is more robust against frame erasures while maintaining coding efficiency for attack frames.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If standard frame error concealment techniques use information from the last correctly received frame to conceal lost frames, then the implementation is simple, but the perceptual impact is very annoying during voiced onsets and transitions

Engineering Contradiction:
Improveimplementation simplicityVSAvoidperceptual distortion during attacks
Core Design Contradiction:
Ease of manufactureVSObject-affected harmful factors

Solution Approach 1:

The patent applies different error concealment strategies locally depending on the frame type. For normal frames, standard concealment is used, but for detected attack frames (voiced onsets, plosives, transitions), the system forces glottal-shape codebook usage which is specifically designed to handle these critical transitions, thereby improving local quality where it matters most without complicating the overall system.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

By detecting attacks in advance and preliminarily marking these frames for special handling, the system prepares the appropriate concealment strategy before frame loss occurs. This preliminary identification allows the decoder to apply the correct glottal-shape codebook reconstruction for attack frames, avoiding the perceptual distortion that would result from applying generic concealment methods.

Inventive Principle:
Principle #10Preliminary action

3Duration of action of moving object

If the adaptive codebook is updated using noise-like excitation contribution from unvoiced frames, then the coding process is continuous, but the periodic part of the excitation is completely missing causing long recovery time

Engineering Contradiction:
Improvecoding continuityVSAvoidexcitation reconstruction accuracy
Core Design Contradiction:
Duration of action of moving objectVSReliability

Solution Approach 1:

The patent changes the excitation model parameter from adaptive codebook to glottal-shape codebook during detected attack frames. This parameter change ensures that even when the adaptive codebook is updated with noise-like unvoiced excitation, the glottal-shape codebook provides the necessary periodic excitation structure, maintaining reconstruction accuracy while preserving coding continuity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3966818B1Methods and devices for detecting an attack in a sound signal to be coded and for coding the detected attack
Publication Date: 2025.11.19 VOICEAGE CORPORATION
  • EP3966818B1 patent drawingFigure 1
  • EP3966818B1 patent drawingFigure 2
  • EP3966818B1 patent drawingFigure 3

AI summary

A method and device for detecting an attack in a sound signal to be coded wherein the sound signal is processed in successive frames each including a number of sub-frames. The device comprises a first-stage attack detector for detecting the attack in a last sub-frame of a current frame, and a second-stage attack detector for detecting the attack in one of the sub-frames of the current frame, including the sub-frames preceding the last sub-frame. No attack is detected when the current frame is not an active frame previously classified to be coded using a generic coding mode. A method and device for coding an attack in a sound signal are also provided. The coding device comprises the above mentioned attack detecting device and an encoder of the sub-frame comprising the detected attack using a transition coding mode using a glottal-shape codebook populated with glottal impulse shapes.