Mixed Codebook Excitation for Low-Bitrate Generic Speech Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Low bit rate speech coding technologies face challenges in efficiently encoding GENERIC class speech signals, which are between VOICED and UNVOICED classes, often resulting in spiky or noisy outputs due to the use of either pulse-like or noise-like codebooks, failing to effectively capture both periodic and noise components.

Innovation Solution

A pulse-noise mixed codebook structure is introduced, combining pulse-like and noise-like codebook entries to generate a mixed codebook vector, which is used to create an encoded audio signal, allowing for better representation of both periodic and noise components in speech signals, thereby improving perceptual quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pulse-like codebook entries are used for encoding GENERIC class speech signals, then periodic components are captured, but the output becomes spiky and perceptual quality deteriorates

Engineering Contradiction:
Improveperiodic component captureVSAvoidspiky output
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent combines pulse-like codebook entries and noise-like codebook entries to form a mixed codebook structure. This merging allows the encoder to simultaneously capture periodic components (from pulse-like entries) and smooth the output (from noise-like entries), resolving the contradiction between periodic component capture and spiky output reduction

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The mixed codebook is constructed as a composite structure containing both pulse-like and noise-like codebook entries. This composite approach enables the system to leverage the strengths of both codebook types: pulse-like entries for periodicity representation and noise-like entries for output smoothing, thereby improving perceptual quality while maintaining periodic component capture

Inventive Principle:
Principle #40Composite materials

2Object-affected harmful factors

If noise-like codebook entries are used for encoding GENERIC class speech signals, then output smoothness is improved, but periodic components are not effectively captured

Engineering Contradiction:
Improveoutput smoothnessVSAvoidperiodic component capture
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The mixed codebook structure merges noise-like codebook entries with pulse-like codebook entries, enabling the system to achieve both output smoothness (from noise-like entries) and effective periodic component capture (from pulse-like entries), thus resolving the contradiction between these two requirements

Inventive Principle:
Principle #5Merging (Combining)

3Object-affected harmful factors

If a mixed codebook structure combining pulse-like and noise-like entries is used, then perceptual quality is improved, but codebook search complexity increases

Engineering Contradiction:
Improveperceptual qualityVSAvoidcodebook search complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The codebook search process is segmented into multiple stages: first identifying candidate entries from pulse-like codebook, then selecting complementary entries from noise-like codebook. This segmentation reduces the overall search complexity compared to an exhaustive search of all possible mixed codebook combinations, while still achieving improved perceptual quality

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3214619B1System and method for mixed codebook excitation for speech coding
Publication Date: 2018.11.14 HUAWEI TECH CO LTD
  • EP3214619B1 patent drawingFigure 1~2
  • EP3214619B1 patent drawingFigure 3~4
  • EP3214619B1 patent drawingFigure 5~6

AI summary

In accordance with an embodiment, a method of encoding an audio/speech signal includes determining a mixed codebook vector based on an incoming audio/speech signal, where the mixed codebook vector includes a sum of a first codebook entry from a first codebook and a second codebook entry from a second codebook. The method further includes generating an encoded audio signal based on the determined mixed codebook vector, and transmitting a coded excitation index of the determined mixed codebook vector.