Dual-Pulse Excitation Model for Speech Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech coding methods, such as Code-Excited Linear Prediction (CELP), face challenges in achieving optimal perceptual quality, reducing computational complexity, and managing memory requirements, particularly at moderate to high bit rates, due to limitations in existing excitation models like random noise, Multi-Pulse, and ACELP.

Innovation Solution

The Dual-Pulse Excitation Model, where two adjacent pulses are used, requiring only one position index to be sent, allowing for varied magnitudes that produce different frequency effects, reducing bit rate and computational complexity by limiting magnitude patterns and enabling local error minimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional excitation models (random noise, Multi-Pulse, ACELP) are used, then speech coding can be implemented, but perceptual quality is insufficient and computational complexity is high at moderate to high bit rates

Engineering Contradiction:
Improveperceptual qualityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The excitation signal is segmented into pairs of adjacent pulses, where each pair is represented by a single position index. This segmentation reduces the number of parameters to be coded and processed, thereby reducing computational complexity while maintaining perceptual quality through the structured pulse pairing approach

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention changes the parameter representation from individual pulse positions and magnitudes to paired pulse positions with limited magnitude patterns. By limiting magnitude patterns and using adjacent pulse pairs, the parameter space is reduced, decreasing computational complexity while preserving speech quality

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional excitation models are used, then speech coding can be implemented, but bit rate requirements are high for optimal performance

Engineering Contradiction:
Improveperceptual qualityVSAvoidbit rate
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The invention extracts and codes only the essential information needed for excitation generation - specifically, the position indices of adjacent pulse pairs and their magnitude patterns. By taking out only the necessary parameters rather than coding all individual pulse characteristics, the bit rate is reduced while maintaining perceptual quality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The paired pulse structure serves multiple functions simultaneously: it provides temporal structure, frequency content through magnitude patterns, and reduces coding requirements. This multi-functionality allows the system to achieve optimal perceptual quality at lower bit rates compared to traditional models

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8175870B2Dual-pulse excited linear prediction for speech coding
Publication Date: 2012.05.08 HUAWEI TECH CO LTD
  • US8175870B2 patent drawing
  • US8175870B2 patent drawing
  • US8175870B2 patent drawing

AI summary

The invention proposed a Dual-Pulse Excitation Model; wherein two pulses of each pair of pulses are always adjacent each other. Only one position index for each pair of pulses needs to be sent to the decoder, which saves bits to code all pulse positions. The magnitudes of each pair of pulses have limited number of patterns. Because the two pulses are adjacent each other, each pair of pulses with different magnitudes can produce different high-pass and/or low-pass effect. Since the magnitudes have enough variation, it is possible to assign the candidate positions of each pair of pulses within a small range in order to save the searching complexity.