Source-Dependent Codebook Segmentation for Speech Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech signal coding methods, such as Code Excited Linear Prediction (CELP), suffer from imperfect reconstruction due to speaker-independent modeling of excitation and vocal tract components, leading to reduced quality of the reconstructed signal and inefficient memory/transmission bandwidth usage.

Innovation Solution

The method involves grouping data into frames and classes, transforming frames into filter and source parameter vectors, computing codebooks based on these vectors, and segmenting data into subframes to optimize coding, allowing for source-dependent coding that improves the quality of the reconstructed signal while reducing memory and transmission bandwidth requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If speaker-independent modeling is used for excitation and vocal tract components, then memory occupation is reduced, but the quality of the reconstructed signal deteriorates

Engineering Contradiction:
Improvememory occupationVSAvoidquality of reconstructed signal
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments the speech signal into multiple classes based on speaker characteristics. Instead of using a single speaker-independent codebook, the system creates multiple class-specific codebooks (e.g., male voice, female voice, child voice) that are optimized for each speaker category. This segmentation allows the system to achieve better reconstruction quality for each class while maintaining efficient memory usage through selective codebook usage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by creating codebooks with different properties for different speaker classes. Each class-specific codebook is optimized with appropriate dimensions and characteristics tailored to that speaker category. For example, male voice codebooks may have different vector dimensions compared to female voice codebooks, allowing each to achieve optimal reconstruction quality for its specific target group while maintaining overall system efficiency.

Inventive Principle:
Principle #3Local quality

2Device complexity

If fixed and independent codebooks are used for excitation and vocal tract, then device complexity is reduced, but the representation quality of the speech signal is limited

Engineering Contradiction:
Improvecodebook structureVSAvoidrepresentation quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces dynamics by making the codebook selection adaptive rather than fixed. The system dynamically selects which class-specific codebook to use based on the characteristics of the input speech signal. This dynamic adaptation allows the system to achieve high representation quality by matching the appropriate codebook to the speaker class, while maintaining manageable complexity through automated classification and selection mechanisms.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes parameters by creating codebooks with different dimensions and properties for different speaker classes. Instead of using a single fixed codebook structure, the system varies the codebook parameters (such as vector dimensions, quantization levels, and basis functions) according to the speaker class. This parameter adaptation enables high representation quality for each class while the overall system complexity remains controlled through systematic parameter management.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8447594B2Multicodebook source-dependent coding and decoding
Publication Date: 2013.05.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8447594B2 patent drawing
  • US8447594B2 patent drawing
  • US8447594B2 patent drawing

AI summary

A method for coding data, includes: grouping data into frames; classifying the frames into classes; for each class, transforming the frames belonging to the class into filter parameter vectors, which are extracted from the frames by applying a first mathematical transformation; for each class, computing a filter codebook based on the filter parameter vectors belonging to the class; segmenting each frame into subframes; for each class, transforming the subframes belonging to the class into source parameter vectors, which are extracted from the subframes by applying a second mathematical transformation based on the filter codebook computed for the corresponding class; for each class, computing a source codebook based on the source parameter vectors belonging to the class; and coding the data based on the computed filter and source codebooks.