Speech Encoding via Inverse Pitch Correlation Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech coding algorithms face challenges in minimizing the bit rate required for transmitting encoded speech signals, particularly in reducing the energy of the residual signal to optimize encoding efficiency.

Innovation Solution

A method and encoder system that add a predetermined noise signal to the input speech signal to generate a simulated signal, determine linear predictive coding coefficients, and form an encoded signal based on these coefficients and the residual signal, with noise shaping techniques to minimize bitrate by optimizing quantization noise distribution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional speech coding algorithms are used, then the encoding process is simpler, but the bit rate required for transmitting encoded speech signals is higher

Engineering Contradiction:
Improveencoding process complexityVSAvoidbit rate
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by adding a predetermined noise signal to the input speech signal before performing linear predictive coding. This preprocessing step generates a simulated signal that, when processed through LPC, produces a residual signal with minimized energy. The noise signal is added in advance to shape the quantization noise distribution, thereby reducing the bit rate required for transmission without significantly increasing encoding complexity.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the residual signal energy is not minimized, then the encoding process is faster, but the bit rate required for transmission is higher

Engineering Contradiction:
Improveencoding speedVSAvoidbit rate
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent changes the parameter of the input signal by adding a predetermined noise signal with specific characteristics (variance equal to the variance of quantization noise) before LPC processing. This parameter modification transforms the signal in such a way that the resulting residual signal has minimized energy. The transformation is computationally efficient and does not significantly slow down the encoding process, thus achieving both fast encoding and reduced bit rate.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If noise shaping techniques are applied, then the bit rate is reduced, but the encoding process becomes more complex

Engineering Contradiction:
Improvebit rateVSAvoidencoding process complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements noise shaping through preliminary action by adding a predetermined noise signal to the input speech signal before LPC analysis. This approach shapes the quantization noise distribution in advance, ensuring that the residual signal contains minimized energy. The technique achieves noise shaping without requiring complex iterative optimization or adaptive filtering, thus reducing bit rate while maintaining relatively simple encoding process complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9530423B2Speech encoding by determining a quantization gain based on inverse of a pitch correlation
Publication Date: 2016.12.27 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9530423B2 patent drawing
  • US9530423B2 patent drawing
  • US9530423B2 patent drawing

AI summary

A method, system and program for encoding and decoding speech according to a source-filter model whereby speech is modelled to comprise a source signal filtered by a time-varying filter. The method comprises: receiving a speech signal comprising successive frames. For each of a plurality of frames of the speech signal: adding a predetermined noise signal generated by a quantization gain multiplied by 0.5 times an inverse of a pitch correlation to the speech signal to generate a simulated signal, determining linear predictive coding coefficients based on the simulated signal frame, and determining a linear predictive coding residual signal based on the linear predictive coding coefficients and one of the speech signal and the simulated signal. Then forming an encoded signal representing said speech signal, based on the linear predictive coding coefficients and the linear predictive coding residual signal.