CELP Speech Decoder Noise Masking Quantization Distortion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech coding schemes based on models like CELP struggle to reproduce natural sound when input signals contain background noise, leading to perceivable quantization distortion.

Innovation Solution

A decoding method that includes a controlling part in the encoding apparatus to generate control information for distinguishing noise-superimposed speech, and a noise appending part in the decoding apparatus to add noise to higher frequency bands, masking quantization distortion and enhancing sound quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If CELP-based speech coding is used to reduce information amount, then encoding efficiency is improved, but quantization distortion becomes perceivable in noise-superimposed speech

Engineering Contradiction:
Improveinformation amountVSAvoidsound quality
Core Design Contradiction:
Loss of informationVSManufacturing precision

Solution Approach 1:

The patent converts the harmful quantization distortion into a beneficial effect by adding artificial noise to mask the distortion. The noise appending part generates noise signals and adds them to the decoded speech signal, transforming the unwanted quantization artifacts into acceptable sound quality by exploiting the masking effect where noise masks the quantization distortion in noise-superimposed speech regions

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent applies different processing to different regions of the speech signal. The noise-superimposed speech determining part identifies specific frames containing background noise, and the noise appending part selectively adds noise only to those identified regions. This local application ensures that quantization distortion is masked where it matters most (in noise-superimposed regions) while maintaining coding efficiency elsewhere

Inventive Principle:
Principle #3Local quality

2Measurement precision

If speech is separated into periodic and non-periodic components for CELP encoding, then encoding precision is improved, but natural sound reproduction deteriorates in noise-superimposed speech

Engineering Contradiction:
Improveencoding precisionVSAvoidnatural sound reproduction
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces noise as an intermediary element between the decoded speech signal and the final output. The noise appending part generates noise signals that act as a mediator to mask the quantization distortion caused by the CELP encoding process. This intermediary noise component bridges the gap between the precision of source separation encoding and the naturalness of sound reproduction in noisy environments

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method effectively reduces the perception of uncomfortable sounds in noise-superimposed speech, producing a more natural sound by masking quantization distortion in CELP-based speech coding schemes.

Implementation Method 1

adding noise to higher frequency bands, masking quantization distortion and enhancing sound quality

Methodology Applied
Scientific EffectMasking effect:

Data Source

PatentEP2869299B1Decoding method, decoding apparatus, program, and recording medium therefor
Publication Date: 2021.07.21 NIPPON TELEGRAPH & TELEPHONE CORP
  • EP2869299B1 patent drawingFigure 1
  • EP2869299B1 patent drawingFigure 2
  • EP2869299B1 patent drawingFigure 3

AI summary

In a speech coding scheme based on a speech production model, such as a CELP-based scheme, an object of the present invention is to provide a decoding method that can reproduce natural sound even if the input signal is a noise-superimposed speech. The decoding method includes a speech decoding step of obtaining a decoded speech signal from an input code, a noise generating step of generating a noise signal that is a random signal, and a noise adding step of outputting a noise-added signal, the noise-added signal being obtained by summing the decoded speech signal and a signal obtained by performing, on the noise signal, a signal processing that is based on at least one of a power corresponding to a decoded speech signal for a previous frame and a spectrum envelope corresponding to the decoded speech signal for the current frame.