Speech Encoding Spectral Slope Control via Tilt Compensation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech encoding methods fail to adaptively adjust the spectral slope of quantization noise while minimizing formant weighting and perform suitable perceptual weighting filtering during noise-speech superposition periods.
Innovation Solution
A speech encoding apparatus and method that include a linear prediction analysis section, a quantizing section, a perceptual weighting section with a tilt compensation coefficient for adjusting the spectral slope of quantization noise, and an excitation search section, which uses a signal-to-noise ratio in a first frequency band to control the tilt compensation coefficient.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If formant weighting coefficient γ2 is changed based on general feature of input signal spectrum, then spectral slope adjustment is achieved, but detailed changes in spectrum cannot be responded to and formant sharpness cannot be adjusted separately
Solution Approach 1:
The patent divides the spectral adjustment function into two separate coefficients: formant weighting coefficient γ2 for formant sharpness control and tilt compensation coefficient γ3 for spectral slope control. This segmentation allows independent adjustment of formant characteristics and overall spectral tilt, resolving the contradiction between general spectral adjustment and detailed spectral response.
Solution Approach 2:
The patent introduces adaptive control mechanisms that dynamically adjust γ2 and γ3 based on different signal conditions (speech periods vs. noise periods). The coefficients are no longer fixed but vary according to the spectral characteristics and noise levels, enabling precise control over both formant sharpness and spectral slope under different operating conditions.
2Reliability
If perceptual weighting filter characteristics are switched between background noise period and speech period, then suitable filtering for each period is achieved, but noise-speech superposition period cannot be handled
Solution Approach 1:
The patent transitions from discrete switching between speech and noise periods to continuous adaptive control that handles all intermediate states including noise-speech superposition. The coefficients γ2 and γ3 are dynamically adjusted based on the actual spectral characteristics and noise levels, allowing smooth transitions and appropriate filtering for mixed conditions rather than forcing discrete category selections.
Solution Approach 2:
The patent modifies the perceptual weighting filter parameters (γ2 and γ3) continuously based on measured signal characteristics including noise levels and spectral slope. This parameter adaptation enables the filter to respond appropriately to noise-speech superposition conditions by adjusting its characteristics according to the actual mixture ratio and spectral properties, rather than using fixed switching thresholds.
3Object-generated harmful factors
If formant weighting coefficient γ2 is used for spectral slope adjustment, then quantization noise shaping is achieved, but formant weighting and spectral slope adjustment cannot be controlled separately
Solution Approach 1:
The patent separates the single coefficient γ2 into two distinct coefficients: γ2 for formant weighting and γ3 for tilt compensation (spectral slope adjustment). This segmentation eliminates the coupling between formant sharpness control and spectral slope control, allowing independent optimization of each function without interfering with the other, thus reducing control complexity while maintaining noise shaping capability.
Data Source
AI summary
Disclosed is an audio encoding device capable of adjusting a spectrum inclination of a quantized noise without changing the Formant weight. The device includes: an HPF (131) which extracts a high-frequency component of the frequency region from an input audio signal; a high-frequency energy level calculation unit (132) which calculates an energy level of the high-frequency component in a frame unit; an LPF (133) which extracts a low-frequency component of the frequency region from the input audio signal; a low-energy level calculation unit (134) which calculates an energy level of a low-frequency component in a frame unit; an inclination correction coefficient calculation unit (141) multiplies the difference between SNR of the high-frequency component and SNR of the low-frequency component inputted from an adder (140) by a constant and adds a bias component to the product so as to calculate an inclination correction coefficient ?3. The inclination correction coefficient is used for adjusting the spectrum inclination of a quantized noise.


