Speech Coding via Time-Varying Interpolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech coders, such as MBE vocoders, face challenges in reducing bit rate while maintaining speech quality, particularly in accurately conveying spectral formants and adapting to time-varying signal characteristics at low to medium bit rates.
Innovation Solution
The technique involves estimating log spectral magnitudes at fixed intervals, downsampling them data-dependently, and quantizing and reconstructing them, with omitted magnitudes interpolated to refine error measurement and optimize interpolation points for reduced bit rate without compromising time resolution or formant characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If spectral magnitudes are quantized at fixed intervals in MBE vocoders, then the system operates at medium to low data rates, but the ability to accurately convey spectral formants and adapt to time-varying signal characteristics deteriorates
Solution Approach 1:
The patent applies dynamics by making the quantization interval variable rather than fixed. The quantization interval adapts to the local spectral characteristics and rate of change, allowing finer resolution when spectral formants are present or changing rapidly, and coarser resolution when the signal is stationary. This dynamic adjustment resolves the contradiction by enabling accurate spectral representation only when necessary, maintaining low data rates otherwise.
Solution Approach 2:
The patent changes the parameter of quantization interval from a fixed value to a variable parameter that adjusts based on signal characteristics. By monitoring spectral magnitude changes and adapting the quantization interval accordingly, the system achieves high spectral accuracy when needed while maintaining low overall data rates, thus resolving the contradiction between data rate and spectral formant accuracy.
2Quantity of substance
If the quantization interval is increased to reduce bit rate, then the data rate decreases, but the time resolution and ability to capture rapid speech transitions deteriorates
Solution Approach 1:
The patent uses dynamic adjustment of the quantization interval based on the local rate of change in spectral magnitudes. When rapid speech transitions are detected (high temporal variation), the quantization interval is reduced to preserve time resolution. When the signal is stationary (low temporal variation), the interval is increased to reduce bit rate. This dynamic behavior resolves the contradiction between bit rate and time resolution.
Solution Approach 2:
The patent applies local quality by allowing different quantization intervals for different time regions based on their specific characteristics. Regions with rapid speech transitions receive finer time resolution (smaller intervals), while stationary regions use coarser resolution (larger intervals). This localized adaptation resolves the contradiction by optimizing both bit rate and time resolution in their respective domains.
3Quantity of substance
If fewer bits per frame are used to reduce bit rate, then the data rate decreases, but the ability to accurately represent spectral parameters deteriorates
Solution Approach 1:
The patent changes the parameter of quantization interval from fixed to variable, allowing the system to achieve accurate spectral parameter representation with fewer bits per frame on average. By concentrating bits in regions requiring high precision (rapid transitions, spectral formants) and using coarser quantization elsewhere, the system maintains representation accuracy while reducing overall bit rate.
Solution Approach 2:
The patent applies partial action by providing high-precision spectral parameter representation only where necessary (in regions with rapid changes or important spectral features) rather than uniformly across all frames. This selective precision maintains overall accuracy while significantly reducing the average bits per frame required.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Encoding a sequence of digital speech samples into a bit stream includes dividing the digital speech samples into frames including N subframes (where N is greater than 1); computing subframe model parameters including spectral parameters; and generating a representation of the frame that includes information representing the spectral parameters of P subframes (where P < N) and information identifying the P subframes. The representation excludes information representing the spectral parameters of the N-P subframes not included in the P subframes. Generating the representation includes selecting the P subframes by, for multiple combinations of P subframes, determining an error induced by representing the frame using the spectral parameters for the P subframes and using interpolated spectral parameter values for the N-P subframes. A combination of P subframes is selected based on the determined error for the combination of P subframes.