Wideband Speech Encoder Using Highband Spectral Reversal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current wideband speech coders are computationally intensive and require significant bandwidth, making them impractical for mobile and embedded applications, especially when encoding the entire wideband signal to desired quality, and transcoding is necessary for systems that only support narrowband coding.
Innovation Solution
A method and apparatus that divide a wideband speech signal into lowband and highband signals, with spectral reversal and downmixing operations applied to the highband signal before decimation, reducing system complexity and allowing for flexible coding of different highband configurations by maintaining a sample rate of 32 kHz or below.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a wideband CELP coder is used to encode the entire wideband signal to desired quality, then speech quality and intelligibility are improved, but computational complexity and processing cycles increase to unacceptable levels
Solution Approach 1:
The wideband speech signal is divided into two separate sub-bands: narrowband (0-4kHz) and highband (4-8kHz). Each sub-band is independently encoded using separate coders (narrowband CELP coder and highband coder), allowing the system to achieve wideband quality without the prohibitive computational complexity of encoding the entire wideband signal with a single CELP coder.
2Measurement precision
If a wideband CELP coder is used to encode the entire wideband signal, then speech quality is improved, but bandwidth consumption increases to unacceptable levels
Solution Approach 1:
The wideband signal is segmented into narrowband and highband components that are encoded separately. The narrowband portion uses efficient CELP coding at lower bitrates, while the highband is encoded with a dedicated highband coder. This segmentation allows the system to achieve wideband quality with significantly reduced total bandwidth compared to encoding the entire wideband signal with full CELP.
3Device complexity
If the entire wideband signal is encoded using narrowband coding techniques, then processing complexity is reduced, but transcoding is required for systems supporting only narrowband coding
Solution Approach 1:
By segmenting the wideband signal into narrowband and highband components encoded separately, the system maintains compatibility with narrowband-only systems. The narrowband encoded portion can be transmitted and decoded by traditional narrowband systems without requiring transcoding, while the highband portion can be optionally transmitted for enhanced quality when system capabilities permit.
4Device complexity
If highband signal processing is performed without spectral reversal and downmixing, then processing steps are reduced, but system flexibility in coding different highband configurations is limited
Solution Approach 1:
Spectral reversal inverts the frequency spectrum of the highband signal, allowing flexible mapping of highband content to different target configurations. Downmixing then combines the reversed spectrum with the narrowband signal. This inversion technique enables the system to adaptively code different highband configurations (such as different bandwidth extensions or quality levels) by manipulating the reversed spectrum, providing greater flexibility than direct processing would allow.
Data Source
Figure 1~3
Figure 4
Figure 5
AI summary
A method and apparatus for encoding a signal is provided herein. During operation a wideband signal that is to be encoded enters a filter bank. A highband signal and a lowband signal are output from the filter bank. Each signal is separately encoded. During the production of the highband signal, a downmixing operation is implemented after preprocessing, and prior to decimating. The downmixing operation greatly reduces system complexity. In fact, it will be observed that the highest sample rate in the prior-art implementation is 64 kHz whereas the sample rate in the system described above remains at 32 kHz or below. This represents a significant complexity saving, as do the reduced number of processing blocks.