Speech Bandwidth Expansion for Sampling-Rate-Compatible Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech coding and decoding technologies are limited by the sampling rates supported by the speech coder and decoder, restricting the acquisition and playback of speech signals to specific ranges, thereby limiting flexibility and compatibility.
Innovation Solution
A method and apparatus for speech coding and decoding that involves obtaining initial frequency bandwidth feature information, performing feature compression to reduce the sampling rate, and coding the speech signal to a rate supported by the module, while allowing for subsequent expansion to the original sampling rate for playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech coding is performed using conventional speech coders, then the speech signal can be coded and transmitted, but the sampling rate must be within the limited range supported by the speech coder
Solution Approach 1:
The speech signal processing is divided into two independent stages: frequency bandwidth compression (converting high sampling rate to low sampling rate) and speech coding. This segmentation allows the frequency bandwidth converter to handle sampling rate conversion independently, enabling the speech coder to work with its native supported sampling rates while accepting various input sampling rates.
Solution Approach 2:
A frequency bandwidth converter is introduced as an intermediary component between the speech acquisition device and the speech coder. This converter transforms the speech signal from high sampling rate to low sampling rate, serving as a bridge that enables compatibility between different sampling rate requirements without modifying the speech coder itself.
2Adaptability or versatility
If the speech coder supports only specific sampling rates, then the coding process is simplified, but the acquisition and playback of speech signals are restricted to specific sampling rate ranges
Solution Approach 1:
The frequency bandwidth conversion is performed in advance before speech coding. By pre-converting the speech signal to the sampling rate required by the speech coder, the system eliminates the need for complex real-time sampling rate adaptation during coding operations, maintaining operational simplicity while achieving sampling rate flexibility.
Solution Approach 2:
The system changes the sampling rate parameter of the speech signal through frequency bandwidth conversion before processing. This parameter transformation enables the speech signal to match the sampling rate requirements of the speech coder, achieving adaptability without requiring the speech coder to support multiple sampling rates.
3Manufacturing precision
If high sampling rate speech signals are processed, then speech quality is improved, but the speech coder cannot handle signals exceeding its supported sampling rate
Solution Approach 1:
The patent converts the potential harm of high sampling rate signals (which would cause coding errors or failures in conventional systems) into a benefit by using frequency bandwidth compression. The high sampling rate signal is transformed into a low sampling rate signal that matches the speech coder's requirements, allowing high-quality speech processing while ensuring reliable coding operation.
Data Source
AI summary
This application relates to a speech decoding method performed by a computer device. The method includes: obtaining coded speech data corresponding to an original speech signal; decoding the coded speech data to obtain a decoded speech signal; generating target frequency bandwidth feature information corresponding to the decoded speech signal; performing feature extension on target feature information corresponding to a compressed band in the target frequency bandwidth feature information to obtain extended feature information corresponding to a second band, and a frequency interval of the compressed band being less than a frequency interval of the second band; and obtaining, based on the extended feature information corresponding to the second band, a target speech signal corresponding to the original speech signal.


