Speech Coding With Frequency Bandwidth Compression for Rate Compatibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech coding and decoding technologies are limited by the sampling rates supported by the speech coder and decoder, restricting the acquisition and playback of speech signals to specific ranges, thereby limiting flexibility and compatibility.
Innovation Solution
A method and apparatus for speech coding and decoding that involves obtaining initial frequency bandwidth feature information, performing feature compression to reduce the sampling rate, and coding the speech signal to a rate supported by the coder, while allowing for subsequent expansion to the original sampling rate for playback, using frequency bandwidth compression and extension techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech coding is performed using conventional speech coders, then the coding process is simple and reliable, but the sampling rate range is limited and flexibility is reduced
Solution Approach 1:
The speech signal processing is divided into multiple stages: original speech signal → frequency bandwidth compression → compressed speech signal → speech coding. This segmentation allows each stage to handle specific tasks, with the frequency bandwidth compression module reducing the sampling rate before coding, thereby expanding the adaptable sampling rate range while keeping each individual module relatively simple
Solution Approach 2:
A frequency bandwidth compression module is introduced as an intermediary between the speech acquisition device and the speech coder. This intermediary component transforms the original speech signal into a compressed speech signal with reduced sampling rate, enabling the conventional speech coder to handle a wider range of sampling rates without requiring the coder itself to be complex
2Manufacturing precision
If speech signals with high sampling rates are processed, then the speech quality is improved, but the data transmission load increases
Solution Approach 1:
The frequency bandwidth compression module extracts and retains only the essential frequency components of the speech signal while discarding redundant high-frequency information. This extraction process reduces the sampling rate and data transmission load while preserving sufficient speech quality for intelligible communication
Solution Approach 2:
The sampling rate parameter is changed from the original high value to a compressed lower value through the frequency bandwidth compression module. This parameter change reduces the data transmission load while the module is designed to maintain acceptable speech quality by preserving important frequency characteristics
3Adaptability or versatility
If frequency bandwidth compression is applied, then the sampling rate is reduced and compatibility is improved, but the processing complexity increases
Solution Approach 1:
The frequency bandwidth compression module is designed with multi-functionality: it performs frequency analysis, bandwidth compression, and signal reconstruction in a single integrated component. This universal design allows the module to handle different input sampling rates and produce compatible output for conventional speech coders, improving compatibility while avoiding the need for multiple separate processing components
Solution Approach 2:
The frequency bandwidth compression is performed as a preliminary action before speech coding. By pre-compressing the speech signal and reducing its sampling rate beforehand, the subsequent speech coding process becomes simpler and more compatible with conventional coders, while the overall processing complexity is managed through this upfront preparation step
Data Source
AI summary
This application relates to a speech coding method performed by a computer device. The method includes: obtaining initial frequency bandwidth feature information corresponding to a speech signal; performing feature compression on initial feature information corresponding to a second band in the initial frequency bandwidth feature information to obtain target feature information corresponding to a compressed band, a frequency interval of the second band being greater than a frequency interval of the compressed band; obtaining, based on the target feature information corresponding to the compressed band, a compressed speech signal corresponding to the speech signal; and coding the compressed speech signal to obtain coded speech data corresponding to the speech signal, a target sampling rate corresponding to the compressed speech signal being less than a sampling rate corresponding to the speech signal.


