Speech Coding Subframe Segmentation for Low Bit Rate
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech coding technologies face challenges in achieving efficient low bit rate coding while maintaining good subjective quality, particularly in half-rate modes of source-controlled variable-rate speech coding systems, where bit rate reduction is essential to improve system capacity without compromising sound quality.
Innovation Solution
The method involves dividing speech frames into subframe units and conducting searches for fixed and adaptive codebook contributions, with at least one subframe unit encoded without the fixed codebook contribution, optimizing the pitch gain and error minimization to reduce bit rate while maintaining performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional CELP coding is used with both fixed and adaptive codebook contributions in all subframes, then speech quality is maintained, but bit rate is too high for efficient half-rate operation
Solution Approach 1:
The frame is divided into multiple subframes, and the invention selectively applies different coding strategies to different subframes. Specifically, some subframes are coded with both fixed and adaptive codebooks while others use only the adaptive codebook, allowing bit rate reduction while maintaining overall speech quality through selective segmentation of coding resources.
Solution Approach 2:
The invention dynamically changes the coding parameters by selecting whether to include fixed codebook contribution on a per-subframe basis. This parameter change is controlled by a decision mechanism that evaluates speech characteristics (voiced/unvoiced detection) to determine the optimal coding mode for each subframe, thereby adapting bit rate to actual speech content requirements.
2Productivity
If bit rate is reduced by eliminating fixed codebook contribution in some subframes, then system capacity improves, but speech quality may deteriorate
Solution Approach 1:
The invention applies different coding quality levels to different subframes based on local speech characteristics. Subframes containing voiced speech segments retain both fixed and adaptive codebook contributions to maintain high quality, while unvoiced or less critical subframes use only adaptive codebook contribution. This local quality adjustment allows system capacity improvement without uniform speech quality degradation.
Solution Approach 2:
The coding structure is made dynamic by allowing the fixed codebook contribution to be selectively enabled or disabled on a per-subframe basis. This dynamic adaptation is controlled by real-time analysis of speech activity and voiced/unvoiced detection, enabling the system to flexibly adjust between quality and bit rate based on instantaneous speech characteristics, thereby improving overall system capacity while maintaining acceptable quality.
3Loss of information
If full CELP model is used for all frames, then speech quality is preserved, but complexity increases and efficiency decreases in half-rate modes
Solution Approach 1:
Instead of applying the complete CELP model (with both fixed and adaptive codebooks) to all subframes, the invention uses partial action by selectively applying only the adaptive codebook contribution to certain subframes. This partial application of the full model reduces computational complexity and bit rate requirements while maintaining sufficient speech quality for those subframes where the reduced model is adequate, thereby improving efficiency in half-rate modes.
Data Source
AI summary
A method for coding speech or other generic signals includes dividing a speech signal into a plurality of frames, and dividing at least one of the plurality of frames into at least two subframe units. A search for a fixed codebook contribution and an adaptive codebook contribution for subframe units is conducted. At least one subframe unit is selected to be coded without the fixed codebook contribution. The encoder may iteratively arrange and encode subframes differently for the same frame, and select for transmission that arrangement that minimizes an error measure across the frame. Various embodiments are shown, as are embodied computer programs, a decoder, and a communication system.


