Codebook Search Optimization for Voice Codecs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice coding technologies face challenges in reducing computational complexity while maintaining high speech quality, particularly in mobile applications where power consumption and efficiency are critical, due to the high complexity of fixed codebook searches in voice codecs like SMV.
Innovation Solution
The implementation of a Selective Joint Search method that restricts the codebook search to a reduced number of tracks contributing least to the Cost function, reducing computational complexity by up to 66% without affecting perceptual quality, by refining underperforming tracks and optimizing pulse position selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a full codebook search is performed to maintain high speech quality, then speech quality is improved, but computational complexity increases
Solution Approach 1:
The codebook search process is segmented into two phases: a coarse search phase that identifies promising track combinations, and a refined search phase that optimizes pulse positions within those combinations. This segmentation reduces the overall search space and computational complexity while maintaining speech quality.
Solution Approach 2:
The invention applies different search strategies to different track combinations based on their contribution to speech quality. Track combinations that contribute more significantly to speech quality undergo more thorough optimization, while less critical combinations receive minimal processing. This local quality approach optimizes computational resources efficiently.
2Measurement precision
If a full codebook search is performed to maintain high speech quality, then speech quality is improved, but power consumption increases
Solution Approach 1:
The codebook search is divided into coarse and refined phases, reducing the total number of computations required. This segmentation directly lowers power consumption by minimizing the active processing time while maintaining speech quality through targeted optimization in the refined phase.
Solution Approach 2:
The invention performs partial searches on less critical track combinations and full optimization only on the most promising combinations. This partial action approach reduces overall power consumption by avoiding exhaustive searches on all track combinations, while still maintaining acceptable speech quality.
3Productivity
If the codebook search is restricted to reduce computational complexity, then productivity is improved, but speech quality may deteriorate
Solution Approach 1:
The coarse search phase performs preliminary identification of promising track combinations before the refined search phase optimizes pulse positions. This preliminary action ensures that the reduced search space still contains the optimal or near-optimal solutions, maintaining speech quality while improving processing efficiency.
Solution Approach 2:
The invention uses feedback from the coarse search phase to guide the refined search phase. The results of the preliminary track combination identification inform which combinations require further optimization, ensuring that speech quality is maintained in the critical regions while reducing overall computational complexity.
Data Source
AI summary
An electronic circuit (1100) including a processor circuit (1110) and a storage circuit establishing a speech coder (1170) for execution by said processor (1110), the speech coder (1170) for approximating speech by pulses having pulse positions selectable from a codebook (550), the speech coder (1170) operable to obtain (1310) a set of estimated pulse positions having a first number of pulse tracks of the estimated pulse positions, use (1320) a cost function (epsilon tilde {tilde over (ε)}) relating to approximation to speech to find a first subset including a second number of one or more pulse tracks fewer in number than the first number wherein the first subset of pulse tracks contributed a lower contribution to the cost function relative to a second subset of pulse tracks, and control (1330) a subsequent pulse position search beginning with the lower-contributing subset of pulse tracks to yield pulse positions to provide a value of the cost function representing a better approximation to speech. Other forms of the invention involve systems, circuits, devices, processes and processes of operation, as disclosed and claimed.


