Embedded CELP Speech Coder Bitrate Scalability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional scalable speech coders require a large number of bit rates to provide bitrate scalability, particularly lacking in 1 kbit/s step scalability, which is inadequate for ensuring consistent speech quality in fluctuating network conditions.
Innovation Solution
An embedded code-excited linear prediction speech coding and decoding apparatus that models error signals based on channel transmission rates using a multiple pulse search mode or gain compensation mode, allowing for optimal bit allocation and improved speech quality by encoding residual excitation signals with additional bits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional scalable speech coders are used, then bitrate scalability is provided, but the speech quality fluctuation is severe due to packet loss in packet-based networks
Solution Approach 1:
The speech coder dynamically adjusts the transmission bitrate based on network conditions and packet loss characteristics. The system transitions between different operating modes (standard mode, packet loss concealment mode, and core layer only mode) to adapt to varying channel conditions, ensuring consistent speech quality despite network fluctuations.
Solution Approach 2:
The speech coder is divided into a core layer and an enhancement layer. The core layer provides basic speech coding functionality that can operate independently, while the enhancement layer adds additional quality. This segmentation allows the system to maintain minimum speech quality through the core layer even when packet loss occurs in the enhancement layer.
2Loss of energy
If the transmission bitrate is reduced to handle channel load, then channel load decreases, but speech quality degradation occurs
Solution Approach 1:
The system changes multiple parameters simultaneously including transmission bitrate, coding mode, and packet loss concealment strategy. By adjusting these parameters based on network conditions, the system optimizes the balance between channel load and speech quality, reducing bitrate only when necessary while maintaining quality through enhanced PLC algorithms.
Solution Approach 2:
The speech coder implements feedback mechanisms that monitor network conditions and speech quality metrics. This feedback is used to dynamically adjust coding parameters and packet loss concealment strategies, ensuring that bitrate reduction does not lead to unacceptable quality degradation.
3Reliability
If packet loss concealment algorithms are used, then some speech quality is maintained, but severe degradation occurs during burst packet loss
Solution Approach 1:
The system prepares multiple packet loss concealment strategies in advance and selects the appropriate one based on the detected packet loss pattern. For burst packet loss, pre-configured PLC algorithms with higher robustness are activated, providing cushioning against severe quality degradation before it occurs.
Solution Approach 2:
The packet loss concealment mechanism dynamically switches between different concealment algorithms based on the severity and pattern of packet loss. During normal conditions, standard PLC is used, but during burst packet loss, more aggressive concealment strategies are activated to maintain speech quality.
4Adaptability or versatility
If separate scalable coding method is used, then bitrate scalability is achieved, but the system complexity increases
Solution Approach 1:
The speech coder merges the core layer and enhancement layer into a unified coding structure with shared processing components. This merging reduces redundancy and simplifies the overall system architecture while maintaining the benefits of scalable coding, making the implementation more practical for real-world applications.
Data Source
AI summary
Provides is an embedded code-excited linear prediction speech coding/decoding apparatus and method that can deal with the capacity change of speech transmission channel by modeling an error signal not coded at a core speech coder based on a transmission rate in a multiple pulse search mode or gain compensation mode and then transmitting it in an optimum mode. The apparatus includes a core speech coding unit for coding an input speech signal with spectral envelop and an excitation signal, a transmission rate determination unit for allocating the number of bits additionally allowed depending on a capacity of a transmission channel, and an embedded excitation signal coding unit for coding a residual excitation signal that is not coded in the core speech coding unit based on the number of additionally allowed bits using one of a multiple pulse excitation coding mode and a gain compensation mode.


