Audio Decoder Context Reset for Random Access and Bit Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio and speech signal decoding technologies face challenges in efficient bit reduction and random access decoding, particularly when previous frame information is lacking, leading to potential wrong decoding or system failure during random access scenarios.
Innovation Solution
The method involves switching between frequency domain coding and linear prediction domain coding modes, using algebraic code excited linear prediction (ACELP) or transform coded excitation (TCX) for low frequency bands and enhanced spectral band replication (eSBR) for high frequency bands, while incorporating context reset information and random access availability to ensure reliable decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If context-based lossless coding is performed using information from previous frames, then coding efficiency is improved, but random access decoding becomes difficult when previous frame information is unavailable
Solution Approach 1:
The patent applies preliminary action by resetting the arithmetic coding context to initial values at the beginning of each frame. This preparation in advance ensures that decoding can start independently at any frame without requiring previous frame information, thereby enabling random access while maintaining coding efficiency through context-based lossless coding when frames are processed sequentially.
2Reliability
If ACELP coding mode is always used, then speech coding reliability is improved, but bit rate flexibility and audio quality are reduced
Solution Approach 1:
The patent applies dynamics by making the coding mode adaptive rather than fixed. The system dynamically switches between ACELP mode (for speech-like signals) and TCX mode (for music-like signals) based on the characteristics of the input signal and available bit rate. This allows the encoder to optimize performance for different signal types and bit rate conditions, combining speech coding reliability with audio quality and bit rate flexibility.
3Manufacturing precision
If detailed fine structure coding is applied to high frequency band, then audio quality is improved, but bit rate consumption increases significantly
Solution Approach 1:
The patent applies local quality by differentiating the coding approach between low and high frequency bands. In the low frequency band, detailed fine structure coding is applied to preserve speech intelligibility. In the high frequency band, where human hearing is less sensitive to fine structures, the patent uses spectral band replication (SBR) technology that requires fewer bits. This localized differentiation of coding precision optimizes audio quality while controlling bit rate consumption.
4Quantity of substance
If stereo signal is converted to mono signal for compression, then bit rate is reduced, but stereo information is lost
Solution Approach 1:
The patent applies the extraction principle by separating the stereo signal into mono signal and stereo parameter components. The mono signal is fully decoded to reconstruct the base audio, while the stereo parameters (such as inter-channel level difference and inter-channel time difference) are extracted and decoded separately. These parameters are then used to reconstruct the stereo effect, allowing bit rate reduction through mono coding while preserving essential stereo information.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for coding and decoding an audio signal or speech signal and an apparatus adopting the method are provided.