Audio Encoding Bit Rate Prediction for Consistent Frame Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding methods using fixed encoding bit rates fail to maintain consistent quality due to the variability of speech signals over time, leading to inconsistent encoding quality.
Innovation Solution
An audio encoding method that dynamically adjusts the encoding bit rate based on audio feature parameters using an encoding bit rate prediction model, allowing for personalized bit rate settings for each audio frame.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a fixed encoding bit rate is used in audio encoding, then the encoding process is simple and efficient, but the encoding quality varies greatly for different speech signals
Solution Approach 1:
The patent applies dynamics by transitioning from a fixed encoding bit rate to a dynamic bit rate that changes according to the characteristics of each audio frame. The system calculates an encoding bit rate for each audio frame based on its features (such as signal complexity, noise level, or spectral characteristics), allowing the bit rate to adapt to the varying requirements of different speech signals in real-time, thus resolving the contradiction between encoding efficiency and quality consistency
Solution Approach 2:
The patent implements parameter changes by modifying the encoding bit rate parameter dynamically based on audio frame characteristics. Instead of using a constant bit rate parameter, the system adjusts this parameter for each audio frame according to its specific features, enabling the encoding system to optimize between compression efficiency and quality maintenance across different signal conditions
2Reliability
If a higher encoding bit rate is used, then the encoding quality improves, but the bandwidth and storage space consumption increase
Solution Approach 1:
The patent applies local quality by assigning different encoding bit rates to different audio frames based on their individual characteristics. Instead of using a uniform high bit rate for the entire audio stream, the system allocates higher bit rates only to frames that require them (such as frames with complex speech patterns or low noise levels) and uses lower bit rates for frames that can tolerate more compression, thus optimizing the balance between overall encoding quality and total bandwidth/storage consumption
Solution Approach 2:
The patent implements partial action by applying high encoding bit rates only partially to specific audio frames rather than universally to all frames. The system evaluates each frame's characteristics and selectively applies higher bit rates only when necessary to maintain quality, avoiding excessive bandwidth and storage consumption for frames that do not require such high quality preservation
Data Source
AI summary
An audio decoding method is performed by a computer device. The method includes: obtaining an audio feature parameter corresponding to audio frames in an original audio; performing encoding bit rate prediction processing on the audio feature parameter through an encoding bit rate prediction model, to obtain an audio encoding bit rate of the audio frames, wherein the encoding bit rate prediction model is configured to predict the audio encoding bit rate according to a target encoding quality score; performing audio encoding on the sample audio frames based on the corresponding sample encoding bit rates to generate sample audio data corresponding to the sample audio frames; and generating target audio data from the original audio by performing audio encoding on the audio frames based on the audio encoding bit rate.


