Audio Coding Residual Weighting for High Pitch Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-resolution audio files are cumbersome to stream due to their large file sizes, which can exceed storage capacity and require compression, resulting in a trade-off between sound quality and file size.
Innovation Solution
The method involves receiving an audio signal comprising subband signals, generating a residual signal, determining if a subband signal is a high pitch signal, and performing weighting on the residual signal to generate a weighted residual signal, thereby optimizing audio coding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution audio is stored without compression, then sound quality is maintained, but file size becomes large
Solution Approach 1:
The patent extracts and removes perceptually insignificant audio information components through psychoacoustic modeling. By identifying and eliminating frequencies that are masked by louder sounds or fall outside human hearing ranges, the system reduces file size while preserving perceived sound quality.
Solution Approach 2:
The patent applies different compression strategies to different frequency bands and time segments of the audio signal. By analyzing local characteristics of the audio signal and applying adaptive quantization and masking thresholds specific to each region, the system optimizes the balance between quality and compression ratio locally rather than uniformly across the entire signal.
2Quantity of substance
If high-resolution audio is compressed to reduce file size, then storage and streaming become efficient, but sound quality deteriorates
Solution Approach 1:
The patent employs iterative optimization where the encoded audio is decoded and compared with the original signal. Psychoacoustic masking thresholds are dynamically adjusted based on the reconstructed signal characteristics, allowing the system to refine compression parameters to maintain perceptual quality while achieving target compression ratios.
Solution Approach 2:
The patent dynamically adjusts encoding parameters such as bit allocation, quantization step size, and transform window length based on the temporal and spectral characteristics of the input signal. This adaptive parameter adjustment allows optimal compression efficiency while preserving perceptually important audio features.
3Productivity
If compression is applied to high-resolution audio for streaming, then streaming efficiency improves, but audio fidelity is reduced
Solution Approach 1:
The patent implements dynamic bitrate adaptation and variable compression ratios based on scene complexity and perceptual importance. During stationary passages with fewer perceptual details, higher compression is applied, while during complex transient passages, lower compression is used to preserve fidelity, optimizing overall streaming efficiency.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing audio coding are described. One example of the methods includes receiving an audio signal that includes one or more subband signals. A residual signal of at least one of the one or more subband signals is generated based on the at least one of the one or more subband signals. It is determined that the at least one of the one or more subband signals is a high pitch signal. In response to determining that the at least one of the one or more subband signals is a high pitch signal, weighting is performed on the residual signal of the at least one of the one or more subband signal to generate a weighted residual signal.


