Audio Encoding Scale Factor Adjustment for Computational Load Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding systems, such as those using the Advanced Audio Coding (AAC) standard, face computational intensity challenges in real-time encoding due to the need for calculating masking thresholds, which is laborious and time-consuming, especially for consumer electronics with fixed-point digital signal processors.
Innovation Solution
The method involves transforming time-domain audio signals into frequency-domain signals, determining scale factors for each frequency band based on energy comparisons with adjacent bands, and adjusting these scale factors to reduce computational load, allowing for real-time encoding even in devices with limited processing capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If masking threshold calculation is performed to remove masked audio portions, then audio compression ratio is improved, but computational complexity increases
Solution Approach 1:
The patent extracts only the essential features needed for compression (scale factors and channel coupling information) rather than performing complete masking threshold calculations. By taking out only the critical redundancy-reducing elements, the system achieves compression without the full computational burden of traditional psychoacoustic modeling.
Solution Approach 2:
Instead of performing complete masking threshold analysis across all frequency bands, the patent applies partial action by selectively processing only those bands where interchannel or temporal redundancy is significant. This partial processing approach maintains compression effectiveness while reducing overall computational complexity.
2Manufacturing precision
If masking threshold calculation is performed to ensure audio fidelity, then audio quality is improved, but encoding time increases
Solution Approach 1:
The patent performs preliminary action by pre-determining scale factors and coupling information before full encoding. These preliminary calculations capture the essential redundancy patterns, allowing the main encoding process to proceed quickly without repeated masking threshold calculations, thus reducing encoding time while preserving fidelity.
Solution Approach 2:
The patent changes the parameters being processed from complete spectral masks to simplified scale factors and coupling indicators. This parameter transformation reduces the computational workload significantly while maintaining the ability to preserve audio fidelity through targeted redundancy removal.
3Loss of substance
If complete psychoacoustic processing is performed, then compression effectiveness is improved, but processing speed decreases
Solution Approach 1:
The patent segments the audio processing into distinct stages: scale factor determination, interchannel redundancy detection, and temporal redundancy detection. Each segment handles a specific aspect of compression, allowing parallel processing and optimizing speed without sacrificing overall compression effectiveness.
Solution Approach 2:
The patent substitutes the traditional mechanical masking threshold calculation mechanism with a more efficient system based on scale factor analysis and interchannel correlation. This substitution replaces computationally intensive operations with simpler mathematical transformations, dramatically improving processing speed while maintaining compression ratios.
Data Source
AI summary
A method of encoding a time-domain audio signal is presented. A device transforms the time-domain signal into a frequency-domain signal including a sequence of sample blocks, wherein each block includes a coefficient for each of multiple frequencies. The coefficients of each block are grouped into frequency bands. For each frequency band of each block, a scale factor is estimated for the band, and the energy of the band for the block is compared with the energy of the band of an adjacent sample block, wherein the blocks may be adjacent to each other in either or both of an interchannel and a temporal sense. If the ratio of the band energy for the first block to the band energy for the adjacent block is less than some value, the scale factor of the band for the first block is increased. The coefficients of the band for each block are quantized based on the resulting scale factor. The encoded audio signal is generated based on the quantized coefficients and the scale factors.


