Coded Audio Loudness Measurement Without Full Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for measuring the loudness of low-bitrate coded audio, such as Dolby Digital, Dolby Digital Plus, and Dolby E, require full decoding of audio signals, which is computationally intensive and inefficient, especially when only a loudness measurement is needed without decoding the audio.
Innovation Solution
A method that partially decodes the audio signal to derive an approximation of the power spectrum from the bitstream, using exponents or scale factors, allowing for a computationally economical measurement of loudness without fully decoding the audio, employing weighted power or psychoacoustic measures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full decoding of audio signals is performed to measure loudness, then measurement precision is improved, but computational complexity and processing time increase significantly
Solution Approach 1:
The patent extracts only the exponent data from the coded audio bitstream, which contains sufficient information for loudness measurement without requiring full decoding. This extraction approach isolates the necessary component (exponents) from the complete decoding process, achieving accurate loudness measurement while avoiding the computational burden of reconstructing the entire audio signal.
Solution Approach 2:
The patent applies partial action by performing only the necessary portion of the decoding process - extracting exponents and calculating power spectrum approximation - rather than completing the full decoding sequence. This partial processing provides sufficient information for loudness measurement without the excessive computational cost of complete audio reconstruction.
2Measurement precision
If full decoding is performed for loudness measurement, then measurement accuracy is improved, but processing speed decreases
Solution Approach 1:
The patent extracts only the exponent data from the coded audio bitstream, which contains sufficient information for loudness measurement without requiring full decoding. This extraction approach isolates the necessary component (exponents) from the complete decoding process, achieving accurate loudness measurement while avoiding the computational burden of reconstructing the entire audio signal.
Solution Approach 2:
The patent skips the computationally intensive steps of full audio decoding (inverse transform, bit allocation, mantissa processing) and rushes directly to the essential measurement task by using exponents to calculate power spectrum approximation. This skipping of unnecessary intermediate steps maintains measurement accuracy while dramatically improving processing speed.
3Reliability
If full decoding processes (bit allocation, inverse transformation) are performed, then loudness measurement reliability is improved, but energy consumption increases
Solution Approach 1:
The patent extracts only the exponent data from the coded audio bitstream, which contains sufficient information for loudness measurement without requiring full decoding. This extraction approach isolates the necessary component (exponents) from the complete decoding process, achieving accurate loudness measurement while avoiding the computational burden of reconstructing the entire audio signal.
Solution Approach 2:
The patent uses a simplified, low-cost approach by relying solely on exponent data for power spectrum approximation rather than investing energy in complete audio reconstruction. This disposable-like approach uses minimal computational resources (just the exponents) to achieve the measurement goal without the expensive overhead of full decoding processes.
Data Source
AI summary
Measuring the loudness of audio encoded in a bitstream that includes data from which an approximation of the power spectrum of the audio can be derived without fully decoding the audio is performed by deriving the approximation of the power spectrum of the audio from said bitstream without fully decoding the audio, and determining an approximate loudness of the audio in response to the approximation of the power spectrum of the audio. The data may include coarse representations of the audio and associated finer representations of the audio, the approximation of the power spectrum of the audio being derived from the coarse representations of the audio. In the case of subband encoded audio, the coarse representations of the audio may comprise scale factors and the associated finer representations of the audio may comprise sample data associated with each scale factor.


