Frequency-Domain Loudness Modification Across Variable Audio Block Sizes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining and modifying the perceived loudness of audio signals, particularly in perceptually coded audio, face challenges when dealing with varying block sizes, as they require complex operations to maintain constant frequency and temporal resolution, leading to difficulties in sound quality for transient signals.
Innovation Solution
The method involves combining frequency domain data from multiple block sizes to form a longest block size, interpolating gain information, and applying it to shorter block sizes to maintain consistent loudness processing, using techniques like interleaving or replication, and delaying data for accurate loudness modification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If perceptual coding methods use varying block sizes to improve sound quality for transient signals, then sound quality is improved, but determining and modifying perceived loudness becomes complex and computationally intensive
Solution Approach 1:
The patent segments the loudness processing task by creating separate processing paths for long blocks and short blocks. Each path uses appropriate block sizes and processing parameters optimized for its specific block type, avoiding the need to handle all varying block sizes through a single complex processing pipeline. This segmentation reduces overall processing complexity while maintaining sound quality.
Solution Approach 2:
The patent changes processing parameters based on block size. For long blocks, it uses one set of processing parameters, while for short blocks, it uses different parameters optimized for transient signals. This parameter adaptation allows efficient loudness processing for each block type without requiring a universally complex processing system.
2Measurement precision
If constant frequency and temporal resolution is maintained in loudness processing, then measurement precision is improved, but processing efficiency decreases due to complex operations on varying block sizes
Solution Approach 1:
The patent implements dynamic processing where the processing path adapts based on block size. Long blocks follow one processing route that maintains high frequency and temporal resolution, while short blocks follow an optimized route. This dynamic adaptation ensures measurement precision is maintained where needed while improving overall processing efficiency.
Solution Approach 2:
The patent performs preliminary classification of blocks by size before processing. This preliminary action allows the system to prepare and apply appropriate processing parameters in advance, ensuring constant frequency and temporal resolution for long blocks while efficiently handling short blocks without requiring complex real-time adjustments during processing.
3Adaptability or versatility
If loudness processing is applied to each varying block size independently, then adaptability is improved, but device complexity increases
Solution Approach 1:
The patent segments adaptability into discrete processing paths for different block sizes. Rather than creating a complex system that handles all possible block size variations independently, it creates specific optimized paths for long and short blocks, reducing structural complexity while maintaining adaptability to the most common block size scenarios.
Solution Approach 2:
The patent creates a universal processing framework that handles both long and short blocks through a common architecture with branching paths. This multi-functional approach allows the same processing structure to adapt to different block sizes without requiring entirely separate processing systems for each block type, thereby reducing overall device complexity.
Data Source
AI summary
Methods of, apparatuses for, and computer readable media having instructions thereon that when executed cause carrying out methods of determining and modifying the perceived loudness of a frequency domain audio signal where the frequency resolution, and corresponding temporal coverage of the frequency domain information is not constant. The frequency (and thus temporal) resolution of the perceived loudness processing is maintained constant at the longest block size. One method includes a block combiner and a loudness modification interpolator.


