Audio Stem Masking Detection Through Partial Loudness Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for addressing sound masking in audio mixes are imprecise, error-prone, and labor-intensive, relying heavily on user recognition and trial-and-error to identify and correct frequency ranges where masking occurs.
Innovation Solution
The development of methods and systems that model sound masking as a function of energy and relative energy between audio stems, using psychoacoustic models to compute loudness and partial loudness, and identify frequency ranges with significant loudness loss for targeted corrective measures like equalization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional trial-and-error methods are used to identify sound masking, then user flexibility is maintained, but identification precision and efficiency deteriorate
Solution Approach 1:
The patent replaces the mechanical trial-and-error process with an automated computational system that uses psychoacoustic models and signal processing algorithms to automatically identify masking frequency ranges, thereby improving precision and eliminating time consumption associated with manual iteration
Solution Approach 2:
The system performs self-analysis by automatically computing loudness, partial loudness, and loudness loss metrics without requiring user intervention, enabling the system to identify masking issues independently and efficiently
2Productivity
If automated psychoacoustic modeling is implemented, then identification efficiency improves, but system complexity increases
Solution Approach 1:
The patent segments the audio signal into multiple stems and analyzes each stem's contribution to masking separately, breaking down the complex mixing problem into manageable frequency ranges and stem combinations that can be processed systematically
Solution Approach 2:
The system introduces intermediate computational metrics (loudness, partial loudness, loudness loss) that serve as mediators between the raw audio signals and the final masking identification, simplifying the analysis process through standardized measurement steps
3Measurement precision
If comprehensive stem analysis is performed, then masking identification accuracy improves, but computational energy consumption increases
Solution Approach 1:
The patent applies local quality analysis by focusing computational resources on specific frequency ranges where masking is most likely to occur, rather than uniformly analyzing all frequency bands, thereby improving accuracy while reducing overall computational energy consumption
Solution Approach 2:
The system performs partial analysis by computing loudness metrics only for relevant stem combinations and frequency ranges identified as potential masking sources, avoiding exhaustive analysis of all possible combinations and reducing computational overhead
Data Source
AI summary
Some embodiments of the invention are directed to enabling a user to easily identify the frequency range(s) at which sound masking occurs, and addressing the masking, if desired. In this respect, the extent to which a first stem is masked by one or more second stems in a frequency range may depend not only on the absolute value of the energy of the second stem(s) in the frequency range, but also on the relative energy of the first stem with respect to the second stem(s) in the frequency range. Accordingly, some embodiments are directed to modeling sound masking as a function of the energy of the stem being masked and of the relative energy of the masked stem with respect to the masking stem(s) in the frequency range, such as by modeling sound masking as loudness loss, a value indicative of the reduction in loudness of a stem of interest caused by the presence of one or more other stems in a frequency range.


