Audio Stem Masking Analysis Using Loudness Loss Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for addressing sound masking in audio mixes are imprecise, error-prone, and time- and labor-intensive, relying heavily on user recognition and trial-and-error to identify and correct frequency ranges affected by masking.
Innovation Solution
The system models sound masking as a function of energy and relative energy between audio stems, using psychoacoustic models to compute loudness and partial loudness, and identifies frequency ranges with significant loudness loss to enable systematic and efficient correction through graphical user interfaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional trial-and-error methods are used to identify and correct sound masking, then users can eventually achieve desired audio quality, but the process becomes time-consuming and labor-intensive
Solution Approach 1:
The patent replaces the manual trial-and-error mechanical process with an automated computational system that uses psychoacoustic models and signal processing algorithms to automatically identify and correct sound masking, eliminating the need for iterative manual adjustment
Solution Approach 2:
The system performs self-analysis by automatically computing loudness loss, identifying masking frequency ranges, and suggesting corrective equalization parameters without requiring user intervention or subjective listening tests
2Measurement precision
If manual sound masking correction is performed without systematic analysis, then the process is simple to operate, but the precision and accuracy of correction are compromised
Solution Approach 1:
The patent segments the audio spectrum into discrete frequency ranges and analyzes loudness loss independently in each segment, allowing precise identification of specific frequency ranges affected by masking while maintaining manageable computational complexity
Solution Approach 2:
The system introduces psychoacoustic models and excitation patterns as intermediary computational layers that bridge the raw audio signal and the final masking identification, enabling accurate measurement without directly complex user interaction
3Reliability
If users rely on subjective recognition of sound masking, then the approach is easy to implement, but errors and inconsistencies increase
Solution Approach 1:
The system provides objective feedback through computed loudness loss values and masking probability metrics that guide users in making informed mixing decisions, replacing subjective judgment with quantifiable data while maintaining user control
Data Source
AI summary
Some embodiments of the invention are directed to enabling a user to easily identify the frequency range(s) at which sound masking occurs, and addressing the masking, if desired. In this respect, the extent to which a first stem is masked by one or more second stems in a frequency range may depend not only on the absolute value of the energy of the second stem(s) in the frequency range, but also on the relative energy of the first stem with respect to the second stem(s) in the frequency range. Accordingly, some embodiments are directed to modeling sound masking as a function of the energy of the stem being masked and of the relative energy of the masked stem with respect to the masking stem(s) in the frequency range, such as by modeling sound masking as loudness loss, a value indicative of the reduction in loudness of a stem of interest caused by the presence of one or more other stems in a frequency range.


