Audio Stem Masking Detection Through Partial Loudness Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for addressing sound masking in audio mixes are imprecise, error-prone, and labor-intensive, relying heavily on user recognition and trial-and-error to identify and correct frequency ranges where masking occurs.

Innovation Solution

The development of methods and systems that model sound masking as a function of energy and relative energy between audio stems, using psychoacoustic models to compute loudness and partial loudness, and identify frequency ranges with significant loudness loss for targeted corrective measures like equalization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional trial-and-error methods are used to identify sound masking, then user flexibility is maintained, but identification precision and efficiency deteriorate

Engineering Contradiction:
Improveidentification precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical trial-and-error process with an automated computational system that uses psychoacoustic models and signal processing algorithms to automatically identify masking frequency ranges, thereby improving precision and eliminating time consumption associated with manual iteration

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-analysis by automatically computing loudness, partial loudness, and loudness loss metrics without requiring user intervention, enabling the system to identify masking issues independently and efficiently

Inventive Principle:
Principle #25Self-service

2Productivity

If automated psychoacoustic modeling is implemented, then identification efficiency improves, but system complexity increases

Engineering Contradiction:
Improveidentification efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the audio signal into multiple stems and analyzes each stem's contribution to masking separately, breaking down the complex mixing problem into manageable frequency ranges and stem combinations that can be processed systematically

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediate computational metrics (loudness, partial loudness, loudness loss) that serve as mediators between the raw audio signals and the final masking identification, simplifying the analysis process through standardized measurement steps

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If comprehensive stem analysis is performed, then masking identification accuracy improves, but computational energy consumption increases

Engineering Contradiction:
Improvemasking identification accuracyVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality analysis by focusing computational resources on specific frequency ranges where masking is most likely to occur, rather than uniformly analyzing all frequency bands, thereby improving accuracy while reducing overall computational energy consumption

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs partial analysis by computing loudness metrics only for relevant stem combinations and frequency ranges identified as potential masking sources, avoiding exhaustive analysis of all possible combinations and reducing computational overhead

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10396744B2Systems and methods for identifying and remediating sound masking
Publication Date: 2019.08.27 NATIVE INSTR USA INC
  • US10396744B2 patent drawing
  • US10396744B2 patent drawing
  • US10396744B2 patent drawing

AI summary

Some embodiments of the invention are directed to enabling a user to easily identify the frequency range(s) at which sound masking occurs, and addressing the masking, if desired. In this respect, the extent to which a first stem is masked by one or more second stems in a frequency range may depend not only on the absolute value of the energy of the second stem(s) in the frequency range, but also on the relative energy of the first stem with respect to the second stem(s) in the frequency range. Accordingly, some embodiments are directed to modeling sound masking as a function of the energy of the stem being masked and of the relative energy of the masked stem with respect to the masking stem(s) in the frequency range, such as by modeling sound masking as loudness loss, a value indicative of the reduction in loudness of a stem of interest caused by the presence of one or more other stems in a frequency range.