Audio Stem Masking Detection Using Relative Loudness Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for addressing sound masking in audio mixes are imprecise, error-prone, and labor-intensive, relying heavily on user recognition and trial-and-error to identify and correct frequency ranges where masking occurs.

Innovation Solution

A system and method that model sound masking as a function of energy and relative energy between audio stems, using psychoacoustic models to compute loudness and partial loudness, and identify frequency ranges with significant loudness loss to enable systematic and efficient correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional trial-and-error methods are used to identify and correct sound masking, then users can eventually achieve desired audio quality, but the process becomes labor-intensive and time-consuming

Engineering Contradiction:
Improveaccuracy of sound masking identificationVSAvoidtime required for audio production
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the manual trial-and-error mechanical process with an automated computational system. The system uses a computer to automatically analyze audio stems, compute loudness and partial loudness values, identify masking frequency ranges, and generate correction recommendations, eliminating the need for manual trial-and-error adjustment by users.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically performing the complete sound masking identification and correction process without requiring user intervention for each adjustment. The computer system independently analyzes the audio, identifies problems, and provides corrective measures, allowing users to simply review and apply recommendations rather than manually troubleshoot.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If users manually recognize and correct sound masking frequency ranges, then audio quality can be improved, but the process becomes error-prone and labor-intensive

Engineering Contradiction:
Improveprecision of frequency range identificationVSAvoidease of sound masking correction
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces manual auditory recognition and manual correction operations with automated computer-based analysis and recommendation generation. The system objectively computes loudness values and identifies masking frequency ranges algorithmically, eliminating human error in detection and providing precise, data-driven correction guidance.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If systematic automated methods are implemented to identify sound masking, then accuracy and efficiency improve, but device complexity increases

Engineering Contradiction:
Improvespeed of sound masking identificationVSAvoidcomplexity of audio analysis system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the audio analysis process into distinct computational steps: separating audio into individual stems, computing loudness for each stem, identifying masking frequency ranges, and generating correction recommendations. This segmentation allows the complex task to be broken down into manageable, automated operations that can be performed systematically by the computer system.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10763812B2Systems and methods for identifying and remediating sound masking
Publication Date: 2020.09.01 NATIVE INSTR USA INC
  • US10763812B2 patent drawing
  • US10763812B2 patent drawing
  • US10763812B2 patent drawing

AI summary

Some embodiments of the invention are directed to enabling a user to easily identify the frequency range(s) at which sound masking occurs, and addressing the masking, if desired. In this respect, the extent to which a first stem is masked by one or more second stems in a frequency range may depend not only on the absolute value of the energy of the second stem(s) in the frequency range, but also on the relative energy of the first stem with respect to the second stem(s) in the frequency range. Accordingly, some embodiments are directed to modeling sound masking as a function of the energy of the stem being masked and of the relative energy of the masked stem with respect to the masking stem(s) in the frequency range, such as by modeling sound masking as loudness loss, a value indicative of the reduction in loudness of a stem of interest caused by the presence of one or more other stems in a frequency range.