Audio Envelope Separation Using a Precomputed Spill Mix Matrix

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound separation technologies require a large processing load to estimate transmission characteristics of spill sound between sound sources, which is unnecessary when only sound levels of individual sound sources are needed.

Innovation Solution

An audio processing method that generates output envelopes using a mix matrix to determine the mix proportions of spill sounds in sound signals, reducing the need for complex estimation by employing Non-negative Matrix Factorization to separate target sounds from spill sounds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If transmission characteristics of spill sound are estimated using conventional methods, then sound separation accuracy is improved, but processing load increases significantly

Engineering Contradiction:
Improvesound separation accuracyVSAvoidprocessing load
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the envelope information from the sound signals rather than processing the complete transmission characteristics. By taking out just the essential envelope data and using a pre-acquired mix matrix, the system achieves adequate sound level measurement without the heavy computational burden of full transmission characteristic estimation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The mix matrix is acquired in advance through preliminary measurements of spill sound transmission characteristics between sound sources. This pre-computed matrix stores the spatial relationships and spill sound proportions, allowing the system to perform rapid sound level calculations during actual use without repeating complex estimation procedures.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If complete sound separation for each sound source is performed, then individual sound source identification is improved, but processing complexity increases

Engineering Contradiction:
Improvesound source identification accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies partial action by focusing only on obtaining sound levels rather than performing complete sound separation. The system processes envelope information and uses the mix matrix to calculate sound levels directly, achieving the necessary information without the excessive complexity of full spectral and temporal sound separation.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If spill sound transmission characteristics are fully estimated, then spill sound removal accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespill sound removal accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The mix matrix containing spill sound transmission characteristics is computed in advance and stored. During actual sound level measurement, the system simply queries this pre-computed matrix and performs straightforward calculations with envelope data, eliminating the need for time-consuming real-time transmission characteristic estimation while maintaining spill sound removal accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12475872B2Audio processing method, audio processing system, and computer-readable medium
Publication Date: 2025.11.18 YAMAHA CORP
  • US12475872B2 patent drawing
  • US12475872B2 patent drawing
  • US12475872B2 patent drawing

AI summary

An audio processing method obtains observed envelopes of picked-up sound signals including a first observed envelope representing a contour of a first sound signal including a first target sound from a first sound source and a second spill sound from a second sound source and a second observed envelope representing a contour of a second sound signal including a second target sound from the second sound source and a first spill sound from the first sound source; and generates, based on the observed envelopes, output envelopes including a first output envelope representing a contour of the first target sound in the first observed envelope and a second output envelope representing a contour of the second target sound in the second observed envelope, using a mix matrix including a mix proportion of the second spill sound in the first sound signal and a mix proportion of the first spill sound in the second sound signal.