Stereo Audio Separation and Spatial Remixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Legacy audio content recorded in stereo often results in imperfect separation of audio content classes, leading to audible artifacts and distortion when separated into stems like dialogue, music, and sound effects, which affects the spatial sound experience in reproduction systems.

Innovation Solution

A trained machine is configured to separate stereo audio signals into multiple audio content classes while conserving summed levels and minimizing distortion, using a separation module that includes neural networks to process the audio signals in the time-frequency domain and adjust gains for binaural reproduction, ensuring spatial localization without cross-talk and preserving phase information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If stereo audio signals are separated into multiple audio content classes using traditional methods, then audio content can be categorized into dialogue, music, and sound effects, but audible artifacts and distortion occur during the separation process

Engineering Contradiction:
Improveaudio content class separationVSAvoidseparation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent transforms the audio signal from time domain to frequency domain using Short-Time Fourier Transform (STFT), converting the separation problem into a parameter-based optimization in the frequency domain. This parameter transformation enables more precise control over separation quality and reduces artifacts by operating on spectral components rather than raw waveforms.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical/audio signal processing methods with a trained machine learning model (neural network) that processes time-frequency representations. This substitution enables the system to learn optimal separation parameters from training data, significantly improving separation accuracy while maintaining audio fidelity across different content types.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If separated audio signals are spatially localized into multiple output channels, then binaural reproduction quality improves, but distortion artifacts from separation are amplified

Engineering Contradiction:
Improvebinaural reproduction qualityVSAvoidseparation artifacts
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent performs gain optimization and phase alignment on the separated audio signals before spatial localization and mixing. By pre-processing the separated stems to correct gain imbalances and phase relationships, the system prevents artifacts from being amplified during the spatial remixing process, ensuring high-quality binaural reproduction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs a trained machine that was trained on paired data of mixed audio and corresponding separated stems, enabling it to predict optimal separation parameters. This feedback mechanism allows the system to learn from training examples and adjust separation parameters dynamically, reducing artifacts while maintaining separation quality during spatial localization.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If gain is adjusted of output channels to conserve summed levels, then audio level consistency is maintained, but phase distortion may occur during separation

Engineering Contradiction:
Improvelevel consistencyVSAvoidphase information
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent processes audio signals by segmenting them into overlapping time frames and transforming each frame to the frequency domain independently. This segmentation allows gain adjustment to be applied to individual frequency bins and time frames, enabling precise level control while preserving phase relationships through proper windowing and overlap-add reconstruction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent optimizes gain parameters in the frequency domain rather than in the time domain, allowing independent control of magnitude and phase parameters. By operating on spectral parameters separately, the system can adjust gains to conserve summed levels while explicitly preserving or restoring phase information through phase alignment algorithms.

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If a trained machine uses neural networks to process audio in the time-frequency domain, then separation accuracy improves, but computational complexity increases

Engineering Contradiction:
Improveseparation accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the audio processing task into discrete time frames and frequency bins, allowing the neural network to process smaller, manageable segments independently. This segmentation reduces the computational burden on the trained machine while maintaining high separation accuracy through frame-by-frame analysis and efficient reconstruction using overlap-add methods.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11979723B2Content based spatial remixing
Publication Date: 2024.05.07 WAVES AUDIO
  • US11979723B2 patent drawing
  • US11979723B2 patent drawing
  • US11979723B2 patent drawing

AI summary

A trained machine configured to input a stereo sound track and separate the stereo sound track into multiple N separated stereo audio signals respectively characterized by multiple N audio content classes. All stereo audio as input in the stereo sound track is included in the N separated stereo audio signals. A mixing module is configured to spatially localize symmetrically and without cross-talk, between left and right, the N separated stereo audio signals into multiple output channels. The output channels include respective mixtures of one or more of the N separated stereo audio signals. Gain is adjusted of the output channels into left and right binaural outputs to conserve summed levels of the N separated stereo audio signals distributed over the output channels.