Multi-Modal Audio Source Channelization for Overlapping Speakers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio source separation technologies struggle to accurately and efficiently isolate multiple target audio sources from complex, dynamic environments with overlapping audio sources such as noise, music, and reverberations, often requiring manual supervision and consuming significant computational resources.

Innovation Solution

A multi-modal audio source channelization system that utilizes a trained model to generate source-separated channel audio samples by combining audio and video signal features, reducing the need for manual supervision and computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing audio source separation technologies are used to isolate multiple target audio sources from complex environments, then the separation accuracy can be improved, but the computational resources required increase significantly

Engineering Contradiction:
Improveseparation accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system segments the audio separation task by processing audio signals in overlapping frames with hop intervals, dividing the continuous processing into discrete manageable units. This allows efficient computational resource utilization while maintaining separation accuracy across the entire audio signal.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary action by pre-processing audio signals through framing and windowing operations before separation. The audio signal is divided into frames with overlapping windows, and padding is applied to handle edge cases, preparing the data structure for efficient separation processing.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If existing audio source separation technologies are used to isolate multiple target audio sources, then the separation quality can be improved, but manual supervision requirements increase

Engineering Contradiction:
Improveseparation qualityVSAvoidmanual supervision
Core Design Contradiction:
Manufacturing precisionVSExtent of automation

Solution Approach 1:

The system implements self-service through automated speaker diarization that identifies and labels different speakers without manual intervention. The model automatically processes audio frames, separates sources, and generates labeled outputs, eliminating the need for manual supervision while maintaining high separation quality.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If complex processing methods are used to separate audio sources in dynamic environments, then the separation accuracy can be improved, but the processing time increases

Engineering Contradiction:
Improveseparation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies periodic action by processing audio signals in regular frames with consistent hop intervals. This periodic processing approach allows the model to efficiently handle dynamic environments through structured, time-discretized processing while maintaining accuracy through the consistent application of separation operations.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20250259639A1Audio source separation using multi-modal audio source channalization system
Publication Date: 2025.08.14 SHURE ACQUISITION HLDG INC
  • US20250259639A1 patent drawing
  • US20250259639A1 patent drawing
  • US20250259639A1 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatuses, systems, and/or devices that are configured to separate multi-source audio signal samples into discrete source separated channel audio samples using trained multi-modal audio source channelization models. Multi-source audio signal samples are difficult to separate because they can include multiple target audio sources (e.g., individual speakers) that are often inter-mixed and overlaid with other audio sources such as noise, music, reverberations, and other audio artifacts. The multi-modal audio source channelization models discussed herein are trained to generate source separated channel audio samples based on audio signal samples and on video signal samples.