AI Video Audio Source Separation with Mouth Movement Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video processing technologies fail to effectively separate multiple sound sources from mixed audio signals, especially in diverse environments with overlapping voices, leading to reduced separation performance and user interaction limitations.

Innovation Solution

An electronic device employing AI models to generate audio-related information from image and audio signals, using a first AI model to analyze mouth movement information and a second AI model with an encoder and bottleneck layer to separate sound sources, along with number-of-speakers related information for improved separation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple sound sources are separated from mixed audio signals using conventional methods, then the separation process can be performed, but the separation performance deteriorates in environments with overlapping voices

Engineering Contradiction:
Improvesound source separation performanceVSAvoidoverlapping voices interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the sound source separation task into multiple processing stages: generating initial separation results, detecting overlapping voice regions, and performing targeted refinement. The audio signal is divided into multiple frames, with overlapping regions identified and processed separately to improve overall separation accuracy in complex acoustic environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary overlapping voice detection module that bridges the initial separation process and final refinement. This detector identifies regions where multiple voices overlap, allowing the system to apply specialized processing to these problematic areas while leaving clear regions unchanged, thus improving overall separation performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If AI models are used to generate audio-related information and separate sound sources, then separation accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improveseparation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-generating audio-related information such as spectrograms and voice activity detection results before the actual separation process. Mouth movement information from video frames is also extracted in advance, providing pre-processed features that reduce the computational burden during the main separation operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial processing by focusing computational resources only on regions identified as having overlapping voices. The overlapping voice detector identifies specific time-frequency regions that require refinement, allowing the system to apply complex AI-based separation only where necessary rather than processing the entire audio signal uniformly.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If audio-related information including mouth movement information is generated and applied to the separation model, then separation performance in overlapping environments is enhanced, but processing time increases

Engineering Contradiction:
Improveseparation performanceVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges audio signal processing with visual information processing by combining mouth movement information from video frames with audio spectrograms. This multi-modal fusion allows the system to leverage visual cues about speaker positioning and lip movements to enhance audio separation, particularly in overlapping voice scenarios where auditory information alone is insufficient.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs partial processing by applying the comprehensive AI-based separation model only to identified overlapping voice regions, while leaving non-overlapping regions to be processed by simpler methods or left unchanged, thus reducing overall processing time while maintaining high accuracy where needed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240127847A1Apparatus for processing video, and operation method of the apparatus
Publication Date: 2024.04.18 SAMSUNG ELECTRONICS CO LTD
  • US20240127847A1 patent drawing
  • US20240127847A1 patent drawing
  • US20240127847A1 patent drawing

AI summary

An electronic device for processing a video including an image signal and a mixed audio signal, includes: a memory configured to store at least one program for processing the video; and at least one processor configured to: generate, from the image signal and the mixed audio signal, audio-related information indicating a degree of overlap in a plurality of sound sources included in the mixed audio signal by using a first artificial intelligence (AI) model; and separate at least one of the plurality of sound sources included in the mixed audio signal from the mixed audio signal, by applying the audio-related information to a second AI model.