AI Video Audio Source Separation with Mouth Movement Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing technologies fail to effectively separate multiple sound sources from mixed audio signals, especially in diverse environments with overlapping voices, leading to reduced separation performance and user interaction limitations.
Innovation Solution
An electronic device employing AI models to generate audio-related information from image and audio signals, using a first AI model to analyze mouth movement information and a second AI model with an encoder and bottleneck layer to separate sound sources, along with number-of-speakers related information for improved separation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple sound sources are separated from mixed audio signals using conventional methods, then the separation process can be performed, but the separation performance deteriorates in environments with overlapping voices
Solution Approach 1:
The patent segments the sound source separation task into multiple processing stages: generating initial separation results, detecting overlapping voice regions, and performing targeted refinement. The audio signal is divided into multiple frames, with overlapping regions identified and processed separately to improve overall separation accuracy in complex acoustic environments.
Solution Approach 2:
The patent introduces an intermediary overlapping voice detection module that bridges the initial separation process and final refinement. This detector identifies regions where multiple voices overlap, allowing the system to apply specialized processing to these problematic areas while leaving clear regions unchanged, thus improving overall separation performance.
2Measurement precision
If AI models are used to generate audio-related information and separate sound sources, then separation accuracy is improved, but computational complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-generating audio-related information such as spectrograms and voice activity detection results before the actual separation process. Mouth movement information from video frames is also extracted in advance, providing pre-processed features that reduce the computational burden during the main separation operation.
Solution Approach 2:
The patent applies partial processing by focusing computational resources only on regions identified as having overlapping voices. The overlapping voice detector identifies specific time-frequency regions that require refinement, allowing the system to apply complex AI-based separation only where necessary rather than processing the entire audio signal uniformly.
3Measurement precision
If audio-related information including mouth movement information is generated and applied to the separation model, then separation performance in overlapping environments is enhanced, but processing time increases
Solution Approach 1:
The patent merges audio signal processing with visual information processing by combining mouth movement information from video frames with audio spectrograms. This multi-modal fusion allows the system to leverage visual cues about speaker positioning and lip movements to enhance audio separation, particularly in overlapping voice scenarios where auditory information alone is insufficient.
Solution Approach 2:
The system performs partial processing by applying the comprehensive AI-based separation model only to identified overlapping voice regions, while leaving non-overlapping regions to be processed by simpler methods or left unchanged, thus reducing overall processing time while maintaining high accuracy where needed.
Data Source
AI summary
An electronic device for processing a video including an image signal and a mixed audio signal, includes: a memory configured to store at least one program for processing the video; and at least one processor configured to: generate, from the image signal and the mixed audio signal, audio-related information indicating a degree of overlap in a plurality of sound sources included in the mixed audio signal by using a first artificial intelligence (AI) model; and separate at least one of the plurality of sound sources included in the mixed audio signal from the mixed audio signal, by applying the audio-related information to a second AI model.


