Audiovisual Source Separation Using GAN Optical Flow
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audiovisual source separation systems, such as PixelPlayer, fail to produce natural-sounding output and do not effectively utilize temporal information from video frames, leading to suboptimal sound localization and separation.
Innovation Solution
A method utilizing a Generative Adversarial Network (GAN) with Deep Neural Networks (DNNs) that incorporates optical flow data to localize sound sources on video pixels, enhancing the separation process by generating natural-sounding outputs through the integration of video frame data and motion information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional audiovisual source separation systems are used, then sound sources can be separated into components, but the output does not sound natural enough
Solution Approach 1:
The patent introduces an intermediary representation called 'pixel-level sound sources' that acts as a bridge between visual video data and audio separation. Each pixel is associated with a sound source vector, serving as an intermediary that connects spatial visual information with temporal audio information, enabling more natural-sounding separation while maintaining accuracy
Solution Approach 2:
The patent adds a spatial dimension to traditional audio separation by mapping sound sources to individual pixels in video frames. This transforms the separation problem from purely temporal audio processing to a spatio-temporal problem, where each pixel contributes to the overall sound separation, improving both naturalness and precision
2Manufacturing precision
If visual cues from video frames are utilized, then separation performance is improved, but temporal information from motion is not effectively used
Solution Approach 1:
The patent merges visual video data with optical flow data in a unified processing framework. Both video frames and optical flow sequences are fed into the network simultaneously, allowing the system to leverage both spatial visual cues and temporal motion information together, preventing loss of motion information while maintaining separation performance
Solution Approach 2:
The patent performs preliminary processing of optical flow data to extract motion-related features before combining them with visual cues. By pre-processing the optical flow to identify motion patterns and temporal changes, the system prepares motion information in advance for effective integration with visual data, ensuring no temporal information is lost
Data Source
AI summary
A method (and structure and computer product) for an audiovisual source separation processing, including receiving video data including images of a plurality of sound sources, receiving an optical flow data of the video data, the optical flow data indicating motions of pixels between frames of the video data, and encoding the received video data into video localization data comprising information associating pixels in the frames of video data with different channels of sound.


