Audiovisual Source Separation Using GAN Optical Flow

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audiovisual source separation systems, such as PixelPlayer, fail to produce natural-sounding output and do not effectively utilize temporal information from video frames, leading to suboptimal sound localization and separation.

Innovation Solution

A method utilizing a Generative Adversarial Network (GAN) with Deep Neural Networks (DNNs) that incorporates optical flow data to localize sound sources on video pixels, enhancing the separation process by generating natural-sounding outputs through the integration of video frame data and motion information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional audiovisual source separation systems are used, then sound sources can be separated into components, but the output does not sound natural enough

Engineering Contradiction:
Improvenaturalness of output soundVSAvoidseparation accuracy
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary representation called 'pixel-level sound sources' that acts as a bridge between visual video data and audio separation. Each pixel is associated with a sound source vector, serving as an intermediary that connects spatial visual information with temporal audio information, enabling more natural-sounding separation while maintaining accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent adds a spatial dimension to traditional audio separation by mapping sound sources to individual pixels in video frames. This transforms the separation problem from purely temporal audio processing to a spatio-temporal problem, where each pixel contributes to the overall sound separation, improving both naturalness and precision

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If visual cues from video frames are utilized, then separation performance is improved, but temporal information from motion is not effectively used

Engineering Contradiction:
Improveseparation performanceVSAvoidtemporal motion information
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent merges visual video data with optical flow data in a unified processing framework. Both video frames and optical flow sequences are fed into the network simultaneously, allowing the system to leverage both spatial visual cues and temporal motion information together, preventing loss of motion information while maintaining separation performance

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary processing of optical flow data to extract motion-related features before combining them with visual cues. By pre-processing the optical flow to identify motion patterns and temporal changes, the system prepares motion information in advance for effective integration with visual data, ensuring no temporal information is lost

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12079731B2Audiovisual source separation and localization using generative adversarial networks
Publication Date: 2024.09.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12079731B2 patent drawing
  • US12079731B2 patent drawing
  • US12079731B2 patent drawing

AI summary

A method (and structure and computer product) for an audiovisual source separation processing, including receiving video data including images of a plurality of sound sources, receiving an optical flow data of the video data, the optical flow data indicating motions of pixels between frames of the video data, and encoding the received video data into video localization data comprising information associating pixels in the frames of video data with different channels of sound.