Audio-Visual Correspondence for Sound Source Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio-visual systems in augmented and virtual reality environments face challenges in accurately localizing sound sources due to the lack of effective correspondence analysis between visual and audio data, which limits the precision of sound source identification and spatialized audio presentation.

Innovation Solution

A method and system that perform correspondence analysis between video and audio content by identifying pixels with changes in pixel values and associating specific audio with those pixels, using a controller that captures video and audio data and applies beamforming and semantic segmentation to enhance sound source localization, allowing for accurate attribution of audio to visual objects and environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If correspondence analysis between video and audio content is performed to identify sound sources, then sound source localization precision is improved, but device complexity increases due to the need for pixel tracking and audio-visual synchronization systems

Engineering Contradiction:
Improvesound source localization precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the video content into individual pixel tracks and associates each pixel with audio characteristics. By dividing the visual content into discrete pixel elements and tracking their movement independently, the system can perform correspondence analysis at a granular level, improving sound source localization precision while managing computational complexity through structured data organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary correspondence analysis mechanism that bridges video content and audio content. This intermediary process analyzes the relationship between pixel movements and audio signals to identify sound sources, acting as a mediator that connects visual and auditory data streams without requiring direct integration of complex audio-visual synchronization systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If pixel value changes are tracked to identify sound sources, then sound source identification accuracy is improved, but loss of information increases due to the selective filtering of pixel data

Engineering Contradiction:
Improvesound source identification accuracyVSAvoidloss of information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system applies local quality analysis by focusing computational resources on specific pixels that exhibit value changes, rather than processing all pixels uniformly. By identifying and analyzing only the pixels with significant value changes, the system improves sound source identification accuracy while minimizing information loss by preserving the characteristics of the selected pixels and discarding only the irrelevant data.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If audio is assigned to specific pixels based on correspondence analysis, then spatialized audio presentation is improved, but processing time increases due to the computational requirements of analyzing pixel-audio relationships

Engineering Contradiction:
Improvespatialized audio presentation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing video content to identify pixels with value changes before conducting the full correspondence analysis with audio. By preparing and filtering pixel data in advance, the system reduces the computational burden during the actual audio-visual matching process, thereby decreasing processing time while maintaining spatialized audio presentation accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11234090B2Using audio visual correspondence for sound source identification
Publication Date: 2022.01.25 META PLATFORMS TECHNOLOGIES LLC
  • US11234090B2 patent drawing
  • US11234090B2 patent drawing
  • US11234090B2 patent drawing

AI summary

An audio-video correspondence system performs a correspondence analysis between audio and video content. The system obtains video content that includes audio content, the video content comprising a plurality of frames. The system identifies a first set of pixels within the video content associated with changes in pixel values over a first set of frames of the video content. The system subsequently identifies, from the audio content, first audio associated with the first set of frames corresponding to changes in pixel values of the first set of pixels. The system finally assigns the first audio over the first set of frames to the first set of pixels.