Egocentric Audio-Visual Object Localization via Geometry-Aware Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Egocentric audio-visual object localization in videos is challenging due to egomotion and out-of-view sounds, which cause difficulties in associating visual content with audio representations and localizing sounding objects accurately in dynamic and limited field-of-view scenarios.

Innovation Solution

A geometry-aware temporal context aggregation module and a cascaded feature enhancement module are proposed to handle egomotion and out-of-view sounds, respectively, using audio-visual temporal synchronization for self-supervised training and disentangling visually indicated audio representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio-visual object localization is performed in egocentric videos, then object localization capability is enabled, but egomotion causes visual content to change making accurate localization difficult

Engineering Contradiction:
Improvelocalization precisionVSAvoidvisual content stability
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

The system dynamically adapts to egomotion by continuously updating visual features across video frames and adjusting the association between audio and visual content based on changing perspectives. The localization model incorporates temporal dynamics to track objects despite viewpoint changes caused by wearer movement.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces an audio-visual association module that acts as an intermediary to bridge the gap between audio representations and visual content. This intermediary component correlates audio signals with visual features to disambiguate objects, especially when visual information alone is insufficient due to egomotion-induced changes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the field of view is limited in egocentric videos, then wearable device comfort is improved, but out-of-view sounds cannot be localized

Engineering Contradiction:
Improvewearable comfortVSAvoidout-of-view sound information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The localization system handles both in-view and out-of-view sound sources using the same audio-visual association framework. The system is designed to process audio signals regardless of whether the sound source is visible in the current frame, utilizing temporal context and audio characteristics to localize sounds from any direction within the wearer's environment.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system extends localization beyond the two-dimensional visual frame by incorporating spatial audio information and temporal context. It uses audio representations that encode directional and spatial cues, allowing the system to localize sound sources in three-dimensional space even when they fall outside the limited camera field of view.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If audio-visual association is performed without geometry awareness, then processing speed is improved, but egomotion cannot be handled

Engineering Contradiction:
Improveprocessing speedVSAvoidegomotion handling capability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary extraction of visual features and audio representations from the input video and audio data. By pre-processing and organizing this information before the association step, the system maintains processing efficiency while preparing the data structures needed for geometry-aware egomotion handling in subsequent processing stages.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240305944A1First-person audio-visual object localization systems and methods
Publication Date: 2024.09.12 UNIVERSITY OF ROCHESTER
  • US20240305944A1 patent drawing
  • US20240305944A1 patent drawing
  • US20240305944A1 patent drawing

AI summary

A localization system may include an image input that receives images from a video source and an audio input that receives, from the video source, audio synchronized with the images. The localization system may also include an audio feature disentanglement network that correlates distinct audio elements from the audio input with corresponding visual features from the image input. Additionally, the localization system may include a geometry-based feature aggregation module that estimates a geometric transformation between two or more images from the video source and aggregates the visual features. Various other devices, systems, and methods are also disclosed.