Sound Image Object Extraction Using Cross-Modal Association

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques fail to effectively extract and isolate a desired object's image area and sound from a moving image with sound, particularly when multiple objects emit sound in the same direction, making it difficult to focus on specific sounds or objects based on their characteristics.

Innovation Solution

The technology detects sound and image objects from a moving image with sound by using sound source separation and image information, associating the detected objects to extract the sound image object, which allows for the isolation of the desired object's sound and image, even in challenging conditions like darkness or unclear subjects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If sound source separation is performed based on sound direction alone, then sound extraction is simplified, but multiple objects in the same direction cannot be distinguished

Engineering Contradiction:
Improvesound extraction processVSAvoidobject identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines sound source separation results with image object detection results to identify target objects. When multiple sound sources exist in the same direction, the system cross-references image data to distinguish between different objects, merging audio and visual information streams to achieve accurate object identification that neither modality could accomplish alone.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If object selection is based on image position only, then selection process is simplified, but objects cannot be selected by semantic concepts like person or car

Engineering Contradiction:
Improveobject selection processVSAvoidselection flexibility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system provides multiple object selection methods including position-based selection, voice-based semantic selection, and automatic selection. Users can select objects by pointing at the screen, by speaking commands like 'focus on the car,' or by automatic detection, making the system adaptable to different user needs and contexts while maintaining ease of operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If sound objects are detected without image information, then audio processing is independent, but sound objects cannot be reliably associated with visual objects

Engineering Contradiction:
Improveprocessing systemVSAvoidobject association accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent uses image objects as an intermediary to bridge sound objects and visual objects. The system detects both sound objects and image objects separately, then uses the image object information as a mediator to associate the two, enabling reliable matching between audio and visual data even when sound direction alone is insufficient for identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If voice recognition is implemented without object definition, then voice input is simple, but commands like 'focus on girl in red shirt' cannot be processed

Engineering Contradiction:
Improvevoice inputVSAvoidobject identification
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary object detection and classification before processing voice commands. Image objects are detected and categorized in advance (e.g., identifying a girl in a red shirt), creating a prepared structure that enables subsequent voice recognition to efficiently match user commands with predefined object categories, reducing the difficulty of both voice input and object identification.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3829161B1Information processing device and method, and program
Publication Date: 2023.08.30 SONY GROUP CORP
  • EP3829161B1 patent drawingFigure 1
  • EP3829161B1 patent drawingFigure 2
  • EP3829161B1 patent drawingFigure 3

AI summary

The present technology relates to an information processing device and method, and a program that enable extraction of a desired object from a moving image with sound. An information processing device includes an image object detection unit that detects an image object on the basis of a moving image with sound, a sound object detection unit that detects a sound object on the basis of the moving image with sound, and a sound image object detection unit that detects a sound image object on the basis of a detection result of the image object and a detection result of the sound object. The present technology can be applied to an information processing device.