Sound Image Object Extraction Using Cross-Modal Association
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to effectively extract and isolate a desired object's image area and sound from a moving image with sound, particularly when multiple objects emit sound in the same direction, making it difficult to focus on specific sounds or objects based on their characteristics.
Innovation Solution
The technology detects sound and image objects from a moving image with sound by using sound source separation and image information, associating the detected objects to extract the sound image object, which allows for the isolation of the desired object's sound and image, even in challenging conditions like darkness or unclear subjects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If sound source separation is performed based on sound direction alone, then sound extraction is simplified, but multiple objects in the same direction cannot be distinguished
Solution Approach 1:
The patent combines sound source separation results with image object detection results to identify target objects. When multiple sound sources exist in the same direction, the system cross-references image data to distinguish between different objects, merging audio and visual information streams to achieve accurate object identification that neither modality could accomplish alone.
2Ease of operation
If object selection is based on image position only, then selection process is simplified, but objects cannot be selected by semantic concepts like person or car
Solution Approach 1:
The system provides multiple object selection methods including position-based selection, voice-based semantic selection, and automatic selection. Users can select objects by pointing at the screen, by speaking commands like 'focus on the car,' or by automatic detection, making the system adaptable to different user needs and contexts while maintaining ease of operation.
3Device complexity
If sound objects are detected without image information, then audio processing is independent, but sound objects cannot be reliably associated with visual objects
Solution Approach 1:
The patent uses image objects as an intermediary to bridge sound objects and visual objects. The system detects both sound objects and image objects separately, then uses the image object information as a mediator to associate the two, enabling reliable matching between audio and visual data even when sound direction alone is insufficient for identification.
4Ease of operation
If voice recognition is implemented without object definition, then voice input is simple, but commands like 'focus on girl in red shirt' cannot be processed
Solution Approach 1:
The system performs preliminary object detection and classification before processing voice commands. Image objects are detected and categorized in advance (e.g., identifying a girl in a red shirt), creating a prepared structure that enables subsequent voice recognition to efficiently match user commands with predefined object categories, reducing the difficulty of both voice input and object identification.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present technology relates to an information processing device and method, and a program that enable extraction of a desired object from a moving image with sound. An information processing device includes an image object detection unit that detects an image object on the basis of a moving image with sound, a sound object detection unit that detects a sound object on the basis of the moving image with sound, and a sound image object detection unit that detects a sound image object on the basis of a detection result of the image object and a detection result of the sound object. The present technology can be applied to an information processing device.