Audio-Visual Object Retrieval via Context Keyword Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Object retrieval from visual data in real-world applications is challenging due to varying object poses, illumination, occlusion, and the need to search large amounts of video data, with unclear time periods for queries like 'Where did I put my car keys?'.

Innovation Solution

Integrating audio cues with video frames by annotating visual data with context keywords recognized from audio data, using learned appearance models and query keywords to identify and search for objects within specific timestamps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio cues are integrated with video frames to annotate visual data with context keywords, then object retrieval accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveobject retrieval accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines audio data and visual data into a unified annotated dataset where audio cues (context keywords) are merged with corresponding video frames. This integration allows the system to leverage both audio and visual information simultaneously, improving object retrieval accuracy while managing complexity through unified processing pipelines.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces context keywords as an intermediary element that bridges audio data and visual data. These keywords serve as mediators that connect audio cues to relevant video frames, enabling the system to retrieve objects more accurately without directly complexly integrating all audio-visual data processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the entire video dataset is searched for object retrieval, then retrieval completeness is improved, but processing time increases

Engineering Contradiction:
Improveretrieval completenessVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and isolates relevant video frames by annotating them with context keywords derived from audio data. Instead of searching the entire video dataset, the system extracts only those frames associated with specific audio cues, maintaining retrieval completeness while dramatically reducing processing time through selective frame identification.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary annotation of video frames with context keywords before the actual object retrieval process. By pre-identifying and marking relevant frames based on audio cues, the system prepares the data in advance, allowing for faster retrieval operations without compromising completeness during the actual search phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10108617B2Using audio cues to improve object retrieval in video
Publication Date: 2018.10.23 TEXAS INSTRUMENTS INC
  • US10108617B2 patent drawing
  • US10108617B2 patent drawing
  • US10108617B2 patent drawing

AI summary

A method of object retrieval from visual data is provided that includes annotating at least one portion of the visual data with a context keyword corresponding to an object, wherein the annotating is performed responsive to recognition of the context keyword in audio data corresponding to the at least one portion of the visual data, receiving a query to retrieve the object, wherein the query includes a query keyword associated with both the object and the context keyword, identifying the at least one portion of the visual data based on the context keyword, and searching for the object in the at least one portion of the visual data using an appearance model corresponding to the query keyword.