Scene Retrieval Platform Using Vector Representations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for managing and retrieving visual information, such as videos, are inefficient and require significant user intervention, making it difficult to identify relevant images or videos from large databases, especially in applications like sporting events where analysts need to find similar scenes.

Innovation Solution

A scene retrieval platform that uses a neural network trained with unsupervised learning to generate vector representations of scenes, allowing for the identification and retrieval of similar scenes based on user input, by accessing and processing timestamped coordinate data from video frames, and ranking similar scenes in a vector space.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual searching and user intervention are used to retrieve visual information, then retrieval accuracy can be maintained, but the time and effort required increases significantly

Engineering Contradiction:
Improveretrieval accuracyVSAvoidsearching time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating metadata, tags, and vector representations for visual content during ingestion, rather than requiring manual annotation at retrieval time. This pre-processing enables rapid similarity search without compromising accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces vector representations and metadata as intermediary elements between the visual content and the search query. These intermediaries enable automated similarity comparison without requiring direct manual intervention, thus reducing time while maintaining retrieval quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated retrieval systems are implemented, then searching time is reduced, but the complexity of the system increases

Engineering Contradiction:
Improveretrieval efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical processes (human analysts watching and categorizing video content) with automated computational processes including vector representation, metadata generation, and similarity algorithms. This substitution dramatically improves productivity while the modular architecture manages complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms visual content into different parameter spaces (vector representations, metadata attributes) that enable automated comparison and search. By changing the representation parameters rather than maintaining raw visual data for comparison, the system achieves efficient automated retrieval.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive metadata and tags are generated for all visual content, then retrieval accuracy improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveretrieval accuracyVSAvoidprocessing resources
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The system generates metadata and vector representations selectively based on the needs of similarity search, rather than creating all possible types of annotations for every piece of content. This partial action approach maintains sufficient accuracy while reducing unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10783377B2Visually similar scene retrieval using coordinate data
Publication Date: 2020.09.22 SAP SE
  • US10783377B2 patent drawing
  • US10783377B2 patent drawing
  • US10783377B2 patent drawing

AI summary

Aspects of the present disclosure therefore involve systems and methods for identifying a set of visually similar scenes to a target scene selected or otherwise identified by a match analyst. A scene retrieval platform performs operations for: receiving an input that comprises an identification of a scene; retrieving a set of coordinates based on the scene identified by the input, where the set of coordinates identify positions of the entities depicted within the frames; generating a set of vector values based on the coordinates of the entities depicted within each of the frames; concatenating the set of vector values to generate a concatenated vector value that represents the scene; generating a visual representation of the concatenated vector value; and identifying one or more similar scenes to the scene identified by the input based on the visual representation of the concatenated vector value.