Video Interaction Graphs for Context-Aware Scene Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video analysis systems struggle with identifying relevant information across multiple frames due to lack of context and causal relationships, requiring significant computing resources and causing high latency in information retrieval.

Innovation Solution

Generating interaction graphs that represent entities and their interactions in videos, allowing for efficient information retrieval by filtering based on query context and timestamps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If descriptions are generated for each frame of videos, then information retrieval can be performed, but the descriptions lack context and causal relationships between frames reducing retrieval quality

Engineering Contradiction:
Improveretrieval qualityVSAvoidcontext information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent merges multiple frame descriptions into a unified video-level description that captures context and causal relationships across frames. Instead of treating each frame independently, the system combines temporal and contextual information from multiple frames to generate a comprehensive representation that improves retrieval quality while preserving essential context.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from a single-frame dimensional view to a multi-frame temporal dimension by generating descriptions that span multiple frames. This dimensional change allows the system to capture causal relationships and contextual information that exist across time, enabling more accurate information retrieval.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If descriptions are generated for each frame of videos, then information can be stored for retrieval, but large amounts of computing resources and memory are required

Engineering Contradiction:
Improveinformation retrieval capabilityVSAvoidmemory resources
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent combines multiple frame descriptions into a single video-level representation, significantly reducing the total quantity of data that needs to be stored. By merging redundant and overlapping information from individual frames into a consolidated video description, the system maintains retrieval capability while reducing memory requirements.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a compressed abstract representation (video-level description) that serves as a efficient copy or summary of the full frame-by-frame data. This copied representation retains the essential information needed for retrieval while occupying significantly less memory space than storing complete frame descriptions.

Inventive Principle:
Principle #26Copying

3Loss of information

If descriptions are generated for each frame of videos, then comprehensive information is available, but large latencies occur when searching through stored data

Engineering Contradiction:
Improveinformation completenessVSAvoidsearch latency
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts the most salient and relevant information from multiple frames to create a condensed video-level description. By taking out only the essential contextual and causal information needed for retrieval rather than storing all frame details, the system reduces search space and latency while maintaining information completeness for retrieval purposes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent organizes information in a temporal dimension across frames, allowing the system to search and retrieve based on video-level context rather than individual frames. This dimensional organization enables more efficient searching by leveraging temporal relationships and reducing the search space from O(n frames) to O(1 video-level description).

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20260080659A1Scene graphs for video scene understanding and information retrieval
Publication Date: 2026.03.19 NVIDIA CORP
  • US20260080659A1 patent drawing
  • US20260080659A1 patent drawing
  • US20260080659A1 patent drawing

AI summary

In various examples, generating and using interaction graphs for video information retrieval systems and applications is described herein. Systems and methods are disclosed that process videos generated using one or more image sensors in order to generate a graph that represents at least interactions between entities depicted by the videos. For instance, nodes of the graph may be associated with the entities—such as people and/or other objects—as well as attributes associated with the entities. Additionally, edges of the graph may be associated with interactions between the entities, times that the interactions occurred, and/or indications of which videos depict the interactions. Systems and methods are then further disclosed that use the graph to perform information retrieval associated with the videos. For instance, the graph may be used to identify relevant information associated with a query, where the information may then be used to generate a response.