Transformer Network for Object-Centric Video Reasoning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video processing systems struggle with high-level spatio-temporal reasoning, predictive, counterfactual, and causal reasoning, particularly in understanding object dynamics and permanence, often requiring hand-engineered neuro-symbolic approaches and large amounts of labeled data.

Innovation Solution

A video processing neural network system that employs an object segmentation subsystem for unsupervised learning, combining disentangled object representations with video frame position encoding, and a transformer neural network with attention layers to process sequences of video frames and respond to queries, enabling efficient handling of queries requiring predictive, counterfactual, and causal reasoning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If hand-engineered neuro-symbolic approaches are used for high-level spatio-temporal reasoning, then reasoning capability is improved, but system complexity and data requirements increase

Engineering Contradiction:
Improvereasoning capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces hand-engineered neuro-symbolic approaches with an end-to-end differentiable physics simulation system. The mechanical system of manually designed reasoning modules is substituted with a learned continuous relaxation of physics simulations that can be trained using standard gradient-based optimization, thereby reducing system complexity while maintaining reasoning capabilities.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces continuous relaxation parameters that allow discrete physics simulation operations to be differentiable. By changing the parameterization of physics operations from discrete to continuous, the system enables gradient flow through physics simulations, allowing end-to-end training without hand-engineered reasoning components.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If supervised learning with large labeled data sets is used for object-centric video reasoning, then model accuracy is improved, but training time and computational requirements increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent enables the system to learn from unlabeled video data through self-supervised learning. The differentiable physics simulation framework allows the model to automatically learn object dynamics and interactions without requiring manual annotation, thereby reducing dependency on large labeled datasets and decreasing training time.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a universal framework that can handle multiple reasoning tasks (spatio-temporal reasoning, predictive reasoning, counterfactual reasoning, causal reasoning) using a single end-to-end trainable system. This multi-functional approach eliminates the need for separate supervised learning pipelines for each task, reducing overall training time and computational requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If supervised learning with large labeled data sets is used for object-centric video reasoning, then model accuracy is improved, but computational requirements increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational requirements
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent replaces computationally intensive supervised learning pipelines with a differentiable physics simulation framework that uses standard gradient-based optimization. This substitution reduces computational requirements by eliminating the need for separate training pipelines for each reasoning task and leveraging efficient automatic differentiation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the optimization landscape from discrete, non-differentiable operations to continuous, differentiable operations. This parameter change enables the use of efficient gradient-based optimizers rather than requiring complex supervised learning infrastructure, thereby reducing computational requirements while maintaining model accuracy.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If discrete physics simulation operations are used for object dynamics modeling, then physical accuracy is improved, but differentiability and trainability decrease

Engineering Contradiction:
Improvephysical accuracyVSAvoiddifferentiability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies continuous relaxation to discrete physics simulation operations, transforming them into differentiable counterparts. By changing the parameterization from discrete to continuous, the system maintains physical accuracy while enabling gradient flow through the physics simulation operations, making the entire pipeline trainable with standard optimization methods.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces continuous relaxation as an intermediary layer between discrete physics operations and the neural network. This intermediary enables the translation of discrete physical concepts into continuous, differentiable representations that can be optimized using gradient-based methods, bridging the gap between physical accuracy and trainability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240020972A1Neural networks implementing attention over object embeddings for object-centric visual reasoning
Publication Date: 2024.01.18 GDM HOLDING LLC
  • US20240020972A1 patent drawing
  • US20240020972A1 patent drawing
  • US20240020972A1 patent drawing

AI summary

A video processing system configured to analyze a sequence of video frames to detect objects in the video frames and provide information relating to the detected objects in response to a query. The query may comprise, for example, a request for a prediction of a future event, or of the location of an object, or a request for a prediction of what would happen if an object were modified. The system uses a transformer neural network subsystem to process representations of objects in the video.