Robotic Long-Term Perception With Spatio-Temporal Memory Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robotic perception systems fail to capture dynamic changes in environments and have short spatio-temporal memory durations, limiting their ability to recall key events and perceive objects not in predefined lists, especially in long-term deployments.

Innovation Solution

A robot generates a stream of sensor data into segments, using machine learning models to create embedded representations stored as memories, and retrieves relevant information through iterative querying to answer questions and perform tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional metric maps and scene graphs are used for environmental representation, then static elements are captured, but dynamic changes to objects and phenomena cannot be captured

Engineering Contradiction:
Improveaccuracy of environmental representationVSAvoidability to capture dynamic changes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transitions from static metric maps and scene graphs to dynamic spatio-temporal memory representations that continuously update to reflect changing environmental conditions. The system captures temporal evolution of objects and phenomena, enabling the robot to track dynamic changes while maintaining spatial accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent combines multiple data types (sensor readings, object states, temporal information, spatial coordinates) into a composite spatio-temporal memory structure. This composite representation integrates both static and dynamic elements, achieving comprehensive environmental modeling that surpasses conventional single-type representations.

Inventive Principle:
Principle #40Composite materials

2Loss of time

If conventional spatio-temporal robotic memory is used, then short-term recall is possible, but long-term deployment performance degrades due to limited memory duration

Engineering Contradiction:
Improvememory durationVSAvoidtask performance over time
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent implements continuous memory consolidation and periodic summarization of environmental observations. By preemptively organizing and compressing sensory data into structured spatio-temporal memories during deployment, the system ensures long-term recall capability without degrading performance, allowing robots to operate effectively over hours to days.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manually prescribed object lists are used for perception, then predefined objects can be detected, but objects not in the list cannot be perceived

Engineering Contradiction:
Improveobject detection accuracyVSAvoidperception of unknown objects
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a dual-mode perception system that handles both known and unknown objects. The spatio-temporal memory structure stores detailed sensory observations for all detected objects, enabling the robot to recognize predefined objects with high accuracy while also capturing and reasoning about novel objects through their visual and contextual features, making the system universally applicable to diverse object types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260042204A1Long-term perception for robotics systems and applications
Publication Date: 2026.02.12 NVIDIA CORP
  • US20260042204A1 patent drawing
  • US20260042204A1 patent drawing
  • US20260042204A1 patent drawing

AI summary

In various examples, a technique for performing a task includes converting one or more sensory inputs obtained using one or more sensors of a machine into a plurality of segments. The technique also includes, for each segment included in the plurality of segments, generating, via execution of a machine learning model, a caption for the segment, and storing, in a data store, a representation of the caption in association with the segment. The technique further includes performing, by the machine, one or more actions based at least on one or more queries of the data store.