Embodied AR Agent Spatial Understanding for Natural Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is a lack of interactive agents with spatial understanding in mixed reality and spatial computing environments, limiting user interaction capabilities.

Innovation Solution

Systems and methods provide intelligent embodied interactive agents with spatial understanding by integrating conversational AI engines with augmented reality headsets, utilizing large and visual language models to generate animations and speech based on user inputs, location information, and scene awareness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional computing interfaces are used in mixed reality environments, then device complexity is reduced, but spatial understanding and interactive capabilities are limited

Engineering Contradiction:
Improvespatial understanding capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a conversational AI engine as an intermediary layer between the user and the spatial computing environment. This engine processes natural language queries, interprets spatial context from camera feeds and sensor data, and generates appropriate responses with animations and speech. The intermediary handles the complexity of spatial understanding internally while presenting a simple natural language interface to users, thus improving adaptability without increasing perceived device complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical or manual interaction mechanisms (buttons, switches, direct manipulation) with voice-based conversational interfaces. Users can query the spatial environment naturally without physically interacting with complex controls. The system substitutes mechanical interaction complexity with automated AI processing, maintaining simplicity in user operation while achieving sophisticated spatial understanding.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If multiple input modalities (audio, video, location) are integrated for spatial understanding, then interaction accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespatial perception accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of spatial data by continuously capturing and pre-processing camera feeds, location information, and environmental data in the background before user queries are submitted. The conversational AI engine maintains ready-state spatial context models that can be quickly queried without requiring real-time processing of all sensor data from scratch. This preliminary preparation reduces response time while maintaining high spatial perception accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the processing of multiple input modalities into separate specialized modules: audio processing for voice queries, visual processing for camera feeds and object recognition, location processing for spatial positioning, and contextual processing for task inference. Each module processes its specific data type independently and efficiently, then integrates results in the conversational AI engine. This segmentation prevents bottlenecks and reduces overall processing time while maintaining comprehensive spatial understanding.

Inventive Principle:
Principle #1Segmentation

3Extent of automation

If large language models and visual language models are used for generating agent responses, then interaction intelligence is improved, but computational resources increase

Engineering Contradiction:
Improveinteractive agent intelligenceVSAvoidcomputational energy consumption
Core Design Contradiction:
Extent of automationVSUse of energy by moving object

Solution Approach 1:

The patent implements a partial processing approach where the conversational AI engine uses large language models and visual language models selectively based on query complexity and context availability. For simple spatial queries, the system uses streamlined processing with reduced model engagement. For complex multi-step tasks requiring deep reasoning, the full capabilities of the large language models are activated. This partial action approach maintains high interactive intelligence for complex tasks while reducing computational energy consumption for simpler interactions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20260024286A1Systems and methods for providing intelligent embodied interactive agents with spatial understanding
Publication Date: 2026.01.22 JPMORGAN CHASE BANK NA
  • US20260024286A1 patent drawing
  • US20260024286A1 patent drawing
  • US20260024286A1 patent drawing

AI summary

A method may include: a conversational artificial intelligence engine receiving from an augmented reality headset worn by a user, a query including audio and images/video that is made to an embodied interactive agent that is displayed in a display of the augmented reality headset; the conversational artificial intelligence engine generating a prompt for a large language model based on the user utterance and the images or video; the conversational artificial intelligence engine providing the prompt to the large language model; the conversational artificial intelligence engine receiving an output of the large language model, wherein the output comprises text and gestures for the embodied interactive agent; the conversational artificial intelligence engine generating animations for the embodied interactive agent from the gestures and speech for the embodied interactive agent based on the text; and the conversational artificial intelligence engine outputting the animations and the speech to the augmented reality headset.