Vision-Language Model Planner for Mixed Reality Task Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current large language models (LLMs) lack spatial awareness of their environment and objects within it, making them ineffective in providing real-time guidance for tasks that require interaction with physical objects.

Innovation Solution

A spatially and semantically aware generative AI system that uses a vision-language model planner to provide instructions and guidance for tasks involving complex multipart objects, by integrating visual perception, language understanding, memory, affordance understanding, and multi-agent reasoning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If LLMs are used to assist users generate text and perform complex tasks, then text generation capability is improved, but spatial awareness is lost

Engineering Contradiction:
Improvetext generation capabilityVSAvoidspatial awareness
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent merges LLMs with vision-language models to create a hybrid system that combines strong text generation capabilities with spatial awareness. The vision-language model processes visual inputs to understand spatial relationships, while the LLM generates coherent text responses, resolving the contradiction between text generation productivity and spatial awareness.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements a multi-functional AI assistant that can perform both text generation tasks and spatial understanding tasks. The vision-language model serves multiple functions: analyzing images, understanding spatial relationships, and providing guidance, while the LLM handles text generation, making the system universally applicable to diverse tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If LLMs provide guidance for tasks requiring physical object interaction, then task completion is improved, but spatial understanding is insufficient

Engineering Contradiction:
Improvetask guidance capabilityVSAvoidspatial understanding
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The vision-language model acts as an intermediary between the user and the physical environment. It processes visual information about objects and spatial relationships, then translates this into actionable guidance instructions. This intermediary component enables the system to provide accurate task guidance while maintaining spatial understanding.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system incorporates feedback loops where the vision-language model continuously analyzes the current state of the environment and adjusts guidance instructions accordingly. This feedback mechanism ensures that task guidance remains accurate and up-to-date with the actual spatial conditions, improving both ease of operation and spatial understanding.

Inventive Principle:
Principle #23Feedback

3Loss of information

If spatial awareness is added to LLMs through vision-language models, then spatial understanding is improved, but computing resources increase

Engineering Contradiction:
Improvespatial awarenessVSAvoidcomputing resources
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system segments the AI functionality into specialized components: vision-language models handle spatial understanding and visual analysis, while LLMs handle text generation. This segmentation allows each component to be optimized for its specific function, improving spatial awareness while managing computing resource consumption more efficiently than a monolithic approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies vision-language models selectively only when spatial understanding is required, rather than continuously. This partial action approach allows the system to maintain spatial awareness capabilities while reducing overall computing resource consumption during tasks that do not require visual analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250191307A1Three dimensional spatial instructions for artifical intelligence assistance authoring
Publication Date: 2025.06.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250191307A1 patent drawing
  • US20250191307A1 patent drawing
  • US20250191307A1 patent drawing

AI summary

A data processing system implements receiving a first request to collaboratively author a mixed reality experience with a vision-language model planner, the mixed reality experience comprising an interactive guide for performing a task involving a complex multipart object; obtaining 3D object geometry information for the complex multipart object; obtaining a description of the task to be performed including a plurality of subtasks each associated with a user action to be performed on a respective part of the complex multipart object; constructing a prompt to the model using a prompt construction unit, the prompt instructing the model to generate a task list based on the geometry information and the description of the task to be performed; providing the prompt as an input to the model to obtain the task list; and generating content for the mixed reality experience using the task list in response to a second request to execute the mixed reality experience.