Vision-Language Model Planner for Mixed Reality Task Guidance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current large language models (LLMs) lack spatial awareness of their environment and objects within it, making them ineffective in providing real-time guidance for tasks that require interaction with physical objects.
Innovation Solution
A spatially and semantically aware generative AI system that uses a vision-language model planner to provide instructions and guidance for tasks involving complex multipart objects, by integrating visual perception, language understanding, memory, affordance understanding, and multi-agent reasoning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If LLMs are used to assist users generate text and perform complex tasks, then text generation capability is improved, but spatial awareness is lost
Solution Approach 1:
The patent merges LLMs with vision-language models to create a hybrid system that combines strong text generation capabilities with spatial awareness. The vision-language model processes visual inputs to understand spatial relationships, while the LLM generates coherent text responses, resolving the contradiction between text generation productivity and spatial awareness.
Solution Approach 2:
The system implements a multi-functional AI assistant that can perform both text generation tasks and spatial understanding tasks. The vision-language model serves multiple functions: analyzing images, understanding spatial relationships, and providing guidance, while the LLM handles text generation, making the system universally applicable to diverse tasks.
2Ease of operation
If LLMs provide guidance for tasks requiring physical object interaction, then task completion is improved, but spatial understanding is insufficient
Solution Approach 1:
The vision-language model acts as an intermediary between the user and the physical environment. It processes visual information about objects and spatial relationships, then translates this into actionable guidance instructions. This intermediary component enables the system to provide accurate task guidance while maintaining spatial understanding.
Solution Approach 2:
The system incorporates feedback loops where the vision-language model continuously analyzes the current state of the environment and adjusts guidance instructions accordingly. This feedback mechanism ensures that task guidance remains accurate and up-to-date with the actual spatial conditions, improving both ease of operation and spatial understanding.
3Loss of information
If spatial awareness is added to LLMs through vision-language models, then spatial understanding is improved, but computing resources increase
Solution Approach 1:
The system segments the AI functionality into specialized components: vision-language models handle spatial understanding and visual analysis, while LLMs handle text generation. This segmentation allows each component to be optimized for its specific function, improving spatial awareness while managing computing resource consumption more efficiently than a monolithic approach.
Solution Approach 2:
The system applies vision-language models selectively only when spatial understanding is required, rather than continuously. This partial action approach allows the system to maintain spatial awareness capabilities while reducing overall computing resource consumption during tasks that do not require visual analysis.
Data Source
AI summary
A data processing system implements receiving a first request to collaboratively author a mixed reality experience with a vision-language model planner, the mixed reality experience comprising an interactive guide for performing a task involving a complex multipart object; obtaining 3D object geometry information for the complex multipart object; obtaining a description of the task to be performed including a plurality of subtasks each associated with a user action to be performed on a respective part of the complex multipart object; constructing a prompt to the model using a prompt construction unit, the prompt instructing the model to generate a task list based on the geometry information and the description of the task to be performed; providing the prompt as an input to the model to obtain the task list; and generating content for the mixed reality experience using the task list in response to a second request to execute the mixed reality experience.


