Multimodal State Tracking via Scene Graphs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing assistant systems face challenges in efficiently processing and storing large amounts of data from client systems, leading to resource-intensive analysis and memory requirements.

Innovation Solution

The assistant system relies on user input to indicate objects of interest, incrementally generating a scene graph based on user interactions, and storing only relevant information, rather than analyzing and storing all incoming data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the assistant system analyzes and stores all incoming data from client systems, then the completeness of information is improved, but the processing resources and memory requirements increase significantly

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing resources
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the relevant information from incoming data streams based on user-indicated objects of interest. The scene graph construction process selectively captures and stores only those data elements that pertain to user-specified objects, filtering out unnecessary information to reduce processing load while maintaining completeness of relevant data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing quality levels to different portions of incoming data. User-indicated objects of interest receive full analysis and detailed storage, while other data receives minimal or no processing. This localized high-quality processing approach maintains information completeness for critical elements while reducing overall resource consumption.

Inventive Principle:
Principle #3Local quality

2Loss of information

If the assistant system analyzes and stores all incoming data from client systems, then the completeness of information is improved, but the memory requirements increase significantly

Engineering Contradiction:
Improveinformation completenessVSAvoidmemory requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts and stores only the essential attributes and relationships of user-indicated objects in the scene graph, rather than storing complete raw data. This selective extraction maintains the completeness of object information while dramatically reducing the quantity of stored data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the incoming data into distinct objects of interest and stores them as separate nodes in a scene graph structure. This segmentation allows the system to store only necessary information for each object and its relationships, reducing overall memory requirements compared to storing monolithic data sets.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If the assistant system processes all incoming data, then the accuracy of object identification is improved, but the processing time increases

Engineering Contradiction:
Improveobject identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary identification of user-indicated objects of interest before conducting detailed analysis. This preliminary action allows the system to focus subsequent processing only on relevant objects, maintaining high identification accuracy while reducing overall processing time by avoiding analysis of irrelevant data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies full analytical processing only to user-indicated objects of interest rather than all incoming data. This partial action approach achieves accurate object identification for critical elements while significantly reducing processing time compared to comprehensive analysis of all data.

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If the assistant system stores only user-relevant information, then the efficiency of resource usage is improved, but the complexity of data filtering increases

Engineering Contradiction:
Improveresource usage efficiencyVSAvoiddata filtering complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a scene graph as an intermediary data structure that simplifies the filtering process. The scene graph serves as a mediator between raw incoming data and stored information, automatically organizing and filtering data based on user-indicated objects of interest, thereby reducing the complexity of direct data filtering while improving resource efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary organization of incoming data into scene graph structures before final storage. This preliminary action pre-filters and structures data according to user-indicated objects, reducing the complexity of subsequent filtering operations and improving overall resource usage efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250148784A1Multimodal State Tracking via Scene Graphs for Assistant Systems
Publication Date: 2025.05.08 META PLATFORMS INC
  • US20250148784A1 patent drawing
  • US20250148784A1 patent drawing
  • US20250148784A1 patent drawing

AI summary

In one embodiment, a method includes receiving, from a client system associated with a user, a first user request that includes a reference to a target object and one or more of an attribute or a relationship of the target object. Visual data including one or more images portraying the target object may then be accessed, and the reference may be resolved to the target object portrayed in the one or more images. Object information of the target object that corresponds to the referenced attribute or relationship of the first user request may be determined based on a visual analysis of the one or more images. Finally, responsive to receiving the first user request, the object information of the target object may be stored in a multimodal dialog state.