Multimodal State Tracking via Scene Graphs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing assistant systems face challenges in efficiently processing and storing large amounts of data from client systems, leading to resource-intensive analysis and memory requirements.
Innovation Solution
The assistant system relies on user input to indicate objects of interest, incrementally generating a scene graph based on user interactions, and storing only relevant information, rather than analyzing and storing all incoming data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If the assistant system analyzes and stores all incoming data from client systems, then the completeness of information is improved, but the processing resources and memory requirements increase significantly
Solution Approach 1:
The patent extracts only the relevant information from incoming data streams based on user-indicated objects of interest. The scene graph construction process selectively captures and stores only those data elements that pertain to user-specified objects, filtering out unnecessary information to reduce processing load while maintaining completeness of relevant data.
Solution Approach 2:
The patent applies different processing quality levels to different portions of incoming data. User-indicated objects of interest receive full analysis and detailed storage, while other data receives minimal or no processing. This localized high-quality processing approach maintains information completeness for critical elements while reducing overall resource consumption.
2Loss of information
If the assistant system analyzes and stores all incoming data from client systems, then the completeness of information is improved, but the memory requirements increase significantly
Solution Approach 1:
The system extracts and stores only the essential attributes and relationships of user-indicated objects in the scene graph, rather than storing complete raw data. This selective extraction maintains the completeness of object information while dramatically reducing the quantity of stored data.
Solution Approach 2:
The patent segments the incoming data into distinct objects of interest and stores them as separate nodes in a scene graph structure. This segmentation allows the system to store only necessary information for each object and its relationships, reducing overall memory requirements compared to storing monolithic data sets.
3Measurement precision
If the assistant system processes all incoming data, then the accuracy of object identification is improved, but the processing time increases
Solution Approach 1:
The patent performs preliminary identification of user-indicated objects of interest before conducting detailed analysis. This preliminary action allows the system to focus subsequent processing only on relevant objects, maintaining high identification accuracy while reducing overall processing time by avoiding analysis of irrelevant data.
Solution Approach 2:
The system applies full analytical processing only to user-indicated objects of interest rather than all incoming data. This partial action approach achieves accurate object identification for critical elements while significantly reducing processing time compared to comprehensive analysis of all data.
4Productivity
If the assistant system stores only user-relevant information, then the efficiency of resource usage is improved, but the complexity of data filtering increases
Solution Approach 1:
The patent introduces a scene graph as an intermediary data structure that simplifies the filtering process. The scene graph serves as a mediator between raw incoming data and stored information, automatically organizing and filtering data based on user-indicated objects of interest, thereby reducing the complexity of direct data filtering while improving resource efficiency.
Solution Approach 2:
The system performs preliminary organization of incoming data into scene graph structures before final storage. This preliminary action pre-filters and structures data according to user-indicated objects, reducing the complexity of subsequent filtering operations and improving overall resource usage efficiency.
Data Source
AI summary
In one embodiment, a method includes receiving, from a client system associated with a user, a first user request that includes a reference to a target object and one or more of an attribute or a relationship of the target object. Visual data including one or more images portraying the target object may then be accessed, and the reference may be resolved to the target object portrayed in the one or more images. Object information of the target object that corresponds to the referenced attribute or relationship of the first user request may be determined based on a visual analysis of the one or more images. Finally, responsive to receiving the first user request, the object information of the target object may be stored in a multimodal dialog state.


