System and method for robust image-query understanding based on contextual features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current robotic systems face challenges in learning new knowledge from users and transferring that knowledge when the context changes, such as from one type of flooring to another, due to limitations in image-query understanding systems that cannot capture changes in image contexts and interpret new user queries.
Innovation Solution
A system and method for robust image-query understanding based on weighted contextual features, where an electronic device retrained using a correlation between a target image area and a target phrase in a user query, allowing it to learn and transfer knowledge across different contexts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the image-query understanding model is trained on fixed datasets, then the model achieves stable performance on known contexts, but the model cannot adapt when the context changes (e.g., different flooring types)
Solution Approach 1:
The patent implements dynamic retraining of the image-query understanding model based on user feedback. When users correct the model's misunderstandings about target areas, the system retrains the model with updated correlation data between image features and query meanings. This dynamic adaptation allows the model to maintain reliability on known contexts while gaining adaptability to new contexts such as different flooring types.
Solution Approach 2:
The system incorporates user feedback loops where users can correct the model's interpretations of target areas in images. These corrections are fed back into the training process, allowing the model to learn from its mistakes and improve its understanding of context-specific meanings. This feedback mechanism enables the model to adapt to changing contexts while maintaining stable performance on established knowledge.
2Measurement precision
If the model uses general image understanding, then it can process a wide variety of images, but it cannot accurately interpret specific target areas in context-dependent queries
Solution Approach 1:
The patent applies local quality by focusing the model's attention on specific target areas within images rather than treating all image regions uniformly. The system learns context-dependent correlations between specific image regions and their meanings in different queries. For example, it learns that a particular area in a bathroom image may represent a shower while the same visual features in a kitchen image represent a sink, thereby improving target area identification accuracy without requiring complete model redesign.
Solution Approach 2:
The system performs preliminary analysis to identify potential target areas in images before processing the full query. By pre-segmenting and analyzing relevant image regions based on initial query understanding, the system reduces the complexity of the main processing task while maintaining high accuracy in target area identification.
3Measurement precision
If the system requires extensive retraining data, then the model achieves higher accuracy on new contexts, but the retraining process becomes time-consuming and resource-intensive
Solution Approach 1:
The patent implements partial retraining by updating only the specific parameters and correlations related to the context changes rather than retraining the entire model from scratch. When users provide corrections about target areas, the system retrains only the relevant portions of the model that pertain to those specific contexts and image-query pairs, significantly reducing retraining time while maintaining interpretation accuracy.
Solution Approach 2:
The system changes specific parameters in the model based on user feedback and context changes rather than modifying the entire model structure. By adjusting only the correlation parameters between specific image features and query meanings that are affected by context changes, the system achieves high interpretation accuracy with minimal retraining time and computational resources.
Data Source
AI summary
A method includes obtaining, using at least one processor of an electronic device, an image-query understanding model. The method also includes obtaining, using the at least one processor, an image and a user query associated with the image, where the image includes a target image area and the user query includes a target phrase. The method further includes retraining, using the at least one processor, the image-query understanding model using a correlation between the target image area and the target phrase to obtain a retrained image-query understanding model.


