Moving Object Control With Depth-Aware Language Region Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques struggle to achieve high accuracy in predicting image regions based on user instructions that include relative positional relationships, such as 'front of the vehicle on the right', especially when fusing language features with image features.
Innovation Solution
A moving object control system that utilizes a machine learning model to fuse image and depth features with language features using a pixel-wise attention mechanism, enabling accurate prediction of image regions based on user instructions by concatenating these features for each predetermined unit region.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If language features are fused to image features of an RGB image, then the system can process user instructions, but prediction accuracy is insufficient when the instruction includes relative positional relationships
Solution Approach 1:
The patent introduces depth information as an additional dimension to the traditional RGB image data. By constructing a depth map and extracting depth features, the system transforms 2D image processing into 3D spatial understanding, enabling accurate interpretation of relative positional relationships in user instructions without significantly increasing system complexity
Solution Approach 2:
The patent introduces a position relationship determination module as an intermediary that bridges the gap between language features and image features. This module specifically processes relative positional relationships by determining spatial relationships between objects, acting as a mediator that enhances the fusion process and improves prediction accuracy for position-related instructions
Data Source
Figure 1A~1B
Figure 2
Figure 3
AI summary
A moving object control system in the present disclosure performs to acquire an image, acquire a user instruction in a natural language including a relative positional relationship; and predict a region in the image corresponding to a position in a scene indicated by the user instruction based on a fused feature obtained by fusing an image feature indicating a feature of the scene captured in the image, a depth of the scene captured in the image, and a language feature indicating a linguistic feature related to the user instruction by using one or more machine learning models.