Moving Object Control With Depth-Aware Language Region Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques struggle to achieve high accuracy in predicting image regions based on user instructions that include relative positional relationships, such as 'front of the vehicle on the right', especially when fusing language features with image features.

Innovation Solution

A moving object control system that utilizes a machine learning model to fuse image and depth features with language features using a pixel-wise attention mechanism, enabling accurate prediction of image regions based on user instructions by concatenating these features for each predetermined unit region.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If language features are fused to image features of an RGB image, then the system can process user instructions, but prediction accuracy is insufficient when the instruction includes relative positional relationships

Engineering Contradiction:
Improveprediction accuracyVSAvoidfeature fusion complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces depth information as an additional dimension to the traditional RGB image data. By constructing a depth map and extracting depth features, the system transforms 2D image processing into 3D spatial understanding, enabling accurate interpretation of relative positional relationships in user instructions without significantly increasing system complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces a position relationship determination module as an intermediary that bridges the gap between language features and image features. This module specifically processes relative positional relationships by determining spatial relationships between objects, acting as a mediator that enhances the fusion process and improves prediction accuracy for position-related instructions

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4625351A1Moving object control system, information processing apparatus, method for a moving object control system, method for generating one or more machine learning models
Publication Date: 2025.10.01 HONDA MOTOR CO LTD
  • EP4625351A1 patent drawingFigure 1A~1B
  • EP4625351A1 patent drawingFigure 2
  • EP4625351A1 patent drawingFigure 3

AI summary

A moving object control system in the present disclosure performs to acquire an image, acquire a user instruction in a natural language including a relative positional relationship; and predict a region in the image corresponding to a position in a scene indicated by the user instruction based on a fused feature obtained by fusing an image feature indicating a feature of the scene captured in the image, a depth of the scene captured in the image, and a language feature indicating a linguistic feature related to the user instruction by using one or more machine learning models.