Moving Object Control for Relative Position Region Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques struggle to achieve high accuracy in predicting a region in an image based on user instructions that include relative positional relationships, such as 'front of the vehicle on the right', especially when fusing language features with image features.

Innovation Solution

A moving object control system that utilizes a combination of machine learning models to fuse image features, depth features, and language features to predict regions in an image, using a first model for image feature extraction, a second for depth prediction, a third for language feature extraction, and a fourth for region prediction based on a fused feature.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If language features are fused to image features of an RGB image, then the system can process natural language instructions, but prediction accuracy deteriorates when the instructions include relative positional relationships

Engineering Contradiction:
Improvelanguage feature fusion capabilityVSAvoidregion prediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces depth information as an additional dimension to the traditional 2D RGB image features. By fusing image features with depth features, the system creates a 3D spatial representation that enables accurate interpretation of relative positional relationships in natural language instructions, thereby resolving the accuracy deterioration problem while maintaining language feature fusion capability

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If only image features from RGB images are used, then the system structure remains simple, but prediction accuracy deteriorates for instructions with relative positional relationships

Engineering Contradiction:
Improvefeature extraction system complexityVSAvoidregion prediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple types of features (image features from RGB images and depth features from depth maps) into a composite feature representation. This composite approach leverages the complementary information from different feature sources, enabling accurate prediction of regions with relative positional relationships while maintaining a manageable system structure through standardized fusion processes

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20250308190A1Moving object control system, information processing apparatus, method for a moving object control system, method for generating one or more machine learning models
Publication Date: 2025.10.02 HONDA MOTOR CO LTD
  • US20250308190A1 patent drawing
  • US20250308190A1 patent drawing
  • US20250308190A1 patent drawing

AI summary

A moving object control system in the present disclosure performs to acquire an image, acquire a user instruction in a natural language including a relative positional relationship; and predict a region in the image corresponding to a position in a scene indicated by the user instruction based on a fused feature obtained by fusing an image feature indicating a feature of the scene captured in the image, a depth of the scene captured in the image, and a language feature indicating a linguistic feature related to the user instruction by using one or more machine learning models.