Mobile Body Positioning Using Scene Graphs for Ambiguous Spatial Cues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods struggle to accurately interpret ambiguous spatial instructions, such as 'right' or 'left', leading to difficulties in guiding mobile bodies like robots to appropriate areas around a target location due to the lack of unique coordinate expression.
Innovation Solution
A mobile body assistance device that generates scene graphs from environmental images to identify suitable areas based on user instructions, using a trained model to determine precise locations by analyzing positional relationships and obstacle information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional object detection methods are used to interpret spatial instructions, then the system can process basic object detection, but it cannot accurately determine the intended area from ambiguous spatial expressions like 'right' or 'left'
Solution Approach 1:
The patent introduces scene graphs as an intermediary representation layer between the mobile body's position and the user's spatial instructions. The scene graph captures semantic relationships between objects and spaces, enabling the system to interpret ambiguous expressions like 'right' by analyzing the spatial context and relationships in the graph structure, rather than directly mapping to coordinates.
Solution Approach 2:
The patent transitions from traditional 2D image coordinates to a multi-dimensional scene graph representation that includes semantic relationships, spatial positions, and contextual information. This dimensional expansion allows the system to disambiguate spatial instructions by considering multiple factors simultaneously (object relationships, space attributes, mobile body position) rather than relying solely on coordinate geometry.
2Measurement precision
If the system uses detailed scene graphs to accurately interpret spatial instructions, then area selection accuracy improves, but processing time and computational load increase
Solution Approach 1:
The system performs preliminary scene graph construction and spatial relationship analysis before the mobile body reaches the designated place. By pre-processing the environmental information and organizing it into a structured scene graph with pre-computed spatial relationships, the system reduces the computational burden during real-time decision-making, enabling faster area candidate identification when needed.
3Measurement precision
If the mobile body stops at precisely determined coordinates, then positioning accuracy improves, but it may stop in inappropriate areas like crosswalks where stopping is prohibited
Solution Approach 1:
The patent changes the parameter representation from precise coordinate values to area candidates with semantic attributes. Instead of determining a single stop coordinate, the system identifies multiple candidate areas with different attributes (e.g., sidewalk, crosswalk, parking area) and selects the most appropriate one based on the instruction context and area properties, thereby ensuring both precision and appropriateness.
Data Source
AI summary
In view of an intension of an instructor whose space designation based on a target place is an ambiguous instruction, provided is a mobile body that can search for an appropriate area around the target place for the mobile body to realize a designated state according to the instruction. The model is constructed using the scene graphs SG1 to SG3 created based on the user's instruction and the environmental image corresponding to the position of the mobile body 20 and the direction facing the designated place as input data. The feature amount of the primary node constituting the state scene graph SG1 is defined according to relative arrangement relationship (distance and angle) with each object with respect to the position of the mobile body 20. The feature amount of the primary node constituting the state scene graph SG1 is defined according to the space occupancy mode of each object.


