Moving Body Meeting Position Estimation From Utterances and Visual Marks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for ultra-compact mobility lack the ability to dynamically adjust meeting positions based on user utterances, especially when the user and vehicle are not at a designated meeting point, leading to difficulties in coordinating movements.
Innovation Solution
An information processing apparatus that acquires utterance information and captured images to determine an object region corresponding to a visual mark, estimating the instruction position for the moving body using machine learning models for voice and image recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a designated meeting position is predetermined for the user and moving body, then the coordination is simplified, but the system lacks flexibility when congestion or other issues prevent meeting at the designated position
Solution Approach 1:
The meeting position is transformed from a static predetermined location to a dynamic adjustable position. The moving body can modify its meeting position based on real-time conditions by interpreting user utterances and adjusting the position coordinates dynamically, allowing adaptation when congestion or other issues arise at the original designated position
Solution Approach 2:
The system incorporates feedback loops where the moving body continuously monitors its position relative to the user, receives utterance information from the user about desired position adjustments, and modifies its meeting position accordingly. This feedback mechanism enables real-time coordination adjustments without requiring complete redesign of the meeting protocol
2Adaptability or versatility
If the user designates a rough area or building instead of a specific position, then flexibility is improved, but the precision of the meeting position determination deteriorates
Solution Approach 1:
The meeting position determination is segmented into two stages: first, the user designates a rough area or building (low precision requirement), then the moving body refines this to a specific meeting position within that area using utterance interpretation (high precision requirement). This segmentation allows the system to balance flexibility in area selection with precision in final position determination
Solution Approach 2:
The user's rough area designation serves as a preliminary action that narrows down the search space for the final meeting position. The moving body uses this preliminary information to focus its position determination efforts within a constrained area, improving both user convenience and position accuracy
3Device complexity
If the system requires the user to visit a predetermined renting place, then the vehicle management is simplified, but the user convenience deteriorates when the user cannot reach the renting place
Solution Approach 1:
The meeting position transitions from a fixed predetermined location to a dynamic position that can be adjusted based on user accessibility. The moving body interprets user utterances about desired positions and modifies its target position dynamically, ensuring the user can actually reach the meeting point even when predetermined locations are inaccessible due to congestion or other barriers
Data Source
AI summary
An information processing apparatus that estimates an instruction position for a moving body used by a user acquires utterance information regarding the instruction position including a visual mark from a communication device used by the user. The information processing apparatus acquires a captured image captured by the moving body and determines an object region in the captured image corresponding to the visual mark included in the utterance information. The information apparatus estimates the instruction position based on the object region.


