Mobile Object Control Using Bird's Eye View Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mobile object control systems face challenges in detecting travelable space efficiently due to the need for complex hardware configurations and large amounts of training data, especially when using multiple ranging sensors or simple camera setups.
Innovation Solution
A mobile object control device that uses a camera to capture images, converts them into bird's eye view coordinates, and employs a trained model to detect three-dimensional objects and travelable spaces, reducing the complexity of hardware and training data requirements by utilizing a combination of radial pattern and color pattern annotations, as well as temporal variation analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple ranging sensors are used to detect obstacles, then detection reliability is improved, but device complexity and cost increase
Solution Approach 1:
The patent combines multiple sensing functions (obstacle detection, distance measurement, and travelable space detection) into a single camera system. By integrating these functions and using image processing algorithms, the system achieves reliable detection without requiring separate ranging sensors, thus reducing hardware complexity while maintaining detection reliability.
Solution Approach 2:
The camera system is designed to perform multiple functions: capturing images for obstacle detection, generating bird's eye view images for spatial analysis, and identifying travelable spaces. This multi-functional approach eliminates the need for dedicated ranging sensors, reducing device complexity while preserving detection capabilities.
2Device complexity
If only cameras are used to reduce hardware complexity, then device complexity is reduced, but large amounts of training data are required to ensure robustness
Solution Approach 1:
The system performs preliminary processing by converting captured images into bird's eye view images before analysis. This preprocessing step transforms the data into a format that highlights spatial relationships and makes training more efficient, reducing the amount of training data needed while maintaining robustness across various scenes.
Solution Approach 2:
The patent changes the representation parameters of the input data by converting standard camera images into bird's eye view images. This parameter transformation enhances the visibility of spatial patterns and travelable spaces, allowing the model to learn more efficiently from smaller datasets while maintaining generalization capability.
3Ease of manufacture
If simple camera configuration is used, then system cost is reduced, but robustness to various scenes deteriorates without large training data
Solution Approach 1:
By performing preliminary conversion of images to bird's eye view representation, the system enhances spatial pattern recognition capabilities. This preprocessing enables the simple camera configuration to achieve robust performance across various scenes by emphasizing geometric relationships that are invariant to camera position and orientation.
Solution Approach 2:
The patent transforms two-dimensional camera images into bird's eye view representations, effectively adding a dimensional transformation that enhances spatial understanding. This dimensional change allows the simple camera system to achieve scene-robust detection by representing obstacles and travelable spaces in a coordinate system that naturally highlights navigable areas.
Data Source
AI summary
Provided is a mobile object control device comprising a storage medium storing computer-readable commands and a processor connected to the storage medium, the processor executing the computer-readable commands to: acquire a subject bird's eye view image obtained by converting an image, which is photographed by a camera mounted in a mobile object to capture a surrounding situation of the mobile object, into a bird's eye view coordinate system; input the subject bird's eye view image into a trained model, which is trained to receive input of a bird's eye view image to output at least a three-dimensional object in the bird's eye view image, to detect a three-dimensional object in the subject bird's eye view image; detect a travelable space of the mobile object based on the detected three-dimensional object; and cause the mobile object to travel so as to pass through the travelable space.


