Single Camera Environmental Map Construction Using 3D Object Dictionary
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for creating environmental maps using a single camera struggle to accurately recognize and analyze various objects, such as 'wall', 'table', and 'sofa', and do not perform detailed analysis of three-dimensional shapes, limiting the map's detail and usefulness for navigation.
Innovation Solution
An information processing apparatus that includes a camera, a self-position detecting unit, an image-recognition processing unit, and a data constructing unit, which uses dictionary data with three-dimensional shape information to detect objects, calculate their positions and postures, and update the environmental map, enabling detailed object arrangement and recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If environmental map creation is performed using only one camera without distance image generation, then device complexity is reduced, but measurement precision of object positions and three-dimensional shapes deteriorates
Solution Approach 1:
The system performs preliminary action by pre-registering three-dimensional shape data of objects in a dictionary before actual environmental map creation. When an object is detected in the captured image, the system retrieves and applies the corresponding pre-stored three-dimensional shape data to calculate accurate spatial information, eliminating the need for complex real-time depth calculation from multiple cameras.
Solution Approach 2:
The system uses copying by creating a virtual three-dimensional representation of objects based on two-dimensional image data. The image recognition unit detects objects in the captured image, and the environmental map creation unit copies the pre-stored three-dimensional shape data from the dictionary to construct accurate three-dimensional environmental maps, effectively transforming 2D image information into 3D spatial data.
2Loss of information
If detailed three-dimensional shape analysis of objects is performed, then environmental map information quality is improved, but processing time and computational load increase
Solution Approach 1:
The system performs preliminary action by pre-registering three-dimensional shape data of objects in a dictionary before actual environmental map creation. When an object is detected in the captured image, the system retrieves and applies the corresponding pre-stored three-dimensional shape data to calculate accurate spatial information, eliminating the need for complex real-time depth calculation from multiple cameras.
Solution Approach 2:
The system applies parameter changes by switching between different data representations. Instead of performing complex real-time three-dimensional calculations, the system changes the approach by retrieving pre-calculated three-dimensional shape parameters from the dictionary, effectively transforming the computational problem from real-time calculation to data retrieval and application.
3Loss of information
If object recognition and three-dimensional shape detection are performed from single camera images, then environmental map detail is improved, but reliability of object detection deteriorates due to ambiguity in depth and shape perception
Solution Approach 1:
The system uses copying by creating a virtual three-dimensional representation of objects based on two-dimensional image data. The image recognition unit detects objects in the captured image, and the environmental map creation unit copies the pre-stored three-dimensional shape data from the dictionary to construct accurate three-dimensional environmental maps, effectively transforming 2D image information into 3D spatial data.
Solution Approach 2:
The dictionary storing pre-registered three-dimensional shape data acts as an intermediary between the two-dimensional camera images and the three-dimensional environmental map. This intermediary provides the missing depth and shape information that cannot be directly obtained from single-camera images, enabling accurate three-dimensional reconstruction without requiring multiple cameras or complex depth sensing.
Data Source
AI summary
An information processing apparatus that executes processing for creating an environmental map includes a camera that photographs an image, a self-position detecting unit that detects a position and a posture of the camera on the basis of the image, an image-recognition processing unit that detects an object from the image, a data constructing unit that is inputted with information concerning the position and the posture of the camera and information concerning the object and executes processing for creating or updating the environmental map, and a dictionary-data storing unit storing dictionary data in which object information is registered. The image-recognition processing unit executes processing for detecting an object from the image with reference to the dictionary data. The data constructing unit applies the three-dimensional shape data to the environmental map and executes object arrangement on the environmental map.


