Occupancy Map Inpainting for Occluded Robot Path Planning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Path planning for mobile robots in indoor spaces is challenging due to limited visual information from ground-level cameras and occlusions from obstacles, making it difficult to navigate effectively without complex and costly sensor installations.
Innovation Solution
A method involving training an inpainting network using depth images from both upper and lower 3D mapping sensors to predict occupancy of unseen areas, generating occupancy maps, and updating global maps for efficient path planning and navigation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If ground-level cameras are used for navigation, then the robot structure remains simple and compact, but visual information is severely limited due to occlusions from obstacles
Solution Approach 1:
The patent introduces a virtual dimension through inpainting network that synthesizes occluded visual information from the ground-level camera perspective. Instead of physically elevating the camera to overcome occlusions, the system uses machine learning to reconstruct what would be visible from higher vantage points, effectively adding an informational dimension without physical complexity
Solution Approach 2:
The inpainting network creates virtual copies of occluded regions by synthesizing image data that resembles what the camera would capture from elevated positions. This copying approach allows the ground-level camera to access information equivalent to higher vantage points without requiring physical camera elevation
2Loss of information
If sensors are placed in the environment to overcome limited visual information, then navigation information improves, but system complexity, installation difficulty, and cost increase
Solution Approach 1:
The system uses the robot's own ground-level camera and computational resources to generate enhanced navigation information. The inpainting network processes the camera's own output to synthesize additional visual data, making the system self-sufficient without requiring external sensor infrastructure or environmental modifications
Solution Approach 2:
The patent replaces the mechanical approach of adding physical sensors with a computational approach using machine learning. Instead of mechanically adding sensors to the environment, the system uses software-based inpainting to generate navigation information, substituting computational processing for physical sensor deployment
3Loss of information
If cameras are positioned high above the ground to replicate human visual field, then field of view improves, but base size and system weight increase
Solution Approach 1:
The patent creates a virtual elevated viewpoint through computational inpainting rather than physical camera elevation. The inpainting network generates image data that represents what would be visible from high vantage points, adding the informational benefit of elevated positioning without the physical weight penalty
Solution Approach 2:
The system creates virtual copies of elevated viewpoint images through the inpainting network. Instead of physically mounting cameras high above the ground, the network synthesizes copies of what such cameras would capture, allowing the lightweight ground-level robot to access elevated perspective information
Data Source
AI summary
A method of predicting occupancy of unseen areas in a region of interest (ROI) includes obtaining a depth image of the ROI, the depth image being captured from a first height; generating an occupancy map based on the obtained depth image, the occupancy map comprising an array of cells corresponding to locations in the ROI; and generating an inpainted map by inputting the occupancy map into a trained inpainting network, the inpainted map comprising an array of cells corresponding to the ROI, and wherein the inpainting network is trained by comparing an output of the inpainting network, based on inputting a training depth image taken from the first height, to a ground truth map, the ground truth map being based on a combination of the training depth image and a depth image taken at a height different than the first height.


