Landmark Localization Using Auto-Labeled Visual Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for visual localization and mapping in robotics lack efficiency and robustness, particularly in large environments, due to the need for extensive manual labeling of images to identify landmarks, which is time-consuming and costly.
Innovation Solution
A method and system that uses a labeling object placed in a specific spatial relation to the desired landmark, allowing automatic generation of labeled training data for neural networks to detect landmarks without manual labeling, utilizing pre-trained object detectors and spatial relation information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling of images is used to identify landmarks, then localization accuracy can be achieved, but the process is time-consuming and costly
Solution Approach 1:
The system enables self-service labeling by using the robot itself to capture images and automatically determine landmark positions through feature detection algorithms. The robot captures images during its autonomous navigation, and the system automatically processes these images to identify and label landmarks without requiring external manual annotation, thus eliminating the time-consuming manual labeling process while maintaining localization accuracy
Solution Approach 2:
The patent replaces the mechanical/manual labeling process with an automated computational system. Instead of manually annotating images, the system uses computer vision algorithms (feature detection, descriptor extraction, and machine learning classifiers) to automatically identify and label landmarks. This substitution of manual mechanical labeling with automated computational processing significantly reduces time and cost while achieving comparable or superior localization accuracy
2Reliability
If extensive manual labeling is performed, then training data quality improves, but the complexity and cost increase
Solution Approach 1:
The system performs self-service data collection and labeling by utilizing the robot's own sensor suite and processing capabilities. The robot captures images during autonomous operation, and the system automatically processes these images through feature detection and machine learning algorithms to generate training data. This self-service approach eliminates the need for complex external labeling infrastructure and manual intervention, reducing overall system complexity while maintaining high training data quality
Solution Approach 2:
The patent implements preliminary action by pre-processing images during the robot's navigation phase before the actual training is needed. The system captures and processes images in advance, extracting features and generating labeled data that can be used for training localization models. This preliminary data collection and preparation eliminates the need for complex post-training labeling processes and reduces overall system complexity
3Measurement precision
If traditional SLAM methods are used, then localization can be achieved, but semantic information of the scene is lacking
Solution Approach 1:
The patent merges traditional SLAM localization capabilities with semantic object detection and recognition. The system combines feature-based SLAM algorithms with deep learning-based object detectors and semantic segmenters to simultaneously achieve precise localization and rich semantic understanding of the environment. This merging of localization and semantic processing modules enables the system to maintain localization precision while gaining comprehensive semantic information about scenes, objects, and their relationships
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to a method for generating detecting information for detecting a new landmark in an image, comprising the steps of a) obtaining (Si) images including the landmark and a labeling object placed in a certain spatial relation to the landmark; b) obtaining (S2, S4) labeling object information for detecting the labeling object in the obtained images and spatial relation information indicating the certain spatial relation; c) detecting (S3), in each of the obtained images, the labeling object based on the labeling object information; d) determining (S5), in each of the obtained images, an image region including the landmark based on the spatial relation information and the detected labeling object to provide a plurality of image regions including the landmark; and e) generating (S6) the detecting information by training a neural network to detect the landmark without the labeling object using the plurality of image regions as training data.