Single-Image Pole Extraction Using Keypoints and Monocular Depth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mapping and navigation service providers face challenges in determining the geolocations and attributes of poles and other objects across large geographic areas due to their ubiquity and large numbers, which often requires labor-intensive ground-based surveys and human interpretation of high-resolution remote sensing imagery.
Innovation Solution
A system utilizing machine learning and photogrammetry to automatically extract pole-like objects from optical imagery, employing deep learning to detect bounding boxes and semantic keypoints, followed by photogrammetric triangulation to determine their geolocations and attributes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If ground-based surveys and human interpretation of high-resolution remote sensing imagery are used to determine geolocations and attributes of poles, then measurement precision can be maintained, but productivity is significantly reduced due to labor-intensive processes
Solution Approach 1:
The patent replaces manual ground-based surveying and human interpretation of remote sensing imagery with an automated computer vision system. The system uses machine learning models to detect pole-like objects in aerial imagery and automatically calculates their geolocations and attributes, substituting mechanical human labor with automated computational processes while maintaining measurement precision through algorithmic accuracy.
Solution Approach 2:
The patent introduces an intermediary automated processing system that bridges aerial imagery and geolocation data. This intermediary system comprises training data preparation, model training, and automated inference components that translate visual imagery into precise geolocation information without requiring direct human intervention in the measurement process.
2Productivity
If automated image processing techniques are used to extract pole geolocations and attributes, then productivity is improved through efficient processing, but device complexity increases due to machine learning and photogrammetry systems
Solution Approach 1:
The patent segments the complex automated processing system into distinct functional modules: data preparation module for organizing training data, model training module for developing the machine learning algorithm, and automated inference module for executing pole detection and geolocation calculation. This segmentation manages device complexity by creating modular, independently manageable components that collectively achieve high productivity.
3Ease of operation
If deep learning models are trained to detect pole-like objects in aerial imagery, then ease of operation is improved through automated detection, but loss of time increases during the model training phase
Solution Approach 1:
The patent performs preliminary actions by preparing and organizing training data in advance, creating a structured dataset with labeled pole-like objects before model training begins. This preliminary data preparation includes collecting aerial imagery, identifying pole locations, and creating annotation files, which streamlines the subsequent model training process and reduces overall time loss by avoiding ad-hoc data processing during deployment.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An approach is provided for pole extraction from a single image. The approach involves, for instance, processing an image using a machine learning model to detect one or more semantic keypoints associated with a pole-like object and to determine two-dimensional coordinate data for the one or more semantic keypoints. The approach also involves performing a monocular depth estimation to determine depth information for the one or more semantic keypoints based on the image. The approach further involves determining three-dimensional coordinate data for the one or more semantic keypoints based on the monocular depth information, the two-dimensional coordinate data, and camera parameter data. The approach yet further involves providing the three-dimensional coordinate data as an output.