Monocular Camera Localisation Using Synthetic 3D Meshes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-cost laser-based sensors, such as those used in state-of-the-art localisation systems for autonomous vehicles, pose a significant barrier for mainstream adoption due to their expense, particularly for robotics applications.
Innovation Solution
The use of cameras fitted to transportable apparatus in conjunction with a prior environmental model allows for the localization of vehicles using monocular cameras, shifting expensive sensing equipment to survey vehicles and enabling the generation of fully textured 3D meshes, which can then be used to simulate images and determine vehicle pose through Normalised Information Distance (NID) analysis, leveraging GPUs for real-time operation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If laser-based sensors are used for localisation, then measurement precision is improved, but device cost increases
Solution Approach 1:
The patent creates a synthetic copy of the environment as a textured 3D mesh model using laser scanner data collected during survey phases. This virtual model serves as a reference that can be reused multiple times without additional survey costs. The system then compares live camera images against rendered views from this synthetic model to achieve precise localisation, effectively replacing expensive laser sensors with inexpensive cameras for ongoing operation.
Solution Approach 2:
The patent performs environmental mapping and mesh generation in advance during survey phases using expensive laser sensors only when needed. This preliminary action creates a reusable prior model that eliminates the need for continuous use of expensive sensors. The survey vehicle collects data infrequently, while subsequent localisation operations rely on the pre-built model combined with cheap camera footage.
2Ease of manufacture
If monocular cameras are used instead of laser sensors, then device cost decreases, but measurement precision deteriorates
Solution Approach 1:
The patent introduces a textured 3D mesh model as an intermediary between the monocular camera and the environment. The mesh model encodes geometric and textural information that the camera alone cannot provide. By rendering synthetic images from the mesh and comparing them with actual camera images, the system enables precise pose estimation using only a monocular camera, bridging the gap between cheap sensing and precise measurement.
Solution Approach 2:
The patent transforms the localisation problem from direct sensor measurement to image similarity comparison. Instead of relying on laser range data, the system changes the measurement parameter from geometric distance to photometric similarity between rendered and captured images. This parameter transformation allows monocular cameras to achieve localisation precision previously only attainable with expensive sensors.
3Area of stationary object
If wide angle lenses are used to increase field of view, then area of detection is improved, but image distortion increases
Solution Approach 1:
Instead of trying to correct distorted wide-angle images, the patent inverts the approach by distorting the rendered images from the synthetic model to match the camera's distortion characteristics. This allows direct comparison between undistorted mesh geometry and distorted camera images, enabling the use of wide-angle lenses without requiring complex undistortion procedures.
Solution Approach 2:
The patent incorporates lens distortion parameters directly into the image rendering process. By changing the rendering parameters to account for wide-angle distortion, the system can utilise the full field of view of distorted images for localisation, transforming what would normally be a disadvantage into a usable feature that expands the detection area.
Data Source
Figure 1~2
Figure 3(a)~3(c)
Figure 4~6
AI summary
A method of localising portable apparatus (100) in an environment, the method comprising obtaining captured image data representing an image captured by an imaging device (106) associated with the portable apparatus, and obtaining mesh data representing a 3-dimensional textured mesh of at least part of the environment. The mesh data is processed to generate a plurality of synthetic images, each synthetic image being associated with a pose within the environment and being a simulation of an image that would be captured by the imaging device from that associated pose. The plurality of synthetic images is analysed to find a said synthetic image similar to the captured image data, and an indication is provided of a pose of the portable apparatus within the environment based on the associated pose of the similar synthetic image.