Monocular BEV Semantic Mapping for Tight-Space Maritime Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies have not effectively addressed the need for detailed, semantically segmented bird's eye view (BEV) semantic mapping in maritime environments, particularly in environments where data is scarce and costly, and the lack of data for effective spatial representation in maritime environments, particularly in maritime environments, where data is scarce and costly, and the lack of effective spatial representation for maneuvering vehicles in tight spaces.
Innovation Solution
A method and system using a monocular camera and artificial neural network (ANN) to generate a semantic map with a uniform scale from a first point of view (POV), and combining multiple semantic maps from different POV's to create a unified bird's eye view (BEV) semantic mapping semantic maps from different points of view (POV) semantic maps with a shared POV, and a system including a monocular camera and ANN to generate a semantic map with a uniform scale from a first point of view (POV) and combining semantic maps from different points of view (POV) semantic maps with a shared POV.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular cameras are used to capture images in maritime environments, then device complexity is reduced, but measurement precision of spatial relationships deteriorates due to lack of depth information
Solution Approach 1:
The patent transforms 2D image data from monocular cameras into a 3D bird's eye view representation by introducing a virtual third dimension (depth/elevation). The neural network processes sequential 2D images and reconstructs 3D spatial relationships, converting the limitation of monocular vision into a comprehensive spatial map that includes depth information without requiring multiple cameras.
Solution Approach 2:
The patent introduces an intermediary computational model (neural network) that acts as a mediator between the monocular camera input and the desired 3D spatial understanding. This intermediary processes the limited 2D visual information and generates intermediate representations (optical flow, depth maps, BEV features) that reconstruct the missing depth information through computational inference.
2Measurement precision
If multiple cameras are used to generate comprehensive environmental maps, then measurement precision improves, but device complexity and cost increase
Solution Approach 1:
The patent segments the complex task of 3D environmental mapping into multiple processing stages: first extracting 2D features from individual monocular images, then computing optical flow between sequential frames, generating depth estimates, and finally assembling these into a bird's eye view representation. This segmentation allows a single camera to achieve multi-camera-level mapping precision through systematic processing.
Solution Approach 2:
The patent achieves comprehensive 3D environmental understanding by transforming sequential 2D images into a 3D bird's eye view representation. By processing images over time and introducing the temporal dimension, the system reconstructs depth and spatial relationships that would normally require multiple simultaneous cameras, thereby reducing device complexity while maintaining mapping precision.
3Ease of operation
If bird's eye view representations are generated from monocular images, then navigation utility is improved, but information loss occurs during perspective transformation
Solution Approach 1:
The patent employs feedback mechanisms where the neural network continuously refines its bird's eye view predictions by comparing expected features with actual observed features in sequential frames. The optical flow computation and depth estimation are adjusted based on feedback from matching features across frames, ensuring that the transformation from monocular view to bird's eye view preserves critical spatial information while maximizing navigation utility.
Solution Approach 2:
The patent performs preliminary extraction of spatial features and depth information from monocular images before completing the bird's eye view transformation. By pre-computing optical flow, depth maps, and feature correspondences, the system preserves essential spatial information that would otherwise be lost during perspective transformation, ensuring accurate navigation utility in the final BEV representation.
Data Source
AI summary
Bird's eye view (BEV) semantic mapping systems and methods are provided. A method includes receiving an image captured by a monocular camera having a first point of view (POV) of an environment including a plurality of features. The method further includes processing, by an artificial neural network (ANN), the captured image to generate a semantic map for the captured image, the semantic map associated with a second POV different from the first POV. The features exhibit a uniform scale in the semantic map. Additional methods and associated systems are also provided.


