Monocular BEV Semantic Mapping for Tight-Space Maritime Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies have not effectively addressed the need for detailed, semantically segmented bird's eye view (BEV) semantic mapping in maritime environments, particularly in environments where data is scarce and costly, and the lack of data for effective spatial representation in maritime environments, particularly in maritime environments, where data is scarce and costly, and the lack of effective spatial representation for maneuvering vehicles in tight spaces.

Innovation Solution

A method and system using a monocular camera and artificial neural network (ANN) to generate a semantic map with a uniform scale from a first point of view (POV), and combining multiple semantic maps from different POV's to create a unified bird's eye view (BEV) semantic mapping semantic maps from different points of view (POV) semantic maps with a shared POV, and a system including a monocular camera and ANN to generate a semantic map with a uniform scale from a first point of view (POV) and combining semantic maps from different points of view (POV) semantic maps with a shared POV.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If monocular cameras are used to capture images in maritime environments, then device complexity is reduced, but measurement precision of spatial relationships deteriorates due to lack of depth information

Engineering Contradiction:
Improvecamera system complexityVSAvoidspatial relationship measurement precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transforms 2D image data from monocular cameras into a 3D bird's eye view representation by introducing a virtual third dimension (depth/elevation). The neural network processes sequential 2D images and reconstructs 3D spatial relationships, converting the limitation of monocular vision into a comprehensive spatial map that includes depth information without requiring multiple cameras.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an intermediary computational model (neural network) that acts as a mediator between the monocular camera input and the desired 3D spatial understanding. This intermediary processes the limited 2D visual information and generates intermediate representations (optical flow, depth maps, BEV features) that reconstruct the missing depth information through computational inference.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple cameras are used to generate comprehensive environmental maps, then measurement precision improves, but device complexity and cost increase

Engineering Contradiction:
Improveenvironmental mapping precisionVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex task of 3D environmental mapping into multiple processing stages: first extracting 2D features from individual monocular images, then computing optical flow between sequential frames, generating depth estimates, and finally assembling these into a bird's eye view representation. This segmentation allows a single camera to achieve multi-camera-level mapping precision through systematic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent achieves comprehensive 3D environmental understanding by transforming sequential 2D images into a 3D bird's eye view representation. By processing images over time and introducing the temporal dimension, the system reconstructs depth and spatial relationships that would normally require multiple simultaneous cameras, thereby reducing device complexity while maintaining mapping precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of operation

If bird's eye view representations are generated from monocular images, then navigation utility is improved, but information loss occurs during perspective transformation

Engineering Contradiction:
Improvenavigation utilityVSAvoidspatial information loss
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent employs feedback mechanisms where the neural network continuously refines its bird's eye view predictions by comparing expected features with actual observed features in sequential frames. The optical flow computation and depth estimation are adjusted based on feedback from matching features across frames, ensuring that the transformation from monocular view to bird's eye view preserves critical spatial information while maximizing navigation utility.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary extraction of spatial features and depth information from monocular images before completing the bird's eye view transformation. By pre-computing optical flow, depth maps, and feature correspondences, the system preserves essential spatial information that would otherwise be lost during perspective transformation, ensuring accurate navigation utility in the final BEV representation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12633126B2Bird's eye view (BEV) semantic mapping systems and methods using monocular camera
Publication Date: 2026.05.19 RAYMARINE UK
  • US12633126B2 patent drawing
  • US12633126B2 patent drawing
  • US12633126B2 patent drawing

AI summary

Bird's eye view (BEV) semantic mapping systems and methods are provided. A method includes receiving an image captured by a monocular camera having a first point of view (POV) of an environment including a plurality of features. The method further includes processing, by an artificial neural network (ANN), the captured image to generate a semantic map for the captured image, the semantic map associated with a second POV different from the first POV. The features exhibit a uniform scale in the semantic map. Additional methods and associated systems are also provided.