Ship navigation sensing method and system based on multi-source information fusion

By employing techniques such as Lie algebra registration, cross-modal Transformer attention, and strong tracking unscented Kalman filtering, the problems of accurate spatiotemporal registration and deep semantic fusion of multi-source heterogeneous data in ship navigation perception systems have been solved, improving perception accuracy and robustness, and ensuring real-time performance and decision security.

CN121934097APending Publication Date: 2026-04-28WUCHUANG TIANHENG ZHIHANG TECHNOLOGY (WUHAN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUCHUANG TIANHENG ZHIHANG TECHNOLOGY (WUHAN) CO LTD
Filing Date
2026-01-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing ship navigation perception systems are inadequate in terms of accurate spatiotemporal registration of multi-source heterogeneous data, cross-modal feature fusion, and robust decision-making, resulting in poor accuracy, low robustness, and insufficient real-time performance of environmental perception results.

Method used

We employ a Lie algebra-based time registration algorithm and a cross-modal Transformer attention mechanism, combined with strong tracking unscented Kalman filtering and information entropy correction DS evidence theory, to achieve spatiotemporal benchmark unification and deep semantic fusion of multi-source heterogeneous data. The system's reliability and real-time performance are ensured through ROS distributed architecture and Docker containerization technology.

Benefits of technology

It achieves high-precision spatiotemporal unification and deep semantic fusion, which improves the ability to perceive weak targets in harsh sea conditions, reduces the false alarm rate, and enhances the system's fault tolerance and self-healing capabilities as well as the decision-making safety of assisted driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934097A_ABST
    Figure CN121934097A_ABST
Patent Text Reader

Abstract

The invention provides a ship navigation sensing method and system based on multi-source information fusion. The system comprises a multi-source heterogeneous sensor group, an edge computing terminal and a ship-shore communication gateway. The method comprises the steps of collecting multi-source data of a laser radar, a camera, an AIS and the like; performing space-time registration by using a Lie algebra-based continuous trajectory interpolation algorithm, and eliminating point cloud distortion caused by ship motion; a cross-modal Transform neural network is constructed, and the two-dimensional texture and the three-dimensional geometric features are fused; carrying out nonlinear state estimation by adopting strong tracking unscented Kalman filtering with a self-adaptive attenuation factor; and processing sensor conflicts through a D-S evidence theory based on information entropy correction, and generating an environment perception result. According to the method, ship-shore-cloud-side cooperation is realized, the problems of data space-time asynchronization and sensor conflict in a complex water area are effectively solved, and the intelligence and safety of ship navigation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ship navigation perception, and in particular to a ship navigation perception method and system based on multi-source information fusion. Background Technology

[0002] With the rapid development of intelligent shipping technology, unmanned and assisted navigation systems for ships place higher demands on their ability to perceive the surrounding environment. Modern intelligent ships typically carry multiple types of sensors, such as millimeter-wave radar, LiDAR, high-definition cameras, AIS (Automatic Identification System), and GPS / IMU integrated navigation systems, to acquire comprehensive navigation environment information. However, the data generated by these sensors is multi-source, heterogeneous, asynchronous, and non-coplanar. How to efficiently and accurately fuse this data to construct a unified and reliable navigation situational awareness result is a major challenge currently facing the technology field.

[0003] In existing technical solutions, some attempts have been made to fuse multi-source information for ships. For example, prior art 1 (publication number CN111507429B) proposes a ship-side fusion method for multi-source sensing data of intelligent ships. Although this method combines ship-side and shore-based information and uses Kalman filtering for state estimation, it often uses relatively simple linear interpolation for time synchronization when processing sensor data. Moreover, the filtering algorithm lacks the ability to adaptively adjust for system model deviations. When the ship makes large-scale maneuvers or encounters severe sea conditions that cause abrupt changes in the statistical characteristics of sensor noise, filtering divergence or tracking loss can easily occur.

[0004] Prior art 2 (publication number CN117109588B) discloses a multi-source detection and multi-target information fusion method for intelligent navigation, which introduces interactive multi-model unscented Kalman filtering (IMMUKF) to improve tracking accuracy. However, this type of technology mainly relies on back-end track association at the feature fusion level, lacking deep feature interaction at the front end. When facing complex scenes with drastic changes in illumination or weak radar cross-sections, relying solely on single-modal feature extraction often makes it difficult to accurately identify target types and contours, and it does not fully utilize the complementarity between vision and point clouds in semantic and geometric spaces.

[0005] Existing technology 3 (publication number CN112857360B) relates to a method for fusing multiple information about ship navigation, focusing on the data storage structure and logical judgment. However, when dealing with "evidence conflicts" between different sensors (e.g., radar detects a target but the camera does not identify it), it often adopts a simple weighted average or fixed priority strategy, which makes it difficult to make accurate judgments when the confidence of the sensors changes dynamically, resulting in a high false alarm rate of the system.

[0006] Existing technology 4 (publication number CN117152572B) proposes a multi-radar image fusion method suitable for ship navigation, which mainly focuses on pixel-level image overlay and colorization. Although this method is intuitive, it is a shallow fusion method, lacking a deep understanding of environmental semantics, and fails to perform high-dimensional feature-level fusion of radar data with AIS and visual semantic information. It is difficult to meet the needs of intelligent ships for refined description of the surrounding situation (such as three-dimensional contours and precise categories).

[0007] In summary, existing ship navigation perception systems still have significant shortcomings in areas such as accurate spatiotemporal registration of multi-source heterogeneous data (especially the elimination of nonlinear errors under high-speed motion), deep feature fusion of cross-modal data (such as semantic complementarity between point clouds and images), and robust decision-making in high-conflict environments. There is an urgent need for a comprehensive perception method and system that can achieve high-precision spatiotemporal unification, deep semantic fusion, and adaptive robust tracking. Summary of the Invention

[0008] The main objective of this invention is to provide a ship navigation perception method and system based on multi-source information fusion, which solves the technical problems of poor accuracy, low robustness and insufficient real-time performance of existing ship navigation perception technologies in complex dynamic environments. These problems are caused by low spatiotemporal reference alignment accuracy of multi-source heterogeneous sensors, insufficient cross-modal feature fusion depth and weak ability of multi-target tracking algorithms to handle abrupt changes and evidence conflicts.

[0009] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a ship navigation perception method based on multi-source information fusion, the method comprising: S1. Collect multi-source heterogeneous data on the navigation environment through various sensors deployed on the ship; S2. Construct a distributed message communication architecture based on the ROS robot operating system, encapsulate the data streams of multiple types of sensors into independent ROS nodes, and aggregate the data to the central computing unit through a publish and subscribe mechanism; S3. Using a time registration algorithm based on Lie algebra and a spatial alignment algorithm based on extrinsic calibration matrix, spatiotemporal benchmarks are unified for multi-source heterogeneous data. S4. Construct a deep neural network model that includes a cross-modal Transformer attention mechanism. Input the registered LiDAR point cloud data and camera image data into the model, extract and fuse environmental features, and generate an intermediate perception vector that includes target category, 3D position and geometric contour. S5. Input the intermediate sensing vector, millimeter-wave radar data, and AIS data into the strong tracking unscented Kalman filter algorithm module with adaptive attenuation factor to perform multi-target state estimation and output a target dynamic tracking list. S6. Introduce the DS evidence theory based on information entropy correction, calculate the confidence level of each target in the target dynamic tracking list, and generate a navigable area raster map with confidence level gradient by combining electronic river map data.

[0010] In the preferred embodiment, the various types of sensors in step S1 include millimeter-wave radar, lidar, camera, AIS receiver, integrated navigation system, ultrasonic sensor, and anemometer. The specific execution steps for building the distributed message communication architecture based on the ROS robot operating system in step S2 are as follows: Initialize the ROS Master node in a Linux kernel-based edge computing device; Load the hardware drivers for each sensor and instantiate a Publisher node for each sensor; Define a custom message type Msg, where LiDAR data is defined as PointCloud2 format, image data is defined as ImageTransport format, and AIS data is defined as a structure format containing MMSI code, latitude and longitude, and heading angle. Configure a multi-threaded Spinner listener to achieve zero-copy data transmission between different nodes through a shared memory mechanism, and map all sensor data stream topics to Ethernet interfaces for full-duplex transmission.

[0011] In the preferred embodiment, the method also includes a system deployment step based on Docker containerization technology: Write a Dockerfile configuration file to pull a base image containing the CUDA parallel computing library, cuDNN neural network acceleration library, ROSNoetic middleware, and OpenCV vision library; Compile and install custom perception algorithm dependency libraries in the image to build an independent perception service container; Write a docker-compose.yml orchestration file to encapsulate the Publisher node and Spinner listener instantiated in step S2 into an independent container service, and define the container startup order and CPU / memory resource limits in the file; Configure the Host Network mode to allow container processes to directly reuse the host machine's network protocol stack and access the physical network interface in a bypass manner, enabling low-latency communication with shipboard sensor hardware.

[0012] In a preferred embodiment, the method further includes an online fault diagnosis and self-recovery step for the sensor: Start a separate watchdog daemon process to poll the heartbeat signals of each ROS node established in step S2 at a frequency of 10Hz; Statistically analyze the timestamps of data frames in each sensor data stream topic, calculate the time interval between adjacent frames, and deduce the frame drop rate; When the frame loss rate calculated by a certain sensor exceeds 15%, the sensor is determined to be in a sub-healthy state, and the weight coefficient of the sensor's data source in subsequent fusion calculations is automatically reduced. If a sensor node becomes unresponsive for more than 2 seconds, the Kill instruction of that node's process is executed and the Docker container running that node is restarted, thus achieving self-recovery from the fault.

[0013] In the preferred scheme, the specific steps of the Lie algebra-based time registration algorithm in step S3 are as follows: Read the high-frequency pose data output by the integrated navigation system and construct the continuous-time SE(3) Lie group trajectory; For each sampling point in the lidar point cloud, B-spline interpolation is performed on the SE(3) Lie group trajectory based on the timestamp of its collection time to calculate the instantaneous pose of the ship at that moment; The pose transformation on SE(3) is converted into a tangent vector on the Lie algebra of SE(3) using logarithmic mapping. After linear interpolation in the tangent space, the pose is restored to SE(3) through exponential mapping. The instantaneous pose obtained by interpolation is used to eliminate the point cloud distortion caused by ship motion, and all sensor data are uniformly projected onto the carrier coordinate system of the current frame.

[0014] In the preferred embodiment, the method further includes a step of calling the PCL point cloud library for data preprocessing: Obtain the distortion-free point cloud data stream output after spatiotemporal registration in step S3; Instantiate a PassThrough filter object in the PCL library, set the coordinate threshold ranges for the X, Y, and Z axes, and filter out invalid point cloud data caused by the obstruction of the ship's own structure; Instantiate a StatisticalOutlierRemoval filter object, and iterate through the point cloud data to calculate the Euclidean distance from each point to its K nearest neighbors. Calculate the mean and standard deviation of the distance distribution, and remove isolated noise points whose distance is greater than the mean plus twice the standard deviation; The VoxelGrid filter is invoked, and the voxel leaf size is set to 0.05 meters to 0.1 meters to downsample the point cloud, thereby reducing the data throughput of the subsequent deep learning model.

[0015] In the preferred embodiment, the specific steps in step S4 for constructing a deep neural network model incorporating a cross-modal Transformer attention mechanism are as follows: Deploy the TensorRT inference acceleration engine to convert the weights of the PyTorch-trained model into a Plan file in FP16 half-precision format; The image feature extraction branch and the point cloud feature extraction branch are loaded in parallel using CUDA streaming technology; The image branch uses a ResNet-50 backbone network to extract two-dimensional texture features, while the point cloud branch uses a VoxelNet voxel network to extract three-dimensional geometric features. Construct a cross-modal Transformer fusion layer to flatten two-dimensional texture features into query vectors, project three-dimensional geometric features, and flatten them into key and value vectors; Calculate the multi-head self-attention matrix, normalize it using the Softmax function to obtain the attention weights, perform weighted aggregation on heterogeneous features, and output the fused multimodal feature map.

[0016] In a preferred embodiment, the method further includes the step of constructing semantic constraints using a high-precision map: Parse high-precision map files in OpenDRIVE format and extract data on channel centerlines, no-navigation zone polygons, and water depth contour lines; The vector data of the high-precision map is rasterized and mapped to the same coordinate system as the locally perceptual raster map; Construct a Bayesian inference network, using the static semantic information provided by the map as prior probabilities; The target category probability output by the deep learning model in step S4 is corrected using prior probability. When the model detection result conflicts with the map semantic constraints, the map semantic constraints are given priority to suppress false detections.

[0017] In the preferred embodiment, the method further includes historical data playback and algorithm iterative optimization steps: Enable rosbag data recording service, configure LZ4 compression mode, and write all topic data to the ship's SSD hard drive in real time. When a ship docks and connects to a high-bandwidth network, the rsync incremental synchronization script is automatically triggered to upload the rosbag file to the shore-based training server. Data is decompressed on the shore-based server, and difficult sample examples are labeled using manual annotation tools; Add the labeled new samples to the training set and fine-tune the deep neural network model in step S4. The updated model weight file is pushed to the ship via OTA (Over-The-Air) download technology, replacing the old version of the Plan file.

[0018] In the preferred embodiment, the execution logic of the strong tracking unscented Kalman filter algorithm module with adaptive attenuation factor in step S5 is as follows: Establish the nonlinear state equations and measurement equations for ship motion; Sigma point sets are generated using UT transformation, and state prediction and measurement prediction values ​​are calculated. The orthogonality of the residual sequences is calculated in real time, and the new information covariance matrix is ​​constructed. By introducing a strong tracking fading factor, when the orthogonality of the residual sequence deteriorates, the filter is forced to track the abrupt change in the system by increasing the weight of the prediction error covariance matrix. The Kalman gain matrix is ​​modified by using a fading factor to update the state estimate and the posterior error covariance matrix, thereby suppressing filter divergence.

[0019] In the preferred scheme, the specific steps of the DS evidence theory based on information entropy correction in step S6 are as follows: Construct a multi-sensor recognition framework and assign a basic probability assignment function (BPA) to each sensor; Calculate the Jousselme distance between any two sources of evidence and construct an evidence distance matrix; The support of each piece of evidence is calculated based on the evidence distance matrix, and the support is normalized to the credibility weight of the evidence. Calculate the information entropy of each evidence source. For evidence sources whose information entropy is greater than a preset threshold, correct their BPA using a credibility weight. The Dempster composition rule is applied to perform orthogonal summation on the modified BPA. When the conflict coefficient k approaches 1, the Yager composition rule is automatically switched to perform fusion, and the final target attribute determination result is output.

[0020] In a preferred embodiment, the method further includes a visualization rendering step based on OpenCV and OpenGL: Create a Qt graphical user interface application main thread, and initialize the OpenCV environment for image processing and the OpenGL context for 3D rendering in the main thread; The cv::Mat constructor of the OpenCV library is called to instantiate a matrix container, read the navigable area raster map data with confidence gradient generated in step S6, map the confidence values ​​in the raster map to single-channel grayscale values ​​from 0 to 255, and write them into the matrix container. Parse the target dynamic tracking list output in step S6, and extract the category ID, 3D center coordinates, and bounding box size parameters for each target in the list; Write vertex and fragment shaders using the OpenGL shader language GLSL; Convert the raster map data stored in the cv::Mat matrix container into a texture map, and convert the extracted target 3D center coordinates and bounding box size parameters into a vertex coordinate array; During the rendering loop, GPU hardware acceleration is used to blend texture maps and vertex arrays and draw them into the frame buffer, which is then output to the ship's bridge display via an HDMI interface.

[0021] In the preferred embodiment, the method further includes a cloud-edge collaboration step based on the MQTT protocol: Start the MQTT client on the ship's edge computing terminal and connect to the MQTT Broker agent server of the shore-based cloud platform; Define various perception data packets in JSON format, including target ID, latitude and longitude, heading, speed, and confidence level; The LZ4 compression algorithm is used to perform binary compression on the data packets; The frequency of data transmission is dynamically adjusted based on the QoS (Quality of Service) level of the current network bandwidth. Subscribe to the global traffic situation topics published by the shore-based cloud platform, analyze the long-distance ship position data issued by the shore-based platform, and input it as a virtual sensor into the Kalman filter algorithm in step S5.

[0022] A ship navigation perception system based on multi-source information fusion, the system comprising: The multi-source heterogeneous sensor group includes a lidar for acquiring point cloud data, a camera for acquiring visual data, an AIS receiver for acquiring target navigation data, and a GPS / IMU integrated navigation system for acquiring the ship's attitude. The edge computing terminal has a built-in NVIDIA GPU accelerator card and FPGA preprocessing chip. The memory of the edge computing terminal stores computer programs. When the computer programs are executed by the processor, they implement the steps of ship navigation perception based on multi-source information fusion. The ship-to-shore communication gateway is equipped with a 5G communication module and a satellite communication module to enable data interaction with the shore-based cloud platform; The human-machine interaction terminal connects to the edge computing terminal via an HDMI interface and is used to display environmental perception results and receive driver commands.

[0023] A computer-readable storage medium storing a computer program that, when executed by a processor, implements a ship navigation perception method based on multi-source information fusion.

[0024] This invention provides a ship navigation perception method and system based on multi-source information fusion. Compared with existing technologies, this application introduces a Lie algebra-based time registration algorithm and performs B-spline interpolation on the Lie group manifold, effectively eliminating point cloud distortion caused by high-speed nonlinear motion of the ship and achieving microsecond-level precision in spatiotemporal reference unification. It employs a cross-modal Transformer attention mechanism, overcoming the limitations of traditional shallow stitching and achieving deep interaction between visual texture and point cloud geometric features, significantly improving the perception capability of weak targets in adverse sea conditions. It utilizes a strong tracking unscented Kalman filter algorithm with an adaptive decay factor, and solves the target loss problem under abrupt maneuvers by online correction of the prediction covariance through a fading factor. Combined with DS evidence theory based on information entropy correction, it scientifically quantifies sensor uncertainty and dynamically reduces the weight of high-conflict evidence, effectively lowering the false alarm rate. Furthermore, the system design based on ROS distributed architecture and Docker containerization, coupled with intuitive confidence gradient visualization, significantly enhances the system's fault tolerance and self-healing capabilities and the safety of assisted driving decision-making. Attached Figure Description

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart of the ship navigation perception method of the present invention; Figure 2 This is a block diagram of the ship navigation perception system of the present invention; Figure 3 This invention relates to the observation of ships navigating in inland waterways. Figure 1 ; Figure 4 This invention relates to the observation of ships navigating in inland waterways. Figure 2 ; Figure 5 This is a system interface diagram of the vessel navigating in inland waterways according to the present invention. Detailed Implementation

[0026] Example 1 like Figure 1-5 As shown, a ship navigation perception method based on multi-source information fusion is proposed, which includes: S1. Collect multi-source heterogeneous data on the navigation environment through various sensors deployed on the ship; The various sensors in step S1 include millimeter-wave radar, lidar, camera, AIS receiver, integrated navigation system, ultrasonic sensor and anemometer. S2. Construct a distributed message communication architecture based on the ROS robot operating system, encapsulate the data streams of multiple types of sensors into independent ROS nodes, and aggregate the data to the central computing unit through a publish and subscribe mechanism; S3. Using a time registration algorithm based on Lie algebra and a spatial alignment algorithm based on extrinsic calibration matrix, spatiotemporal benchmarks are unified for multi-source heterogeneous data. S4. Construct a deep neural network model that includes a cross-modal Transformer attention mechanism. Input the registered LiDAR point cloud data and camera image data into the model, extract and fuse environmental features, and generate an intermediate perception vector that includes target category, 3D position and geometric contour. S5. Input the intermediate sensing vector, millimeter-wave radar data, and AIS data into the strong tracking unscented Kalman filter algorithm module with adaptive attenuation factor to perform multi-target state estimation and output a target dynamic tracking list. S6. Introduce the DS evidence theory based on information entropy correction, calculate the confidence level of each target in the target dynamic tracking list, and generate a navigable area raster map with confidence level gradient by combining electronic river map data.

[0027] The following is a detailed description of the implementation of the technical solution you provided. This description aims to fully disclose the technical details, algorithm principles, and parameter meanings of the present invention, ensuring that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification. All explanations in parentheses have been removed, and a direct statement approach has been adopted. Formulas are output in LaTeX format.

[0028] This invention proposes a ship navigation perception method and system based on multi-source information fusion. First, it collects multi-source heterogeneous data about the navigation environment using various sensors deployed on the ship. These sensors specifically include millimeter-wave radar for detecting mid-to-long-range targets, lidar for constructing a high-precision 3D environment, cameras for acquiring visual semantic information, an AIS receiver for receiving dynamic information from other ships, a combined navigation system for providing high-frequency pose data of the ship, ultrasonic sensors for short-range blind spot detection, and anemometers for weather assistance. To achieve efficient data flow and processing, this embodiment constructs a distributed message communication architecture based on the ROS robot operating system. In this architecture, the system loads hardware drivers for each sensor and encapsulates the data stream of each type of sensor into an independent ROS node. Utilizing the publish / subscribe mechanism provided by the ROS system, each sensor node publishes data to a designated topic and aggregates it to the central computing unit via shared memory or Ethernet transmission. This distributed architecture design effectively reduces system coupling; when one sensor fails, it does not cause the entire system to collapse, thereby improving system reliability.

[0029] After data aggregation, it is necessary to unify the spatiotemporal references of the multi-source heterogeneous data. To address the error problem of traditional linear interpolation under high-speed or rolling motions of ships, this embodiment utilizes a Lie algebra-based time registration algorithm for high-precision interpolation. Specifically, the ship's pose belongs to a Lie group. In spatial cases, direct linear interpolation of the rotation matrix leads to loss of orthogonality. Therefore, this method first reads the high-frequency pose data output by the integrated navigation system to construct a continuous-time trajectory. For each sampling point in the lidar point cloud, the trajectory is constructed based on its precise acquisition time timestamp. The instantaneous pose at that moment is calculated using the B-spline interpolation algorithm. The interpolation process in Lie algebras The calculation is performed in the tangent space, and the formula is as follows: ; in, Indicates time pose matrix, express B-order spline basis functions, This indicates the position and orientation of the control points in the integrated navigation system. Describes the logarithmic mapping operator from Lie groups to Lie algebras. This represents the exponential mapping operator from Lie algebras to Lie groups. This algorithm can accurately recover the ship's attitude at the moment of point cloud acquisition, eliminating point cloud motion distortion caused by ship movement. Simultaneously, combined with a spatial alignment algorithm based on extrinsic calibration matrices, all sensor data are uniformly projected onto the ship's coordinate system using pre-calibrated rotation and translation matrices, laying a precise geometric foundation for subsequent feature fusion.

[0030] After completing spatiotemporal alignment, this embodiment constructs a deep neural network model incorporating a cross-modal Transformer attention mechanism to address the challenges of sparse and semantically unsound LiDAR point clouds and lack of depth information in camera images. The model receives registered LiDAR point cloud data and camera image data as input. Internally, a two-stream backbone network is used to extract geometric features from the point cloud and texture features from the image, which are then fused through a cross-modal Transformer module. In this module, image features are flattened and mapped to query vectors. Point cloud features are mapped to key vectors Sum value vector The formula for calculating cross-modal attention weights is as follows: ; in, The query matrix generated for image features The key matrix generated for point cloud features. The value matrix generated for point cloud features. The scaling factor for the feature dimension. This is a normalized exponential function. This formula adaptively assigns weights to different modal features by calculating the inner product correlation between image features and point cloud features, thereby achieving deep complementarity between semantic and geometric information at the feature level. The model ultimately outputs an intermediate perception vector containing the target category, 3D center location, and geometric contour dimensions (length, width, and height). This step significantly improves the system's robustness under low-light conditions such as nighttime, rain, and fog, effectively preventing missed detections and false detections.

[0031] After obtaining the intermediate sensing vector, the system inputs it, along with millimeter-wave radar data and AIS data, into the ST-UKF algorithm module, which employs a strong tracking unscented Kalman filter with an adaptive attenuation factor. This module aims to address the target tracking divergence problem caused by model mismatch during ship maneuvers. The algorithm introduces a strong tracking attenuation factor based on the standard unscented Kalman filter. The prediction error covariance matrix is ​​corrected in real time. The corrected prediction covariance matrix is ​​then applied. The calculation formula is as follows: ; in, Here is the state transition matrix. Let $\mathbf{a}$ be the posterior error covariance of the previous time step. The process noise covariance. The fading factor. The calculation is based on the principle of orthogonality of residual sequences. When a significant increase in the mean of the residuals is detected, i.e., a sudden change in the system, It automatically increases, forcing the filter to rely more on current observation data than historical predictions. This mechanism ensures that the algorithm can quickly respond to the dynamic behaviors of ships, such as steering and speed changes, and output a high-precision target dynamic tracking list.

[0032] Finally, to address potential conflicting evidence among multiple sensors, this embodiment introduces a DS evidence theory based on information entropy correction. The system calculates the confidence level of each target in the target dynamic tracking list and combines it with electronic river map data to generate the final navigable area grid map. During the fusion process, a basic probability assignment function (BPA) is first constructed for each sensor. To measure the uncertainty of the evidence, information entropy is introduced. The calculation formula is as follows: ; in, This indicates that the evidence points to the proposition. Basic probability assignment, To identify the framework, the system automatically reduces the weight of sensor evidence with high information entropy in the fusion rule. Then, Dempster's synthesis rule is used for fusion; if the conflict coefficient between sensors... If the value is too large, the system automatically switches to the Yager synthesis rule. This step effectively eliminates false targets and, based on the static water depth and channel boundaries provided by the electronic river chart, overlays the confidence distribution of dynamic targets to generate a navigable area grid map with confidence gradients. This provides an intuitive and reliable safety boundary reference for ship navigation decisions. This entire process, from data acquisition to final decision-making, is progressive and solves the technical problems of low ship perception accuracy and poor robustness in complex environments.

[0033] In the preferred embodiment, the specific execution steps for constructing the distributed message communication architecture based on the ROS robot operating system in step S2 are as follows: Initialize the ROS Master node in a Linux kernel-based edge computing device; Load the hardware drivers for each sensor and instantiate a Publisher node for each sensor; Define a custom message type Msg, where LiDAR data is defined as PointCloud2 format, image data is defined as ImageTransport format, and AIS data is defined as a structure format containing MMSI code, latitude and longitude, and heading angle. Configure a multi-threaded Spinner listener to achieve zero-copy data transmission between different nodes through a shared memory mechanism, and map all sensor data stream topics to Ethernet interfaces for full-duplex transmission.

[0034] In step S2 of this embodiment, an efficient and modular distributed message communication architecture is first constructed. The basic operating environment for this architecture uses Linux kernel-based edge computing devices, such as industrial control computers or embedded AI computing modules running the Ubuntu operating system. During system startup, the ROSMaster master node is first initialized in the operating system's user space. The ROS Master, acting as a name service and parameter server, is responsible for recording the registration information of all active nodes in the system, including node names, published or subscribed topic names, and message types, thereby establishing a point-to-point network connection topology between nodes. Subsequently, the system loads the underlying hardware drivers for various types of sensors. These drivers convert physical signals into digital signals by calling the SDK interfaces provided by sensor manufacturers or reading raw binary data from serial ports and network ports. For each independent sensor device, the system instantiates a Publisher node within the ROS framework. The role of the Publisher node is to encapsulate the converted sensor data into a standard message format and broadcast the data to a specific Topic, thereby decoupling data acquisition from data processing.

[0035] To address the data characteristics of different sensors, this embodiment strictly defines a custom message type, Msg. For LiDAR data with large volumes and containing 3D spatial information, the system defines it as PointCloud2 format. PointCloud2 format internally stores point cloud data using binary large objects, including spatial coordinates (x, y, z) and reflection intensity fields, supporting the storage of both unordered and structured point clouds. For visual data acquired by cameras, the system uses ImageTransport format for encapsulation. ImageTransport not only supports the transmission of raw RGB image data but also integrates an image compression transmission mechanism, automatically switching to JPEG or PNG compression streams when bandwidth is limited. For AIS data, the system defines a structure format containing MMSI code, latitude and longitude coordinates, and heading angle. The MMSI code uniquely identifies the ship, latitude and longitude are stored as double-precision floating-point numbers to ensure positioning accuracy, and the heading angle describes the ship's instantaneous direction of motion. This strongly typed data definition method ensures the consistency and standardization of heterogeneous data transmission within the system.

[0036] To meet the stringent real-time requirements of ship navigation sensing, the system employs a multi-threaded Spinner listener. Traditional single-threaded listeners are prone to blocking when processing high-frequency sensor data, leading to data backlog or frame drops. The multi-threaded Spinner, by enabling multiple execution threads, can process message queues from different callback functions in parallel, significantly improving concurrent processing capabilities. At the inter-node data transmission level, this embodiment adopts a zero-copy data transmission mechanism based on shared memory. In traditional network protocol stack communication, data needs to be copied multiple times between user space and kernel space, consuming significant CPU resources. The zero-copy mechanism maps the virtual address spaces of different nodes to the same physical memory region, allowing the receiving node to directly read data written by the sending node, avoiding the serialization and deserialization processes, thereby reducing data transmission latency to the microsecond level. Simultaneously, the system maps all sensor data stream topics to Ethernet interfaces and configures the network adapter to operate in full-duplex mode, ensuring that data transmission and reception do not interfere with each other, achieving high-throughput data exchange.

[0037] The beneficial effects of this embodiment are as follows: First, by constructing a ROS-based distributed communication architecture, complete decoupling of sensor hardware drivers and upper-layer perception algorithms is achieved, greatly improving the system's scalability and maintainability. When the sensor model needs to be changed, only the corresponding publisher node needs to be updated, without modifying the core algorithm code. Second, by adopting a multi-threaded Spinner listener combined with a shared memory zero-copy mechanism, the latency and congestion problems of high-bandwidth sensor data such as LiDAR and cameras during transmission are solved, significantly reducing the CPU load and ensuring the real-time response capability of the perception system in high-speed navigation scenarios. Third, by defining standardized PointCloud2, ImageTransport, and custom structure message types, the interface standard for multi-source heterogeneous data is standardized, providing a unified and high-quality data input for subsequent data fusion algorithms and avoiding parsing errors caused by data format chaos.

[0038] In the preferred embodiment, the method also includes a system deployment step based on Docker containerization technology: Write a Dockerfile configuration file to pull a base image containing the CUDA parallel computing library, cuDNN neural network acceleration library, ROSNoetic middleware, and OpenCV vision library; Compile and install custom perception algorithm dependency libraries in the image to build an independent perception service container; Write a docker-compose.yml orchestration file to encapsulate the Publisher node and Spinner listener instantiated in step S2 into an independent container service, and define the container startup order and CPU / memory resource limits in the file; Configure the Host Network mode to allow container processes to directly reuse the host machine's network protocol stack and access the physical network interface in a bypass manner, enabling low-latency communication with shipboard sensor hardware.

[0039] In practice, the first step is to write a Dockerfile configuration file, which serves as the blueprint for building the image and defines the basic environment required for the perception system to run. The system pulls an official NVIDIA base image from the image repository, containing the CUDA parallel computing library and the cuDNN neural network acceleration library, providing underlying GPU computing power support for the deep learning model described in the claims. Based on this, the ROS Noetic middleware and OpenCV vision library are overlaid to build a full-stack development and runtime environment. During the image building phase, developers copy the custom perception algorithm source code and its dependent third-party library files to the image file system and execute compilation instructions, thereby building an independent perception service container that does not depend on the host operating system. This encapsulation method ensures the consistency of the algorithm environment and avoids dependency conflicts, i.e., the DLL hell problem, caused by differences in the onboard computer's operating system version.

[0040] To achieve efficient collaborative management of multiple nodes, this embodiment uses a `docker-compose.yml` orchestration file. In this file, the Publisher node instantiated in step S2, the Spinner listener, and the subsequent core algorithm modules are each encapsulated as independent services. The orchestration file uses the `depends_on` parameter to strictly define the container startup order, ensuring that the underlying sensor driver container starts first and successfully connects to the hardware before the upper-layer perception algorithm container begins initialization, preventing process crashes due to missing data sources. Simultaneously, the `resources` property under the `deploy` field in the file hard-limits the number of CPU cores and memory size used by each container. For example, the non-critical logging container is limited to using only 10% of CPU resources, preventing individual processes from exhausting system resources due to abnormal infinite loops and ensuring the stability of the core navigation and collision avoidance process.

[0041] Regarding network communication configuration, this embodiment specifically configures a host network mode. In the default bridged mode, communication between containers and external networks requires going through a Network Address Translation (NAT) layer, which introduces additional network latency and increases CPU overhead. By setting `network_mode` to `host` in `docker-compose.yml`, the container process directly reuses the host machine's network protocol stack, no longer having its own independent IP address, but sharing the IP with the host machine. This means that the ROS nodes inside the container can directly listen to multicast or broadcast data on the physical network interface, just like native processes on the host machine. This is crucial for receiving high-frequency LiDAR UDP packets, enabling direct access to the physical network interface in a bypass manner, eliminating packet forwarding loss caused by the virtual bridge, and achieving low-latency communication with shipboard sensor hardware.

[0042] In a preferred embodiment, the method further includes an online fault diagnosis and self-recovery step for the sensor: Start a separate watchdog daemon process to poll the heartbeat signals of each ROS node established in step S2 at a frequency of 10Hz; Statistically analyze the timestamps of data frames in each sensor data stream topic, calculate the time interval between adjacent frames, and deduce the frame drop rate; When the frame loss rate calculated by a certain sensor exceeds 15%, the sensor is determined to be in a sub-healthy state, and the weight coefficient of the sensor's data source in subsequent fusion calculations is automatically reduced. If a sensor node becomes unresponsive for more than 2 seconds, the Kill instruction of that node's process is executed and the Docker container running that node is restarted, thus achieving self-recovery from the fault.

[0043] The following is a detailed description of the implementation method and an analysis of its beneficial effects for the "online fault diagnosis and self-recovery steps for sensors". This section aims to fully disclose the technical features in the claims, ensuring that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification.

[0044] To ensure the continuity and reliability of sensing during system operation, this embodiment integrates a comprehensive online sensor fault diagnosis and self-recovery mechanism. Specifically, the system starts a watchdog process in the background, independent of the main sensing thread. This process, as a high-priority monitoring task, utilizes the node manager interface provided by ROS, setting the sampling frequency to 10Hz, that is, initiating a heartbeat detection request every 100 milliseconds to all ROS nodes established in step S2. If the detected node returns a response signal within the specified time, it is determined that the node is alive; otherwise, a heartbeat loss event is recorded.

[0045] In addition to monitoring node liveness, this embodiment also delves into data quality, performing statistical analysis on the timestamps of data frames in each sensor data stream topic. The system reads the timestamp field from the header of each received message frame in real time, recording the current frame timestamp as... The timestamp of the previous frame is Calculate the time interval between two adjacent frames. Combined with the standard period corresponding to the theoretical sampling frequency of the sensor. Statistics during the sliding window time frame drop rate within The formula for calculating the frame drop rate is as follows: ; in, This represents the total number of frames within the window. This is an indicator function used to determine whether the time interval significantly exceeds the limit. The system uses the calculated frame drop rate... The system dynamically assesses the health of the sensors. When the calculated frame drop rate exceeds a preset safety threshold of 15%, the sensor is deemed to be in a sub-healthy state. At this point, although the sensor is still outputting data, the continuity and integrity of the data can no longer meet the requirements of high-precision sensing. The system will automatically trigger a degradation strategy. In subsequent Kalman filtering or DS evidence theory fusion calculations, an attenuation coefficient inversely proportional to the frame drop rate is introduced to automatically reduce the weight coefficient of the sensor's data source, preventing low-quality data from contaminating the overall fusion result.

[0046] For severe failures where sensor nodes completely fail, this embodiment employs strict self-recovery logic. When the watchdog daemon detects that a sensor node has not responded for a continuous two-second time window, it determines that the node has deadlocked or crashed. At this point, the daemon calls the operating system's process management interface to forcibly execute the kill instruction of the node's process, eliminating zombie processes. Subsequently, using the Docker container orchestration engine's API, a restart command is sent to the daemon to restart the Docker container running the node. Leveraging the second-level startup capability and environment consistency of Docker containers, the failed node can complete initialization and rejoin the ROS communication network in a very short time, achieving automatic fault repair and seamless business continuity.

[0047] In the preferred scheme, the specific steps of the Lie algebra-based time registration algorithm in step S3 are as follows: Read the high-frequency pose data output by the integrated navigation system and construct the continuous-time SE(3) Lie group trajectory; For each sampling point in the lidar point cloud, B-spline interpolation is performed on the SE(3) Lie group trajectory based on the timestamp of its collection time to calculate the instantaneous pose of the ship at that moment; The pose transformation on SE(3) is converted into a tangent vector on the Lie algebra of SE(3) using logarithmic mapping. After linear interpolation in the tangent space, the pose is restored to SE(3) through exponential mapping. The instantaneous pose obtained by interpolation is used to eliminate the point cloud distortion caused by ship motion, and all sensor data are uniformly projected onto the carrier coordinate system of the current frame.

[0048] The following is a detailed description of the implementation method and analysis of the beneficial effects of step S3, "Time Registration Algorithm Based on Lie Algebra". This section aims to fully disclose the technical features in the claims to ensure that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification.

[0049] In step S3 of this embodiment, to address the spatiotemporal asynchrony of sensor data caused by the ship's violent six-degree-of-freedom swaying under complex sea conditions, a high-precision continuous-time trajectory interpolation method based on Lie group and Lie algebra theory is employed. Although the integrated navigation system can output high-frequency pose data, its sampling time often does not coincide with the acquisition time of environmental perception sensors such as lidar and cameras. Furthermore, the ship's pose transformation matrix lies in a special Euclidean group. The manifold surface is a curved space, not a flat Euclidean space. Directly performing conventional linear interpolation on the rotation matrix elements in the pose matrix would violate the orthogonality constraints of the rotation matrix, leading to distortion of the calculated intermediate pose, making it physically unrealizable. Therefore, this embodiment first reads the time-series discrete pose data output by the integrated navigation system, using it as control points to construct a continuous-time... Li Queue Trajectory.

[0050] For a frame of point cloud data acquired by lidar, due to the mechanical or solid-state scanning characteristics of the laser beam, the actual acquisition time of each laser point within the frame is different. The system extracts the precise acquisition timestamp of each sampling point. The instantaneous pose of the ship at that moment is calculated using the B-spline interpolation algorithm. The specific interpolation operations are not performed directly in the curved Lie group space, but rather transformed into tangent space using the logarithmic and exponential mapping relationships between Lie groups and Lie algebras. First, the logarithmic mapping operator is used to... Transformation of pose matrix on a group into Tangent vectors on Lie algebras. The space of Lie algebras is a linear vector space that allows addition, subtraction, and scalar multiplication. Within the tangent space, based on timestamps... The time distance from adjacent control points is used to calculate weights using B-spline basis functions and then linear interpolation is performed to obtain the time interval. The corresponding Lie algebraic tangent vector. Then, the interpolated tangent vector is losslessly restored using the exponential mapping operator. The pose matrix on the group. The mathematical expression of this process is as follows: ; In the formula, Representative moment The instantaneous pose matrix of the ship obtained by interpolation, The pose of the reference control point, It is an exponential mapping function. This is the cumulative form of the B-spline basis functions. Let be the difference vector of the relative poses between control points in the Lie algebra space. The symbol represents the operation of mapping a six-dimensional vector to an antisymmetric matrix. Through the above calculations, the system obtains the precise pose of each laser point at the moment of acquisition. Next, the system calculates the instantaneous pose. With reference pose of the point cloud in that frame The relative transformation matrix between the two is used to project the laser points back to the coordinate system of the reference time, thereby eliminating the point cloud layering, ghosting and distortion caused by the ship's own motion, and completing high-precision motion distortion correction and spatiotemporal unification.

[0051] In the preferred embodiment, the method further includes a step of calling the PCL point cloud library for data preprocessing: Obtain the distortion-free point cloud data stream output after spatiotemporal registration in step S3; Instantiate a PassThrough filter object in the PCL library, set the coordinate threshold ranges for the X, Y, and Z axes, and filter out invalid point cloud data caused by the obstruction of the ship's own structure; Instantiate a StatisticalOutlierRemoval filter object, and iterate through the point cloud data to calculate the Euclidean distance from each point to its K nearest neighbors. Calculate the mean and standard deviation of the distance distribution, and remove isolated noise points whose distance is greater than the mean plus twice the standard deviation; The VoxelGrid filter is invoked, and the voxel leaf size is set to 0.05 meters to 0.1 meters to downsample the point cloud, thereby reducing the data throughput of the subsequent deep learning model.

[0052] The following is a detailed description of the implementation method and an analysis of its beneficial effects for the "step of calling the PCL point cloud library for data preprocessing". This section aims to fully disclose the technical features in the claims to ensure that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification.

[0053] After acquiring the distortion-free point cloud data stream output after spatiotemporal registration in step S3, a series of standardized preprocessing operations based on the PCL point cloud library were performed to further improve data quality and reduce the computational load of subsequent deep learning models. First, the system instantiates a PassThrough filter object from the PCL library in memory. Due to the physical limitations of the LiDAR installation location, the acquired raw point cloud often contains echoes from the ship's deck, mast, or antenna structures, which constitute background noise for environmental perception. The PassThrough filter constructs a cuboid region of interest (ROI) by setting coordinate threshold ranges for the X, Y, and Z axes. The algorithm iterates through each point in the input point cloud, determining whether its coordinate values ​​are within the set range. Points falling outside the coordinate threshold range, especially those within the ship's own geometric contour, are marked as invalid and directly removed from the point cloud list. This step quickly removes background data irrelevant to the navigation environment, significantly reducing the search space for subsequent algorithms.

[0054] After completing the region clipping, to eliminate sensor measurement noise and outliers caused by water mist and dust in the air, the system instantiates a StatisticalOutlierRemoval filter object. This filter cleans the point cloud based on statistical principles. The algorithm traverses the point cloud data, and for each query point... Using the Kd-Tree data structure to quickly search its neighborhood Find the nearest neighbor and calculate the distance from the queried point to this nearest neighbor. The average Euclidean distance between neighboring points The calculation formula is as follows: ; in, for The The nearest neighbor point, This represents the L2 norm. Next, the system statistically analyzes the average distance distribution of all points and calculates its global mean. and standard deviation Based on the Gaussian distribution assumption, the average distance of the vast majority of valid measurement points should fall near the mean. The system sets a threshold criterion to determine the average distance... Greater than Points that are not clearly defined as outliers are identified and removed. This statistical filtering method can effectively filter out sparse isolated points caused by radar multipath effects or airborne particles, while preserving the dense point cloud on the object's surface.

[0055] Finally, to address the issue of massive and unevenly distributed LiDAR point cloud data, the system utilizes a VoxelGrid filter for downsampling. This filter divides the 3D space into a series of tiny cubic units, or voxels, with the leaf size (side length) of each voxel set to between 0.05 meters and 0.1 meters. For each voxel containing all the point cloud data... The algorithm no longer retains all the original points, but instead calculates the geometric centroids of these points. This approximates all points within the voxel. The formula for calculating the centroid is as follows: ; in This represents the number of points within the voxel. In this way, the originally extremely dense near-range point cloud is sparsified, while far-range point clouds are preserved, resulting in a more uniform spatial distribution of point cloud density. The processed point cloud data volume can typically be reduced to 10% to 20% of the original data volume, thus matching the input tensor dimensionality constraints of subsequent deep neural network models.

[0056] The beneficial effects are as follows: First, the PassThrough filtering effectively removes the self-obstruction interference of the ship's own structure, preventing the ship's own components from being mistakenly detected as obstacles, thus improving the accuracy of perception.

[0057] Second, statistical filtering was used to remove false noise caused by sea surface moisture or waves, which significantly improved the signal-to-noise ratio of point cloud data, making the boundaries generated by subsequent target clustering and segmentation algorithms clearer.

[0058] Third, VoxelGrid voxel filtering was used to achieve data dimensionality reduction, which significantly reduced data throughput while preserving the geometric features of the environment, alleviating the memory pressure and computing bottleneck of the shipborne edge computing device, and ensuring the real-time operating frame rate of the perception system.

[0059] In the preferred embodiment, the specific steps in step S4 for constructing a deep neural network model incorporating a cross-modal Transformer attention mechanism are as follows: Deploy the TensorRT inference acceleration engine to convert the weights of the PyTorch-trained model into a Plan file in FP16 half-precision format; The image feature extraction branch and the point cloud feature extraction branch are loaded in parallel using CUDA streaming technology; The image branch uses a ResNet-50 backbone network to extract two-dimensional texture features, while the point cloud branch uses a VoxelNet voxel network to extract three-dimensional geometric features. Construct a cross-modal Transformer fusion layer to flatten two-dimensional texture features into query vectors, project three-dimensional geometric features, and flatten them into key and value vectors; Calculate the multi-head self-attention matrix, normalize it using the Softmax function to obtain the attention weights, perform weighted aggregation on heterogeneous features, and output the fused multimodal feature map.

[0060] The following is a detailed description of the implementation method and an analysis of its beneficial effects for step S4, "Constructing a deep neural network model including a cross-modal Transformer attention mechanism." This section aims to fully disclose the technical features in the claims, ensuring that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification.

[0061] In step S4, to address the challenge of depth alignment between visual and LiDAR data in the feature space during ship navigation, a high-performance heterogeneous feature fusion inference system was constructed. During the model deployment phase, the system first incorporates the NVIDIA TensorRT inference acceleration engine. Developers export the original model weight files trained using the PyTorch deep learning framework to the ONNX universal format. Then, TensorRT's builder optimizes and reconstructs the model graph. Through layer fusion and automatic kernel tuning, the model weights are converted into FP16 half-precision floating-point Plan serialization files. Compared to the traditional FP32 format, using FP16 reduces GPU memory usage by half and lowers memory access bandwidth requirements, thus significantly improving inference throughput while maintaining perception accuracy. In the data loading phase, the system utilizes CUDA streaming technology for parallel processing. The main program creates two independent CUDA streams in the GPU, one for image data preprocessing and transmission, and the other for voxelization and transmission of point cloud data. This parallel pipeline design allows the CPU's data preparation work and the GPU's computation work to overlap, maximizing hardware resource utilization.

[0062] In the design of the feature extraction branches, the image branch adopts a ResNet-50 backbone network. This network abstracts the input 2D camera image layer by layer through a series of convolutional layers, batch normalization layers, and residual connection structures, extracting a 2D semantic feature map containing texture, color, and edge information. The point cloud branch adopts a VoxelNet voxel network. This network first divides the unstructured LiDAR point cloud into a regular voxel grid, extracts the local geometric features within each voxel through a voxel feature encoding layer (VFE), and then further aggregates them through a 3D convolutional layer to output a 3D geometric feature tensor containing spatial structure information. To achieve deep fusion of heterogeneous features, the system constructs a cross-modal Transformer fusion layer. This layer abandons simple channel concatenation operations and instead adopts a feature interaction method based on an attention mechanism. The system flattens the 2D texture feature map output from the image branch in the spatial dimension and generates a query vector sequence through linear mapping. Simultaneously, the 3D geometric features output from the point cloud branch are projected onto the corresponding image plane and flattened, and key vector sequences are generated through linear mapping. Sum value vector sequence .

[0063] Within the cross-modal Transformer fusion layer, the core computation lies in calculating the multi-head self-attention matrix. The system calculates the query vector... With key vector The dot product measures the consistency between visual and geometric features in the semantic space. To prevent the dot product from becoming too large and causing gradient vanishing, a scaling factor is introduced. The matching scores are then normalized using the Softmax function to obtain the attention weight matrix. Finally, this weight matrix is ​​used to adjust the value vector. Weighted aggregation is performed to generate a fused multimodal feature map. The mathematical expression of the single-head attention mechanism is as follows: ; in, The query matrix is ​​obtained by image feature mapping. The key matrix is ​​obtained from point cloud feature mapping. The value matrix obtained from point cloud feature mapping. Let be the dimension of the key vector. This represents the matrix transpose operation. To capture feature correlations across different subspaces, the system employs a multi-head attention mechanism. , , The image is segmented into multiple sub-vectors, and attention is computed in parallel. Finally, the outputs of each head are concatenated and a linear transformation is performed to obtain the final result. This fused feature map contains both rich semantic category information of the image and retains the precise spatial location information of the point cloud. It is then fed into the detection head network for object classification and bounding box regression.

[0064] The beneficial effects are as follows: First, by using the FP16 half-precision acceleration of the TensorRT engine and CUDA streaming parallel technology, the problem of high inference latency of complex deep learning models on shipboard edge computing devices is solved, ensuring that the perception system can run at a real-time frame rate and meet the fast response requirements under high-speed navigation.

[0065] Second, by utilizing the dual-stream architecture of ResNet-50 and VoxelNet, the advantages of visual sensors in object recognition and the advantages of lidar in ranging and direction finding are fully leveraged, achieving complementary perception capabilities.

[0066] Third, an innovative cross-modal Transformer attention mechanism is introduced, fundamentally solving the problem of heterogeneous data fusion. Through adaptive allocation of attention weights, the model can automatically focus on the point cloud features corresponding to the target region in the image, effectively suppressing the interference of background noise and invalid point clouds. Especially when low light or rain / fog interference at night causes the degradation of a certain modality feature, the attention mechanism can automatically enhance the weight of another modality feature, thereby significantly improving the robustness and accuracy of target detection under complex weather conditions.

[0067] In a preferred embodiment, the method further includes the step of constructing semantic constraints using a high-precision map: Parse high-precision map files in OpenDRIVE format and extract data on channel centerlines, no-navigation zone polygons, and water depth contour lines; The vector data of the high-precision map is rasterized and mapped to the same coordinate system as the locally perceptual raster map; Construct a Bayesian inference network, using the static semantic information provided by the map as prior probabilities; The target category probability output by the deep learning model in step S4 is corrected using prior probability. When the model detection result conflicts with the map semantic constraints, the map semantic constraints are given priority to suppress false detections.

[0068] The following is a detailed description of the implementation method and an analysis of its beneficial effects for the "steps of constructing semantic constraints using high-precision maps". This section aims to fully disclose the technical features in the claims to ensure that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification.

[0069] After obtaining the initial results of environmental perception, this embodiment introduces a high-precision map (HD Map) as a static semantic constraint source to further enhance the logical rationality and security of the perception. First, the system loads and parses a high-precision electronic waterway map file in the standard OpenDRIVE (Open Dynamic Road Information for Vehicle Environment) format. The OpenDRIVE format uses an XML hierarchical structure to describe the road / waterway network. The system uses an XML parser to traverse the document tree, focusing on extracting three core geometric elements: the waterway centerline (ReferenceLine) defining the recommended driving trajectory, the Forbidden Zone Polygon enclosed by a sequence of latitude and longitude coordinates, and the depth contour describing the underwater terrain gradient. The system reads the geometric attributes and topological connections of these vector data and stores them as an object structure in memory.

[0070] Since the perception results output in step S4 are typically based on the ship's local coordinate system (e.g., the forward right-lower coordinate system), while OpenDRIVE map data is typically based on a geodetic coordinate system (e.g., WGS-84), the system needs to rasterize and map the extracted map vector data. The system first converts the geodetic coordinates to local Cartesian coordinates, and then maps the continuous vector graphics to a grid with the same resolution as the local perception raster map in step S6 (e.g., 0.5m × 0.5m). For each grid cell, a corresponding semantic label (e.g., "navigable," "shallow water danger," "absolutely prohibited") is assigned based on whether it is located within a channel, in a restricted area, or whether the water depth meets the ship's draft requirements.

[0071] After completing the spatiotemporal alignment, the system constructs a Bayesian Inference Network to fuse map priors and real-time perception. In this network, the static semantic information provided by the high-precision map is modeled as prior probabilities. For example, if a location is within the "pier" or "shore foundation" area defined by OpenDRIVE, the prior probability of a static obstacle at that location is close to 1; if a location is located in the "center of the main channel," the prior probability of a "building" at that location is close to 0. The target category probability output by the deep learning model in step S4 (e.g., determined as "ship," "navigation mark," or "shoreline") is used as the likelihood probability. .

[0072] The system uses Bayes' theorem to calculate the posterior probability. : ; The calculated posterior probability is used to correct the original detection results. When the model's detection results severely conflict with the map's semantic constraints—for example, the deep learning model detects a "moving ship" in the "inland land" region defined by OpenDRIVE, or a "shore building" in the center of the "deep-water main channel"—the system will determine that the false detection is due to sensor noise or false reflections. In this case, the logic layer prioritizes the map's strong semantic constraints and suppresses these false detections that violate geographical common sense by setting masks or forcing zeros.

[0073] In the preferred embodiment, the method further includes historical data playback and algorithm iterative optimization steps: Enable rosbag data recording service, configure LZ4 compression mode, and write all topic data to the ship's SSD hard drive in real time. When a ship docks and connects to a high-bandwidth network, the rsync incremental synchronization script is automatically triggered to upload the rosbag file to the shore-based training server. Data is decompressed on the shore-based server, and difficult sample examples are labeled using manual annotation tools; Add the labeled new samples to the training set and fine-tune the deep neural network model in step S4. The updated model weight file is pushed to the ship via OTA (Over-The-Air) download technology, replacing the old version of the Plan file.

[0074] The following is a detailed description of the implementation method and an analysis of its beneficial effects for the "historical data playback and algorithm iteration optimization steps". This section aims to fully disclose the technical features in the claims, ensuring that the technical solution meets the requirements of patent law regarding sufficient disclosure and that the claims are supported by the specification.

[0075] During routine ship navigation, the system automatically activates the rosbag data recording service in the background. To preserve massive amounts of sensor data within limited onboard storage space, the system is configured with the LZ4 lossless data compression algorithm. The LZ4 algorithm is renowned for its extremely high compression and decompression speeds, capable of processing millions of data points per second from the LiDAR in real time without consuming excessive CPU resources. The compressed data stream covers all core topics, including LiDAR point clouds, camera images, millimeter-wave radar target lists, and the ship's own attitude, and is sequentially written to the ship's SSD, which boasts high sequential write speeds, forming a complete historical navigation database.

[0076] When a ship docks or enters a 5G network coverage area, and the system detects that the network bandwidth meets the high-throughput transmission conditions, it automatically triggers the rsync incremental synchronization script. The rsync tool uses a rolling checksum algorithm to compare the data differences between the shipborne and shore-based files, transmitting only the changed data blocks instead of retransmitting the entire file. After the data is uploaded to the shore-based training server, the server automatically executes a decompression script to restore the rosbag file. The data engineering team uses a manual annotation tool to finely annotate the selected difficult examples. The so-called difficult examples refer to scenario data with low confidence during online operation, those that have experienced missed detections, or those that seriously conflict with the results of other sensors, such as a semi-submersible vessel that is almost completely submerged, a small fishing boat in dense fog, or an oddly shaped navigation beacon.

[0077] After labeling, the new samples are merged into the original training dataset to form the augmented dataset. The shore-based server loads the deep neural network model weights defined in step S4 and performs fine-tuning training using the augmented dataset. Fine-tuning uses a small learning rate and backpropagation to update network parameters, enabling the model to learn the feature distribution of the new samples. After training is complete and passes validation set testing, the server uses the TensorRT builder to recompile the new PyTorch model into a Plan inference engine file optimized for shipboard GPU hardware. Finally, via OTA (Over-The-Air) download technology, an encrypted and secure transmission channel is established to push the updated Plan file to the shipboard edge computing device, automatically replacing the old weight file and completing the online upgrade of the perception system.

[0078] The beneficial effects are as follows: First, it establishes a data-driven, self-evolving algorithm loop. Through an automated process of "collection, transmission, training, and deployment," the ship's perception system can continuously accumulate experience as its voyage mileage increases, effectively addressing the deficiency of insufficient generalization ability of artificial intelligence models when facing unseen complex scenarios, i.e., long-tail problems. Second, it achieves efficient data transmission and storage management. By adopting a strategy combining LZ4 compression and rsync incremental synchronization, the data transmission traffic between ship and shore is reduced by more than 40% while ensuring data integrity, significantly saving expensive satellite or mobile communication traffic costs. Third, it reduces system operation and maintenance costs and upgrade risks. OTA technology allows technicians to update and iterate core algorithms without boarding the ship. Combined with targeted optimization for difficult examples, it ensures that the ship maintains optimal perception performance and navigation safety throughout its entire lifecycle.

[0079] In the preferred embodiment, the execution logic of the strong tracking unscented Kalman filter algorithm module with adaptive attenuation factor in step S5 is as follows: Establish the nonlinear state equations and measurement equations for ship motion; Sigma point sets are generated using UT transformation, and state prediction and measurement prediction values ​​are calculated. The orthogonality of the residual sequences is calculated in real time, and the new information covariance matrix is ​​constructed. By introducing a strong tracking fading factor, when the orthogonality of the residual sequence deteriorates, the filter is forced to track the abrupt change in the system by increasing the weight of the prediction error covariance matrix. The Kalman gain matrix is ​​modified by using a fading factor to update the state estimate and the posterior error covariance matrix, thereby suppressing filter divergence.

[0080] In step S5, to address the issue of conventional filters easily diverging when ships perform nonlinear maneuvers in complex waters, a strong tracking unscented Kalman filter algorithm module with an adaptive attenuation factor is designed. First, the system establishes the nonlinear state equations and measurement equations for the ship's motion. The state vector is then defined. , respectively representing the ship at time The x-coordinate, y-coordinate, resultant velocity, heading angle, and angular velocity are given. A constant turning rate and velocity model is used to describe the ship's motion evolution, and the state equation is expressed as follows: ,in It is a nonlinear state transition function. This is process noise. The measurement equation is expressed as follows: ,in For sensor observation vectors, It is a nonlinear measurement function. For measuring noise.

[0081] During the filtering iteration process, the algorithm utilizes the Unscented Transform (UT) to handle nonlinear transit problems. The system uses the posterior state estimate from the previous time step... and posterior error covariance matrix Generate according to the proportionally modified symmetric sampling strategy There are Sigma points, among which Let be the dimension of the state vector. Substitute these Sigma points into the nonlinear function. Propagation is performed, and the weighted average of the predicted state values ​​is calculated. and the state prediction covariance matrix Similarly, the Sigma point is used through a nonlinear measurement function. Calculate the predicted value of the measurement .

[0082] To enable the filter to quickly track abrupt changes, this embodiment introduces the core mechanism of the strong tracking principle, namely the orthogonality principle of the residual sequences. The algorithm calculates the innovation, i.e., the residual vector, in real time. Theoretically, when the filter is in its optimal state, the residual sequence should be Gaussian white noise and orthogonal to each other. If the residual sequence is no longer orthogonal, it indicates that the filter's estimation of the system state has deviated. The system estimates the actual covariance matrix of the residuals using the sliding window method. And introduce a strong tracking fading factor This factor adjusts the weights of the prediction error covariance matrix in real time, and the calculation criteria are as follows: ; In the formula, Represents the trace of a matrix. , .in To measure the noise covariance, For process noise covariance, For the Jacobian approximation or statistical linearization matrix of the measurement matrix, This is a weakening factor. When the system detects that the ship is undergoing violent maneuvers that cause the residuals to increase, the calculated... It will be significantly greater than 1.

[0083] The system utilizes the calculated fading factor The prior prediction covariance matrix is ​​forcibly corrected using the following formula: This operation artificially increases the system's uncertainty assessment of the current prediction model, forcing the filter to calculate the Kalman gain matrix more precisely. It relies more heavily on current real-time measurement data. Instead of past predictions, the state estimate is updated using the corrected gain matrix. and the posterior error covariance matrix This completes the filtering loop at the current moment.

[0084] In the preferred scheme, the specific steps of the DS evidence theory based on information entropy correction in step S6 are as follows: Construct a multi-sensor recognition framework and assign a basic probability assignment function (BPA) to each sensor; Calculate the Jousselme distance between any two sources of evidence and construct an evidence distance matrix; The support of each piece of evidence is calculated based on the evidence distance matrix, and the support is normalized to the credibility weight of the evidence. Calculate the information entropy of each evidence source. For evidence sources whose information entropy is greater than a preset threshold, correct their BPA using a credibility weight. The Dempster composition rule is applied to perform orthogonal summation on the modified BPA. When the conflict coefficient k approaches 1, the Yager composition rule is automatically switched to perform fusion, and the final target attribute determination result is output.

[0085] In step S6, to address the issue of conflicting or uncertain identification results of multiple sensors for the same target in complex navigation environments, a DS evidence theory fusion algorithm based on information entropy correction is designed. The system first constructs a multi-sensor identification framework, denoted as the identification framework. The framework includes all possible assumptions, such as the target being a passenger ship, cargo ship, navigational aid, island, or false clutter. For each sensor... The system assigns a basic probability assignment function (BPA) to each classifier based on its internal classifier confidence output, denoted as BPA. This function satisfies two fundamental conditions: the probability of an empty set is zero, i.e. And the sum of the probabilities of all the hypothetical propositions is one, that is... .in Indicates sensor On the proposition The level of trust.

[0086] After acquiring the raw evidence from each sensor, the system needs to evaluate the degree of mutual support between the evidence to identify sensors that may be malfunctioning or interfered with. To this end, the system calculates the correlation between any two evidence sources. and Jousselme distance between The Jousselme distance is a metric that effectively measures the difference between two evidence vectors. Its calculation formula is as follows: ; In the formula, and Assign values ​​to the basic probabilities in column vector form. The Jaccard similarity coefficient matrix describes the intersection relationships between different subsets within the identification framework. Based on the calculated distance matrix, the system further calculates the support for each piece of evidence. Support reflects the degree of consistency between this evidence and all other evidence, and is calculated using the following formula: The system then normalizes the support scores to obtain the credibility weight for each piece of evidence. The calculation formula is: Evidence with higher credibility weights indicates consistency with the judgments of most sensors and should be given higher fusion priority.

[0087] To further eliminate low-quality evidence that, while consistent, possesses extremely high inherent uncertainty, the system introduces the concept of information entropy. The system calculates the information entropy for each source of evidence. Here, a variant of Dunn entropy or Shannon entropy is used to quantify the dispersion of evidence. The calculation formula is as follows: ; Where $|A|$ represents a proposition The number of elements contained therein. When the calculated information entropy... When the value exceeds a preset threshold, it indicates that the sensor's judgment of the target is ambiguous and contains low information. In this case, the system utilizes the confidence weight calculated above. The original BPA is weighted and corrected to generate the corrected basic probability assignment. The corrective strategy is to replace highly conflicting or high-entropy evidence with weighted average evidence, thereby reducing its negative impact on the final result.

[0088] Finally, the system applies the Dempster synthesis rule to perform orthogonal summation on the modified BPA. The Dempster synthesis rule achieves fusion by calculating the joint probability distribution of multiple independent pieces of evidence, and its core calculation formula is as follows: ; In the formula, The conflict coefficient is calculated using the following formula: .coefficient This reflects the degree of conflict between pieces of evidence. When When the value is close to 1, it indicates a strong conflict between the evidence. Directly applying Dempster's rule at this point would lead to counterintuitive conclusions similar to the "Zadeh Paradox." Therefore, this embodiment includes an automatic switching mechanism: when a conflict coefficient is detected... When the probability exceeds the threshold of 0.8, the system automatically switches to the Yager synthesis rule. The Yager rule no longer normalizes the collision probability allocation; instead, it assigns a mass probability of collision occurrence. All allocated to the complete set This is represented as "unknown". Finally, based on the fused BPA value, the system outputs the target attribute determination result according to the maximum membership principle, and combines it with the electronic river map to generate a navigable area raster map with confidence gradient.

[0089] In a preferred embodiment, the method further includes a visualization rendering step based on OpenCV and OpenGL: Create a Qt graphical user interface application main thread, and initialize the OpenCV environment for image processing and the OpenGL context for 3D rendering in the main thread; The cv::Mat constructor of the OpenCV library is called to instantiate a matrix container, read the navigable area raster map data with confidence gradient generated in step S6, map the confidence values ​​in the raster map to single-channel grayscale values ​​from 0 to 255, and write them into the matrix container. Parse the target dynamic tracking list output in step S6, and extract the category ID, 3D center coordinates, and bounding box size parameters for each target in the list; Write vertex and fragment shaders using the OpenGL shader language GLSL; Convert the raster map data stored in the cv::Mat matrix container into a texture map, and convert the extracted target 3D center coordinates and bounding box size parameters into a vertex coordinate array; During the rendering loop, GPU hardware acceleration is used to blend texture maps and vertex arrays and draw them into the frame buffer, which is then output to the ship's bridge display via an HDMI interface.

[0090] After obtaining the final fusion perception results, a high-performance visualization rendering engine based on a hybrid OpenCV and OpenGL programming approach was designed to present the abstract data intuitively to the ship's navigators. The system first creates a Qt graphical user interface application main thread. The Qt framework was chosen as the system's UI container due to its powerful cross-platform capabilities and signal-slot mechanism. During the main thread initialization phase, the system loads the OpenCV dynamic link library environment for matrix operations and image processing, as well as the OpenGL context environment for hardware-accelerated 3D rendering. The OpenGL context acts as a bridge between the OpenGL state machine and the underlying graphics card driver, ensuring that subsequent drawing instructions can be directly executed by the GPU.

[0091] During the data conversion phase, the system calls the core data structure constructor of the OpenCV library to instantiate a matrix container, namely a cv::Mat object. This object is used to receive the navigable area raster map data with confidence gradients generated in step S6. Since each cell in the raster map stores a normalized confidence probability value... Its value ranges from 0 to 1. To adapt to the texture format of the graphics rendering pipeline, the system performs a linear mapping operation, converting the floating-point confidence value into an 8-bit unsigned integer single-channel grayscale value ranging from 0 to 255. The mapping formula is as follows: ; in This indicates the rounding down sign. The converted data is written into a cv::Mat matrix container, forming a grayscale heatmap, where bright areas represent navigable waters with high confidence, and dark areas represent low-confidence or non-navigable areas. Simultaneously, the system parses the target dynamic tracking list output in step S6. This list is a structured array, and the system iterates through the array to extract the attribute information of each tracked target, including a unique category ID to distinguish the target's identity and the three-dimensional center coordinates describing the target's spatial position in the carrier coordinate system. And the bounding box dimensions that describe the geometric contour of the target, namely length, width, and height. .

[0092] During the rendering phase, the system utilizes a custom rendering pipeline program written in the OpenGL Shading Language (GLSL), primarily consisting of vertex shaders and fragment shaders. The vertex shader handles the spatial transformation of 3D vertices, projecting the target's 3D coordinates onto the 2D screen space through multiplication of the model matrix, view matrix, and projection matrix. The calculation formula is as follows: ; in This refers to the vertex vectors in the local coordinate system. The fragment shader is responsible for calculating the pixel colors ultimately displayed on the screen, including texture sampling and lighting calculations. The system utilizes OpenGL's texture mapping technology to bind the raster map data stored in the cv::Mat matrix container into 2D texture objects, which are then drawn on a rectangular plane representing the water surface. For dynamic targets, the system constructs a cube wireframe model consisting of 8 vertices based on the extracted 3D center coordinates and bounding box size, and passes the vertex data to a vertex buffer object (VBO) in video memory. During the rendering loop, the system calls GPU hardware acceleration instructions to enable depth testing and blending modes, blending the underlying nautical map texture with the upper-layer target 3D wireframe and rendering it into the frame buffer. Finally, the rendered image frame signal is transmitted in real-time to the ship's bridge display via the HDMI physical interface.

[0093] The beneficial effects are as follows: First, it achieves simultaneous display of 2D raster data and 3D vector data. Through a hybrid architecture that uses OpenCV to process the underlying heatmap and OpenGL to draw the upper-level 3D bounding boxes, it retains global right-of-way information for environmental perception while highlighting the spatial three-dimensionality of obstacles, greatly reducing the driver's cognitive load on complex situations. Second, it significantly improves visualization efficiency by utilizing GPU hardware acceleration. Offloading the heavy tasks of graphics rasterization and texture sampling from the CPU to the GPU ensures a smooth refresh rate of over 60 frames per second even at high-resolution displays, avoiding visual fatigue and operational delays caused by interface lag. Third, it establishes an intuitive risk expression mechanism through a mathematical mapping from confidence levels to grayscale values. Drivers no longer need to read tedious probability values; they can quickly judge the safety level of an area simply by the depth of the map's color, providing the most direct visual reference for emergency collision avoidance decisions.

[0094] In the preferred embodiment, the method further includes a cloud-edge collaboration step based on the MQTT protocol: Start the MQTT client on the ship's edge computing terminal and connect to the MQTT Broker agent server of the shore-based cloud platform; Define various perception data packets in JSON format, including target ID, latitude and longitude, heading, speed, and confidence level; The LZ4 compression algorithm is used to perform binary compression on the data packets; The frequency of data transmission is dynamically adjusted based on the QoS (Quality of Service) level of the current network bandwidth. Subscribe to the global traffic situation topics published by the shore-based cloud platform, analyze the long-distance ship position data issued by the shore-based platform, and input it as a virtual sensor into the Kalman filter algorithm in step S5.

[0095] To overcome the physical limitations of line-of-sight sensing by shipborne sensors and achieve efficient ship-shore data exchange, this embodiment designs a cloud-edge collaborative communication mechanism based on the MQTT (Message Queuing Telemetry) protocol. During system operation, the ship's edge computing terminal acts as a client, initiating network service processes by calling interfaces of open-source client libraries such as Paho MQTT or Mosquitto. The client initiates a connection request to the MQTT Broker proxy server deployed on the shore-based cloud platform through a TCP / IP network link established by the shipborne satellite communication terminal or 4G / 5G mobile communication module. During the connection process, a Keep Alive heartbeat mechanism is configured, with a heartbeat interval set to 30 to 60 seconds to maintain the stability of the long connection and prevent unexpected link interruption in weak network environments. Once the connection is successfully established, the edge computing terminal initiates the data serialization and publishing process. The system defines a lightweight JSON (JavaScript object representation) format as the data exchange carrier. When constructing the perception data packet, the system traverses the target dynamic tracking list output in step S5 and encapsulates key attributes such as the unique identifier ID, double-precision floating-point latitude and longitude coordinates, heading angle, ground speed, and fusion confidence level of each target into key-value pairs.

[0096] To reduce the cost of expensive satellite communication traffic and improve transmission efficiency, the system introduces the LZ4 compression algorithm before data packets are sent. LZ4 is an extremely fast lossless compression algorithm with significantly faster compression and decompression speeds than the traditional Gzip algorithm, making it particularly suitable for embedded devices with limited computing power. The system converts the generated JSON string into a byte stream and calls the LZ4 compression interface to encode it into binary compressed blocks. Actual testing shows that this step can compress the original text data volume to 30% to 40% of its original size. Subsequently, the system dynamically adjusts the data transmission frequency based on the QoS (Quality of Service) level of the current network environment. The system backend monitors the round-trip time (RTT) and packet loss rate of the network link in real time. When a high-bandwidth near-shore 5G network environment is detected, the system sets the transmission frequency to 5Hz to 10Hz and uses QoS level 1 to ensure that messages are delivered at least once. When a low-bandwidth offshore satellite network environment is detected, the system automatically reduces the transmission frequency to 0.5Hz to 1Hz and switches to QoS level 0 (at most one transmission) to prevent network congestion caused by data backlog.

[0097] While achieving data uplink, this embodiment also implements a data downlink closed loop through a subscription mechanism. The ship's edge computing terminal subscribes to global traffic situation topics published by the shore-based cloud platform. The shore-based cloud platform aggregates macroscopic data from satellite remote sensing, shore-based radar stations, and the VTS (Vessel Traffic Management System), enabling it to perceive information about distant vessels beyond the ship's line of sight. The edge terminal receives and parses these binary data packets from the cloud, extracting the position and motion status information of distant vessels. The system treats this cloud data as a virtual sensor input, formats it into a state vector consistent with the local sensor, and assigns a corresponding measurement noise covariance matrix. Since cloud data typically has a large communication delay, the system calculates the delay time based on the timestamp of the data packet, increases the noise term in its covariance matrix, and then inputs it into the Kalman filtering algorithm described in step S5 for asynchronous updating. This process mathematically extends the ship's perception range.

[0098] The beneficial effects are as follows: First, it establishes a low-bandwidth, low-power ship-shore collaborative communication link. By combining the lightweight MQTT protocol with the LZ4 compression algorithm, it effectively solves the problems of narrow bandwidth and high cost in maritime satellite communication, enabling ships to report high-precision situational awareness in real time at extremely low data costs. Second, it extends perception capabilities beyond line-of-sight. By introducing macroscopic data from the shore-based cloud platform as virtual sensors into the local filtering algorithm, it overcomes the limitations of physical line-of-sight, allowing ships to perceive potential collision risks several kilometers or even tens of kilometers away in advance, reserving sufficient decision-making time for long-distance path planning. Third, it enhances the system's adaptability in fluctuating network environments. The dynamic frequency adjustment strategy ensures that critical data can still be transmitted preferentially when network quality deteriorates, avoiding data interruptions due to buffer overflows and guaranteeing the robustness of the communication link.

[0099] Example 2 Further explanation in conjunction with Example 1, such as Figure 1-5 As shown, a ship navigation perception system based on multi-source information fusion includes: The multi-source heterogeneous sensor group includes a lidar for acquiring point cloud data, a camera for acquiring visual data, an AIS receiver for acquiring target navigation data, and a GPS / IMU integrated navigation system for acquiring the ship's attitude. The edge computing terminal has a built-in NVIDIA GPU accelerator card and FPGA preprocessing chip. The memory of the edge computing terminal stores computer programs. When the computer programs are executed by the processor, they implement the steps of ship navigation perception based on multi-source information fusion. The ship-to-shore communication gateway is equipped with a 5G communication module and a satellite communication module to enable data interaction with the shore-based cloud platform; The human-machine interaction terminal connects to the edge computing terminal via an HDMI interface and is used to display environmental perception results and receive driver commands.

[0100] A computer-readable storage medium storing a computer program that, when executed by a processor, implements a ship navigation perception method based on multi-source information fusion.

[0101] This embodiment proposes a ship navigation perception system based on multi-source information fusion, employing a modular and highly integrated design approach in its hardware architecture. The core of the system comprises a multi-source heterogeneous sensor group, which serves as the system's sensory interface for perceiving the external environment. This includes a lidar system for emitting high-frequency laser beams and receiving echoes to construct a high-precision 3D point cloud model of the ship's surroundings, acquiring geometric shape and distance information of obstacles; a camera system for acquiring high-resolution video stream data, providing rich texture, color, and semantic information to assist the system in recognizing details such as navigation mark colors and ship hull numbers; an AIS receiver system for demodulating dynamic and static data broadcast from surrounding ships, acquiring the target ship's MMSI code, latitude and longitude, heading, and speed, providing reliable prior verification data for multi-source fusion; and a GPS and IMU integrated navigation system for real-time measurement of the ship's absolute latitude and longitude position, three-axis velocity, and roll, pitch, and heading attitude angles, providing a unified carrier coordinate system reference for the spatiotemporal registration of multi-source data.

[0102] The core data processing tasks are handled by the edge computing terminal. This terminal employs a heterogeneous computing architecture, integrating an NVIDIA GPU accelerator card and an FPGA preprocessing chip. The FPGA chip leverages its parallel pipeline processing capabilities to primarily handle hard synchronization triggering, timestamp alignment, and distortion correction preprocessing of the LiDAR point cloud, ensuring the spatiotemporal consistency of the input data. The NVIDIA GPU accelerator card utilizes its powerful floating-point computing capabilities to run parallel computing tasks based on the CUDA architecture, primarily responsible for performing inference of deep neural network models, state estimation using Kalman filtering, and complex matrix operations. The edge computing terminal's memory contains a computer program containing all the instruction code of the method described in any one of claims 1 to 13. When loaded and executed by the processor, this program drives the hardware to complete the entire computational process from data acquisition to environmental model generation.

[0103] To break down information silos and achieve ship-shore collaboration, the system is equipped with a dual-mode redundant ship-shore communication gateway. This gateway integrates both a 5G communication module and a satellite communication module. When the vessel is navigating in near-shore waters or port areas, the gateway prioritizes locking onto the 5G base station signal, utilizing its high bandwidth and low latency to upload massive amounts of sensing data and download high-precision map update packages. When the vessel enters deep-sea areas and leaves the base station coverage, the gateway automatically switches to the satellite communication link, maintaining a heartbeat connection for critical status data via maritime satellite, achieving seamless data interaction with the shore-based cloud platform. In addition, the system is equipped with a human-machine interface terminal, directly connected to an edge computing terminal via an HDMI high-definition multimedia interface. This terminal not only displays real-time augmented reality environmental perception results rendered using OpenGL, but also integrates touch or button input devices to receive commands from the driver for mode switching, alarm confirmation, and system configuration, enabling two-way information flow between humans and the intelligent system.

[0104] This embodiment also relates to a computer-readable storage medium, such as a solid-state drive (SSD), flash memory (Flash), or read-only memory (ROM). The storage medium non-volatilely stores compiled and optimized computer program code. This program code is logically divided into a data acquisition module, a spatiotemporal registration module, a deep learning inference module, a state estimation module, and a visualization rendering module. When this program is read and executed by the processor of the ship's edge computing terminal, it can accurately implement all algorithmic steps, including Lie algebraic time registration, cross-modal Transformer feature fusion, strong-tracking unscented Kalman filtering, and DS evidence theory fusion based on information entropy, transforming the abstract mathematical model into a concrete physical signal processing flow.

[0105] A heterogeneous edge computing architecture combining FPGA and GPU was adopted. The FPGA was used to handle high-frequency signal alignment tasks, while the GPU was used to handle high-density matrix operation tasks. This hardware-software combination design significantly reduced the end-to-end processing latency of the system, ensuring that complex fusion algorithms could complete closed-loop operations within milliseconds, thus meeting the real-time requirements of ships at high speeds.

[0106] An all-weather, all-sea communication support system has been established. The complementary design of 5G and satellite communication allows the system to benefit from nearshore big data while ensuring basic safety monitoring for ocean voyages, effectively supporting the implementation of cloud-edge collaboration and OTA remote upgrade functions.

[0107] It provides an intuitive and user-friendly interactive experience. By outputting visualizations containing confidence gradients through the HDMI interface, it transforms the complex black box of algorithms into a visual language that is easy for drivers to understand. This not only improves the transparency of the system but also provides a reliable decision support interface for assisted driving.

[0108] Example 3 Further explanation in conjunction with Example 1, such as Figure 1-5 As shown, this embodiment provides a specific implementation of a ship navigation perception system and method based on multi-source information fusion. The system first constructs a complete hardware environment at the physical level, deploying various types of sensors around the mast and hull of the ship under test. These include a 128-line mechanical rotating lidar for acquiring high-precision 3D point cloud data, an industrial-grade high-definition camera for capturing environmental texture information, a millimeter-wave radar for detecting distant targets, an AIS receiver for receiving ship dynamic data, and a GPS / IMU integrated navigation system for providing the ship's high-frequency attitude reference. It is also equipped with ultrasonic sensors and anemometers to perceive nearby obstacles and weather conditions. All sensor signal lines are connected to an edge computing terminal located inside the ship's cabin. This terminal integrates an NVIDIA Jetson AGX Orin GPU module for deep learning inference and a Xilinx series FPGA chip for underlying signal synchronization. It is connected to the human-machine interface display on the bridge via an HDMI interface and interconnected with the cloud via a ship-to-shore communication gateway integrating 5G and BeiDou satellite links.

[0109] During the system software startup phase, the edge computing terminal runs on a Linux kernel, first constructing a distributed message communication architecture based on the ROS robot operating system. The system initializes the ROS Master node and loads the hardware drivers for each sensor. LiDAR data is encapsulated in PointCloud2 format, image data in ImageTransport format, and AIS data in a structure format containing MMSI and latitude / longitude. Independent Publisher nodes are instantiated for each sensor. To ensure efficient data transmission, a multi-threaded Spinner listener is configured, utilizing shared memory to achieve zero-copy communication between nodes, and Topics are mapped to Ethernet interfaces for full-duplex transmission. To simplify deployment and maintenance, the software environment employs Docker containerization technology. A Dockerfile is written to pull base images containing CUDA, ROS Noetic, and OpenCV, and the startup order and CPU / memory resource limits of each container are arranged in the docker-compose.yml file. Host Network mode is configured to reuse the host network protocol stack in a bypass manner, ensuring microsecond-level low latency communication. Meanwhile, a separate watchdog process runs in the background of the system, which polls the heartbeat of the nodes at a frequency of 10Hz and counts the frame loss rate. Once it detects that the frame loss rate of a certain sensor exceeds 15% or the node is unresponsive for more than 2 seconds, it automatically reduces its fusion weight or executes the Kill command to restart the corresponding container, thus realizing the self-diagnosis and self-recovery of faults.

[0110] As sensor data continuously flows into the system, the FPGA chip triggers a synchronization signal, and the system uses a time registration algorithm based on Lie algebra to unify the spatiotemporal reference. A continuous SE(3) Lie group trajectory is constructed for the high-frequency pose output by the integrated navigation system. For each point acquired by the lidar, B-spline interpolation is performed on the SE(3) trajectory based on its precise timestamp. The instantaneous pose is calculated in the tangent space using logarithmic and exponential mappings, thereby eliminating point cloud distortion caused by ship swaying. Subsequently, the PCL point cloud library is called for preprocessing. A PassThrough filter is instantiated to filter out points obscured by the ship's hull, and a StatisticalOutlierRemoval filter is instantiated to remove rain and fog noise based on the average distance of K-nearest neighbors. A VoxelGrid filter is used to downsample at leaf sizes of 0.05 to 0.1 meters, significantly reducing data throughput while preserving environmental geometric features, thus preparing for subsequent algorithms.

[0111] In the feature perception stage, the system deploys a TensorRT inference acceleration engine and utilizes CUDA streaming technology to load image and point cloud data in parallel. The image branch extracts 2D texture features using ResNet-50, while the point cloud branch extracts 3D geometric features using VoxelNet. Both are then merged into a cross-modal Transformer fusion layer. In this layer, the 2D features are flattened into query vectors, and the 3D features are projected as key and value vectors. By calculating a multi-head self-attention matrix and performing Softmax normalization, a deep-weighted aggregation of heterogeneous features is achieved, outputting an intermediate perception vector containing the target category and 3D contour. To suppress false detections, the system parses a high-precision OpenDRIVE map, extracts channel centerlines and no-navigation zone data, constructs a Bayesian inference network, and uses the static semantic prior probabilities of the map to correct the output of the deep learning model, such as forcibly suppressing false water targets located in land areas of the map. In addition, the system starts the rosbag recording service in the background, uses LZ4 compression format to store the full data, and synchronizes it to the shore server via rsync when the ship docks and connects to a high-bandwidth network. After manually annotating difficult examples, the model is fine-tuned, and the updated Plan weight file is pushed back to the ship via OTA technology, realizing continuous iteration of the algorithm.

[0112] After obtaining intermediate sensing results, the data enters the state estimation and decision fusion module. The system inputs the sensing vectors and millimeter-wave radar and AIS data into a strong-tracking unscented Kalman filter algorithm module with an adaptive attenuation factor. The algorithm monitors the orthogonality of the residual sequence in real time. When it detects a deterioration in residual orthogonality (meaning the target is performing a sudden maneuver), it increases the weight of the prediction error covariance matrix using a strong-tracking attenuation factor, corrects the Kalman gain, and thus tightly locks onto the nonlinearly moving target. Subsequently, the DS evidence theory based on information entropy correction is introduced to handle conflicts between multiple sensors. The system calculates the Jousselme distance and information entropy of each evidence source, reduces the weight of evidence with high entropy values ​​(uncertainty), and applies the Dempster synthesis rule for fusion. If the conflict coefficient k is detected to be close to 1, it automatically switches to the Yager synthesis rule, finally generating a dynamic target list with confidence and a navigable area grid map.

[0113] Finally, the system transforms the perceived results into user value through visualization rendering and cloud-edge collaboration. On the control panel, a Qt main thread is created to initialize the OpenCV and OpenGL environments. The system calls the cv::Mat constructor to read raster map data, linearly mapping the confidence level values ​​to grayscale values ​​from 0 to 255 and storing them in a matrix. Simultaneously, the target list is parsed to extract 3D coordinates and bounding boxes. Using GLSL shaders, and with GPU hardware acceleration, the grayscale raster map is converted into texture maps, and the 3D wireframe models of the targets are blended and drawn onto the textures, outputting intuitive augmented reality visuals via HDMI. For cloud interaction, the edge device initiates an MQTT client to connect to the shore-based broker, packaging core target data into JSON and compressing it with LZ4. The QoS level and broadcast frequency are dynamically adjusted based on whether the network is 5G or satellite. Simultaneously, it subscribes to global traffic situation topics published by the shore-based network, analyzes the positions of distant ships, and uses these positions as inputs to the local filtering algorithm as virtual sensors, thereby achieving beyond-line-of-sight perception extension and integrated ship-shore collaboration.

[0114] The system first constructs a complete hardware environment at the physical level, deploying a multi-source heterogeneous sensor array at the highest point of the mast and key locations around the hull of the tested vessel. This includes a 128-line mechanical rotating LiDAR providing high-precision 3D point cloud data, industrial-grade high-definition cameras capturing environmental color and texture information, a 77GHz millimeter-wave radar for long-range detection, an AIS receiver receiving dynamic data from the vessel, and a GPS / IMU integrated navigation system providing the vessel's spatiotemporal reference. All sensors are connected to an edge computing terminal with built-in GPUs and FPGAs, and interconnected with the cloud through a dual-mode ship-to-shore communication gateway. During the software startup phase, the edge computing terminal constructs a ROS-based distributed message communication architecture, loads drivers, and standardizes and encapsulates the data from each sensor into independent topics for release. Docker containerization is adopted for deployment, with container order and resources orchestrated using docker-compose, Host Network mode configured to ensure low-latency communication, and a background watchdog process used to automatically restart faulty nodes. Upon data ingestion, the FPGA triggers a synchronization signal. The system utilizes a Lie algebra-based time registration algorithm for spatiotemporal unification, eliminating point cloud distortion caused by ship motion. It also calls the PCL library for preprocessing including pass-through filtering, statistical filtering, and voxel downsampling. In the feature perception stage, a cross-modal Transformer deep neural network is run using the TensorRT acceleration engine to deeply fuse 2D image textures with 3D point cloud geometric features. Semantic constraints from the OpenDRIVE high-precision map are then used to suppress false detections. After obtaining intermediate perception results, a robust state estimation is performed using a strong tracking unscented Kalman filter algorithm module with an adaptive decay factor. Subsequently, a DS evidence theory based on information entropy correction is introduced to handle multi-sensor conflicts, generating a dynamic target list with confidence and a navigable area grid map.

[0115] Finally, the system transforms complex calculation results into an intuitive user interface through visualization rendering, such as... Figure 3 , Figure 4 As shown, Figure 5 As shown. The human-computer interaction terminal on the control panel uses the Qt framework, initializes the OpenCV and OpenGL environment, maps the raster map confidence level to a grayscale texture, and then blends and draws the target 3D wireframe onto it. For example... Figure 3The image shows the augmented reality situational awareness interface output by the system during the straight-line navigation phase. The center of the interface clearly displays the ship's speed and heading information, and different colored sectors visually represent the coverage areas of the lidar, cameras, and millimeter-wave radar. The colored grid overlaid on the satellite map is a navigable area risk heatmap generated based on DS theory; green represents high-confidence safe areas, and red represents high-risk areas near the shore or obstacles. The system integrates and identifies multiple target types and provides clear confidence level labels, such as 95% confidence level for navigation marks and 88% confidence level for approaching ships. Specifically, for an obstacle with indistinct features, the system provides a "low confidence level: 40%" prompt based on evidence theory, reflecting the quantification of uncertainty. Simultaneously, the interface also displays the "static structure" constructed using high-precision map data, as well as the cloud connection status and the overall system fusion confidence trend. Figure 4 As shown, a more complex scenario of intersections at bends is demonstrated. The system stably tracked fishing boats crossing the channel (92% confidence) and key channel bifurcation buoys (98% confidence), proving the effectiveness of the strong tracking filter algorithm in dynamic scenarios. Floating low-confidence obstacles were also mapped in a timely manner. Through these intuitive AR interfaces, the driver can clearly understand the dynamic and static risks of the surrounding environment and the system's operational status, achieving safe and efficient assisted driving. Simultaneously, the edge device initiates an MQTT client, compressing this core perception data and dynamically sending it to the shore-based cloud platform according to network conditions, while also receiving beyond-line-of-sight virtual sensor data from the cloud, achieving ship-shore collaborative perception.

[0116] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A ship navigation perception method and system based on multi-source information fusion, characterized by: The method includes: S1. Collect multi-source heterogeneous data on the navigation environment through various sensors deployed on the ship; S2. Construct a distributed message communication architecture based on the ROS robot operating system, encapsulate the data streams of multiple types of sensors into independent ROS nodes, and aggregate the data to the central computing unit through a publish and subscribe mechanism; S3. Using a time registration algorithm based on Lie algebra and a spatial alignment algorithm based on extrinsic calibration matrix, spatiotemporal benchmarks are unified for multi-source heterogeneous data. S4. Construct a deep neural network model that includes a cross-modal Transformer attention mechanism. Input the registered LiDAR point cloud data and camera image data into the model, extract and fuse environmental features, and generate an intermediate perception vector that includes target category, 3D position and geometric contour. S5. Input the intermediate sensing vector, millimeter-wave radar data, and AIS data into the strong tracking unscented Kalman filter algorithm module with adaptive attenuation factor to perform multi-target state estimation and output a target dynamic tracking list. S6. Introduce the DS evidence theory based on information entropy correction, calculate the confidence level of each target in the target dynamic tracking list, and generate a navigable area raster map with confidence level gradient by combining electronic river map data.

2. The ship navigation perception method based on multi-source information fusion according to claim 1, characterized in that: The various types of sensors in step S1 include millimeter-wave radar, lidar, camera, AIS receiver, integrated navigation system, ultrasonic sensor and anemometer; The specific execution steps for building the distributed message communication architecture based on the ROS robot operating system in step S2 are as follows: Initialize the ROS Master node in a Linux kernel-based edge computing device; Load the hardware drivers for each sensor and instantiate a Publisher node for each sensor; Define a custom message type Msg, where LiDAR data is defined as PointCloud2 format, image data is defined as ImageTransport format, and AIS data is defined as a structure format containing MMSI code, latitude and longitude, and heading angle. Configure a multi-threaded Spinner listener to achieve zero-copy data transmission between different nodes through a shared memory mechanism, and map all sensor data stream topics to Ethernet interfaces for full-duplex transmission.

3. The ship navigation perception method based on multi-source information fusion according to claim 2, characterized in that: This method also includes system deployment steps based on Docker containerization technology: Write a Dockerfile configuration file to pull a base image containing the CUDA parallel computing library, cuDNN neural network acceleration library, ROSNoetic middleware, and OpenCV vision library; Compile and install custom perception algorithm dependency libraries in the image to build an independent perception service container; Write a docker-compose.yml orchestration file to encapsulate the Publisher node and Spinner listener instantiated in step S2 into an independent container service, and define the container startup order and CPU / memory resource limits in the file; Configure the Host Network mode to allow container processes to directly reuse the host machine's network protocol stack and access the physical network interface in a bypass manner, thereby achieving low-latency communication with shipboard sensor hardware.

4. The ship navigation perception method based on multi-source information fusion according to claim 2, characterized in that: The method also includes online fault diagnosis and self-recovery steps for sensors: Start a separate watchdog daemon process to poll the heartbeat signals of each ROS node established in step S2 at a frequency of 10Hz; Statistically analyze the timestamps of data frames in each sensor data stream topic, calculate the time interval between adjacent frames, and deduce the frame drop rate; When the frame loss rate calculated by a certain sensor exceeds 15%, the sensor is determined to be in a sub-healthy state, and the weight coefficient of the sensor's data source in subsequent fusion calculations is automatically reduced. If a sensor node becomes unresponsive for more than 2 seconds, the Kill command of that node's process is executed and the Docker container running that node is restarted, thus achieving self-recovery from the fault.

5. The ship navigation perception method based on multi-source information fusion according to claim 1, characterized in that: The specific steps of the Lie algebra-based time registration algorithm in step S3 are as follows: Read the high-frequency pose data output by the integrated navigation system and construct the continuous-time SE(3) Lie group trajectory; For each sampling point in the lidar point cloud, B-spline interpolation is performed on the SE(3) Lie group trajectory based on the timestamp of its collection time to calculate the instantaneous pose of the ship at that moment; The pose transformation on SE(3) is converted into a tangent vector on the Lie algebra of SE(3) using logarithmic mapping. After linear interpolation in the tangent space, the pose is restored to SE(3) through exponential mapping. The instantaneous pose obtained by interpolation is used to eliminate the point cloud distortion caused by ship motion, and all sensor data are uniformly projected onto the carrier coordinate system of the current frame.

6. The ship navigation perception method based on multi-source information fusion according to claim 5, characterized in that: The method also includes a step of calling the PCL point cloud library for data preprocessing: Obtain the distortion-free point cloud data stream output after spatiotemporal registration in step S3; Instantiate a PassThrough filter object from the PCL library, set the coordinate threshold ranges for the X, Y, and Z axes, and filter out invalid point cloud data caused by obstruction from the ship's own structure; Instantiate a StatisticalOutlierRemoval filter object, and iterate through the point cloud data to calculate the Euclidean distance from each point to its K nearest neighbors. Calculate the mean and standard deviation of the distance distribution, and remove isolated noise points whose distance is greater than the mean plus twice the standard deviation; The VoxelGrid filter is invoked, and the voxel leaf size is set to 0.05 meters to 0.1 meters to downsample the point cloud, thereby reducing the data throughput of the subsequent deep learning model.

7. The ship navigation perception method based on multi-source information fusion according to claim 1, characterized in that: The specific steps in step S4 for constructing a deep neural network model that includes a cross-modal Transformer attention mechanism are as follows: Deploy the TensorRT inference acceleration engine to convert the weights of the PyTorch-trained model into a Plan file in FP16 half-precision format; The image feature extraction branch and the point cloud feature extraction branch are loaded in parallel using CUDA streaming technology; The image branch uses a ResNet-50 backbone network to extract two-dimensional texture features, while the point cloud branch uses a VoxelNet voxel network to extract three-dimensional geometric features. Construct a cross-modal Transformer fusion layer to flatten two-dimensional texture features into query vectors, project three-dimensional geometric features, and flatten them into key and value vectors; Calculate the multi-head self-attention matrix, normalize it using the Softmax function to obtain the attention weights, perform weighted aggregation on heterogeneous features, and output the fused multimodal feature map.

8. The ship navigation perception method based on multi-source information fusion according to claim 7, characterized in that: The method also includes the step of constructing semantic constraints using high-precision maps: Parse high-precision map files in OpenDRIVE format and extract data on channel centerlines, no-navigation zone polygons, and water depth contour lines; The vector data of the high-precision map is rasterized and mapped to the same coordinate system as the locally perceptual raster map; Construct a Bayesian inference network, using the static semantic information provided by the map as prior probabilities; The target category probability output by the deep learning model in step S4 is corrected using prior probability. When the model detection result conflicts with the map semantic constraints, the map semantic constraints are given priority to suppress false detections.

9. The ship navigation perception method based on multi-source information fusion according to claim 7, characterized in that: This method also includes historical data playback and algorithm iterative optimization steps: Enable rosbag data recording service, configure LZ4 compression mode, and write all topic data to the ship's SSD hard drive in real time. When a ship docks and connects to a high-bandwidth network, the rsync incremental synchronization script is automatically triggered to upload the rosbag file to the shore-based training server. Data is decompressed on the shore-based server, and difficult sample examples are labeled using manual annotation tools; Add the labeled new samples to the training set and fine-tune the deep neural network model in step S4. The updated model weight file is pushed to the ship via OTA (Over-The-Air) download technology, replacing the old version of the Plan file.

10. The ship navigation perception method based on multi-source information fusion according to claim 1, characterized in that: The execution logic of the strong tracking unscented Kalman filter algorithm module with adaptive attenuation factor in step S5 is as follows: Establish the nonlinear state equations and measurement equations for ship motion; Sigma point sets are generated using UT transformation, and state prediction and measurement prediction values ​​are calculated. The orthogonality of the residual sequences is calculated in real time, and the new information covariance matrix is ​​constructed. By introducing a strong tracking fading factor, when the orthogonality of the residual sequence deteriorates, the filter is forced to track the abrupt change in the system by increasing the weight of the prediction error covariance matrix. The Kalman gain matrix is ​​modified by using a fading factor to update the state estimate and the posterior error covariance matrix, thereby suppressing filter divergence.

11. The ship navigation perception method based on multi-source information fusion according to claim 1, characterized in that: The specific steps of the DS evidence theory based on information entropy correction in step S6 are as follows: Construct a multi-sensor recognition framework and assign a basic probability assignment function (BPA) to each sensor; Calculate the Jousselme distance between any two sources of evidence and construct an evidence distance matrix; The support of each piece of evidence is calculated based on the evidence distance matrix, and the support is normalized to the credibility weight of the evidence. Calculate the information entropy of each evidence source. For evidence sources whose information entropy is greater than a preset threshold, correct their BPA using a credibility weight. The Dempster composition rule is applied to perform orthogonal summation on the modified BPA. When the conflict coefficient k approaches 1, the Yager composition rule is automatically switched to perform fusion, and the final target attribute determination result is output.

12. The ship navigation perception method based on multi-source information fusion according to claim 11, characterized in that: This method also includes visualization and rendering steps based on OpenCV and OpenGL: Create a Qt graphical user interface application main thread, and initialize the OpenCV environment for image processing and the OpenGL context for 3D rendering in the main thread; The cv::Mat constructor of the OpenCV library is called to instantiate a matrix container. The navigable area raster map data with confidence gradient generated in step S6 is read, the confidence values ​​in the raster map are mapped to single-channel grayscale values ​​from 0 to 255, and written into the matrix container. Parse the target dynamic tracking list output in step S6, and extract the category ID, 3D center coordinates, and bounding box size parameters for each target in the list; Write vertex and fragment shaders using the OpenGL shader language GLSL; Convert the raster map data stored in the cv::Mat matrix container into a texture map, and convert the extracted target 3D center coordinates and bounding box size parameters into a vertex coordinate array; During the rendering loop, GPU hardware acceleration is used to blend texture maps and vertex arrays and draw them into the frame buffer, which is then output to the ship's bridge display via an HDMI interface.

13. The ship navigation perception method based on multi-source information fusion according to claim 1, characterized in that: This method also includes a cloud-edge collaboration step based on the MQTT protocol: Start the MQTT client on the ship's edge computing terminal and connect to the MQTT Broker agent server of the shore-based cloud platform; Define various perception data packets in JSON format, including target ID, latitude and longitude, heading, speed, and confidence level; The LZ4 compression algorithm is used to perform binary compression on the data packets; The frequency of data transmission is dynamically adjusted based on the QoS (Quality of Service) level of the current network bandwidth. Subscribe to the global traffic situation topics published by the shore-based cloud platform, analyze the long-distance ship position data issued by the shore-based platform, and input it as a virtual sensor into the Kalman filter algorithm in step S5.

14. A ship navigation perception system based on multi-source information fusion, characterized in that: the system include: The multi-source heterogeneous sensor group includes a lidar for acquiring point cloud data, a camera for acquiring visual data, an AIS receiver for acquiring target navigation data, and a GPS / IMU integrated navigation system for acquiring the ship's attitude. An edge computing terminal has a built-in NVIDIA GPU accelerator card and FPGA preprocessing chip. The memory of the edge computing terminal stores a computer program. When the computer program is executed by a processor, it implements the steps of any one of claims 1 to 13. The ship-to-shore communication gateway is equipped with a 5G communication module and a satellite communication module to enable data interaction with the shore-based cloud platform; The human-machine interaction terminal connects to the edge computing terminal via an HDMI interface and is used to display environmental perception results and receive driver commands.

15. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the method as claimed in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Intelligent ship multi-source sensing data ship-side fusion method, device and decision system

    CN111507429B

  • A method for fusing multiple information about ship navigation

    CN112857360B

  • A multi-source detection and multi-target information fusion method for intelligent navigation

    CN117109588B

  • A multi-radar image fusion method suitable for ship navigation and shore-based monitoring

    CN117152572B