Ship-shore cooperative tracking and positioning method based on multi-modal sensor fusion

By optimizing ship-shore cooperative positioning through multimodal sensor fusion and unscented Kalman filtering algorithm, the problems of insufficient accuracy and real-time performance in traditional systems are solved, and high-precision, real-time ship positioning and tracking are achieved.

CN120877042APending Publication Date: 2025-10-31WUHAN UNIV OF TECH
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510972633.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional ship-shore collaborative monitoring systems are insufficient in terms of data coupling and spatiotemporal reference alignment, making it difficult to meet the requirements of high-precision operations. They suffer from low ship positioning accuracy and poor real-time performance, and are easily affected by weather in complex environments, resulting in problems such as missed detections, false detections, and positioning drift.

Method used

A multimodal sensor fusion method is adopted, which includes deploying multiple sensors on the ship and shore, performing calibration and data fusion, and combining unscented Kalman filtering algorithm and GPS factor to establish an adaptive motion state model and observation model to optimize the pose of the target ship.

Benefits of technology

It improves the accuracy and real-time performance of ship positioning and tracking, enhances dynamic adaptability, reduces navigation risks, and improves operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877042A_ABST
    Figure CN120877042A_ABST
Patent Text Reader

Abstract

The invention discloses a ship-shore cooperative tracking and positioning method based on multi-modal sensor fusion, and belongs to the technical field of target positioning. The method comprises the steps that multi-sensor layout and sensor fusion calibration are carried out on a target ship set and a shore end respectively, multi-view image sequence data and three-dimensional point cloud data are obtained, and the target ship set comprises a plurality of target ships; performing data fusion based on the multi-view image sequence data and the three-dimensional point cloud data to obtain fusion data of the target ship set, establishing an adaptive motion state model and an adaptive observation model based on the fusion data, and performing state prediction, state updating, data association and tracking management on the target ship set by adopting an unscented Kalman filtering algorithm; and establishing a space-time diagram model based on the Kalman filtering fusion observation factor and the Kalman filtering state prediction factor, and performing pose optimization on the target ship set in combination with the GPS factor. According to the method, the accuracy of ship and ship-shore cooperative positioning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of target positioning technology, and in particular relates to a ship-shore cooperative tracking and positioning method based on multimodal sensor fusion. Background Technology

[0002] With the rapid development of the global shipping industry and the continuous increase in ship traffic density, the demand for safety and efficiency in waterway operations and berthing / unberthing is becoming increasingly urgent. In waterway navigation, ships must cope with complex environments, such as narrow channels, dynamic obstacles, variable hydrological and meteorological conditions, and shoreline obstruction. In port berthing / unberthing scenarios, ships need to achieve high-precision positioning and obstacle avoidance within limited spaces, while being constrained by multiple factors such as wharf structure, mooring facilities, and the movement of nearby vessels. Traditional collaborative monitoring systems suffer from architectural layer separation defects, manifested in insufficient coordination between shore-based monitoring units and shipborne sensing modules in terms of data coupling and spatiotemporal reference alignment. This makes it difficult to meet the requirements of high-precision operations, resulting in inconsistent data spatiotemporal references and weak moving target tracking and positioning capabilities, hindering the support of high-precision real-time collaborative operations. Therefore, how to achieve high-precision ship-shore collaborative tracking has become a crucial problem that urgently needs to be solved.

[0003] In ship-shore scenarios, dynamic and static targets are intertwined. Existing ship-shore perception of waterways, berths, and surrounding vessels primarily relies on traditional radar and automatic identification systems. However, these systems are susceptible to weather conditions, have blind spots, and suffer from insufficient multi-data fusion and interaction, leading to incomplete or ineffective environmental perception. Furthermore, their ability to track and locate multiple moving vessels is weak, often resulting in missed or false detections and tracking drift. Traditional algorithms also struggle to predict the intentions of other vessels. In addition, when global navigation satellite systems deny navigation or sensors partially fail, ships relying on inertial navigation systems are prone to cumulative errors, especially at low speeds, where positioning drift becomes prominent, increasing the risk of collisions. Existing methods for ship-shore cooperative tracking and positioning suffer from low positioning accuracy, poor real-time performance, and high computational complexity. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a ship-shore cooperative tracking and positioning method based on multimodal sensor fusion, which improves the accuracy of ship-shore cooperative positioning.

[0005] In a first aspect, this application provides a ship-shore cooperative tracking and positioning method based on multimodal sensor fusion, the method comprising:

[0006] Multi-sensor deployment and sensor fusion calibration were performed on the target vessel set and the shore end to obtain multi-view image sequence data and three-dimensional point cloud data. The target vessel set includes multiple target vessels.

[0007] Data fusion is performed based on multi-view image sequence data and 3D point cloud data to obtain fused data of the target ship set. The fused data includes the center coordinates, length, width, height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set.

[0008] Based on the fused data, an adaptive motion state model and an adaptive observation model are established. The unscented Kalman filter algorithm is used to perform state prediction, state update, data association, and tracking management of the target ship set.

[0009] A spatiotemporal graph model is established based on the fusion of observation factors and state prediction factors using Kalman filtering, and the pose of the target vessel set is optimized by combining GPS factors.

[0010] According to one embodiment of this application, the step of performing multi-sensor deployment and sensor fusion calibration on the target vessel assembly and the shore end respectively to obtain multi-view image sequence data and three-dimensional point cloud data includes:

[0011] A first visual sensor, a first point cloud sensor, and a combined inertial navigation sensor are deployed on the target vessel assembly, and a second visual sensor and a second point cloud sensor are deployed on the shore end.

[0012] Multiple viewpoint images were acquired in the common viewing area of ​​the first and second vision sensors using a checkerboard calibration board. The extrinsic parameters of adjacent vision sensor modules were calibrated based on the Open uniform linear motion state model, and the intrinsic parameter matrix and relative extrinsic parameter matrix of each vision sensor module in the first and second vision sensors were obtained.

[0013] Using a checkerboard calibration plate with reflective material, the pose transformation is calculated by matching the edge of the point cloud with the image feature points, and the extrinsic parameter matrices of the first and second point cloud sensors are obtained.

[0014] Based on the global trajectory acquired by the combined inertial navigation sensor and the local trajectory of the point cloud acquired by the first point cloud sensor, a trajectory alignment algorithm is used to obtain the rigid transformation relationship between the global trajectory and the local trajectory of the point cloud.

[0015] Based on the intrinsic parameter matrix and relative extrinsic parameter matrix of each visual sensor module in the first and second visual sensors, the extrinsic parameter matrix of the first and second point cloud sensors, and the rigid transformation relationship between the global trajectory and the local trajectory of the point cloud, multi-view image sequence data and three-dimensional point cloud data are obtained.

[0016] The first and second vision sensors each include multiple vision sensor modules.

[0017] According to one embodiment of this application, the fused data of the target ship set obtained by fusing multi-view image sequence data and three-dimensional point cloud data includes:

[0018] YOLOv5-Tiny was used to extract the two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set in the multi-view image sequence data;

[0019] The 3D point cloud data is quantized into Pillar format using Point Pillars and forward inference is performed to obtain Pillar features.

[0020] The two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set are mapped to Pillar features and then concatenated column by column to obtain fused features.

[0021] The fusion features are filtered using the NMS filtering method to obtain the fusion data of the target ship set.

[0022] According to one embodiment of this application, the step of establishing an adaptive motion state model and an adaptive observation model based on the fused data, and using an unscented Kalman filter algorithm to perform state prediction, state update, data association, and tracking management of the target ship set includes:

[0023] Based on the fused data, an adaptive motion state model and an adaptive observation model are established. The adaptive motion state model includes a uniform linear motion state model and a constant turning rate and velocity motion state model.

[0024] Based on the model probability weighting, the uniform linear motion state model and the constant turning rate and velocity motion state model are fused to obtain a fused motion model;

[0025] Based on the fused motion model, the unscented Kalman filter algorithm is used to predict the state of the target ship set;

[0026] The state of the target ship set is updated based on the adaptive observation model.

[0027] The motion and appearance features of the target ship set are obtained, a joint cost matrix is ​​constructed based on the motion and appearance features, and data association is performed using the Hungarian algorithm.

[0028] The design incorporates a target initialization mechanism, a target elimination mechanism, and a target trajectory aggregation mechanism to track and manage the target vessel set.

[0029] According to one embodiment of this application, the step of predicting the state of the target ship set using an unscented Kalman filter algorithm based on the fused motion model includes:

[0030] A Sigma point set is generated based on the initial state mean and covariance matrix of the fusion motion model, and the Sigma point set includes multiple Sigma points;

[0031] Substitute each Sigma point into the uniform linear motion state model and the constant turning rate and velocity motion state model respectively to obtain the prediction results of uniform linear motion state and constant turning rate and velocity motion state.

[0032] Based on the prediction results of uniform linear motion and constant turning rate and velocity motion, the state prediction results of the target ship set are obtained through a weighted fusion formula.

[0033] According to one embodiment of this application, updating the state of the target ship set based on the adaptive observation model includes:

[0034] The state prediction results of the target ship set are updated by unscented transformation to obtain a new Sigma point set;

[0035] The new Sigma point set is input into the adaptive observation model to obtain the observation prediction value corresponding to each Sigma point in the new Sigma point set;

[0036] The predicted observations corresponding to each Sigma point are weighted and summed to obtain the predicted observations of the adaptive observation model.

[0037] The Kalman gain is calculated based on the predicted observations from the adaptive observation model, and the state of the target ship set is updated.

[0038] According to one embodiment of this application, the step of establishing a spatiotemporal graph model based on Kalman filter fusion of observation factors and Kalman filter state prediction factors, and combining GPS factors to optimize the pose of the target vessel set, includes:

[0039] A spatiotemporal graph model is established based on the fusion of observation factors and state prediction factors using Kalman filtering. The nodes of the spatiotemporal graph model represent the positions of the target ship set at different times, and the edges of the spatiotemporal graph model include fusion observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints.

[0040] An objective function is established based on fusion observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints. The LM algorithm is used to solve the objective function, and the pose optimization results of the target ship set are obtained.

[0041] Secondly, this application provides a ship-shore cooperative tracking and positioning device based on multimodal sensor fusion, the device comprising:

[0042] The calibration module is used to perform multi-sensor layout and sensor fusion calibration on the target vessel set and the shore end respectively, to obtain multi-view image sequence data and three-dimensional point cloud data. The target vessel set includes multiple target vessels.

[0043] The first processing module is used to perform data fusion based on multi-view image sequence data and three-dimensional point cloud data to obtain fused data of the target ship set. The fused data includes the center coordinates, length, width, height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set.

[0044] The second processing module is used to establish an adaptive motion state model and an adaptive observation model based on the fused data, and to perform state prediction, state update, data association and tracking management of the target ship set using the unscented Kalman filter algorithm.

[0045] The optimization module is used to establish a spatiotemporal graph model based on the fusion of observation factors and state prediction factors using Kalman filtering, and to optimize the pose of the target ship set by combining GPS factors.

[0046] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in the first aspect above.

[0047] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in the first aspect above.

[0048] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in the first aspect.

[0049] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in the first aspect above.

[0050] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application.

[0051] The present invention provides a ship-shore cooperative tracking and positioning method based on multimodal sensor fusion, which has the following advantages over the prior art:

[0052] (1) This invention obtains multi-view image sequence data and 3D point cloud data by deploying and fusion-calibrating multiple sensors on the target vessel set and shore end, and then fuses the data to obtain fused data, which effectively improves the accuracy of target vessel positioning and tracking. By establishing an adaptive motion state model and an observation model, and using an unscented Kalman filter algorithm for target vessel state prediction, state update, and data association, the dynamic adaptability of target vessel tracking is further enhanced. By combining the observation factors and state prediction factors fused by Kalman filter, and combining GPS factors to optimize the pose of the target vessel, the real-time positioning and tracking capabilities of the vessel are significantly improved. The trajectory and positioning information of the target vessel can be determined more accurately, reducing navigation risks and improving operational efficiency, thus providing a guarantee for automated and efficient vessel navigation.

[0053] (2) This invention effectively improves the state prediction accuracy of the target ship set by establishing an adaptive motion state model and an adaptive observation model based on fused data, and by combining a uniform linear motion state model with a constant turning rate and velocity motion state model. State prediction, updating, data association, and tracking management are performed using an unscented Kalman filter algorithm. A joint cost matrix constructed from motion and appearance features is used, and data association is performed using a Hungarian algorithm, enabling better identification and tracking of target ships. The design of target initialization, elimination, and trajectory aggregation mechanisms makes the tracking and management of target ships more efficient and stable. Attached Figure Description

[0054] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0055] Figure 1 This is one of the flowcharts illustrating the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application embodiment;

[0056] Figure 2 This is a schematic diagram of the layout of the ship sensors provided in the embodiments of this application;

[0057] Figure 3 This is a schematic diagram of the layout of the shore-end sensor provided in an embodiment of this application;

[0058] Figure 4 This is a schematic diagram of the structure for ship sensor calibration provided in an embodiment of this application;

[0059] Figure 5 This is a schematic diagram of the spatiotemporal graph model provided in the embodiments of this application;

[0060] Figure 6 This is the second flowchart illustrating the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application.

[0061] Figure 7 This is a schematic diagram of the structure of the ship-shore cooperative tracking and positioning device based on multimodal sensor fusion provided in the embodiments of this application;

[0062] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0063] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0064] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0065] The following description, in conjunction with the accompanying drawings, details the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion, the ship-shore cooperative tracking and positioning device based on multimodal sensor fusion, the electronic device, and the readable storage medium provided in this application, through specific embodiments and application scenarios.

[0066] Among them, the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion can be applied to the terminal, specifically executed by the hardware or software in the terminal.

[0067] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0068] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0069] The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The following uses an electronic device as the execution subject to illustrate the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application embodiment.

[0070] Figure 1 This is one of the flowcharts illustrating the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application embodiment, such as... Figure 1 As shown, the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion includes steps 110, 120, 130 and 140.

[0071] Step 110: Perform multi-sensor deployment and sensor fusion calibration on the target vessel set and the shore end respectively to obtain multi-view image sequence data and three-dimensional point cloud data. The target vessel set includes multiple target vessels.

[0072] In some embodiments, the step of performing multi-sensor deployment and sensor fusion calibration on the target vessel assembly and the shore end respectively to obtain multi-view image sequence data and three-dimensional point cloud data includes:

[0073] A first visual sensor, a first point cloud sensor, and a combined inertial navigation sensor are deployed on the target vessel assembly, and a second visual sensor and a second point cloud sensor are deployed on the shore end.

[0074] Multiple viewpoint images were acquired in the common viewing area of ​​the first and second vision sensors using a checkerboard calibration board. The extrinsic parameters of adjacent vision sensor modules were calibrated based on the Open uniform linear motion state model, and the intrinsic parameter matrix and relative extrinsic parameter matrix of each vision sensor module in the first and second vision sensors were obtained.

[0075] Using a checkerboard calibration plate with reflective material, the pose transformation is calculated by matching the edge of the point cloud with the image feature points, and the extrinsic parameter matrices of the first and second point cloud sensors are obtained.

[0076] Based on the global trajectory acquired by the combined inertial navigation sensor and the local trajectory of the point cloud acquired by the first point cloud sensor, a trajectory alignment algorithm is used to obtain the rigid transformation relationship between the global trajectory and the local trajectory of the point cloud.

[0077] Based on the intrinsic parameter matrix and relative extrinsic parameter matrix of each visual sensor module in the first and second visual sensors, the extrinsic parameter matrix of the first and second point cloud sensors, and the rigid transformation relationship between the global trajectory and the local trajectory of the point cloud, multi-view image sequence data and three-dimensional point cloud data are obtained.

[0078] The first and second vision sensors each include multiple vision sensor modules.

[0079] Figure 2 This is a schematic diagram of the layout of the ship sensors provided in the embodiments of this application, as shown below. Figure 2 As shown, a first visual sensor, a first point cloud sensor, and a combined inertial navigation sensor are deployed on each ship in the target ship set. The first visual sensor is a panoramic surround-view optical array, which includes multiple visual sensor modules, such as cameras. The first point cloud sensor is a multi-axis rotating three-dimensional environmental perception device, such as lidar. The installation attitude of all sensors is measured mechanically and preliminarily leveled before startup to ensure that the installation error is less than ±5°.

[0080] Figure 3 This is a schematic diagram of the layout of the shore-end sensor provided in an embodiment of this application, as shown below. Figure 3 As shown, a second visual sensor and a second point cloud sensor are arranged on the shore end. The second visual sensor is a binocular visual sensor, such as a camera, and the second point cloud sensor is a lidar. Each sensor is fixed by a rigid bracket.

[0081] It is easy to understand that, in order to achieve spatiotemporal alignment of multi-source sensor data in a unified coordinate system, multi-sensor platforms are built at both the ship and shore ends, and the coordinate transformation matrix between each sensor is calculated through a multi-sensor extrinsic parameter calibration module.

[0082] The first vision sensor comprises multiple vision sensor modules. Due to the generally large size of ships, the distance between non-adjacent vision sensors is large, and their fields of view are difficult to overlap. By acquiring data from multiple locations, a relative transformation relationship is constructed based on the overlapping fields of view between every two adjacent vision sensors. Figure 4 This is a schematic diagram of the structure of the ship sensor calibration provided in the embodiments of this application, as shown below. Figure 4 As shown, two calibration plates are arranged on the shore. Camera 1 and camera 2 correspond to one calibration plate, and camera 2 and camera 3 correspond to one calibration plate. The ship is rotated and moved until the calibration plate appears in the same field of view of camera 4, camera 1 and camera 3. At this time, the position of the two calibration plates is the optimal position, which can achieve a panoramic view effect.

[0083] For example, a checkerboard calibration board is used to acquire multiple viewpoint images in the shared viewing area of ​​adjacent visual sensors. Based on open-source tools such as the Open Uniform Linear Motion State Model, the extrinsic parameters of the adjacent visual sensors are calibrated to obtain the intrinsic parameter matrix and relative extrinsic parameter matrix of each visual sensor module in the first and second visual sensors.

[0084] Using a checkerboard calibration plate with reflective material, the pose transformation is calculated by matching the edge of the point cloud with the image feature points, and the extrinsic parameter matrix from the coordinate system of the first point cloud sensor and the second point cloud sensor to the coordinate system of the main vision sensor module is obtained.

[0085] It is easy to understand that when a ship is sailing at a constant speed along different paths, the global trajectory output by the combined inertial navigation system and the local trajectory of the point cloud output by the first point cloud sensor are collected respectively, and the rigid transformation relationship between the two coordinate systems is solved by using the trajectory alignment method.

[0086] The ship's various sensors establish spatial transformation relationships with the combined inertial navigation coordinate system through extrinsic parameter calibration. The sensing results of each sensor are uniformly represented under the combined inertial navigation reference system for subsequent data fusion and target tracking.

[0087] In some embodiments, geometric features of target vessels in the common viewing area of ​​the water scene are extracted, multi-view data are collected through vessel motion and pose transformation is calculated by matching point cloud and image feature points. The second point cloud sensor is used as the main reference coordinate system of the shore-based observation platform, and the second visual sensor is mapped to this coordinate system through static calibration to obtain the external parameter matrix from the second point cloud sensor coordinate system to the second visual sensor coordinate system.

[0088] In this embodiment, by deploying multiple visual sensors, point cloud sensors, and combined inertial navigation sensors at the target vessel assembly and the shore, and using a checkerboard calibration plate to calibrate the extrinsic parameters of the visual sensors and a checkerboard calibration plate with reflective material to calibrate the extrinsic parameters of the point cloud sensors, the intrinsic and extrinsic parameter matrices are obtained, effectively improving the accuracy and fusion effect of the sensor data. Through multi-dimensional data complementarity of heterogeneous sensor matrices, the fusion process of multi-view image sequence data and 3D point cloud data is optimized, enhancing the continuous target tracking capability under complex sea conditions and strengthening the positioning and tracking capabilities of the target vessel assembly.

[0089] Step 120: Perform data fusion based on multi-view image sequence data and 3D point cloud data to obtain fused data of the target ship set. The fused data includes the center coordinates, length, width, height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set.

[0090] In some embodiments, the fusion of data based on multi-view image sequence data and 3D point cloud data to obtain fused data of the target ship set includes:

[0091] YOLOv5-Tiny was used to extract the two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set in the multi-view image sequence data;

[0092] The 3D point cloud data is quantized into Pillar format using Point Pillars and forward inference is performed to obtain Pillar features.

[0093] The two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set are mapped to Pillar features and then concatenated column by column to obtain fused features.

[0094] The fusion features are filtered using the NMS filtering method to obtain the fusion data of the target ship set.

[0095] It is easy to understand that both the ship and the shore end perform multimodal perception and 3D target detection fusion on the target ship set to provide initial trajectory input for subsequent tracking and prediction. Multimodal feature extraction is achieved through a deep learning target detection framework. Multi-view image sequence data is processed by a convolutional neural network to obtain target semantic features, while 3D point cloud data is rasterized to extract spatial geometric features. Finally, information fusion is achieved through cross-modal feature mapping. The specific process is as follows:

[0096] (1) Acquire multi-view image sequence data and 3D point cloud data;

[0097] (2) YOLOv5-Tiny was used to extract the two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set in the multi-view image sequence data;

[0098] (3) Quantize the 3D point cloud data into Pillar format using Point Pillars and perform forward inference to obtain Pillar features;

[0099] (4) Map the two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set to Pillar features, and then perform column-by-column splicing to obtain fused features;

[0100] (5) Further regress the 3D bounding box of the fusion features and apply NMS to filter to obtain the fusion data of the target ship set, including the center coordinates, length, width and height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set.

[0101] In this embodiment, YOLOv5-Tiny is used to extract the 2D bounding box and semantic segmentation features of each ship in the target ship set, and the 3D point cloud data is quantized into Pillar format. Pillar features are obtained by combining forward inference, effectively improving the accuracy of ship detection and tracking. By fusing the 2D bounding box, semantic segmentation features, and Pillar features column-by-column, and filtering the fused features using the NMS filtering method, fused data of the target ship set is obtained. This enables more accurate identification and positioning of target ships, improving the efficiency of ship tracking.

[0102] Step 130: Based on the fused data, establish an adaptive motion state model and an adaptive observation model, and use the unscented Kalman filter algorithm to perform state prediction, state update, data association and tracking management of the target ship set;

[0103] Furthermore, an adaptive motion state model and an adaptive observation model are established. Based on fused data as input, an unscented Kalman filter algorithm is used to perform state prediction, state update, data association, and tracking management of the target ship set.

[0104] It should be noted that when a vessel in the target vessel set is used as a collaborative sensing vessel, the observations between the collaborative sensing vessel and the shore end are kept in spatiotemporal synchronization, and both the collaborative sensing vessel and the shore end can observe other vessels in the target vessel set.

[0105] Step 140: Establish a spatiotemporal graph model based on the Kalman filter fusion observation factor and the Kalman filter state prediction factor, and optimize the pose of the target ship set by combining GPS factors.

[0106] In some embodiments, the step of establishing a spatiotemporal graph model based on the fusion of Kalman filter observation factors and Kalman filter state prediction factors, and combining GPS factors to optimize the pose of the target vessel set, includes:

[0107] A spatiotemporal graph model is established based on the fusion of observation factors and state prediction factors using Kalman filtering. The nodes of the spatiotemporal graph model represent the positions of the target ship set at different times, and the edges of the spatiotemporal graph model include fusion observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints.

[0108] An objective function is established based on fusion observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints. The LM algorithm is used to solve the objective function, and the pose optimization results of the target ship set are obtained.

[0109] It is easy to understand that while continuously tracking the trajectory of the target vessel set, the positioning results are further optimized by constructing a spatiotemporal correlation constraint network model. The positions of target vessels at different times within the ship-shore collaborative sensing area of ​​the berth or waterway constitute the nodes in the graph model. The edges of the spatiotemporal correlation constraint network model are the observations fused from multi-source sensors between the vessel and the shore, the unscented Kalman filter prediction, the GPS positioning data of each target vessel, the inter-ship distance constraints, and the ship motion model constraints. In the solution process, the LM algorithm is used to solve the objective function to achieve the optimal state estimation of the vessel positions in the sensing area, realizing ship-shore collaborative positioning optimization. The specific process is as follows:

[0110] The multi-ship localization under ship-shore collaboration is also defined as a maximum a posteriori estimation problem. It is assumed that from time 0 to K, the observation values ​​of N target ships are obtained by the collaborative sensing of multiple sensors on the ship and shore. Simultaneously, prior information is obtained from its adaptive nonlinear motion model. By combining these two factors, the conditional probability distribution of the target ship's state is obtained. Furthermore, Bayes' rule is used to transform this into solving for the product of the maximum likelihood estimate and the prior information. The calculation formula is shown below:

[0111] X * =argmaxP(Z|X)P(X|X) - )

[0112] Among them, X * P(Z|X) is the maximum a posteriori probability estimate, P(Z|X) is the maximum likelihood estimate, and P(Z|X) is the prior probability.

[0113] Noise in the observation equation and motion estimation equation is usually modeled in the form of a Gaussian distribution. Taking the negative logarithm of the Gaussian density function of the two transforms the product of the maximum likelihood estimate and the prior into minimizing the negative logarithm, thus turning it into a least squares problem. The optimal state estimate is obtained by minimizing the Mahalanobis distance between various constraints and the true values ​​of all target ships at all times using multiple constraints. Figure 5 This is a schematic diagram of the spatiotemporal graph model provided in the embodiments of this application, such as... Figure 5 As shown, the left side of each node represents the GPS positioning constraint, shore-side sensor fusion observation constraint, and unscented Kalman filter prediction constraint for each target vessel. The dashed line represents the vessel motion model constraint, and the arrow represents the inter-ship distance constraint.

[0114] For GPS positioning constraints, residual terms are constructed by the difference between each ship's actual pose and the position measurement provided by GPS. Where x i,k ,z i,k Let Q be the actual position and GPS measurement of the i-th ship at time k. gThis is the GPS observation covariance matrix.

[0115] For the motion model and prediction constraints, since the ship is analyzed using nonlinear motion, it is impossible to directly correlate the positions of the same ship at previous and subsequent times. Therefore, a secondary nonlinear prediction is required for the current time. Taking times k-2 to k as an example, this requires adaptive fusion of nonlinear motion equations and unscented Kalman filtering from x... i,k-2 predict We still need to start from Secondary prediction The formula for calculating the residual term is as follows:

[0116]

[0117] in, Let Q be the position of the i-th ship at times k-1 and k-2. m Let be the motion covariance matrix.

[0118] For ship-to-ship distances and observation constraints, a multi-sensor fusion observation framework can be constructed. The 3D bounding boxes of the target ships detected by the fusion sensors of both the ship and the shore end can be used to obtain the relative distances between any two target ships, while also incorporating adaptive observation weights. and The residual term is calculated by allocating the two components, as shown in the following formula:

[0119]

[0120] in, and denoted as and , respectively, the relative distances between the i-th and j-th target ships obtained from ship observations and shore observations at time k.

[0121] Since adding all vertices and edges from the initial time step to the current time step to the graph optimization would continuously increase the size of the graph model and increase the optimization time, a sliding window approach is used, and a minimum of 3 frames is set to limit the number of frames required to construct the motion model constraints. All constraints within the sliding window, i.e., the edges of the graph optimization, are grouped together. The calculation formula of the optimization function is as follows:

[0122]

[0123] Where, N x Let be the set of all points to be optimized.

[0124] The objective functions with different constraints are processed separately. The Levenberg-Marquardt Algorithm (LM) method is used to solve the problem, and the global optimum is obtained by iteratively finding local optima.

[0125] In this embodiment, by fusing observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints, and combining them with the LM algorithm to optimize the pose of the target ship set, the positioning accuracy of the target ships can be effectively improved. Optimizing the pose of the ship set through a spatiotemporal graph model significantly improves the positioning and tracking accuracy in dynamic ship environments, and enables real-time adjustment of the target ship's position, reducing the positioning error of the target ship in tracking scenarios and obtaining more robust and accurate positioning results.

[0126] The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application improves the accuracy of target ship positioning and tracking by fusing multi-sensor deployment and calibration of the target ship set and shore end to obtain multi-view image sequence data and 3D point cloud data. This data fusion effectively enhances the dynamic adaptability of target ship tracking. Furthermore, by establishing an adaptive motion state model and observation model, and employing an unscented Kalman filter algorithm for target ship state prediction, state update, and data association, the dynamic adaptability of target ship tracking is further enhanced. By combining the observation factors and state prediction factors fused by Kalman filtering, and incorporating GPS factors to optimize the target ship's pose, the real-time positioning and tracking capabilities of the ship are significantly improved. This allows for more accurate determination of the target ship's trajectory and positioning information, reducing navigation risks, improving operational efficiency, and providing assurance for automated and efficient ship navigation.

[0127] In some embodiments, the step of establishing an adaptive motion state model and an adaptive observation model based on the fused data, and using an unscented Kalman filter algorithm to perform state prediction, state update, data association, and tracking management of the target vessel set includes:

[0128] Based on the fused data, an adaptive motion state model and an adaptive observation model are established. The adaptive motion state model includes a uniform linear motion state model and a constant turning rate and velocity motion state model.

[0129] Based on the model probability weighting, the uniform linear motion state model and the constant turning rate and velocity motion state model are fused to obtain a fused motion model;

[0130] Based on the fused motion model, the unscented Kalman filter algorithm is used to predict the state of the target ship set;

[0131] The state of the target ship set is updated based on the adaptive observation model.

[0132] The motion and appearance features of the target ship set are obtained, a joint cost matrix is ​​constructed based on the motion and appearance features, and data association is performed using the Hungarian algorithm.

[0133] The design incorporates a target initialization mechanism, a target elimination mechanism, and a target trajectory aggregation mechanism to track and manage the target vessel set.

[0134] It is easy to understand that, Figure 6 This is the second flowchart illustrating the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application embodiment, as shown below. Figure 6 As shown, considering that the actual motion of different target ships is not fixed under complex maritime conditions during navigation, it is difficult to describe their operation with a single linear motion model. An adaptive fusion-switching unscented Kalman filter is used to predict the state of multiple target ships under multiple motion models. Within the area detected by the shore and the ship, the model primarily uses the ship's uniform linear motion state model and the constant turning rate and velocity motion state model. Model switching and adaptive motion model fusion are performed based on scene feature discrimination indicators. The specific process is as follows:

[0135] For the uniform linear motion state model, it is mainly conducted under conditions of open waterways and good water conditions. In the unified coordinate system established by the fused data, a two-dimensional plane model of uniform linear motion is established, with its state vector set as [x...]. t y t v xt v yt ] T , , respectively represent the x-axis coordinate, y-axis coordinate, x-axis linear velocity, and y-axis linear velocity of the target ship in a unified coordinate system at time t.

[0136] Based on the ship's motion relationship, we can obtain: x t =x t-1 +v x(t-1) ·Δt, y t =y t-1 +v y(t-1) ·Δt, where x t-1 With y t-1 This represents the x-coordinate and y-coordinate of the target ship at the previous moment.

[0137] The formula for the ship's equation of motion is shown below:

[0138]

[0139] Where, x t-1 With y t-1 This represents the x-coordinate and y-coordinate of the target ship at the previous moment, v x(t-1) With v y(t-1) This represents the target ship's linear velocity in the x-direction and y-direction at the previous moment, and Δt is the time interval between the tracking of the previous and next frames.

[0140] The formula for calculating the noise covariance matrix during motion is shown below:

[0141]

[0142] in, Q represents the acceleration noise variance, reflecting the uncertainty of velocity abrupt changes. cv This is the noise covariance matrix of the uniform linear motion state model.

[0143] For the constant turning rate and speed motion state model, considering situations where the ship needs to adjust its course or the berthing area is complex, a state vector [x] is also established. t y t v t ψ t ω t ] T Where ψ t For the yaw angle, ω t Let be the yaw rate. The formula for calculating the target ship's motion equation at time t is shown below:

[0144]

[0145] Where, ψ t For the yaw angle, ω t x is the yaw angular velocity. t-1 With y t-1 This represents the x-coordinate and y-coordinate of the target ship at the previous moment.

[0146] The formula for calculating the noise covariance matrix of the constant steering ratio and velocity motion state model is as follows:

[0147]

[0148] in, and These represent the distance noise covariance in the horizontal and vertical directions, respectively. For linear velocity noise covariance, Steering angle noise covariance The variance of the steering rate noise is denoted as .

[0149] When switching from a uniform linear motion state model to a constant turning rate and velocity motion state model, the parameter relationships under the state extension are calculated as follows:

[0150]

[0151] Among them, v x and v y These represent the linear velocities in the horizontal and vertical directions, respectively.

[0152] The state variable covariance matrix needs to be reinitialized due to the addition of a new dimension. The calculation formula is shown below:

[0153]

[0154] Among them, P ctrv Let P be the covariance matrix of the motion state with constant steering rate and velocity. add Add ψ and ω dimension matrices.

[0155] When switching from a constant steering rate and velocity motion model to a uniform linear motion model, the ψ and ω dimensions are removed, and simultaneously v x =v·cosψ,v y =v·sinψ.

[0156] Calculate the current average heading angle change rate based on the most recent N frames of historical data. The calculation formula is as follows:

[0157]

[0158] Where, ψ t Let |ψ be the heading angle at time t. t-k -ψ t-k-1 | represents the absolute value of the difference in heading angles at times tk and tk-1.

[0159] Set a threshold range for the rate of change of heading angle while maintaining motion state, and the minimum average rate of change ω. low and the highest average rate of change ω high .

[0160] when At that time, the target ship's motion state model was set as a uniform linear motion state model.

[0161] when At that time, the target ship's motion state model is set as a constant turning rate and speed motion state model.

[0162] when At that time, the target ship maintains the merged motion state model

[0163] A smooth transition of the fused motion state model is achieved through model probability weighting, and the weight calculation is also based on the rate of change of heading angle. From adaptive allocation, as follows:

[0164]

[0165] in, and These are the fusion weights for the motion states of the constant turning rate and velocity motion state model and the uniform linear motion state model, respectively.

[0166] Simultaneously, based on the fusion motion model, the state dimension is unified, therefore the state vector of the fusion model is [x t y t v t ψ t ω t ] T At this point, the uniform linear motion state model includes ω = 0.

[0167] The formula for calculating state fusion is as follows:

[0168]

[0169] Among them, X cv and X ctrv These represent the state values ​​of the uniform linear motion state and the constant turning rate and velocity motion state models, respectively.

[0170] In some embodiments, the step of predicting the state of the target vessel set using an unscented Kalman filter algorithm based on the fused motion model includes:

[0171] A Sigma point set is generated based on the initial state mean and covariance matrix of the fusion motion model, and the Sigma point set includes multiple Sigma points;

[0172] Substitute each Sigma point into the uniform linear motion state model and the constant turning rate and velocity motion state model respectively to obtain the prediction results of uniform linear motion state and constant turning rate and velocity motion state.

[0173] Based on the prediction results of uniform linear motion and constant turning rate and velocity motion, the state prediction results of the target ship set are obtained through a weighted fusion formula.

[0174] It is easy to understand that the motion state prediction of the uniform linear motion state model can be directly made using linear estimation, while the nonlinear motion state model and the fusion model with constant turning rate and velocity are predicted based on unscented Kalman filtering.

[0175] Based on the initial state mean (here assumed to be the posterior state estimate at time t-1). Covariance Generate 2n+1 Sigma sampling points, calculated as follows:

[0176]

[0177] in, The state value of the first sampling point. The state values ​​of the sampling points between the 2nd and 2n+1th sampling points. This is the covariance matrix for the corresponding sampling points. The formula for calculating the mean weights is shown below:

[0178]

[0179] in, The weight value for the first sampling point. Here, λ represents the weight of the sampling points from the 2nd to the 2n+1th sampling point, and λ is the degree to which the sampling points other than the mean deviate from the mean. The larger the value, the further away from the mean, and the smaller the weight. The formula for calculating the covariance weight is as follows:

[0180]

[0181]

[0182] in, The covariance weight value for the first sampling point. α represents the covariance weight value of the sampling points from the 2nd to the 2n+1th, α represents the distribution range of the Sigma points, and β represents the prior knowledge of the state distribution.

[0183] For the constant steering ratio and velocity motion state model, the Sigma sampling points are directly substituted into the motion state equation of the constant steering ratio and velocity motion state model.

[0184] For the fusion model, each Sigma point Substituting the state equations of the uniform linear motion model and the constant turning rate and velocity motion model respectively, we obtain the prediction results of the two models at each point. and Then use the weighted fusion formula The fusion prediction results of all sample points can be obtained. The formula for predicting the overall state mean is as follows:

[0185]

[0186] The formula for calculating covariance prediction is shown below:

[0187]

[0188] Among them, Q t The noise covariance matrix is ​​calculated using the following formula:

[0189]

[0190] In this embodiment, a Sigma point set is generated and substituted into uniform linear motion and constant turning rate and velocity motion state models respectively for state prediction. The state prediction result of the target ship set is obtained by combining a weighted fusion formula, effectively improving the accuracy of ship target tracking. By fusing the prediction results of different motion models, the ship's trajectory can be predicted more accurately, improving the accuracy of target ship state prediction and reducing prediction errors.

[0191] In some embodiments, updating the state of the target vessel set based on the adaptive observation model includes:

[0192] The state prediction results of the target ship set are updated by unscented transformation to obtain a new Sigma point set;

[0193] The new Sigma point set is input into the adaptive observation model to obtain the observation prediction value corresponding to each Sigma point in the new Sigma point set;

[0194] The predicted observations corresponding to each Sigma point are weighted and summed to obtain the predicted observations of the adaptive observation model.

[0195] The Kalman gain is calculated based on the predicted observations from the adaptive observation model, and the state of the target ship set is updated.

[0196] Based on the influence of different maritime scenarios and weather conditions, multi-source sensor weight analysis is used to calculate and form a globally adaptive fused observation, which also serves as the input for the unscented Kalman filter update. The calculation formula for observation fusion is as follows:

[0197]

[0198] Among them, Z ship Z shore They are separate integrated observations from both the ship and the shore. The observation weights for ships and shore ends are respectively calculated using the following formulas:

[0199]

[0200] Among them, Confidence ship Confidence shore The confidence scores of multi-sensor fusion for both the ship and the shore end are calculated using the following formulas:

[0201] Confidence ship = (a1·b1·c1·d1)·Lidar_q ship +[(1-a1)·(1-b1)·(1-a1)·(1-a1)]·Camera_qship

[0202] Confidence shore = (a²·b²·c²·d²)·Lidar_q shore +[(1-a2)·(1-b2)·(1-a2)·(1-a2)]·Camera_q shore

[0203] Among them, Lidar_q ship Camera_q ship ,Lidar_q shore Camera_q shore The values ​​represent the observation quality of the point cloud sensor and the visual sensor at the ship and shore respectively. The pairs a, b, c, and d represent the weights of the point cloud sensor at the same end under different scene conditions. 1-a, 1-b, 1-c, and 1-d are the weights of the visual sensor at the same end.

[0204] Table 1 illustrates the multi-sensor weighting strategy under various scenario conditions.

[0205] Table 1

[0206] Scene Ship point cloud sensor Ship vision sensors shore-based cloud sensor shore-side visual sensors storm <![CDATA[a1]]> <![CDATA[1-a1]]> <![CDATA[a2]]> <![CDATA[1-a2]]> Fog and rain <![CDATA[b1]]> <![CDATA[1-b1]]> <![CDATA[b2]]> <![CDATA[1-b2]]> Light and shadow <![CDATA[c1]]> <![CDATA[1-c1]]> <![CDATA[c2]]> <![CDATA[1-c2 <!-- 14 -->]]> Near and far <![CDATA[d1]]> <![CDATA[1-d1]]> <![CDATA[d2]]> <![CDATA[1-d2]]>

[0207] Under different conditions, such as in foggy or rainy weather, the visibility of both ships and shore-based visual sensors is significantly reduced, while the penetration capability of point cloud sensors is relatively less affected, so the weights of b1 and b2 should be increased. At the same time, in the case of large waves, shore-based sensors are more stable than ships, so the a2 value should be higher than the a1 value. However, in clear weather and near-shore conditions, the weights of 1-c2 and 1-d2 for shore-based visual sensors should be higher.

[0208] The observation quality of ship and shore-based visual sensors and point cloud sensors is calculated as follows:

[0209]

[0210] Where ρ0 is the ideal point cloud density, denoted as the variance of the historical frame rate estimate, Sharp0 is the ideal sharpness (e.g., the gradient energy baseline under clear, fog-free conditions), f(L,DR) is the illumination adaptation function, and Occ is the percentage of occluded pixels.

[0211] To determine the observation quality of a point cloud sensor, the proportion of target point cloud density and target velocity stability are calculated. This is achieved by voxelizing the point cloud within the detection frame and then calculating the proportion of non-empty voxels. The formula for calculating the target point cloud density is shown below:

[0212]

[0213] To determine the target velocity stability, first record the linear velocity estimates {v1,…,v5} for the most recent 5 frames, and then calculate their variance.

[0214] The observation quality of the visual sensor is determined by calculating the illumination conditions and occlusion rate separately. Regarding illumination conditions, the average brightness parameter L and the brightness variation range parameter DR under illumination conditions are first constructed, and the calculation formula is shown below:

[0215]

[0216] Where N is the number of valid pixels in the target detection bounding box, and Pixel i Let Pixel be the grayscale value of the i-th pixel. 95% Pixel 5% These are the 95th and 5th percentiles used when calculating the grayscale histogram of the target ROI, respectively.

[0217] The illumination adaptation function is constructed using a Gaussian attenuation model as follows:

[0218]

[0219] in, These are the standard deviations of brightness tolerance and the standard deviations of variation range tolerance, respectively.

[0220] In calculating the occlusion rate, the proportion of occluded pixels is calculated using the corresponding visual target detection bounding box.

[0221]

[0222] For state updates, the mean of the predicted state in one step. Covariance The unscented transformation is applied again to generate a new Sigma point set, and the weights of each Sigma point are calculated. The newly generated Sigma point set is then substituted into the observation prediction equation. For each sample point, the observation prediction equation is as follows:

[0223]

[0224] in, For each Sigma point, the predicted observations are fused together, H t The observation matrix is ​​determined based on the state vector of the fused motion model and the fused observation vector, with the fused observation vector set as Z. t(fuse) =[x t y t ψ t ] T The formula for calculating the observation matrix is ​​as follows:

[0225]

[0226] The formulas for calculating the mean and covariance of the predicted observations obtained by weighted summation of the predicted observations at each Sigma point are shown below:

[0227]

[0228] in, To fuse the mean of predicted observations, S t To correspond to the prediction covariance matrix, R t To observe the noise covariance matrix

[0229] Meanwhile, the formula for calculating the cross-covariance between the predicted state values ​​and the predicted observations is as follows:

[0230]

[0231] The formulas for calculating Kalman gain and state update are shown below:

[0232]

[0233] in, This represents the updated posterior state estimation result at time t. This corresponds to the updated covariance matrix.

[0234] In this embodiment, the state prediction results of the target vessel set are updated through unscented transformation to obtain a new Sigma point set, which is then input into the adaptive observation model to obtain the observation prediction value corresponding to each Sigma point. The predicted observation values ​​of the adaptive observation model are obtained through weighted summation, and then the Kalman gain is calculated for state updating. This effectively improves the accuracy of the target vessel set state update, achieving more precise vessel state updates and providing more efficient and stable technical support for target vessel tracking and navigation.

[0235] It should be noted that, in order to ensure the alignment of ship and shore-based sensing data on the time axis, each sensor module has a time synchronization mechanism to ensure the comparability of timestamps of multi-frame observation data and support the cross-frame target association and trajectory generation process.

[0236] Regarding motion characteristics, based on the continuous frame point clouds acquired by the ship and shore-based point cloud sensors, the translational velocity v and angular velocity ω of the target ship are calculated by registering adjacent frames using ICP. The cost matrix is ​​calculated using the motion similarity between the ship-sensing target i and the shore-sensing target j in terms of translational velocity and angular velocity, as shown in the following formula:

[0237]

[0238] The weight φ is further adjusted based on sensor confidence analysis. and These are the translational velocities of target i at the ship's end and target j at the shore's end, respectively. and These are the angular velocities of target i at the ship's end and target j at the shore's end, respectively.

[0239] In terms of appearance characteristics, the first step is to calculate the ship target. With shore-side targets The volume ratio of the intersection to the union of 3D bounding boxes between detection boxes reflects the consistency of geometric dimensions, position, and orientation. The formula for calculating the intersection-union ratio is as follows:

[0240]

[0241] in, This represents the 3D detection bounding box for the i-th target vessel detected at the ship's end. This is the 3D detection frame for the j-th target vessel detected at the ship's end.

[0242] The cost matrix is ​​constructed based on the intersection-union ratio, and the calculation formula is as follows:

[0243]

[0244] A joint cost matrix combining motion and appearance features is constructed, and the Hungarian algorithm is used for optimized matching. The cost matrix is ​​calculated as follows:

[0245] C1=γ·D motion1 +(1-γ)·D appearance1

[0246] Where γ is the dynamic weight

[0247] After obtaining the target state prediction result for the next frame (assumed to be t) based on the adaptive motion model, the predicted state is measured by Mahalanobis distance. With the current fused observation Z t(fuse) The statistical distance between them is calculated as follows:

[0248]

[0249] Where i and j are the predicted target observation target numbers, respectively. The result of one-step prediction of the covariance of the target unscented Kalman filter. Let be the predicted state of the i-th target ship at time t. Let t represent the fused observation status of the j-th target ship at time t.

[0250] In terms of appearance cost calculation, the 3DIoU matching between detection boxes is also used to calculate the cost of the target in the previous frame. With the target in the next frame The cost matrix is ​​constructed by the volume ratio of the intersection and union of the 3D bounding boxes between the detection boxes as follows:

[0251]

[0252] in, Let be the 3D detection bounding box of the i-th target ship at time t-1. Let be the 3D detection bounding box of the j-th target ship at time t.

[0253] The formula for calculating the joint cost matrix is ​​as follows:

[0254] C2=η·D motion2 +(1-η)·D appearance2

[0255] Where η is the weighting coefficient

[0256] Based on the joint cost matrix, the observed associated targets in consecutive frames can be obtained through ship-shore cooperation. The weighting coefficients can be adjusted according to environmental complexity. For example, in high-density complex berthing areas, the appearance weight is increased by decreasing the value η to avoid ID switching of similar moving targets. In open waterway areas, the value of η is increased to increase the motion weight to accommodate high-speed ships.

[0257] To achieve continuous tracking and stable management of multiple targets on the water surface, a unified target data management module is introduced, based on the multimodal sensing and tracking performed by both the vessel and the shore. This module organizes, judges, and updates the multi-source tracking results, ensuring the consistency and effectiveness of the sensed information. Specifically, the process includes the following:

[0258] (1) Target initialization mechanism

[0259] When a new target is detected by the shore-based sensing system, a temporary trajectory is assigned to it and observed for verification. By analyzing the spatiotemporal consistency of the target's observations in the ship and shore-based sensors, combined with motion trends and sensing confidence levels, the target is upgraded to a formal tracking object after meeting certain persistence and stability conditions. This process avoids interference from short-term noise or false alarms to the tracking system.

[0260] (2) Target Removal Mechanism

[0261] For targets that are not effectively observed by any sensor within a continuous time period, they are marked as "out of observation." If the predicted location of the target is also outside the coverage area of ​​the current multi-sensor system, it is determined that it has exited the effective observation space, and its corresponding trajectory recording is terminated. This mechanism ensures the stable operation of the tracking system and the efficiency of resource management.

[0262] (3) Target trajectory aggregation mechanism

[0263] In actual sensing processes, differences in perspective, resolution, or target fragmentation may lead to multiple observations corresponding to the same real target. To address this, trajectory aggregation is performed on spatially adjacent targets with consistent motion trends from multi-source data to merge duplicate or fragmented target observations, thereby improving the accuracy and consistency of the overall tracking data.

[0264] In this embodiment, by establishing an adaptive motion state model and an adaptive observation model based on fused data, and combining the fusion of a uniform linear motion state model and a constant turning rate and velocity motion state model, the accuracy of state prediction for the target vessel set is effectively improved. State prediction, updating, data association, and tracking management are performed using an unscented Kalman filter algorithm. A joint cost matrix constructed from motion and appearance features is used, and data association is performed using a Hungarian algorithm, enabling better identification and tracking of target vessels. The design of target initialization, elimination, and trajectory aggregation mechanisms makes the tracking and management of target vessels more efficient and stable.

[0265] The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application can be executed by a ship-shore cooperative tracking and positioning device based on multimodal sensor fusion. This application uses the example of a ship-shore cooperative tracking and positioning device based on multimodal sensor fusion executing the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion to illustrate the ship-shore cooperative tracking and positioning device based on multimodal sensor fusion provided in this application.

[0266] This application also provides a ship-shore cooperative tracking and positioning device based on multimodal sensor fusion, such as... Figure 6 As shown, the ship-shore cooperative tracking and positioning device based on multimodal sensor fusion includes: an acquisition module 610, a first processing module 620, a second processing module 630, and an optimization module 640.

[0267] The calibration module 610 is used to perform multi-sensor layout and sensor fusion calibration on the target vessel set and the shore end respectively, to obtain multi-view image sequence data and three-dimensional point cloud data. The target vessel set includes multiple target vessels.

[0268] The first processing module 620 is used to perform data fusion based on multi-view image sequence data and three-dimensional point cloud data to obtain fused data of the target ship set. The fused data includes the center coordinates, length, width and height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set.

[0269] The second processing module 630 is used to establish an adaptive motion state model and an adaptive observation model based on the fused data, and to perform state prediction, state update, data association and tracking management of the target ship set using the unscented Kalman filter algorithm.

[0270] The optimization module 640 is used to establish a spatiotemporal graph model based on the fusion of observation factors and state prediction factors using Kalman filtering, and to optimize the pose of the target ship set by combining GPS factors.

[0271] The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion provided in this application improves the accuracy of target ship positioning and tracking by fusing multi-sensor deployment and calibration of the target ship set and shore end to obtain multi-view image sequence data and 3D point cloud data. This data fusion effectively enhances the dynamic adaptability of target ship tracking. Furthermore, by establishing an adaptive motion state model and observation model, and employing an unscented Kalman filter algorithm for target ship state prediction, state update, and data association, the dynamic adaptability of target ship tracking is further enhanced. By combining the observation factors and state prediction factors fused by Kalman filtering, and incorporating GPS factors to optimize the target ship's pose, the real-time positioning and tracking capabilities of the ship are significantly improved. This allows for more accurate determination of the target ship's trajectory and positioning information, reducing navigation risks, improving operational efficiency, and providing assurance for automated and efficient ship navigation.

[0272] The ship-shore cooperative tracking and positioning device based on multimodal sensor fusion provided in this application embodiment can achieve... Figures 1 to 5 The various processes implemented in the embodiment of the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion will not be described in detail here to avoid repetition.

[0273] In some embodiments, such as Figure 7 As shown, this application embodiment also provides an electronic device 700, including a processor 701, a memory 702, and a computer program stored in the memory 702 and executable on the processor 701. When the program is executed by the processor 701, it implements the various processes of the above-described ship-shore cooperative tracking and positioning method embodiment based on multimodal sensor fusion and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0274] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0275] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described ship-shore cooperative tracking and positioning method embodiment based on multimodal sensor fusion and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0276] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0277] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described ship-shore cooperative tracking and positioning method based on multimodal sensor fusion.

[0278] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0279] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described embodiments of the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0280] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a device-level chip, device chip, chip device, or on-chip device chip, etc.

[0281] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0282] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion of the various embodiments of this application.

[0283] In the description of this application, "first feature" and "second feature" may include one or more of the features.

[0284] In the description of this application, "multiple" means two or more.

[0285] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0286] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0287] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A ship-shore cooperative tracking and positioning method based on multimodal sensor fusion, characterized in that, The method includes: Multi-sensor deployment and sensor fusion calibration were performed on the target vessel set and the shore end to obtain multi-view image sequence data and three-dimensional point cloud data. The target vessel set includes multiple target vessels. Data fusion is performed based on multi-view image sequence data and 3D point cloud data to obtain fused data of the target ship set. The fused data includes the center coordinates, length, width, height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set. Based on the fused data, an adaptive motion state model and an adaptive observation model are established. The unscented Kalman filter algorithm is used to perform state prediction, state update, data association, and tracking management of the target ship set. A spatiotemporal graph model is established based on the fusion of observation factors and state prediction factors using Kalman filtering, and the pose of the target vessel set is optimized by combining GPS factors.

2. The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion according to claim 1, characterized in that, The process involves performing multi-sensor deployment and sensor fusion calibration on the target vessel assembly and ship-shore areas, respectively, to obtain multi-view image sequence data and 3D point cloud data, including: A first visual sensor, a first point cloud sensor, and a combined inertial navigation sensor are deployed on the target vessel assembly, and a second visual sensor and a second point cloud sensor are deployed on the shore end. Multiple viewpoint images were acquired in the common viewing area of ​​the first and second vision sensors using a checkerboard calibration board. The extrinsic parameters of adjacent vision sensor modules were calibrated based on the Open uniform linear motion state model, and the intrinsic parameter matrix and relative extrinsic parameter matrix of each vision sensor module in the first and second vision sensors were obtained. Using a checkerboard calibration plate with reflective material, the pose transformation is calculated by matching the edge of the point cloud with the image feature points, and the extrinsic parameter matrices of the first and second point cloud sensors are obtained. Based on the global trajectory acquired by the combined inertial navigation sensor and the local trajectory of the point cloud acquired by the first point cloud sensor, a trajectory alignment algorithm is used to obtain the rigid transformation relationship between the global trajectory and the local trajectory of the point cloud. Based on the intrinsic parameter matrix and relative extrinsic parameter matrix of each visual sensor module in the first and second visual sensors, the extrinsic parameter matrix of the first and second point cloud sensors, and the rigid transformation relationship between the global trajectory and the local trajectory of the point cloud, multi-view image sequence data and three-dimensional point cloud data are obtained. The first and second vision sensors each include multiple vision sensor modules.

3. The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion according to claim 1, characterized in that, The data fusion based on multi-view image sequence data and 3D point cloud data to obtain fused data of the target ship set includes: YOLOv5-Tiny was used to extract the two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set in the multi-view image sequence data; The 3D point cloud data is quantized into Pillar format using Point Pillars and forward inference is performed to obtain Pillar features. The two-dimensional bounding box and semantic segmentation features of each target ship in the target ship set are mapped to Pillar features and then concatenated column by column to obtain fused features. The fusion features are filtered using the NMS filtering method to obtain the fusion data of the target ship set.

4. The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion according to claim 1, characterized in that, The process of establishing an adaptive motion state model and an adaptive observation model based on the fused data, and using an unscented Kalman filter algorithm to perform state prediction, state update, data association, and tracking management of the target vessel set, includes: Based on the fused data, an adaptive motion state model and an adaptive observation model are established. The adaptive motion state model includes a uniform linear motion state model and a constant turning rate and velocity motion state model. Based on the model probability weighting, the uniform linear motion state model and the constant turning rate and velocity motion state model are fused to obtain a fused motion model; Based on the fused motion model, the unscented Kalman filter algorithm is used to predict the state of the target ship set; The state of the target ship set is updated based on the adaptive observation model. The motion and appearance features of the target ship set are obtained, a joint cost matrix is ​​constructed based on the motion and appearance features, and data association is performed using the Hungarian algorithm. The design incorporates a target initialization mechanism, a target elimination mechanism, and a target trajectory aggregation mechanism to track and manage the target vessel set.

5. The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion according to claim 4, characterized in that, The process of predicting the state of the target ship set using the unscented Kalman filter algorithm based on the fused motion model includes: A Sigma point set is generated based on the initial state mean and covariance matrix of the fusion motion model, and the Sigma point set includes multiple Sigma points; Substitute each Sigma point into the uniform linear motion state model and the constant turning rate and velocity motion state model respectively to obtain the prediction results of uniform linear motion state and constant turning rate and velocity motion state. Based on the prediction results of uniform linear motion and constant turning rate and velocity motion, the state prediction results of the target ship set are obtained through a weighted fusion formula.

6. The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion according to claim 4, characterized in that, The state update of the target ship set based on the adaptive observation model includes: The state prediction results of the target ship set are updated by unscented transformation to obtain a new Sigma point set; The new Sigma point set is input into the adaptive observation model to obtain the observation prediction value corresponding to each Sigma point in the new Sigma point set; The predicted observations corresponding to each Sigma point are weighted and summed to obtain the predicted observations of the adaptive observation model. The Kalman gain is calculated based on the predicted observations from the adaptive observation model, and the state of the target ship set is updated.

7. The ship-shore cooperative tracking and positioning method based on multimodal sensor fusion according to claim 1, characterized in that, The process of establishing a spatiotemporal graph model based on the fusion of Kalman filter observation factors and Kalman filter state prediction factors, and optimizing the pose of the target vessel set by combining GPS factors, includes: A spatiotemporal graph model is established based on the fusion of observation factors and state prediction factors using Kalman filtering. The nodes of the spatiotemporal graph model represent the positions of the target ship set at different times, and the edges of the spatiotemporal graph model include fusion observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints. An objective function is established based on fusion observation constraints, unscented Kalman filter prediction constraints, GPS positioning constraints, ship-to-ship distance constraints, and motion model constraints. The LM algorithm is used to solve the objective function, and the pose optimization results of the target ship set are obtained.

8. A ship-shore cooperative tracking and positioning device based on multimodal sensor fusion, implemented using the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in any one of claims 1 to 7, characterized in that, The device includes: The calibration module is used to perform multi-sensor layout and sensor fusion calibration on the target vessel set and the shore end respectively, to obtain multi-view image sequence data and three-dimensional point cloud data. The target vessel set includes multiple target vessels. The first processing module is used to perform data fusion based on multi-view image sequence data and three-dimensional point cloud data to obtain fused data of the target ship set. The fused data includes the center coordinates, length, width, height, orientation data, corresponding confidence score and category label of the 3D bounding box of each target ship in the target ship set. The second processing module is used to establish an adaptive motion state model and an adaptive observation model based on the fused data, and to perform state prediction, state update, data association and tracking management of the target ship set using the unscented Kalman filter algorithm. The optimization module is used to establish a spatiotemporal graph model based on the fusion of observation factors and state prediction factors using Kalman filtering, and to optimize the pose of the target ship set by combining GPS factors.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the ship-shore cooperative tracking and positioning method based on multimodal sensor fusion as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Ship flow detection and analysis method based on laser radar

    CN121170722A

  • Map construction method and device, computer equipment and storage medium

    CN121409215A

  • Ship berthing pose estimation method and system based on observation confidence self-calibration

    CN122149500A

  • Ship berthing pose estimation method and system based on observation confidence self-calibration

    CN122149500B

  • Ship multi-mode sensing intelligent monitoring system and method based on data analysis

    CN122176648A