Marine ship detecting and tracking method and system

Through the improved YOLOv10 model and DIoU loss function combined with SE attention mechanism, the real-time refined detection and identification of ship targets in distant ocean environments is solved, and the ship detection and tracking with high accuracy, robustness and low computing cost are achieved, which is suitable for marine surveillance scenarios.

CN120298651APending Publication Date: 2025-07-11HANGZHOU NORMAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510178747.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing marine ship target tracking technology is difficult to achieve real-time refined detection and identification in far-sea and ocean environments. It is affected by factors such as irregular ship distribution, structural differences, marine environmental interference and complex weather. Traditional IoU cannot provide effective gradient information when the target frame does not overlap.

Method used

The improved YOLOv10 model is adopted to combine the SE attention mechanism and DIoU loss function. By introducing the SE attention mechanism module and DIoU loss function, the detection accuracy and robustness of the model in complex marine environments are improved, and the target matching and tracking is used by the ByteTrack tracker.

Benefits of technology

In complex marine environments, the detection accuracy of small targets and ships of different sizes is significantly improved, the adaptability to occlusion and background interference is enhanced, and the detection stability and tracking accuracy is ensured, and it is suitable for real-time application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298651A_ABST
    Figure CN120298651A_ABST
Patent Text Reader

Abstract

The invention discloses an ocean ship detection tracking method and system, and the method comprises the following steps: obtaining an ocean ship navigation video frame or image, and carrying out the preprocessing; inputting the preprocessed marine ship navigation video frame or image into a pre-trained marine ship detection tracking model to obtain a marine ship detection result; and inputting the sea surface ship detection result into a ByteTrack tracker to obtain an ocean ship tracking result. According to the invention, improvement is carried out based on the YOLOv10 model, the C2f module and the SCConv module of the backbone network are combined, and the SE attention mechanism module is introduced into the backbone network to obtain the marine ship detection and tracking model, so that the reliability and accuracy of the model in practical application are remarkably improved, and tracking and detection of marine ships are effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection and tracking, and particularly relates to a method and system for detecting and tracking marine vessels. Background Art

[0002] The monitoring of marine vessel targets plays an important role in maritime traffic control, shipping safety, fishery monitoring, maritime law enforcement, combating illegal immigration, etc. However, the bottleneck problem faced in this field currently is the real-time and refined detection and identification of targets in the open sea and the high seas. The detection and tracking of marine point cloud perception targets refer to the signal processing technology for detecting and tracking marine targets within a certain range.

[0003] The marine environment is complex and changeable, which has a great impact on the detection accuracy of sensors. Usually, higher target detection and positioning accuracy can be obtained when detecting targets at close range. Using sensor point cloud data to detect targets in complex sea areas can effectively improve the detection accuracy and reduce safety risks.

[0004] However, there are still many influencing factors and difficulties in the existing marine vessel target tracking technology. For example, the distribution of vessels is irregular, the external shapes and structures of vessels are significantly different, the interference of the marine environment is strong, and the weather is complex. In addition, vessel targets have different visual features at different scales and angles, and effective feature extraction and fusion are required. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for detecting and tracking marine vessels in view of the deficiencies of the prior art.

[0006] In a first aspect, the present invention provides a method for detecting and tracking marine vessels, including the following steps:

[0007] Obtain video frames or images of marine vessels in navigation and perform preprocessing;

[0008] Input the preprocessed video frames or images of marine vessels in navigation into a pre-trained marine vessel detection and tracking model to obtain marine vessel detection results; wherein, the marine vessel detection and tracking model is improved based on the YOLOv10 model, the C2f module of the backbone network is combined with the SCConv module, and the SE (Squeeze-and-Excitation) attention mechanism module is introduced into the backbone network;

[0009] Input the sea vessel detection results into the ByteTrack tracker to obtain marine vessel tracking results.

[0010] Furthermore, the loss function of the marine vessel detection and tracking model adopts the DIoU loss function, and the calculation formula is as follows:

[0011]

[0012] Among them, A represents the predicted bounding box, B represents the ground truth bounding box, and IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box; b and b gt are the center points of A and B respectively, ρ(·) represents the Euclidean distance; c is the diagonal length of the smallest enclosing rectangle that contains both A and B.

[0013] DIoU has significant adaptability and advantages in the task of marine vessel target detection and tracking. The marine environment is complex, and vessel targets may be sparsely distributed and have diverse motion trajectories. Traditional IoU cannot provide effective gradient information when the target bounding boxes do not overlap, while DIoU can better optimize the regression of the bounding boxes by introducing a center point distance penalty term, and can provide effective gradient feedback even when the target bounding boxes do not overlap. In addition, DIoU can converge quickly, improving the target localization accuracy, especially suitable for the detection and tracking of vessel targets at long distances or in occlusion situations. Its robustness and efficiency make DIoU perform excellently in the task of marine vessel target detection and tracking, and can effectively cope with the challenges in the complex marine environment.

[0014] Furthermore, the SE attention mechanism module includes a squeeze module and an excitation module, where the squeeze module uses the following calculation formula:

[0015]

[0016] where x c is the feature map of the c-th channel, i and j are the row index and column index of the feature map respectively, H and W are the height and width of the feature map respectively, and z c is the global feature descriptor compressed for each channel;

[0017] The excitation module uses the following calculation formula:

[0018] s = σ(W2δ(W1z))

[0019] where W1 and W2 are the weights of the fully connected layers, δ(·) is the ReLU activation function, σ(·) is the Sigmoid activation function, s is the attention weight for each channel, and z is the global feature vector containing all channel information;

[0020] Finally, the attention weight of each channel is weighted to the features of each channel to obtain the final output of the SE attention mechanism module.

[0021] In the task of marine vessel target detection and tracking, the SE attention mechanism can significantly improve the performance of the model. The marine environment is complex, and vessel targets may be affected by background interference, occlusion, or lighting changes. The SE module helps the model more accurately locate and identify vessel targets by enhancing the feature responses of key channels. At the same time, the SE attention mechanism has a small computational overhead and can improve the accuracy and robustness of detection and tracking without significantly increasing the computational burden, especially suitable for marine monitoring scenarios with high real-time requirements.

[0022] Further, the tracking result of the marine vessel is specifically obtained through the following method:

[0023] Obtain the detection result of the marine vessel, and store the positions, categories, and confidences of all targets in the detection result of the marine vessel in the form of detection boxes (boxes);

[0024] Split the detection boxes into high-confidence detection boxes and low-confidence detection boxes according to the confidence scores (conf);

[0025] Match the tracking result of the previous frame with the high-confidence detection boxes of the current frame;

[0026] For the high-confidence detection boxes that match the tracking result, inherit the tracking ID of the previous frame and update the position; for the high-confidence detection boxes that do not match the tracking result, they may be newly emerged targets, initialize them as new tracking targets, and assign a new tracking ID;

[0027] Match the low-confidence detection boxes with the tracking trajectories that do not match the high-confidence detection boxes, and try to recover the occluded or unstable detected targets to avoid tracking interruption; if the match is successful, restore the tracking ID of the target;

[0028] Combine the matching results of the high-confidence detection boxes and the matching results of the low-confidence detection boxes to generate the tracking result of the current frame.

[0029] In a second aspect, the present invention provides a marine vessel detection and tracking system, including the following modules:

[0030] An image processing module, configured to obtain a marine vessel navigation video frame or image and perform preprocessing;

[0031] A vessel detection module, which processes the preprocessed marine vessel navigation video frame or image by using a pre-trained marine vessel detection and tracking model to detect vessels therein;

[0032] A vessel tracking module, configured to perform vessel tracking on the detection result of the vessel detection module.

[0033] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to execute the method described above.

[0034] In a fourth aspect, the present invention provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described above is implemented.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] The present invention combines the efficient object detection ability of YOLOv10, the multi-scale feature fusion advantage of SCconv, the feature enhancement ability of the SE attention mechanism, and the precise positioning ability of the DIoU loss function, enabling the model to significantly improve the detection accuracy of small targets and vessels of different sizes in complex marine environments, enhancing the adaptability to occlusion and background interference, and ensuring detection stability in complex environments. Among them, the SE attention mechanism can dynamically adjust channel weights, highlight key features, suppress irrelevant background information, and improve the detection effect of the model in marine environments; the DIoU loss function introduces a center point distance penalty on the basis of the traditional IoU, effectively optimizing bounding box regression and improving the accuracy of target positioning; the integrated ByteTrack tracker ensures the information flow and synchronization between detection and tracking, optimizes the target matching algorithm, and effectively deals with the situation of target disappearance or reappearance, thereby achieving smoother and more accurate vessel tracking.

[0037] In summary, the present invention provides a high-precision, high-robustness, and low-computation-cost vessel detection and tracking solution in complex marine environments, which has important application value and is suitable for real-time application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of marine vessel detection according to the present invention.

[0039] Figure 2 It is a flowchart of marine vessel tracking according to the present invention.

[0040] Figure 3 It is a structural diagram of a marine vessel detection and tracking model according to the present invention.

[0041] Figure 4 It is a structural diagram of the ByteTrack tracker according to the present invention.

[0042] Figure 5 It is a schematic diagram of marine vessel detection according to the present invention.

[0043] Figure 6 It is a schematic diagram of marine vessel tracking according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0045] As Figure 1-2 shown, the present invention provides a method for detecting and tracking marine vessels, including a marine vessel detection step and a marine vessel tracking step, wherein the marine vessel detection step is specifically as follows:

[0046] Obtain video frames or images of marine vessel navigation and perform preprocessing.

[0047] Input the preprocessed video frames or images of marine vessel navigation into a pre-trained marine vessel detection and tracking model to obtain a marine vessel detection result; wherein, the marine vessel detection and tracking model is improved based on the YOLOv10 model, combines the C2f module of the backbone network with the SCConv module, and introduces the SE (Squeeze-and-Excitation) attention mechanism module in the backbone network.

[0048] The marine vessel tracking step first executes the above marine vessel detection step, and then inputs the sea surface vessel detection result into the ByteTrack tracker to obtain a marine vessel tracking result.

[0049] During the pre-training process of the marine vessel detection and tracking model, first obtain a labeled video frame training set; in order to improve the generalization ability and adaptability of the model, targeted data augmentation is performed on the marine vessel data in the training set, including a series of data augmentation techniques such as light and shadow changes, image flipping, size scaling, and random weather condition simulation, which can effectively expand the diversity of the data set, expand the original training set to three times the original, ensure that the model can stably identify the target under various environmental conditions, and significantly improve the reliability and accuracy of the model in practical applications. Then, input the augmented training set into the marine vessel detection and tracking model for training; use the test set to detect the trained marine vessel detection and tracking model to obtain the accuracy of the marine vessel detection and tracking model. If the accuracy of the marine vessel detection and tracking model is less than the accuracy threshold, repeat the data augmentation and training until the accuracy of the marine vessel detection and tracking model is greater than the accuracy threshold.

[0050] As Figure 3 shown, the marine vessel detection and tracking model of the present invention adopts a Backbone (backbone network)-Neck (neck)-Head (detection head) architecture, wherein:

[0051] The backbone is responsible for extracting features at different levels from the input image and providing feature maps at different levels for subsequent use, including a first convolutional layer, a second convolutional layer, a first C2fScConv module, a third convolutional layer, a second C2fScConv module, a first SE attention mechanism module, a downsampling layer, a third C2fScConv module, a second SE attention mechanism module, a fourth convolutional layer, a fourth C2fScConv module, a third SE attention mechanism module, an SPPF (SpatialPyramidPoolingFast) module, and a PSA (PyramidSplitAttention) attention mechanism module that are cascaded in sequence.

[0052] The SE attention mechanism module includes a squeeze module and an excitation module, where the squeeze module uses the following calculation formula:

[0053]

[0054] where x c is the feature map of the c-th channel, i and j are the row index and column index of the feature map respectively, H and W are the height and width of the feature map respectively, and z c is the global feature descriptor compressed for each channel;

[0055] The excitation module uses the following calculation formula:

[0056] s = σ(W2δ(W1z))

[0057] where W1 and W2 are the weights of the fully connected layers, δ(·) is the ReLU activation function, σ(·) is the Sigmoid activation function, s is the attention weight for each channel, and z is the global feature vector containing information of all channels;

[0058] Finally, the attention weight of each channel is weighted to the features of each channel to obtain the final output of the SE attention mechanism module.

[0059] The Neck part adopts a multi-scale feature fusion strategy to integrate feature maps of different resolutions, mainly including a first upsampling layer, a first feature concatenation layer, a first C2f module, a second upsampling layer, a second feature concatenation layer, a second C2f module, a convolutional layer, a third feature concatenation layer, a third C2f module, a downsampling layer, a fourth feature concatenation layer, and a feature fusion layer cascaded in sequence; among them, the first feature concatenation layer receives the output of the third C2fScConv module in the Backbone, the second feature concatenation layer receives the output of the first SE attention mechanism module in the Backbone, the third feature concatenation layer receives the output of the first C2f module, and the fourth feature concatenation layer receives the output of the third SE attention mechanism module in the Backbone; the feature fusion layer adopts a C2fCIB module, and while maintaining the overall structure, the CIB (Compact Inverted Block) module replaces the bottleneck module in C2f.

[0060] The Head adopts a multi-branch design and is divided into three detection heads according to different target sizes: small (small targets), medium (medium targets), and large (large targets); each detection head contains a One-to-oneHead (one-to-one head) or a One-to-manyHead (one-to-many head) for processing targets of corresponding sizes.

[0061] The loss function of the marine vessel detection and tracking model adopts the DIoU loss function, and the calculation formula is as follows:

[0062]

[0063] where A represents the predicted bounding box, B represents the ground truth bounding box, IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box; b and b gt are the center points of A and B respectively, ρ(·) represents the Euclidean distance; c is the diagonal length of the smallest enclosing rectangle that contains both A and B.

[0064] As Figure 4 shown, the present invention uses a ByteTrack tracker to track marine vessels, specifically including the following steps:

[0065] Obtain the marine vessel detection results, and store the positions, categories, and confidences of all targets in the marine vessel detection results in the form of detection boxes (boxes);

[0066] Split the detection boxes into high-confidence detection boxes and low-confidence detection boxes according to the confidence score (conf) with a threshold of 0.5;

[0067] Match the tracking results of the previous frame with the high-confidence detection boxes of the current frame;

[0068] For the high-confidence detection boxes that match the tracking results, inherit the tracking ID of the previous frame and update the position; for the high-confidence detection boxes that do not match the tracking results, they may be newly emerged targets, initialize them as new tracking targets, and assign a new tracking ID;

[0069] Match the low-confidence detection boxes with the tracking trajectories that do not match the high-confidence detection boxes, attempt to recover the occluded or unstable detected targets, and avoid tracking interruption; if the match is successful, restore the tracking ID of the target;

[0070] Combine the matching results of the high-confidence detection boxes and the matching results of the low-confidence detection boxes to generate the tracking results of the current frame.

[0071] To further verify the effectiveness of the method of the present invention, comparative tests are respectively carried out on the public dataset seaship using the model of the present invention and the existing models of the same type, and the object detection results are evaluated from two aspects of accuracy and speed, as shown in Table 1 and Table 2 specifically.

[0072] Table 1

[0073]

[0074] Table 2

[0075]

[0076] In terms of accuracy, the accuracy of this ship detection and tracking model reaches 97.6%, which is higher than the detection models of the same type such as faster-rcnn (91.2%) and SSD (80.5%).

[0077] In terms of speed, the speed of this ship detection and tracking model far exceeds the general required rate of 24fps for real-time detection and tracking, and there is a huge improvement compared with the previous similar models.

[0078] In terms of scalability, this model can perform transfer learning on new datasets and perform targeted retraining for different scenarios to expand the application scenarios, and has good extensibility.

[0079] Figure 5-6 They are respectively schematic diagrams of actually applying the method of the present invention to detect and track marine ships.

[0080] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and retouches can still be made, and these improvements and retouches are also regarded as the protection scope of the present invention.

Claims

1. A method for detecting and tracking marine vessels, characterized in that, It includes the following steps: Obtain video frames or images of ocean vessels in navigation and perform preprocessing; Input the preprocessed video frames or images of ocean vessels in navigation into a pre-trained ocean vessel detection and tracking model to obtain ocean vessel detection results; wherein, the ocean vessel detection and tracking model is improved based on the YOLOv10 model, combines the C2f module and the SCConv module of the backbone network, and introduces the SE attention mechanism module into the backbone network; Input the sea vessel detection results into the ByteTrack tracker to obtain ocean vessel tracking results.

2. The marine vessel detection and tracking method according to claim 1, characterized in that, The loss function of the ocean vessel detection and tracking model adopts the DIoU loss function, and the calculation formula is as follows: Among them, A represents the predicted bounding box, B represents the ground truth bounding box, and IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box; b and b gt are the center points of A and B respectively, ρ(·) represents the Euclidean distance; c is the length of the diagonal of the smallest bounding rectangle that contains both A and B.

3. The marine vessel detection and tracking method according to claim 1, wherein The SE attention mechanism module includes a compression module and an excitation module, wherein the compression module adopts the following calculation formula: where x c is the feature map of the c-th channel, i and j are the row index and column index of the feature map respectively, H and W are the height and width of the feature map, and z c is the global feature descriptor after compression for each channel; The excitation module adopts the following calculation formula: s = σ(W2δ(W1z)) where, W1 and W2 are the weights of the fully connected layer, δ(·) is the ReLU activation function, σ(·) is the Sigmoid activation function, s is the attention weight of each channel, and z is the global feature vector containing all channel information; Finally, the attention weight of each channel is weighted to the features of each channel to obtain the final output of the SE attention mechanism module.

4. The marine vessel detection and tracking method according to claim 1, wherein The ocean vessel tracking results are specifically obtained through the following method: Obtain ocean vessel detection results, and store the positions, categories, and confidences of all targets in the ocean vessel detection results in the form of detection boxes; Split the detection boxes into high-confidence detection boxes and low-confidence detection boxes according to the confidence scores; Match the tracking results of the previous frame with the high-confidence detection boxes of the current frame; For the high-confidence detection boxes that do not match the tracking results, initialize them as new tracking targets; Match the low-confidence detection boxes with the tracking results that do not match the high-confidence detection boxes; Combine the matching results of the high-confidence detection boxes and the matching results of the low-confidence detection boxes to generate the tracking results of the current frame.

5. The marine vessel detection and tracking method according to claim 4, characterized in that The matching is as follows: Use Kalman filtering to predict the position of the tracking trajectory of the current frame in the next frame, and the IoU between the predicted box and the actual detection box is used as the similarity during the two matches, and the matching is completed through the Hungarian algorithm.

6. The marine vessel detection and tracking method according to claim 1, characterized in that, The training process of the ocean vessel detection and tracking model is as follows: Collect vessel pictures as a dataset, and divide the vessel pictures into a training set and a test set; Perform data augmentation on the vessel pictures to expand the diversity of the dataset; Input the training set into the ocean vessel detection and tracking model for training; Use the test set to detect the trained ocean vessel detection and tracking model to obtain the accuracy of the ocean vessel detection and tracking model. If the accuracy of the ocean vessel detection and tracking model is less than the accuracy threshold, repeat data augmentation and training until the accuracy of the ocean vessel detection and tracking model is greater than the accuracy threshold.

7. An ocean vessel detection and tracking system, characterized in that, It includes the following modules: An image processing module for obtaining video frames or images of ocean vessels in navigation and performing preprocessing; A vessel detection module that uses a pre-trained ocean vessel detection and tracking model to process the preprocessed video frames or images of ocean vessels in navigation and detect the vessels therein; A vessel tracking module for tracking vessels based on the detection results of a vessel detection module.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1-6.

9. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Water surface small target detection and classification method based on deep learning

    CN113591617A

  • Plant stem growth geometric parameter measurement method based on machine vision

    CN117522950A

  • SNAU-Net-based liver and tumor segmentation method

    CN118196113A

  • Infrared ship target detection tracking method based on deep learning

    CN118298161A

  • Fish target length detection method based on binocular vision

    CN118397074A