Unmanned aerial vehicle maritime search and rescue target detection method based on improved YOLOV8 algorithm

By improving the YOLOV8 algorithm, the deformable large core selection attention mechanism and coordination attention mechanism are integrated, and a small object detection layer is added to the detection head, which solves the problem of low detection accuracy of small sea targets under the drone's field of vision, achieving higher detection accuracy and fewer missed detection and false detection.

CN120182858APending Publication Date: 2025-06-20HENAN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510103323.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has low detection accuracy for small sea targets under the vision of a drone, and is prone to missed and misdetected targets.

Method used

The improved YOLOV8 algorithm is adopted, and the deformable large-core selection attention mechanism DLKS and the coordinate attention mechanism Coordinate Attention are fused in the backbone network, and a small object detection layer is added to the detection head, and the WIoU loss function is used for training to improve the detection accuracy of the model for small objects.

Benefits of technology

It significantly improves the detection accuracy of small targets under the drone's vision, reduces the missed and misdetected targets, and enhances the model's attention to small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182858A_ABST
    Figure CN120182858A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle maritime search and rescue target detection method based on an improved YOLOV8 algorithm, facilitates water search and rescue deployment when maritime affairs such as drowning or shipwreck occur, and belongs to the field of maritime rescue. Comprising the following steps: a training stage: training a maritime target detection network by using a training set, training samples in the training set being collected by an unmanned aerial vehicle, and labels in a data set comprising drowning persons, motorboats, steamships, life buoys and life-saving equipment (life jackets); the improved YOLOv8-based network architecture of the marine target detection network comprises a skeleton, a neck and a detection head. According to the method, the problem of low target detection precision caused by small size of a detection object in an aerial photography scene of the unmanned aerial vehicle can be well solved, the negative influence of seawater ripples and illumination factors on target detection can be well overcome, and the method is good in real-time performance and high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology and water rescue, and particularly to a method for detecting unmanned aerial vehicle (UAV) maritime search and rescue targets based on an improved YOLOV8 algorithm. Background Art

[0002] The ocean covers more than 70% of the Earth's surface, and global economic trade depends on maritime transportation. With the continuous development of technology, more and more engineers are applying robot and camera technologies to the field of maritime monitoring. For the field of water search and rescue, such as sunken ships and maritime disasters, quickly identifying and locating maritime targets can effectively ensure personnel safety and prevent property losses. Engineers in the marine field are trying to use various sensor technologies to collect visual data in the ocean. At the same time, more reliable detection, classification, and video understanding methods are required. Traditional maritime target detection methods usually require the application of radar, but due to its working principle, it is usually difficult to accurately identify maritime targets. With the development of UAV technology, UAVs play an important role in maritime target detection with their advantages of flexibility, portability, and aerial accessibility. Summary of the Invention

[0003] The purpose of the present invention is to propose a method for detecting unmanned aerial vehicle (UAV) maritime search and rescue targets based on an improved YOLOV8 algorithm, which is used to enhance the feature expression ability of the feature extraction network of the backbone network under the UAV's field of view, alleviate the serious problem of small target information loss, strengthen the network's attention to small targets, and improve the mean average precision of small targets.

[0004] The technical solution is as follows:

[0005] Step 1: The maritime patrol UAV takes pictures of maritime target images to collect a picture dataset containing target features, where the target features include contours, textures, and colors;

[0006] Step 2: The picture dataset is labeled and divided into a training set and a test set, and then preprocessed;

[0007] Step 3: Based on the network architecture of the improved YOLOV8, construct a C2F module with a deformable large kernel selection attention mechanism module DLKS and an SPPF module that fuses the position coordination attention mechanism Coordinate Attention as the backbone network of the DLCS-YOLOV8 network model.

[0008] Step 4: Use WIoU as the loss function of the DLCS-YOLOV8, and use the training set to train the network model to obtain an optimized DLCS-YOLOV8 model;

[0009] Step 5: Use the optimized DLCS-YOLOV8 model to test the image test set, obtain the final accuracy of the model, and apply it in practice.

[0010] Furthermore, the structure of the DLCS-YOLOV8 network model described in the step also includes processing the images obtained in step 2 through the backbone network, feature fusion network, and detection head of the model in sequence.

[0011] The backbone network system includes an initial layer convolution Conv module, four feature fusion units, and a C-SPPF module that fuses the coordinated attention mechanism. Each feature unit includes a DLKSA-C2F module and a 3×3 convolution.

[0012] The detection head network is composed of a 3×3 convolution module and a DLKS-C2F module.

[0013] Step 3.1: The size of the initial layer convolution Conv module is 3×3. The Conv module first extracts features from the training images in the image training set to obtain an initial feature map, and the output feature map has 64 channels.

[0014] Step 3.2: After the feature map with 64 channels output by the initial convolution passes through a 3×3 convolution module for downsampling, it is then passed into the DLKSA-C2F module for feature fusion, and this is repeated four times. The final output feature map has 1024 channels.

[0015] Since the size of the detection target is small and its shape can change from the perspective of the drone, the DLKSA attention mechanism module enhances the feature extraction ability of small targets by using deformable large kernel convolution. This is a simplified attention mechanism using a large convolution kernel that can fully understand the volume context. This mechanism operates in a receptive field similar to self-attention while avoiding computational overhead. The DLKSA attention mechanism benefits from deformable convolution and can flexibly change the sampling grid, enabling the model to appropriately adapt to different data patterns.

[0016] The SA module of the DLKSA attention mechanism can dynamically adjust the spatial receptive field of the convolution kernel to better simulate the ranging context of various objects from the perspective of the drone.

[0017] The large kernel convolution of the DLKSA module constructs a large convolution kernel with fewer parameters and calculations by using depth convolution, depth dilated convolution, and 1×1 convolution.

[0018] The equation for the kernel size of the deconvolution and deconvolution dilated convolution used to construct a K×K kernel for an input of dimension H×W and channels C is:

[0019] DW=(2d - 1)×(2d - 1),(1)

[0020]

[0021] The size of the convolution kernel is K, and the dilation rate is d. The number of parameters P(K, d) and floating-point operations (FLOP) F(K, d) is calculated by the following formula:

[0022]

[0023] F(K,d) = P(K,d) × H × W, (4)

[0024] The DLKA module can be formulated as:

[0025] Attention = Conv1×1(DDW - D - Conv(DDW - Conv(F′))), (5)

[0026]

[0027] The kernel selection mechanism spatially selects feature maps from large convolution kernels of different scales, and stitches together features with different receptive field ranges obtained by different kernel functions

[0028]

[0029] Then, by applying channel-based average and max pooling (denoted as Pavg() and Pmax()) to the PIU, the spatial relationship is effectively extracted:

[0030]

[0031] Connect the spatially pooled features and use the convolutional layer F 2→N Convert the merged features (with 2 channels) into N spatial attention maps:

[0032]

[0033] For each spatial attention map, use the sigmoid activation function to obtain a separate spatial selection mask for each decomposed large kernel, and perform spatial weighting, and fuse through the convolutional layer F(k) to obtain the attention feature S:

[0034]

[0035] The final output of the kernel spatial selection module is the element-wise product between the input feature X and S:

[0036] Y = X · S, (12)

[0037] Step 3.3: Input the feature map with 1024 output channels finally output by the DLKSA-C2F module into the C-SPPF module of the fusion coordinated attention mechanism for feature aggregation.

[0038] Step 3.4: Coordinate Attention decomposes the channel attention into two one-dimensional feature encoding processes, aggregating features along two spatial directions respectively. It captures long-range dependencies along one spatial direction while preserving precise position information along the other spatial direction. Then the generated feature maps are encoded into a pair of direction-aware and position-sensitive attention maps respectively.

[0039] Step 3.5: In the small object detection layer of the neck network, the parameter-free attention mechanism SimAM is used to enhance the feature map while reducing the model's computational cost.

[0040] The SimAM attention mechanism is based on neuroscience theory. By minimizing the energy function, the weight of each neuron in the feature map can be obtained to discriminate the linear separability between the current target neuron and other neurons.

[0041] The minimum energy function of the i-th neuron of the SimAM attention mechanism As shown in Equation (1):

[0042]

[0043] In the formula: λ—regular term;

[0044] i—neuron index;

[0045] t i —the i-th neuron on a single channel of the input feature map,

[0046] —the mean value of all neurons on a single channel, used to reduce the computational cost;

[0047] —the variance of all neurons on a single channel, used to reduce the computational cost.

[0048] In the above formula The smaller the value, the more separable the current target neuron and other neurons in the feature map, the more important this neuron is. The weight of each neuron on the feature map can be obtained through The final result of the output feature map X' is shown in Equation (5), where E is the set of all neuron values of the citrus feature map.

[0049]

[0050] The weight normalization is performed by the Sigmoid function, and the obtained neuron weights are multiplied by the features of the original feature map to obtain the final output feature map.

[0051] Step 3.6: The feature maps of different sizes are respectively input into the head network for feature matching at different scales.

[0052] The loss function described in Step 4 uses WIoU as the loss function for calculating the regression of the target box. WIoU is based on a dynamic non-monotonic focusing mechanism. By using the outlier degree as an index for quality evaluation, the focusing on bounding boxes of different qualities is realized. The WIoU is divided into three different levels: V1, V2, and V3. WIoU constructs distance attention according to the distance metric and obtains WIoU v1 with a two-layer attention mechanism. The calculation formula is:

[0053] L WIoUv1 =R WIoU L IoU ,(1)

[0054]

[0055] In the loss function formula, x, x gt , y, y gt are the central point coordinates of the predicted box and the ground truth box, and W and H are the width and height of the minimum bounding box of the ground truth box and the predicted box.

[0056] A monotonic focusing mechanism WIoU v2 for cross-entropy is designed with reference to Focal Loss, effectively reducing the contribution of simple examples to the loss value. This enables the model to focus on difficult examples and obtain an improvement in classification performance. The calculation formula is:

[0057]

[0058]

[0059]

[0060] A non-monotonic focusing coefficient is constructed using β and applied to WIoU v1 to obtain WIoU v3 with dynamic non-monotonic FM. The calculation formula is:

[0061]

[0062]

[0063] The WIoU loss function described in Step 4 uses WIoU V3.

[0064] Step 5 uses the training set and validation set in Step 2 to train, validate, and test the improved YOLOv8 model.

[0065] The present invention has the following beneficial effects:

[0066] Compared with the original YOLOV8 model, the present invention integrates the deformable large kernel selection attention mechanism DLSKA in the backbone network C2F module, introduces the coordinate attention in the feature pyramid network, and adds a detection head for small targets to the head network.

[0067] 1. Aiming at the problem that in the perspective of drones, the detected targets present as small or occluded objects, and there are still a large number of target missed detections and false detections in practical applications. The deformable large kernel selection attention mechanism DLSKA is integrated in the C2F module of YOLOV8. The deformable large kernel convolution can improve the adaptive ability of the model to the object size, and the kernel selection attention mechanism can dynamically adjust its large spatial receptive field and improve the model's understanding ability of the picture context. Thus, the combination of the two improves the detection accuracy of the model for small targets in the drone perspective.

[0068] 2. Aiming at the problem that due to the small object size, the spatial pyramid pooling network SPPF still has insufficient spatial attention to objects. The coordinate attention is introduced into the SPPF network, decomposing the channel attention into two one-dimensional feature encoding processes, aggregating features along two spatial directions respectively, capturing long-range dependencies along one spatial direction, and at the same time retaining accurate position information along the other spatial direction. Finally, the generated feature maps are encoded into a pair of direction-aware and position-sensitive attention maps, which can be complementarily applied to the input feature maps to enhance the representation of the objects of interest. In the neck feature fusion module, the parameter-free attention mechanism SimAM is introduced to reduce the computational amount while enhancing the features. By introducing WIOU, the BBR balance problem between samples with better and worse quality can be solved, and the negative impact on the model caused by the target size problem can be alleviated. Description of the Drawings

[0069] Figure 1 It is a flowchart of a drone maritime rescue target detection method based on the improved YOLOV8 algorithm;

[0070] Figure 2 It is the network structure diagram of the original YOLOV8 model;

[0071] Figure 3 It is the network structure diagram of the improved YOLOV8;

[0072] Figure 4For different sampling points of deformable convolution and schematic diagrams;

[0073] Figure 5 For the network structure diagram of the deformable large kernel selection attention mechanism DLSKA

[0074] Figure 6 For the network structure diagram of Coordinate Attention

[0075] Figure 7 For the SimAM attention mechanism structure diagram. Specific implementation manners

[0076] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0077] In order to make the purpose, technical solutions, innovation points and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.

[0078] See Figure 1 , the present invention provides a maritime rescue target detection method based on improved YOLOv8, and the specific implementation manners are as follows:

[0079] Step 1: Construct a water rescue target dataset, and the specific operations are as follows:

[0080] Define five categories of maritime search and rescue target datasets, namely swimmer, boat, jetski, buoy, and life vest, and select 400 valid images containing target categories from the maritime search and rescue target images captured by the camera.

[0081] Step 2: Divide the valid images into a training set and a validation set, with a division ratio of 4:1, and use the labelme annotation tool to annotate the categories of the training set and the validation set;

[0082] Step 3: Based on the network architecture of improved YOLOV8, construct a C2F module with a deformable large kernel selection attention mechanism module DLKS added and an SPPF module integrating a position coordinate attention mechanism Coordinate Attention as the backbone network of the DLCS-YOLOV8 network model.

[0083] Step 3.1: The size of the initial layer convolution Conv module is 3×3. The Conv module first extracts features from the training images in the image training set to obtain the initial feature map, and the number of output feature map channels is 64.

[0084] Step 3.2: The feature map with 64 channels output by the initial convolution first passes through a 3×3 convolution module for downsampling and then is fed into the DLKSA-C2F module for feature fusion, which is repeated four times. The final output feature map has 1024 channels.

[0085] Step 3.3: The feature map with 1024 channels finally output by the DLKSA-C2F module is fed into the C-SPPF module that integrates the coordinated attention mechanism for feature aggregation.

[0086] Step 3.4: Coordinate Attention decomposes the channel attention into two one-dimensional feature encoding processes, aggregating features along two spatial directions respectively. It captures long-range dependencies along one spatial direction while preserving precise position information along the other spatial direction. Then the generated feature maps are encoded into a pair of direction-aware and position-sensitive attention maps respectively.

[0087] Step 3.5: In the small object detection layer of the neck network, the parameter-free attention mechanism SimAM is used to enhance the feature map.

[0088] Step 3.6: Feature maps of different sizes are fed into the head network respectively for feature matching at different scales.

[0089] Step 4: WIoUV3 is used as the loss function for calculating the target box regression, and the network model is trained with the training set to obtain the optimized DLCS-YOLOV8 model. WIoU constructs distance attention based on distance metrics to obtain WIoU v1 with a two-layer attention mechanism, and the calculation formula is:

[0090] L WIoUv1 =R WIoU L IoU ,(1)

[0091]

[0092] In the loss function formula, x, x gt , y, y gt are the central point coordinates of the predicted box and the ground truth box, and W and H are the width and height of the minimum bounding box of the ground truth box and the predicted box.

[0093] A monotonic focusing mechanism WIoU v2 for cross-entropy is designed with reference to Focal Loss, effectively reducing the contribution of simple examples to the loss value. This enables the model to focus on difficult examples and achieve an improvement in classification performance. The calculation formula is as follows:

[0094]

[0095]

[0096]

[0097] A non-monotonic focusing coefficient is constructed using β and applied to WIoU v1 to obtain WIoU v3 with dynamic non-monotonic FM. The calculation formula is as follows:

[0098]

[0099]

[0100] Step 5: Use the optimized DLCS-YOLOV8 model to test the picture test set obtained in Step 2, obtain the final accuracy of the model, and apply it in practice.

[0101] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0102] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that The following steps are involved: Step 1: Use a maritime patrol drone to take images of maritime targets and collect a data set of images containing target features, including contours, textures, and colors; Step 2: Label the image dataset and divide it into training set and test set in a ratio of 8:2 for preprocessing; Step 3: Based on the improved YOLOV8 network architecture, the C2F module with the deformable large core selection attention mechanism module DLKS and the SPPF module with the position coordination attention mechanism Coordinate Attention are constructed as the backbone network of the DLCS-YOLOV8 network model. A detection layer for small targets is introduced in the shallow layer of the backbone network, and the SimAM attention mechanism is introduced in the neck network for this layer. WIoU is introduced as the loss function in the loss calculation part of the model detection head; Step 4: Use the training set to train the network model and obtain the average accuracy of the model in training. Use the optimized DLCS-YOLOV8 model to test in the test set to obtain the final detection accuracy of the model.

2. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The backbone network of the unmanned aerial vehicle maritime search and rescue target detection model based on the improved YOLOV8 algorithm performs feature extraction processing on the processed image data set to obtain maritime search and rescue image features; The neck network of the unmanned aerial vehicle maritime search and rescue target detection model based on the improved YOLOV8 algorithm performs feature fusion processing on the features extracted by the backbone network; The detection head of the unmanned aerial vehicle maritime search and rescue target detection model based on the improved YOLOV8 algorithm calculates the fused features and finally outputs the detection results of the maritime search and rescue.

3. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The backbone network includes four groups of convolutions and an improved DLKS-C2F module, and is connected to an SPPF module that integrates the coordinate attention mechanism.

4. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The deformable large kernel selection attention mechanism module DLKS first decomposes a super large kernel deformable large kernel convolution into a series of deformable large kernel convolutions to extract input features; different feature maps are divided into 1×1 Convolution filters the features in the feature map, and then concatenates these feature maps in sequence to obtain a fused feature map; channel-based average and maximum pooling are used to effectively extract the spatial relationship in the feature map, and the convolution layer is used to convert the merged features into a spatial attention map; finally, these spatial attention maps are spatially weighted to obtain the final attention map, and the product is made with the original input.

5. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The position coordination attention mechanism Coordinate Attention module obtains fused maritime search and rescue target detection image features by combining the feature information and position information of the maritime search and rescue images of different scales, including: The maritime search and rescue image features are averagely pooled in the horizontal and vertical directions to obtain two feature mapping blocks, and then the two feature mapping blocks are spliced ​​and connected through a 1×1 Convolution reduces dimension to obtain feature information in two directions in the image; The feature information is processed in blocks and input into the activation function Sigmoid function to obtain two final feature maps in the horizontal direction and the vertical direction; The feature maps in the two directions are multiplied with the original feature map to obtain the final output feature map enhanced by position coordinated attention.

6. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The SimAM attention mechanism uses an energy function to give higher weights to neurons with spatial suppression, automatically obtains the importance of neurons, gives more weight to small target information in the feature map, enhances the feature representation of small target information, obtains the enhanced feature map and inputs it into the corresponding detection head, thereby improving detection accuracy.

7. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The WIoU loss function improves the positioning performance of the model for small targets by using a dynamic non-monotonic focusing mechanism and a gradient gain allocation strategy.

8. According to claim 1, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: The specific model training process includes the following steps: Step 1. Select the YOLOV8 model depth size Step 2. Adjust YOLOV8 parameters.

9. According to claim 7, a method for detecting unmanned aerial vehicle maritime search and rescue targets based on an improved YOLOV8 algorithm, characterized in that: YOLOV8n is selected as the baseline model, and the number of training rounds is set to 1000.

10. The method for detecting unmanned aerial vehicle maritime search and rescue targets based on the improved YOLOV8 algorithm according to claim 1, characterized in that: After completing the training of the model, the optimized model is used to detect images with drowning people and determine the target location information.

Citation Information

Cited By

  • Intelligent life buoy unmanned aerial vehicle

    CN120964076A