Improved YOLOv5-based unmanned aerial vehicle target detection method and system, equipment and medium

By introducing SwinTransformerV2 and C3 modules in YOLOv5, the model structure and loss function are improved, and the problems of poor scale adaptability, dense target miss detection, low computing efficiency and insufficient positioning accuracy in drone target detection are solved, achieving more efficient multi-scale modeling and stronger position perception capabilities.

CN120220004APending Publication Date: 2025-06-27BLUE SKY LABORATORY +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510435152.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the detection of drone targets, traditional YOLOv5 has problems such as poor scale adaptability, intensive target missed detection, low computing efficiency and insufficient positioning accuracy.

Method used

By introducing the SwinTransformerV2 module and C3 module, the YOLOv5 model is improved, and the depth separation convolution, dynamic window attention module and convolution attention module are adopted to enhance feature extraction and multi-scale modeling capabilities, and the WIoU loss function is used to improve the accuracy of small object detection.

Benefits of technology

It improves the scale adaptability of drone target detection, prevents intensive target miss detection, and improves computing efficiency and positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220004A_ABST
    Figure CN120220004A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual processing, and provides an improved YOLOv5-based unmanned aerial vehicle target detection method, which comprises the steps of S1, acquiring a training image set; s2, performing target labeling on each training image to form a labeling file; s3, training the training image set and the annotation file based on the improved YOLOv5 to obtain an improved YOLOv5 model; the improved YOLOv5 model comprises an input end, a backbone network, a neck network and a detection head which are connected in sequence; the backbone network is used for extracting features; the backbone network comprises a multi-stage depth separable convolution module, a multi-stage C3 module and a dynamic window attention module; each stage of C3 module is located between the two stages of depth separable convolution modules; the neck network is used for carrying out feature fusion; the detection head is used for outputting detection information; and S4, inputting a to-be-detected image set to the improved YOLOv5 model to obtain a detection target. According to the scheme, the scale adaptability can be improved, dense target leak detection is prevented, and the calculation efficiency and the positioning precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vision processing, and in particular to a drone target detection method, system, device, and medium based on improved YOLOv5. Background Art

[0002] In recent years, drone technology has developed rapidly and has been applied in a wider and wider range of fields. Its applications in scenarios such as plant protection and urban surveillance are of great significance. However, the images captured by drones have problems such as drastic changes in target scale, dense distribution, and complex backgrounds. Traditional target detection models (YOLOv5) have the following limitations:

[0003] 1. Poor scale adaptability: The change in the flight altitude of the drone leads to a large difference in target scale, and the existing model has insufficient detection ability for small targets.

[0004] 2. Missed detection of dense targets: High-density targets are difficult to distinguish due to occlusion and motion blur.

[0005] 3. Low computational efficiency: The traditional convolutional structure has a large number of parameters and is difficult to be deployed in real time on the drone side.

[0006] 4. Insufficient positioning accuracy: The traditional IoU loss function has insufficient optimization for difficult samples.

[0007] Therefore, there is an urgent need to provide a drone target detection method, system, device, and medium based on improved YOLOv5, which can improve scale adaptability, prevent missed detection of dense targets, and improve computational efficiency and positioning accuracy.

[0008] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present application. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0009] The main object of the present invention is to overcome the problems of poor scale adaptability, missed detection, low computational efficiency, and insufficient accuracy in target detection, and to provide a drone target detection method, system, device, and medium based on improved YOLOv5, which can improve scale adaptability, prevent missed detection of dense targets, and improve computational efficiency and positioning accuracy.

[0010] To achieve the above object, the first aspect of the present invention provides a drone target detection method based on improved YOLOv5, including the following steps:

[0011] S1: Obtain a training image set;

[0012] S2: Perform target annotation on each training image to form an annotation file;

[0013] S3: Train the training image set and annotation files based on the improved YOLOv5 to obtain an improved YOLOv5 model;

[0014] The improved YOLOv5 model includes an input end, a backbone network, a neck network, and a detection head connected in sequence;

[0015] The backbone network is used to extract features: The backbone network includes multiple levels of depthwise separable convolution modules, multiple levels of C3 modules, and a dynamic window attention module; Each level of C3 module is located between two levels of depthwise separable convolution modules; Each level of C3 module includes 3 or 9 C3 modules;

[0016] The neck network is used for feature fusion;

[0017] The detection head is used to output detection information;

[0018] S4: Input the image set to be detected into the improved YOLOv5 model to obtain detection targets.

[0019] According to an exemplary embodiment of the present invention, in step S3, each level of depthwise separable convolution module realizes downsampling through strided convolution. The depthwise separable convolution module is used for depth convolution and pointwise convolution, and formula 1 is adopted:

[0020] Conv_DS(X) = PointwiseConv(DepthwiseConv(X)) Formula 1;

[0021] Where Conv_DS(X) represents depthwise separable convolution, Depthwise Conv represents depth convolution, and Pointwise Conv represents pointwise convolution.

[0022] According to an exemplary embodiment of the present invention, each C3 module includes convolution modules of various different sizes. The neck network cross-scale fuses the deep semantic features and shallow detail features output by the C3 module through depthwise separable convolution, upsampling, and channel splicing; The number of channels of the deep semantic features is 512, and the number of channels of the shallow detail features is 256 or 128.

[0023] According to an exemplary embodiment of the present invention, the Inception structure of GoogleNet is introduced into the C3 module.

[0024] According to an exemplary embodiment of the present invention, after the neck network cross-scale fuses through depthwise separable convolution, the fused result is fused with the dynamic window attention module, and each level of dynamic window attention module outputs to the detection head. The dynamic window attention module is a SwinTransformerV2 module.

[0025] According to an exemplary embodiment of the present invention, after each level of the dynamic window attention module except the last level, a convolutional attention module is connected; the convolutional attention module enhances feature selection through channel attention and spatial attention.

[0026] According to an exemplary embodiment of the present invention, in step S3, the detection head includes a dynamic window partitioning strategy module, a scaled cosine attention mechanism module, and a hierarchical residual connection module;

[0027] The dynamic window partitioning strategy module is used to adjust the window size according to the resolution of the input feature map;

[0028] The scaled cosine attention mechanism module is used to calculate the cosine similarity of the attention weights and dynamically adjust the weight distribution of different feature channels through a scaling factor;

[0029] The hierarchical residual connection module is used to fuse deep semantic features and shallow detail features.

[0030] According to an exemplary embodiment of the present invention, the cosine similarity calculation of the attention weights of the scaled cosine attention mechanism module adopts formula 4:

[0031]

[0032] Where λ is the learnable channel scaling amount, B is the relative position bias matrix within the window; Attention(Q, K, V) represents the cosine similarity calculation formula of the attention weights of the scaled cosine attention mechanism module, Q represents the query, K represents the key, V represents the value matrix, and T represents the matrix transpose.

[0033] As a second aspect of the present invention, the present invention provides a drone target detection system based on an improved YOLOv5, including: a training image set acquisition module, an annotation module, and an improved YOLOv5 model;

[0034] The training image set acquisition module is used to acquire a training image set;

[0035] The annotation module is used to perform target annotation on each training image to form an annotation file;

[0036] The improved YOLOv5 model is used to train the training image set and the annotation file based on the improved YOLOv5; it is also used to obtain detection targets according to the input image set to be detected.

[0037] As a third aspect of the present invention, the present invention provides an electronic device, including:

[0038] One or more processors;

[0039] A storage device for storing one or more programs;

[0040] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the drone target detection method based on the improved YOLOv5.

[0041] As a fourth aspect of the present invention, the present invention provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the drone target detection method based on the improved YOLOv5 is implemented.

[0042] The advantageous effects of the present invention are:

[0043] By introducing the Swin Transformer V2 module, this solution has more efficient multi-scale modeling capabilities, stronger position perception capabilities, and adaptability to large-size inputs; by introducing Google ResNet through the C3 module, the feature extraction and calculation efficiency are further improved, and the WIoU loss function is used to enhance the progress of small target detection. Therefore, this solution can improve scale adaptability, prevent missed detection of dense targets, and improve calculation efficiency and positioning accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objects, features, and advantages of the present application will become more apparent. The following described drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0045] Figure 1 Schematically shows the structural diagram of the improved YOLOv5 model.

[0046] Figure 2 Schematically shows the step diagram of the drone target detection method based on the improved YOLOv5.

[0047] Figure 3 Schematically shows the structural diagram of the C3 module.

[0048] Figure 4 Schematically shows the first image to be detected.

[0049] Figure 5 Schematically shows the schematic diagram of the target detected by the drone target detection method based on the improved YOLOv5.

[0050] Figure 6 Schematically shows the schematic diagram of the target in other scenarios detected by the drone target detection method based on the improved YOLOv5.

[0051] Figure 7 Schematically shows the second image to be detected.

[0052] Figure 8 Schematically shows a schematic diagram of a target detected by a drone target detection method using YOLOv5.

[0053] Figure 9 Schematically shows a schematic diagram of a target detected by a drone target detection method based on improved YOLOv5.

[0054] Figure 10 Schematically shows a structural diagram of an electronic device.

[0055] Figure 11 Schematically shows a structural diagram of a computer medium. Detailed implementation manners

[0056] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0057] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.

[0058] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0059] The flowcharts shown in the drawings are only illustrative and not necessarily include all the contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0060] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below may be referred to as the second component without departing from the teachings of the concept of this application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0061] Those skilled in the art can understand that the drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing this application, so they cannot be used to limit the protection scope of this application.

[0062] According to the first specific embodiment of the present invention, the present invention provides a drone target detection system based on improved YOLOv5, including: a training image set acquisition module, an annotation module, and an improved YOLOv5 model.

[0063] The training image set acquisition module is used to acquire a training image set.

[0064] The annotation module is used to perform target annotation on each training image to form an annotation file.

[0065] The improved YOLOv5 model is used to train the training image set and the annotation file based on the improved YOLOv5; it is also used to obtain detection targets according to the input image set to be detected.

[0066] As Figure 1 shown, the improved YOLOv5 model includes an input end (not shown), a backbone network, a neck network, and a detection head connected in sequence.

[0067] The backbone network is located Figure 1 on the left side of, and is used to extract features. The backbone network includes multiple levels of depthwise separable convolution modules, multiple levels of C3 modules, a spatial pyramid pooling module, and a dynamic window attention module (STV2 module, i.e., SwinTransformer V2 module). Each level of C3 module is located between two levels of depthwise separable convolution modules; each level of C3 module includes 3 or 9 C3 modules. Each C3 module includes various convolution modules of different sizes. Specifically, the structure of the backbone network is in sequence: the first-level depthwise separable convolution module, the first-level C3 module (including 3 C3 modules), the second-level depthwise separable convolution module, the second-level C3 module (including 9 C3 modules), the third-level depthwise separable convolution module, the third-level C3 module (including 9 C3 modules), the fourth-level depthwise separable convolution module, the spatial pyramid pooling module, the first-level dynamic window attention module (including 3 dynamic window attention modules), where the first-level C3 module, the second-level C3 module, the third-level C3 module, and the first-level dynamic window attention module output to the neck network.

[0068] The neck network is located in Figure 1 the middle part, and is used for feature fusion. The neck network cross-scale fuses the deep semantic features and shallow detail features output by the C3 module through depthwise separable convolution, upsampling, and channel concatenation; the number of channels of the deep semantic features is 512, and the number of channels of the shallow detail features is 256 or 128. After cross-scale fusion by the neck network through depthwise separable convolution, the fused result is fused with the dynamic window attention module, and each level of the dynamic window attention module outputs to the detection head. The dynamic window attention module is a Swin Transformer V2 module. Figure 1 In, the dynamic window attention module of the four-level neck network outputs to the detection head. After each level of the dynamic window attention module except the last level is connected to the convolutional attention module, it is connected to the detection head; the convolutional attention module enhances feature selection through channel attention and spatial attention. Specifically, the output of the first-level C3 module passes through channel concatenation and outputs the first-level feature fusion result through the dynamic window attention module; the output of the dynamic window attention module passes through the convolutional attention module and the depthwise separable convolution module, and is concatenated with the output of the second-level C3 module through channel concatenation, C3 module, convolutional attention module, and depthwise separable convolution module, and outputs the second-level feature fusion result through the dynamic window attention module; the output of the second-level C3 module passes through channel concatenation, C3 module, convolutional attention module, and depthwise separable convolution module, and is upsampled and concatenated with the first-level C3 module through channels; the output of the second-level feature fusion result passes through the convolutional attention module and the depthwise separable convolution module, and is concatenated with the output of the third-level C3 module through channel concatenation, C3 module, convolutional attention module, and depthwise separable convolution module, and outputs the third-level feature fusion result through the dynamic window attention module; the output of the first-level dynamic window attention module passes through the depthwise separable convolution module and the upsampling module and is concatenated with the output of the third-level C3 module through channels; the output of the first-level dynamic window attention module passes through the output of the depthwise separable convolution module and is concatenated with the output of the third-level feature fusion result through the convolutional attention module and the depthwise separable convolution module, and outputs the fourth-level feature fusion result through the dynamic window attention module.

[0069] The detection head is used to output detection information.

[0070] According to the second specific embodiment of the present invention, the present invention provides a drone target detection method based on improved YOLOv5, and uses the drone target detection system based on improved YOLOv5 of the second specific embodiment, as Figure 2 shown, including the following steps:

[0071] S1: Obtain a training image set.

[0072] First, additional aerial data is added based on the VisDrone2019 dataset, and the original aerial video stream is converted into the MP4 format with a resolution of 1080×720 and a frame rate of 30 FPS using the FFmpeg tool.

[0073] For each processed video, images are extracted at an interval of one minute / frame, and 40 images are generated for each video. The images are named in the format of "video name_sequence number" and stored in folders according to the video source.

[0074] S2: Target annotation is performed on each training image to form an annotation file.

[0075] The LabelImg or CVAT tool is used to perform target annotation on each image, and the annotation format is the YOLO standard format (class index x_center y_center width height). The annotation file has the same name as the image and the extension.txt.

[0076] The annotation categories need to be exactly the same as the 11 categories (pedestrians, vehicles, drones, etc.) defined in the VisDrone2019 dataset to ensure compatibility when merging multiple datasets. The 6471 training images and 548 validation images of the official VisDrone2019 dataset are imported in the original annotation format, retaining their original directory structure (images / train / , labels / train / , etc.).

[0077] The newly generated aerial images and their annotation files are respectively merged into the corresponding images / train / and labels / train / directories to form an extended complete dataset. Then, the dataset is preprocessed and sent to step S3 for training.

[0078] S3: The training image set and annotation files are trained based on the improved YOLOv5 to obtain an improved YOLOv5 model.

[0079] As Figure 1 shown, the improved YOLOv5 model includes an input end ( Figure 1 not shown in

[0080] The backbone network is used to extract features: The backbone network includes multi-level depthwise separable convolution modules, multi-level C3 modules, and dynamic window attention modules; Each level of C3 module is located between two levels of depthwise separable convolution modules; Each level of C3 module includes 3 or 9 C3 modules. Each C3 module includes various convolution modules of different sizes. Specifically, the structure of the backbone network is as follows: The first-level depthwise separable convolution module, the first-level C3 module (including 3 C3 modules), the second-level depthwise separable convolution module, the second-level C3 module (including 9 C3 modules), the third-level depthwise separable convolution module, the third-level C3 module (including 9 C3 modules), the fourth-level depthwise separable convolution module, the spatial pyramid pooling module, the first-level dynamic window attention module (including 3 dynamic window attention modules), where the first-level C3 module, the second-level C3 module, the third-level C3 module, and the first-level dynamic window attention module are output to the neck network.

[0081] The convolutional part of the backbone network is replaced with depthwise separable convolution, which decomposes the standard convolution into two steps: depth convolution and pointwise convolution.

[0082] Each level of depthwise separable convolution module realizes downsampling through strided convolution (Stride = 2). The depthwise separable convolution module is used for depth convolution and pointwise convolution, and formula 1 is adopted:

[0083] Conv_DS(X) = PointwiseConv(DepthwiseConv(X)) Formula 1;

[0084] Among them, Conv_DS(X) represents depthwise separable convolution, DepthwiseConv represents depth convolution, and PointwiseConv represents pointwise convolution.

[0085] Depth convolution: Independently perform spatial filtering on each input channel, with a convolution kernel size of 3×3 and a stride of 1 or 2 (during downsampling).

[0086] Pointwise convolution: 1×1 convolution integrates cross-channel information, and the number of output channels is dynamically adapted (such as 256→512).

[0087] Lightweight feature extraction of the backbone network: Use multi-level depthwise separable convolution (Conv_DP) to replace the standard convolution (as Figure 1 shown). Each level contains: Conv_DP module (depthwise separable convolution module): Consisting of depth convolution (3×3, number of groups = input channels) and pointwise convolution (1×1), the number of parameters is reduced to 35% of the original YOLOv5, and the storage of the existing operations is reduced.

[0088] The C3 module is based on the multi-branch Inception idea of GoogleNet, introducing the Inception structure of GoogleNet and optimizing it with cross-stage connections, significantly enhancing the feature expression ability and computational efficiency. As Figure 3 shown, the input feature map is divided into two paths, namely Path 1 and Path 2. Path 2 is directly concatenated with the final output to retain the original information. The multi-branch Inception processing is Path 1, which has a total of 4 branches. From left to right, they are Branch 1, Branch 2, Branch 3, and Branch 4. Branch 1 is a 1×1 convolution to extract local details, and the number of output channels = the number of input channels / 4. Branch 2 is a 3×3 convolution to extract spatial features in a lightweight manner, and the number of output channels = the number of input channels / 4; Branch 3 is a 5×5 convolution to expand the receptive field, and the number of output channels = the number of input channels / 4. Branch 4 is a combination of global average pooling and 1×1 convolution to capture global context, and the number of output channels = the number of input channels / 4. The multi-branches perform batch normalization, and the results of the 4 branches are concatenated along the channel dimension, and the number of output channels = the number of input channels. Then ReLU activation is performed, and convolution is carried out in three paths and then merged. The three paths are 1×1 convolution, 1×1 convolution and 3×3 convolution, 1×1 convolution and 1×1 convolution. Among them, 1×1 convolution can effectively reduce or increase the number of channels, that is, change the depth (number of channels) of the feature map.

[0089] Based on the residual connection optimization, that is, the ResNet idea, after concatenating Path 1 and Path 2, the number of channels is adjusted through 1×1 convolution and added to the original input residually, using Formula 2:

[0090] X out =X in +Conv 1×1 (BatchNorn(PathAK, PathB)) Formula 2;

[0091] Among them, X out represents the output of the C3 module, X in represents the original input, Conv 1×1 represents 1×1 convolution, PathA represents Path 1, and PathB represents Path 2.

[0092] By enhancing the multi-scale feature extraction ability through the Inception multi-branch structure, features under different receptive fields can be captured simultaneously; combined with the ResNet residual connection to optimize gradient propagation, while retaining the lightweight characteristics of the CSP cross-stage partial connection, alleviating the problem of small gradients in deep networks, the C3-Inception-Res module, that is, the improved C3 module of this scheme, significantly improves the accuracy and efficiency of UAV target detection.

[0093] Spatial pyramid pooling is used to output multi-scale receptive field features, enhancing the robustness of the target scale.

[0094] The neck network is used for feature fusion.

[0095] The neck network cross-scale fuses the deep semantic features (such as the target shape) and shallow detail features (such as edges and textures) output by the C3 module through depthwise separable convolution, upsampling, and channel concatenation; the number of channels of the deep semantic features is 512, and the number of channels of the shallow detail features is 256 or 128.

[0096] After the neck network cross-scale fuses through depthwise separable convolution, the fused result is fused with the dynamic window attention module, and each level of the dynamic window attention module outputs to the detection head. The dynamic window attention module is the SwinTransformerV2 module.

[0097] The implementation steps of SwinTransformerV2 (STV2) are as follows:

[0098] When the matrix is fed in, the input shape is [batch size, number of channels, feature width, feature height].

[0099] 1. First, according to the size of the image and the window size, a mask for the window attention mechanism is generated. This mask determines the interaction pattern between different windows. Here, a local window mechanism is adopted, where each window only calculates the self-attention of the elements within the window, and by setting the masking relationship between different windows, the high computational complexity of global calculation is avoided.

[0100] 2. Then, the input matrix is subjected to a Padding operation to ensure that the size of the image is divisible by the window size. The Padding attribute is used to set the spacing between the element content and its border. This step is also a dynamic window partitioning strategy.

[0101] 3. For each window, calculate the self-attention within the window. This step is to obtain the weighted representation of each element by calculating the relationship between each element within the window and other elements. Apply the mask of the attention mechanism obtained in step 1 to each window to ensure that there is no interference between windows during the attention calculation. Obtain the attention output.

[0102] 4. The calculated attention output is converted into the shape of the window by recombining the elements in the window back into the spatial shape of the original matrix.

[0103] 5. After the attention calculation, use a residual connection to add the input and the calculated output, and perform random dropout (for regularization). This helps the model avoid overfitting during training.

[0104] 7. After passing through the attention mechanism and residual connection, the input undergoes a non-linear transformation, which includes a linear layer and an activation function.

[0105] 8. Finally, after the residual connection and non-linear transformation, the output feature map undergoes a transpose operation to restore it to the shape of [batch size, number of channels, feature width, feature height], which is used as the output of this layer.

[0106] In each SwinTransformerV2 module, a residual connection is used to connect the input and the output of the attention layer and the feed-forward neural network layer. Specifically, the role of the residual connection is to directly add the input to the output, avoiding the situation of gradient disappearance or information loss in deep neural networks.

[0107] Swin TransformerV2 (STV2) has the following advantages:

[0108] 1. More efficient multi-scale modeling ability

[0109] The hierarchical window mechanism of STV2: Through hierarchical window attention, STV2 supports multi-scale feature fusion. In drone images, there are large differences in object scales (such as small targets captured at high altitude and large targets at low altitude). The hierarchical structure of STV2 can capture context information at different scales more naturally.

[0110] Limitations of traditional TPH: TPH-YOLOv5 addresses the multi-scale problem by adding prediction heads, but the feature interaction of each scale head relies on traditional convolution, while the cross-window attention of SwinV2 can fuse global and local information more flexibly.

[0111] 2. Stronger position perception ability

[0112] Relative Position Bias: The relative position encoding within the window introduced by SwinV2 can better model the spatial relationship between objects, which is crucial for distinguishing densely arranged objects (such as crowds and vehicles) in drone images.

[0113] Shortcomings of TPH: The absolute position encoding of the standard Transformer may not be sensitive enough to the positioning of dense small targets, while the local position bias of SwinV2 is more suitable for high-density scenarios.

[0114] 3. Adaptability to large-size inputs

[0115] Continuous window partitioning strategy: SwinV2 constructs pyramid features through continuous downsampling (such as 4x4 → 8x8 → 16x16), which is more compatible with the characteristics of "large field of view at high altitude and small targets at low altitude" in drone images.

[0116] Fixed resolution limitation of TPH: TPH-YOLOv5 relies on a prediction head with a fixed number of layers, which may be insufficiently adaptable to extreme scale changes (such as tiny targets at high altitudes).

[0117] After each level of the dynamic window attention module except the last level, a convolutional attention module is connected; the convolutional attention module enhances feature selection through channel attention and spatial attention. After each level of the dynamic window attention module, a convolutional attention module (i.e., the CBAM module) is connected to dynamically calibrate the feature response through channel attention and spatial attention, improving the target discrimination in complex backgrounds.

[0118] The detection head is used to output detection information.

[0119] The feature fusion result of each level of the neck network is connected to the detection head after depthwise separable convolution. Figure 1 In this scheme, the detection head has 4 detection heads.

[0120] The detection head is an object detection head improved based on the SwinTransformerV2 architecture, which is optimized for the multi-scale and dense small target detection problems in the UAV aerial photography scenario. Therefore, the detection head contains ST2PH features, where ST2PH means SwinTransformerV2 prediction head. The default detection head of the original YOLOv5 is the C3+Conv layer, and this scheme is a hierarchical STV2-Head attention head, forming a five-level feature pyramid (P2-P6).

[0121] The detection head includes a dynamic window partitioning strategy module, a scaled cosine attention mechanism module, and a hierarchical residual connection module.

[0122] The dynamic window partitioning strategy module is used to adjust the window size according to the resolution of the input feature map, with the size ranging from 4×4 to 6×6.

[0123] The dynamic window partitioning strategy introduces a resolution adaptive window partitioning algorithm for the input feature Figure X ∈R B×H×W×C (X is the feature matrix, R represents any value, B represents the batch size, H represents the height, W represents the width, and C represents the number of channels) The window size s is dynamically calculated using Equation 3:

[0124]

[0125] Among them, s represents the calculation window size, k represents the basic scaling factor, l represents the pyramid level, and the pyramid level includes levels P2 to P6, with l corresponding to 0 to 4 respectively; H represents the height and W represents the width. The low level (P2) uses a 4×4 window to capture pixel-level detail information (such as drone rotors, vehicle edges), and the high level (P6) uses a 16×16 window to model the global context relationship of large targets.

[0126] The scaled cosine attention mechanism module is used to calculate the cosine similarity of the attention weights (i.e., attention calculation), and dynamically adjust the weight distribution of different feature channels through the scaling factor.

[0127] The Swin transformer first gives an input feature map [batch size, number of channels, feature width, feature height]. We use a linear layer to project the input into Q, K, V (query, key, value), and group it using window partitioning.

[0128] The cosine similarity calculation of the attention weights of the scaled cosine attention mechanism module adopts Equation 4:

[0129]

[0130] Among them, λ is the learnable channel scaling amount (initial value 1.0), B ∈ R (2s-1)×(2s-1) is the relative position bias matrix within the window; Attention(Q, K, V) represents the cosine similarity calculation formula of the attention weights of the scaled cosine attention mechanism module, Q represents the query, K represents the key, V represents the value matrix, which is linearly mapped from the input feature, and T represents the matrix transpose.

[0131] After the cosine similarity calculation of the attention weights, there is a linear layer for projecting the output of the attention.

[0132] By introducing the scaled cosine attention mechanism, it replaces the standard multi-head self-attention (MSA) in the traditional Transformer. Its core lies in calculating the cosine similarity of the attention weights and dynamically adjusting the weight distribution of different feature channels. For example, in the drone target detection scenario, this mechanism can effectively distinguish dense small targets from background noise, such as dealing with the problem of feature confusion caused by motion blur during the low-altitude flight of drones.

[0133] The hierarchical residual connection module is used to fuse deep semantic features and shallow detail features.

[0134] Construct a cross-layer feature fusion path to fuse the shallow detail features (i.e., shallow high-resolution features) F P2 with the deep semantic features F P5 through gated residual connection, adopting Equation 5:

[0135] F fuse = α·Conv 1×1 (F P2 ) + (1 - α)·UpSample(F P5 ) Equation 5;

[0136] where F fuse represents the final feature map fused through the gated residual connection mechanism, Conv1×1 represents a 1×1 convolution, Upsample represents upsampling, F P2 represents the shallow detail features, F P5 represents the deep semantic features, and the gating coefficient α is generated by channel attention. The pyramid levels range from the low level to the high level as P2 to P6.

[0137] For the scaled cosine attention mechanism module, in the original self-attention calculation, the similarity term of pixel pairs is calculated as the dot product of the query vector and the key vector. A scaled cosine attention method is adopted, which calculates the attention of pixel pairs i and j through the scaled cosine function:

[0138] Sim(q i , k j ) = cos(q i , k j ) / τ + B ij ,

[0139] where the above formula is the scaled cosine attention formula, Sim represents the cosine similarity formula, q i represents the query vector, which is used to determine the positions that need to be focused on, k j represents the key vector, which is used to measure its correlation with the current query and is used to judge the parameters that are more important for the classification decision, B ij is the relative position deviation between pixels i and j; τ is a learnable scalar that is not shared across heads and layers. τ is set to be greater than 0.01. The cosine function is naturally normalized, so it has more moderate attention values.

[0140] The above formula is an improvement over the traditional self-attention calculation, which solves the problem that the traditional self-attention calculates the similarity between the query vector and the key vector through the dot product, but the dot product value may be too large due to the high vector dimension, resulting in unstable gradients. The combination of the scaled cosine attention mechanism module and the hierarchical residual connection module enhances feature diversity and solves the feature collapse problem of deep Transformers.

[0141] Specifically, run the train.py file for training.

[0142] The input data is the preprocessed UAV aerial photography dataset (VisDrone2019 + self-built extended set), which is divided into training set, validation set, and test set according to the ratio of 7:2:1.

[0143] The data augmentation strategies include Mosaic splicing (probability 60%), random rotation (±15°), and HSV color gamut perturbation (hue ±0.1, saturation / brightness ±0.5).

[0144] Model configuration: The backbone network uses the improved C3-Inception-Res module to replace the C3 module of the original YOLOv5.

[0145] The loss function is the dynamic weighted WIoU Loss (the weight coefficient of small targets is increased to 1.5).

[0146] The optimizer is selected as AdamW, the initial learning rate is 0.001, and the cosine annealing strategy is adopted (cycle 60 epochs).

[0147] The trained weight model (best.pt) is obtained, that is, the improved YOLOv5 model. The best.pt is input into the test network, and the detection results are saved to runs / detect / exp / , and the labeled box images (JPEG format).

[0148] During the training process, the dynamic weighted WIOU loss function is introduced to improve the detection accuracy of the model on targets of different scales (especially small targets).

[0149] In terms of the loss function, the dynamic weighted WIOU loss is adopted. The dynamic weighted WIoU loss function is defined as:

[0150]

[0151] Among them, L WIoU represents the weighted intersection over union loss, N is the number of predicted boxes (usually multiple candidate boxes in each image), and IoU i is the (intersection over union) value of the i-th predicted box and the ground truth box, which is used to measure the overlap degree between the predicted box and the ground truth box; β i is the dynamic weight of the i-th predicted box, which is used to adjust the contribution of this predicted box in the loss function. The intersection over union (IoU) is an important metric in computer vision for measuring the overlap degree between the predicted result and the ground truth result, and is widely used in object detection and image segmentation tasks.

[0152] Among them, the dynamic weight β i The calculation formula is:

[0153]

[0154] w i 、h i represent the width and height of the i-th predicted bounding box in pixels. w j and h j are the widths and heights of all candidate bounding boxes.

[0155] Compress the scale advantage of large objects through logarithmic transformation and enhance the weight contribution of small objects. After weighting, the loss function pays more attention to objects at specific scales, especially small objects that are more difficult in the detection task. The weight is negatively correlated with the object area (w i h i ), and small objects (w i h i small) obtain higher weights, and their localization accuracy is preferentially optimized during backpropagation. The combination of and the logarithmic function is used to balance the gradient distribution of objects at different scales and avoid over-amplifying the weights of small objects.

[0156] The loss function of the original YOLOv5 usually calculates the overlap between the predicted bounding box and the ground truth box based on IoU (Intersection over Union). The YOLOv5 loss function mainly trains through localization loss, confidence loss, and classification loss, and usually treats large and small objects equally. This means that in the case of small objects, since small objects account for a small proportion in the image, the model may not be able to focus well on the localization and classification of small objects, easily resulting in the failure to detect small objects.

[0157] Compared with the traditional IoU loss, the WIOU loss function has been improved, and the contribution of small objects during training is strengthened through a weighting method. Specifically, WIOU assigns larger weights to small objects, so that small objects receive more attention during training. This weighting method helps the model better learn the features and localization information of small objects, balance the influence of large and small objects, and thus improve the detection accuracy of small objects. In the perspective of group aerial photography, small objects and overlapping objects can be captured more precisely.

[0158] S4: Input the image set to be detected into the improved YOLOv5 model to obtain the detection targets.

[0159] As Figure 4 shown, inputting the image set to be detected is the test set, as Figure 5 and Figure 6As shown, the coordinates and class labels (TXT format) are obtained; the performance report (JSON format, including indicators such as mAP, FPS, and video memory occupancy). After detection, Figure 5 multiple cars, vans, and trucks are detected in Figure 5 . The numbers after the English words indicate the confidence level. Figure 6 multiple cars, people, motorcycles, vans, and pedestrians are detected on the road in Figure 6 .

[0160] Compare this solution with the original YOLOv5. Figure 7 is the image to be detected. Figure 8 is the detection result of the original YOLOv5. Figure 9 is the detection result of the improved YOLOv5 of this solution. From Figure 8 and Figure 9 By comparison, it can be seen that the original YOLOv5 solution has missed detections and low confidence levels. This solution can improve scale adaptability, prevent missed detections of dense targets, and improve computational efficiency and positioning accuracy.

[0161] It can be seen that by introducing the SwinTransformerV2 module, this solution has more efficient multi-scale modeling capabilities, stronger position perception capabilities, and adaptability to large-size inputs; by introducing GoogleResNet through the C3 module, the feature extraction and computational efficiency are further improved, and the WIoU loss function is used to enhance the progress of small target detection. Therefore, this solution can improve scale adaptability, prevent missed detections of dense targets, and improve computational efficiency and positioning accuracy.

[0162] According to the third specific implementation manner of the present invention, the present invention provides an electronic device, as Figure 10 shown. Figure 10 is a block diagram of an electronic device shown according to an exemplary embodiment.

[0163] Next, refer to Figure 10 to describe the electronic device 100 according to this implementation manner of the present application. Figure 10 The electronic device 100 shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.

[0164] As Figure 10 shown, the electronic device 100 is presented in the form of a general-purpose computing device. The components of the electronic device 100 may include, but are not limited to: at least one processing unit 110, at least one storage unit 120, a bus 130 connecting different system components (including the storage unit 120 and the processing unit 110), a display unit 140, etc.

[0165] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 110, so that the processing unit 110 executes the steps according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 110 can execute the steps shown in the second specific embodiment.

[0166] The storage unit 120 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 3201 and / or a cache storage unit 1202, and may further include a read-only storage unit (ROM) 1203.

[0167] The storage unit 120 may further include a program / utility 1204 having a set (at least one) of program modules 1205. Such program modules 1205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0168] The bus 130 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0169] The electronic device 100 can also communicate with one or more external devices 100' (such as a keyboard, a pointing device, a Bluetooth device, etc.), so that the device that enables a user to interact with the electronic device 100 communicates, and / or the electronic device 100 can communicate with any device that can communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 150. And, the electronic device 100 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 160. The network adapter 160 can communicate with other modules of the electronic device 100 through the bus 130. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0170] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a way of software combined with necessary hardware.

[0171] Therefore, according to the fourth specific embodiment of the present invention, the present invention provides a computer-readable medium. AsFigure 11 As shown, the technical solution according to an embodiment of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, server, or network device, etc.) to execute the above method according to an embodiment of the present invention.

[0172] The software product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0173] The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0174] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0175] The above computer-readable medium carries one or more programs, which, when executed by the device, cause the computer-readable medium to implement the functions of the second specific embodiment.

[0176] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are different from this embodiment only. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.

[0177] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.

[0178] The above specifically shows and describes the exemplary embodiments of the present invention. It should be understood that the present invention is not limited to the detailed structures, settings or implementation methods described herein; on the contrary, the present invention is intended to cover various modifications and equivalent settings included in the spirit and scope of the appended claims.

Claims

1. A drone target detection method based on improved YOLOv5, characterized in that: The following steps are involved: S1: Get the training image set; S2: label each training image and generate a labeling file; S3: training the training image set and the annotation file based on the improved YOLOv5 to obtain an improved YOLOv5 model; The improved YOLOv5 model includes an input end, a backbone network, a neck network and a detection head connected in sequence; The backbone network is used to extract features; the backbone network includes a multi-level depth-separable convolution module, a multi-level C3 module and a dynamic window attention module; each level of C3 module is located between two levels of depth-separable convolution modules; each level of C3 module includes 3 or 9 C3 modules; The neck network is used for feature fusion; The detection head is used to output detection information; S4: Input the image set to be detected into the improved YOLOv5 model to obtain the detection target.

2. The unmanned aerial vehicle target detection method based on improved YOLOv5 according to claim 1, characterized in that: In step S3, each level of depth-wise separable convolution module implements downsampling through strided convolution. The depth-wise separable convolution module is used to perform depth-wise convolution and point-wise convolution, using formula 1: Conv_DS(X)=PointwiseConv(DepthwiseConv(X)) Formula 1; Among them, Conv_DS(X) represents depthwise separable convolution, Depthwise Conv represents depthwise convolution, and PointwiseConv represents pointwise convolution.

3. The unmanned aerial vehicle target detection method based on improved YOLOv5 according to claim 2, characterized in that: Each C3 module includes a variety of convolution modules of different sizes. The neck network integrates the deep semantic features and shallow detail features output by the C3 module across scales through depthwise separable convolution, upsampling, and channel splicing. The number of channels for deep semantic features is 512, and the number of channels for shallow detail features is 256 or 128.

4. The unmanned aerial vehicle target detection method based on improved YOLOv5 according to claim 3, characterized in that: After the neck network is fused across scales through deep separable convolution, the fused result is fused with the dynamic window attention module. Each level of the dynamic window attention module is output to the detection head. The dynamic window attention module is a Swin Transformer V2 module.

5. The unmanned aerial vehicle target detection method based on improved YOLOv5 according to claim 4, characterized in that: Each level of dynamic window attention module except the last level is followed by a convolutional attention module; the convolutional attention module enhances feature selection through channel attention and spatial attention.

6. The unmanned aerial vehicle target detection method based on improved YOLOv5 according to claim 1, characterized in that: In step S3, the detection head includes a dynamic window division strategy module, a scaled cosine attention mechanism module, and a hierarchical residual connection module; The dynamic window partitioning strategy module is used to adjust the window size according to the input feature map resolution; The scaled cosine attention mechanism module is used to calculate the cosine similarity of the attention weights and dynamically adjust the weight distribution of different feature channels through the scaling factor; The hierarchical residual connection module is used to fuse deep semantic features with shallow detail features.

7. The unmanned aerial vehicle target detection method based on improved YOLOv5 according to claim 6, characterized in that: The cosine similarity calculation of the attention weight of the scaled cosine attention mechanism module adopts Formula 4: Among them, λ is the learnable channel scaling amount, B is the relative position bias matrix within the window; Attention(Q, K, V) represents the cosine similarity calculation formula of the attention weight of the scaled cosine attention mechanism module, Q represents the query, K represents the key, V represents the value matrix, and T represents the matrix transpose.

8. A drone target detection system based on improved YOLOv5, characterized in that: include: Training image set acquisition module, annotation module, and improved YOLOv5 model; The training image set acquisition module is used to acquire the training image set; The annotation module is used to annotate each training image to form an annotation file; The improved YOLOv5 model is used to train the training image set and the annotation file based on the improved YOLOv5; and is also used to obtain the detection target according to the input image set to be detected.

9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the drone target detection method based on the improved YOLOv5 as described in any one of claims 1 to 6.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, the drone target detection method based on the improved YOLOv5 as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Insulating material deterioration early warning and service life prediction system

    CN120927550A

  • Deep learning-based target detection method and system under view angle of low-altitude unmanned aerial vehicle

    CN121545082A

  • Target detection method and system under low-altitude unmanned aerial vehicle perspective based on deep learning

    CN121545082B

  • Underwater sonar image target detection method and device

    CN121746898A