Improved YOLOv11 unmanned aerial vehicle small target detection method

By introducing the bidirectional feature pyramid structure, RepVitBlock module, channel-space attention mechanism and Inner-WIoUv3 loss function on YOLOv11, the problem of low accuracy and missed detection and missed detection under complex backgrounds and target scales is solved, and significant performance improvement is achieved.

CN119992524APending Publication Date: 2025-05-13JIMEI UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510066511.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the complex context and the small target scale of existing drone target detection technologies, there are problems of low accuracy, missed detection and missed detection.

Method used

On the basis of YOLOv11, Bidirectional feature pyramid structure BiFPN was introduced, combined with P2 small object detection layer, and S-BiFPN feature fusion structure was constructed; RepVitBlock module and channel-space attention mechanism CSA improved the C3k2 module; Inner-WIoUv3 loss function was constructed, and a small object detection model was trained based on the CBD data set.

Benefits of technology

It significantly improves the accuracy of drone target detection, improves the problems of missed and missed detection, and performs well especially in cases of complex backgrounds and large changes in target scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992524A_ABST
    Figure CN119992524A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle small target detection method based on an improved YOLOv11, and the method comprises the steps: introducing a bidirectional feature pyramid structure BiFPN (Bidirectional Feature Pyramid Network) on the basis of the YOLOv11 to improve an original top-down structure of a feature fusion layer, and constructing an S-BiFPN feature fusion structure (S represents Small object) in combination with a P2 small target detection layer; the method comprises the following steps: introducing a RepVitBlock module and a channel-space attention mechanism (CSA) (Channel-Space Attention) to improve a C3k2 module; an Inner-WIoUv3 loss function is constructed; training the small target detection model based on the CBD data set to obtain an optimal small target detection model; and inputting the test set into the optimal small target detection model to obtain a detection result of the model. The method is mainly applied to an anti-unmanned aerial vehicle target detection task, can improve missed detection and false detection of the unmanned aerial vehicle, and has remarkable performance improvement and guarantees safe development of low-altitude economy under the conditions of a complex background, large target scale change and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of drone detection, image recognition, etc., and specifically relates to a drone small target detection method based on improved YOLOv11. Background Art

[0002] In recent years, with the rapid development of the drone industry, the threshold for drones has continued to decrease. With their low cost, high flexibility and versatility, drones are widely used in agriculture, communications, transportation or film shooting, and have become an important part of the low-altitude economy. At present, the detection method for low-altitude drones is mainly to use radar detection combined with photoelectric recognition, ultrasound and radio to detect all flying targets within the visible range of the naked eye. However, these traditional methods often have problems such as limited detection range, insufficient accuracy and high cost. For example, the use of visible light images to identify low-altitude drones is easily interfered, and birds in the air are mistakenly detected as drone targets, requiring additional manual secondary identification, resulting in a waste of resources. To solve such problems, deep learning target detection technology has been widely used, among which common typical deep learning frameworks include YOLO, Faster R-CNN and SSD. These methods can automatically extract image features through neural networks, effectively distinguish targets from background information, and achieve the purpose of accurately identifying and locating drone targets. Compared with traditional drone target detection technology, deep learning methods show better adaptability and robustness in the face of various complex environmental backgrounds, and can cope with more diverse challenges.

[0003] Although the above-mentioned deep learning-based target detection research has proposed a good detection algorithm for drone targets, the drone targets still have low accuracy, missed detection and false detection problems when the scale of the drone targets varies greatly under various complex backgrounds and the targets are too small. Therefore, studying an efficient and accurate drone small target detection technology to solve the defects of the above-mentioned existing methods is an important topic to ensure the safety of airport operations and maintain urban safety. Summary of the invention

[0004] In view of the defects and shortcomings of the prior art, the purpose of the present invention is to provide a method for detecting small targets of drones by improving YOLOv11, which can improve the accuracy of the detection model under complex backgrounds and when the drone target scale is too small, and can improve the problems of missed detection and false detection of drones.

[0005] Based on YOLOv11, the present invention introduces a bidirectional feature pyramid structure BiFPN (Bidirectional Feature Pyramid Network) to improve the original top-down structure of the feature fusion layer, and combines the P2 small target detection layer to construct an S-BiFPN feature fusion structure (S represents Small object); introduces the RepVitBlock module and the channel-spatial attention mechanism CSA (Channel-SpatialAttention) to improve the C3k2 module; constructs the Inner-WIoUv3 loss function; trains the above small target detection model based on the CBD data set to obtain the optimal small target detection model; inputs the test set into the optimal small target detection model to obtain the detection result of the model. The present invention can improve the situation of missed detection and false detection of drones, and has significant performance improvement in drone target detection under complex backgrounds and large changes in target scale.

[0006] The technical solution specifically adopted by the present invention to solve the technical problem is:

[0007] A method for detecting small targets of unmanned aerial vehicles (UAVs) by improving YOLOv11 is proposed: based on YOLOv11, a bidirectional feature pyramid structure BiFPN is introduced, and combined with the P2 small target detection layer, feature information flows from the P5 layer to the P2 layer of the network, and then from the P2 layer to the P5 layer, so as to construct a bidirectional information flow and form a feature fusion structure S-BiFPN; in the neck network, a RepVitBlock module is introduced based on the C3k2 module to replace the original Bottleneck module to form a C3k2-RVB module; in the backbone network, a channel-spatial attention mechanism CSA is introduced based on the C3k2-RVB module to form a C3k2-RC module to replace the C3k2 module; in the detection head, a loss function Inner-WIoUv3 constructed by introducing an Inner-IoU auxiliary frame based on the Wise-IoUv3 loss function is adopted; thereby, a target detection model is improved and trained; and the optimal UAV small target detection model obtained by training is used to realize UAV small target detection.

[0008] Furthermore, the modules of the neck network are connected via a feature fusion structure S-BiFPN.

[0009] Furthermore, in the RepVitBlock module, the input feature map first passes through a 3×3 deep convolution module DWConv to extract local spatial features, then passes through a 1×1 deep convolution module to perform channel adjustment on the features, and then passes through two 1×1 convolution modules Conv in sequence for feature fusion; the input feature map is also added to the processed features through a jump connection to form a residual structure, and finally the refined and fused feature map is output for use by subsequent modules; in the inference stage, the two parallel deep convolution modules are simplified into a 3×3 deep convolution module using the structural reparameterization technology.

[0010] Furthermore, the feature fusion structure S-BiFPN adopts a weighted feature fusion strategy in the process of feature fusion: different weights are assigned when fusing feature maps of different scales. The weights are learnable parameters and are continuously adjusted and optimized during the model training process to achieve the best fusion effect.

[0011] Furthermore, the channel-spatial attention mechanism CSA is a composite attention module, which is divided into a channel attention branch and a spatial attention branch: in the channel attention branch, the input feature map is first globally averaged pooled in the spatial dimensions H and W, and a 1×1×C feature map representing the global features of each channel is output; then, through a linear transformation, the number of channels is mapped to C / r, where r is the channel scaling factor; then the activation function GELU is applied to introduce nonlinear characteristics; after a linear transformation, the number of channels is restored to C; finally, the weight at each channel position is obtained through the Sigmoid activation function, and the obtained weight is used to weight the feature map in the channel dimension; in the spatial attention branch, a H×W×1 feature map is output through a 7×7 convolution module, and then the Sigmoid activation function is applied to generate weights at each spatial position, and the input feature map is weighted in the spatial dimension using the obtained spatial weights.

[0012] Furthermore, the construction process of the loss function Inner-WIoUv3 is specifically as follows:

[0013] The calculation formula of Wise-IoUv3 is as follows:

[0014] L IoU =1-IoU (1)

[0015] L WIoUv3 = rRL IoU (2)

[0016]

[0017] In the formula, IoU represents the ratio of the intersection and union of the real box and the predicted box, and the loss function L IoUis the corresponding loss function; L WIoUv3 is the dynamic loss function; r is the non-monotonic focus factor; β is the outlier point, which is L IoU is the mean of the anchor boxes, used to evaluate the quality of the anchor boxes; δ and α are adjustable hyperparameters; R is the weight coefficient used to adjust the IoU loss; x, y and x gt ,y gt Represent the center coordinates of the predicted box and the real box respectively; W g ,H g Represents the width and height of the minimum bounding box of the predicted box and the real box respectively;

[0018] The auxiliary box is introduced through Inner-IoU to assist in calculating the loss. The size of the auxiliary box is adjusted according to the actual needs of the model through a scale factor ratio to construct the Inner-WIoUv3 loss function; the calculation formula is as follows:

[0019]

[0020] union=(w gt ×h gt )×(ratio) 2 + (w×h)×(ratio) 2 -inter (8)

[0021]

[0022] L innerIoU =1-IoU inner (10)

[0023]

[0024] Among them, b gt and b represent the true box and predicted box respectively. Indicates the center point coordinates of the real frame and the auxiliary frame, (x c ,y c ) is represented by the center point coordinates of the prediction box and the prediction auxiliary box; w gt and h gt They represent the width and height of the real box respectively, w and h represent the width and height of the predicted box respectively; the ratio ranges from 0.5 to 1.5.

[0025] An edge computing terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of an improved YOLOv11 method for detecting small targets in unmanned aerial vehicles are implemented as described above.

[0026] A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of an improved YOLOv11 method for detecting small targets in unmanned aerial vehicles as described above.

[0027] Compared with the prior art, the present invention and its preferred solution are aimed at detecting drone targets under complex backgrounds and with large changes in target scales, and can effectively realize accurate identification and positioning of drone targets, have high detection accuracy, and at the same time improve the problems of missed detection and false detection of drones to a certain extent. At the same time, compared with the YOLOv11 model, the present invention also has certain advantages in parameter quantity indicators. Therefore, the present invention has a good detection effect for drone targets, and can effectively ensure the safety management of low-altitude areas under the background of rapid development of the low-altitude economy, and has good application prospects. It is mainly used in anti-drone target detection tasks, which can improve the situation of missed detection and false detection of drones, and has significant performance improvement under complex backgrounds and large changes in target scales, ensuring the safe development of the low-altitude economy. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0029] Figure 1 An overall flow chart of a method for detecting small targets of unmanned aerial vehicles using an improved YOLOv11 model proposed in an embodiment of the present invention;

[0030] Figure 2 This is a diagram of the improved RSB-YOLOv11 network structure according to an embodiment of the present invention;

[0031] Figure 3 This is a structural diagram of the RepVitBlock module in an embodiment of the present invention;

[0032] Figure 4 A schematic diagram of the training and reasoning structure under the structure re-parameterization technology of an embodiment of the present invention;

[0033] Figure 5 This is a structural diagram of the channel-spatial attention mechanism CSA of an embodiment of the present invention;

[0034] Figure 6 The improved C3k2 module structure diagram of the embodiment of the present invention, wherein (a) is the improved C3k2-RVB structure diagram, and (b) is the improved C3k2-RC structure diagram;

[0035] Figure 7 This is a structural diagram of a bidirectional feature pyramid BiFPN according to an embodiment of the present invention;

[0036] Figure 8 This is a schematic diagram of Inner-IoU according to an embodiment of the present invention;

[0037] Fig. 9 The diagrams show the effect of improving the loss function according to the embodiment of the present invention, wherein (a) is a comparison diagram of the improved loss function and other loss function evaluation index Map@0.5, and (b) is a comparison diagram of the improved loss function and other loss function evaluation index F1 score;

[0038] Fig.10 1 is a diagram showing the detection effect of an embodiment of the present invention; wherein (a), (b), and (c) are comparisons of the detection effect of the original model and the improved model of the present invention on UAV targets under different complex backgrounds. DETAILED DESCRIPTION

[0039] In order to make the features and advantages of the present invention more clearly understood, the following embodiments are specifically described in detail as follows:

[0040] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0041] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0042] In order to cope with the actual challenge of frequent "illegal flying" phenomenon, the embodiment of the present invention designs a new UAV small target detection method for anti-UAV target detection tasks, which can improve the situation of missed detection and false detection of UAVs, and has significant performance improvement under complex backgrounds and large changes in target scale. Experiments were conducted on the public dataset CBD, and the results showed that the improved YOLOv11n algorithm reached 94.8% and 57.7% on Map@0.5 and Map@0.5-0.95, respectively, which is 4.2% and 4.4% higher than the original algorithm.

[0043] like Figure 1 As shown, this embodiment provides a specific implementation process of a method for detecting small targets of unmanned aerial vehicles by improving YOLOv11, including the following steps:

[0044] Step S1: Use the public data set CBD (the Complex Background Dataset) and convert it into the format for YOLO model training, dividing it into a training set and a test set;

[0045] Step S2, introduce the bidirectional feature pyramid structure BiFPN (Bidirectional Feature Pyramid Network) to improve the original top-down structure of the feature fusion layer, and combine it with the P2 small target detection layer to construct the S-BiFPN feature fusion structure (S here stands for Small object);

[0046] Step S3, introduce the lightweight module RepVitBlock and the channel-spatial attention mechanism CSA (Channel-Spatial Attention) to improve the C3k2 module;

[0047] Step S4, construct a new loss function: in the Wise-IoUv3 loss function, the Inner-WIoUv3 loss function is constructed by combining the auxiliary box of the Inner-IoU loss function;

[0048] Step S5: Use the training set in the CBD data set mentioned in S1 to train the target detection model constructed by S2-S4 to obtain the optimal UAV small target detection model;

[0049] Step S6: Input the test set in the data set into the small target detection model in S5 to obtain the detection result of the model.

[0050] As a preferred solution of this embodiment, in step S1, the CBD dataset includes two categories: drone and bird, with a total of 12,490 pictures, of which 9,725 are training sets and 2,765 are test sets, including various complex scenes, such as urban buildings, outdoor, sunny days, and mountains, etc., which are in line with actual detection scenarios. Compared with public drone datasets such as RealWorld and MIDGARD, this dataset has rich scenes, varied perspectives, and many types of drones with large scales of variation. The present invention uses homemade Python code to convert it into a form that can be trained by the YOLO model.

[0051] As a preferred solution of this embodiment, in step S2, the feature fusion structure S-BiFPN is improved on the basis of the bidirectional feature pyramid structure BiFPN. The structure of the bidirectional feature pyramid structure BiFPN is as follows: Figure 7As shown in the figure, the feature fusion structure S-BiFPN adds reverse feature information transmission, specifically, the feature information flows from the P5 layer to the P2 layer of the network, and then from the P2 layer to the P5 layer, to build a two-way information flow. This two-way information transmission can make the feature maps of different levels more fully integrated, so as to improve the performance of the model. At the same time, in the process of feature fusion, a weighted feature fusion strategy is adopted, that is, when the feature maps of different scales are fused, different weights will be assigned to them. These weights are learnable parameters, which will be continuously adjusted and optimized during the model training process to achieve the best fusion effect. Combined with the P2 small target detection layer, while maintaining the original multi-scale feature fusion capability, the perception ability of small target detection of drones is improved.

[0052] As a preferred solution of this embodiment, in step S3, the improved C3k2 modules are respectively C3k2-RVB modules and C3k2-RC modules.

[0053] The C3k2-RVB module is based on the C3k2 module and introduces a lightweight module RepVitBlock to replace the original Bottleneck module. Figure 6 (a). Figure 3 As shown in the figure, the RepVitBlock module uses the MetaFormer architecture design of MobileNetV3 to implement a separate token mixer and channel mixer. The input feature map first passes through a 3×3 deep convolution module DWConv to extract local spatial features, then passes through a 1×1 deep convolution module to adjust the channel of the features, and then passes through two 1×1 convolution modules Conv in sequence for feature fusion. At the same time, the input feature map is added to the processed features through jump connections to form a residual structure, and finally the refined and fused feature map is output for use by subsequent modules. This structural composition effectively extracts multi-scale feature information, enabling the model to integrate the global information and context information of the drone target.

[0054] In the inference phase, structural reparameterization techniques such as Figure 4 As shown in the figure, two parallel deep convolution modules are streamlined into a 3×3 deep convolution module, which greatly reduces the amount of computation while maintaining the high accuracy of the detection model.

[0055] The C3k2-RC module introduces a channel-spatial attention mechanism CSA (Channel-SpatialAttention) based on the C3k2-RVB module. Figure 6(b) shows the attention mechanism. This attention mechanism is a composite attention module, which is divided into a channel attention branch and a spatial attention branch. Its structure is as follows: Figure 5 As shown in the figure. In the channel attention branch, firstly, the input feature map is global average pooled (GAP) in the spatial dimension (H and W), and a 1×1×C feature map representing the global features of each channel is output; secondly, the number of channels is mapped to C / r through linear transformation, where r is the channel scaling factor to reduce the amount of calculation and capture more compact features; then the activation function GELU is applied to introduce nonlinear characteristics and enhance the expressiveness of the target characteristics; then the number of channels is restored to C through linear transformation; finally, the weight at each channel position is obtained through the Sigmoid activation function, and the obtained weight is used to weight the feature map in the channel dimension. In the spatial attention branch, it mainly outputs a H×W×1 feature map through a 7×7 convolution module, and then the Sigmoid activation function is applied to generate weights at each spatial position, and the obtained spatial weight is used to weight the input feature map in the spatial dimension. By introducing the CSA attention mechanism, the model can more accurately enhance key target features when performing feature extraction, while suppressing irrelevant or interfering feature information such as complex background and other targets.

[0056] As a preferred solution of this embodiment, the loss function Inner-WIoUv3 in step S4 is constructed by introducing the Inner-IoU auxiliary frame on the basis of the Wise-IoUv3 loss function, specifically:

[0057] Wise-IoUv3 has a dynamic non-monotonic focusing mechanism. By introducing this focusing mechanism, the original IoU is replaced by evaluating the quality of the anchor frame based on the outlier degree. By assigning lower gradient gains to low-quality and high-quality anchor frames, the model pays more attention to the anchor frames with moderate quality. The calculation formula is as follows:

[0058] L IoU =1-IoU (1)

[0059] L WIoUv3 = rRL IoU (2)

[0060]

[0061] In the formula, IoU represents the ratio of the intersection and union of the real box and the predicted box. When the sizes of the two are close, the intersection and union ratio is approximately equal to 1, and the loss function L IoU Close to 0; L WIoUv3 is the dynamic loss function; r is the non-monotonic focus factor; β is the outlier point, which is L IoUis used to evaluate the quality of the anchor box; δ and α are hyperparameters that can be adjusted according to the specific use of the model; R is a weight coefficient used to adjust the IoU loss; x, y and x gt ,y gt Represent the center coordinates of the predicted box and the real box respectively; W g ,H g Represents the width and height of the minimum bounding box of the predicted box and the true box respectively.

[0062] Then, the auxiliary frame is introduced through Inner-IoU to assist in calculating the loss, such as Figure 8 As shown in the figure, a scale factor ratio can be used to adjust the size of the auxiliary box according to the actual needs of the model and construct the Inner-WIoUv3 loss function. The calculation formula is as follows:

[0063]

[0064] union=(w gt ×h gt )×(ratio) 2 + (w×h)×(ratio) 2 -inter (8)

[0065]

[0066] Among them, b gt and b represent the true box and predicted box respectively. Indicates the center point coordinates of the real frame and the auxiliary frame, (x c ,y c ) is represented by the center point coordinates of the prediction box and the prediction auxiliary box; w gt and h gt They represent the width and height of the real box respectively, w and h represent the width and height of the predicted box respectively; the ratio ranges from 0.5 to 1.5.

[0067] Through the above steps S2 to S4, the YOLOv11 model is optimized to obtain a target detection model. The improved YOLOv11 model (namely RSB-YOLO11 in the following text) is as follows: Figure 2 As shown, it includes a backbone part Backbone, a neck part Neck and a detection head Head.

[0068] Among them, the neck network mainly includes convolution module, aggregation module Bi_FPN, upsampling module and C3k2-RVB module, which are connected by the S-BiFPN feature fusion structure proposed based on the bidirectional feature pyramid structure combined with the P2 small target detection layer;

[0069] The backbone network includes a first convolution module, a second convolution module, a first C3k2-RC module, a third convolution module, a second C3k2-RC module, a fourth convolution module, a third C3k2-RC module, a fifth convolution module, a fourth C3k2-RC module, an SPPF module, and a C2PSA module, which are connected in sequence.

[0070] Step S5 is to input the training set of step S1 into the improved target detection model for training, and obtain the optimal UAV small target detection model. Step S6 is to input the test set of step S1 into the improved UAV small target detection model trained in step S5, and obtain the detection result of the model.

[0071] In order to more intuitively illustrate the effectiveness and feasibility of the above method proposed in the embodiment of the present invention, this embodiment uses the YOLOv11n model as the benchmark model, and the experimental environment used is the ubuntu20.04 system, equipped with an Intel(R) Xeon(R) Platinum 8352V CPU@2.10GHz processor, 120G running memory, and a graphics card of NVIDIA GeForceRTX4090GPU (24GB). The deep learning framework uses PyTorch 2.0.0, Python 3.8.20 and CUDA 12.1 environment. During the model training process, data enhancement was turned off in the last 10 rounds, and pre-trained weights were not turned on in all ablation experiments. The remaining training parameters are shown in Table 1.

[0072] Table 1. Training hyperparameter settings

[0073] parameter Value parameter Value Input size 640×640 Optimizer SGD Training rounds 350 momentum 0.937 Batch size 16 Weight decay 0.0005 Initial learning rate 0.01 Early Stop Round 50 Cosine annealing parameters 0.01 Number of threads 18

[0074] The evaluation indicators used by the model are accuracy (Precision, P), recall (Recall, R), mean average precision (meanAverage Precision, mAP), F1 score, etc. The calculation formula is as follows. Based on Map@0.5 and F1 score It is used as the main evaluation index to evaluate the detection effect.

[0075]

[0076] In order to verify the effectiveness and feasibility of the improved model in this example, ablation experiments and comparative experiments were set up to explore the impact of the improved parts of the proposed improved model on the model detection effect.

[0077] First, a comparative experiment is conducted to explore the impact of introducing the attention mechanism at different positions of the backbone on the model. Based on the improvement of the YOLOv11n network structure using the S-BiFPN feature fusion structure and the C3k2-RVB module, according to the shallow and deep layers of the backbone network, the following methods are mainly used: Method 1 does not introduce the attention mechanism to improve the backbone network; Method 2 introduces the attention mechanism to improve the shallow part of the backbone network (i.e., P2 and P3 layers); Method 3 introduces the attention mechanism to improve the deep part of the backbone network (i.e., P4 and P5 layers); Method 4 introduces the attention mechanism to improve all C3k2 modules of the backbone network. The results are shown in Table 2.

[0078] Table 2. Comparative experiments of different improved positions

[0079] method FLOPs / G Params / M Map@0.5 / % F1 v11n+S-Bifpn 10.7 2.72 93.9 90.30 +Method 1 9.8 2.28 94.1 90.64 +Method 2 10.7 2.33 94.2 90.39 +Method 3 9.8 2.29 93.5 91.04 +Method 4 10.7 2.34 94.3 91.30

[0080] From the experimental results, it can be seen that the number of parameters of the four methods has decreased significantly. Method 1 does not introduce the attention mechanism, but Map@0.5 has increased to 94.1%, showing good global and local feature extraction capabilities. Method 2 introduces the attention mechanism in the shallow layer of the backbone network, and compared with method 1, there are slight improvements in Map@0.5 and Map@0.5:0.95. Although method 3 introduces the attention mechanism in the deep layer of the backbone network, Map@0.5 drops to 93.5%, which does not achieve the expected performance improvement. Method 4 introduces the attention mechanism in both the shallow and deep layers, and has obvious improvements in Map@0.5 and Map@0.5:0.95, showing its superior global and local feature extraction capabilities. Therefore, on the whole, while maintaining a low number of parameters and calculations, method 4 has excellent feature extraction capabilities, significantly improves the accuracy of the detection model, and is the best choice for improving the C3k2-RVB module.

[0081] Based on the above experimental results, considering that different attention mechanisms have different effects on model performance, this example further designed and conducted a series of comparative experiments to systematically evaluate the performance differences and advantages of various attention mechanisms in the model, and provide a more comprehensive and scientific basis for model improvement. It mainly includes the channel-space combined attention mechanism CBAM; the space-channel coordinated attention mechanism SCSA; the mixed local attention MLCA; and the multi-scale convolutional attention MSCA. The final experimental results are shown in Table 3.

[0082] Table 3. Comparative experiments of different attention mechanisms

[0083]

[0084]

[0085] The experimental results show that, except for MLCA, the accuracy of other attention mechanisms has decreased to varying degrees, and the accuracy of attention mechanisms CBAM and MSCA has decreased significantly, with Map@0.5 decreasing by 0.8% and 0.7% respectively. Compared with other attention mechanisms, CSA performs better, with Map@0.5 reaching 94.3%, an increase of 0.2%. Therefore, after comprehensively considering various evaluation indicators, the attention mechanism CSA is the best choice for this example model.

[0086] Secondly, in order to determine the impact of the hyperparameter ratio in the loss function on the model performance, this example conducts a set of comparative tests based on this, and the results are shown in Table 4. The results show that the hyperparameter ratio increases from 0.6 to 1.1, and the ratio is equal to 0.7 and 0.8, which have an optimization effect on the model. When ratio = 0.7, the model achieves the best performance, and Map@0.5 reaches 94.8%, which is 0.5% higher than before the loss function is replaced, further improving the accuracy of the detection model. Therefore, the best choice for the hyperparameter ratio in this example is 0.7.

[0087] Table 4. Comparison experiment of ratio coefficient

[0088] Ratio Map@0.5 / % Map@0.5:0.95 / % F1 0.6 93.8 56.3 90.67 0.7 94.8 57.7 91.62 0.8 94.6 57.4 91.28 0.9 93.8 57.1 89.71 1.0 94.1 57.4 90.97 1.1 94.0 57.1 90.45

[0089] In order to further evaluate the performance of the improved loss function, this example compares the improved loss function with other loss functions (including Inner-MPDIoU, EIoU, and SIoU), such as Fig. 9 As shown in Figure 1, Figure (a) shows the comparison results of the evaluation index Map@0.5, and Figure (b) shows the comparison results of the F1 score. The results show that the Inner-WIoUv3 loss function outperforms other loss functions in both the evaluation index Map@0.5 and the F1 score, further verifying the superior performance of the improved loss function.

[0090] In this example, while keeping all environmental parameters and data sets the same, ablation experiments are performed on all improved strategies using the YOLOv11n model as the benchmark. The results are shown in Table 5.

[0091] Table 5. Ablation experiments of improved models

[0092]

[0093]

[0094] The first group of experiments A represents the original unimproved YOLOv11n model, with a Map@0.5 of 90.6. The second group of experiments B represents the introduction of the C3k2-RVB module to replace the C3k2 module in the original model, and the third group of experiments C represents the introduction of the C3k2-RC module to replace the C3k2 module in the backbone of the original model. It can be seen that when used alone, B and C have the effect of lightweighting and reducing the number of parameters for model improvement, but have no effect on improving the Map@0.5 index. However, in the fifth group of experiments E, B and C are combined at the same time, and Map@0.5 and Map@0.5-0.95 reach 91.2% and 54.3%, respectively, an increase of 0.6% and 1.0%. The fourth group of experiments D introduces S-BiFPN to improve the neck structure of the original model, so that the model can better capture the feature information of small targets, and the model performance has been greatly improved. Map@0.5 and Map@0.5-0.95 reach 93.9% and 57.2%, respectively, an increase of 3.3% and 3.9%. The sixth group of experiments, F, combined the improvements of experiment B on the basis of experiment E, and further improved the accuracy of the detection model while lightweighting the model and reducing the number of parameters. Map@0.5 reached 94.1%, an increase of 3.5%. The seventh group of experiments, G, combined the three improvements of B, C and D at the same time, and Map@0.5 and Map@0.5-0.95 increased by 3.7% and 4.4% respectively, with significant increases. It fully shows that the improved YOLOv11n model has significantly improved the detection effect of small drone targets and is suitable for low-altitude drone detection tasks. The ninth group of experiments, I, introduced the Inner-WIoUv3 loss function as a whole, and the model performance was further improved. Finally, compared with the baseline model, the improved models Map@0.5 and Map@0.5-0.95 proposed in this example increased by 4.2% and 4.4% respectively, verifying the effectiveness of each improvement point and showing that each improvement point has good compatibility.

[0095] In order to further verify the detection performance of the improved algorithm for drones under backgrounds of different complexity, this example uses graphical visualization analysis to compare the detection effects of the improved algorithm and the unimproved algorithm under small-scale targets and various complex backgrounds. Fig.10 As shown. Obviously, in Figure (b), the detection confidence score of the unimproved YOLOv11n model is 0.51, while the detection confidence score of the improved model proposed in this example is 0.72, which is an improvement of 0.21. In Figures (b) and (c), the original model has missed detection, while the detection confidence of the improved model of this example is 0.51 and 0.78 respectively, which is a significant improvement. Therefore, it is further verified that the model of this example can not only improve the detection accuracy of the model, but also improve the missed detection phenomenon.

[0096] In summary, the improved model proposed in this example has significant improvements in all aspects compared to the original model. Therefore, in the case of complex background, large target scale changes and bird interference, the method of this example has certain advantages and is very suitable for low-altitude UAV target detection tasks.

[0097] Under the same experimental environment configuration, the target detection algorithm proposed in this example is compared with other mainstream detection algorithms to further verify the effectiveness of this method. Other algorithms mainly include RT-DETR, YOLOv3, YOLOv5, YOLOv7, YOLOv8, and YOLOv10. The experimental results are shown in Table 6.

[0098] Table 6. Comparative experiments with other algorithms

[0099] Model FLOPs / G Map@0.5 / % Map@0.5-0.95 / % F1 RT-DETR 103.4 78.6 40.0 73.99 YOLOv3 12.9 88.7 49.2 87.42 YOLOv5s 15.8 92.4 55.2 90.20 YOLOv7-tiny 13.2 91.7 53.4 89.10 YOLOv8n 8.1 89.6 53.5 87.27 YOLOv8s 28.4 91.1 56.5 88.65 YOLOv10n 8.2 88.3 51.7 86.23 YOLOv10s 24.4 90.7 54.2 88.04 YOLOv11n 6.4 90.6 53.3 86.98 YOLOv11s 21.3 91.2 55.7 88.17 RSB-YOLO11 10.7 94.8 57.7 91.62

[0100] The results show that compared with other algorithms, the improved algorithm proposed in this example achieved the best detection effect: Map@0.5 and Map@0.5-0.95 reached 94.8% and 57.7% respectively, and the F1 score reached 91.62. It can be seen that the improved algorithm greatly improves the model performance without significantly increasing the complexity of the model. Compared with YOLOv8s and YOLOv11s, the computational complexity of the improved algorithm is reduced by 62.32% and 49.77% respectively, but Map@0.5 is increased by 3.7% and 3.6% respectively, which further verifies that the improved algorithm can effectively balance model performance and computational complexity, and is superior to most common algorithms and has good application value.

[0101] The experimental results show that the improved algorithm proposed in this example has a significant improvement in Map@0.5, Map@0.5-0.95 and F1 scores, and the number of model parameters has also been reduced. Through actual scene tests, the improved algorithm has a significantly better detection effect than the original model in complex backgrounds and small target scales, and has improved the problems of missed detection and false detection, indicating that the improved algorithm can effectively improve the model's detection accuracy for small drone targets and has a certain degree of robustness.

[0102] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0103] Based on the same inventive concept, the present invention also provides an edge computing terminal device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor is a central processing unit (CPU) and a graphics processing unit (GPU), which is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0104] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program executes the above method when it is executed by a processor. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. In the present invention, the computer-readable storage medium includes a hard disk, a solid-state hard disk, a cloud storage device, etc., which can be used to store the training set and computer program used for training, so that the algorithm can run smoothly and complete the detection task.

[0105] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0106] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0107] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

[0108] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive other various forms of an improved YOLOv11 drone small target detection method under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.

Claims

1. A method for detecting small drone targets by improving YOLOv11, characterized by: Based on YOLOv11, a bidirectional feature pyramid structure BiFPN is introduced, and combined with the P2 small target detection layer, the feature information flows from the P5 layer of the network to the P2 layer, and then from the P2 layer to the P5 layer, to build a bidirectional information flow and form a feature fusion structure S-BiFPN; in the neck network, the RepVitBlock module is introduced on the basis of the C3k2 module to replace the original Bottleneck module to form a C3k2-RVB module; in the backbone network, the channel-spatial attention mechanism CSA is introduced on the basis of the C3k2-RVB module to form a C3k2-RC module to replace the C3k2 module; in the detection head, the loss function Inner-WIoUv3 constructed by introducing the Inner-IoU auxiliary box on the basis of the Wise-IoUv3 loss function is adopted; thereby improving the target detection model and training it; the optimal UAV small target detection model obtained by training is used to realize UAV small target detection.

2. According to claim 1, a method for detecting small targets of unmanned aerial vehicles by improving YOLOv11 is characterized in that: The modules of the neck network are connected through a feature fusion structure S-BiFPN.

3. The method for detecting small targets of unmanned aerial vehicles by improving YOLOv11 according to claim 1, characterized in that: In the RepVitBlock module, the input feature map first passes through a 3×3 deep convolution module DWConv to extract local spatial features, then passes through a 1×1 deep convolution module to adjust the channels of the features, and then passes through two 1×1 convolution modules Conv in sequence for feature fusion; the input feature map is also added to the processed features through jump connections to form a residual structure, and finally the refined and fused feature map is output for use by subsequent modules; in the inference stage, the two parallel deep convolution modules are simplified to a 3×3 deep convolution module using the structural reparameterization technology.

4. The method for detecting small targets of unmanned aerial vehicles by improving YOLOv11 according to claim 1, characterized in that: The feature fusion structure S-BiFPN adopts a weighted feature fusion strategy in the process of feature fusion: different weights are assigned when fusing feature maps of different scales. The weights are learnable parameters and are continuously adjusted and optimized during the model training process to achieve the best fusion effect.

5. The method for detecting small targets of unmanned aerial vehicles by improving YOLOv11 according to claim 1, characterized in that: The channel-spatial attention mechanism CSA is a composite attention module, which is divided into a channel attention branch and a spatial attention branch: in the channel attention branch, the input feature map is first globally averaged pooled in the spatial dimensions H and W, and a 1×1×C feature map representing the global features of each channel is output; then the number of channels is mapped to C / r through linear transformation, where r is the channel scaling factor; then the activation function GELU is applied to introduce nonlinear characteristics; after linear transformation, the number of channels is restored to C; finally, the weight at each channel position is obtained through the Sigmoid activation function, and the obtained weight is used to weight the feature map in the channel dimension; in the spatial attention branch, a H×W×1 feature map is output through a 7×7 convolution module, and then the Sigmoid activation function is applied to generate weights at each spatial position, and the input feature map is weighted in the spatial dimension using the obtained spatial weight.

6. The method for detecting small targets of unmanned aerial vehicles by improving YOLOv11 according to claim 1, characterized in that: The construction process of the loss function Inner-WIoUv3 is specifically as follows: The calculation formula of Wise-IoUv3 is as follows: L IoU =1-IoU (1) L WIoUv3 =rRL IoU (2) In the formula, IoU represents the ratio of the intersection and union of the real box and the predicted box, and the loss function L IoU is the corresponding loss function; L WIoUv3 is the dynamic loss function; r is the non-monotonic focus factor; β is the outlier point, which is L IoU is the mean of the anchor boxes, used to evaluate the quality of the anchor boxes; δ and α are adjustable hyperparameters; R is the weight coefficient used to adjust the IoU loss; x, y and x gt ,y gt Represent the center coordinates of the predicted box and the real box respectively; W g ,H g Represents the width and height of the minimum bounding box of the predicted box and the real box respectively; The auxiliary box is introduced through Inner-IoU to assist in calculating the loss. The size of the auxiliary box is adjusted according to the actual needs of the model through a scale factor ratio to construct the Inner-WIoUv3 loss function; the calculation formula is as follows: L innerIoU =1-IoU inner (10) Among them, b gt and b represent the true box and predicted box respectively. Indicates the center point coordinates of the real frame and the auxiliary frame, (x c ,y c ) is represented by the center point coordinates of the prediction box and the prediction auxiliary box; w gt and h gt They represent the width and height of the real box respectively, w and h represent the width and height of the predicted box respectively; the ratio ranges from 0.5 to 1.

5.

7. An edge computing terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of an improved YOLOv11 drone small target detection method as described in any one of claims 1 to 6 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of an improved YOLOv11 method for detecting small targets in unmanned aerial vehicles are implemented as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Photovoltaic panel dust retention detection method and device, electronic equipment and storage medium

    CN120298402A

  • Improved YOLO11-based water hyacinth target rapid detection method

    CN120635395A

  • Urban low-altitude weak unmanned aerial vehicle identification method based on self-attention mechanism

    CN121582552A

  • Vehicle target detection method and system based on improved YOLOV12N

    CN122473438A

  • Urban road vehicle rapid detection method based on unmanned aerial vehicle application

    CN122493408A