Target detection method for multi-rotor unmanned aerial vehicle
By improving the YOLO11 network structure and combining the RVB-CSP module and Wise-IoU loss function, the accuracy and efficiency problems in multi-rotor UAV detection were solved, achieving efficient and accurate UAV target detection.
Patent Information
- Application Number
- CN202510808874.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-11-07
AI Technical Summary
Existing UAV target detection technologies suffer from insufficient detection box positioning accuracy, poor small target detection performance, and low detection efficiency when detecting multi-rotor UAVs, making it difficult to meet the high-precision requirements in complex environments.
An improved method based on the YOLO11 network is adopted. By introducing the RVB-CSP module, MAFPN structure and Wise-IoU loss function, feature processing and feature fusion are optimized to improve detection accuracy and efficiency. The model is deployed on the Nvidia Jetson Xavier NX computing board for real-time detection.
It improves the accuracy and real-time performance of multi-rotor UAV target detection, reduces computational overhead and parameter count, enhances the ability to detect small targets, and improves detection efficiency and accuracy.
Smart Images

Figure CN120912848A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of target detection in computer vision and deep learning, and particularly relates to a multi-rotor unmanned aerial vehicle target detection method based on an improved YOLO11 network model. BACKGROUND
[0002] In recent years, multi-rotor unmanned aerial vehicles occupy more than 50% of the unmanned aerial vehicle market due to their simple structure, affordable price and vertical take-off and landing function. However, as the application scenarios of multi-rotor unmanned aerial vehicles continue to expand, the existing unmanned aerial vehicle safety supervision system has been difficult to meet the complex management needs, and needs to be upgraded systematically. As the first step of the supervision process, the performance and effect of the unmanned aerial vehicle target detection technology directly determine the execution effect of subsequent tracking, countermeasures and other tasks.
[0003] Compared with traditional radar detection and radio sensing technology, the deployment and maintenance cost of the unmanned aerial vehicle detection technology based on vision is relatively high, and it needs to rely on complex hardware devices and professional operation, and is also susceptible to electromagnetic interference in complex scenes such as shielding. In comparison, the unmanned aerial vehicle detection technology based on vision has a high cost performance, and only needs ordinary cameras and other low-cost imaging devices to detect various types of multi-rotor unmanned aerial vehicles, and shows flexible adaptability in complex environments such as urban low altitude.
[0004] Although the current target detection technology based on deep learning has achieved many outstanding results in general scenarios such as pedestrians and vehicles, the research on multi-rotor unmanned aerial vehicle targets is still relatively rare. At the same time, since the height and attitude of the multi-rotor unmanned aerial vehicle change constantly when it is flying in the air, the appearance features of the target also change obviously, which makes it difficult for general detection models to achieve high-precision detection, and further optimization needs to be combined with the characteristics of the unmanned aerial vehicle. SUMMARY
[0005] In order to solve the problems of insufficient detection frame positioning accuracy, poor small target detection effect and low detection efficiency of the traditional detection model in detecting multi-rotor unmanned aerial vehicles, the application discloses a multi-rotor unmanned aerial vehicle target detection network structure and detection method based on an improved YOLO11 network, and deploys the trained model on an Nvidia Jetson Xavier NX computing board to detect the multi-rotor unmanned aerial vehicle target in the real-time image of the camera.
[0006] The overall implementation process of the application is as shown in the figure. Figure 1 A method for multi-rotor unmanned aerial vehicle target detection, specifically comprising the following steps:
[0007] Step 1: establishing a detection network;
[0008] The detection network comprises a backbone network, a neck network and a head network.
[0009] The backbone network comprises in sequence: a first CBS module, a second CBS module, a first RVB-CSP module, a third CBS module, a second RVB-CS module, a fourth CBS module, a third RVB-CS module, a fifth CBS module, a fourth RVB-CSP module, an SPPF module, and a C2PSA module, wherein outputs of the first RVB-CSP module, the second RVB-CSP module, the third RVB-CSP module, and the C2PSA module simultaneously serve as a first input, a second input, a third input, and a fourth input of the neck network;
[0010] The first input of the neck network is connected to a sixth CBS module, and an input of the sixth CBS module serves as an input of a first Fusion module; the second input of the neck network is connected in sequence to a seventh CBS module, the first Fusion module, and a fifth RVB-CSP module, an output of the fifth RVB-CSP module is divided into two paths, one path is connected to a tenth CBS module, and the other path is connected in sequence to a second Fusion module, a sixth RVB-CSP module, and an eleventh CBS module, and an output of the sixth CBS module simultaneously serves as a first output of the neck network; an output of the seventh CBS module simultaneously serves as an input of an eighth CBS module, and an input of the eighth CBS module serves as an input of a third Fusion module; the third input of the neck network is connected in sequence to a ninth CBS module, the third Fusion module, and a seventh RVB-CSP module; an output of the seventh RVB-CSP module is divided into three paths, a first path is connected to a first upsample module, an output of the first upsample module also serves as an input of the first Fusion module and the second Fusion module, a second path and outputs of the tenth CBS module and the eleventh CBS module are jointly input into a fourth Fusion module, and a third path is connected to a fourteenth CBS module; an output of the fourth Fusion module is connected in sequence to an eighth RVB-CSP module and a fifteenth CBS module; an output of the ninth CBS module is also input into a fifth Fusion module after passing through a twelfth CBS module, an output of the eighth RVB-CSP module simultaneously serves as a second output of the neck network; the fourth input of the neck network is connected in sequence to a thirteenth CBS module, the fifth Fusion module, and a ninth RVB-CSP module, an output of the ninth RVB-CSP module includes two paths, a first path is input into a second upsample module, the second upsample module includes two outputs, and the two outputs are input into the third Fusion module and the fourth Fusion module, respectively, a second path and outputs of the fourteenth CBS module and the fifteenth CBS module are jointly input into a sixth Fusion module, and the sixth Fusion module is input into a tenth RVB-CSP module; an output of the tenth RVB-CSP module serves as a third output of the neck network;
[0011] The head network comprises a first Detect module, a second Detect module, and a third Detect module, and the inputs of the three modules correspond to the first output, the second output, and the third output of the neck network in sequence.
[0012] The CBS module is a basic convolution unit, which is composed of a standard convolution layer, a batch normalization, and a SiLU activation function connected in sequence, and is used for extracting local features of an image.
[0013] The RVB-CSP module is a feature processing module.
[0014] The Fusion module is a multi-scale feature fusion module, which is used for fusing the multi-path feature branches of the MAFPN neck network.
[0015] The upsample module is an up-sampling module.
[0016] The SPPF module is a fast spatial pyramid pooling module.
[0017] The C2PSA module is a convolution block with an attention mechanism.
[0018] The Detect module is an output layer, which is used for converting a feature map into a final detection result.
[0019] Step 2: using the DUT Anti-UAV public dataset, training the detection network in step 1;
[0020] Step 3: deploying the trained detection network.
[0021] Step 4: using the deployed detection network to detect images.
[0022] Further, the loss function for training the detection network in step 2 comprises an IoU loss Wise-IoU v3 loss
[0023] The IoU loss is:
[0024]
[0025] wherein, and b l ,b r ,b t ,b b are the coordinates of the left, right, top and bottom boundaries of the auxiliary real box and the auxiliary predicted box respectively, and (x c ,y c ) are the center point coordinates of the real box and the predicted box respectively, wgt , h gt and w, h are the width and height of the real and predicted boxes respectively; ratio is the scaling factor, when ratio < 1, the smaller scale auxiliary box is obtained; when ratio > 1, the larger scale auxiliary box is obtained; inter represents the intersection area of the auxiliary real box and the auxiliary predicted box, d represents the center point distance of the real box and the predicted box, and c represents the diagonal distance of the minimum bounding box of the real box and the predicted box;
[0026] The Wise-IoU v3 loss is
[0027]
[0028] wherein, is the distance attention term, represents the IoU loss, represents the basic version v1 of Wise-IoU, W represents the width of the minimum bounding box of the real box and the predicted box, H represents the height of the minimum bounding box of the real box and the predicted box, β represents the abnormality degree, and r represents the non-monotonic focusing coefficient; and α and δ are hyperparameters;
[0029] The final loss is:
[0030]
[0031] The RVB-CSP module is adopted as the feature processing module, so that the appearance features of the unmanned aerial vehicle can be more efficiently captured; the Fusion module is adopted as the multi-scale feature fusion module, the multi-path feature branches of the MAFPN neck network are fused, so that the fused features contain rich semantic information and the model has a smaller parameter amount; the upsample module is adopted to restore or promote the low-resolution feature map to a higher resolution; the loss function of the Inner-WIoU is adopted to improve the positioning accuracy of the target through a dynamic focusing mechanism and to accelerate the convergence speed; the SPPF module is adopted to improve the efficiency and speed of multi-scale feature extraction while maintaining the detection accuracy; the C2PSA module combines the CSPNet structure and the pyramid compression attention mechanism, enhances the feature extraction capability and the attention to key areas, and thus improves the detection accuracy and real-time performance of the quad-rotor unmanned aerial vehicle. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is the overall implementation process of the present application.
[0033] Figure 2 is the network structure of the improved YOLO11 model.
[0034] Figure 3 For C3K2 module structure.
[0035] Figure 4 For MobileNet Block and RepViT Block comparison.
[0036] Figure 5 For RVB_CSP module structure.
[0037] Figure 6 For structure re-parameterization.
[0038] Figure 7 For YOLO11 Neck structure.
[0039] Figure 8 For MAFPN network structure.
[0040] Figure 9 For Wise-IoU distance attention mechanism schematic diagram.
[0041] Figure 10 For improved loss function calculation flowchart.
[0042] Figure 11 For different ablation experiment setting model in the training process loss change comparison chart.
[0043] Figure 12 For original YOLO11 model and improved model in the target positioning problem on the heat map comparison; wherein, (a) is the test sample; (b) is the original model heat map; (c) is the improved model heat map.
[0044] Figure 13 For original YOLO11 model and improved model in the wrong detection, the heat map comparison of the missed detection problem. Wherein, (a) is the test sample; (b) is the original model heat map; (c) is the improved model heat map.
[0045] Figure 14 For the hardware deployment device of the present application.
[0046] Figure 15 For detection result display. DETAILED DESCRIPTION
[0047] The implementation process of the present application is as shown in Figure 1 The following will be described in conjunction with the schematic diagram and implementation steps
[0048] S1. Model optimization:
[0049] The application optimizes the feature processing module, feature fusion structure and loss function of the YOLO11 network model, so that the model can more accurately identify and locate the multi-rotor unmanned aerial vehicle with lower computational overhead. Figure 2 The steps can be divided into three parts.
[0050] S11. Optimize the C3k2 feature processing module
[0051] The C3K2 module is the main module responsible for feature processing tasks in the Backbone structure, and its structure is shown in Figure 3 The ordinary C3K and Bottleneck model used in it still has the problem of limited information flow when processing features, which will affect the fusion effect of deep features.
[0052] Therefore, the application introduces RepViT Block and optimizes it on this basis. The structure of RepViT Block is shown in Figure 4 It moves the DWConv and Squeeze and Excitation (SE) attention layer in MobileNetV3 Block, and sets optional operations for SE, which realizes the separation of spatial dimension information fusion task (tokenmixer) and channel information interaction task (channel mixer), so that the model can process information of different dimensions at a lower computational cost.
[0053] Then, the application further improves RepViT Block by referring to the idea of CSPNet structure, and names it RVB_CSP module, and restructures C3K2 module based on it. Its structure is shown in Figure 5 First, replace the standard Bottleneck block in C3K with RepViT Block. Then change the configuration parameter c3k to rvb. When rvb=True, use the RVB block based on the CSPNet architecture; otherwise, use the standard RepViT Block directly. This optimization not only reduces the computational amount and parameter amount of the model, but also enables the model to better process spatial and channel dimension information, thereby more efficiently capturing the features of the multi-rotor unmanned aerial vehicle target.
[0054] Finally, the application also introduces structure reparameterization technology to reduce the computational complexity and memory occupation of the model in the inference stage, and its schematic diagram is shown in Figure 6As shown, during the training phase, the model constructs multiple DWConv branches to extract features of different scales and types, enhancing its learning ability for different types of drones. During the inference phase, the inference code merges the convolutional layers and batch normalization layers from multiple branches, transforming them into a single-branch structure. This allows the model to better balance detection accuracy and speed during inference.
[0055] S12. Optimize the Neck structure
[0056] Neck is a core component in YOLOv11 responsible for multi-scale feature fusion, employing a Path Aggregation Network (PAN). Its structure is as follows: Figure 7 As shown, it fuses the high-level and low-level features extracted by Backbone through a bidirectional path from top to bottom and bottom to top, so that each layer of features can contain more information. However, PAN only splices two sets of features, so the degree of fusion is relatively limited, making it difficult to effectively detect small targets. Furthermore, too many feature splices will lead to a large amount of computation in subsequent convolution operations and reduce inference efficiency.
[0057] To address this, the present invention introduces the MAFPN structure and makes lightweight improvements based on it. The network structure of MAFPN is as follows: Figure 8 As shown, it designs Superficial Assisted Fusion (SAF) and Advanced Assisted Fusion (AAF) modules. SAF combines the features output by the backbone network (P2-P5) with the shallow features of the neck network, and preserves the detailed information of small targets through cross-layer connections. AAF, on the other hand, transmits gradient information in the deep neck layer through multi-directional connections. In this way, the input of the feature concatenation module of each layer not only contains the information of the same layer, but also adds the features of both high and low layers. At the same time, it can also utilize the information of the P2 layer without adding an additional fusion path to the P2 layer, thus avoiding a significant increase in the number of parameters and computation.
[0058] Building upon this, this invention further optimizes the design by drawing on the lightweight principles of BIFPN, replacing the original feature concatenation operation in MAFPN with a feature-weighted fusion strategy. Taking the P3 node as an example, the fusion formula is as follows:
[0059]
[0060] Where ε is the local minimum, used to prevent the denominator from being zero; w i These are the weights corresponding to the features. They are learnable and need to be non-negative through ReLU activation and restricted to a reasonable range through normalization to maintain training stability.
[0061] After introducing the weighted fusion, the application also adds a 1*1 convolution module after the P3, P4 and P5 layers to uniformly adjust the number of channels, ensure that the number of channels remains consistent each time the fusion is performed, and name the new Neck structure as LMSF_Neck, which has the structure as shown in the middle part. Figure 2
[0062] S13. Optimizing the loss function
[0063] The bounding box regression loss can optimize the positioning ability of the model to the target in many general scenarios, but in the case of small target detection and complex background, the ordinary IoU loss function may be difficult to cover various actual situations, and it is difficult to accurately regress the target boundary.
[0064] To this end, the application introduces Inner-IoU and Wise-IoU loss functions and fuses them.
[0065] The Inner-IoU loss considers that smaller bounding boxes can increase the gradient absolute value of high IoU samples to speed up the bounding box regression, and for low IoU samples, larger bounding boxes are needed. Therefore, it uses an auxiliary bounding box to optimize the IoU loss term itself, and the formula is as follows:
[0066]
[0067] wherein, and b l ,b r ,b t ,b b are the coordinates of the left, right, top and bottom boundaries of the real box and the predicted box (auxiliary box), and (x c ,y c ) are the center point coordinates of the real box and the predicted box, w gt ,h gt and w, h are the width and height of the real box and the predicted box; ratio is a scaling factor, when ratio < 1, a smaller scale auxiliary box can be obtained; when ratio > 1, a larger scale auxiliary box can be obtained.
[0068] The traditional IoU loss function usually does not consider the quality of the data, but in the actual training process, low-quality samples (such as samples with inaccurate labeling or blurred targets) will be excessively punished, which will damage the positioning performance of the model. To this end, the Wise-IoU loss evaluates the quality of the anchor box through a dynamic non-monotonic focusing mechanism. As shown in Figure 9 As shown, when the overlap between the ground truth bounding box (gt) and the predicted bounding box is high, Wise-IoU v1 uses a distance attention mechanism to reduce the loss function's focus on center point distance and geometry, thereby reducing factors hindering convergence. Building on this, Wise-IoU v3 introduces the anomaly factor β of the anchor box to evaluate its quality and constructs a non-monotonic focusing coefficient r. When the anomaly factor β is large, a smaller gradient gain can be allocated, allowing the model to mitigate the negative impact of low-quality samples and focus more on optimizing samples of average quality. The formula is as follows:
[0069]
[0070] in, It is a distance attention term; This is the standard IoU loss, used to control whether the loss value is increased or decreased; It is the exponential moving average of IoU loss; α and δ are hyperparameters.
[0071] This invention integrates Inner-IoU and Wise-IoU v3 to optimize the original bounding box regression loss function. The specific implementation process is as follows: Figure 10 As shown. During the initialization phase, a global variable is created. The mean loss of the focusing mechanism is recorded. Then, in the forward propagation stage, various intermediate variables of the ground truth bounding box and the predicted bounding box (such as center point coordinates, width and height, overlapping and union area, etc.) are calculated to obtain the Inner-IoU loss. And use it as a reference value to update the global loss. Then, the non-monotonic focusing mechanism stage begins, where the loss of the current sample is calculated first. With global loss The ratio is used to measure the sample anomaly β, and then the non-monotonic focusing coefficient r is calculated from it; finally, Substituting β and r into the Wise-IoU v3 loss formula, we obtain the final loss value.
[0072] By using the Inner-WIoU composite loss, the model can dynamically allocate gradients according to the degree of anomaly of the samples, reducing the adverse effects of low-quality samples and improving the convergence speed.
[0073] S2. Model Training:
[0074] The DUT Anti-UAV public dataset is selected for training and verification in the experiment, in which the proportion of small targets is more than 70%, the average area ratio of targets to images in the dataset is only 0.013, and the minimum area ratio is as low as 1.9e-06. The evaluation indexes of the experiment are selected as the five COCO indexes of AP50, AP50:95, AP50:95[small], parameter quantity (Params) and calculation quantity (GFLOPs).
[0075] Firstly, four ablation experiments are set, and their performances in the training process are as shown in Figure 11 It can be seen that, compared with the baseline model YOLO11, the loss of the model after adding the RVB_CSP module decreases obviously faster, and a lower loss can be reached in the early training; further combining the RVB_CSP+LMSF_Neck model, the loss curve is smoother, and better stability is shown in the middle training period, and the loss value is lower than the previous two; the final experimental model after adding the Inner-WIou loss function has a loss value obviously lower than the other three in the late training period.
[0076] The performance results of the four ablation experiments on the test set are shown in the following table:
[0077] Table 1 Ablation experiment results of the original model and the improved model on the DUT Anti-UAV test set
[0078]
[0079] From the AP50 index, with the gradual addition of improved modules, AP50 gradually increased from 91.5% of the YOLO11 baseline model to 94.3%, indicating that the optimized model has improved the detection accuracy of common targets. The improvement of the AP50:95 index is more significant. The final improved model reaches 68.7% on this index, which is 5.2% higher than the 63.5% of the YOLO11 baseline model, indicating that the improved model not only has a greater improvement in the average detection accuracy under different IoU thresholds, but also performs better in the detection scene with higher accuracy requirements. The detection accuracy of small targets AP50:95[small] is also improved by 2.0% compared with the baseline model, especially after the introduction of RVB_CSP and LMSF_Neck, the model's ability to detect small targets has improved most significantly. In terms of parameter quantity (Params), the parameter quantity of the model gradually decreases from 9.413M to 6.814M, indicating that the improved network structure can effectively reduce the parameter quantity and reduce the storage requirements of the model. The calculation amount (GFLOPs) is also optimized. Although additional modules are added, the calculation amount is still lower than that of the baseline model, and the improved model remains at 19.6 GFLOPs, indicating that the improved structure effectively controls the calculation amount and can ensure the efficiency during running.
[0080] The heat map comparison of the original model and the final improved model on part of the data set is shown in Figure 11 and 12 .
[0081] Figure 12 (b) shows a case where the target positioning of the YOLO11 original model is inaccurate. As can be seen from the figure, the high attention area identified by the original model deviates from the actual position of the unmanned aerial vehicle, resulting in inaccurate detection frame position and size, which affects the overall detection effect and even adversely affects the tracking accuracy of the subsequent target tracking model.
[0082] Figure 13 (b) shows a case where the YOLO11 original model misses detection and mis-detection. As can be seen from the figure, in some areas where there are actually multi-rotor unmanned aerial vehicles, the color of the heat map is blue, indicating that the original model failed to identify the targets in these areas, resulting in missed detection; while in some areas where there are no multi-rotor unmanned aerial vehicle targets, there are red high attention areas, indicating that the original model has mis-detection.
[0083] Figure 12 As can be seen from (c) and 13(c), the improved model effectively solves the above problems. For areas where multi-rotor UAV targets actually exist, the heatmap shows a distinct red color, indicating that the improved model can accurately focus on the target and reduce the problem of missed detections. At the same time, in areas where multi-rotor UAVs do not exist, the heatmap color is significantly darker, indicating that the false detection situation has also been improved. For the localization problem, in the heatmap of the improved model, the high-attention areas closely match the actual location of the UAV, and the accuracy of localization is greatly improved.
[0084] Finally, you will obtain the pt (PyTorch) file after training is complete.
[0085] S3. Model Deployment:
[0086] S31. Install the appropriate CUDA, TensorRT packages, Python environment, and related dependency libraries such as PyTorch on the Nvidia Jetson Xavier NX computing board.
[0087] S32. Use the torch.onnx.export() function to export the pt model as ONNX format.
[0088] S33. Using TensorRT's Python API, load the ONNX model generated in the previous step, generate the TensorRTEngine model, and save the generated model to the specified directory of Jetson Xavier NX.
[0089] S34. Write inference code to load and run the TensorRT Engine model, providing an external API. Obtain input data, pass it to the model for inference, and output the inference results.
[0090] S4. Hardware Device Connection and Data Processing:
[0091] S41. For example Figure 14 The diagram shows the hardware deployment device of this invention, which mainly includes an Nvidia Jetson Xavier NX computing board, a camera, and some display devices.
[0092] S42. Connect the camera to the computing board, obtain the real-time image of the camera through the SDK, and then call the model call interface in step S34 to pass the image to the model for processing, and obtain information such as the bounding box and confidence level of the multi-rotor UAV target in the image.
[0093] S5. Results Display:
[0094] Based on the output, the target is marked using OpenCV and displayed through a visualization interface, such as...Figure 14 as shown.
Claims
1. A method for multi-rotor unmanned aerial vehicle target detection, specifically comprising the following steps: Step 1: Establish a detection network; The detection network includes: backbone network, neck network, head network; The backbone network includes in turn: first CBS module, second CBS module, first RVB-CSP module, third CBS module, second RVB-CS module, fourth CBS module, third RVB-CS module, fifth CBS module, fourth RVB-CSP module, SPPF module, C2PSA module, wherein the outputs of the first RVB-CSP module, the second RVB-CSP module, the third RVB-CSP module and the C2PSA module are simultaneously the first input, the second input, the third input and the fourth input of the neck network; The first input of the neck network is connected to the sixth CBS module, and the input of the sixth CBS module is the input of the first Fusion module; the second input of the neck network is connected in turn to the seventh CBS module, the first Fusion module, the fifth RVB-CSP module, the output of the fifth RVB-CSP module is divided into two paths, one path is connected to the tenth CBS module, and the other path is connected in turn to the second Fusion module, the sixth RVB-CSP module and the eleventh CBS module, the output of the sixth CBS module is simultaneously the first output of the neck network; the output of the seventh CBS module is simultaneously input to the eighth CBS module, and the input of the eighth CBS module is the input of the third Fusion module; the third input of the neck network is connected in turn to the ninth CBS module, the third Fusion module and the seventh RVB-CSP module; the output of the seventh RVB-CSP module is divided into three paths, the first path is connected to the first upsample module, the output of the first upsample module is also input to the first Fusion module and the second Fusion module respectively, the second path and the outputs of the tenth CBS module and the eleventh CBS module are jointly input to the fourth Fusion module, and the third path is connected to the fourteenth CBS module; the output of the fourth Fusion module is connected in turn to the eighth RVB-CSP module and the fifteenth CBS module; the output of the ninth CBS module is also input to the fifth Fusion module after passing through the twelfth CBS module, and the output of the eighth RVB-CSP module is simultaneously the second output of the neck network; the fourth input of the neck network is connected in turn to the thirteenth CBS module, the fifth Fusion module and the ninth RVB-CSP module, the output of the ninth RVB-CSP module includes two paths, the first path is input to the second upsample module, the second upsample module includes two outputs which are input to the third Fusion module and the fourth Fusion module respectively, the second path is input to the sixth Fusion module together with the outputs of the fourteenth CBS module and the fifteenth CBS module, and the sixth Fusion module is input to the tenth RVB-CSP module; the output of the tenth RVB-CSP module is the third output of the neck network; The head network comprises a first Detect module, a second Detect module and a third Detect module, and the inputs of the three modules correspond to the first output, the second output and the third output of the neck network in sequence; The CBS module is a basic convolution unit, which is composed of a standard convolution layer, a batch normalization and a SiLU activation function connected in sequence, and is used for extracting local features of an image; The RVB-CSP module is a feature processing module; The Fusion module is a multi-scale feature fusion module, which is used for fusing the multi-path feature branches of the MAFPN neck network; The upsample module is an up-sampling module; The SPPF module is a fast spatial pyramid pooling module; The C2PSA module is a convolution block with an attention mechanism; The Detect module is an output layer, which is used for converting a feature map into a final detection result; Step 2: using the DUT Anti-UAV public dataset, training the detection network in step 1; Step 3: deploying the trained detection network; Step 4: using the deployed detection network to detect an image.
2. The method for multi-copter drone target detection of claim 1, wherein, The loss function for training the detection network in step 2 comprises: IoU loss Wise-IoU v3 loss 3. The method for multi-copter drone target detection of claim 2, wherein, The IoU loss is: wherein, and b l , r , t , b are the coordinates of the left, right, top and bottom boundaries of the auxiliary real box and the auxiliary predicted box respectively, and (x c , y c ) are the center point coordinates of the real box and the predicted box respectively, w gt , h gt and w, h are the width and height of the real box and the predicted box respectively; ratio is a scaling factor, when ratio < 1, a smaller scale auxiliary box is obtained; when ratio > 1, a larger scale auxiliary box is obtained; inter represents the intersection area of the auxiliary real box and the auxiliary predicted box, d represents the center point distance of the real box and the predicted box, and c represents the diagonal distance of the minimum bounding box of the real box and the predicted box.
4. The method for multi-copter drone target detection of claim 3, wherein, The Wise-IoU v3 loss is: wherein, is the distance attention term, denotes the IoU loss, denotes the base version v1 of Wise-IoU, W denotes the width of the minimum bounding box of the real box and the predicted box, H denotes the height of the minimum bounding box of the real box and the predicted box, β denotes the abnormality degree, r denotes the non-monotonic focusing coefficient; α, δ are hyperparameters.
5. The method for multi-copter drone target detection of claim 4, wherein, Final loss Is: wherein