Real-time unmanned aerial vehicle detection method based on four-head detection network
Through the improved YOLOv8 model and four-scale fusion structure, combined with the adaptive spatial feature fusion mechanism, the problem of small target features being masked in complex backgrounds is solved, the accuracy and robustness of drone detection are improved, and the drone detection is adapted to drone detection in complex environments.
Patent Information
- Application Number
- CN202510755026.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
AI Technical Summary
In scenarios with complex backgrounds and large changes in target sizes, existing drone detection methods are difficult to effectively retain small target features, resulting in reduced detection accuracy and high missed detection rates. Traditional multi-scale feature fusion methods often lead to dilution of small target features.
The real-time drone detection method based on the four-head detection network is adopted. Through the improved YOLOv8 model, the four-scale fusion structure and the adaptive spatial feature fusion mechanism are combined to enhance the sensitivity of small target features, avoiding small target features being masked by large-scale features during the scale fusion process, and the detection accuracy is improved through the optimized loss function.
It improves the accuracy and robustness of small-object detection, reduces the missed detection rate, adapts to drone detection in complex environments, and provides a more accurate and efficient solution.
Smart Images

Figure CN120259990A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV monitoring, and particularly to a real-time UAV detection method based on a four-head detection network. Background Art
[0002] Unmanned aerial vehicles (UAVs) have been widely used in recent years in fields such as surveillance, agriculture, and logistics, significantly improving the development efficiency of these industries. However, the abuse of UAVs in illegal activities also poses a major threat to public safety and privacy. To address these threats, there is an urgent need to establish an effective UAV surveillance network. Current monitoring technologies include radar, audio, radio, and visual surveillance, etc. Due to the high cost and complex installation requirements of radar, audio, and radio technologies, their deployment faces huge challenges. In contrast, visual target detection has become one of the most promising and feasible UAV surveillance methods due to its low cost and wide application range.
[0003] Using deep learning technology for UAV detection is crucial for ensuring airspace safety and improving monitoring efficiency. In actual UAV detection scenarios, to improve detection accuracy and reduce the missed detection rate, the detection algorithm has high requirements for detection accuracy, detection speed, model complexity, and generalization. Therefore, to meet the above requirements, there is an urgent need for an improved UAV detection method. Summary of the Invention
[0004] To solve the technical problem of how to provide an improved UAV detection method that can improve detection accuracy and reduce the missed detection rate in actual UAV detection scenarios, the present invention provides a real-time UAV detection method based on a four-head detection network. By inputting the real-time acquired image data into the improved UAV detection model, the sensitivity of small target features is enhanced, and the small target features are prevented from being masked by large-scale features during the scale fusion process, thereby achieving the technical effect of improving the accuracy of small target detection.
[0005] To solve the above technical problem, the following technical solutions are now proposed: A real-time UAV detection method based on a four-head detection network, comprising the following steps: Acquire the image data of the target area in real time; Input the image data into the trained UAV detection model for UAV target detection; Wherein, the UAV detection model includes a backbone network for feature extraction, a neck network with a multi-scale fusion structure being a four-scale fusion structure including four different resolution feature layers, and a head network with four-scale detection heads.
[0006] In this embodiment, the training method of the UAV detection model includes: Obtain a UAV dataset and divide it into a training set and a test set after preprocessing; Build a UAV detection model and use the training set and the test set for the UAV detection model.
[0007] In this embodiment, inputting the image data into the trained UAV detection model for UAV target detection includes: The backbone network extracts features from the image data to obtain four feature maps of different scales; The neck network performs feature fusion on the four feature maps of different scales through the four-scale fusion structure to obtain four fusion features of different scales; The head network processes the four fusion features of different scales output by the neck network to obtain the target detection result.
[0008] In this embodiment, the backbone network extracts features from the image data to obtain four feature maps of different scales, including: Preprocess the image data to adjust the size and pixel value of the image data; Input the preprocessed image data into the initial convolutional layer to obtain the feature map C1; Extract features from the first feature map through a four-stage feature extraction module to obtain the feature maps C2, C3, C4, and C5. The feature map scales of the feature maps C2, C3, C4, and C5 decrease in sequence; Among them, the four-stage feature extraction module includes 4 downsampling modules for halving the feature map size and doubling the number of channels, and 4 feature extraction modules for extracting features in the feature map. The first feature map passes through the downsampling module and the feature extraction module in sequence.
[0009] In this embodiment, the neck network performs feature fusion on the four feature maps of different scales through the four-scale fusion structure to obtain four fusion features of different scales, including: Obtain the feature map C5, and adjust the scale of the feature map C5 through an upsampling module so that the scale of the adjusted feature map C5 is aligned with the scale of the feature map C4; Concatenate the adjusted feature map C5 with the feature map C4 through a concat block and generate the feature map M2 through the feature extraction module; Obtain the feature map M2, and adjust the scale of the feature map M2 through an upsampling module so that the scale of the adjusted feature map M2 is aligned with the scale of the feature map C3; The adjusted feature map M2 is concatenated with the feature map C3 through a concat block, and the feature map M3 is generated through the feature extraction module; The feature map M3 is obtained, and the scale of the feature map M3 is adjusted through an upsampling module so that the scale of the adjusted feature map M3 is aligned with the scale of the feature map C2; The adjusted feature map M3 is concatenated with the feature map C2 through a concat block, and the fused feature P2 is generated through the feature extraction module; The fused feature P2 is obtained, and the scale of the fused feature P2 is adjusted through a downsampling module so that the scale of the adjusted fused feature P2 is aligned with the scale of the feature map M3; The adjusted fused feature P2 is concatenated with the feature map M3 through a concat block, and the fused feature P3 is generated through the feature extraction module; The fused feature P3 is obtained, and the scale of the fused feature P3 is adjusted through a downsampling module so that the scale of the adjusted fused feature P3 is aligned with the scale of the feature map M2; The adjusted fused feature P3 is concatenated with the feature map M2 through a concat block, and the fused feature P4 is generated through the feature extraction module; The fused feature P4 is obtained, and the scale of the fused feature P4 is adjusted through a downsampling module so that the scale of the adjusted fused feature P4 is aligned with the scale of the feature map C5; The adjusted fused feature P2 is concatenated with the feature map C5 through a concat block, and the fused feature P5 is generated through the feature extraction module; Among them, the fused feature P2, fused feature P3, fused feature P4, and fused feature P5 are four fused features with different scales.
[0010] In this embodiment, the four fused features with different scales output by the neck network are processed through the head network to obtain the object detection result, including: The fused feature P2, fused feature P3, fused feature P4, and fused feature P5 are respectively input into the detection heads corresponding to their scales; The detection heads adjust the scales of the fused feature P2, fused feature P3, fused feature P4, and fused feature P5 through an upsampling module or a downsampling module so that the adjusted fused feature P2, fused feature P3, fused feature P4, and fused feature P5 are aligned; The adjusted fused feature P2, fused feature P3, fused feature P4, and fused feature P5 are weighted and fused through a weighted summation method to generate a prediction feature for prediction; The prediction feature is input into the prediction layer to obtain the object detection result.
[0011] In this embodiment, the adjusted fusion features P2, P3, P4, and P5 are weighted and fused through a weighted summation formula to generate a prediction feature for prediction; The weighted summation formula includes: where, and , represents the feature of layer is the feature map scaled from layer to layer The parameter is trainable and represents the weights of different feature layers, respectively represent the values at positions at .
[0012] In this embodiment, the feature extraction module includes a first convolutional block, a spllt block, two partial convolutional blocks, a concat block, and a second convolutional block. Among them, the first convolutional block is used to convolve the input feature map to generate an intermediate feature map; The spllt block is used to split the intermediate feature map into a first feature map and a second feature map. The first feature map is directly transmitted to the concat block, and the second feature map is transmitted to the partial convolutional block; The two partial convolutional blocks are connected in series in sequence, and the second feature map is processed by the two partial convolutional blocks to obtain a third feature map; The concat block is used to receive the first feature map and the third feature map, and splice the first feature map and the third feature map to obtain a fourth feature map; The second convolutional block is used to perform convolutional processing on the fourth feature map to generate an output feature map.
[0013] In this embodiment, the downsampling module is an SCD module. The SCD module includes a third convolutional block and a fourth convolutional block. The input feature map is input into the third convolutional block for pointwise convolution to adjust the channel dimension, and then spatial downsampling is performed through depth convolution of the fourth convolutional block, thereby obtaining a feature map with a reduced image size.
[0014] In this embodiment, the loss function of the UAV detection model is STIoU, specifically as follows: Among them, represents the Inner-IoU value, is the Focaler-IoU loss.
[0015] The definition of Focaler-IoU is as follows: The corresponding loss formula is: The mathematical expression of Inner-IoU is: Among them, and respectively represent the left, right, top, and bottom boundary coordinates of the target and the anchor box. is the scale factor of the auxiliary box, and its value range is [0.5, 1.5].
[0016] Distance loss Δ The calculation formula is as follows: The calculation formula of the shape loss Ω is: .
[0017] Beneficial effects: 1. Based on the YOLOv8 detection head, the present invention adds a small target detection layer and combines it with the ASFF4 mechanism to form a QDH detection head, which solves the limitations of traditional three-scale feature fusion methods in dealing with small target detection. In traditional methods, it is difficult to effectively retain the details of small targets in scenarios with complex backgrounds and large variations in target sizes. Especially during the scale change process, it is easy to cause the loss of small target features, affecting the detection accuracy. Therefore, QDH enhances the feature retention ability of small targets and improves the detection accuracy by introducing four-scale feature processing. In addition, to address the problem of information loss in complex environments, the present invention also combines an adaptive spatial feature fusion improvement mechanism (ASFF4). This mechanism dynamically adjusts the contributions of different scale features according to the context importance of the feature layers to ensure the effective retention of key details. Through weighted summation, QDH can integrate information from different scales, thereby optimizing multi-scale fusion and alleviating the common feature confusion problem in traditional methods. During the training process, QDH learns the spatial weights of each scale feature map through backpropagation, effectively filtering out conflicting information from other layers, reducing spatial inconsistency, and avoiding multi-scale conflicts, thus improving the accuracy of drone detection in variable scenarios.
[0018] 2. In the Neck network of the present invention, the multi-scale feature fusion structure is set as a four-scale feature fusion structure (FSF), which improves the small target detection accuracy and effectively reduces the network computational complexity. Traditional multi-scale feature fusion methods, although they can improve the overall detection accuracy, often lead to the dilution of small target features due to the dominance of large target features, thereby affecting the detection accuracy. Therefore, the FSF structure of the present invention is optimized based on the original PANet network, adopting a refined feature layer configuration of five, four, and four layers, and specifically adding a high-resolution feature layer as the fourth scale for small targets, enhancing the sensitivity to small target features. The addition of this high-resolution layer effectively avoids the risk of small target features being covered by large-scale features, alleviates the feature confusion problem, and significantly improves the accuracy of small target detection. Secondly, in the drone detection model of the present invention, as the core component of the QDH module, FSF further enriches the feature hierarchy through multi-scale feature fusion, ensuring that each scale can comprehensively capture detailed information and enhancing the recognition ability of drone features. Through the design of this structure, FSF effectively enhances the small target detection ability, reduces the computational overhead, and provides a more accurate and efficient solution for target detection in complex environments.
[0019] 3. Based on the C2F module, the present invention constructs a new LCF module. The LCF module allows for replacing the original coarse-to-fine processing structure through a more refined feature layer configuration, thereby reducing the computational load and simplifying the model architecture. Through partial convolution technology, the LCF module reduces computational redundancy and effectively reduces the memory access requirements, significantly improving the model's ability to process small targets. This optimization not only enhances the sensitivity to small target features but also avoids the coverage of small target features by large-scale features during the feature fusion process, greatly improving the accuracy of small target detection and providing effective technical support for precise target detection in complex environments.
[0020] 4. Based on Focaler-IoU and Inner-IoU, the present invention constructs the loss function STIoU. Traditional IoU loss functions often cause the optimization to bias towards large targets due to the small target space in small target detection, thereby affecting the detection accuracy. To solve this problem, STIoU combines the advantages of Focaler-IoU and Inner-IoU, adjusts the regression sample loss using a linear interval mapping, and further optimizes the regression efficiency and accuracy of small targets by considering the internal overlap between the target box and the anchor box. This method effectively improves the model's detection ability for small targets in complex environments and provides a more accurate and efficient solution for small target detection.
[0021] 5. To address the adaptation problem of different computational requirements and application scenarios, the present invention designs two variants: extremely small (N) and medium (M). These variants provide different balances between model complexity, computational cost, and detection accuracy by adjusting the network depth and width. It enables efficient real-time detection under different size configurations and meets the multi-scenario application requirements. Brief Description of the Drawings
[0022] Figure 1 It is the model structure diagram of the drone detection model provided by the embodiment of the present invention; Figure 2 It is the structure diagram of the four-scale feature fusion structure provided by the embodiment of the present invention; Figure 3 It is the QDH module structure diagram of the present invention; Figure 4 (a) It is the SCD module structure diagram of the present invention; Figure 4 (b) It is the LCF module structure diagram of the present invention; Figure 5 It is the schematic diagram of the STIoU loss function of the present invention; Figure 6 It is the comparison diagram of the training effects of the N version of the present invention and various models; Figure 7 It is the comparison diagram of the training effects of the M version of the present invention and various models; Figure 8 This is a comparison chart of the detection effects of the present invention and each model. Detailed implementation manners
[0023] Next, embodiments of the technical solution of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present application, so they are only examples and cannot be used to limit the protection scope of the present application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above accompanying drawing descriptions are intended to cover non-exclusive inclusion.
[0025] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "a plurality" is more than two, unless otherwise clearly and specifically defined.
[0026] Referring to "embodiments" herein means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0027] In the description of the embodiments of this application, the term "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0028] In the description of the embodiments of this application, the term "a plurality" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0029] In the description of the embodiments of the present application, the orientation or positional relationship indicated by technical terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the embodiments of the present application.
[0030] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installation", "connection", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can also be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.
[0031] Unmanned aerial vehicles (UAVs) have been widely used in recent years in fields such as surveillance, agriculture, and logistics, significantly improving the development efficiency of these industries. However, the abuse of UAVs in illegal activities has also posed a major threat to public safety and privacy. To address these threats, there is an urgent need to establish an effective UAV surveillance network. Current surveillance technologies include radar, audio, radio, and visual surveillance, etc. Due to the high cost and complex installation requirements of radar, audio, and radio technologies, their deployment faces huge challenges. In contrast, visual target detection has become one of the most promising and feasible UAV surveillance methods due to its low cost and wide application range.
[0032] Recently, many researchers have been dedicated to developing visual methods based on deep convolutional neural networks (DCNNs) to address various UAV detection challenges in the real world. In actual UAV detection scenarios, in order to improve detection accuracy and reduce the missed detection rate, the detection algorithm has high requirements for detection accuracy, detection speed, model complexity, and generalization. Therefore, to meet the above requirements, an improved UAV detection method is urgently needed.
[0033] To solve the technical problem of how to provide an improved drone detection method that can improve the detection accuracy and reduce the missed detection rate in actual drone detection scenarios, the present invention provides a real-time drone detection method based on a four-head detection network. By inputting the real-time acquired image data into the improved drone detection model, the sensitivity of small target features is enhanced, and the small target features are prevented from being covered by large-scale features during the scale fusion process, thereby achieving the technical effect of improving the detection accuracy of small targets.
[0034] A real-time drone detection method based on a four-head detection network provided by an embodiment of the present invention specifically includes the following steps: Acquire image data of the target area in real time; Input the image data into the trained drone detection model for drone target detection.
[0035] Specifically, in this embodiment, the real-time image data of the target area is transmitted to the trained drone detection model. After receiving the image data, the trained drone detection model processes the image data for drone target detection. In this embodiment, by inputting the real-time acquired image data into the improved drone detection model, the sensitivity of small target features is enhanced, and the small target features are prevented from being covered by large-scale features during the scale fusion process, thereby achieving the technical effect of improving the detection accuracy of small targets.
[0036] As Figure 1 shown, Figure 1 is the model structure diagram of the drone detection model provided by an embodiment of the present invention. Exemplarily, in this embodiment, the drone detection model adopts an improved YOLOv8 model structure. By improving the YOLOv8 model structure, the sensitivity of small target features is enhanced, and the small target features are prevented from being covered by large-scale features during the scale fusion process, thereby achieving the technical effect of improving the detection accuracy of small targets.
[0037] The drone detection model that improves the YOLOv8 model structure includes a backbone network for feature extraction, a neck network with a multi-scale fusion structure that is a four-scale fusion structure including four different resolution feature layers, and a head grid with four-scale detection heads.
[0038] Among them, the backbone network (Backbone) is partly responsible for feature extraction. A series of convolutional or deconvolutional layers are adopted, and at the same time, residual connections and bottleneck structures are used to reduce the size of the network and improve performance.
[0039] Neck network (Neck). The Neck network part is responsible for multi-scale feature fusion. By fusing feature maps from different stages of the Backbone, the feature representation ability is enhanced. In this embodiment, the multi-scale fusion structure in the Neck network is set as a four-scale fusion structure (FSF) including four different resolution feature layers. By introducing a higher-resolution feature layer with higher resolution on the basis of the original multi-scale fusion structure, the sensitivity of small target features is enhanced, and small target features are prevented from being covered by large-scale features during the scale fusion process.
[0040] Head network (Head). The Head network is mainly used to be responsible for the final object detection and classification tasks. By using the features extracted by the backbone network and the Neck network, predictions are made, and then the network output content is obtained to detect the UAV target. In this embodiment, the Head network adopts a four-scale detection head (QDH), which enhances the detail retention and detection accuracy of small target features by processing the four-scale features from the Neck network.
[0041] Exemplarily, in this embodiment, after inputting the image data into the trained UAV detection model for UAV target detection, the processing steps of the trained UAV detection model for the image data are as follows: S1. After receiving the above image data, the backbone network in the UAV detection model will perform feature extraction on the image data to obtain four feature maps with different scales, namely feature map C2, feature map C3, feature map C4, and feature map C5, where the feature map scales of the feature map C2, the feature map C3, the feature map C4, and the feature map C5 decrease in sequence.
[0042] Exemplarily, as Figure 2 shown, when the backbone network performs feature extraction on the image data to obtain four feature maps with different scales, the backbone network mainly includes an initial convolutional layer, and a four-stage feature extraction module composed of 4 downsampling modules for halving the feature map size and doubling the number of channels at the same time and 4 feature extraction modules for extracting features in the feature map; Exemplarily, the image data input to the backbone network is a 640×640×3 RGB image; when the image data is input to the initial convolutional layer, the initial convolutional layer performs fast downsampling and edge feature extraction on the image data, and then obtains a feature map C1 with an image size of 320×320×64; After obtaining the feature map C1, the size of the feature map C1 is adjusted by the upsampling module, and the adjusted feature map C1 is subjected to feature extraction by the feature extraction module, and then a feature map C2 with an image size of 160×160×128 is obtained; Obtain the feature map C2, adjust the size of the feature map C2 through the upsampling module, and perform feature extraction on the resized feature map C2 through the feature extraction module, thereby obtaining the feature map C3 with an image size of 80×80×256; Obtain the feature map C3, adjust the size of the feature map C3 through the upsampling module, and perform feature extraction on the resized feature map C3 through the feature extraction module, thereby obtaining the feature map C4 with an image size of 40×40×512; Obtain the feature map C4, adjust the size of the feature map C4 through the upsampling module, and perform feature extraction on the resized feature map C4 through the feature extraction module, thereby obtaining the feature map C5 with an image size of 20×20×1024.
[0043] S2. The neck network performs feature fusion on the four feature maps of different scales through the four-scale fusion structure to obtain four fusion features of different scales, namely the fusion feature P2, the fusion feature P3, the fusion feature P4, and the fusion feature P5; Exemplarily, obtain the feature map C5, and perform three-level pooling on the feature map C5 through the Spatial Pyramid Pooling Fast (SPPF) module to generate the feature map M1; Obtain the feature map M1, and adjust the scale of the feature map M1 through the upsampling module so that the scale of the adjusted feature map M1 is aligned with the scale of the feature map C4; Concatenate the adjusted feature map M1 with the feature map C4 through the concat block, and generate the feature map M2 through the feature extraction module; Obtain the feature map M2, and adjust the scale of the feature map M2 through the upsampling module so that the scale of the adjusted feature map M2 is aligned with the scale of the feature map C3; Concatenate the adjusted feature map M2 with the feature map C3 through the concat block, and generate the feature map M3 through the feature extraction module; Obtain the feature map M3, and adjust the scale of the feature map M3 through the upsampling module so that the scale of the adjusted feature map M3 is aligned with the scale of the feature map C2; Concatenate the adjusted feature map M3 with the feature map C2 through the concat block, and generate the fusion feature P2 with an image size of 160×160×128 through the feature extraction module; Obtain the fusion feature P2, and adjust the scale of the fusion feature P2 through the downsampling module so that the scale of the adjusted fusion feature P2 is aligned with the scale of the feature map M3; Concatenate the adjusted fusion feature P2 with the feature map M3 through the concat block, and generate the fusion feature P3 with an image size of 80×80×256 through the feature extraction module; Obtain the fused feature P3, and adjust the scale of the fused feature P3 through the downsampling module so that the scale of the adjusted fused feature P3 is aligned with the scale of the feature map M2; Concatenate the adjusted fused feature P3 and the feature map M2 through the concat block, and generate a fused feature P4 with an image size of 40×40×512 through the feature extraction module; Obtain the fused feature P4, and adjust the scale of the fused feature P4 through the downsampling module so that the scale of the adjusted fused feature P4 is aligned with the scale of the feature map M1; Concatenate the adjusted fused feature P2 and the feature map M1 through the concat block, and generate a fused feature P5 with an image size of 20×20×1024 through the feature extraction module.
[0044] In this embodiment, the multi-scale fusion structure in the neck network is set to a four-scale fusion structure (FSF) including four different resolution feature layers. The four-scale fusion structure (FSF) introduces a higher-resolution feature layer with higher resolution on the basis of the original multi-scale fusion structure to enhance the sensitivity of small target features and avoid small target features being covered by large-scale features during the scale fusion process. The four-scale fusion structure adopts a feature layer configuration of five layers, four layers, and four layers, that is, it includes a five-layer backbone network downsampling feature map structure composed of the feature maps C1, C2, C3, C4, and C5; a feature pyramid network structure of the top-down path of the four-layer neck network composed of the feature maps M1, M2, M3, and M4; and a feature pyramid network structure of the bottom-up path of the four-layer neck network composed of the fused features P2, P3, P4, and P5. The four-scale fusion structure adopts a feature layer configuration of five layers, four layers, and four layers, and improves the feature expression ability through cross-layer path aggregation, thereby improving the accuracy of small target detection. In addition, through the detail enhancement of the low-level feature map, the four-scale fusion structure reduces feature confusion and improves the capture ability of UAV features.
[0045] S3. Process the four different-scale fused features output by the neck network through the head network to obtain the target detection result.
[0046] As Figure 3As shown in the figure, in this embodiment, a four-scale detection head (QDH) is used to process the fused features of four different scales from the neck network, enhancing the detail retention and detection accuracy of small object features. At the same time, to address the feature fusion challenges in complex environments, an adaptive spatial feature fusion mechanism (ASFF4) is introduced. This mechanism dynamically adjusts the contributions of features at each scale according to the context importance of the features, effectively avoiding the detail loss caused by standard fusion methods under scale changes and resolution changes, thereby improving the recognition accuracy and robustness of small objects.
[0047] Exemplarily, the fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are respectively input into the detection heads corresponding to their scales; The detection heads adjust the scales of the fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 through an upsampling module or a downsampling module, so that the adjusted fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are aligned; The adjusted fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are weighted and fused by a weighted summation method to integrate multi-scale information and generate a prediction feature for prediction; The prediction feature is input into a prediction layer to obtain the object detection result.
[0048] In this embodiment, the adjusted fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are weighted and fused by a weighted summation formula to generate a prediction feature for prediction; The weighted summation formula includes: Among them, And , Represents The feature of layer Is the feature map scaled from layer To layer The parameter Is trainable and represents the weights of different feature layers, Respectively represent the position At Value.
[0049] In this embodiment, the downsampling module is an SCD module, and the SCD module includes a third convolutional block and a fourth convolutional block. The input feature map is input into the third convolutional block for pointwise convolution to adjust the channel dimension, and then spatial downsampling is performed through depth convolution of the fourth convolutional block, thereby obtaining a feature map with a reduced image size. It can be understood that the SCD module performs spatial downsampling through spatial and channel decoupling operations, using pointwise convolution and depth convolution, reducing the amount of computation and the number of parameters, while retaining information to the greatest extent, thus achieving the goal of improving the model efficiency while reducing the computational overhead. As Figure 4 shown in (a), SCD uses point convolution to adjust the channel dimension and then performs spatial downsampling through depth convolution. The computational cost is reduced to O (2 HWC ² + 9 / 2 HWC ), and the number of parameters is reduced to O (2 C ² + 18 C ).
[0050] In this embodiment, the feature extraction module uses an LCF module. The LCF module includes a first convolutional block, a spllt block, two partial convolutional blocks, a concat block, and a second convolutional block. Among them, the first convolutional block is used to generate an intermediate feature map by convolving the input feature map; The spllt block is used to split the intermediate feature map into a first feature map and a second feature map. The first feature map is directly passed to the concat block, and the second feature map is passed to the partial convolutional block; The two partial convolutional blocks are connected in series in sequence, and the second feature map is processed by the two partial convolutional blocks to obtain a third feature map; The concat block is used to receive the first feature map and the third feature map, and splice the first feature map and the third feature map to obtain a fourth feature map; The second convolutional block is used to perform convolution processing on the fourth feature map to generate an output feature map.
[0051] It can be understood that the LCF module retains the original double-branch structure on the basis of C2F and introduces partial convolution to replace traditional convolution. As Figure 4 shown in (b), and respectively represent the width and height of the input and output feature maps, represents the number of channels participating in the convolution, Denotes the convolutional kernel size. This method extracts spatial features by selectively applying convolutions to a subset of the input channels while keeping the remaining channels unchanged. To efficiently utilize memory, the entire feature map is represented by the first or last consecutive channels during calculation. Assuming the number of channels in the input and output feature maps is the same, the FLOPs calculation formula for partial convolution is as follows: In addition, partial convolution requires the minimum memory access amount, which can be specifically expressed by the following formula: In this embodiment, to improve the accuracy of object detection and optimize the positioning and size of the bounding boxes, the loss function of the UAV detection model is set to STIoU. It can be understood that this embodiment combines Focaler-IoU and Inner-IoU to form the loss function STIoU. When dealing with small objects, the traditional IoU loss function is prone to causing the optimization to bias towards large objects because its contribution to the loss calculation is relatively small. Focaler-IoU improves the traditional IoU by adjusting the loss value through linear interval mapping, focusing on small objects and difficult samples, thereby improving the detection accuracy of small objects.
[0052] The definition of Focaler-IoU is as follows: The corresponding loss formula is: This method effectively reduces the optimization bias of small objects and improves the overall detection accuracy. To further improve the detection accuracy of small objects, this study combines Focaler-IoU and Inner-IoU and proposes the STIoU loss function. STIoU not only considers traditional factors such as overlapping area, distance, angle, and direction, but also optimizes the processing method of diverse IoU samples by adjusting the scale ratio of the auxiliary box, thereby improving the detection accuracy of small objects and occluded objects.
[0053] The loss formula of STIoU is as follows: Among them, Denotes the Inner-IoU value, Is the Focaler-IoU loss.
[0054] The mathematical expression of Inner-IoU is: Among them, and respectively represent the left, right, top, and bottom boundary coordinates of the target and the anchor box. is the scale factor of the auxiliary box, and its value range is [0.5, 1.5]. A larger auxiliary box helps to include more context information and reduce the error caused by pixel deviation, thereby improving the detection accuracy.
[0055] Distance loss Δ The calculation formula is as follows: The calculation formula for the shape loss Ω is: Among them, Ψ represents the degree of attention to the shape loss, and its value range is [2, 6], and the default value is 4. The STIoU loss function not only considers the angular factor but also adapts to IoU samples of different scales through auxiliary boundary training, further improving the target detection accuracy in complex backgrounds or at long distances. It performs particularly well in the drone detection task, and its principle is illustrated as Figure 5 shown.
[0056] Exemplarily, to solve the adaptation problem of different computing requirements and application scenarios, this embodiment designs two variants, Nano (N) and Medium (M), based on the drone detection model. These variants provide different balances between model complexity, computational cost, and detection accuracy by adjusting the network depth and width. The smallest N model contains approximately 2.8M parameters and is suitable for low-latency real-time deployment scenarios; M has 17.9M parameters and is suitable for applications with high requirements for accuracy. The comprehensive selection of these variants meets the needs of various operating conditions, from resource-constrained edge deployments to performance-critical tasks with sufficient computing power. Exemplarily, in this embodiment, before processing the image data through the UAV detection model for UAV target detection, it is also necessary to train the UAV detection model obtained after improving the YOLOv8 model structure, so that the UAV detection model can learn features from the data, adjust parameters to optimize performance.
[0057] Exemplarily, the training method of the UAV detection model specifically includes: A1. Obtain a UAV dataset, and after preprocessing it, divide it into a training set and a test set; Use web crawlers to crawl UAV pictures in various scenarios from the network, perform size normalization processing on the collected images, and expand 10,000 UAV pictures in the public dataset DUT_Anti-UAV to 21,308 pictures. Use the labelimg label annotation software to annotate the UAV targets in the expanded UAV dataset, and generate an xml file storing target anchor information. Finally, divide the constructed UAV dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0058] A2. Construct a UAV detection model, and use the training set and the test set for the UAV detection model.
[0059] (1) Construct an experimental environment The algorithm of the present invention is written based on the Pytorch framework, and the experiment is carried out in the environment of Table 1.
[0060] Table 1 Experimental environment The hyperparameter settings for the model training of the present invention are shown in Table 2: Table 2 Hyperparameter settings (2) Model training Send the training set into the UAV detection model for training to obtain the trained UAV detection model, and use the trained UAV detection model for testing.
[0061] (3) Evaluation index To ensure a fair comparison with the latest real-time detectors on the dataset, the applicant uses standard evaluation metrics such as average precision AP, AP50, and AP75. Among them, AP is calculated through the intersection over union (IoU) between the predicted bounding box and the ground truth bounding box. The IoU threshold ranges from 0.5 to 0.95 with a step of 0.05, which can reflect the comprehensive performance of the detector; AP50 and AP75 are evaluated at IoU thresholds of 0.5 and 0.75 respectively. In addition, we also evaluate the implementation efficiency through common performance metrics such as the floating-point value (FLOPs) of the model, the number of parameters (Params), and the latency (Lantency). The model size is measured by the number of parameters. To evaluate the detection effect, metrics such as the mean average precision (mAP), the average precision, precision, and recall at an IoU threshold of 50% (mAP50) and 75% (mAP75) are used. Precision and recall are used to measure the ability of the model to miss detections and false alarms in drone detection, thus truly reflecting the distribution of drones. The formulas are shown as follows: Among them, TP, FP, and FN represent the correctly identified positive samples, the wrongly identified positive samples, and the wrongly identified negative samples respectively, and these samples are all within the real device area.
[0062] (4)Experimental Results and Analysis 4.1) Ablation Experiments To evaluate the actual effect of the network structure optimization of the present invention, the applicant carried out a series of ablation experiments with the YOLOv8-M model as the benchmark. The experimental comparison data is shown in Table 3, where M represents mAP (mean average precision), Pr represents Precision, R represents Recall, Pa represents Params (number of parameters), and F represents FLOPs (number of floating-point operations).
[0063] Table 3 Ablation Experiment Data (bold underlined represents the best) Without pre-trained weights, the AP (Average Precision), Precision, Recall, and Params (number of parameters) of the baseline network are 58.3%, 93.9%, 79.5%, and 25.9M respectively, and the FLOPs (floating point operations) are 79.3G. After adding the FSF structure, the mAP is improved by 4.10%, indicating that this structure can better focus on target information. By replacing the detection head with the QDH head adapted to FSF, although the accuracy is improved, the model complexity increases, with Params increasing by 8.1M and FLOPs increasing by 40.8G. After introducing the LCF and SCD modules, although the detection accuracy decreases slightly, the model size is significantly reduced, with Params decreasing by 16.1M and FLOPs decreasing by 50.4G. The optimized loss function makes the mAP, Precision, Recall, and Params of the model be 62.1%, 97.2%, 87.1%, and 17.9M respectively. Although the mAP decreases compared with the unoptimized model, the significant reduction in Params, FLOPs, and inference time significantly improves the detection speed and real-time performance, meeting the actual application requirements.
[0064] Compared with the baseline network, after the structure optimization, the AP is increased by 3.8% and the accuracy is increased by 3.3%. The number of model parameters is reduced by 8.0M, indicating that the network improves the accuracy of UAV detection while reducing the complexity and computational burden. At the same time, the FLOPs are reduced to 71.1G, a reduction of 8.2G, verifying that the algorithm enhances the detection ability while effectively reducing the spatial complexity. These ablation experiment results fully prove that the proposed algorithm improves the accuracy of UAV detection while simplifying the complexity and computational burden.
[0065] 4.2) Comparative experiments To evaluate the model performance of the present invention, the applicant conducted comprehensive comparative experiments with a variety of advanced detection algorithms. The experimental comparison data are shown in Table 4, where M represents mAP, M50 represents mAP50, M75 represents mAP75, Pa represents Params, F represents FLOPs, and L represents Latency.
[0066] Table 4 Comparative experiment data (bold underlined represents the best) The experimental results show that the present invention has achieved an optimal balance in terms of detection accuracy and implementation efficiency. In the extremely small version (N), the mAP of the present invention on the dataset is 57.1%, leading the sub-optimal model YOLOv9-N by 6.0%. In addition, the present invention also outperforms all other models in mAP50 and mAP75, reaching 88.9% and 65.9% respectively. Although models such as YOLOv10-N and YOLOv6-N have lower computational overheads (6.79G and 13.1G FLOPs respectively), their detection performances (mAP of 49.5% and 49.1% respectively) are significantly behind the present invention. In the medium version (M), the present invention continues to demonstrate excellent detection performance. For example, the present invention reaches 62.1% in mAP, and mAP50 and mAP75 are 92.8% and 71.4% respectively, significantly exceeding YOLO-MS-S and DAMO-YOLO-S. Although the latency of the present invention is relatively high (20.02 ms), the improvement in accuracy makes this trade-off highly practical. The experimental comparison of the N and M versions of the present invention is as Figure 6 , 7 shown.
[0067] 4.3) Visual comparison To demonstrate the advantages of the present invention in the UAV detection task, the applicant adopted visualization means to compare it with other detection models. This qualitative evaluation assesses the detection effects in challenging scenarios such as long-distance detection and complex backgrounds through representative visual examples. The visual results highlight the advantages of the present invention in accurately locating and focusing on the target UAV, demonstrating its robustness and feature extraction ability compared with other methods.
[0068] As Figure 8 shown, the applicant presented the detection results of each model in long-distance UAV detection, UAV detection in complex backgrounds, and detection of overlapping structures such as power lines. The present invention always shows a relatively high detection confidence. The detected UAVs are shown by red bounding boxes with relatively high confidence scores. For example, in the first row, the confidence levels of the UAVs detected by the present invention are 0.88, 0.87, and 0.90 respectively, while YOLOv8 and YOLO-MS often miss detections or report relatively low confidence levels, about 0.72 or lower. In contrast, models such as YOLOv8 and DAMO-YOLO often miss detections (MISS) in complex scenarios. The present invention provides robust detection with compact and well-aligned bounding boxes, reducing false positives and accurately locating the UAV.
[0069] It should be noted that this application is not limited to the above-described embodiments. The above-described embodiments are merely examples, and embodiments having the same constitution as the technical idea and achieving the same effects within the scope of the technical solution of this application are all included in the technical scope of this application. In addition, within the scope not departing from the gist of this application, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways constructed by combining some of the constituent elements in the embodiments are also included in the scope of this application.
Claims
1. A real-time UAV detection method based on a four-head detection network, characterized in that, It includes the following steps: Obtain the image data of the target area in real time; Input the image data into the trained UAV detection model for UAV target detection; Among them, the UAV detection model includes a backbone network for feature extraction, a neck network with a multi-scale fusion structure as a four-scale fusion structure including four different resolution feature layers, and a head grid with a four-scale detection head.
2. The real-time UAV detection method based on a four-head detection network according to claim 1, wherein The training method of the UAV detection model includes: Obtain a UAV data set and divide it into a training set and a test set after preprocessing; Construct a UAV detection model and use the training set and the test set for the UAV detection model.
3. The real-time UAV detection method based on a four-head detection network according to claim 1, characterized in that The inputting the image data into the trained UAV detection model for UAV target detection includes: The backbone network extracts features from the image data to obtain four feature maps with different scales; The neck network performs feature fusion on the four feature maps with different scales through the four-scale fusion structure to obtain four fusion features with different scales; The head network processes the four fusion features with different scales output by the neck network to obtain the target detection result.
4. The real-time UAV detection method based on the four-head detection network according to claim 3, wherein The backbone network extracts features from the image data to obtain four feature maps with different scales, including: Preprocess the image data to adjust the size and pixel value of the image data; Input the preprocessed image data into the initial convolutional layer to obtain the feature map C1; Perform feature extraction on the feature map C1 through a four-stage feature extraction module to obtain the feature maps C2, C3, C4, and C5. The feature map scales of the feature maps C2, C3, C4, and C5 decrease in sequence; Among them, the four-stage feature extraction module includes 4 downsampling modules for halving the feature map size and doubling the number of channels, and 4 feature extraction modules for extracting features in the feature map. The feature map C1 passes through the downsampling module and the feature extraction module in sequence.
5. The real-time drone detection method based on a four-head detection network according to claim 4, wherein The neck network performs feature fusion on the four feature maps with different scales through the four-scale fusion structure to obtain four fusion features with different scales, including: Obtain the feature map C5 and adjust the scale of the feature map C5 through an upsampling module so that the scale of the adjusted feature map C5 is aligned with the scale of the feature map C4; Concatenate the adjusted feature map C5 with the feature map C4 through a concat block and generate the feature map M2 through the feature extraction module; Obtain the feature map M2 and adjust the scale of the feature map M2 through an upsampling module so that the scale of the adjusted feature map M2 is aligned with the scale of the feature map C3; Concatenate the adjusted feature map M2 with the feature map C3 through a concat block and generate the feature map M3 through the feature extraction module; Obtain the feature map M3 and adjust the scale of the feature map M3 through an upsampling module so that the scale of the adjusted feature map M3 is aligned with the scale of the feature map C2; The adjusted feature map M3 is concatenated with the feature map C2 through a concat block, and the fused feature P2 is generated through the feature extraction module; The fused feature P2 is obtained, and the scale of the fused feature P2 is adjusted through the downsampling module so that the scale of the adjusted fused feature P2 is aligned with the scale of the feature map M3; The adjusted fused feature P2 is concatenated with the feature map M3 through a concat block, and the fused feature P3 is generated through the feature extraction module; The fused feature P3 is obtained, and the scale of the fused feature P3 is adjusted through the downsampling module so that the scale of the adjusted fused feature P3 is aligned with the scale of the feature map M2; The adjusted fused feature P3 is concatenated with the feature map M2 through a concat block, and the fused feature P4 is generated through the feature extraction module; The fused feature P4 is obtained, and the scale of the fused feature P4 is adjusted through the downsampling module so that the scale of the adjusted fused feature P4 is aligned with the scale of the feature map C5; The adjusted fused feature P2 is concatenated with the feature map C5 through a concat block, and the fused feature P5 is generated through the feature extraction module; Among them, the fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are fused features of four different scales.
6. The real-time UAV detection method based on a four-head detection network according to claim 5, characterized in that, The four different scales of the fused features output by the neck network are processed through the head network to obtain the object detection result, including: The fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are respectively input into the detection heads of the corresponding scales; The detection heads adjust the scales of the fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 through the upsampling module or the downsampling module so that the adjusted fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are aligned; The adjusted fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are weighted and fused through the method of weighted summation to generate a prediction feature for prediction; The prediction feature is input into the prediction layer to obtain the object detection result.
7. The real-time UAV detection method based on the four-head detection network according to claim 6, characterized in that, The adjusted fused feature P2, the fused feature P3, the fused feature P4, and the fused feature P5 are weighted and fused through the weighted summation formula to generate a prediction feature for prediction; The weighted summation formula includes: Among them, and , represents the feature of the layer, is the feature map scaled from layer to layer , and the parameter is trainable and represents the weights of different feature layers, respectively represent the values at the position at .
8. The real-time UAV detection method based on a four-head detection network according to any one of claims 1-7, characterized in that, The feature extraction module includes a first convolution block, a spllt block, two partial convolution blocks, a concat block, and a second convolution block, Among them, the first convolution block is used to convolve the input feature map to generate an intermediate feature map; The spllt block is used to split the intermediate feature map into a first feature map and a second feature map. The first feature map is directly transmitted to the concat block, and the second feature map is transmitted to the partial convolution block; The two partial convolution blocks are connected in series in sequence, and the second feature map is processed through the two partial convolution blocks to obtain a third feature map; The concat block is used to receive the first feature map and the third feature map, and splice the first feature map and the third feature map to obtain a fourth feature map; The second convolutional block is used to perform convolutional processing on the fourth feature map.
9. The real-time UAV detection method based on a four-head detection network according to any one of claims 1-7, characterized in that, The downsampling module is an SCD module. The SCD module includes a third convolutional block and a fourth convolutional block. The input feature map is input into the third convolutional block for pointwise convolution to adjust the channel dimension, and then depth convolution is performed through the fourth convolutional block for spatial downsampling, thereby obtaining a feature map with a reduced image size.
10. The real-time drone detection method based on a four-head detection network according to claim 1, characterized in that, The loss function of the UAV detection model is STIoU, specifically as follows: Among them, represents the Inner-IoU value, is the Focaler-IoU loss; The definition of Focaler-IoU is as follows: The corresponding loss formula is: The mathematical expression of Inner-IoU is: Among them, and respectively represent the left, right, top, and bottom boundary coordinates of the target and the anchor box, is the scale factor of the auxiliary box, and its value range is [0.5, 1.5]; Distance loss Δ The calculation formula is as follows: The calculation formula of the shape loss Ω is: 。
Citation Information
Patent Citations
Unmanned aerial vehicle multi-scale target detection and identification method
CN113420607A
Small target detection method for images acquired by unmanned aerial vehicle based on improved YOLOv8 algorithm
CN118628939A
Unmanned aerial vehicle aerial photography target detection method based on improved YOLOv8
CN119832456A