Unmanned aerial vehicle aerial photography small target detection method and device based on FLEX-YOLO network
By adopting the FLEX-YOLO network in the aerial target detection of drones, combining the multi-path feedback perception module, joint feature interaction network and GhostPlus lightweight module, the problems of small object detection accuracy and calculation cost are solved, and efficient and real-time detection performance is achieved.
Patent Information
- Application Number
- CN202510112626.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing drone aerial target detection method has limited detection accuracy for small targets in complex contexts and is highly computationally cost-effective, making it difficult to achieve lightweight and efficient detection of models.
Using the FLEX-YOLO network-based method, the multi-path feedback perception module MPFP, joint feature interaction network JFIN and GhostPlus lightweight module are introduced, combining deep separation convolution, double-layer routing attention and multi-layer perceptron to enrich feature extraction capabilities and reduce computing complexity, and compress the model through LAMP pruning strategy.
It significantly improves the detection accuracy of small targets, reduces the computational complexity of the model, and achieves efficient and real-time detection performance, which is suitable for drone aerial photography scenarios.
Smart Images

Figure CN119942085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method and device for detecting small targets in unmanned aerial photography based on a FLEX-YOLO (Feedback-Driven Lightweight and Enhanced Exchange Network) network. Background Art
[0002] In recent years, drones have shown their important value in the field of target detection due to their advantages of lightness, flexibility, strong concealment, and long flight time. Their ability to flexibly adjust flight altitude and viewing angle can achieve efficient observation and target tracking in wide-area scenes, significantly improving the efficiency and accuracy of detection tasks, and becoming an important tool for intelligent visual perception.
[0003] Object detection is an important task in computer vision. In recent years, detection algorithms based on deep learning have made significant progress in accuracy and efficiency. Currently popular object detection methods are mainly divided into two-stage detection algorithms and single-stage detection algorithms. Two-stage algorithms (such as R-CNN and its improved versions) achieve precise positioning by generating candidate regions, but the computational overhead is large; while single-stage algorithms (such as the YOLO series) achieve faster detection speeds by directly regressing bounding boxes and class probabilities. The YOLO series of models has become a popular choice in the field of drone aerial photography object detection due to its real-time and high efficiency.
[0004] However, the targets in drone aerial photography scenes usually have characteristics such as small size, dense distribution, frequent occlusion and complex background, which poses higher challenges to target detection methods. Small targets are easily disturbed by complex backgrounds due to their weak feature expression ability, and dense distribution and frequent occlusion further increase the difficulty of detection. Although existing target detection algorithms have made certain progress in feature extraction, the detection accuracy of small targets in complex backgrounds is still limited, and the multi-scale feature fusion mechanism also has obvious deficiencies in the interaction of non-adjacent layer information. In addition, the high computational cost further limits the feasibility of existing methods in practical applications. Therefore, how to achieve lightweight models while improving the performance of small target detection has become a core issue that needs to be solved urgently. Summary of the invention
[0005] The present invention aims to overcome the problems of insufficient feature extraction capability and excessive computational overhead in small target detection in existing methods, and proposes a method and device for detecting small targets in drone aerial photography based on the FLEX-YOLO network.
[0006] The present invention improves the existing network structure and combines efficient feature extraction strategies with lightweight design, which can effectively improve the detection accuracy of small targets and reduce the calculation complexity of the model, thereby meeting the requirements for high efficiency and real-time performance in practical applications.
[0007] The present invention achieves the above-mentioned purpose through the following technical scheme: a method for detecting small targets in drone aerial photography based on FLEX-YOLO network, comprising the following steps:
[0008] S1: Obtain the public drone aerial image dataset Visdrone2019 and perform preprocessing;
[0009] S2: Build a FLEX-YOLO network model, which is based on the YOLOv11 network. The designed multi-path feedback perception module MPFP (Multi-Path Feedback Perception) is introduced into the backbone network of the YOLOv11. The improved backbone network is obtained by combining the depthwise separable convolution DWConv (Depthwise Separable Convolution), the bi-level routing attention BRA (Bi-level Routing Attention) and the multi-layer perceptron MLP (Multi-Layer Perceptron) through the multi-path design.
[0010] S3: The neck network of YOLOv11 is improved to the Joint Feature Interaction Network (JFIN). By establishing a fully connected information flow channel between different feature layers to enrich the information interaction between non-adjacent layers, an improved neck network is obtained.
[0011] S4: The designed GhostPlus lightweight module is introduced into the backbone network and neck network of YOLOv11. The traditional channel splicing operation is replaced by point-by-point addition operation, and the "front activation" strategy is introduced to achieve lightweight calculation of the model.
[0012] S5: The constructed model is compressed using the global scoring-based pruning strategy LAMP (Layer-Adaptive Magnitude Pruning);
[0013] S6: Use the FLEX-YOLO network as the detection model, train and verify it using the training set and validation set to generate the final detection model;
[0014] S7: Use the final trained detection model and the test set as input to test the FLEX-YOLO model.
[0015] Preferably, in step S1, the pre-processing step includes:
[0016] S11: Convert the original annotation format to YOLO format, where each row of data represents a target information, including the target category number, the center position coordinates x and y of the target in the image, and the width w and height h of the target;
[0017] S12: All images are cropped or filled to a fixed size pixel, while maintaining the aspect ratio of the image, and the unfilled area is processed with zero filling;
[0018] S13: The data set is divided reasonably. The training set is used for model training, the validation set is used for parameter optimization and performance evaluation, and the test set is used for testing the final detection effect.
[0019] Preferably, in step S2, the design method of the multi-path feedback perception module MPFP is:
[0020] S21: DWConv is used instead of standard convolution. The standard convolution is decomposed into a depthwise convolution with a convolution kernel of 3×3, and a convolution operation is performed independently on each channel to extract spatial features, and a point-by-point convolution with a convolution kernel of 1×1. All channels are linearly combined to fuse the information between channels.
[0021] S22: Introducing the double-layer routing attention BRA, the core of which is to retain the most relevant Top-k connections between each region and other regions, and to generate the adjacency matrix A r Perform the Top-k operation row by row to generate the routing index matrix I r , where each row contains the index of the current region and its k most relevant regions to avoid redundant calculations;
[0022] I r =TopKIndex(A r ) (1)
[0023] S23: Add a multi-layer perceptron (MLP) to enhance the expressiveness of channel features through nonlinear transformation.
[0024] Preferably, in step S3, the design method of the joint feature interaction network JFIN is:
[0025] S31: The feature maps extracted by Backbone in YOLOv11 are marked from C1 to C5 according to the scale from large to small. C1 belongs to the initial stage of feature extraction and does not directly participate in the multi-scale fusion process. First, the middle-level features of C3 and C4 are fused to start the multi-scale fusion sequence:
[0026]
[0027] S32: A fully connected information flow channel is established between different feature layers to enrich the information interaction between non-adjacent layers:
[0028]
[0029] in, represents the output feature after fusion at position i, j in feature layer k, or represents the feature vector extracted from the original feature map, Used to model the feature interactions between C3 and C4, Represents the weight parameter and satisfies the normalization constraint
[0030] S33: Based on the multi-scale detection mechanism of JFIN, a small target detection head with a resolution of 160×160 is added to the network structure.
[0031] Preferably, in step S4, the design method of the GhostPlus lightweight module is:
[0032] S41: Use Add operation to replace Concat operation, and replace the traditional channel splicing operation with point-by-point addition operation:
[0033]
[0034] Among them, s represents the number of feature channels merged by the addition operation;
[0035] S42: The “pre-activation” strategy is introduced, which places the normalization and activation function before the convolution operation, and replaces the traditional ReLU activation function with the SiLU activation function with better nonlinear expression ability.
[0036] Preferably, in step S5, the implementation method of the LAMP pruning algorithm is:
[0037] S51: Based on the weight magnitude, an importance score definition is introduced. For the u-th index in the weight tensor W, the LAMP score is defined as follows:
[0038]
[0039] Among them, W[u] represents the element with index u in the weight tensor, and the denominator is the sum of the squares of all weights from index u to the end;
[0040] S52: Based on the LAMP score, the relative importance of the target weight in the remaining weights of its layer is measured, and the weights with relatively low LAMP scores are pruned according to the set pruning rate.
[0041] Preferably, in step S6, the FLEX-YOLO network is used as the detection model, and the training set and the verification set are used to train and verify the detection model, and the step of generating the final detection model includes:
[0042] S61: The experimental environment is set on the Ubuntu 20.04.4 LTS system, based on the PyTorch 2.5.0 framework, and accelerated using CUDA 11.4. The hardware environment is NVIDIA RTX A6000 GPU with 48G video memory;
[0043] S62: The initial learning rate of the model is set to 0.01, the momentum parameter is set to 0.937, the weight decay coefficient is 0.0005, and the batch size is set to 16;
[0044] S63: Input the preprocessed training set into the FLEX-YOLO network, calculate the classification loss, positioning loss and confidence loss through forward propagation, and then perform back propagation and update the parameters. The training process continues for 200 epochs;
[0045] S64: After each epoch, the validation set is input into the current model for inference, and the performance indicators are evaluated in real time. After the training is completed, the model weight with the highest mAP of the validation set is selected as the final detection model.
[0046] The second aspect of the present invention relates to a small target detection device for drone aerial photography based on a FLEX-YOLO network, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the small target detection method for drone aerial photography based on a FLEX-YOLO network of the present invention.
[0047] A third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the method for detecting small targets in drone aerial photography based on a FLEX-YOLO network of the present invention.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] Based on the YOLOv11 network, the present invention designs a multi-path feedback perception module MPFP, which significantly enhances the feature extraction capability by fusing deep separable convolution DWConv, double-layer routing attention mechanism BRA and multi-layer perceptron MLP. At the same time, a joint feature interaction network JFIN is designed to achieve more efficient multi-scale feature fusion by establishing a fully connected information flow channel between different feature layers. In addition, the GhostPlus lightweight module is adopted to replace channel splicing by point-by-point addition, and the "pre-activation" strategy is introduced to effectively reduce the computational complexity. In order to further improve the efficiency of the model, the LAMP pruning strategy is used to compress the model, thereby accelerating the inference speed and improving the real-time performance. Experimental results show that the present invention surpasses most mainstream target detection algorithms and some advanced small target detection methods on the VisDrone2019 dataset, demonstrating its advantages in performance and efficiency, and providing an accurate, efficient and deployable solution for drone visual perception tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of the overall process of the method of the present invention;
[0051] Figure 2 FLEX-YOLO network structure diagram designed for an embodiment of the present invention;
[0052] Figure 3 It is a structural diagram of the MPFP module of the present invention;
[0053] Figure 4 This is the JFIN feature fusion network structure diagram of the present invention;
[0054] Figure 5 It is a structural diagram of the GhostPlus module of the present invention;
[0055] Figure 6 It is a comparison diagram of LAMP pruning of the present invention;
[0056] Figure 7 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION
[0057] The present invention is further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:
[0058] Example 1
[0059] like Figure 1 As shown, a method for detecting small targets in UAV aerial photography based on FLEX-YOLO network includes the following steps:
[0060] Step S1: Obtain the public drone aerial image dataset Visdrone2019 and perform preprocessing.
[0061] This embodiment uses the VisDrone2019 dataset collected and collated by the AISKYEYE team of Tianjin University. The dataset is designed for drone vision tasks and aims to promote the deep integration of drones and visual perception technology. The VisDrone2019 dataset contains 8629 static images collected from various scenes in different cities and villages, covering 10 target categories, including pedestrians, vehicles, bicycles, etc., with a total of 343,205 object instances annotated. It has rich visual challenges such as large target size differences, complex backgrounds, and frequent occlusions. It is an important benchmark in the field of drone vision. In the data preprocessing process, the present invention converts the annotation format of the VisDrone2019 dataset, converting the original annotation format into a YOLO format adapted to the FLEX-YOLO network, where each row of data represents a target information, specifically including the target category number, the target center position coordinates x and y in the image, and the target width w and height h. In addition, all images are cropped or padded to a fixed size of 640×640 pixels, while maintaining the aspect ratio of the image, and the unfilled area is processed with zero padding. In order to ensure the effectiveness of model training, the data set was reasonably divided, with 6471 images for the training set, 548 images for the validation set, and 1610 images for the test set, so as to be used for model training, hyperparameter optimization and performance evaluation, and the final detection effect test.
[0062] Step S2: Construct the FLEX-YOLO network model, whose network structure is as follows Figure 2 The network is based on the YOLOv11 network. The designed multi-path feedback perception module MPFP is introduced into the backbone network of YOLOv11. The improved backbone network is obtained by combining the deep separable convolution DWConv, the double-layer routing attention BRA and the multi-layer perceptron MLP through multi-path design. The implementation principle of BRA and the overall architecture of the MPFP module are shown in Figure 3 shown.
[0063] Although the C3K2 module in YOLOv11 has been optimized in feature extraction, it still mainly relies on traditional convolution operations. This operator based on local receptive fields has natural limitations in capturing global information and modeling long-distance dependencies, especially when dealing with small targets and complex scenes. To this end, the present invention designs a multi-path feedback perception module MPFP, which comprehensively improves the feature extraction capability through multi-path design combined with deep separable convolution DWConv, double-layer routing attention BRA and multi-layer perceptron MLP, and realizes efficient target detection in complex scenes.
[0064] In the MPFP module, DWConv is used to replace the standard convolution. It significantly reduces the computational complexity and number of parameters by decomposing the standard convolution into deep convolution and point-by-point convolution. Specifically, the deep convolution uses a 3×3 convolution kernel to perform convolution operations independently on each channel to extract spatial features; the point-by-point convolution uses a 1×1 convolution kernel to linearly combine all channels to fuse information between channels. This decomposition method not only reduces the computational complexity, but also retains the feature extraction capability of the local receptive field.
[0065] At the same time, the MPFP module introduces the BRA dynamic sparse attention mechanism to capture long-distance context information in the spatial dimension. The key to BRA lies in the implementation of its dynamic sparse attention mechanism, which greatly avoids redundant calculations through flexible regional routing. The implementation process is as follows:
[0066] First, for the input feature map X∈R H×W×C , it can be divided into S×S regions, each region contains feature vectors. Then we can get the query, key, and value tensors The linear change is as follows:
[0067] Q=X r W q ,K=X r W k ,V=X r W v (1)
[0068] Among them, represents X r The feature map representing the region division, W q ,W k ,W v They are the linear projection matrices used to generate Q, K, and V respectively.
[0069] On this basis, a directed graph is constructed to determine the association between each region and other regions, and Q and K are averagely pooled at the regional level to obtain regional-level queries and keys. By performing matrix operations on them, an adjacency matrix describing the semantic association between regions is generated.
[0070] A r =Q r (K r ) T (2)
[0071] BRA preserves the top-k most relevant connections of each region to other regions. r Perform the Top-k operation row by row to generate the routing index matrix Ir , where each row contains the index of the current region and its k most related regions for subsequent calculations.
[0072] I r =TopKIndex(A r ) (3)
[0073] Based on the regional routing index matrix I r , BRA implements fine-grained Token-to-Token attention. For the query token in region i, its attention range is limited to I r The key-value pairs in the specified k related regions. Since the related regions are scattered, in order to improve the calculation efficiency, we first use I r Collect keys and values:
[0074] K g =gather(K,I r ),V g =gather(V,I r ) (4)
[0075] in, And use gather to select key-value pairs from the top-k windows most relevant to it for attention calculation. Then, the attention is calculated by the following formula:
[0076] O=Attention(Q,K g ,V g )+LCE(V) (5)
[0077] Here, LCE(·) indicates parameterization using depthwise convolution.
[0078] In addition, in order to further optimize the feature expression of the channel dimension, the present invention adds a multi-layer perceptron MLP to the MPFP module. MLP enhances the expression ability of channel features through nonlinear transformation and promotes the interaction of information between channels.
[0079] The MPFP module takes into account both spatial and channel dimension feature expression. By dynamically weighting and fusing global context information with local receptive fields, it not only enhances the model's ability to extract features of small targets in complex backgrounds, but also effectively alleviates the detection difficulties caused by dense distribution and frequent occlusion of targets, significantly improving the accuracy and efficiency of feature extraction. Figure 3 shown.
[0080] Step S3: Improve the neck network of YOLOv11 into a joint feature interaction network JFIN. By establishing a fully connected information flow channel between different feature layers to enrich the information interaction between non-adjacent layers, an improved neck network is obtained. Its network structure is as follows: Figure 4 shown.
[0081] The specific fusion method is as follows: The present invention marks the feature graphs extracted by Backbone in YOLOv11 as C1 to C5 from large to small scales, where C1 belongs to the initial stage of feature extraction and does not directly participate in the multi-scale fusion process. First, C3 and C4 are fused with mid-level features to start the multi-scale fusion sequence. In this process, not only are the features weighted and integrated, but also the high-resolution details of shallow features and the semantic information of deep features are fully shared through the cross-scale interaction mechanism:
[0082]
[0083] Then, a fully connected information flow channel is established between different feature layers to enrich the information interaction between non-adjacent layers:
[0084]
[0085] in, represents the output feature after fusion at position (i, j) in feature layer k, or represents the feature vector extracted from the original feature map, Used to model the feature interactions between C3 and C4, Represents the weight parameter and satisfies the normalization constraint
[0086] JFIN achieves full integration of multi-scale features by efficiently connecting shallow details with deep semantics, avoiding the weakening problem of feature information in layer-by-layer transmission. In addition, based on JFIN's multi-scale detection mechanism, a 160×160 resolution small target detection head is added to the network structure, further improving the perception of small targets.
[0087] Step S4: The designed GhostPlus lightweight module is introduced into the backbone network and neck network of YOLOv11. The traditional channel splicing operation is replaced by point-by-point addition operation, and the “front activation” strategy is introduced to achieve lightweight calculation of the model.
[0088] YOLOv11 widely uses channel concatenation (concat) operation for feature connection in module design. Although this method can integrate multiple layers of features, it will significantly increase the computational overhead in practical applications, as shown in the following formula, where x is the input feature, s represents the number of feature channels merged by addition operation, and Φ i (x) represents the feature map generated by different branches:
[0089] y=Concat([Φ1(x),Φ2(x),...,Φ s-1 (x)]) (9)
[0090] The GhostPlus module designed by the present invention uses the Add operation to replace the Concat operation. The Add operation replaces the traditional channel splicing operation with a point-by-point addition operation, which greatly reduces redundant calculations and improves the efficiency of feature connection:
[0091]
[0092] At the same time, the module introduces a "pre-activation" strategy, placing normalization and activation functions before convolution operations, effectively improving the efficiency of feature extraction and the stability of gradient propagation. In addition, the present invention also replaces the traditional ReLU activation function with the SiLU activation function. Compared with ReLU, SiLU has better nonlinear expression ability and can enhance the feature learning ability of the network without introducing significant additional computational overhead. This "pre-activation" design can effectively avoid the loss of feature information and improve the efficiency of gradient propagation in deep networks. The final optimized module is as follows Figure 5 As shown, wherein the dotted line represents the improved part.
[0093] Step S5: The constructed model is subjected to the global scoring-based pruning strategy LAMP to achieve model compression.
[0094] LAMP is a new scoring metric for global pruning that aims to automatically select layer-by-layer sparsity by minimizing output distortion at the model level without requiring additional hyperparameter tuning. LAMP introduces an importance score definition based on weight magnitude. For the u-th index in the weight tensor W, the LAMP score is defined as follows:
[0095]
[0096] Among them, W[u] represents the element with index u in the weight tensor, and the denominator is the sum of the squares of all weights from index u to the end.
[0097] The LAMP score measures the relative importance of the target weight among the remaining weights in its layer, ensuring that larger weights have higher importance scores and are retained first. Through this scoring criterion, LAMP can globally sort the weights of all layers and select the least important weights for pruning. In addition, LAMP has an important property: if the square value of the weight W[u] is greater than the square value of the weight W[v], then its corresponding LAMP score will also be higher:
[0098] (W[u]) 2 >(W[v]) 2 →score(u;W)>score(v;W) (12)
[0099] This feature ensures that the importance of weights is highly correlated with their magnitude during LAMP pruning, and the global pruning strategy automatically balances the layer-by-layer sparsity rate, thereby effectively reducing output distortion. The present invention prunes weights with lower scores based on the LAMP score, thereby achieving model compression and acceleration while minimizing the impact on model performance. This pruning strategy based on global scoring makes the selection of weights more accurate, ensuring a balance between performance and computational efficiency after model pruning. The comparison of the effects before and after pruning is as follows: Figure 6 shown.
[0100] Step S6: Use the FLEX-YOLO network as the detection model, train and verify it using the training set and the validation set to generate the final detection model.
[0101] The experimental environment is set on the Ubuntu 20.04.4 LTS system, based on the PyTorch 2.5.0 framework, and accelerated using CUDA 11.4. The hardware environment is NVIDIA RTX A6000 GPU with 48G video memory.
[0102] The initial learning rate of the model is set to 0.01, the momentum parameter is set to 0.937, the weight decay coefficient is 0.0005, and the batch size is set to 16.
[0103] The VisDrone2019 training set divided in step S1 is input into the FLEX-YOLO network, and the model is iteratively trained using the training data until the loss function converges. The training process lasts for 200 epochs. After the training is completed, the precision (Precision), recall (Recall) and mean average precision (mAP) of the validation set are used as the main evaluation indicators of the model performance, and the number of model parameters and GFLOPs (billion floating-point operations per second) are used as the measurement criteria for lightweight evaluation. The hyperparameters of the model are adjusted according to the verification results, and finally the detection model is generated. The specific definitions of the above performance evaluation indicators are as follows:
[0104]
[0105] In the above formula, true positives (TP) refer to the number of actual positive samples that are correctly identified as positive, and false positives (FP) refer to the number of negative samples that are misclassified as positive.
[0106]
[0107] In the above formula, false negatives (FN) refer to the number of samples that are actually positive but are misclassified as negative.
[0108]
[0109] In the above formula, AP represents the average precision of a single category, and mAP represents the average precision of all categories.
[0110] Step S7: Use the final trained detection model and the test set as input to test the FLEX-YOLO model.
[0111] Load the weight file trained in step S6 into the detection program, and use the VisDrone2019 test set divided in step S1 to fully test the FLEX-YOLO model. During the test, the model performs target detection on each drone aerial image in the test set, outputs detection results including target category, bounding box and confidence score, and superimposes the detection results on the original image in a visual form to intuitively show the actual detection effect of the model. At the same time, key performance indicators are calculated through the test set, including precision, recall, mean average precision (mAP), and the lightweight performance of the model is evaluated in combination with the number of model parameters and GFLOPs. These test results not only fully verify the detection performance and robustness of the FLEX-YOLO model in drone aerial photography scenes, but also provide important data support for the optimization and improvement of subsequent models.
[0112] In order to evaluate the performance of the improved model, the present invention conducted a comparative experiment on the model, and the experimental results are shown in Table 1.
[0113] Table 1 Comparison of detection methods on the VisDrone2019 validation set
[0114]
[0115] Experimental results show that FLEX-YOLO achieves 45.3% mAP on the VisDrone2019 validation set, surpassing most mainstream target detection algorithms and improving by 4.9% over the strongest baseline YOLOv11-S. With minimal parameters and low GFLOPs, FLEX-YOLO achieves an excellent balance between efficiency and performance. Similarly, compared with advanced small object detection methods, FLEX-YOLO performs best overall.
[0116] Example 2
[0117] like Figure 7 This embodiment relates to a UAV aerial photography small target detection device based on a FLEX-YOLO network, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the UAV aerial photography small target detection method based on a FLEX-YOLO network of Example 1.
[0118] Example 3
[0119] The present embodiment relates to a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for detecting small targets in drone aerial photography based on a FLEX-YOLO network of Embodiment 1 is implemented.
[0120] The above description is a specific embodiment of the present invention and the technical principles used. If the changes made according to the concept of the present invention do not exceed the spirit covered by the description and drawings, they should still fall within the scope of protection of the present invention.
Claims
1. A method for detecting small targets in drone aerial photography based on FLEX-YOLO network, characterized in that: The steps include: S1: Obtain the public drone aerial image dataset Visdrone2019 and perform preprocessing; S2: Build a FLEX-YOLO network model, which is based on the YOLOv11 network. The designed multi-path feedback perception module MPFP is introduced into the backbone network of the YOLOv11. The improved backbone network is obtained by combining the deep separable convolution DWConv, the double-layer routing attention BRA and the multi-layer perceptron MLP through the multi-path design. S3: The neck network of YOLOv11 is improved to the joint feature interaction network JFIN. By establishing a fully connected information flow channel between different feature layers to enrich the information interaction between non-adjacent layers, an improved neck network is obtained; S4: The designed GhostPlus lightweight module is introduced into the backbone network and neck network of YOLOv11. The traditional channel splicing operation is replaced by point-by-point addition operation, and the "front activation" strategy is introduced to achieve lightweight calculation of the model. S5: The constructed model is subjected to the global scoring-based pruning strategy LAMP to achieve model compression; S6: Use the FLEX-YOLO network as the detection model, train and verify it using the training set and validation set to generate the final detection model; S7: Use the final trained detection model and the test set as input to test the FLEX-YOLO model.
2. The method for detecting small targets in drone aerial photography based on FLEX-YOLO network according to claim 1, characterized in that: In step S1, the pre-processing step includes: S11: Convert the original annotation format to YOLO format, where each row of data represents a target information, including the target category number, the center position coordinates x and y of the target in the image, and the width w and height h of the target; S12: All images are cropped or filled to a fixed size pixel, while maintaining the aspect ratio of the image, and the unfilled area is processed with zero filling; S13: The data set is divided reasonably. The training set is used for model training, the validation set is used for parameter optimization and performance evaluation, and the test set is used for testing the final detection effect.
3. The method for detecting small targets in drone aerial photography based on FLEX-YOLO network according to claim 1, characterized in that: In step S2, the design method of the multi-path feedback perception module MPFP is: S21: DWConv is used instead of standard convolution. The standard convolution is decomposed into a depthwise convolution with a convolution kernel of 3×3, and a convolution operation is performed independently on each channel to extract spatial features, and a point-by-point convolution with a convolution kernel of 1×1. All channels are linearly combined to fuse the information between channels. S22: Introducing the double-layer routing attention BRA, the core of which is to retain the most relevant Top-k connections between each region and other regions, and to generate the adjacency matrix A r Perform the Top-k operation row by row to generate the routing index matrix I r , where each row contains the index of the current region and its k most relevant regions to avoid redundant calculations; I r =TopKIndex(A r ) (1) S23: Add a multi-layer perceptron (MLP) to enhance the expressiveness of channel features through nonlinear transformation.
4. The method for detecting small targets in drone aerial photography based on FLEX-YOLO network according to claim 1, characterized in that: In step S3, the design method of the joint feature interaction network JFIN is: S31: The feature maps extracted by Backbone in YOLOv11 are marked from C1 to C5 according to the scale from large to small. C1 belongs to the initial stage of feature extraction and does not directly participate in the multi-scale fusion process. First, the middle-level features of C3 and C4 are fused to start the multi-scale fusion sequence: S32: A fully connected information flow channel is established between different feature layers to enrich the information interaction between non-adjacent layers: in, represents the output feature after fusion at position i, j in feature layer k, or represents the feature vector extracted from the original feature map, Used to model the feature interactions between C3 and C4, Represents the weight parameter and satisfies the normalization constraint S33: Based on the multi-scale detection mechanism of JFIN, a small target detection head with a resolution of 160×160 is added to the network structure.
5. The method for detecting small targets in drone aerial photography based on FLEX-YOLO network according to claim 1, characterized in that: In step S4, the design method of the GhostPlus lightweight module is: S41: Use Add operation to replace Concat operation, and replace the traditional channel splicing operation with point-by-point addition operation: Among them, s represents the number of feature channels merged by the addition operation; S42: The "pre-activation" strategy is introduced, which places the normalization and activation function before the convolution operation, and replaces the traditional ReLU activation function with the SiLU activation function with better nonlinear expression ability.
6. The method for detecting small targets in drone aerial photography based on FLEX-YOLO network according to claim 1, characterized in that: In step S5, the implementation method of the LAMP pruning algorithm is: S51: Based on the weight magnitude, an importance score definition is introduced. For the u-th index in the weight tensor W, the LAMP score is defined as follows: Among them, W[u] represents the element with index u in the weight tensor, and the denominator is the sum of the squares of all weights from index u to the end; S52: Based on the LAMP score, the relative importance of the target weight in the remaining weights of its layer is measured, and the weights with relatively low LAMP scores are pruned according to the set pruning rate.
7. The method for detecting small targets in drone aerial photography based on FLEX-YOLO network according to claim 1, characterized in that: In step S6, the FLEX-YOLO network is used as the detection model, and the training set and the verification set are used to train and verify the detection model, and the steps of generating the final detection model include: S61: The experimental environment is set on the Ubuntu 20.04.4 LTS system, based on the PyTorch 2.5.0 framework, and accelerated using CUDA 11.
4. The hardware environment is NVIDIA RTX A6000 GPU with 48G video memory; S62: The initial learning rate of the model is set to 0.01, the momentum parameter is set to 0.937, the weight decay coefficient is 0.0005, and the batch size is set to 16; S63: Input the preprocessed training set into the FLEX-YOLO network, calculate the classification loss, positioning loss and confidence loss through forward propagation, and then perform back propagation and update the parameters. The training process continues for 200 epochs; S64: After each epoch, the validation set is input into the current model for inference, and the performance indicators are evaluated in real time. After the training is completed, the model weight with the highest mAP of the validation set is selected as the final detection model.
8. A UAV aerial photography small target detection device based on FLEX-YOLO network, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the unmanned aerial vehicle aerial photography small target detection method based on the FLEX-YOLO network as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by the processor, the method for detecting small targets in drone aerial photography based on the FLEX-YOLO network described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Foot sole image recognition method based on convolutional neural network
CN115294599A
Target detection method for YOLOv8 partial convolutional network
CN117173395A
Lightweight steel surface defect detection method based on improved YOLOv8n
CN118196529A
Neural architecture searching method for long-tail data set
CN119047519A
AU2020102091A4
Cited By
Aerial photography real-time small target detection method based on model compression and hardware acceleration
CN121545067A