Unmanned aerial vehicle aerial photography small target detection method based on branch weighted fusion

By constructing the PBE-YOLO network model, combining the BWAFPN module, C2f-PKBFormer module and EWD loss function, the accuracy and missed detection problems of small target detection from the perspective of the drone are solved, and efficient target recognition and positioning are achieved.

CN120339884APending Publication Date: 2025-07-18YANSHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510492325.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing drone target detection methods are difficult to effectively capture and identify small targets from the perspective of drones, and conventional methods are prone to missed detection, resulting in poor detection results.

Method used

The aerial small object detection method of branch-weighted fusion is adopted. By constructing a PBE-YOLO network model, combining the BWAFPN module, C2f-PKBFormer module and EWD loss function, feature fusion and target positioning capabilities are enhanced and small object detection is optimized.

Benefits of technology

It significantly improves the detection accuracy of small targets from the perspective of the drone, reduces the missed detection rate, improves the feature capture capability of small targets, and optimizes the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339884A_ABST
    Figure CN120339884A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle aerial photography small target detection method based on branch weighted fusion, and belongs to the technical field of image processing, and the method comprises the following steps: S1, setting the size of an input image needed by training to be 640 * 640, guaranteeing the consistency of model input data, and carrying out the Mosia data enhancement of a picture in a data set; s2, a PBE-YOLO network model is constructed based on the YOLOv8 model, and model training is carried out; the PBE-YOLO network model comprises a backbone network used for feature extraction, and a branch weighted auxiliary fusion feature pyramid network BWAFPN module, a C2f-PKBFormer module and an exponential normalized Gaussian Wasserstein distance loss function EWD module which are connected to the backbone network; and S3, model testing: inputting a test image into the trained PBE-YOLO network model, carrying out target detection, and outputting the category and bounding box position of each target by the model, so that the detection precision of the small target under the view angle of the unmanned aerial vehicle can be effectively improved, the feature capture capability of the small target can be improved, and the detection effect of the small target can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for detecting small targets in UAV aerial photography based on branch weighted fusion. Background Art

[0002] With the rapid development of deep learning, object detection based on UAV aerial photography images has shown great application potential. As a lightweight, convenient, and easy-to-operate device, UAVs can adapt to the complex and dangerous environments of various tasks and are widely used in industries such as traffic supervision, environmental monitoring, and emergency rescue and disaster relief. This benefits from the object detection technology of UAVs, which uses the image acquisition devices carried by UAVs and, through advanced image processing technologies and object detection algorithms, quickly and accurately identifies targets on the ground or in the air. Usually, object detection is mainly used to detect medium or large objects in nature, but the core of UAV object detection is to quickly and accurately identify smaller targets such as people and vehicles from the images collected from the complex and changing perspectives of UAVs and clarify their specific positions in the images in the form of bounding boxes. This requires a high level of accuracy for the algorithm and ensures that accurate detection results can be obtained under the conditions of the UAV flying quickly and the scene constantly changing.

[0003] However, due to the high position of the UAV camera and the long shooting distance, and the fact that the flight altitude of the UAV changes during use, the images collected often have the problem of large target scale changes, especially there are many small-sized targets, which increases the difficulty of the UAV object detection task. Although significant progress has been made in object detection methods based on deep learning, the previous models are mainly object detection methods for natural scenes. These methods work well in detecting medium or large targets but are difficult to handle complex small targets in the images. Therefore, conventional detection methods lack the ability to capture the features of small targets in object detection from the UAV perspective, have insufficient information processing and fusion capabilities for small targets in the feature fusion stage, and the traditional IoU-based loss function is sensitive to the position offset of the bounding boxes of small targets, prone to the problem of missed detection of small targets, resulting in poor detection effects, which undoubtedly poses a huge challenge to the accurate identification of UAVs and the positioning of small targets.

[0004] Therefore, there is a need for a method for detecting small targets in UAV aerial photography based on branch weighted fusion that can effectively improve the detection accuracy of small targets from the UAV perspective, enhance the ability to capture the features of small targets, and avoid missed detection of small targets. Summary of the Invention

[0005] The object of the present invention is to provide a method for detecting small targets in UAV aerial photography based on branch weighted fusion, which can effectively improve the detection accuracy of small targets from the UAV perspective, enhance the ability to capture small target features, effectively avoid missing detection of small targets, and optimize the detection effect of small targets.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A method for detecting small targets in UAV aerial photography based on branch weighted fusion includes the following steps:

[0008] Step S1: Data preprocessing: First, set the size of the input image required for training to 640×640 to ensure the consistency of the model input data, and then perform Mosiac data augmentation on the pictures in the dataset;

[0009] Step S2: Construct a PBE-YOLO network model based on the YOLOv8 model and perform model training; the PBE-YOLO network model includes a backbone network for feature extraction and a branch weighted auxiliary fusion feature pyramid network BWAFPN module, a C2f-PKBFormer module, and an exponential normalized Gaussian Wasserstein distance loss function EWD module connected to the backbone network;

[0010] Step S3: Model testing: Input the test image into the trained PBE-YOLO network model for object detection, and the model outputs the category and bounding box position of each object.

[0011] A further improvement of the technical solution of the present invention lies in that step S1 includes the following steps:

[0012] Step S101: Crop or pad the pictures in the VisDrone dataset required for training so that the size of the input image is 640×640;

[0013] Step S102: Perform Mosiac data augmentation on the pictures in the dataset. Randomly select four images from the dataset, randomly crop each image to obtain four image patches, and then splice the four image patches to form a new image.

[0014] A further improvement of the technical solution of the present invention lies in that step S2 includes the following steps:

[0015] Step S201: Construct a Branch Weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module, including a Shallow Auxiliary Weighted Fusion module (SWAF) for initially auxiliarily fusing multi-scale features output by the backbone network in the shallow layer of the neck, a Deep Auxiliary Weighted Fusion module (AWAF) for transmitting diverse gradient information to the detection head in the deep layer of the neck, and a Small Object Detection layer (SODL) for enhancing the learning and fusion ability of small object feature information; the Branch Weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module is tightly connected after the backbone network, and through the collaborative action of the Shallow Auxiliary Weighted Fusion module (SWAF) and the Deep Auxiliary Weighted Fusion module, the model's feature expression and fusion ability for small objects are enhanced;

[0016] Step S202: Construct a C2f-PKBFormer module to enhance the model's ability to capture local context information by arranging multiple depth convolutions of different sizes in parallel;

[0017] Step S203: Construct an Exponential Normalized Gaussian Wasserstein Distance Loss function (EWD) module and apply it to the loss function of the detection head. By Gaussian distribution modeling and Wasserstein distance calculation, the model's localization ability for small objects is enhanced; during the object detection process, the detection head predicts the category and location of the objects in the image based on the feature information extracted and processed by the previous network layers. The EWD loss function measures the difference between the prediction result and the ground truth label, optimizing the training process of the model;

[0018] Step S204: Input the processed data of the VisDrone dataset and train the PBE-YOLO network model.

[0019] A further improvement of the technical solution of the present invention lies in: in Step S201, the Shallow Auxiliary Weighted Fusion module (SWAF) utilizes the feature information output by different levels of the backbone network. By fusing the high-resolution feature P n-1 with the same-resolution feature P n output by the backbone network and the low-resolution feature P' n+1 output from the previous SWAF module, the optimization of the features is achieved.

[0020] A further improvement of the technical solution of the present invention lies in: during the shallow fusion process, the high-resolution feature P n-1 passes through a downsampling module to make its resolution and number of channels consistent with this layer. At the same time, a 1×1 standard convolution is used to control the channels of the output w n of the backbone network at the same resolution to adjust the dimension and information distribution of the feature channels. The low-resolution feature P' n+1Through the upsampling module, the resolution and the number of channels are made the same as those of the other two features. Finally, through fast normalization fusion, the high-resolution feature P n-1 , the output P n , and the low-resolution feature P' n+1 are weighted and fused. The shallow fusion formula is as follows:

[0021]

[0022] Among them, w1, w2, and w3 are learnable weights, which are automatically learned and adjusted by the network during the training process; ∈ takes the value of 0.0001 to avoid the error of division by zero during the calculation process; Down(·) represents the downsampling operation, which reduces the resolution of the feature map, reduces the amount of data while retaining the main feature information; C(·) represents the operation of using 1×1 convolution for channel control, which adjusts the channel dimension of the feature map, filters and integrates channel information; Up(·) represents the upsampling operation, which increases the resolution of the feature map, enabling feature maps with different resolutions to be fused at the same scale.

[0023] A further improvement of the technical solution of the present invention lies in: in step S201, the deep auxiliary weighted fusion module AWAF further enhances the interaction and utilization between feature layers, and fuses the high-resolution feature P' n-1 , the low-resolution feature P' n+1 , the same-resolution feature P' n transferred from the shallow auxiliary weighted fusion module SWAF with the high-resolution feature P″ n-1 output by the previous AWAF module; before fusion, the high-resolution feature P' n-1 and the high-resolution feature P″ n-1 output by the previous AWAF module are downsampled, and the low-resolution feature P' n+1 is upsampled, so that the four-part feature maps input to the fusion module are consistent in resolution and the number of channels.

[0024] A further improvement of the technical solution of the present invention lies in: the deep fusion formula is as follows:

[0025]

[0026] Among them, w1, w2, w3, and w4 are learnable weights, which are optimized and adjusted by the network according to data characteristics and model performance during the training process; ∈ takes the value of 0.0001 to avoid the error of division by zero during the calculation process; Down(·) represents the downsampling operation, C(·) represents the operation of using 1×1 convolution for channel control, and Up(·) represents the upsampling operation.

[0027] A further improvement of the technical solution of the present invention lies in: in step S202, the PKBFormer module in the C2f-PKBFormer module includes two cooperating residual sub-modules. One of the residual modules consists of a normalization operation and a multi-core inverted bottleneck module PKB, and its formula is expressed as:

[0028] Y = PKB(Norm(X)) + X

[0029] Among them, Norm(·) represents the normalization operation, and through normalization, the distribution of the input data is standardized; PKB(·) represents the multi-core inverted bottleneck module;

[0030] The other residual module consists of a two-layer MLP with a non-linear activation function, and its formula is expressed as:

[0031] Z = StarRelu(Norm(Y)W1)W2 + Y

[0032] Among them, Star Relu is the activation function; W1 and W2 are the learnable parameters of the MLP.

[0033] A further improvement of the technical solution of the present invention lies in: the multi-core inverted bottleneck module PKB uses multiple parallel deep convolutional kernels to capture multi-scale texture features: first, through a linear convolutional module, the dimension of the input feature map is expanded, and then a group of parallel deep convolutions are used to capture context information across multiple scales. The kernel sizes of the deep convolutions are set to 5×5, 7×7, 9×9, and 11×11 respectively; the formula of the multi-core inverted bottleneck module PKB is expressed as:

[0034] I′ = StarRelu(PWConv(I))

[0035] I″ = DWConv 5×5 (I′) + DWConv 7×7 (I′) + DWConv 9×9 (I′) + DWConv 11×11 (I′) + Id(I′)

[0036] I″′ = PWConv(I″)

[0037] Among them, PWConv(·) represents the linear convolutional module, which is used to expand the dimension of the feature map and perform preliminary feature extraction; DWConv n×n (·) represents the depthwise separable convolution with different kernel sizes, which is used to extract texture features at different scales; Id(·) represents the identity mapping, which is used to directly output the input, fuse with the convolved features while retaining the original feature information, and enrich the feature representation.

[0038] A further improvement of the technical solution of the present invention lies in: in step S203, the exponential normalized Gaussian Wasserstein distance loss function EWD module models the bounding box as a two-dimensional Gaussian distribution, uses the Wasserstein distance between two Gaussian distributions to measure the similarity between bounding boxes, and performs exponential normalization transformation to obtain a more reasonable loss value. The calculation formula of the exponential normalized Gaussian Wasserstein distance loss function EWD is as follows:

[0039]

[0040] The final loss function is:

[0041]

[0042] Loss = α × L CIoU +(1 - α) × L EW

[0043] Wherein,

[0044]

[0045] Wherein, C represents a constant closely related to the data set, taking the average size of the targets in the data set; α represents a balance factor, used to adjust the weight ratio of and in the total loss function. When there are more small targets in the data set, the value of α can be lowered to increase the proportion of in the loss function. In this model, α is set to 0.5.

[0046] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:

[0047] The small target detection method for UAV aerial photography based on branch weighted fusion of the present invention can effectively improve the detection accuracy of small targets from the UAV perspective, enhance the small target feature capture ability, effectively avoid missing detection of small targets, and optimize the detection effect of small targets.

[0048] By designing the branch weighted auxiliary fusion feature pyramid network BWAFPN, C2f - PKBFormer module and exponential normalized Gaussian Wasserstein distance loss function EWD, the present invention significantly improves the detection accuracy of small targets in UAV aerial photography images.

[0049] The branch weighted auxiliary fusion feature pyramid network BWAFPN adopted by the present invention strengthens the network feature fusion ability through the synergistic effect of the shallow weighted auxiliary fusion module and the deep weighted auxiliary fusion module. On the basis of the three detection layers of the baseline model, a detection layer SODL more suitable for small targets is added, which solves the problem of feature loss caused by insufficient fusion of small target features in the shallow layer and effectively improves the ability of the model to detect small targets.

[0050] The PKBFormer module adopted in the present invention replaces the Bottleneck part in the baseline model C2f. First, the input feature channels are expanded, and then operations are performed on depth convolutions of different sizes arranged in parallel, enhancing the model's ability to understand features of different scales. The features are fused along the channel dimension, thereby enabling the collection of local context information.

[0051] In the loss function part of the present invention, a balance factor α and an exponential normalized Gaussian Wasserstein distance loss function EWD are introduced. First, the bounding boxes are modeled as two-dimensional Gaussian distributions, and the Wasserstein distance is used to evaluate the similarity between two bounding boxes. This loss function can solve the problem that the baseline model is more sensitive to the position change of the bounding boxes of small targets during detection, better measure the differences between small targets, and improve the detection accuracy of small targets. Description of the Drawings

[0052] Figure 1 is the network structure diagram of the present invention;

[0053] Figure 2 is the structural flowchart of the shallow auxiliary weighted fusion module of the present invention;

[0054] Figure 3 is the structural flowchart of the deep auxiliary weighted fusion module of the present invention;

[0055] Figure 4 is the structural flowchart of the C2f-PKBFormer module of the present invention;

[0056] Figure 5 is the result comparison diagram of the confusion matrices of the baseline model and the PBE-YOLO network model on the VisDrone test dataset;

[0057] Figure 6 is the different scenario comparison diagram of the baseline model and the PBE-YOLO network model on the VisDrone test dataset. Detailed Embodiments

[0058] The following further describes the present invention in detail with reference to the embodiments:

[0059] The present invention provides a method for detecting small targets in UAV aerial photography based on branch weighted fusion, including the following steps:

[0060] Step S1: Data preprocessing: First, set the size of the input image required for training to 640×640 to ensure the consistency of the model input data, and then perform Mosiac data augmentation on the pictures in the dataset;

[0061] Specifically, it includes the following steps:

[0062] Step S101: Crop or pad the images in the VisDrone dataset required for training so that the size of the input images is all 640×640;

[0063] Step S102: Perform Mosiac data augmentation on the images in the dataset. Randomly select four images from the dataset, randomly crop each image to obtain four image patches, and then splice the four image patches to form a new image;

[0064] Step S2: As Figures 1-4 shown, build the PBE - YOLO network model based on the YOLOv8 model and perform model training; The PBE - YOLO network model includes a backbone network for feature extraction and a Branch - weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module, a C2f - PKBFormer module, and an Exponential Normalized Gaussian Wasserstein Distance Loss Function (EWD) module connected to the backbone network;

[0065] Specifically, it includes the following steps:

[0066] Step S201: Build the Branch - weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module, which includes a Shallow - layer Auxiliary Weighted Fusion Module (SWAF) for initially auxiliary - fusing the multi - scale features output by the backbone network in the shallow layer of the neck, a Deep - layer Auxiliary Weighted Fusion Module (AWAF) for transmitting diverse gradient information to the detection head in the deep layer of the neck, and a Small - Object Detection Layer (SODL) for strengthening the learning and fusion ability of small - object feature information; The Branch - weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module is closely connected after the backbone network. Through the collaborative action of the Shallow - layer Auxiliary Weighted Fusion Module (SWAF) and the Deep - layer Auxiliary Weighted Fusion Module, the model's feature expression and fusion ability for small objects are enhanced;

[0067] Among them, the Shallow - layer Auxiliary Weighted Fusion Module (SWAF) utilizes the feature information output by different levels of the backbone network. By fusing the high - resolution feature P n-1 with the same - resolution feature P n output by the backbone network and the low - resolution feature P′ n+1 output from the previous SWAF module, the optimization of the features is achieved;

[0068] In the shallow - layer fusion process, the high - resolution feature P n-1 passes through a down - sampling module to make its resolution and number of channels consistent with this layer. At the same time, a 1×1 standard convolution is used to control the channels of the output P n of the backbone network at the same resolution to adjust the dimension and information distribution of the feature channels. The low - resolution feature P′n+1 Then, through the upsampling module, the resolution and the number of channels requirements are made the same as those of the other two features. Finally, through fast normalization fusion, the high-resolution feature P n-1 and the output P n and the low-resolution feature P' n+1 These three parts of features are weighted and fused, and the shallow fusion formula is as follows:

[0069]

[0070] Among them, w1, w2, and w3 are learnable weights, which are automatically learned and adjusted by the network during the training process. They can assign reasonable weights according to the importance of different features to achieve more effective feature fusion; ∈ takes the value of 0.0001, which is a very small value. Its purpose is to avoid the error of division by zero during the calculation process and ensure the stability and accuracy of the calculation; Down(·) represents the downsampling operation, which reduces the resolution of the feature map through a specific algorithm, reduces the amount of data while retaining the main feature information; C(·) represents the operation of channel control using 1×1 convolution, which can adjust the channel dimension of the feature map, screen and integrate channel information; Up(·) represents the upsampling operation, which improves the resolution of the feature map through methods such as interpolation, so that feature maps with different resolutions can be fused at the same scale; Through such a fusion method, the SWAF module can retain rich spatial information and integrate high-level semantic information, providing more valuable feature data for the subsequent network layers;

[0071] The deep auxiliary weighted fusion module AWAF can further enhance the interactive utilization between feature layers. The high-resolution feature P' passed from the shallow auxiliary weighted fusion module SWAF n-1 and the low-resolution feature P' n+1 and the same-resolution feature P' n are fused with the high-resolution feature P″ of the output of the previous AWAF module n-1 These four parts of features; before fusion, the high-resolution feature P' n-1 and the high-resolution feature P″ of the output of the previous AWAF module n-1 , need to perform the downsampling operation, while the low-resolution feature P' n+1 needs to perform the upsampling operation, so that the four parts of feature maps input to the fusion module are consistent in resolution and the number of channels;

[0072] Through this multi-feature fusion method, the interactive utilization of feature information can be further enhanced, the information loss during the feature transmission process can be compensated, and the output layer can obtain multi-scale output information spanning three resolutions, thus significantly improving the fusion effect of the neck network. The deep fusion formula is as follows:

[0073]

[0074] Among them, w1, w2, w3, and w4 are learnable weights, which are optimized and adjusted by the network according to data characteristics and model performance during the training process; ∈ takes a value of 0.0001, and its purpose is to avoid the error situation of division by zero during the calculation process and ensure the stability and accuracy of the calculation; the operation meanings of Down(·), C(·), and Up(·) are the same as those in the SWAF module. Down(·) represents the downsampling operation, C(·) represents the operation of using 1×1 convolution for channel control, and Up(·) represents the upsampling operation;

[0075] In drone aerial images, small targets are likely to be missed due to their small size and unobvious features. In order to retain the shallow features of higher resolution, the small target detection layer SODL introduces an additional small target detection layer in the three-layer pyramid architecture of the neck network through special design and parameter configuration. This will generate a larger-scale feature map that can distinguish the finer features of small targets. The dedicated small target detection layer can better adapt to the features of small targets, capture the detailed information of small targets, and can perform targeted processing on the features fused by the SWAF module and the AWAF module, paying more attention to the feature information of small targets, thereby effectively reducing the missed detection rate of small targets and improving the detection accuracy of the model for small targets;

[0076] Step S202: Construct the C2f-PKBFormer module. By arranging multiple depth convolutions of different sizes in parallel, the ability of the model to capture local context information is enhanced; the C2f-PKBFormer module is used to replace the C2f module in the baseline model. The design of this module aims to enhance the ability of the model to capture context information across multiple scales to cope with the complex backgrounds and diverse target scales in drone aerial images;

[0077] The PKBFormer module in the C2f-PKBFormer module includes two cooperating residual sub-modules. One of the residual modules consists of a normalization operation and a multi-core inverted bottleneck module PKB, and its formula is expressed as:

[0078] Y = PKB(Norm(X)) + X

[0079] Among them, Norm(·) represents the normalization operation, which plays a crucial role in the model training process. Through normalization, the distribution of the input data can be standardized, so that the mean and variance of the data are maintained within a certain range, effectively alleviating the problem of gradient disappearance or explosion and ensuring the stability of the model training process; PKB(·) represents the multi-core inverted bottleneck module, which is the key component for realizing multi-scale feature extraction;

[0080] The multi-core inverted bottleneck module PKB, as a core component of PKBFormer, adopts multiple depth convolution kernels arranged in parallel to capture multi-scale texture features: First, through a linear convolution module, this module expands the dimension of the input feature map, providing more feature dimensions for the network to better capture the detailed features of objects in the image. Then, a set of parallel depth convolutions are used to capture context information across multiple scales. The kernel sizes of these depth convolutions are set to 5×5, 7×7, 9×9, and 11×11 respectively. Different-sized convolution kernels can sense feature information at different scales. Among them, the 5×5 convolution kernel is suitable for capturing smaller-scale texture features, while the 11×11 convolution kernel is better at capturing larger-scale context information. Through this multi-scale convolution method, the PKB module can capture extensive context information within a larger scale range, effectively solving the problem of difficult small target feature extraction in the complex background of drones.

[0081] The formula of the multi-core inverted bottleneck module PKB is expressed as:

[0082] I′ = StarRelu(PWConv(I))

[0083] I″ = DWConv 5×5 (I′) + DWConv 7×7 (I′) + DWConv 9×9 (i′) + DWConv 11×11 (I′) + Id(i′)

[0084] i″′ = PWConv(i″)

[0085] Among them, PWConv(·) represents the linear convolution module, which is used for dimension expansion and preliminary feature extraction of the feature map; DWConv n×n (·) represents depthwise separable convolutions with different kernel sizes, which are used to extract texture features at different scales; Id(·) represents the identity mapping, which is used to directly output the input, fuse with the convolved features while retaining the original feature information, and enrich the feature representation;

[0086] Another residual module is mainly composed of a two-layer MLP with a non-linear activation function, and its formula is expressed as:

[0087] Z = StarRelu(Norm(Y)W1)W2 + Y

[0088] Among them, Star Relu is a specially designed activation function that can effectively alleviate the problem of distribution shift, enabling the model to better learn the features of data during the training process; W1 and W2 are the learnable parameters of the MLP, which are continuously optimized and adjusted during the model training process to adapt to different data features and task requirements;

[0089] By stacking two cooperating residual sub-modules, the C2f-PKBFormer module can effectively capture more complex features, enhance the feature expression ability of the model, and thus better capture the information of multi-scale targets;

[0090] Step S203: Construct an exponential normalized Gaussian Wasserstein distance loss function EWD module, which is applied to the loss function of the detection head. By Gaussian distribution modeling and Wasserstein distance calculation, it enhances the model's ability to locate small targets, solves the problem that the traditional IoU loss is sensitive to the position deviation of small targets, and is used for the final object detection task; during the object detection process, the detection head predicts the category and position of the targets in the image based on the feature information extracted and processed by the previous network layers. The introduction of the EWD loss function can better measure the difference between the prediction result and the true label, and optimize the model training process;

[0091] The exponential normalized Gaussian Wasserstein distance loss function EWD module models the bounding box as a two-dimensional Gaussian distribution, uses the Wasserstein distance between two Gaussian distributions to measure the similarity between bounding boxes, and performs exponential normalization transformation to obtain a more reasonable loss value. The calculation formula of the exponential normalized Gaussian Wasserstein distance loss function EWD is as follows:

[0092]

[0093] The final loss function is:

[0094]

[0095] Loss = α × L CIoU +(1 - α) × L EW

[0096] Among them,

[0097]

[0098] Among them, C represents a constant closely related to the dataset, which is taken as the average size of the targets in the dataset; α represents a balance factor used to adjust the weight ratio of and in the total loss function. When there are more small targets in the dataset, the value of α can be appropriately decreased to increase the proportion in the loss function, thereby achieving a better regression effect. In this model, α is set to 0.5. Through this optimized loss function, the model can more accurately learn the features and location information of the targets during the training process, improving the detection accuracy of small targets.

[0099] Step S204: Input the processed data of the VisDrone dataset and train the PBE-YOLO network model.

[0100] Step S3: Model testing: Input the test images into the trained PBE-YOLO network model for object detection, and the model outputs the category and bounding box position of each object.

[0101] The present invention is trained and tested on the VisDrone dataset, and experiments verify the performance of the model and the effectiveness of the module. Compared with the baseline model YOLOv8, the mAP value of the PBE-YOLO network model is increased by 9.8%, and the AP values of all categories have been improved to varying degrees. Among them, the improvement amplitudes of the AP values of the Pedestrain, People, and Motor categories are all above 10%. The PBE-YOLO network model exceeds models such as YOLOv5s, Faster R-CNN, and CenterNet in terms of mAP, which indicates that the present invention has better performance in comparison with YOLO series models and other classical models. The PBE-YOLO network model also exceeds the baseline model YOLOv8s and improved models such as YOLO-UAV and LUD-YOLO in terms of mAP, which shows the competitiveness of the PBE-YOLO network model in the UAV object detection task. In terms of the overall average precision (mAP), the mAP value of the PBE-YOLO network model is 48.3%, which is higher than other models in the table, indicating that PBE-YOLO performs excellently in the overall performance of integrating multiple categories.

[0102] Specifically, the VisDrone dataset is used for training and testing to test the training effect of this method. As shown in Table 1, it is a comparison of the results of different algorithms on the VisDrone test dataset. From the comparison results in Table 1, it can be seen that the PBE-YOLO network model proposed by the present invention has achieved the best performance in terms of the mAP index and the detection results of various objects, and the improvement of the model performance also confirms the effectiveness of the components proposed by the present invention.

[0103] Table 1 Object Detection Results of PBE-YOLO in the VisDrone Test Set

[0104]

[0105] As Figure 5 shown, it presents a comparison of the results of the confusion matrices of the baseline model and the PBE-YOLO network model on the VisDrone test dataset. The rows of the confusion matrix represent the true classes, the columns represent the predicted classes, the values in the diagonal region represent the proportion of correctly predicted classes, and the classes in other regions represent the proportion of incorrectly predicted classes. The diagonal color of PBE-YOLO is darker than that of YOLOv8s, indicating that the ability of our model to correctly predict classes has been improved compared to YOLOv8s. The model has relatively good recognition effects for cars, buses, and motorcycles; however, for small objects such as bicycles and awning tricycles, the proportion judged as background is higher, and the improved model reduces the missed detection rate of these classes.

[0106] To further demonstrate the prediction effect, four representative scenarios in the VisDrone2019 dataset, namely streets, nights, basketball courts, and intersections, are selected as experimental data to evaluate the detection situations of the baseline model and the PBE-YOLO network model. These scenarios contain a large number of different small objects, and these scenarios are suitable for inference experiments. The detection effects of the two models are as Figure 6 shown.

[0107] Compared with the YOLOv8s model, the PBE-YOLO network model can detect more small targets and distant targets under various shooting angles and complex and dense backgrounds, showing superior detection performance. In a complex background environment, when the features such as the color and shape of the target object are similar to the background, by comparing the images, the baseline model YOLOv8s misidentifies the background as the target. At the same time, due to the dense arrangement of targets in the image, many small and distant targets are missed. The PBE-YOLO network model enhances the network's feature extraction ability for small targets and can more accurately identify targets in complex environments and backgrounds.

[0108] Generally speaking, compared with YOLOv8s, the PBE-YOLO network model has more obvious advantages when dealing with the drone image target detection task. The PBE-YOLO network model is less affected by external conditions, has strong detection capabilities for small objects, distant objects, and objects in complex backgrounds, can effectively avoid missed detections and false detections, shows good generalization ability, and can meet the requirements of actual tasks.

[0109] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will know that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A method for detecting small targets in UAV aerial photography based on branch weighted fusion, characterized in that It includes the following steps: Step S1: Data preprocessing: First, set the size of the input images required for training to 640×640 to ensure the consistency of the model input data, and then perform Mosiac data augmentation on the images in the dataset; Step S2: Build a PBE-YOLO network model based on the YOLOv8 model and perform model training; the PBE-YOLO network model includes a backbone network for feature extraction and a Branch Weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module, a C2f-PKBFormer module, and an Exponential Normalized Gaussian Wasserstein Distance Loss Function (EWD) module connected to the backbone network; Step S3: Model testing: Input the test images into the trained PBE-YOLO network model for object detection, and the model outputs the category and bounding box position of each object.

2. The method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 1, characterized in that: The said Step S1 includes the following steps: Step S101: Crop or pad the images in the VisDrone dataset required for training so that the size of the input images is all 640×640; Step S102: Perform Mosiac data augmentation on the images in the dataset. Randomly select four images from the dataset, randomly crop each image to obtain four image patches, and then splice the four image patches to form a new image.

3. The method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 2, characterized in that: The said Step S2 includes the following steps: Step S201: Build a Branch Weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module, including a Shallow Auxiliary Weighted Fusion Module (SWAF) for initially auxiliary fusing the multi-scale features output by the backbone network in the shallow layer of the neck, a Deep Auxiliary Weighted Fusion Module (AWAF) for transmitting diverse gradient information to the detection head in the deep layer of the neck, and a Small Object Detection Layer (SODL) for strengthening the learning and fusion ability of small object feature information; the Branch Weighted Auxiliary Fusion Feature Pyramid Network (BWAFPN) module is tightly connected after the backbone network, and through the collaborative action of the Shallow Auxiliary Weighted Fusion Module (SWAF) and the Deep Auxiliary Weighted Fusion Module, the model's feature expression and fusion ability for small objects are enhanced; Step S202: Build a C2f-PKBFormer module to enhance the model's ability to capture local context information by arranging multiple depth convolutions of different sizes in parallel; Step S203: Build an Exponential Normalized Gaussian Wasserstein Distance Loss Function (EWD) module and apply it to the loss function of the detection head. By Gaussian distribution modeling and Wasserstein distance calculation, the model's localization ability for small objects is enhanced; during the object detection process, the detection head predicts the category and position of the objects in the image based on the feature information extracted and processed by the previous network layers, and the EWD loss function measures the difference between the prediction result and the true label, optimizing the model training process; Step S204: Input the data of the processed VisDrone dataset and perform training on the PBE-YOLO network model.

4. A method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 3, characterized in that: In step S201, the Shallow Auxiliary Weighted Fusion module SWAF utilizes the feature information output by different levels of the backbone network. By taking the high-resolution feature P n-1 and the same-resolution feature P n output by the backbone network, and the low-resolution feature P' n+1 output from the previous SWAF module for fusion, the optimization of the features is achieved.

5. The method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 4, wherein: During the shallow fusion process, the high-resolution feature P n-1 passes through the downsampling module to make its resolution and number of channels consistent with this layer. At the same time, the output P of the backbone network at the same resolution is processed using a standard 1×1 convolution n for channel control to adjust the dimension and information distribution of the feature channels. The low-resolution feature P' n+1 passes through the upsampling module to meet the requirements of the same resolution and number of channels as the other two features. Finally, through fast normalization fusion, the high-resolution feature P n-1 , the output P n , and the low-resolution feature P' n+1 are weighted and fused. The shallow fusion formula is as follows: Among them, w1, w2, and w3 are learnable weights that are automatically learned and adjusted by the network during the training process; ∈ takes a value of 0.0001 to avoid the error situation of division by zero during the calculation process; Down(·) represents the downsampling operation, which reduces the resolution of the feature map, reduces the amount of data while retaining the main feature information; C(·) represents the operation of using 1×1 convolution for channel control, which adjusts the channel dimension of the feature map, filters and integrates channel information; Up(·) represents the upsampling operation, which increases the resolution of the feature map, enabling feature maps of different resolutions to be fused at the same scale.

6. The method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 5, wherein: In step S201, the deep auxiliary weighted fusion module AWAF further enhances the interaction and utilization between feature layers, and fuses the high-resolution feature P′ passed from the shallow auxiliary weighted fusion module SWAF n-1 , the low-resolution feature P′ n+1 , the same-resolution feature P′ n with the high-resolution feature P″ output by the previous AWAF module n-1 ; before fusion, the high-resolution feature P′ n-1 and the high-resolution feature P″ output by the previous AWAF module n-1 are downsampled, and the low-resolution feature P′ n+1 is upsampled so that the four-part feature maps input to the fusion module are consistent in resolution and number of channels.

7. The method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 6, wherein: The formula for deep fusion is as follows: Among them, w1, w2, w3, and w4 are learnable weights that are optimized and adjusted by the network according to data features and model performance during the training process; ∈ takes a value of 0.0001 to avoid the error situation of division by zero during the calculation process; Down(·) represents the downsampling operation, C(·) represents the operation of using 1×1 convolution for channel control, and Up(·) represents the upsampling operation.

8. A method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 7, characterized in that: In step S202, the PKBFormer module in the C2f-PKBFormer module includes two cooperating residual sub-modules. One of the residual modules consists of a normalization operation and a multi-core inverted bottleneck module PKB, and its formula is expressed as: Y = PKB(Norm(x)) + X Among them, Norm(·) represents the normalization operation, which standardizes the distribution of the input data through normalization; PKB(·) represents the multi-core inverted bottleneck module; The other residual module consists of a two-layer MLP with a non-linear activation function, and its formula is expressed as: Z = StarRelu(Norm(Y)W1)W2 + Y Among them, Star Relu is the activation function; W1 and W2 are the learnable parameters of the MLP.

9. A method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 8, characterized in that: The multi-core inverted bottleneck module PKB uses multiple parallel deep convolutional kernels to capture multi-scale texture features: First, through a linear convolution module, the dimension of the input feature map is expanded, and then a set of parallel deep convolutions are used to capture context information across multiple scales. The kernel sizes of the deep convolutions are set to 5×5, 7×7, 9×9, and 11×11 respectively; the formula for the multi-core inverted bottleneck module PKB is expressed as: I′ = StarRelu(PWConv(I)) I″ = DWConv 5×5 (I′) + DWConv 7×7 (I′) + DWConv 9×9 (I′) + DWConv 11×11 (I′) + Id(I′) I″′ = PWConv(I″) Among them, PWConv(·) represents a linear convolution module, which is used to expand the dimension of the feature map and perform preliminary feature extraction; DWConv n×n (·) represents depthwise separable convolutions with different kernel sizes, which are used to extract texture features at different scales; Id(·) represents the identity mapping, which is used to directly output the input, retain the original feature information while fusing with the convolved features, and enrich the feature representation.

10. The method for detecting small targets in UAV aerial photography based on branch weighted fusion according to claim 9, wherein: In step S203, the exponential normalized Gaussian Wasserstein distance loss function EWD module models the bounding box as a two-dimensional Gaussian distribution, uses the Wasserstein distance between two Gaussian distributions to measure the similarity between bounding boxes, and performs exponential normalization transformation to obtain a more reasonable loss value. The calculation formula for the exponential normalized Gaussian Wasserstein distance loss function EWD is as follows: The final loss function is: Loss=α×L CIoU +(1-α)×L EWD Among them, Among them, C represents a constant closely related to the dataset, which is the average size of the targets in the dataset; α represents a balance factor used to adjust the weight ratio of and in the total loss function. When there are more small targets in the dataset, the value of α can be lowered to increase the proportion of in the loss function. In this model, α is set to 0.5.

Citation Information

Cited By

  • Scoliosis identification method, device and equipment for back image

    CN121353814A