Target detection data enhancement method for unmanned aerial vehicle aerial image
By combining single-sample and multi-sample data augmentation with an aerial image target placement network, the problem of data augmentation in UAV aerial image target detection not conforming to the real world is solved, thus improving the model's detection performance and adaptability.
Patent Information
- Application Number
- CN202111625378.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-12-28
AI Technical Summary
Existing data augmentation methods for target detection in UAV aerial images fail to effectively utilize the relationship between the target and its placement, resulting in data augmentation effects that do not conform to real-world conditions. Furthermore, existing methods lack the ability to extract and sample features from UAV aerial images, making it impossible to generate samples that fit complex scenarios.
We employ single-sample and multi-sample data augmentation methods. Through the Aerial Image Target Placement Network (AITP-Network), we utilize deep learning to construct a target placement network, combining channel and spatial attention modules to predict the optimal position of targets in the environment, thereby generating augmented datasets that are more consistent with the real world.
It improves the performance of the object detection model, enhances the real-world relevance of the dataset, and improves the model's detection accuracy and adaptability.
Smart Images

Figure CN116416139B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technology in the field of image processing, specifically a method for target detection data enhancement of drone aerial images. Background Technology
[0002] When a drone captures images, the resulting images will have different environmental contexts depending on its altitude and viewing angle. Existing techniques (such as the paper "mixup: Beyond Empirical Risk Minimization" published on arXiv:1710.09412 on October 25, 2017) consider data augmentation with other sample images, but the mixing process does not take into account the relationship between the target and its placement, resulting in some data-augmented samples not conforming to real-world conditions and reducing the effectiveness of the data augmentation.
[0003] The existing technology (the paper "Modeling Visual Context is KeytoAugmenting ObjectDetection Datasets" published in Proceedings of the European Conference on ComputerVision (ECCV) in 2018, pages 364-380) considers incorporating environmental information when mixing with other samples for data augmentation. However, the context network it builds is insufficient for feature extraction and sampling of drone aerial images. It only uses the simplest CNN network for feature extraction and does not consider combining single-sample data augmentation methods.
[0004] Chinese patent document CN112215100A, published on 20210112, discloses a target detection method for degraded images under imbalanced training samples. However, this prior art does not utilize information from other samples to generate mixed samples, which fails to bring more complex new scenes to the dataset. At the same time, it does not take into account the environmental information of the samples and is prone to generating samples that do not conform to the real scene. Summary of the Invention
[0005] To address the aforementioned shortcomings of existing technologies, this invention proposes a target detection data augmentation method for UAV aerial images. First, single-sample data augmentation is performed, with a series of data augmentation processes designed for each image. Then, multi-sample data augmentation is conducted using deep learning to construct an Aerial Image Target Placement network (AITP-Network) based on the target's surrounding environment information. This places target instances on suitable backgrounds to create an augmented dataset that better reflects the real world.
[0006] This invention is achieved through the following technical solution:
[0007] This invention relates to a target detection data augmentation method for UAV aerial images. The method involves performing single-sample data augmentation on the original labeled dataset, storing the resulting single-sample augmented dataset, constructing and training an aerial image target placement network (AITP-Network), inputting the single-sample augmented dataset into the aerial image target placement network for training, inputting the single-sample augmented dataset again into the aerial image target placement network for inference, and performing multi-sample augmentation on data from the same viewpoint to obtain the final target detection sample augmentation set.
[0008] The aerial image target placement network includes: a ResNet framework backbone network, a channel attention module, and a spatial attention module. The backbone network extracts feature maps based on the input image information to obtain feature map results. The channel attention module assigns channel information weights based on the feature map information to obtain channel-weighted feature map results. The spatial attention module assigns spatial information weights based on the channel-weighted feature map information to obtain attention-weighted feature map results.
[0009] This invention relates to a system for implementing the above-described method, comprising: a single-sample enhancement unit, an aerial image target placement environment prediction unit, and a target placement unit, wherein: the single-sample enhancement unit performs single-sample enhancement processing based on the information of the original labeled dataset to obtain a single-sample enhancement dataset result; the aerial image target placement environment prediction unit performs target placement environment prediction processing based on the information of the single-sample enhancement dataset to obtain a target placement environment result; and the target placement unit performs target placement processing based on the single-sample enhancement dataset and the target placement environment prediction information to obtain a target detection sample enhancement set result.
[0010] Technical effect
[0011] Compared with the prior art, the technical effects of the present invention include:
[0012] 1) By employing an aerial image target placement network, the most suitable environment and target instance for placement are predicted, making full use of information from multiple samples, thereby achieving the goal of making the enhanced image more consistent with the real world situation.
[0013] 2) By employing an attention mechanism in the aerial image target placement network using CNN, spatial and channel attention modules were added, thereby improving the sampling and inference capabilities of the aerial image target placement network.
[0014] 3) By statistically analyzing the aspect ratio distribution of the bounding boxes in the single-sample augmentation dataset, and creating random preset target boxes based on this distribution, the goal is to make the target detection sample augmentation set more consistent with the data characteristics of the original labeled dataset.
[0015] 4) By adopting a multi-sample data augmentation method on the basis of the single-sample data augmentation method, the effect of making full use of the information of the sample itself and the mixed information of other samples is achieved. Attached Figure Description
[0016] Figure 1 This is a flowchart of the present invention;
[0017] Figure 2 This is a schematic diagram of the network architecture of the present invention. Detailed Implementation
[0018] like Figure 1 As shown in this embodiment, a target detection data augmentation method for UAV aerial images includes:
[0019] Step 1: For each image in the original labeled dataset, perform single-sample data augmentation and store the resulting single-sample augmentation dataset, which specifically includes:
[0020] Step 1.1: For each image in the original labeled dataset, while keeping the aspect ratio unchanged, perform random scaling with the long side ranging from [1333, 768] pixels and the wide side ranging from [1333, 1280] pixels to obtain a randomly scaled dataset;
[0021] The original labeled dataset mentioned above is specifically from the UAV aerial target detection dataset Visdrone-DET2019;
[0022] Step 1.2: Randomly select half of the images in the random scaling dataset, flip them horizontally, and add the newly generated images to the random scaling dataset to obtain the flipped dataset;
[0023] Step 1.3: Randomly select images from the flipped dataset and randomly perform hue, saturation, and brightness transformations (i.e., HSV transformation). All relevant parameters are obtained by taking a random natural number between [0, 100] to obtain the HSV dataset.
[0024] Step 1.4: Pad all images in the HSV dataset, setting the padding size_divider to 32 to obtain the single-sample augmentation dataset.
[0025] Step 2: Construct and train the aerial image target placement network. Input the single-sample augmentation dataset into the aerial image target placement network for training, specifically including:
[0026] Step 2.1: In the single-sample data augmentation dataset, pre-set 100 context boxes of different sizes. In each image, when the context box can completely wrap the label box, set the context box as a positive sample. When the IOU between the context box and the label box does not exceed 0.2, set the context box as a negative sample.
[0027] Step 2.2: For each context box, set all pixels in the labeled box within it to (0,0,0). Then, use these context boxes as training samples and input them into the aerial image target placement network for training. The label values are set to a k+1 dimensional vector, where k is the target category contained in the current dataset, i.e., the target category contained in the current training sample.
[0028] like Figure 2 As shown, the aerial image target placement network includes: a ResNet backbone network, a channel attention module, and a spatial attention module. The channel attention module uses environmental location information for enhanced encoding, while the spatial attention module uses both global max pooling and average pooling for weight compression and 7*7 convolutional kernels for dimension alignment.
[0029] The enhanced encoding specifically refers to the following: for each channel of a feature map information X (with dimensions H*W*C), the weights of the feature map pixels in the W direction are calculated by average pooling in the W direction, and the weights of the feature map pixels in the H direction are calculated by average pooling in the W direction. Then, z1 and z2 are concatted, and a 1*1 convolution kernel is used to perform convolution dimensionality reduction to make the dimension C. Finally, the ReLU activation function is used for activation.
[0030] The aforementioned dimension alignment specifically refers to the following: For each (h, w) pixel in a feature map information X (with dimensions H*W*C), we sequentially perform global max pooling and average pooling along its channel direction to obtain z3 and z4 respectively. Then, we concat z1 and z2, and then use a convolutional layer with a kernel of 7*7 to perform dimension alignment so that its dimension is C. Finally, we activate it using the sigmoid activation function.
[0031] Step 3: Input the single-sample augmentation dataset back into the aerial image target placement network for inference, perform multi-sample augmentation on the data from the same viewpoint, and obtain the final target detection sample augmentation set, which specifically includes:
[0032] Step 3.1: Statistically analyze the aspect ratio distribution of the sample annotation boxes in the single-sample augmented dataset, and set 100 preset target boxes of random size with aspect ratios that conform to this distribution in all environment boxes according to this ratio;
[0033] Step 3.2: Set the pixel values in the preset target box to (0,0,0) to obtain the environment box to be inferred;
[0034] Step 3.3: For each image, take all the context boxes to be inferred on the image and input them into the aerial image target placement network for inference. Sort the final classification probabilities, select the five highest target categories, and randomly select a context box. Place the target category instance with the highest probability in the context box to obtain the target detection sample augmentation set.
[0035] Through specific practical experiments, under a specific environment setting of Ubuntu 18.04 server, Python 3.6.8, and PyTorch 1.3.1, the above-mentioned target detection data augmentation method for UAV aerial images was run on the Visdrone-DET2019 UAV aerial target detection dataset to obtain an augmented target detection sample set. The classic YOLOv5-x model was then trained on this augmented set with default hyperparameter configuration. The experimental data obtained are as follows: YOLOv5-x achieved a mAP (mean Average Precision) of 36.79 and an AP50 (Average Precision on IoU50, average precision with an IoU threshold of 50) of 58.12.
[0036] Compared to existing YOLOv5-x models, the target detection data augmentation method for UAV aerial images proposed in this invention improves the model's detection performance, increasing the mAP value by 1.05 (a relative improvement of 2.94%) and the AP50 value by 1.4 (a relative improvement of 2.43%). Compared to YOLOv5-x models trained based on traditional random placement data augmentation algorithms, the method proposed in this invention improves the mAP value by 0.47 (a relative improvement of 1.29%) and the AP50 value by 0.8 (a relative improvement of 1.40%).
[0037] In summary, this invention enhances environmental features by adding channel attention and spatial attention modules to the aerial image target placement network, and occludes the target during the training phase. This effectively extracts the regular information between the target and the environment, and then uses the aerial image target placement network for inference to provide environmental regions suitable for placing specific target categories, thereby creating enhanced samples that are more in line with real-world scenarios.
[0038] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.
Claims
1. A method for target detection data augmentation for UAV aerial images, characterized in that, The original annotation dataset is subjected to single sample data enhancement, the obtained single sample enhanced dataset is stored, an aerial image target placement network is constructed and trained, the single sample enhanced dataset is input into the aerial image target placement network for training, the single sample enhanced dataset is input into the aerial image target placement network again for inference, the data under the same perspective is subjected to multi-sample enhancement, and a final target detection sample enhanced set is obtained. The training refers to: in the single sample data enhanced dataset, 100 environment boxes with different sizes are preset, in each picture, when the environment box can completely wrap the annotation box, the environment box is set as a positive sample, when the IOU of the environment box and the annotation box is less than 0.2, the environment box is set as a negative sample; for each environment box, the pixel points in the annotation box are all set as (0, 0, 0), and then the environment boxes are input into the aerial image target placement network for training, and the label value is set as a k+1-dimensional vector, k is the target type contained in the current dataset, that is, the target category contained in the current training sample; The inference refers to: the length-width ratio distribution of the sample annotation box in the single sample enhanced dataset is counted, 100 preset target boxes with random sizes and length-width ratios consistent with the distribution are set in all environment boxes, and then the pixel values in the preset target boxes are set as (0, 0, 0) to obtain the environment boxes to be inferred; For each picture, all the environment boxes to be inferred on the picture are input into the aerial image target placement network for inference, the final classification probability is sorted, the top five target categories are taken and a random environment box is selected, and the target category instance with the highest probability is placed in the environment box to obtain the target detection sample enhanced set.
2. The method of claim 1, wherein the method further comprises: The aerial image target placement network comprises a backbone network of a Resnet framework, a channel attention module and a spatial attention module, wherein: the backbone network performs feature map extraction processing according to the input picture information to obtain a feature map result; the channel attention module performs channel information weighting processing according to the feature map information to obtain a channel weighted feature map result and uses environment position information for enhanced coding; and the spatial attention module performs spatial information weighting processing according to the channel weighted feature map information to obtain an attention weighted feature map result and uses global maximum pooling and average pooling for weight compression and dimension alignment.
3. The method of claim 2, wherein the method further comprises: The enhanced coding specifically refers to: for each channel of a feature map information X with a dimension of H*W*C, the weight of the feature map pixel in the W direction is calculated in the W direction through average pooling, the weight of the feature map pixel in the W direction is calculated in the H direction through average pooling, z1 and z2 are subjected to concat operation, a 1*1 convolution kernel is used for convolution dimension reduction operation, the dimension is C, and then an activation function ReLU is used for activation.
4. The method of claim 2, wherein the method further comprises: The dimension alignment specifically refers to: for each (h, w) pixel point of a feature map information X with a dimension of H*W*C, global maximum pooling and average pooling are sequentially performed on the channel direction of the pixel point, to obtain z3 and z4 respectively, then z1 and z2 are subjected to concat operation, a convolution layer with a convolution kernel of 7*7 is used for dimension alignment, so that the dimension is C, and then an activation function sigmoid is used for activation.
5. The method of claim 1-4, wherein the method comprises: Comprise: Step 1: for each image in the original labeled data set, single sample data augmentation is performed, and the obtained single sample augmented data set is stored, specifically comprising: Step 1.1: for each image in the original labeled data set, random scaling is performed on the long side in the range of [1333, 768] pixels and the width side in the range of [1333, 1280] pixels while keeping the aspect ratio unchanged, to obtain a random scaling data set; Step 1.2: randomly select half of the images in the random scaling data set, perform horizontal flipping, and add the newly generated images to the random scaling data set to obtain a flipping data set; Step 1.3: randomly select images in the flipping data set, and randomly perform HSV transformation of hue, saturation and brightness. All related parameters are obtained by taking a random natural number between 0 and 100, to obtain an HSV data set; Step 1.4: perform padding operation on all images in the HSV data set, set the padding size_divider to 32, and obtain a single sample augmented data set; Step 2: construct and train an aerial image target placement network, input the single sample augmented data set into the aerial image target placement network for training, specifically comprising: Step 2.1: in the single sample data augmented data set, 100 environment boxes with different sizes are preset, in each picture, when the environment box can completely wrap the labeled box, the environment box is set as a positive sample, and when the IOU of the environment box and the labeled box is less than 0.2, the environment box is set as a negative sample; Step 2.2: for each environment box, set all pixel points in the labeled box inside the environment box to (0, 0, 0), and then input these environment boxes as training samples into the aerial image target placement network for training, and set the label value to a k+1-dimensional vector, k is the number of target categories contained in the current data set, i.e. the number of target categories contained in the current training sample; Step 3: input the single sample augmented data set into the aerial image target placement network again for inference, and perform multi-sample augmentation on the data under the same perspective to obtain a final target detection sample augmented set, specifically comprising: Step 3.1: count the length-width ratio distribution of the sample labeled boxes in the single sample augmented data set, and set 100 preset target boxes with random sizes and length-width ratios consistent with the distribution in all environment boxes; Step 3.2: set the pixel values in the preset target boxes to (0, 0, 0) to obtain the inference environment boxes; Step 3.3: For each picture, all the environment boxes to be inferred on the picture are input into the aerial image target placement network for inference, the final classification probability is sorted, the top five target categories are taken, and a random environment box is selected, in which the target category instance with the highest probability is placed, to obtain a target detection sample enhancement set.
6. A system for implementing the target detection data augmentation method for UAV aerial images according to any one of claims 1-5, characterized in that, The method comprises the steps of: The single sample enhancement unit, the aerial image target placement environment prediction unit and the target placement unit, wherein: the single sample enhancement unit performs single sample enhancement processing according to the original annotation dataset information to obtain a single sample enhancement dataset result, the aerial image target placement environment prediction unit performs target placement environment prediction processing according to the single sample enhancement dataset information to obtain a target placement environment result, and the target placement unit performs target placement processing according to the single sample enhancement dataset and the target placement environment prediction information to obtain a target detection sample enhancement set result.
Citation Information
Patent Citations
Target detection method for degraded image under unbalanced training sample
CN112215100A
Method for improving detection performance of instance segmentation model based on data enhancement
CN113033573A
Concrete pavement crack detection method for improving PoolNet network structure
CN113222904A