An Unmanned Aerial Vehicle Small Target Detection Method Based on Unsupervised Adversarial Learning
Through the unsupervised adversarial learning method, using image multi-scale degradation and enhancement technology, a feature extraction and discrimination network is constructed, which solves the problem of insufficient generalization capabilities of the drone object detection algorithm under complex backgrounds and different postures, and achieves high-precision small-object recognition.
Patent Information
- Application Number
- CN202411265268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The existing drone object detection algorithms rely on supervised learning, require a large amount of labeled data, and have limited generalization capabilities under complex backgrounds and different poses, making it difficult to adapt to new fields and environments.
Unsupervised adversarial learning method is adopted to build feature extraction networks and discriminant networks through multi-scale image degradation and enhancement, conduct adversarial learning, generate and distinguish small-objective feature maps, and improve the robustness and generalization ability of the model.
It improves the accuracy and adaptability of drone small target detection, can accurately identify small targets under complex backgrounds and different postures, and enhances the generalization ability of the model.
Smart Images

Figure CN119131364B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target detection, and in particular to a method for detecting small targets of unmanned aerial vehicles based on unsupervised adversarial learning. Background Technique
[0002] The detection of unmanned aerial vehicle (UAV) targets is an important branch in the field of computer vision target detection. With the continuous progress of UAV technology, the application scope of UAVs has been continuously expanded, covering multiple fields such as military, civilian, and commercial. The purpose of UAV target detection is to locate, detect, and classify UAVs from the images or videos to be inspected. However, the potential threats of UAVs have gradually emerged, such as illegal intrusion of UAVs and malicious attacks of UAVs. Nowadays, UAV target detection based on machine vision has been widely applied in daily life, such as in fields like airports, classified places, and video surveillance.
[0003] Early target detection mainly relied on manually designed features and machine learning-based classifiers. Typical methods included Haar features and cascade classifiers (such as the Viola-Jones algorithm), as well as methods based on HOG (Histogram of Oriented Gradients) features. However, these methods performed poorly in cases of complex backgrounds, multiple targets, and scale variations. With the continuous development of the field of computer vision, the emergence of various detection algorithms has changed the above situation. R-CNN (Region with CNN feature) is one of the pioneering works using deep learning methods, which proposed to locate objects in images through region proposals, then use a convolutional neural network (CNN) to extract features for each proposal, and finally classify. Subsequently, Fast R-CNN improved R-CNN by integrating the entire detection process into a single network to improve speed and efficiency, and used the RoI (Region of Interest) pooling layer to extract features from the shared feature map. Nowadays, single-stage object detectors such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) directly predict the category and location of targets through a single forward propagation process, greatly improving the detection speed. The YOLO-v5 network, as an efficient one-stage algorithm, has characteristics such as strong generalization performance and fast inference speed.
[0004] Currently, a large number of UAV target detection algorithms, such as improved algorithms like YOLO-v5 and EfficientDet, have achieved good results in UAV target detection. The essence of these algorithms is supervised learning. However, supervised learning target detection algorithms still have some deficiencies. For example, supervised learning algorithms usually require a large amount of labeled data for training. Especially in the field of deep learning, supervised target detection algorithms have limited adaptability to target occlusion and different poses. Supervised target detection algorithms usually can only learn the patterns in the labeled data during training, and may have poor generalization ability for targets in new fields or different environments. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a UAV small target detection method based on unsupervised adversarial learning to improve the accuracy of UAV target detection.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A UAV small target detection method based on unsupervised adversarial learning includes:
[0008] Obtain the original images containing UAV small targets in the training set, extract the rectangular regions containing the detection targets in the original images to obtain the detection target images and the pure background images without the detection targets;
[0009] Perform image degradation and enhancement operations on the detection target images, including: performing nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation on the detection target images respectively, and then randomly selecting one interpolation method for secondary degradation on the three groups of images with different degradation degrees obtained to obtain another three groups of images with different degradation degrees, totaling six groups of different degraded images; perform enhancement processing in multiple modes on the obtained degraded images to obtain multiple groups of degraded and enhanced images of the detection targets;
[0010] Construct a small target detection model, including a feature extraction network and an unsupervised target discriminant network; the feature extraction network serves as a generator, including a feature extractor composed of the YOLO-v5 backbone network, and a feature adapter composed of multiple fully connected layers and / or multi-layer perceptrons MLP; the unsupervised target discriminant network serves as a discriminator, including a binary classification discriminator.
[0011] Train the small target detection model, that is, perform adversarial learning, including: fusing the degraded enhanced image of the detection target with the pure background image to obtain a synthetic image, and inputting the synthetic image and the pure background image into the feature extraction network respectively to generate an abnormal feature map and a normal feature map; inputting the obtained abnormal feature map and normal feature map into the discriminator, and the discriminator discriminates and distinguishes the differences between the two feature maps, discriminates the abnormal feature map as false, and the normal feature map as true;
[0012] Select the image of the target to be detected in the test set and input it into the trained small target detection model to output the detection result of the small target of the drone.
[0013] Furthermore, the enhancement process includes random scaling and translation transformation operations, operations of adjusting the chroma, saturation and brightness of the image, and operations of fusing two images together with a certain transparency.
[0014] Furthermore, the feature extractor includes: two Focus structures and two CSP structures. The Focus structure is used to perform image slicing operations on the image before semantic feature extraction, obtain four similar complementary images, concentrate the width W and height H information into the channel space, and expand the input channels by 4 times, that is, the spliced image becomes 12 channels relative to the original RGB three channels. Finally, the obtained new image is subjected to convolution operations to finally obtain a two-fold downsampled feature map without information loss;
[0015] The CSP structure is used to split the feature map into two parts, one part is subjected to convolution operations, and the other part is Concate spliced with the result of the previous convolution operation, and then subjected to multiple convolution operations to extract multi-layer semantic features of the processed feature map.
[0016] Furthermore, the binary classifier discriminator is specifically a two-layer MLP structure, which is used as a normality scorer to directly estimate the normality of each position where h is the height parameter of the position where the image block is located, and w is the width parameter of the position where the image block is located;
[0017] The estimation process of the normality is expressed as: , where, is the normality scorer, which directly estimates the normality of each position , is the transformed adaptive feature.
[0018] Furthermore, the loss function of the discriminator adopts the binary cross-entropy loss function, which is used to quantify the accuracy of the discriminator in identifying the mapped image features and the synthetic image features. The lower the value, the better the performance. The calculation formula of the discriminator loss function is as follows:
[0019]
[0020] Among them, G is the generator, D is the discriminator, x represents the real data, y represents the synthetic data, G(x) represents the normal feature map, G(y) represents the abnormal feature map, D[G(x)] represents the discrimination result of the discriminator on the normal feature map, and D[G(y)] represents the discrimination result of the discriminator on the abnormal feature map.
[0021] The beneficial effects of the present invention are as follows:
[0022] (1) The present invention introduces image multi-scale degradation and enhancement into the object detection algorithm, enabling the model to better learn the multivariate structures and patterns in the data, and enabling the detection algorithm to better identify small objects. In this way, the model can assist the object detection framework in learning more distinguishable and universal semantic features for small object recognition. In addition, enhancing the degraded objects in multiple modes can not only make the training data as close as possible to the data with the real distribution, but also avoid problems such as sample imbalance, forcing the model to learn more robust features, thereby effectively improving the generalization ability of the model.
[0023] (2) The present invention introduces the generative adversarial network into the object detection task, uses the feature extractor instead of the generator, generates the corresponding feature maps of the image background and the synthetic image through the feature extractor respectively, and simultaneously inputs the two different feature maps into the discriminator. The discriminator will then identify and distinguish the differences between the two feature maps, and the feature extractor will also continuously learn and strengthen the extraction of image features. In the continuous adversarial learning, the discriminator can focus on learning the differences between the two feature maps for locating and detecting small target objects in the synthetic image feature map. Description of the Drawings
[0024] Figure 1 is the overall network architecture diagram of the present invention;
[0025] Figure 2 is the schematic diagram of image multi-scale degradation and enhancement;
[0026] Figure 3 is the schematic diagram of adversarial learning between the mapped features and the image features. Detailed Embodiments
[0027] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] Refer to Figures 1-3, the present invention provides a technical solution:
[0029] A method for detecting small targets of unmanned aerial vehicles based on unsupervised adversarial learning, comprising:
[0030] S1. Obtain the original images containing small targets of unmanned aerial vehicles in the training set, extract the rectangular regions containing the detection targets in the original images to obtain the detection target images and the pure background images without the detection targets;
[0031] Small target objects usually have small sizes, which makes the number of pixels of them in the image relatively small. The target objects may be submerged by adjacent background pixels, and the detection algorithm is extremely likely to identify non-target regions as small targets. The detection algorithm requires more and more discriminative features to distinguish small target objects from the background. The purpose of image multi-scale degradation and enhancement is to enable the model to fully learn the multivariate and multi-scale data features of small target objects at different scales and different modalities.
[0032] S2. Perform image degradation and enhancement operations on the detection target images, including: performing nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation on the detection target images respectively, and then randomly selecting one interpolation method for the three groups of images with different degradation degrees obtained to perform secondary degradation to obtain another three groups of images with different degradation degrees, totaling six groups of different degraded images; performing enhancement processing in multiple modes on the obtained degraded images to obtain multiple groups of degraded and enhanced images of the detection targets; the schematic diagram of the process of image multi-scale degradation and enhancement is as Figure 2 shown.
[0033] The enhancement processing includes operations of random scaling and translation transformation, adjusting the chromaticity, saturation, and brightness of the image, and fusing two images together with a certain transparency.
[0034] In this embodiment, 3 different degradation scales and 3 degradation methods are adopted to generate a total of 9 groups of different degraded images, generate diverse training samples, complete the expansion and enrichment of the training data set, and enable the detection model to learn more abundant and comprehensive target features. In the actual scenario, small target detection is often affected by multiple factors, such as low resolution, occlusion, and illumination change, resulting in the unclear representation of the target in the image. Adopting image multi-scale degradation can simulate this situation, enabling the model to encounter more diverse data during the training process, thereby better adapting to various complex detection scenarios and enabling the model to better adapt to different scenario changes.
[0035] By using the training sample data of an existing model and applying data augmentation techniques to generate more training data, the augmented training data can be made closer to the data with the real distribution. The data augmentation method adopts a random combination of color channel transformation, image splicing, image scaling, and image blurring. This can force the model to learn more robust features, thereby effectively improving the generalization ability of the model.
[0036] S3. Construct a small target detection model, such as Figure 1 , including a feature extraction network and an unsupervised object discrimination network; the feature extraction network serves as a generator, including a feature extractor composed of a YOLO-v5 backbone network, a plurality of fully connected layers, and / or a feature adapter composed of a multi-layer perceptron MLP; the unsupervised object discrimination network serves as a discriminator, including a binary classifier.
[0037] Specifically, the feature extractor includes: two Focus structures and two CSP structures. The Focus structure is used to perform image slicing operations on the image before semantic feature extraction, obtaining four similar complementary images, concentrating the width W and height H information into the channel space, expanding the input channels by 4 times, that is, the spliced image becomes 12 channels relative to the original RGB three channels. Finally, the obtained new image undergoes a convolution operation to finally obtain a two-fold downsampled feature map without information loss.
[0038] The CSP structure is used to split the feature map into two parts, one part undergoes a convolution operation, and the other part and the result of the convolution operation of the previous part are subjected to a Concate splicing operation, and then multiple convolution operations are performed to extract multi-layer semantic features of the processed feature map.
[0039] In some specific embodiments, the feature extractor is composed of several convolutional layers, pooling layers, activation functions, batch normalization layers, activation functions, global average pooling layers, and fully connected layers, which can gradually refine the shallow semantic information of the image through the feature extractor to obtain the deep high-level semantic information of the image, serving the subsequent image target detection steps.
[0040] S4. Train the small target detection model, that is, perform adversarial learning, including: fusing the degraded enhanced image of the detection target with the pure background image to obtain a synthetic image, inputting the synthetic image and the pure background image into the feature extraction network respectively to generate an abnormal feature map and a normal feature map; inputting the obtained abnormal feature map and normal feature map into the discriminator, and the discriminator discriminates and distinguishes the differences between the two feature maps, discriminating the abnormal feature map as false and the normal feature map as true.
[0041] The feature extractor extracts features from the synthesized image to generate synthesized image features; the feature adapter is used to project the local features obtained from the pure background image by the feature extractor onto the adaptive features during the training phase, so that the training features are transferred to the target domain to generate mapped image features.
[0042] The binary classifier is specifically a two-layer MLP structure, which is used as a normality scorer to directly estimate the normality of each position where h is the height parameter of the position where the image patch is located, and w is the width parameter of the position where the image patch is located; the estimation process of the normality is expressed as: , where, is the normality scorer, which directly estimates the normality of each position ; is the transformed adaptive feature. Since the negative samples are generated together with the normal features, they are all input to the discriminator during training. The discriminator expects a positive output for the normal features and a negative output for the abnormal features.
[0043] The unsupervised target discriminant network feeds the abnormal feature map (synthesized image features) and the normal feature map (mapped image features) into the discriminator. The discriminator continuously learns how to better distinguish these two types of images, while the generator is also learning how to generate more realistic images to deceive the discriminator. In this process, the discriminator loss function calculates the performance of the discriminator in distinguishing the two types of images.
[0044] In this embodiment, the loss function of the discriminator adopts the binary cross-entropy loss function, which is used to quantify the accuracy of the discriminator in identifying the mapped image features and the synthesized image features. The lower the value, the better the performance. The calculation formula of the discriminator loss function is as follows:
[0045]
[0046] where G is the generator, D is the discriminator, x represents the real data, y represents the synthesized data, G(x) represents the normal feature map, G(y) represents the abnormal feature map, D[G(x)] represents the discriminant result of the discriminator for the normal feature map, and D[G(y)] represents the discriminant result of the discriminator for the abnormal feature map.
[0047] During the model training process, the generator and the discriminator learn adversarially. The generator attempts to generate samples similar to the real data, while the discriminator tries to distinguish between the generated samples and the real samples. Through this adversarial learning process, the generator continuously improves the quality of the generated samples, and the discriminator also continuously enhances its ability to identify the generated samples. The present invention introduces the generative adversarial network into the object detection task. The feature extractor is used instead of the generator. The image background and the synthetic image are respectively passed through the feature extractor to generate corresponding feature maps. The two different feature maps are simultaneously input into the discriminator. The discriminator will then identify and distinguish the differences between the two feature maps, and the feature extractor will also continuously learn and strengthen the extraction of image features. Through the above adversarial learning, the feature extractor will have a stronger feature extraction ability, and the discriminator will have a stronger discrimination ability. The discriminator can focus on learning the differences between the two feature maps during continuous adversarial learning, and can achieve the localization and detection of small target objects in the feature map of the synthetic image.
[0048] S5. Select the target image to be detected in the test set and input it into the trained small target detection model to output the small target detection result of the drone.
[0049] The schematic diagram of the discriminator performing adversarial learning based on the mapped image features and the synthetic image features and being used for object detection is as Figure 3 , and the trained model can focus on the target area according to the features of the target image to be detected, and achieve the localization and detection of small targets.
[0050] Since image multi-scale degradation and enhancement and the feature extraction process are added during network training, the trained model has a more obvious difference in identifying the target and the background, has strong adaptability to real-world images, has high real-world generalization, high robustness, and effectively improves the detection accuracy.
[0051] The above is only the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and modifications made by those skilled in the art that do not depart from the spirit and scope of the present invention should all be within the protection scope of the appended claims of the present invention.
Claims
1. An unmanned aerial vehicle small target detection method based on unsupervised adversarial learning, characterized in that: Including: Obtain the original images containing small UAV targets in the training set, extract the rectangular regions containing the detection targets in the original images to obtain the detection target images and the pure background images without detection targets; Perform image degradation and enhancement operations on the detection target images, including: performing nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation on the detection target images respectively, and then randomly selecting one interpolation method for the three groups of images with different degradation degrees obtained to perform secondary degradation to obtain another three groups of images with different degradation degrees, totaling six groups of different degraded images; perform enhancement processing in multiple modes on the obtained degraded images to obtain multiple groups of degraded and enhanced images of the detection targets; Construct a small target detection model, including a feature extraction network and an unsupervised target discrimination network; the feature extraction network serves as a generator, including a feature extractor composed of a YOLO-v5 backbone network, and a feature adapter composed of multiple fully connected layers and / or multi-layer perceptrons MLP; the unsupervised target discrimination network serves as a discriminator, including a binary classifier; Train the small target detection model, that is, perform adversarial learning, including: fusing the degraded and enhanced images of the detection targets with the pure background images to obtain synthetic images, and inputting the synthetic images and the pure background images into the feature extraction network respectively to generate abnormal feature maps and normal feature maps; input the obtained abnormal feature maps and normal feature maps into the discriminator, and the discriminator discriminates and distinguishes the differences between the two types of feature maps, and discriminates the abnormal feature maps as false and the normal feature maps as true; Select the images of the targets to be detected in the test set and input them into the trained small target detection model to output the detection results of small UAV targets.
2. The method for detecting small targets of drones based on unsupervised adversarial learning according to claim 1, wherein: The enhancement processing includes operations such as random scaling and translation transformation, adjusting the chroma, saturation, and brightness of the images, and fusing two images together with a certain transparency.
3. The method for detecting small targets of an unmanned aerial vehicle based on unsupervised adversarial learning according to claim 1, wherein: The feature extractor includes: two Focus structures and two CSP structures. The Focus structure is used to perform image slicing operations on the images before semantic feature extraction, obtain four similar and complementary images, concentrate the width W and height H information into the channel space, and expand the input channels by 4 times, that is, the spliced image becomes 12 channels relative to the original RGB three channels. Finally, the obtained new image undergoes convolution operations to finally obtain a two-fold downsampled feature map without information loss; The CSP structure is used to split the feature map into two parts, one part performs convolution operations, and the other part performs a Concate splicing operation with the result of the previous part's convolution operation, and then performs multiple convolution operations to extract multi-layer semantic features of the processed feature map.
4. A method for detecting small targets of an unmanned aerial vehicle based on unsupervised adversarial learning according to claim 1, characterized in that: The binary classifier is specifically a two-layer MLP structure, which is used as a normality scorer to directly estimate the normality of each position , where h is the height parameter of the position where the image patch is located, and w is the width parameter of the position where the image patch is located; The estimation process of the normality is expressed as: , where is a normality scorer that directly estimates the normality of each position , is the transformed adaptive feature.
5. The method for detecting small targets of an unmanned aerial vehicle based on unsupervised adversarial learning according to claim 1, wherein: The loss function of the discriminator adopts the binary cross-entropy loss function, which is used to quantify the accuracy of the discriminator in identifying the feature maps of the mapped images and the synthetic images. The lower its value, the better the performance. The calculation formula of the discriminator loss function is as follows: Among them, G is the generator, D is the discriminator, x represents real data, y represents synthetic data, G(x) represents the normal feature map, G(y) represents the abnormal feature map, D[G(x)] represents the discrimination result of the discriminator on the normal feature map, and D[G(y)] represents the discrimination result of the discriminator on the abnormal feature map.
Citation Information
Patent Citations
Unmanned aerial vehicle aerial image target detection method based on improved YOLO V5
CN113807464A
Text image confrontation generation system, method and equipment based on multi-modal information prompt and medium
CN118429779A