Unmanned aerial vehicle target detection method based on generative adversarial network
Through the combination of generative adversarial network and Siamese encoder-decoder network, the cross-domain problem in drone target detection is solved, and the detection accuracy and adaptability of the model in different scenarios is improved, and it is suitable for complex environments and real-time target recognition.
Patent Information
- Application Number
- CN202510402755.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
AI Technical Summary
UAV image object detection faces cross-domain problems, and the model performs poorly under different geographical locations and shooting conditions of training and test data, resulting in insufficient generalization capabilities.
The Siamese encoder-decoder network based on the generative adversarial network is adopted to build the Siamese encoder-decoder network through pre-training feature extraction network, and trained with the discriminator to reduce the difference in encoding feature distribution of the source domain and the target domain image, minimize reconstruction error, and conduct joint training of the adversarial loss function to improve the cross-domain adaptability of the model.
It improves the accuracy and generalization capabilities of drone target detection, adapts to complex environments under different geographical locations and shooting conditions, is real-time and robust, and is suitable for emergency rescue and reconnaissance tasks.
Smart Images

Figure CN120279448A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, and particularly to a method for detecting drone targets based on a generative adversarial network. Background Art
[0002] With the development and popularization of drone technology, drones have shown broad application prospects in many fields such as military reconnaissance, environmental monitoring, agricultural plant protection, and logistics distribution. Especially in the field of target detection and classification, drones, due to their unique aerial perspective and flexibility, can quickly cover large areas and obtain high-resolution images of ground targets. However, target detection and classification in drone images face a series of challenges, including large image sizes, small sizes of detection objects, dense distributions, instance overlaps, and insufficient lighting, etc., which all pose higher requirements for the effectiveness of target detection.
[0003] Traditional target detection techniques, such as feature-based methods (e.g., SIFT, SURF) and template matching-based methods, can be used for target detection in drone images to a certain extent. However, these methods usually require manual feature design, have large computational amounts, and are less adaptable to small targets and complex backgrounds in drone images. With the development of deep learning technology, target detection methods based on convolutional neural networks (CNNs) have gradually become a research hotspot due to their powerful feature learning ability and end-to-end training method.
[0004] Deep learning technology, especially convolutional neural networks (CNNs), has made revolutionary progress in the fields of image recognition and target detection. Typical deep learning target detection frameworks include the R-CNN series, YOLO, SSD, etc. These methods can automatically learn the target features in images without manual design of feature extractors, greatly improving the accuracy and efficiency of target detection.
[0005] Another challenge in drone target detection is the cross-domain problem. Since the training data and test data may come from different geographical locations and different shooting conditions, the model performs well on the source domain (training data), but its performance degrades on the target domain (test data). This phenomenon is called domain shift, which seriously affects the generalization ability of the model. Summary of the Invention
[0006] This application proposes a method for detecting drone targets based on a generative adversarial network, aiming to effectively handle the cross-domain problem to improve the accuracy and generalization ability of drone target detection in different scenarios.
[0007] The technical solution adopted in this application is as follows:
[0008] A method for detecting drone targets based on a generative adversarial network, the method comprising the following steps:
[0009] Step 1, construct a Siamese encoder-decoder network using a pre-trained feature extraction network, which includes two encoders and two decoders. Among them, the network structures of the two encoders are both pre-trained feature extraction networks, which are respectively used to extract the encoded features of the source domain image and the target domain image in the drone recognition image; the input of the decoder is the encoded feature output by the encoder, which is used to reconstruct the encoded feature into the corresponding domain image;
[0010] Step 2, train the Siamese encoder-decoder network based on the source domain image and the target domain image in the drone recognition image;
[0011] Input the pre-processed source domain image into a branch including an encoder and a decoder in the Siamese encoder-decoder network; and input the pre-processed target domain image into another branch including an encoder and a decoder;
[0012] Send the encoded features output by the two encoders into a discriminator, which is used to perform domain classification on the input encoded features; and learn and update the network parameters of the discriminator based on the domain classification output by the discriminator and the true domain classification. When the set training convergence condition is satisfied, a trained discriminator is obtained;
[0013] Tune the network parameters of the Siamese encoder-decoder network based on the distribution difference and reconstruction error between the encoded features of the source domain image and the target domain image; when the set tuning convergence condition is satisfied, obtain a trained source domain encoder based on the encoder currently used to process the source domain image, and obtain a trained target domain encoder based on the encoder currently used to process the target domain image;
[0014] Among them, the tuning objective is: minimize the reconstruction error while reducing the distribution difference between the encoded features of the source domain image and the target domain image;
[0015] Step 3, jointly train the source domain encoder, target domain encoder and discriminator trained in Step 2 based on the adversarial loss function (that is, simultaneously train the source domain encoder, target domain encoder and discriminator trained in Step 2), and stop when the set joint training convergence condition (such as the discrimination performance of the discriminator converges or reaches the preset number of training times) is satisfied;
[0016] Step 4, set a target classifier, and train the target classifier based on the source domain encoder after joint training in Step 3 on the source domain image set; when training, the input of the target classifier is the encoded feature output by the source domain encoder;
[0017] Step 5: For the target domain image to be recognized, feature extraction is performed based on the target domain encoder after joint training in Step 3, and then it is input into the target classifier trained in Step 4 to obtain the UAV target detection result of the target domain image.
[0018] Furthermore, the pre-trained feature extraction network can adopt the feature extraction network of VGG16.
[0019] Furthermore, the pre-trained feature extraction network can be directly set as six modules connected in sequence. The first module includes two convolutional layers with a convolution kernel of 3x3 in sequence; the second module includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 5x5 in sequence; the third module includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 3x3 in sequence; the fourth module includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 3x3 in sequence; the fifth module includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 3x3 in sequence; the sixth module is a max pooling layer with a 2x2 pooling window.
[0020] Furthermore, the target classifier includes at least two fully connected layers, and the last fully connected layer is based on the Softmax function to convert the input of this fully connected layer into a probability distribution of target categories.
[0021] Furthermore, the image preprocessing includes: normalizing the image size, normalizing the image pixel values, and color space correction.
[0022] Furthermore, in Step 3, the anti-loss function for joint training includes: the discriminator attempts to maximize its ability to distinguish the encoded features of the source domain and the target domain, and the encoder attempts to minimize the discriminator's ability to distinguish the encoded features of the source domain and the target domain.
[0023] The technical solution provided by this application brings at least the following beneficial effects:
[0024] (1) Cross-domain generalization ability: It can effectively process image data under different geographical locations and different shooting conditions, reduce the distribution difference between the source domain and the target domain, and improve the generalization ability of the model.
[0025] (2) Real-time performance: It is suitable for real-time target detection, can quickly recognize and classify targets during the flight of the UAV, and meets the real-time requirements of scenarios such as emergency rescue and reconnaissance missions.
[0026] (3) High accuracy: The invariant feature representation learned through adversarial training enables the model to maintain high accuracy in target detection tasks, especially on cross-domain datasets.
[0027] (4) Flexibility and adaptability: It can adapt to various complex environments and changes in different target sizes, improving the adaptability and flexibility of the model in practical applications.
[0028] (5) Robustness: The feature representation learned through the Siamese network has good robustness to problems such as target deformation and occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0030] Figure 1 is a schematic process diagram of the drone target detection method based on the generative adversarial network provided by the embodiment of the present application;
[0031] Figure 2 is a schematic network architecture diagram for target recognition implemented by the encoder in the Siamese encoder-decoder network in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present application will be described in detail and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described by referring to the drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0033] The embodiment of the present application proposes a drone target detection method based on the generative adversarial network, which is a drone target detection method based on Siamese-GAN. This method can effectively handle cross-domain problems and improve the accuracy and generalization ability of drone target detection in different scenarios. The method proposed in the present application is particularly suitable for cross-domain target detection problems and can still maintain high-accuracy target detection when there are significant distribution differences between the source domain and the target domain. The present application combines the Siamese network (dual-branch feature extraction network) and the generative adversarial network, and learns cross-domain invariant feature representations through adversarial training, effectively solving the cross-domain problems in drone target detection. In addition, the method proposed in the present application can handle unlabeled data, improving the generalization ability and practicality of the model. Through experimental verification, the method proposed in the present application shows excellent performance on multiple cross-domain datasets and has broad application prospects.
[0034] In one embodiment, referring to Figure 1 , the drone target detection method based on the generative adversarial network provided by the embodiment of the present application includes the following steps:
[0035] Step 1: Construct a Siamese encoder-decoder network using a pre-trained feature extraction network.
[0036] The Siamese encoder-decoder network includes two encoders and two decoders. Among them, the network structures of the two encoders are both pre-trained feature extraction networks, which are respectively used to extract the encoded features of the source domain image (with target recognition labels) and the target domain image (unlabeled image without target recognition labels) in the UAV recognition image; the input of the decoder is the encoded feature output by the encoder, which is used to reconstruct the encoded feature into the corresponding domain image.
[0037] The loss function of the encoder during training can be set as:
[0038]
[0039] where W = |G w (X1) - G W (X2)|, Y is the label indicating whether two samples match. Y = 1 represents that the two samples are similar or match, Y = 0 represents non-match, X1 and X2 are different inputs, P is the number of inputs, and G W represents the network operation with shared parameters. For example, G W (X1) represents the output of the encoder when the input is X1.
[0040] where
[0041] where represents the Euclidean distance between the feature vectors of two samples X1 and X2, P′ represents the dimensionality of the sample features, m is the set threshold, and N is the number of samples.
[0042] Step 2: Combine a discriminator (discriminator network) and train the Siamese encoder-decoder network based on the source domain image and the target domain image in the UAV recognition image.
[0043] Input the pre-processed source domain image into a branch of the Siamese encoder-decoder network that includes one encoder and one decoder; and input the pre-processed target domain image into another branch that includes one encoder and one decoder.
[0044] Send the encoded features output by the two encoders into a discriminator, which is used to perform domain classification on the input encoded features; and update the network parameters of the discriminator based on the domain classification output by the discriminator and the true domain classification. When the set training convergence condition is met, a trained discriminator is obtained.
[0045] Tune the network parameters of the Siamese encoder-decoder network based on the distribution difference and reconstruction error between the encoded features of the source-domain image and the target-domain image; when the set tuning convergence condition is met, obtain the trained source-domain encoder based on the encoder currently used to process the source-domain image, and obtain the trained target-domain encoder based on the encoder currently used to process the target-domain image;
[0046] Among them, the tuning objective is: to minimize the reconstruction error while reducing the distribution difference between the encoded features of the source-domain image and the target-domain image;
[0047] Step 3, jointly train the source-domain encoder, the target-domain encoder, and the discriminator in the Siamese encoder-decoder network trained in Step 2 based on the adversarial loss function, and stop when the set joint training convergence condition is met (such as the discrimination performance of the discriminator converges or reaches the preset number of training times);
[0048] Step 4, set up a target classifier, and train the target classifier on the source-domain image set based on the source-domain encoder after joint training in Step 3; when training, the input of the target classifier is the encoded features output by the source-domain encoder;
[0049] Step 5, for the target-domain image to be recognized, extract features based on the target-domain encoder after joint training in Step 3, and then input it into the target classifier trained in Step 4 to obtain the UAV target detection result of the target-domain image.
[0050] That is, in this application, first, based on a pre-trained feature extraction network (such as the feature extraction network of the VGG16 network), extract features from the source-domain image (labeled image) and the target-domain image (unlabeled image) to obtain high-dimensional feature vectors. Then, build a Siamese encoder-decoder network based on this feature extraction network, and train its network parameters based on the combined discriminator. The network parameters set during training mainly include: the regularization parameter λ, the mini-batch size b, the learning rate of the Adam optimization method, and other related parameters. Then, train the Siamese encoder-decoder network to obtain a trained discriminator and encoder. Then, continue to train the discriminator and encoder simultaneously using the adversarial loss function, that is, regard the encoder as the generator in the generative adversarial network, and optimize the discriminator and encoder, that is, continue to update the network parameters of the discriminator and encoder through backpropagation and the Adam optimization algorithm. Then, train the classifier: use the encoder after joint training to encode the features of the source-domain and target-domain data to obtain the mapped low-dimensional feature representation, and train the classifier network based on the encoded features of the source-domain image to learn to distinguish different categories. Finally, perform target classification based on the trained classifier: use the trained classifier network to classify the encoded target-domain image data.
[0051] In one embodiment, the encoder structure and the target classifier structure adopted in the embodiments of the present application are as follows Figure 2 shown, and it includes a total of 9 parts, which are in sequence:
[0052] (1) Input Layer: The size of the input image is set to 224x224x3, where 3 represents the number of RGB channels; 224x224 represents the spatial size of the input image;
[0053] (2) Convolutional Layer 1: Apply 2 convolutional kernels (filters), each convolutional kernel covering a 3x3 area, and there are a total of 64 feature maps; the size of the output feature maps of Convolutional Layer 1 is: 224x224x64;
[0054] (3) Pooling Layer 1+Convolutional Layer 2: Max pooling, use a 2x2 pooling window with a stride of 2 to reduce the size of the feature maps; then, apply 2 convolutional kernels, each convolutional kernel covering a 5x5 area, and there are a total of 128 feature maps; the size of the output feature maps of this part is: 112x112x128;
[0055] (4) Pooling Layer 2+Convolutional Layer 3: Max pooling, and then apply 3 convolutional kernels, each convolutional kernel covering a 3x3 area, and there are a total of 256 feature maps; the size of the output feature maps of this part is: 56x56x256;
[0056] (5) Pooling Layer 3+Convolutional Layer 4:
[0057] Max pooling, and then apply 3 convolutional kernels, each convolutional kernel covering a 3x3 area, and there are a total of 512 feature maps. The size of the output feature maps of this part is: 28x28x512;
[0058] (6) Pooling Layer 4+Convolutional Layer 5:
[0059] Max pooling, and then apply 3 convolutional kernels, each convolutional kernel covering a 3x3 area, and there are a total of 512 feature maps. The size of the output feature maps of this part is: 14x14x512;
[0060] (7) Pooling Layer 5: Max pooling is used with a 2x2 pooling window to reduce the feature map size to 7x7. The corresponding output feature map size is: 7x7x512;
[0061] (8) Fully Connected Layer 1 and 2: Flatten the output of the pooling layer (Pooling Layer 5) and connect it to a fully connected layer with 4096 neurons; the output feature map size of this part is: 2x4096;
[0062] (9) Output Layer: Use the Softmax function to convert the output of 4096 neurons into a probability distribution of 1000 classes (where 1000 is the set number of target classes, which is set based on the actual application scenario). The output size of the output layer is: 1x1000. Furthermore, the UAV target detection result is obtained based on the target class corresponding to the maximum probability.
[0063] That is, in the embodiments of the present application, the above (1) to (7) constitute the network structure of the encoder, and (8) to (9) are the corresponding target classifiers.
[0064] In one embodiment, the pre-trained feature extraction network of the embodiments of the present application can be directly obtained based on the specified layer of the pre-trained VGG16 network.
[0065] In one embodiment, when performing image preprocessing on the image (source domain or target domain) to be input into the encoder to make it match the input of the encoder, the specific image preprocessing may include:
[0066] A1. Image size adjustment: Resize all images in the source domain and target domain to the fixed size required for the input of the encoder, usually 224x224 pixels;
[0067] A2. Normalize the image pixel values, usually by scaling the pixel values from [0, 255] to the range of [0, 1], or standardize them using the same mean and standard deviation as when VGG16 was trained.
[0068] A3. Color space correction: Convert the image from the RGB color space to the color space used when VGG16 was trained (usually RGB).
[0069] In one embodiment, in step 2, the network parameters set during training mainly include: setting the regularization parameter λ to 1, determining the mini-batch size b to be 100, setting the learning rate of the Adam optimization method to 0.0001, setting the exponential decay rates β1 of the Adam optimization method to 0.9 and β2 to 0.999, and the epsilon value to e -8 .
[0070] In one embodiment, step 2 further includes: randomly shuffling the source domain image data and dividing it into small batches. For each epoch:
[0071] a. Randomly select a small batch of data from the source domain;
[0072] b. At the same time, randomly select a small batch of data from the target domain;
[0073] c. Calculate the encoded features of the two-domain data by the encoder;
[0074] d. Use the discriminator to perform domain classification on the encoded features output by the encoder to update the discriminator network parameters, so that the discriminator can distinguish the features of the source domain and the target domain; that is, send the encoded features output by the source domain encoder and the target domain encoder into the discriminator respectively, and obtain the domain classification result of the current input encoded features through its output;
[0075] e. Try to reconstruct the source domain and target domain data using the encoded features and the decoder network, and calculate the reconstruction error;
[0076] f. Update the network parameters of the encoder-decoder to minimize the reconstruction error and at the same time reduce the distribution difference between the features of the source domain and the target domain.
[0077] In one embodiment, when performing the joint training of the encoder (source domain encoder and target domain encoder) and the discriminator, it specifically includes:
[0078] a. Train according to the set adversarial generation loss function. The discriminator tries to maximize its ability to distinguish the features of the source domain and the target domain, and the encoder (regarded as the generator in the generative adversarial network) tries to minimize the discriminator's ability to distinguish the features of the source domain and the target domain, that is, tries to make the discriminator unable to distinguish the features of the two domains;
[0079] b. The training continues until the performance of the discriminator network no longer improves or reaches the preset number of iterations.
[0080] In one embodiment, when training the target classifier, use the trained Siamese encoder network (source domain encoder) to extract the encoded features of the source domain images, and send them into the target classifier. Based on the predicted labels output by it and the true labels based on the source domain image data, train the target classifier based on the set target recognition loss (such as cross loss, etc.). The target classifier outputs the probabilities of each target category, so that the predicted category can be obtained based on the target category corresponding to the maximum probability among them.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting drone targets based on a generative adversarial network, characterized in that, Including the following steps: Step 1: Construct a Siamese encoder-decoder network using a pre-trained feature extraction network, which includes two encoders and two decoders. Among them, the network structures of the two encoders are both pre-trained feature extraction networks, which are respectively used to extract the encoded features of the source domain image and the target domain image in the UAV recognition image; the input of the decoder is the encoded feature output by the encoder, which is used to reconstruct the encoded feature into the corresponding domain image; Step 2: Train the Siamese encoder-decoder network based on the source domain image and the target domain image in the UAV recognition image; Input the pre-processed source domain image into a branch of the Siamese encoder-decoder network that includes an encoder and a decoder; and input the pre-processed target domain image into another branch that includes an encoder and a decoder; Send the encoded features output by the two encoders into a discriminator, which is used to perform domain classification on the input encoded features; and learn and update the network parameters of the discriminator based on the domain classification output by the discriminator and the true domain classification. When the set training convergence condition is met, obtain the trained discriminator; Optimize the network parameters of the Siamese encoder-decoder network based on the distribution difference and reconstruction error between the encoded features of the source domain image and the target domain image; when the set optimization convergence condition is met, obtain the trained source domain encoder based on the encoder currently used to process the source domain image, and obtain the trained target domain encoder based on the encoder currently used to process the target domain image; Among them, the optimization objective is: minimize the reconstruction error while reducing the distribution difference between the encoded features of the source domain image and the target domain image; Step 3: Jointly train the source domain encoder, target domain encoder, and discriminator trained in Step 2 based on the adversarial loss function, and stop when the set joint training convergence condition is met; Step 4: Set up a target classifier and train the target classifier on the source domain image set based on the source domain encoder after joint training in Step 3; when training, the input of the target classifier is the encoded feature output by the source domain encoder; Step 5: For the target domain image to be recognized, extract features based on the target domain encoder after joint training in Step 3, and then input it into the target classifier trained in Step 4 to obtain the UAV target detection result of the target domain image.
2. The method according to claim 1, characterized in that, The pre-trained feature extraction network uses the feature extraction network of VGG16.
3. The method according to claim 1, wherein The pre-trained feature extraction network is set as six modules connected in sequence; Among them, the first module sequentially includes two convolutional layers with a convolution kernel of 3x3; the second module sequentially includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 5x5; the third module sequentially includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 3x3; the fourth module sequentially includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 3x3; the fifth module sequentially includes a max pooling layer with a 2x2 pooling window and three convolutional layers with a convolution kernel of 3x3; the sixth module is a max pooling layer with a 2x2 pooling window.
4. The method according to claim 1, characterized in that, The target classifier includes at least two layers of fully connected layers, and the last fully connected layer, based on the Softmax function, converts the input of this fully connected layer into a probability distribution of target categories.
5. The method according to claim 1, wherein The image preprocessing includes: normalizing the image size, normalizing the image pixel values, and color space correction.
6. The method according to claim 1, wherein In step 3, the anti-loss function for joint training includes: the discriminator attempts to maximize its discrimination ability to distinguish the encoded features of the source domain and the target domain, and the encoder attempts to minimize the discriminator's discrimination ability to distinguish the encoded features of the source domain and the target domain.
7. The method according to claim 1, wherein In step 3, the convergence condition for joint training is set to the convergence of the discriminative performance of the discriminator or reaching a preset number of training times.
Citation Information
Cited By
Aeromagnetic data noise reduction method of noise reduction auto-encoder based on adversarial regularization
CN122045597A