Unmanned aerial vehicle aerial image deblurring method based on improved DeblGAN
Through the improved DeblurGAN algorithm, combined with FPN MobileNet and Swin Transformer networks, the problem of drone aerial images blurring in complex environments is solved, efficient image debuming effect is achieved, and the effectiveness of drone remote sensing applications is improved.
Patent Information
- Application Number
- CN202510553569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-05
AI Technical Summary
Drone aerial images are easily blurred in complex weather environments, and existing defuzzing algorithms are difficult to effectively restore clear images, affecting information acquisition and application efficiency.
Using the improved DeblurGAN algorithm, by building a dual-branch network structure, combining FPN MobileNet and Swin Transformer networks, the spatial and frequency domain feature information of the image is extracted, and the training effect is optimized through the improved generation adversarial network loss function to restore high-resolution clear images.
It improves the accuracy and robustness of aerial images of drones to deblur, can effectively restore high-resolution clear images in complex environments, and improves the accuracy of information acquisition and application efficiency.
Smart Images

Figure CN120430962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) aerial image processing, and in particular to a deblurring method for UAV aerial images based on an improved DeblurGAN algorithm. Background Art
[0002] Drones, with their significant advantages such as compact size, high maneuverability, wide coverage, and low deployment costs, have become a vital vehicle for modern aerial communication base stations and mobile data collection. As a new type of remote sensing technology, they have achieved remarkable results in urban planning, precision agriculture, environmental monitoring, and other fields. In particular, the rapid information acquisition capabilities enabled by aerial image segmentation technology in scenarios such as disaster response and transportation network monitoring have built an efficient data support system for decision makers.
[0003] At the technical application level, the high-resolution imaging system carried by drones can capture surface spatial information in real time. However, as the flight altitude increases, the stability of the equipment caused by aerodynamic effects gradually becomes apparent. These meteorological factors will significantly affect the flight attitude of the drone, resulting in a decrease in image quality. More importantly, the formation mechanism of image blurring involves multiple physical coupling effects, including: (1) equipment factors: lens optical calibration deviation, sensor resolution limitation; (2) environmental factors: photon noise under low illumination conditions, atmospheric turbulence disturbance; (3) motion factors: dynamic blur caused by platform vibration, motion artifacts caused by high-speed target movement. Such image degradation problems not only restrict the effective extraction of geographic spatial information, but also have a chain effect on social and economic operations at a deeper level: in the field of disaster emergency response, it may lead to misjudgment of disaster situation and delay golden rescue time; in infrastructure monitoring, it is easy to cause structural damage to be missed and increase safety risks; in agricultural production, it will affect the accuracy of crop growth analysis and cause inaccurate resource allocation. Breaking through these technical bottlenecks has become a key research direction for improving the effectiveness of drone remote sensing applications.
[0004] Traditional deblurring algorithms begin with a blur degradation model and employ mathematical modeling or convolutional networks to predict the blur kernel, which is then used to restore the image through deconvolution. However, the blurring process is often complex, variable, and difficult to accurately estimate, making traditional deblurring methods inadequate for restoring a clear image. In recent years, deep learning has rapidly developed, boasting powerful feature learning and mapping capabilities. It can learn deblurring patterns for complex blurry images from large datasets. Examples include convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and transformer networks. Deep learning can continuously adjust and learn features based on training and testing, adaptively adjusting the network to achieve the goal of end-to-end image deblurring. Compared to traditional deblurring algorithms, deep learning demonstrates greater robustness and wider applicability in complex application scenarios. Summary of the Invention
[0005] The purpose of this invention is to propose a deblurring method for drone aerial images based on an improved DeblurGAN, improve the accuracy of the deblurring algorithm, and address the problem of information loss in drone image acquisition caused by complex weather conditions. To achieve this purpose, the specific technical solutions adopted by this invention are as follows:
[0006] Step 1: Prepare the dataset and perform synthetic blurring on the dataset;
[0007] Step 2: Before algorithm processing, perform image preprocessing on the input data set;
[0008] Step 3: Construct the generative network and discriminative network framework of the improved algorithm; the generative network part consists of two parts: the encoder and the decoder. The encoder uses a dual-branch network structure to extract feature maps, and the decoder restores high-resolution images through layer-by-layer upsampling and convolution operations.
[0009] Step 4: Set the target loss function of the improved generative adversarial network model;
[0010] Step 5: Evaluate the performance of the trained deblurring method for UAV aerial images based on the improved DeblurGAN.
[0011] In step 1, an experimental data set is prepared, and manual blur processing is performed based on the UAV aerial image data set Uavid2020. The Albumentations image data enhancement library is used to simulate the random wind conditions that the UAV may encounter; the wind intensity is mainly divided into three levels, namely (15, 30) weak wind, (30, 50) medium wind, and (50, 80) strong wind. The wind direction is set to random in the range of 0-360°; a sufficient number of data sets can be obtained by setting different parameters to meet the experimental requirements; the corresponding blurred images and the corresponding clear images are saved as experimental data; the data set is divided into a training set, a validation set, and a test set in a ratio of 7:2:1;
[0012] In step 2, the input data set is subjected to image preprocessing, and feature information of the image is extracted in the spatial domain and frequency domain respectively; in the spatial domain, the gradient information of the image in the horizontal and vertical directions is calculated using an edge operator (Sobel operator) to obtain spatial domain features reflecting the edge and texture information of the image; in the frequency domain, the distribution information of the image frequency is obtained by performing fast Fourier transform (FFT) and frequency domain centering processing on the blurred image; the extracted spatial domain features and frequency domain features are normalized and scaled to a predetermined interval to eliminate amplitude value differences, and the frequency domain features are further multiplied by a scaling factor to balance the weight relationship between the spatial domain features and the frequency domain features; the processed spatial domain features and frequency domain features are superimposed in the channel dimension according to the set weights to form a fusion feature map containing multiple channels;
[0013] In step 3, the generation network and the discrimination network of the improved algorithm are constructed. For the encoder part of the generation network, a dual-branch network structure is adopted, which is processed by the FPN MobileNet network and the Swin Transformer network in turn.
[0014] The first branch is a MobileNet network based on the FPN feature pyramid structure. It includes an initial convolutional layer that performs a 3×3 convolution operation on the input image, transforming the number of input channels from 3 to the predetermined number of channels 128. Multiple FPN layers are designed, each consisting of a 3×3 convolution, batch normalization, and ReLU activation function. Max pooling is used to achieve layer-by-layer downsampling, generating multi-scale feature maps with shapes such as (H / 2, W / 2), (H / 4, W / 4), and (H / 8, H / 8).
[0015] The second branch uses a Swin Transformer encoder structure, which includes a Stem convolution layer to convert the input blurred image to the same number of channels as the first branch; multi-layer Swin Transformer modules, each of which includes relative position embedding and window-based multi-head self-attention calculation; 1×1 convolution is used to generate the query matrix Q, key matrix K, and value matrix V respectively, and the attention distribution is obtained using the scaling factor and softmax function to achieve feature enhancement within the local window; for the feature maps of each scale output of the second branch, they are normalized using 1×1 convolution to obtain the same number of channels, and then weighted fused with the feature maps of the corresponding scale of the first branch according to the learnable fusion weight;
[0016] For the decoder part of the generative network, layer-by-layer upsampling and convolution operations are used to restore the high-resolution features of the image, and different scale feature extraction and fusion modules are set between the two;
[0017] In step 4, the target loss function of the improved generative adversarial network model is set, and the generator network loss function is composed of a weighted combination of adversarial loss (L_gan), perceptual loss (L_perceive) and pixel loss (L_pixel);
[0018] Set the generator loss function as follows:
[0019] L total =λ gan ·L gan +λ perceive ·L perceive +λ pixel ·L pixel
[0020] Among them, λ gan ,λ perceive ,λ pixel Denote the adversarial loss L gan , Perceptual Loss L perceive and pixel loss L pixel The weight coefficient is set, and the initial weight distribution is set to 2:3:5;
[0021] Set the discriminator loss function as follows:
[0022]
[0023] The goal of the generator is to generate realistic samples, so that the discriminator cannot distinguish between the generated samples and the real samples as much as possible; the first term of the discriminator loss function is to calculate the difference between the discriminator output of the real sample x and the generated sample G(z), and the second term focuses on the discriminator output of the generated sample and tries to minimize D(G(z)) and Ideally, D(x) should be close to 1, indicating that the real sample is identified as real by the discriminator, and D(G(z)) should be close to 0, indicating that the generated sample is identified as false.
[0024] In step 5, the improved generative adversarial network algorithm is trained and experimentally evaluated. The evaluation indicators of the model effect are peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). PSNR can quickly evaluate the global pixel error of the image by calculating the mean square error of two images; SSIM calculates the local similarity block by block through a sliding window and takes the global average value, which can effectively distinguish the changes in the image content structure. The two are used together to better comprehensively evaluate the image quality.
[0025] The present invention provides a deblurring method for UAV aerial images based on an improved DeblurGAN algorithm, which has the following beneficial effects:
[0026] This invention addresses the blurring problem of drone aerial images by improving the generative adversarial network algorithm. Specifically, this method extracts edge texture information and frequency domain information of blurred images in the spatial domain and frequency domain respectively, and performs dynamic weighted fusion through the extrusion excitation module. The preprocessed image is then input into the designed dual-branch network structure, which successively passes through the FPN MobileNet network and the Swin Transformer network to realize dual-branch feature extraction, strengthen the fusion of multi-scale features, and facilitate the restoration of high-resolution clear images. The proposed method has a good effect on the restoration of blurred drone aerial images. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Schematic diagram of the process framework of the method of the present invention;
[0028] Figure 2 To generate the algorithm block diagram of the network;
[0029] Figure 3 To introduce the Swin Transformer Block module structure diagram. DETAILED DESCRIPTION
[0030] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0031] like Figure 1As shown, the embodiment of the present invention discloses a method for processing blurred images of drone aerial photography based on an improved DeblurGAN, and the specific steps are as follows:
[0032] Step 1: Prepare the experimental dataset and synthesize blurred images based on the UAV aerial image dataset Uavid2020;
[0033] Step 2: Preprocess the input image and extract the feature information of the image in the spatial domain and frequency domain respectively;
[0034] Step 3: Build an improved DeblurGAN deblurring algorithm. In the generative network part, a dual-branch network structure combining the FPN MobileNet network and the Swin Transformer network is adopted.
[0035] Step 4: Perform model training for the improved DeblurGAN deblurring algorithm;
[0036] Step 5: After training is complete, save the best model and further evaluate the effectiveness and accuracy of the improved algorithm by inputting test set images.
[0037] In step 1, since there is little blurred image data in the public drone aerial photography dataset, the synthetic blur effect can simulate the high-altitude resistance, wind force, and other conditions that drones encounter during aerial operations. The dataset Uavid2020 captures 4k high-resolution images. The dataset mainly includes two environments: urban street scenes and suburban roads, providing rich scene changes and challenges for moving object recognition. This dataset is selected to synthesize blurred images, and the Albumentations image data enhancement library is used to simulate the random situations that drones may encounter. It is mainly divided into wind_strength wind strength, blur_prob blur probability, distortion_prob distortion probability, and random horizontal displacement d x = int(self.wind_strengthh*w*(2*np.random.rand()-1)), used to simulate the blurring artifacts caused.
[0038] In the synthetic fuzzy dataset, the wind intensity is divided into three levels: (15, 30) for weak wind, (30, 50) for medium wind, and (50, 80) for strong wind. The wind direction is set to random, ranging from 0-360°. By setting different parameters, a sufficient number of datasets can be obtained to meet the experimental requirements. The dataset is organized and paired fuzzy and clear images are obtained as algorithm input.
[0039] In step 2, the input image is preprocessed, that is, the image feature information is extracted in the spatial domain and frequency domain respectively. First, the image IMREAD_GRAYSCALE is read as a grayscale image, and then the Sobel edge detection method is used to calculate the gradient of the image in the x and y directions: grad_x = cv2.Sobel(img,cv2.CV_64F,1,0,ksize=3) and grad_y = cv2.Sobel(img,cv2.CV_64F,0,1,ksize=3). The gradient is then normalized to the range [0,1].
[0040] In the frequency domain, the image is first converted to floating-point type np.float32, and then the image's magnitude spectrum is calculated through Fourier transform (FFT) (magnitude_spectum = 20*np.log(np.abs((fshift)+1e-10)). The magnitude spectrum is then normalized to the range [0, 1]. The frequency domain features also need to be multiplied by the scaling factor scale_factor = 0.5 to balance the weights.
[0041] The two extracted gradients in the spatial domain are merged into two channels, the frequency domain features are expanded to a single channel, and feature fusion is performed on the channel fused_features = torch.cat([grad_x,grad_y,freq_feature],dim=0). At this time, the shape of the feature map is [3,H,W]. The Squeeze-and-Excitation Module (SE) is introduced. The module uses a simple gating mechanism and sigmoid activation function to adjust channel features, learn the nonlinear relationship between channels, and compress the spatial dimension to 1×1 through two fully connected layers. Finally, the module reweights the frequency domain channel and the spatial domain channel according to the excitation results to achieve dynamic adjustment of the features.
[0042] In step 3, an improved DeblurGAN deblurring algorithm is constructed. The specific network structure is as follows: Figure 2As shown. The first branch extracts feature maps of multiple scales by combining the FPN (Feature Pyramid Network) structure with the MobileNet network structure; first, a 128-channel feature map is generated through a 3×3 convolution, batch normalization, and ReLU activation function; the FPN layer simulates feature extraction at 5 scales (from large to small), and each scale uses a 3x3 convolution, batch normalization, and ReLU activation function to form a convolution block; the output of each FPN layer will undergo a 2x2 maximum pooling operation to gradually reduce the spatial size of the feature map; finally, the feature maps of 5 scales are output, and the shapes are: (B, 128, H / 2, W / 2), (B, 128, H / 4, W / 4), (B, 128, H / 8, W / 8), (B, 128, H / 16, W / 16), (B, 128, H / 32, W / 32);
[0043] The second branch uses a Swin Transformer encoder structure and a windowed self-attention mechanism to capture long-range dependencies and global context. A 3x3 convolution is first used to increase the number of channels in the input image from 3 to 128. Each layer consists of a Patch Partition and a Swin Transformer Block. The Patch Partition module divides the input image into non-overlapping patches, each of which is treated as a "token."
[0044] The structure of Swin Transformer Block is as follows: Figure 3 As shown, receiving the input feature tensor Z∈R B×C×H×W , where B is the batch size, C is the number of channels, H and W are the spatial dimensions of the image, and the input feature map is divided into non-overlapping windows of fixed size, each window size is M×M, and the windowed feature W is generated. l On the left side of the structure is the multi-head attention module W-MSA Block (Window-based Multi-head Self-Attention) with a conventional window configuration, which generates query Q, key K, value V matrices through linear projection, calculates the attention score within the window and weights the aggregate features, and outputs the feature Z W-MSA , add the features to the original input element by element to complete the first residual connection Z l =Z+Z W-MSA , realize cross-window information interaction; LN (LayerNormalization) for Z l Perform layer normalization Z of the channel dimension normal=LayerNorm(Z l ), stabilize the numerical distribution; on the right side of the structure is the multi-head attention module SW-MSA Block (Shifted-Window MSA) with a moving window configuration, which performs a window shift operation on the normalized features, and the window offset is Redivide and calculate attention, the output feature is Z SW-MSA ; Then perform element-by-element addition to complete the second residual connection Z l+1 =Z normal +Z SW-MSA ; Apply MLP (Multi-LayerPerception) channel-by-channel feedforward sub-network to Z l+1 Perform two-stage nonlinear transformations, namely linear projection + GULU activation and linear projection + Dropout, to ensure that the output feature dimension is consistent with the input; finally, perform final channel normalization on the MLP output and output the complete feature X out =LayerNorm(X MLP ), used for subsequent network layer processing.
[0045] Swin Transformer Block uses a moving window partitioning method, and the calculation formula is as follows:
[0046]
[0047] Among them, Z l-1 Indicates entering Block token, and Z l Respectively represent The output features of the W-MSA module and MLP module of each Block, and Z l+1 Respectively represent The output features of the SW-MSA module and the MLP module of each block. In the implementation of the code, the moving window size is set to window_size = 7, the number of multi-head attention is set to num_heads = 3, the first-level attention residual, the original input x and the feature after W-MSA processing are added together l =X+X W-MSA , and after two LN processes, the second-level MLP residual, fuses the feature information strengthened before and after, enhances the local attention of the hierarchical level, and improves the ability of subsequent output features to pay attention to global information.
[0048] The two branches of the dual-branch network structure obtain feature maps of different scales respectively and fuse them at the feature layers of corresponding sizes. Then, through the 1×1 Convolution module, the scale features are normalized to 128 channels, and then the multi-scale feature maps are fused, and finally the restored high-resolution image is output.
[0049] In step 4, the pre-processed images are divided into training set, validation set, and test set according to the ratio of 7:2:1, and input into the model for training. When designing the training program, the input image size is cropped to 256×256, and data augmentation operations are performed, including random resection, geometric transformation and other operations to increase data and improve the generalization ability of the model. The number of training rounds epochs is set to 200, the number of batches per training round train_batches_per_epoch is 880, the number of batches per verification round val_batches_per_epoch is 440, the batch processing size batch_size is set to 16, the optimizer uses Adam, the initial learning rate is 0.001, and linear learning rate decay is used, and the decay is set to start from the 50th round. When the loss function in the training process tends to converge and the image similarity value tends to be stable, it is judged that the current training model is good and can be saved for subsequent testing;
[0050] The loss function of the improved DeblurGAN deblurring algorithm is divided into a generator loss function and a discriminator loss function. The generator loss includes the adversarial loss L gan , Perceptual Loss L perceive and pixel loss L pixel , the calculation formula is as follows:
[0051]
[0052] Where B represents the total number of samples in a mini-batch in a forward or backward propagation, that is, the batch size batch_size; b represents the subscript of the sample in the processing batch; D(x) represents the output of the discriminator for the real sample x; H, W, C represent the length, width, height and number of channels of the image respectively; φ(x) represents the feature tensor obtained in the network layer; I truth Represents the real sample image or target image, I false Represents the image generated by the generator. Through the above calculation method, we can obtain the pixel gap and perceived content gap between the real image and the generated image, and continuously optimize the training effect of the model.
[0053] These loss functions are weighted and combined according to the corresponding settings, and the formula is as follows:
[0054] L total=0.2·L gan +0.3·L perceive +0.5·L pixel
[0055] In the specific experimental process, different weighting coefficients may result in different training effects. The weight ratio can be continuously adjusted according to the training effect to achieve the best training model effect.
[0056] The target loss function based on the GAN generation adversarial network can be expressed as follows:
[0057]
[0058] Among them, input sample, sampling real data sample distribution P data A batch of samples {X1,X2,...,X N}, from the noise distribution P z Sampling a batch of noise samples {z1,z2,...,z M}, represents the expected value calculated for the real sample; D(x) represents the output of the discriminator for the real sample x, judging whether the sample comes from real data; when D(x)→1, logD(x)→0, and the loss approaches the minimum value; represents the expectation of the noise distribution calculation; G(z) represents the output of the generator, the goal of which is to generate realistic samples and try to make the discriminator unable to distinguish between generated samples and real samples; when D(G(z))→0, log(1-D(G(z))→0, and the loss is close to the minimum value. The first term of the formula is to calculate the difference between the discriminator output of the real sample x and the generated sample G(z), and the second term focuses on the discriminator output of the generated sample and tries to minimize D(G(z)) and difference.
[0059] In step 5, the evaluation indicators of the model are peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), and the specific formulas are as follows:
[0060]
[0061] Among them, MAX I represents the maximum pixel value of the image (usually 255), MSE represents the mean square error, H and W represent the height and width of the image respectively, and I truth and I false are the real sample image and the image generated by the generative network, (i, j) represents a pixel in the image; μ x ,μ y represents the mean of image X and image Y, σx ,σ y represents the standard deviation, σ xy Denotes covariance, C1 and C2 are constants used to prevent the denominator from being zero. Evaluating images jointly with PSNR and SSIM can provide a more comprehensive assessment of the model's training effectiveness.
[0062] The present invention is not limited to the above-described embodiments. The above description of the specific embodiments is intended to illustrate and explain the technical solutions of the present invention. Without departing from the spirit of the present invention and the scope of protection of the claims, those skilled in the art may make many variations based on the teachings of the present invention, all of which fall within the scope of protection of the present invention.
Claims
1. A deblurring method for UAV aerial images based on improved DeblurGAN, characterized in that: The steps include: S1: Prepare the dataset based on the UAV aerial image dataset Uavid2020. By adding a blur effect to the original UAV aerial image, and saving the blurred image and the corresponding clear image as experimental data; S2: Before the algorithm is processed, the data set is preprocessed. The frequency domain information and spatial domain information of the blurred image are extracted using Fourier transform in the frequency domain and convolution in the spatial domain respectively. The feature information obtained from the two is dynamically weighted and fused through the squeeze excitation module. The fused image is output as the input of the algorithm processing. S3: Improve the DeblurGAN deblurring algorithm. This method for deblurring drone aerial images is based on an improved generative adversarial network. The entire network structure consists of a generator network and a discriminator network. The encoding part of the generator network uses an FPN MobileNet network and a Swin Transformer network to form a dual-branch network structure. The decoding part restores the high-resolution features of the image through layer-by-layer upsampling and convolution operations. S4: Set the target loss function of the generative adversarial network model. The generator network loss function consists of a weighted combination of adversarial loss, perceptual loss, and pixel loss. The discriminator network loss function uses an improved hinge loss function, combined with a gradient penalty term and a stable adversarial training process. The generator and discriminator are iteratively trained on the training set through alternating optimization until the loss function converges well. S5: During model training, the blurred image is first preprocessed in the frequency and spatial domains, and then processed by a dual-branch network. The image undergoes effective feature fusion at multiple scales, and finally a deblurred image is generated. The image output by the generator network is used as the input of the discriminator network. The discriminator network adopts a patch structure to crop the generated image into different patches to improve the network's discrimination speed. The generator and the discriminator constantly compete with each other, and through backpropagation, the generator network is optimized to output better deblurring results. The PSNR and SSIM indicators are used to evaluate the quality of the restored image.
2. The improved DeblurGAN UAV aerial image deblurring method according to claim 1 is characterized in that: In step S1, based on the public UAV aerial photography dataset Uavid2020, the Albumentations image data enhancement library is used to simulate the image blurring effect caused by random wind conditions that the UAV may encounter, including wind_strength wind strength, blur_prob blur probability, distortion_prob distortion probability, and random horizontal displacement d x = int(self.wind_strength*w*(2*np.random.rand()-1)), and by setting different parameters, a sufficient number of data sets are obtained to meet the experimental requirements.
3. The improved DeblurGAN UAV aerial image deblurring method according to claim 1 is characterized in that: In step S2, the specific frequency domain and spatial domain feature extraction process is as follows: In the spatial domain, the edge operator is used to calculate the gradient information of the image in the horizontal direction grad_x and the vertical direction grad_y to obtain the spatial domain features reflecting the edge and texture information of the image; in the frequency domain, the fast Fourier transform FFT and frequency domain centering processing are performed on the blurred image to obtain the distribution information of the image frequency; Next, the extracted spatial domain features and frequency domain features are normalized and scaled to a predetermined interval to eliminate amplitude value differences, and the frequency domain features are further multiplied by a scaling factor to balance the weight relationship between the spatial domain features and the frequency domain features. The two gradients in the spatial domain are merged into two channels, and the frequency domain features are expanded into a single channel. Feature fusion is performed on the channel using fused_features = torch.cat([grad_x, grad_y, freq_feature], dim = 0). The squeeze excitation module is introduced to weight the processed spatial domain features and frequency domain features in the channel dimension according to the dynamic learning weights to form a fused feature map containing multiple channels, and the map is dynamically adjusted to achieve better feature fusion.
4. The improved DeblurGAN UAV aerial image deblurring method according to claim 1 is characterized in that: In the S3 step, the process of building a dual-branch network is as follows: The first branch is a MobileNet network based on the FPN feature pyramid structure, which includes an initial convolutional layer that performs a 3×3 convolution operation on the input image, transforming the number of input channels from 3 to the predetermined number of channels 128. Multiple FPN layers are designed, each consisting of a 3×3 convolution, batch normalization, and ReLU activation function. Maximum pooling is then used to achieve layer-by-layer downsampling, generating multi-scale feature maps with shapes of (H / 2, W / 2), (H / 4, W / 4), and (H / 8, W / 8). The second branch uses the Swin Transformer encoder structure, which includes a Stem convolution layer to convert the input blurred image to the same number of channels as the first branch; multi-layer Swin Transformer modules, each of which includes relative position embedding and window-based multi-head self-attention calculation. 1×1 convolution is used to generate the query matrix Q, key matrix K, and value matrix V respectively, and the scaling factor and softmax function are used to obtain the attention distribution to achieve feature enhancement within the local window; For the feature maps of different scales output by the two branches, they are normalized using 1×1 convolution to obtain the same number of channels, and feature fusion is performed; the fused feature maps are then input into the decoder module, which uses layer-by-layer upsampling and convolution operations to restore high-resolution features and ultimately reconstruct a clear image.
5. The improved DeblurGAN UAV aerial image deblurring method according to claim 1 is characterized in that: In step S4, the generator network loss function is set as follows: L total =λ gan ·L gan +λ perceive ·L perceive +λ pixel ·L pixel Among them, λ gan ,λ perceive ,λ pixel Denote the adversarial loss L gan , Perceptual Loss L perceive and pixel loss L pixel The weight coefficient is set to 2:3:
5. The objective loss function of the generative adversarial network is as follows: First input the sample and sample the real data sample distribution P data A batch of samples {X1,X2,...,X N }, from the noise distribution P z Sampling a batch of noise samples {z1,z2,...,z M }, represents the expected value calculated from the real sample, Represents the expectation of the noise distribution calculation; then calculate the output of the discriminator, D(x) represents the output of the discriminator for the real sample x, G(z) represents the output of the generator, the first term of the formula is to calculate the difference between the discriminator output of the real sample x and the generated sample G(z), the second term focuses on the discriminator output of the generated sample, and tries to minimize D(G(z)) and Ideally, D(x) should be close to 1, indicating that the real sample is identified as real by the discriminator, and D(G(z)) should be close to 0, indicating that the generated sample is identified as false.
Citation Information
Cited By
Double-domain Swin Mama-based generative adversarial network MRI reconstruction method
CN120635248A
Unmanned aerial vehicle image enhancement method and system
CN120953059A
A method and system for image enhancement of a drone
CN120953059B
Image deblurring method and device
CN121582101A
An image deblurring method and apparatus
CN121582101B