Corn disease identification method based on improved generative adversarial network

By improving the generative adversarial network model to perform super-resolution reconstruction of low-resolution maize leaf images collected by UAVs, the problems of low image quality and insufficient recognition accuracy in high-altitude UAV remote sensing were solved, and high-precision maize disease identification was achieved.

CN121236639APending Publication Date: 2025-12-30HENAN UNIV OF ECONOMICS & LAW

Patent Information

Application Number
CN202511359402.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing methods for identifying maize leaf diseases suffer from problems such as low image quality, blurred leaf edges, complex background interference, weak robustness, poor generalization ability, and low recognition accuracy in high-altitude UAV remote sensing images. In particular, the ability to distinguish between soil color and texture features and spectral features of lesion symptoms decreases in low-resolution images, leading to a significant increase in false positive and false negative rates.

Method used

An improved generative adversarial network (ERSCA-WGAN) model is used to perform super-resolution reconstruction on the initial low-resolution image. By fusing RRDB nested residual dense blocks, spatial channel attention mechanism and multi-scale texture enhancement module, image details and texture information are restored. An adaptive hybrid upsampling module is combined to improve image clarity, and an image classification model is used for disease prediction.

Benefits of technology

It significantly improved the accuracy of identifying maize leaf rust, reduced the misclassification rate between mild and moderate disease, optimized the problems of background interference and blurred lesion edges, and improved the accuracy and robustness of disease identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236639A_ABST
    Figure CN121236639A_ABST
Patent Text Reader

Abstract

The invention discloses a corn disease recognition method based on an improved generative adversarial network. The corn disease recognition method comprises the steps that S1, an initial low-resolution image of a corn field is collected and obtained through an unmanned aerial vehicle; s2, performing super-resolution reconstruction on the initial low-resolution image by using an improved generative adversarial network model; the model construction comprises the following steps: S2.1, constructing a shallow feature extraction layer; s2.2, constructing a deep feature extraction network based on a plurality of RRDB nested residual dense blocks; s2.3, constructing an attention module based on a space and channel dual attention mechanism; s2.4, a multi-scale texture enhancement module is constructed through multi-scale convolution and smooth branches; s2.5, constructing a global residual connection layer; s2.6, constructing an adaptive hybrid up-sampling module based on transposed convolution and stable up-sampling; s2.7, performing mapping output on the features after up-sampling; and S3, carrying out disease prediction on the high-resolution reconstructed image. According to the method, details such as spatial resolution and texture of the unmanned aerial vehicle high-altitude flight remote sensing image are improved, and then the corn disease monitoring precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of agricultural remote sensing and artificial intelligence, specifically to a method for identifying maize diseases based on an improved generative adversarial network. Background Technology

[0002] Maize is one of the world's highest-yielding and most widely planted food crops. However, various biological stresses, especially fungal diseases, seriously threaten the healthy growth and yield of maize. Among them, maize leaf rust, a fungal disease that is prevalent worldwide, can infect leaves, hinder photosynthesis, and thus lead to weakened crop growth.

[0003] With the vigorous development of precision agriculture and smart agriculture technologies, crop disease intelligent identification technology based on remote sensing platforms has become an important development direction in the field of modern agricultural disease monitoring due to its significant advantages such as high efficiency, non-contact, macro-monitoring capabilities, and economy, providing a new way to grasp the occurrence of crop diseases in a timely and accurate manner.

[0004] In recent years, research on crop disease monitoring based on UAV remote sensing imagery has received widespread attention and made significant progress. Although UAV imagery has demonstrated significant applicability and feasibility in crop disease detection, its application still faces a series of challenges. One core issue is the trade-off between flight altitude and image quality. The spatial resolution of UAV remote sensing images is mainly determined by the instantaneous field of view (IFOV) and platform altitude; for the same sensor, platform flight altitude and spatial resolution are inversely proportional. Specifically, when UAVs fly at low altitudes, although they can acquire high-resolution images, large geometric distortions, inconsistent brightness differences, and overlapping areas can easily lead to artifacts such as seams and ghosting during image stitching, severely affecting the effectiveness of continuous monitoring of large areas of farmland. Conversely, while high-altitude flight can significantly improve operational coverage efficiency and image stitching quality, it faces the problem of insufficient spatial resolution. This results in blurred lesion details and loss of edge information in the images, making it difficult to meet the needs of accurate disease identification and detailed assessment. Furthermore, the low-resolution problem caused by high-altitude flight can exacerbate interference from complex soil backgrounds. Due to reduced image resolution, the distinguishability between soil color and texture features and the spectral characteristics of lesion symptoms decreases, especially in the early stages of disease, where slight leaf discoloration becomes more difficult to differentiate from the soil background. This leads to a significant increase in false positive and false negative rates in disease identification. Furthermore, blurred leaf edges are more pronounced in aerial imagery. Due to focal length limitations and reduced resolution, the edges of crop leaves become indistinct, resulting in a significant loss of lesion boundary information. This poses a significant challenge to accurate disease grading and quantitative assessment. These challenges underscore the critical importance of improving image quality for accurate disease identification in the context of high-altitude UAV remote sensing.

[0005] In the prior art, Chinese invention patent with publication number CN116630960A discloses a method for identifying maize diseases based on a texture-color multi-scale residual shrinkage network, including the following steps: acquiring mixed maize leaf disease images from experimental and field environments, and performing image enhancement; extracting local binary pattern feature maps from training images through a texture feature extraction module; inputting the RGB features and texture features of the training images into a texture-color two-branch shallow feature extraction module at a certain ratio; connecting the outputs of the two branches after adjusting their weight ratios; establishing a multi-scale residual shrinkage module by combining a soft thresholding function, an attention mechanism, and multi-scale convolution; and sequentially passing the outputs through two interconnected multi-scale residual shrinkage modules; adjusting the network model parameters according to the semantic information of the output images and the corresponding categories of maize leaf disease images to obtain a trained network model. While the aforementioned methods achieve disease identification in maize leaves, they rely on feature extraction, construction, and fusion to improve recognition accuracy. When used for maize lesion identification, they directly employ clear RGB images. However, direct identification of low-resolution RGB images acquired by drones leads to a decrease in the distinguishability between soil color and texture features and the spectral characteristics of lesion symptoms, resulting in a significant increase in false positive and false negative rates. Therefore, the above scheme cannot directly identify diseases from low-resolution images acquired by drones, nor can it resolve the contradiction between high-altitude flight with low resolution and low-altitude flight with low efficiency.

[0006] In summary, existing methods for identifying corn leaf diseases still suffer from problems such as low image quality, blurred leaf edges, complex background interference, weak robustness, poor generalization ability, and low recognition accuracy in actual high-altitude UAV remote sensing image disease detection. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art, improve the resolution of high-altitude UAV remote sensing images, optimize background interference and lesion edge blurring in maize rust monitoring, enrich the details and texture features of maize leaves, and provide a maize disease identification method based on an improved generative adversarial network.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: A method for identifying maize diseases based on an improved generative adversarial network includes the following steps: S1: Use drones to collect and acquire remote sensing images of cornfields, preprocess the remote sensing images to obtain initial low-resolution images with pixel size N×N; S2: Use a pre-built improved generative adversarial network model to perform super-resolution reconstruction on the initial low-resolution image to obtain a high-resolution reconstructed image with a pixel size of R×R; The method for constructing the improved generative adversarial network model is as follows: S2.1: Construct a shallow feature extraction layer to extract shallow features from the initial low-resolution input image; S2.2: Construct a deep feature extraction network based on multiple nested residual dense blocks of RRDB to extract deep features; S2.3: An attention module is constructed based on a dual attention mechanism of spatial and channel, which is used to perform feature recalibration and key region localization on the extracted deep features; S2.4: A multi-scale texture enhancement module is constructed through multi-scale convolution and smoothing branches to enhance the texture of deep features output by the attention module and extract multi-scale texture features. The smoothing branches are used to suppress high-frequency noise. S2.5: Construct a global residual connection layer to perform global residual connections on shallow features and deep features output by the multi-scale texture enhancement module to obtain fused features; S2.6: An adaptive hybrid upsampling module is constructed based on transposed convolution and stable upsampling. The adaptive hybrid upsampling module has two layers, which are used to perform two-level upsampling on the fused features. S2.7: Construct an output layer to map the upsampled features and output the high-resolution reconstructed image; S3: The trained image classification model is used to predict diseases in high-resolution reconstructed images to obtain disease prediction results.

[0009] This invention utilizes an improved generative adversarial network model (ERSCA-WGAN) that integrates nested residual dense blocks of RRDB, spatial channel attention mechanism, and multi-scale texture enhancement to reconstruct the initial low-resolution image of maize leaves collected and processed by UAVs. This generates a clear, high-quality, and detail-rich high-resolution reconstructed image, improving the clarity of the remote sensing image. Based on the integration of nested residual dense blocks of RRDB and spatial channel attention mechanism, image details and texture information are significantly restored. Based on the multi-scale texture enhancement module and adaptive hybrid upsampling module, the accuracy of maize leaf rust identification is improved, and the misclassification rate between mild and moderate diseases can be effectively reduced, thereby improving the accuracy of disease identification.

[0010] Preferably, the preprocessing method in step S1 includes: By performing image stitching, geometric correction, and coordinate transformation on remote sensing images, standard orthophotos are obtained. The standard orthophoto is divided into grids and cropped into an initial low-resolution image.

[0011] Preferably, the deep feature extraction network in step S2.2 includes multiple RRDB nested residual dense blocks in series. Each RRDB nested residual dense block includes multiple residual dense blocks. The multiple residual dense blocks are connected by macroscopic residuals and form a multi-level residual structure through a residual scaling mechanism. The expression for each RRDB nested residual dense block is: ; Where i is the index of the nested residual dense block of RRDB, β is the residual scaling factor, x represents the input image, and RDB1, RDB2 and RDB3 represent the first residual dense block, the second residual dense block and the third residual dense block, respectively.

[0012] Preferably, the attention module in step S2.3 includes a channel attention layer and a spatial attention layer connected in sequence, wherein the channel attention layer is used to recalibrate features and the spatial attention layer is used to locate key regions; The expression for the channel attention layer is: ; Where CA is the channel attention function, σ is the sigmoid activation function, MLP is the multilayer perceptron, and GAP and GMP are the global average pooling and max pooling operations, respectively. The expression for the spatial attention layer is: ; Where SA is the spatial attention function, Conv 7×7 This represents a 7×7 convolution operation. GAP c and GMP c These represent average pooling and max pooling along the channel dimension, respectively, and [;] denotes the feature concatenation operation.

[0013] Preferably, the multi-scale texture enhancement module in step S2.4 is used to fuse the input features after multi-scale convolution and smoothing branching, and perform residual connection with the input features. The expression of the multi-scale texture enhancement module is: ; ; Where TE is the texture enhancement function, α is the residual connection weight coefficient, Fusion is the feature fusion operation, and Conv... k This represents a k×k convolution operation, where k takes values ​​of 1, 3, or 5; Smooth is the smoothing branch function, and AvgPool is the average pooling operation.

[0014] Preferably, the expression for the adaptive hybrid upsampling module in step S2.6 is: ; Where AdaUp is the adaptive hybrid upsampling function. ω 1 and ω Both are dynamically generated adaptive weights, ConvT is transposed convolution, and StableUp is stable upsampling; The expression for dynamically generated adaptive weights is: ; Here, Softmax is the activation function, Conv is the convolution function, and GAP is the global average pooling function.

[0015] Preferably, the image classification model in step S3 adopts the MobileNetV3 model or the InceptionV3 model.

[0016] Preferably, the improved generative adversarial network model is trained using a super-resolution reconstruction dataset, and the image classification model is trained using a disease classification dataset. The methods for obtaining the super-resolution reconstruction dataset and the disease classification dataset are as follows: Historical images acquired at a preset flight altitude are obtained, and preprocessed by correction and cropping to obtain historical high-resolution images with pixel dimensions of R×R; A historical high-resolution image is degraded to generate a low-resolution degraded image with a pixel size of N×N. The degradation process is used to reduce the resolution of the historical high-resolution image. The historical high-resolution image is used as the high-resolution ground truth image after undergoing the first data augmentation, and the low-resolution degraded image corresponding to the historical high-resolution image is used as the low-resolution input image after undergoing the first data augmentation. The high-resolution ground truth image and the low-resolution input image are used as super-resolution reconstruction samples to construct a super-resolution reconstruction dataset. The first data augmentation includes original image preservation, horizontal flipping, vertical flipping and rotation transformation. Disease severity levels are labeled on historical high-resolution images, and second data augmentation is performed to obtain disease classification samples. Disease classification datasets are then constructed from these disease classification samples. The second data enhancement includes original image preservation, brightness enhancement, contrast enhancement, rotation transformation, horizontal flip, and vertical flip.

[0017] Preferably, the total loss function used when training the improved generative adversarial network model includes adversarial loss, content loss, perceptual loss, and edge loss; The adversarial loss is constructed based on the WGAN-GP framework and includes a discriminator loss and a generator loss, with the discriminator loss containing a gradient penalty term; The content loss is calculated using the L1 norm to determine the pixel-level difference between the high-resolution reconstructed image and the high-resolution ground truth image. The perceptual loss is constructed based on a pre-trained VGG-19 network to extract multi-layer features; The edge loss is constructed by combining the Sobel and Laplacian operators and is used to extract multi-scale edge information.

[0018] Preferably, the expression for the discriminator loss is: ; in, For discriminator loss, Let D be the expectation function, and D be the discriminator. x fake and x real These are generated samples and real samples, respectively. The interpolated sample is obtained by random interpolation of the real sample and the generated sample. The gradient penalty coefficient is... This represents the gradient of the discriminator relative to the interpolated samples. Represents the L2 norm; The generator loss for: ; The expression for the perceptual loss is: ; in, Let G be the perceptual loss function and G be the generator. x lr and x hr These are a low-resolution input image and a high-resolution target image, respectively. Let M represent the VGG feature extraction function of the i-th layer, where M is the total number of feature layers in the pre-trained VGG-19 network. Let be the weight coefficients of the features in the i-th layer. Represents the L1 norm; The expression for the edge loss is: ; ; in, Sobel is the edge loss function. x and Sobel y ε represents the horizontal and vertical Sobel operators, respectively; Laplacian is the Laplacian operator; ε is the combination weight; and Edge is the edge detection function.

[0019] The beneficial effects of this invention are: This invention establishes a super-resolution reconstruction dataset to train an improved generative adversarial network model. The sample size is expanded by first data augmentation, and a degradation operation is used to generate pairs of low-resolution degraded images and high-resolution ground truth images. The adversarial loss, content loss, perceptual loss and edge loss complement each other to work together, which can balance the generation quality, perceptual realism and training stability.

[0020] This invention establishes a disease classification dataset to train an image classification model, distinguishes disease levels through manual annotation, and expands the sample size through second data augmentation to improve the recognition ability of the image classification model.

[0021] This invention achieves high precision and high quality in the application of high-altitude UAV remote sensing technology for monitoring corn leaf rust, improves image resolution and disease identification accuracy, and effectively optimizes the problems of background interference and blurred lesion edges in corn rust monitoring by improving the generative adversarial network model (ERSCA-WGAN), thereby significantly enhancing the accuracy and robustness of corn rust identification and providing important technical support for precision agriculture and intelligent plant protection. Attached Figure Description

[0022] The present invention will now be described in further detail with reference to the accompanying drawings: Figure 1 This is a method block diagram of the present invention; Figure 2 This is a schematic diagram of the structure of the improved generative adversarial network model of the present invention; Figure 3 This is a comparison of the reconstructed images generated during training by the improved generative adversarial network model of this invention with those of other models; Figure 4 This invention relates to a contrast confusion matrix between the reconstructed image generated during training of the improved generative adversarial network model and other models. Figure 5 The image classification model of this invention displays the ROC curve, loss curve, and accuracy curve for the four categories of healthy, mild, moderate, and severe. Figure 6 This is a comparison image of the initial low-resolution image and the high-resolution reconstructed image of the present invention; Figure 7 This is the contrast confusion matrix between the initial low-resolution image and the high-resolution reconstructed image of this invention; Figure 8 This is a comparison diagram of the distribution mapping effects of the initial low-resolution image and the high-resolution reconstructed image of the present invention. Detailed Implementation

[0023] like Figure 1As shown, the present invention provides a method for identifying maize diseases based on an improved generative adversarial network, comprising the following steps: S1: Use drones to collect and acquire remote sensing images of cornfields, preprocess the remote sensing images to obtain initial low-resolution images with pixel size N×N; The preprocessing methods include: image stitching, geometric correction and coordinate transformation of remote sensing images to obtain standard orthophotos; grid segmentation of the standard orthophotos and cropping to obtain initial low-resolution images with a pixel size of N×N.

[0024] In this embodiment, a drone equipped with an RGB sensor is used to take aerial photos of a cornfield in the milk-ripe stage to obtain remote sensing images.

[0025] The remote sensing images were mosaicked and geometrically corrected using Pix4Dmapper software to generate a digital orthophoto (DOM). Then, coordinate system transformation was performed using ENVI software to obtain a standard orthophoto projected in WGS84_UTM_Zone 49N. The standard orthophoto was then divided into regular grids, with a sliding window moving sequentially from left to right and from top to bottom without overlap, cropping it into 32×32 pixel TIF format image blocks to obtain the initial low-resolution image. N was set to 32.

[0026] S2: Use a pre-built improved generative adversarial network model (ERSCA-WGAN) to perform super-resolution reconstruction on the initial low-resolution image to obtain a high-resolution reconstructed image with a pixel size of R×R.

[0027] like Figure 2 As shown, the construction method of the improved generative adversarial network model (ERSCA-WGAN) is as follows: S2.1: Construct a shallow feature extraction layer to extract shallow features from the initial low-resolution input image.

[0028] S2.2: Construct a deep feature extraction network based on multiple nested residual dense blocks of RRDB to extract deep features.

[0029] The deep feature extraction network consists of multiple RRDB nested residual dense blocks (abbreviated as RRDB blocks) connected in series. These multiple RRDB nested residual dense blocks are then connected to convolutional layers. Each RRDB nested residual dense block includes multiple residual dense blocks. These multiple residual dense blocks are connected using macroscopic residuals and form a multi-level residual structure through a residual scaling mechanism.

[0030] The expression for each RRDB nested residual dense block is: ; Where i is the index of the nested residual dense block of RRDB, β is the residual scaling factor, x represents the input image, and RDB1, RDB2 and RDB3 represent the first residual dense block, the second residual dense block and the third residual dense block, respectively.

[0031] Deep feature extraction employs a dense connection mechanism, which effectively alleviates the gradient vanishing problem and enhances feature propagation capability through multi-level feature reuse.

[0032] S2.3: An attention module is constructed based on a dual attention mechanism of spatial and channel to perform feature recalibration and key region localization on the extracted deep features.

[0033] The attention module consists of a channel attention layer and a spatial attention layer connected in sequence. The channel attention layer is used to recalibrate features, and the spatial attention layer is used to locate key regions.

[0034] The channel attention layer captures the dependencies between channels through global statistics, and its expression is: ; Where CA is the channel attention function, σ is the sigmoid activation function, MLP is the multilayer perceptron, and GAP and GMP are the global average pooling and max pooling operations, respectively.

[0035] The spatial attention layer generates a spatial attention map by aggregating the spatial information of the feature maps. Its expression is: ; Where SA is the spatial attention function, Conv 7×7 This represents a 7×7 convolution operation. GAP c and GMP c These represent average pooling and max pooling along the channel dimension, respectively, and [;] denotes the feature concatenation operation.

[0036] The attention module adopts a serial fusion spatial-channel attention mechanism. First, the features are recalibrated through the channel attention layer, and then the key areas are located through the spatial attention layer, so that the model can pay more attention to the key areas when inferring the level of rust infection.

[0037] S2.4: A multi-scale texture enhancement module is constructed using multi-scale convolution and smoothing branches to enhance the texture of deep features output by the attention module, extract multi-scale texture features, and suppress high-frequency noise. The multi-scale texture enhancement module uses multiple convolutional kernels with different receptive fields for parallel processing, while introducing smoothing branches to suppress high-frequency noise.

[0038] The multi-scale texture enhancement module is used to fuse the input features processed by multi-scale convolution and smoothing branches, and then perform residual connections with the input features. The expression for the multi-scale texture enhancement module is: ; ; Where TE is the texture enhancement function, α is the residual connection weight coefficient, Fusion is the feature fusion operation, and Conv... k This represents a k×k convolution operation, where k takes values ​​of 1, 3, or 5; Smooth is the smoothing branch function, and AvgPool is the average pooling operation.

[0039] The multi-scale texture enhancement module enhances texture details while maintaining overall structural stability through multi-scale feature fusion, effectively balancing detail restoration and noise suppression.

[0040] S2.5: Construct a global residual connection layer to perform global residual connections between shallow features and deep features output by the multi-scale texture enhancement module, obtaining fused features. In this embodiment, the global residual connection layer adds the outputs of the shallow feature extraction layer and the multi-scale texture enhancement module element-wise.

[0041] S2.6: An adaptive hybrid upsampling module is constructed based on transposed convolution and stable upsampling. The adaptive hybrid upsampling module has two layers, which are used to perform two-level upsampling on the fused features.

[0042] The expression for the adaptive hybrid upsampling module in step S2.6 is: ; Where AdaUp is the adaptive hybrid upsampling function. ω 1 and ω Both are dynamically generated adaptive weights, ConvT is transposed convolution, and StableUp is stable upsampling; The expression for dynamically generated adaptive weights is: ; Here, Softmax is the activation function, Conv is the convolution function, and GAP is the global average pooling function.

[0043] Stable upsampling employs an improved ICNR (Initialization for Checkerboard Removal) initialization strategy and multi-stage processing, effectively reducing checkerboard artifacts and high-frequency noise.

[0044] S2.7: Construct the output layer to map and output the upsampled features.

[0045] The improved generative adversarial network model (ERSCA-WGAN) is trained using a super-resolution reconstruction dataset. The method for obtaining the super-resolution reconstruction dataset is as follows: Historical images acquired at a preset flight altitude are obtained, and preprocessed by correction and cropping to obtain historical high-resolution images with pixel dimensions of R×R. In this embodiment, the preset flight altitude is 25m, and R is set to 128.

[0046] The correction preprocessing includes image stitching, geometric correction and coordinate transformation. The cropping preprocessing uses grid segmentation to crop the historically acquired images after correction preprocessing into historical high-resolution images of 128×128 pixels in TIF format.

[0047] A low-resolution degraded image is generated by degrading a historical high-resolution image. The pixel size is N×N. The degradation process is used to reduce the resolution of the historical high-resolution image. In this embodiment, the degradation process includes Gaussian blur and bicubic downsampling. The pixel size of the low-resolution degraded image is 32×32.

[0048] The historical high-resolution image is used as the high-resolution ground truth image after undergoing the first data augmentation. The low-resolution degraded image corresponding to the historical high-resolution image is used as the low-resolution input image after undergoing the first data augmentation. The high-resolution ground truth image and the low-resolution input image are used as super-resolution reconstruction samples to construct a super-resolution reconstruction dataset. The first data augmentation includes original image preservation, horizontal flipping, vertical flipping and rotation transformation.

[0049] To construct a super-resolution optimization framework suitable for complex agricultural images, this invention employs four complementary loss terms working synergistically. Therefore, the total loss function used when training the improved generative adversarial network model (ERSCA-WGAN) includes adversarial loss, content loss, perceptual loss, and edge loss to balance generation quality, perceptual realism, and training stability.

[0050] Total loss function The expression is: ; in, and These are the adversarial loss function, content loss function, perceptual loss function, and edge loss function, respectively. and These are the weight coefficients corresponding to the adversarial loss function, content loss function, perceptual loss function, and edge loss function, respectively.

[0051] The adversarial loss is constructed based on the WGAN-GP framework and includes discriminator loss and generator loss. The discriminator loss contains a gradient penalty term to satisfy the Lipschitz constraint. The expression for the discriminator loss in adversarial loss is: ; in, For discriminator loss, Let D be the expectation function, and D be the discriminator. x fake and x real These are generated samples and real samples, respectively. The interpolated sample is obtained by random interpolation of the real sample and the generated sample. The gradient penalty coefficient is... This represents the gradient of the discriminator relative to the interpolated samples. This represents the L2 norm. During training, generated samples and ground truth samples refer to the high-resolution reconstructed image and the high-resolution ground truth image, respectively.

[0052] Generator loss for: ; The generator loss is used to maximize the discriminator's score on the generated samples.

[0053] Content loss is calculated using the L1 norm to determine the pixel-level differences between the high-resolution reconstructed image and the high-resolution ground truth image.

[0054] The perceptual loss is constructed based on a pre-trained VGG-19 network to extract multi-layer features.

[0055] The expression for perceived loss is: ; in, Let G be the perceptual loss function and G be the generator. x lr and x hr These are a low-resolution input image and a high-resolution target image, respectively. Let M represent the VGG feature extraction function of the i-th layer, where M is the total number of feature layers in the pre-trained VGG-19 network. Let be the weight coefficients of the features in the i-th layer. This represents the L1 norm. During training, the high-resolution target image is the same as the aforementioned high-resolution ground truth image. In this embodiment, high-level features have greater weight to emphasize semantic information.

[0056] The edge loss is constructed by combining the Sobel and Laplacian operators to extract multi-scale edge information.

[0057] The expression for edge loss is: ; ; in, Sobel is the edge loss function. x and Sobel y ε represents the horizontal and vertical Sobel operators, respectively; Laplacian is the Laplacian operator; ε is the combination weight; and Edge is the edge detection function.

[0058] S3: The trained image classification model is used to predict diseases in high-resolution reconstructed images to obtain disease prediction results.

[0059] Image classification models can be EfficientNetV2, MobileNetV3, VGG16, ResNet50, or InceptionV3.

[0060] The image classification model is trained using a disease classification dataset. The disease classification dataset is obtained by labeling the historical high-resolution images with disease levels and performing a second data augmentation to obtain disease classification samples, which are then used to construct the disease classification dataset.

[0061] The method for labeling disease severity levels is as follows: Local standards and / or authoritative labeling methods are followed. In this embodiment, the Henan Provincial Standard DB41 / T 2668-2024 and existing research (the paper published in 2024 by Lv, Z et al., entitled "Improved monitoring of southern corn rust using UAV-based multi-view imagery and an attention-based deep learning method," journal: Computers and Electronics in Agriculture, 224, 109232) are followed. A quality control mechanism combining independent labeling by two individuals and review by agricultural experts is employed to ensure labeling accuracy. Based on the proportion of lesion area, color depth, boundary clarity, and overall leaf condition, historical high-resolution image samples are divided into four levels: Healthy, Slight, Moderate, and Severe.

[0062] The second data enhancement includes original image preservation, brightness enhancement, contrast enhancement, rotation transformation, horizontal flip, and vertical flip. In this embodiment, the brightness is enhanced by 1.2 times, and the rotation transformation angle is 15 degrees.

[0063] S4: Performance evaluation of the improved generative adversarial network model (ERSCA-WGAN) and the image classification model. Specifically, the performance of the improved generative adversarial network model (ERSCA-WGAN) is evaluated using peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) metrics.

[0064] PSNR is used to measure the pixel-level difference between a high-resolution reconstructed image and a high-resolution ground truth image. Its calculation formula is as follows: ; ; Where W, H, and C represent the width, height, and number of channels of the image, respectively; X represents the high-resolution reconstructed image; Y represents the high-resolution ground truth image; and MAX is the maximum possible value of a pixel.

[0065] SSIM is used to evaluate the structural similarity of images, and its formula is: ; in, μ X and μ Y Let σ be the mean of image X and Y, respectively. X and σ Y Their variances, σ and σ, are respectively. XY For covariance, C1 and C2 are stability constants to avoid division by zero.

[0066] Image classification models are evaluated using classification performance metrics, including precision, accuracy, recall, F1 score, error rate, and error rate reduction rate.

[0067] Precision measures the proportion of actual positive classes out of those predicted as positive; Accuracy represents the proportion of correct predictions out of all predictions; Recall measures the proportion of actual positive classes that are correctly predicted; F1 Score is the harmonic mean of precision and recall, comprehensively reflecting classification performance; Error Rate and Error Reduction are used to more intuitively analyze the model's ability to distinguish between adjacent levels.

[0068] Error rate is used to measure the proportion of errors in a model's predictions, and it is defined as: ; Where ErrorRate is the error rate, and TP, TN, FP, and FN represent true positives, true negatives, false positives, and false negatives, respectively.

[0069] Error rate is calculated by treating any pair of adjacent grades as a binary sub-problem and calculating the error rate between adjacent grades accordingly.

[0070] Error rate reduction rate measures the relative decrease in error rate of the improved model compared to the baseline model, and is defined as follows: ; Wherein, ErrorReduction is the error rate reduction rate. E baseline and E model These represent the error rates of the baseline model and the improved model, respectively.

[0071] By combining the error rate and the error rate reduction rate, the improvement of the model in fine-grained discrimination can be quantitatively reflected in adjacent-level sub-problems of multi-classification tasks.

[0072] This invention uses the PyTorch 2.2.1 deep learning framework to build and train the model, and the verification environment is a Windows 10 x64 operating system. In the optimized optimizer configuration of the improved generative adversarial network model, the generator and discriminator each use independent Adam optimizers. The training parameters are configured as follows: the training batch size and the training ratio of the discriminator to the generator are set to 16 and 2:1, respectively. The generator and discriminator each use independent Adam optimizers, with initial learning rates of 2 × 10⁻⁶. -4 and 1×10 -4 The momentum parameters β1 and β2 are 0.5 and 0.9, respectively. The learning rate adjustment strategy employs a cosine annealing restart method, with an initial restart cycle of 50 epochs (iteration rounds), a cycle multiplication factor of 2, and a minimum learning rate threshold of 1×10⁻⁶. -6 .

[0073] To verify the training performance of the improved generative adversarial network model (ERSCA-WGAN) proposed in this invention, super-resolution reconstruction was performed on low-resolution input images using Bicubic (bicubic interpolation), WGAN-GP (Wasserstein generative adversarial network with gradient penalty), ESRGAN (enhanced super-resolution generative adversarial network), Real-ESRGAN (real-scene enhanced super-resolution generative adversarial network), and the proposed ERSCA-WGAN method. The performance evaluation results are shown in Table 1. As can be seen from Table 1, the ERSCA-WGAN method of this invention achieved the highest peak signal-to-noise ratio (PSNR) (20.96 dB) and structural similarity (66.67%). Real-ESRGAN ranked second with 20.57 dB and 64.36%, respectively, outperforming Bicubic's 19.77 dB and 58.47%. WGAN-GP and ESRGAN had relatively lower values, at 19.40 dB and 55.73%, and 18.41 dB and 52.23%, respectively.

[0074] Table 1. Super-resolution results of deep learning methods (×4x) To further verify the training effect of the improved generative adversarial network model (ERSCA-WGAN) of this invention, low-resolution input images of maize leaves at four rust severity levels—healthy leaves, lightly infected, moderately infected, and severely infected—and their corresponding high-resolution ground truth images were used for evaluation. Figure 3 As shown, in terms of visual presentation, Figure 3This paper presents reconstruction results for four rust severity levels: healthy leaves, lightly infected, moderately infected, and severely infected. The images are presented in the following order: low-resolution input image (LR), high-resolution ground truth image (HR), and reconstructed images obtained using Bicubic, ESRGAN, Real-ESRGAN, WGAN-GP, and ERSCA-WGAN methods. The reconstructed image obtained using the Bicubic method is the most blurry, with significant loss of detail and edge information. The reconstructed image obtained using the WGAN-GP method shows improvement over the Bicubic method, but exhibits noticeable noise, especially on the leaf surface and in the background area. The reconstructed image obtained using the ESRGAN method outperforms the WGAN-GP method in texture restoration, but its overall sharpness is insufficient. The reconstructed image obtained using the Real-ESRGAN method performs well in color and shape preservation, but the accurate reconstruction of leaf boundaries is still inadequate. In contrast, the visual result of the reconstructed image obtained using the ERSCA-WGAN method shows more continuous and clear lines, with fewer jagged edges and blurring, providing a more reliable image basis for the accurate identification and assessment of rust severity.

[0075] To improve the accuracy of corn disease identification, the image classification model in this embodiment adopts the InceptionV3 model.

[0076] To verify the differences in maize disease identification among EfficientNetV2, MobileNetV3, VGG16, ResNet50, and InceptionV3 models, low-resolution input images of 32×32, reconstructed images of 128×128 obtained from the low-resolution input images using the ERSCA-WGAN method, and high-resolution ground truth images (HR) of 128×128 were respectively input into the MobileNetV3, EfficientNetV2, VGG16, ResNet50, and InceptionV3 models. The performance evaluation comparison data are shown in Table 2.

[0077] Table 2. Comparison of maize disease classification performance under different resolution conditions As shown in Table 2, the ERSCA-WGAN method of this invention can significantly improve classification performance. After super-resolution reconstruction of low-resolution input images, the accuracy of all image classification models improved, with InceptionV3 showing the largest increase, from 90.89% to 93.81%, an improvement of 2.92 percentage points; MobileNetV3 followed closely, from 90.03% to 93.71%, an improvement of 3.68 percentage points, accompanied by a significant improvement in F1-Score. VGG16, ResNet50, and EfficientNetV2 models also improved by 1.37, 0.69, and 1.72 percentage points, respectively.

[0078] When the input is a low-resolution image, MobileNetV3 achieves the best accuracy of 90.89%. When the input is a reconstructed image obtained from a low-resolution input image using the ERSCA-WGAN method, both MobileNetV3 and InceptionV3 exceed 93%, showing a significant improvement in overall performance. Compared to high-resolution ground truth images, the classification performance of the reconstructed images is close to that of high-resolution ground truth images in most cases. In particular, InceptionV3 achieves an accuracy of 95.69% on high-resolution ground truth images, indicating that the ERSCA-WGAN method has significantly narrowed the performance gap with high-resolution ground truth images. In conclusion, the ERSCA-WGAN method can effectively reconstruct low-resolution input images of 32×32 resolution and significantly improve the accuracy of disease identification.

[0079] A low-resolution input image of 32×32, a reconstructed image of 128×128 obtained from the low-resolution input image using the ERSCA-WGAN method, and a high-resolution ground truth image (HR) of 128×128 are input into MobileNetV3, EfficientNetV2, VGG16, ResNet50, and InceptionV3 models, respectively. The confusion matrices of the five image classification models in identifying the corn rust level on the reconstructed image are as follows. Figure 4 As shown. From Figure 4It can be seen that MobileNetV3 performs best in identifying severe rust, with an accuracy of 93.52%, and is tied for first place with InceptionV3 in classifying healthy samples, with an accuracy of 98.94%. EfficientNetV2 performs evenly across all levels, with accuracies of 91.67%, 88.24%, 88.80%, and 97.87% for healthy, light, moderate, and severe rust, respectively. ResNet50 has an accuracy of 89.60% in classifying moderate rust, close to InceptionV3, but only 82.75% for light rust, the lowest among the five models. VGG16 has weak overall performance, with accuracy of 85.19% for healthy samples and 80.80% for moderate rust, both the lowest values.

[0080] Meanwhile, in terms of error rate and error reduction rate, InceptionV3's overall error rate is 6.19%, significantly lower than MobileNetV3's 8.42%, EfficientNetV2's 9.45%, ResNet50's 14.43%, and VGG16's 14.78%. Compared to these models, its error reduction rate reaches 26.5%, 34.5%, 57.1%, and 58.1%, respectively, demonstrating the best overall performance.

[0081] In fine-grained differentiation of adjacent disease levels, InceptionV3 has an error rate of 5.23% in the healthy-mild differentiation, which is lower than the other four models; the error rate in the mild-moderate differentiation is only 1.32%, which is particularly significant; and the error rate in the moderate-severe differentiation is 4.57%, which is better than ResNet50 and VGG16, and slightly higher than EfficientNetV2 and MobileNetV3.

[0082] Furthermore, using InceptionV3 as a benchmark, the error rate reduction rate was calculated. In the healthy-mild distinction, it reduced the error rate by 57.4%, 48.9%, 44.6%, and 30.2% compared to VGG16, ResNet50, EfficientNetV2, and MobileNetV3, respectively. The advantage was most prominent in the mild-moderate distinction, with reductions of 82.9%, 85.3%, 67.3%, and 70.8%, respectively. In the moderate-severe distinction, it still reduced the error rate by about 35% compared to ResNet50 and VGG16.

[0083] Overall, the InceptionV3 model demonstrates significant advantages in both global classification and fine-grained recognition of adjacent levels, with the most notable improvement in recognition ability between mild and moderate cases.

[0084] Figure 5 (a) The ROC curves of the InceptionV3 model across four categories: healthy, mild, moderate, and severe. Figure 5 As can be seen, the AUC (Area Under the Curve) values ​​for different disease levels range from 0.988 to 0.998, while the AUC values ​​for the micro-average and macro-average reach 0.991 and 0.993 respectively, which are close to 1. Figure 5 (b) demonstrates that the InceptionV3 model achieves rapid convergence within 100 epochs, with the accuracy on the training and validation sets increasing from 48.97% and 74.40% to 99.89% and 94.50%, respectively. The loss value decreases simultaneously, and all indicators continue to improve without significant overfitting, indicating that the InceptionV3 model has robust learning ability and strong generalization performance.

[0085] In this embodiment, step S1 uses a drone to collect remote sensing images of a cornfield at an altitude of 100 meters, thereby obtaining an initial low-resolution image with a resolution of 32×32. The initial low-resolution image is then input into an improved generative adversarial network model (ERSCA-WGAN), resulting in the image shown below. Figure 6 As shown. Figure 6 Images (a)-(d) are initial low-resolution images (32×32 pixels) showing different grades of corn rust; images (A)-(D) are high-resolution reconstructed images (128×128 pixels) generated after ERSCA-WGAN enhancement showing different grades of corn rust. Figure 6 As can be seen, the visual effect of the high-resolution reconstructed image is significantly better than that of the initial low-resolution image. Furthermore, the high-resolution reconstructed image shows significant improvements in lesion texture restoration and boundary sharpness, and effectively preserves the irregular boundary morphology and internal texture patterns of rust lesions, which is crucial for subsequent disease classification and identification.

[0086] To further quantify the performance of the ERSCA-WGAN model in identifying maize diseases after reconstruction from the initial low-resolution image, the InceptionV3 model was used to classify the disease levels of the initial low-resolution image at a flight altitude of 100 meters and the high-resolution reconstructed image generated after ERSCA-WGAN enhancement. The results are shown in Table 3.

[0087] Table 3. Comparison of classification performance of the InceptionV3 model using images at different resolutions. As shown in Table 3, the classification accuracy of the high-resolution reconstructed image reached 93.13%, which is 3.27 percentage points higher than the 89.86% of the original low-resolution image. The F1 score and recall also improved simultaneously, indicating that the method has achieved comprehensive optimization in multiple performance indicators.

[0088] Figure 7 Figures (a) and (b) show the confusion matrices of the InceptionV3 model between the initial low-resolution images and the high-resolution reconstructed images generated after ERSCA-WGAN enhancement. The figures show that out of the 99 initial low-resolution images actually classified as healthy, 6 were incorrectly predicted as mildly healthy, and 93 were correctly predicted as healthy. This demonstrates that after ERSCA-WGAN super-resolution reconstruction, the overall error rate decreased from 10.14% to 6.87%, representing an error reduction of 32.2%, indicating that ERSCA-WGAN achieved a significant improvement in global classification performance.

[0089] Furthermore, regarding the effectiveness in distinguishing adjacent disease levels, the error rate for healthy to mild disease decreased from 6.00% to 5.31%, a reduction of 11.5%; the error rate for mild to moderate disease decreased from 6.65% to 3.83%, a reduction of 41.67%; and the error rate for moderate to severe disease decreased from 6.28% to 2.87%, a reduction of 54.3%. This indicates that ERSCA-WGAN has a significant effect in mitigating confusion between adjacent disease levels, especially in improving the ability to distinguish between mild and moderate, and moderate and severe diseases.

[0090] Figure 8 This paper demonstrates the mapping effect of a drone on the distribution of cornfield diseases at a flight altitude of 100 meters. (a) shows the mapping result from the high-resolution reconstructed image, and (b) shows the mapping result from the initial low-resolution image. It can be seen that the mapping from the high-resolution reconstructed image exhibits higher spatial resolution, presenting not only a more continuous and accurate distribution of disease patches, but also preserving clearer spatial information in field boundaries and local detail areas. This improvement in spatial continuity and detail is of great significance for determining the location of disease occurrence and analyzing its spread trend in precision agriculture.

[0091] This invention achieves a PSNR of 20.96 dB and a SSIM of 66.67% by improving the super-resolution reconstruction technology of the Generative Adversarial Network (ERSCA-WGAN) model. This improves the spatial resolution and texture details of UAV high-altitude remote sensing images, thereby enhancing the monitoring efficiency of maize diseases. Through the synergistic effect of residual dense blocks and attention mechanisms, ERSCA-WGAN maintains stable reconstruction performance even under realistic remote sensing conditions, significantly improving image interpretation capabilities and visual quality. In maize rust identification, the InceptionV3 model significantly improves accuracy, reducing the misclassification rate of mild and moderate diseases by 41.67%. Therefore, this invention is feasible in preserving disease details and improving classification accuracy, and can be widely applied in high-altitude, large-scale farmland disease monitoring.

[0092] Compared to traditional disease monitoring methods based on spectral features, the technology proposed in this invention exhibits stronger adaptability and higher detection accuracy under high-altitude remote sensing conditions. Specifically targeting crop disease identification under high-altitude flight conditions, this invention achieves a 93.13% identification accuracy rate in real-world scenarios through super-resolution reconstruction technology, demonstrating a significant performance improvement over existing high-altitude monitoring methods. This method effectively compensates for the impact of insufficient spatial resolution in high-altitude imaging on the detailed features of diseases, enabling deep learning models to maintain accurate disease identification capabilities over a larger operational range, thus providing technical support for large-scale crop disease monitoring.

[0093] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A corn disease recognition method based on an improved generative adversarial network, characterized in that, The method comprises the following steps: S1: collecting and acquiring remote sensing images of a corn field by using a UAV, preprocessing the remote sensing images, and obtaining an initial low-resolution image with a pixel size of N*N; S2: performing super-resolution reconstruction on the initial low-resolution image by using a pre-constructed improved generative adversarial network model to obtain a high-resolution reconstructed image with a pixel size of R*R; The construction method of the improved generative adversarial network model is as follows: S2.1: constructing a shallow feature extraction layer for extracting shallow features from the input initial low-resolution image; S2.2: constructing a deep feature extraction network based on multiple RRDB nested residual dense blocks to extract deep features; S2.3: constructing an attention module based on a spatial and channel dual attention mechanism for feature recalibration and key region positioning of the extracted deep features; S2.4: constructing a multi-scale texture enhancement module through multi-scale convolution and a smoothing branch for texture enhancement of the deep features output by the attention module, extracting multi-scale texture features, and the smoothing branch is used to suppress high-frequency noise; S2.5: constructing a global residual connection layer for global residual connection of the shallow features and the deep features output by the multi-scale texture enhancement module to obtain fused features; S2.6: constructing an adaptive hybrid upsampling module based on transposed convolution and stable upsampling, the adaptive hybrid upsampling module is provided with two layers for two-stage upsampling of the fused features; S2.7: constructing an output layer for mapping and outputting the upsampled features to obtain the high-resolution reconstructed image; S3: performing disease prediction on the high-resolution reconstructed image by using a trained image classification model to obtain a disease prediction result.

2. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, The preprocessing method in step S1 comprises: performing image stitching, geometric correction and coordinate conversion on the remote sensing images to obtain a standard orthographic image; performing grid segmentation on the standard orthographic image to crop the initial low-resolution image.

3. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, The deep feature extraction network in step S2.2 comprises multiple RRDB nested residual dense blocks connected in series, each RRDB nested residual dense block comprises multiple residual dense blocks, the multiple residual dense blocks adopt macro residual connection and form a multi-level residual structure through a residual scaling mechanism; The expression of each RRDB nested residual dense block is as follows: ; wherein i is the serial number of the RRDB nested residual dense block, β is a residual scaling factor, x represents an input image, RDB1, RDB2 and RDB3 represent a first residual dense block, a second residual dense block and a third residual dense block respectively.

4. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, The attention module in step S2.3 comprises a channel attention layer and a spatial attention layer connected in series, the channel attention layer is used for recalibrating features, and the spatial attention layer is used for positioning key regions; The expression of the channel attention layer is as follows: ; wherein CA is a channel attention function, σ is a Sigmoid activation function, MLP is a multi-layer perception, GAP and GMP are global average pooling and maximum pooling operations respectively; The expression of the spatial attention layer is as follows: ; where SA is a spatial attention function, Conv 7×7 denotes a 7x7 convolution operation, GAP c and GMP c are average pooling and max pooling along the channel dimension, respectively, [;] denotes a feature concatenation operation.

5. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, The multi-scale texture enhancement module in step S2.4 is used for feature fusion of the input features processed by the multi-scale convolution and smoothing branch, and residual connection with the input features, and an expression of the multi-scale texture enhancement module is as follows: ; ; wherein TE is a texture enhancement function, a is a residual connection weight coefficient, Fusion is a feature fusion operation, Conv k denotes a k x k convolution operation, k takes 1, 3, 5; Smooth is a smoothing branch function, and AvgPool is an average pooling operation.

6. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, An expression of the adaptive mixed up-sampling module in step S2.6 is as follows: ; where AdaUp is an adaptive mixed up-sampling function, ω 1 and ω 2 are dynamically generated adaptive weights, ConvT is a transposed convolution, and StableUp is a stable up-sampling. An expression of the adaptive weight dynamically generated is as follows: ; wherein, Softmax is an activation function, Conv is convolution, and GAP is global average pooling.

7. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, The image classification model in step S3 adopts a MobileNetV3 model or an InceptionV3 model.

8. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, The improved generative adversarial network model is trained by using a super-resolution reconstruction data set, and the image classification model is trained by using a disease classification data set. An acquisition method of the super-resolution reconstruction data set and the disease classification data set is as follows: A historical high-resolution image under a preset flight height is acquired, and correction and cropping preprocessing is performed to obtain the historical high-resolution image with a pixel size of R×R. Degradation processing is performed on the historical high-resolution image to generate a low-resolution degraded image with a pixel size of N×N, and the degradation processing is used to reduce the resolution of the historical high-resolution image. After first data enhancement is performed on the historical high-resolution image, the historical high-resolution image is used as a high-resolution ground truth image, and after first data enhancement is performed on the low-resolution degraded image corresponding to the historical high-resolution image, the low-resolution degraded image is used as a low-resolution input image, and the high-resolution ground truth image and the low-resolution input image are used as super-resolution reconstruction samples to construct a super-resolution reconstruction data set; the first data enhancement includes original image retention, horizontal flipping, vertical flipping and rotation transformation. Disease grade labeling is performed on the historical high-resolution image, and second data enhancement is performed to obtain a disease classification sample, and a disease classification data set is constructed by using the disease classification sample; The second data enhancement includes original image retention, brightness enhancement, contrast enhancement, rotation transformation, horizontal flipping and vertical flipping.

9. The method for corn disease recognition based on improved generative adversarial network according to claim 1, characterized in that, A total loss function used when the improved generative adversarial network model is trained includes an adversarial loss, a content loss, a perception loss and an edge loss; The adversarial loss is constructed based on a WGAN-GP framework, and includes a discriminator loss and a generator loss, and the discriminator loss includes a gradient penalty term; The content loss is calculated by using an L1 norm to calculate a pixel-level difference between a high-resolution reconstructed image and a high-resolution ground truth image; The perception loss is constructed based on a pre-trained VGG-19 network to extract multi-layer features; The edge loss is constructed by combining a Sobel operator and a Laplacian operator to extract multi-scale edge information.

10. The method of claim 9, wherein the improved generative adversarial network-based corn disease recognition method is characterized by, An expression of the discriminator loss is as follows: ; wherein, is a discriminator loss, is a desired function, D is a discriminator, x fake and x real are a generated sample and a real sample, respectively, is an interpolated sample obtained by a random interpolation of the real sample and the generated sample, is a gradient penalty coefficient, denotes a gradient of the discriminator with respect to the interpolated sample, denotes an L2 norm; The generator loss is: ; An expression of the perception loss is as follows: ; wherein, G is a generator, x lr and x hr are a low-resolution input image and a high-resolution target image, respectively, represents an i-th layer VGG feature extraction function, M is a total number of feature layers of a pre-trained VGG-19 network, is a weight coefficient of the i-th layer feature, represents an L1 norm; An expression of the edge loss is as follows: ; ; wherein, is an edge loss function, Sobel x and Sobel y are horizontal and vertical Sobel operators, respectively, Laplacian is a Laplacian operator, ε is a combination weight, and Edge is an edge detection function.

Citation Information

Patent Citations

  • Corn disease identification method based on texture-color multi-scale residual shrinkage network

    CN116630960A

Cited By

  • Unmanned aerial vehicle image-based corn elongation stage water stress identification and segmentation method

    CN121686297A

  • Astronomical image super-resolution method based on dense residual connection and attention mechanism

    CN121724839A