Small sample image expansion method based on improved CycleGAN

By using small-size convolution kernels and extrusion-excitation attention mechanism in the CycleGAN model, the generation and recognition effect of corn disease images is improved, and the problems of insufficient sample data and low recognition accuracy are solved. The generated images are more realistic and diverse, and the recognition accuracy is improved.

CN120219884APending Publication Date: 2025-06-27HENAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510374801.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient sample data and low recognition accuracy in corn leaf disease image recognition, especially in the recognition of corn leaf disease images.

Method used

Under the framework of the CycleGAN model, improving the network structure by replacing the original 7×7 convolution kernel with a 3×3 convolution kernel, increasing the network depth and improving the image feature extraction ability. At the same time, an extrusion-excitation attention mechanism (SE) is introduced into the residual module to enhance attention on the characteristics of micro lesions.

Benefits of technology

The disease images generated by the improved CycleGAN model are more realistic and diverse, with improved recognition accuracy and reduced FID values, indicating that the disease characteristic distribution of the generated image is closer to the real image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219884A_ABST
    Figure CN120219884A_ABST
Patent Text Reader

Abstract

The invention provides a small sample image expansion method based on an improved CycleGAN, and adopts a corn leaf disease image data enhancement method based on an improved cyclic consistency generative adversarial network CycleGAN to solve the problems of difficulty in data set acquisition, insufficient samples, imbalance of different types of disease samples and the like in a corn disease image recognition task. The method comprises the following steps: firstly, optimizing a CycleGAN network structure by using a convolution kernel with a relatively small receptive field; secondly, embedding an SE attention mechanism into a residual module of the generator; compared with the prior art, the FID value of the disease image generated by the improved CycleGAN is reduced, the recognition accuracy on AlexNet, VGGNet and ResNet recognition networks is improved by 3.9%, 4.41% and 3.44% respectively, the problem that a disease image data set is deficient is effectively solved, and reliable data support is provided for training of a disease recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sample image enhancement and expansion, and specifically relates to an improved method for expanding small-sample images by using an adversarial network. Background Art

[0002] For a long time, traditional methods have been used in China to identify crop diseases. Agricultural workers observe with the naked eye, judge diseases through experience, or collect samples for identification in the laboratory. These methods are time-consuming and laborious, and are restricted by multiple factors such as manpower, material resources, and financial resources, resulting in inability to be widely applied in agricultural production, and the problem of diseases still emerging in an endless stream. In particular, corn is rich in nutritional value and has high economic benefits, but it is vulnerable to environmental and pest threats during the growth process, resulting in poor yield and quality, causing serious economic losses. Therefore, early prevention and intervention of corn diseases to ensure the sustainable development of the corn industry is an urgent problem to be solved. At the same time, the country strongly advocates smart agriculture and uses artificial intelligence methods to solve agricultural problems. Currently, deep learning has developed rapidly and made significant contributions in the agricultural field. Using convolutional neural networks such as AlexNet, VGG, GoogLeNet, ResNet, etc. for disease identification has relatively significant effects. Many research scholars have applied the above networks to the agricultural field according to actual problems.

[0003] It is difficult to obtain crop disease images, and the data samples are few. However, deep learning models require a large amount of data support. Therefore, expanding the data has become a difficult problem. Affine transformation is a typical method in traditional data augmentation. It transforms the image by flipping, translating, randomly cropping the image, etc., which can effectively expand the data volume and increase the spatial complexity of the data. However, with the rise of the Generative Adversarial Network (GAN), new possibilities have been brought to data augmentation. Researchers have tried to use the GAN network to generate "realistic" images to achieve the purpose of expanding the images.

[0004] Although the recognition effect of crop diseases has been improved to a certain extent currently, there are still problems such as insufficient sample data and low recognition accuracy in solving the recognition of corn leaf disease images.

[0005] In view of the above, this study proposes a method for expanding small-sample images based on an improved CycleGAN within the framework of the CycleGAN model. This method aims to train and generate high-quality disease images from small-sample disease data, thereby reducing the acquisition cost of corn disease images to solve the above problems. Summary of the Invention

[0006] In view of the above situation, to overcome the defects of the existing technology, the present invention provides a small sample image augmentation method based on improved CycleGAN, and the specific idea is as follows: (1) All 3×3 convolutional kernels are used in the generator for operation, replacing the 7×7 convolutional kernels in the original CycleGAN network, which can increase the network depth and improve the image feature extraction ability, and generate leaf images with obvious disease characteristics; (2) The Squeeze-Excitation (SE) attention mechanism is inserted into the residual module to prevent the loss of position feature information caused by convolutional operations. On the premise of not losing the original feature information, the relationship between channels is fused, the attention to tiny lesion features is increased, and the generation effect of disease images is improved.

[0007] A small sample image augmentation method based on improved CycleGAN, characterized in that it includes: First, use convolutional kernels with a smaller receptive field to optimize the CycleGAN network structure, thereby generating high-quality sample images and reducing the occurrence of overfitting; Second, embed the Squeeze-Excitation, abbreviated as SE, that is, the squeeze-excitation attention mechanism into the residual module of the generator to enhance the ability of CycleGAN to extract disease features, so that the network can more accurately capture small target diseases or features with insignificant domain differences.

[0008] Further, the optimization using convolutional kernels with a smaller receptive field specifically includes: In the improved CycleGAN network, all 3×3 convolutional kernels are used in the generator for operation, and the 7×7 convolutional kernels in the original CycleGAN network are cancelled. The receptive field of the small-size convolutional kernel is smaller and more sensitive to the detailed information of the image, and it can generate small corn leaf lesions. Then, by stacking small-size convolutional kernels, the scanning range consistency with the large-size convolutional kernel is achieved.

[0009] Further, the stacking of small-size convolutional kernels specifically includes: In the same layer of the network, the number of parameters of the 7×7 convolutional kernel is as high as 7×7×N. In this scheme, 3 3×3 convolutional kernels are stacked for replacement, and the number of parameters only needs 3×3×3×N, which is only 55% of the number of parameters of the original 7×7 convolutional kernel. Here, N represents the number of output channels. In terms of network depth, using 3 small convolutional kernels to replace the large convolutional kernel can deepen the network, improve the feature extraction ability, and reduce the occurrence of overfitting.

[0010] Further, the process of embedding SE into the residual module of the generator includes taking the output of the residual module in the optimized CycleGAN network structure as the input for a Squeeze operation, compressing and flattening it into global features, assigning different weights to channel features according to their correlation through an Excitation mechanism, multiplying the excited features by the residual output to achieve the purpose of assigning attention, and finally adding the result to the upper-layer input of the residual module for output, thus forming a Squeeze-Excitation attention mechanism residual module, abbreviated as the SE-ResnetBlock module.

[0011] The beneficial effects of the above technical solutions are as follows:

[0012] (1) In this study, a series of improvements were made within the framework of the original CycleGAN network model. The convolutional kernels in the generator were changed to small-sized convolutional kernels, resulting in a smaller receptive field and greater sensitivity to the detailed information of the image, enabling the generation of detailed lesion features with smaller targets.

[0013] (2) In this study, by stacking small-sized convolutional kernels, the scanning range was made consistent with that of the original large-sized convolutional kernels. Moreover, using 3 small convolutional kernels instead of the original large convolutional kernel could deepen the network, improve the feature extraction ability, and reduce the occurrence of overfitting.

[0014] (3) In this study, a Squeeze-Excitation attention mechanism (SE) module was inserted into the original residual structure. This module helps prevent the loss of positional feature information caused by convolutional operations, increases the attention to minute lesion features, and improves the generation effect of disease images.

[0015] (4) In this study, compared with the original CycleGAN, DCGAN, and WGAN algorithms, the improved CycleGAN reduced the FID values of the generated disease images by 43.33, 32.67, and 19.72 respectively, indicating that the disease feature distribution of the generated images is closer to that of the real images, and the generated images are more diverse.

[0016] (5) When the image dataset obtained by the improved CycleGAN expansion method in this study was used for leaf disease recognition, the recognition accuracies on the AlexNet, VGGNet, and ResNet recognition networks were increased by 3.9%, 4.41%, and 3.44% respectively. Brief Description of the Drawings

[0017] Figure 1 It is a schematic diagram of maize leaf diseases;

[0018] Figure 2 It is a schematic diagram of the structure of the SE-ResnetBlock module of the present invention;

[0019] Figure 3 Schematic diagram of the improved CycleGAN network generator structure of the present invention;

[0020] Figure 4 Schematic diagram of the GAN evaluation index processing flow;

[0021] Figure 5 Schematic diagram of the comparison of the improved CycleGAN network of the present invention;

[0022] Figure 6 Schematic diagram of the statistical results of the recognition accuracy of AlexNet in the specific implementation manner;

[0023] Figure 7 Schematic diagram of the statistical results of the recognition accuracy of VGGNet in the specific implementation manner;

[0024] Figure 8 Schematic diagram of the statistical results of the recognition accuracy of ResNet in the specific implementation manner. Specific implementation manner

[0025] Regarding the foregoing and other technical contents, features and effects of the present invention, they will be clearly presented in the following detailed description of the embodiments with reference to the accompanying drawings of the present application. The contents mentioned in the following embodiments are all referenced to the accompanying drawings of the specification.

[0026] Example 1, collation of the test data set, as Figure 1 shown, the test data set of this study consists of 3 common maize leaf diseases (northern leaf blight of maize, common rust of maize, gray leaf spot of maize) and healthy leaf images. Typical data samples are shown in the figure. Among them, there are 510 healthy images, 554 northern leaf blight of maize images, 578 common rust of maize images, and 511 gray leaf spot of maize images. The images are from the authoritative official website Kaggle in the industry.

[0027] Example 2, introducing the cyclic consistency loss mechanism. The cyclic adversarial network (CycleGAN) is an asymmetric network for image translation. Because it does not require the use of paired data during the training process, it effectively reduces the requirements of the network for image data and enhances the scope of application of the network. The cyclic adversarial network consists of 2 generators and 2 discriminators, and is used to perform asymmetric conversion of the source domain and the target with images. Its loss function mainly consists of 3 parts: adversarial loss, cyclic consistency loss, and identity loss. Among them, the role of the adversarial loss is to force the generator to generate realistic images to deceive the discriminator. The adversarial loss function consists of two parts, and its structure is:

[0028]

[0029] In the formula: is the adversarial loss function, E is the mathematical expectation, S is the source domain, s is the image in the source domain, T is the target domain, t is the image in the target domain, G is the generator from the source domain to the target domain, F is the generator from the target domain to the source domain, D S is the source domain discriminator, D T is the target domain discriminator.

[0030] Since using the adversarial loss function alone cannot guarantee that the generator maps the input to the desired output, there is a problem of "mode collapse", resulting in the loss of a large amount of information of the source domain images in the images generated by the generator. Therefore, this study introduces the cycle consistency loss to constrain the cycle consistency of the mapping function between the target domain and the source domain. The cycle consistency loss mainly improves the ability of the generator to retain the original image by minimizing the difference between the target image and the cycled image, and can effectively improve the accuracy of image conversion. Its overall structure is:

[0031]

[0032] In the formula: is the cycle loss function, E is the mathematical expectation, s is the image in the source domain, t is the image in the target domain, G is the generator from the source domain to the target domain, and F is the generator from the target domain to the source domain.

[0033] In addition to the adversarial loss and the cycle consistency loss, the identity loss is also applied to CycleGAN to constrain the accuracy of the texture, color and other information of the input image and the cycled generated image during the image style transfer process. The identity loss function is:

[0034]

[0035] Example 3, using the improved CycleGAN network for data augmentation. This network learns the disease characteristics of the corn leaf disease dataset through the adversarial training of the generator and the discriminator, and generates new corn leaf disease images. The structure of this network is similar to the original CycleGAN network and consists of 2 recurrent networks. Taking the training process of healthy images and northern leaf blight images of corn as an example, this network maps the images in domain B (healthy images) to the images in domain A (northern leaf blight images of corn), and the images in domain A are restored to domain B through the generator, and so on in a cycle. With the auxiliary role of the discriminator, the network completes the conversion from domain B to domain A. Similarly, the process of converting domain A to domain B images is the same as above.

[0036] When the original CycleGAN network generates images of corn leaf diseases, there are problems such as the uncertain position of lesion generation and the poor effect of lesion generation. Therefore, an improved CycleGAN network is designed, and the following improvements are made to this network: (1) In the initial convolutional blocks of the encoding layer and the decoding layer of the basic CycleGAN network, the convolutional kernel with a size of 7×7 pixels is optimized to a convolutional kernel with a size of 3×3 pixels. In this way, the increase in network depth and the extraction of obvious disease features can be achieved, making the obtained disease images more realistic; (2) In the conversion layer of the basic CycleGAN network, the squeeze-excitation attention mechanism (i.e., Squeeze-Excitation, abbreviated as SE) is embedded into the original residual block to form the SE-ResnetBlock module. This module can make the algorithm focus more on the position of leaf lesions and improve the disease authenticity of the images generated by the generator.

[0037] Example 4, as Figure 3 shown, the structure of the generator network is optimized. In this study, the CycleGAN network uses only 3×3 convolutional kernels for operations in the generator, and the 7×7 convolutional kernels in the original CycleGAN network are cancelled. Figure 3 It is the structure diagram of the generator of the improved CycleGAN network. The receptive field of the small-size convolutional kernel is small and it is more sensitive to the detailed information of the image. It can generate small target corn leaf lesions, and then by stacking small-size convolutional kernels, the consistency with the scanning range of the large-size convolutional kernel can be achieved. In terms of the number of parameters, for the same layer of the network, the number of parameters of the 7×7 convolutional kernel is as high as 7×7×N, and the number of parameters of three 3×3 convolutional kernels is only 3×3×3×N, which is only 55% of the number of parameters of the 7×7 convolutional kernel, where N represents the number of output channels. In terms of network depth, using three small convolutional kernels instead of a large convolutional kernel can deepen the network, improve the feature extraction ability, and reduce the occurrence of overfitting.

[0038] Example 5, as Figure 2 shown, the attention mechanism is embedded. The generator of the original CycleGAN network is similar to an encoder-decoder structure. When the B (healthy) and A (a certain disease) domain images are input, after passing through the encoder structure, they enter the conversion layer, and then the converted features are decoded to output the generated A domain image. In the conversion layer structure, the original CycleGAN network is realized by stacking nine residual structure blocks, but the ability to extract the position features of lesions is limited, resulting in the inability to generate lesions on the leaves. Therefore, in this study, the squeeze-excitation attention mechanism (Squeeze-Excitation, SE) is inserted into the residual module to prevent the loss of position feature information caused by convolutional operations. The structure of the SE-ResnetBlock module is as Figure 2As shown in the figure, the dashed box part in the figure is the SE module. Among them, the SE module takes the output of the residual module as the input for the squeezing operation, compresses and flattens it into global features, assigns different weights to the channel feature information according to the correlation through the excitation mechanism, multiplies the excited features by the residual output to achieve the purpose of assigning attention, and finally adds the result to the upper-layer input of the residual module for output, forming the SE-ResnetBlock module, that is, the squeeze-excitation attention mechanism residual module. The improved residual attention module combines the ideas of shortcut connection and attention mechanism. Without losing the original feature information, it integrates the relationship between channels, increases the attention to the features of small lesions, improves the generation effect of disease images. After introducing this module, the generator will focus more on the leaf lesion area and learn more valuable disease features.

[0039] Example 6, conduct experiments and method evaluation:

[0040] (1) The test environment is configured as follows:

[0041] The operating system is 64-bit Windows 10, using the Python language with version 3.7.0. The deep learning framework is Pytorch. The network is built, and the dataset is trained and tested in this environment. The central processing unit is i5-12400F, the graphics processing unit is RTX3060Ti, and the video memory capacity is 8G.

[0042] (2) Test settings:

[0043] The improved CycleGAN network is a one-to-one mutual generation adversarial network. Therefore, the experiment is divided into 3 groups, namely domain B1 (healthy leaves) and domain A1 (northern corn leaf blight), domain B2 (healthy leaves) and domain A2 (common rust of maize), domain B3 (healthy leaves) and domain A3 (gray leaf spot of maize). Among them, the ratio of the training set to the validation set is set to 10:1. The optimization algorithm uses Adam. The learning rates of the generator and the discriminator are set to 0.0002, the batch size is 32, the generated image size is 256 pixels × 256 pixels, and after 300 epochs of iterative training, the training is terminated. The comparative experiment is to compare the recognition accuracy on 3 recognition networks, namely AlexNet, VGGNet and ResNet, and the number of iterations is 100 times. The settings of each group of experiments are exactly the same.

[0044] (3) Evaluation indicators:

[0045] When evaluating the images generated by GAN series models, the following three aspects are mainly concerned: authenticity, diversity, and structural consistency. Authenticity refers to whether the images generated by the GAN model are as realistic as the images that exist in reality. Diversity means that the generated images are not of only one form, but rather images with various shapes. As for structural consistency, it concerns to what extent the generated images maintain the stability of their local and overall structures. These factors together constitute the evaluation of the quality of the images generated by the GAN series models.

[0046] Therefore, this study selects the GAN-train index, the GAN-test index, and the FID index as the indicators to measure the performance of the images generated by GAN. GAN-train uses the images generated by GAN to train a classifier and then tests it on real images, corresponding to the recall rate (diversity) of GAN. While GAN-test uses real images as the training samples to train a classifier and then tests it on the generated images, corresponding to the precision rate (image quality) of GAN, so as to realize the evaluation of the quality of the images generated by the GAN model. Through these two indicators, it can be understood whether the images generated by GAN are diverse and realistic enough, providing strong support for the improvement and optimization of GAN so that it can be better applied to actual scenarios. The principle of GAN evaluation is as Figure 4 shown.

[0047] Using FID (feature inception distance) to evaluate images is to extract feature vectors using the inception network, compare the differences in the distribution of the generated images and real images in the feature space, and obtain the FID evaluation score. The lower the FID evaluation index score, the smaller the difference between the obtained images and real images, and the closer the authenticity and diversity are to reality. The FID calculation formula is:

[0048]

[0049] In the formula: x is the real sample data distribution, g is the generated sample data distribution, μ and C are the mean and covariance of the feature vectors after feature extraction, and tr is the trace of the matrix.

[0050] (4) Display of generated images:

[0051] The generated effect diagrams of the improved CycleGAN network and the original CycleGAN network are as Figure 5As shown. The first column in the figure shows the lesion images of 3 different diseases generated by the original CycleGAN. The second column (labeled CycleGAN-1) shows the images generated by improving the small convolutional kernel and increasing the network depth on the basis of the original CycleGAN. The third column shows the generated images obtained by adding the SE-ResnetBlock module to the second column. The fourth column shows the real images of each disease.

[0052] Observation Figure 5 , it can be seen that the images generated by the original CycleGAN network are insufficient in terms of authenticity and clarity, the disease characteristics are not obvious enough, and sometimes the images cannot be generated completely; when the network is improved to CycleGAN-1, although the generated images are improved compared with those of the original CycleGAN, the images are still somewhat blurred, and the quality of the generated lesions is poor. In contrast, the images generated by the improved CycleGAN network have been significantly improved in terms of authenticity and layering, and are closer to the quality of real images.

[0053] (5) Comparative experiment:

[0054] In order to scientifically evaluate the performance of the improved CycleGAN network in generating maize disease leaf images, taking the CycleGAN network as a benchmark, DCGAN, WGAN, and the improved CycleGAN network are added for comparison. In addition, to ensure the comprehensiveness and accuracy of the experiment, first, this study expands these 3 kinds of disease images and integrates them into the original dataset to jointly form the training set. Second, experiments and tests are carried out on 3 different pre-trained recognition networks; finally, the disease recognition accuracies of the disease images generated by them are tested on the AlexNet, VGGNet, and ResNet recognition networks for intuitive comparison. The specific experimental results are shown in Table 1. It can be seen from Table 1 that the improved algorithm in this study can achieve higher recognition accuracy compared with the other 3 algorithms, indicating that this algorithm can generate more real and diverse disease images, thus proving the effectiveness of the algorithm in this study. In addition, 500 generated images are selected from various disease images to calculate their GAN-train, GAN-test, and FID scores. The comparison of evaluation indicators is shown in Figure 6 、 7 and Figure 8.

[0055] Table 1 Comparison table of the accuracies of 3 recognition networks

[0056]

[0057] Table 2 Comparison of 4 models under 3 evaluation indicators

[0058]

[0059] Figure 6 、 7 And 8 represents the comparison chart of the accuracy rates of the images generated by 4 model algorithms after 100 epochs of iteration in 3 pre-trained recognition networks. It can be seen from the figure that the highest recognition accuracy of the improved CycleGAN network model on AlexNet is 95.20%, which is 3.9%, 2.78%, and 0.62% higher than those of the CycleGAN, DCGAN, and WGAN models respectively. The highest recognition rate on VGGNet is 96.57%, which is 4.41%, 3.55%, and 0.77% higher than those of the other 3 models. The highest recognition rate on ResNet is 97.15%, which is 3.44%, 3.21%, and 0.48% higher than those of the other 3 models. The accuracy rate in all recognition networks is higher than that of the other 3 models. Moreover, it can be seen from the 3 figures that the improved CycleGAN model converges the fastest and its trend is easy to stabilize. It is proved that the model in this study can generate more real and diverse corn leaf disease images compared with the commonly used CycleGAN, DCGAN, and WGAN models.

[0060] From Table 1, Table 2 and Figure 6 、 Figure 7 、 Figure 8 it can be seen that the improved CycleGAN network model is superior to the other 3 models in terms of the recognition network accuracy rate and evaluation indicators. The Gan-train and Gan-test indicators are respectively increased to 81.25% and 68.75%, indicating that the generated images are more real and diverse. The FID scores are respectively decreased by 43.33, 32.67, and 19.72 compared with the other 3 models, indicating that the disease feature distribution of the images generated by using the improved CycleGAN network model is closer to that of the real images, and the generated images are more diverse.

[0061] (6) Ablation experiment:

[0062] To verify the effectiveness of the improved CycleGAN network model, different optimization strategies were applied on the basis of CycleGAN for training to generate the corresponding augmented dataset of maize disease leaf images, and GAN-train, GAN-test, and FID were used as evaluation metrics. The test results are shown in Table 3. It can be seen from the data in Table 3 that after improving the convolutional kernel and increasing the network depth, all metrics have been improved. After further inserting the SE-ResnetBlock module into the generator, these metrics are improved again. The reason is that the SE-ResnetBlock module enhances the attention mechanism for the features of tiny lesions, thus significantly improving the authenticity of the images. The improved CycleGAN network model reached 81.25% and 68.75% on the GAN-train and GAN-test metrics respectively, while the FID metric decreased to 64.96. These three metrics indicate that the images generated by the improved CycleGAN network model are more realistic, the leaf lesion information is richer, and it can better meet the needs of practical applications.

[0063] Table 3 Comparison results of the ablation experiments of the models in this study

[0064]

[0065] In summary, to solve the problems such as difficult acquisition of the dataset for the maize disease image recognition task, insufficient samples, and imbalance of disease samples in different categories. A series of improvements were made in this study under the framework of the original CycleGAN network model. First, the convolutional kernel in the generator was changed from 7×7 to 3×3 for operation. The receptive field of the small-size convolutional kernel is smaller and more sensitive to the detailed information of the image, and it can generate maize leaf lesions with smaller targets. Then, by stacking small-size convolutional kernels, the scanning range consistency with the large-size convolutional kernel can be achieved. And using three small convolutional kernels instead of a large convolutional kernel can deepen the network, improve the feature extraction ability, and reduce the occurrence of overfitting. Secondly, the squeeze-and-excitation attention mechanism (SE) module was inserted into the original residual structure. This module helps to prevent the loss of position feature information caused by convolutional operations, increases the attention to the features of tiny lesions, and improves the generation effect of disease images.

[0066] In the comparative experiment, the images generated by the improved CycleGAN model had the highest recognition accuracy in the three disease networks. The recognition accuracies of GAN-train and GAN-test increased to 81.25% and 68.75% respectively, while the FID decreased to 64.96, all of which were better than the control model. Therefore, the improved CycleGAN model alleviated the problem of insufficient data in the neural network training process to a certain extent, providing strong support for the enhancement and recognition of maize leaf disease images. The generated disease images provided more real and diverse training samples for the disease detection model, significantly improving the recognition accuracy of the recognition network and the evaluation index of the generated disease images, indicating that the improved CycleGAN model could effectively expand the small sample dataset of maize disease leaves.

[0067] The above description is only for the purpose of illustrating the present invention. It should be understood that the present invention is not limited to the above embodiments, and various equivalent forms conforming to the idea of the present invention are within the protection scope of the present invention.

Claims

1. A small sample image expansion method based on improved CycleGAN, characterized in that: include: Firstly, the convolution kernel with smaller receptive field is used to optimize the CycleGAN network structure to generate high-quality sample images and reduce the overfitting phenomenon. Secondly, the Squeeze-Excitation (SE) attention mechanism is embedded into the residual module of the generator to enhance CycleGAN's ability to extract disease features, so that the network can more accurately capture the features of small target diseases or those with unclear differences between domains.

2. The small sample image expansion method based on improved CycleGAN according to claim 1, characterized in that: The optimization of the convolution kernel using a smaller receptive field specifically includes: the improved CycleGAN network uses 3×3 convolution kernels for all operations in the generator, cancels the 7×7 convolution kernel in the original CycleGAN network, and the small-sized convolution kernel has a smaller receptive field and is more sensitive to the details of the image. It can generate corn leaf spots with smaller targets, and then achieve consistency with the scanning range of the large-sized convolution kernel by stacking small-sized convolution kernels.

3. The small sample image expansion method based on improved CycleGAN according to claim 2, characterized in that: The stacked small-size convolution kernels specifically include: in the same layer of the network, the number of parameters of the 7×7 convolution kernel is as high as 7×7×N. This solution uses three 3×3 convolution kernels stacked to replace it, and the number of parameters is only 3×3×3×N, which is only 55% of the original 7×7 convolution kernel parameter number, where N represents the number of output channels. In terms of network depth, using three small convolution kernels instead of large convolution kernels can deepen the network, improve feature extraction capabilities, and reduce the occurrence of overfitting.

4. The small sample image expansion method based on improved CycleGAN according to claim 1, characterized in that: The process of embedding SE into the residual module of the generator includes taking the output of the residual module in the optimized CycleGAN network structure as input for squeezing operation, compressing and flattening into global features, assigning different weights to channel features according to relevance through an excitation mechanism, multiplying the excitation features with the residual output to achieve the purpose of giving attention, and finally adding the output with the upper layer input of the residual module to form a squeeze-excitation attention mechanism residual module referred to as SE-ResnetBlock module.