Progressive generative adversarial network-based crop disease and insect pest classification method
Through the crop pest classification method based on a progressive generative adversarial network, the problem of traditional methods is solved, and the high accuracy of pest classification is achieved, and the model training efficiency and feature extraction ability are improved.
Patent Information
- Application Number
- CN202510082992.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional crop pest classification methods rely on manual labor, are time-consuming and labor-intensive, and have low accuracy, making it difficult to meet the needs of modern agriculture for efficient and automated testing.
A crop pest classification method based on a progressive generative adversarial network is adopted to achieve high-accurate pest classification through the combination of data preprocessing, feature extraction and classifiers. Specific steps include image data preprocessing, feature extraction and fading operations, combining convolutional attention model and balanced convolution blocks to improve the accuracy of feature extraction and classification.
This method significantly improves the accuracy of crop pest classification, reduces the computational amount of model training, enhances the network's sensitivity to local key information, and improves the ability to extract deep features.
Smart Images

Figure CN119992196A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for classifying crop diseases and insect pests with high accuracy. Background Art
[0002] my country is a large agricultural country, and agriculture occupies an important position in the national economy. However, relevant studies have shown that the losses in agricultural production caused by pests and diseases are countless every year. Crop pests and diseases not only have a serious impact on the growth and development of crops, but also lead to problems such as reduced quality and yield of crops and damage to the agricultural ecological environment. Therefore, it is very important to study the identification of crop pests and diseases, which can enable researchers to prevent and treat crop pests and diseases in the early stage, thereby ensuring agricultural production, which is of great significance and value.
[0003] Traditional classification of crop pests and diseases mainly relies on manual labor, and agricultural researchers need to go to the site in person to collect samples and identify them, which consumes a lot of manpower costs. In recent years, image recognition technology based on deep learning has been widely used in many fields and has achieved remarkable results. In view of the limitations of traditional methods, researchers have explored the application of image recognition technology based on deep learning to the classification and identification of crop pests and diseases, so that crop pests and diseases can be detected efficiently and automatically, overcoming the time-consuming and labor-intensive shortcomings of manual detection. Summary of the invention
[0004] In view of the above problems, the present invention provides a crop pest classification method based on a progressive generative adversarial network, which includes the following process:
[0005] S1, obtaining crop pest and disease image data;
[0006] S2, preprocessing the image data;
[0007] S3, performing image feature processing on the preprocessed image; inputting the image into the left and right linear channels to obtain a feature image, while assigning reasonable weight values to the feature images obtained by the left and right channels, and connecting the two feature images based on a fade-in operation;
[0008] S4, classifying pests and diseases based on the trained crop pest and disease classification model;
[0009] The model includes a feature enhancement module, which inputs the feature image after merging the left and right channels into the balanced convolution block and expands the number of channels of the feature image to twice the original to enrich the feature diversity, so as to facilitate the subsequent feature extraction;
[0010] It includes a feature extraction module, which extracts image features by building a chain convolution structure based on the convolution attention model, balanced convolution blocks, and maximum pooling;
[0011] It includes a target detection module, which sends the feature image obtained by the feature extraction module to the classifier for further processing and classification, and finally outputs the pest and disease category of each image.
[0012] Preferably, the data preprocessing in S2 includes smoothing and sharpening, and the image is uniformly reduced to a size of 256×256 using the Resize method in the Transformer class.
[0013] Preferably, the smoothing and sharpening processing is implemented in the following steps:
[0014] Use a Gaussian filter based on normal distribution to smooth the image and remove noise and interference information from the image;
[0015]
[0016] Among them, I smooth (x, y) represents the pixel value of the image after smoothing, I(x, y) represents the pixel value of the original image, G(i, j) represents the weight of the Gaussian filter, σ is the standard deviation, which controls the width of the Gaussian distribution, and k is the radius of the filter window, which determines the size of the filter;
[0017] Binarize the image and use the Laplacian operator to sharpen the image to highlight the edge features of the image;
[0018] Find the texture, edge and disease location of crop leaves, then add the processed image to the original image through digital image addition to complete the sharpening operation of image features:
[0019]
[0020] Among them, I binary (x, y) represents the binary image, I sharpened (x, y) represents the sharpened image, and L(i, j) represents the weight of the Laplacian operator.
[0021] Preferably, the linear left channel implementation steps in S3 are as follows:
[0022] The image is processed using a 1×1 balanced convolution layer, where the balanced convolution layer is a convolution layer initialized with a balanced learning rate, and the formula is:
[0023]
[0024] Where C represents the number of channels of the input image, C′ represents the number of channels of the output image, W represents the weight of the convolution kernel in the balanced convolution layer, and I c (x, y) represents the value of the input image at position (x, y), Ic′ (x, y) represents the value of the output image at position (x, y) after balanced convolution processing.
[0025] Preferably, the linear right channel implementation steps in S3 are as follows:
[0026] Downsample the image using max pooling:
[0027]
[0028] Among them, m and n represent the height and width of the pooling window, s is the step size of the pooling, and I c (x, y) represents the value of the input image at position (x, y) and channel c, O c (x, y) represents the value of the output image at position (x, y) and channel c;
[0029] The obtained image is mapped through a 1×1 convolution layer to obtain the feature image of the right channel.
[0030] Preferably, the steps for implementing the fade-in operation in S3 are as follows:
[0031] Assign weights to the feature images obtained from the left and right channels, and then perform a dot-add operation on the two feature images:
[0032] F=α×I left +(1-α)×Downsample(I right ) (7)
[0033] Among them, α is the weight assigned to the current left channel feature image, I left represents the feature image of the left channel, Downsample represents the downsampling operation, (1-α) represents the weight assigned to the current feature image of the right channel, and I right Represents the feature image of the right channel.
[0034] Preferably, the feature enhancement module implements the following steps:
[0035] After the feature image is input into the balanced convolution block, it first undergoes a 1×1 balanced convolution to increase the dimension of the feature image in the channel.
[0036] Then the LeakyRelu activation function is used to perform nonlinear transformation on the feature image;
[0037] Normalize the feature vector of the feature image at the pixel level, and divide the pixel value by the standard deviation of the color channel where the pixel value is located, so that the pixel value is within the normal distribution with a mean of 0 and a standard deviation of 1. The formula is as follows:
[0038]
[0039] Among them, μ c represents the mean of channel c, σ c represents the standard deviation of channel c, I c (x, y) represents the value of the input image at position (x, y) and channel c, I norm,c (x, y) represents the value of the output image at position (x, y) and channel c;
[0040] Finally, the feature image is input into a 3×3 balanced convolution layer. Local features are extracted in different areas of the feature image through convolution operations. In addition, a built-in edge area is added to the edge of the feature image before the convolution operation. After that, the LeakyRelu activation function and pixel-level feature vector normalization are performed, and finally the feature image is output.
[0041] Preferably, the convolutional attention module includes a channel attention part and a spatial attention part; the channel attention part is implemented in the following steps:
[0042] The feature image is input into the channel attention module and weighted according to the importance of the feature map of each channel of the feature image;
[0043] The channel attention module uses two fully connected layers to compress and expand the features of each channel to obtain the importance score of each channel;
[0044] Use the Sigmoid function to normalize the importance score to [0, 1] to obtain the attention weight of each channel;
[0045] Apply the attention weights to the input feature map to obtain weighted features;
[0046] The steps for implementing the spatial attention part are as follows:
[0047] The feature image obtained by the channel attention module is input into the spatial attention module to weight the features at different spatial positions;
[0048] Spatial features are weighted through a bilinear pooling operation and two fully connected layers, where the bilinear pooling captures the spatial relationship in the feature map, and the fully connected layer compresses and expands the features of each spatial position to obtain the importance score of each position;
[0049] The importance score of each position is normalized and applied to the input feature map to obtain a weighted feature map.
[0050] Preferably, the implementation steps of the classifier are as follows:
[0051] Perform small batch standard deviation processing on the feature image output by the feature extraction network, calculate the standard deviation of each channel in the current batch of samples, and obtain a four-dimensional tensor;
[0052] The tensor is averaged along the channel dimension to obtain a four-dimensional tensor with one channel, and then the obtained tensor is concatenated to the feature image to obtain a new feature image with one channel added.
[0053] Pass the new feature image through a 3×3 convolution layer and a 1×1 convolution layer, and reduce the number of channels of the new feature image by 1;
[0054] The feature image is converted into a feature vector and input into the three-layer fully connected layer. The fully connected layer then performs a nonlinear transformation on the feature vector according to the weights and finally outputs the probability distribution of the category to obtain the classification result.
[0055] The beneficial effects of the present invention are as follows:
[0056] 1. Reduce the amount of computation during model training: Use the Resize method of the Transformer class to reduce the size of the image and the total number of features, thereby reducing data complexity for subsequent model training.
[0057] 2. Improve the accuracy of network recognition: Use a Gaussian filter based on normal distribution to smooth the image and remove noise and interference information; use the Laplace operator to sharpen the image to improve the image quality, thereby improving the accuracy of network recognition.
[0058] 3. More stable model training: Use balanced convolution layers to process images, so that the dynamic range of the convolution kernel value can be constrained when updating it in the back-propagation phase, avoiding instability in network training caused by large fluctuations.
[0059] 4. Improve sensitivity to local key information: Based on the progressive generative adversarial network, the image is gradually rendered from low resolution to high resolution, and then the layers of adjacent resolutions are connected using a fade-in method, so that the underlying network pays more attention to local key features.
[0060] 5. It can improve the extraction of deep features: The feature extraction network is built with the convolutional block attention module (CBAM) and balanced convolution blocks as the main structure, and multiple CBAM and balanced convolution blocks are stacked to expand the network receptive field and enhance the ability to extract deep features. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0062] Figure 1 The figure is a flow chart of the overall process of the crop pest classification method of the present invention.
[0063] Figure 2 This is a diagram showing the classification results of crop diseases and insect pests of the present invention;
[0064] Figure 3 This is a flow chart of image feature processing of the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] Example 1
[0067] In order to improve the accuracy of crop disease and pest classification, the present invention proposes a crop disease and pest classification method based on a progressive generative adversarial network. This method uses affine transformation, smoothing and sharpening for data preprocessing, wherein affine transformation is used to expand the data set and improve the robustness and generalization ability of the model, and smoothing and sharpening use a normally distributed Gaussian filter to smooth the image, remove noise and interference in the image, and improve the recognition accuracy; then, the image input to the network is adaptively processed, mainly based on the balanced convolution layer to reduce the number of image channels, thereby constraining its dynamic range when backpropagating to update the value of the convolution kernel, and avoiding instability during network training due to large fluctuations; since the image in the progressive generative adversarial network is gradually generated and rendered from low resolution to high resolution, in order to connect the training processes of different resolutions, the layers of adjacent resolutions are connected using a fade-in method. The input image is made more adaptable to the network, and the underlying network pays more attention to local key features; then, the feature image is passed through balanced convolution blocks of different sizes and normalized at the pixel level, so as to reduce the variance of the data and accelerate the inference speed of the model; then, the convolutional attention model is used to extract features from the image, so as to enhance the network's ability to focus on key features of the image and make the network pay more attention to features that are useful for the task of classifying crop pests and diseases; finally, the extracted features are classified using a classifier. Before using the fully connected layer for classification, the feature image is first processed with a small batch standard deviation to enhance the richness of the feature representation, and then the value of the feature image is constrained by convolution kernels of different sizes to prevent large difference values due to category differences. The overall process is as follows: Figure 1 shown.
[0068] In this implementation, we first collected crop pest and disease image datasets for model training, mainly based on the PlantVillage crop pest and disease dataset, which includes 38 disease categories for 14 crops such as apples, grapes, wheat, and potatoes;
[0069] Among them, the steps for data enhancement using affine transformation for the PlantVillage dataset are as follows:
[0070] First, use the Resize method in the Transformer class to reduce the image to a uniform size of 256×256;
[0071] Then, the data is symmetrically transformed horizontally and vertically using horizontal and vertical affine transformations, and each data image is randomly rotated 45° with a probability of 50% to expand the data set.
[0072] The steps for data enhancement using smoothing and sharpening for PlantVillage are as follows:
[0073] (1) Use a Gaussian filter based on normal distribution to smooth the image and remove noise and interference information from the image;
[0074]
[0075] in, smooth (x, y) represents the pixel value of the image after smoothing, I(x, y) represents the pixel value of the original image, G(i, j) represents the weight of the Gaussian filter, σ is the standard deviation, which controls the width of the Gaussian distribution, and k is the radius of the filter window, which determines the size of the filter.
[0076] (2) Binarize the image and use the Laplacian operator to sharpen the image to highlight the edge features of the image;
[0077] (3) Find the texture, edge and disease location of crop leaves, and then add the processed image to the original image through digital image addition to complete the sharpening operation of image features.
[0078]
[0079] Among them, I binary (x, y) represents the binary image, I sharpened (x, y) represents the sharpened image, and L(i, j) represents the weight of the Laplacian operator.
[0080] The image features are processed on the preprocessed image. The process is as follows: Figure 3 As shown in the figure: the image is input into the linear left channel and the linear right channel for processing. In the left channel, the image is processed using a 1×1 balanced convolution layer. Features of different dimensions are extracted by controlling the number of channels. In the right channel, the image is downsampled using the maximum pooling, and the downsampled image is mapped identically through a 1×1 convolution layer. Finally, the images extracted from the left and right channels are spliced as data for subsequent training models.
[0081] The steps to implement the linear left channel are as follows:
[0082] (1) The image is processed using a 1×1 balanced convolutional layer, where the balanced convolutional layer is a convolutional layer initialized with a balanced learning rate.
[0083]
[0084] Where C represents the number of channels of the input image, C′ represents the number of channels of the output image, W represents the weight of the convolution kernel in the balanced convolution layer, and I c (x, y) represents the value of the input image at position (x, y), I c′(x, y) represents the value of the output image at position (x, y) after balanced convolution processing.
[0085] Among them, the linear right channel implementation steps are as follows:
[0086] (1) Use maximum pooling to downsample the image;
[0087]
[0088] Among them, m and n represent the height and width of the pooling window, s is the step size of the pooling, and I c (x, y) represents the value of the input image at position (x, y) and channel c, O c (x, y) represents the value of the output image at position (x, y) and channel c.
[0089] (2) The obtained image is mapped through a 1×1 convolutional layer to obtain the feature image of the right channel.
[0090] The steps for implementing the fade-in operation are as follows:
[0091] (1) Assign weights to the feature images obtained from the left and right channels, and then perform a point addition operation on the two feature images.
[0092] F=α×I left +(1-α)×Downsample(I right ) (7)
[0093] Among them, α is the weight assigned to the current left channel feature image, I left represents the feature image of the left channel, Downsample represents the downsampling operation, (1-α) represents the weight assigned to the current feature image of the right channel, and I right Represents the feature image of the right channel.
[0094] After the above processing, a crop disease and insect pest classification model is constructed.
[0095] The model includes a feature enhancement module, which inputs the feature image after merging the left and right channels into the balanced convolution block and expands the number of channels of the feature image to twice the original to enrich the feature diversity, so as to facilitate the subsequent feature extraction;
[0096] It includes a feature extraction module, which builds a chain convolution structure based on the convolution attention model, balanced convolution blocks, and maximum pooling to extract image features.
[0097] It includes a target detection module, which sends the feature image obtained by the feature extraction module to the classifier for further processing and classification, and finally outputs the pest and disease category of each image.
[0098] Among them, the feature enhancement module uses the balanced convolution block to expand the number of feature image channels as follows:
[0099] (1) After the feature image is input into the balanced convolution block, it first undergoes a 1×1 balanced convolution to increase the dimension of the feature image in terms of channels;
[0100] (2) Use the LeakyRelu activation function to perform nonlinear transformation on the feature image;
[0101] (3) normalizing the feature vector at the pixel level of the feature image by dividing the pixel value by the standard deviation of the color channel where the pixel value is located, so that the pixel value is within a normal distribution with a mean of 0 and a standard deviation of 1;
[0102]
[0103] Among them, μ c represents the mean of channel c, σ c represents the standard deviation of channel c, I c (x, y) represents the value of the input image at position (x, y) and channel c, I norm,c (x, y) represents the value of the output image at position (x, y) and channel c.
[0104] (4) The feature image is input into a 3×3 balanced convolution layer. Local features are extracted from different regions of the feature image through convolution operations. In addition, a built-in edge region is added to the edge of the feature image before the convolution operation. After that, the LeakyRelu activation function and pixel-level feature vector normalization are performed, and finally the feature image is output.
[0105] Among them, the convolutional attention module includes a channel attention part and a spatial attention part;
[0106] The steps to implement channel attention are as follows:
[0107] (1) Input the feature image into the channel attention module and weight it according to the importance of the feature map of each channel of the feature image;
[0108] (2) The channel attention module uses two fully connected layers to compress and expand the features of each channel to obtain the importance score of each channel;
[0109] (3) Use the Sigmoid function to normalize the importance score to [0, 1] to obtain the attention weight of each channel;
[0110] (4) Apply the attention weights to the input feature map to obtain the weighted features.
[0111] Among them, the spatial attention implementation steps of the convolutional attention module are as follows:
[0112] (1) Input the feature image obtained by the channel attention module into the spatial attention module to weight the features at different spatial positions;
[0113] (2) Spatial features are weighted through a bilinear pooling operation and two fully connected layers. The bilinear pooling captures the spatial relationship in the feature map, and the fully connected layer compresses and expands the features of each spatial position to obtain the importance score of each position.
[0114] (3) Normalize the importance score of each position and apply it to the input feature map to obtain the weighted feature map.
[0115] Among them, the target detection module implementation steps are as follows:
[0116] (1) Perform small batch standard deviation processing on the feature image output by the feature extraction network, calculate the standard deviation of each channel in the current batch of samples, and obtain a four-dimensional tensor;
[0117] (2) The tensor is averaged along the channel dimension to obtain a four-dimensional tensor with a channel number of one, and then the obtained tensor is concatenated onto the feature image to obtain a new feature image with a channel number of one added;
[0118] (3) Pass the new feature image through a 3×3 convolution layer and a 1×1 convolution layer, and reduce the number of channels of the new feature image by 1;
[0119] (4) The feature image is converted into a feature vector and input into the three-layer fully connected layer. The fully connected layer then performs a nonlinear transformation on the feature vector according to the weights and finally outputs the probability distribution of the category to obtain the classification result. Figure 2 shown.
[0120] In the description of the present invention, it should be understood that the terms "longitudinal", "lateral", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a mechanical connection or an electrical connection, or it can be the internal communication of two elements, it can be a direct connection, or it can be an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0121] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A crop pest classification method based on progressive generative adversarial networks, characterized in that: The Method The process includes: S1, obtaining crop pest and disease image data; S2, preprocessing the image data; S3, performing image feature processing on the preprocessed image; inputting the image into the left and right linear channels to obtain a feature image, while assigning reasonable weight values to the feature images obtained by the left and right channels, and connecting the two feature images based on a fade-in operation; S4, classifying pests and diseases based on the trained crop pest and disease classification model; The model includes a feature enhancement module, which inputs the feature image after merging the left and right channels into the balanced convolution block and expands the number of channels of the feature image to twice the original to enrich the feature diversity, so as to facilitate the subsequent feature extraction; It includes a feature extraction module, which extracts image features by building a chain convolution structure based on the convolution attention model, balanced convolution blocks, and maximum pooling; It includes a target detection module, which sends the feature image obtained by the feature extraction module to the classifier for further processing and classification, and finally outputs the pest and disease category of each image.
2. The crop pest classification method based on progressive generative adversarial network according to claim 1, characterized in that: The data preprocessing in S2 includes smoothing and sharpening, and the image is uniformly reduced to a size of 256×256 using the Resize method in the Transformer class.
3. The crop pest classification method based on progressive generative adversarial network according to claim 2, characterized in that: The smoothing and sharpening process is implemented as follows: Use a Gaussian filter based on normal distribution to smooth the image and remove noise and interference information from the image; Among them, I smooth (x, y) represents the pixel value of the image after smoothing, I(x, y) represents the pixel value of the original image, G(i, j) represents the weight of the Gaussian filter, σ is the standard deviation, which controls the width of the Gaussian distribution, and k is the radius of the filter window, which determines the size of the filter; Binarize the image and use the Laplacian operator to sharpen the image to highlight the edge features of the image; Find the texture, edge and disease location of crop leaves, then add the processed image to the original image through digital image addition to complete the sharpening operation of image features: Among them, I binary (x, y) represents the binary image, I sharpened (x, y) represents the sharpened image, and L(i, j) represents the weight of the Laplacian operator.
4. The crop pest classification method based on progressive generative adversarial network according to claim 1, characterized in that: The steps for implementing the linear left channel in S3 are as follows: The image is processed using a 1×1 balanced convolution layer, where the balanced convolution layer is a convolution layer initialized with a balanced learning rate, and the formula is: Where C represents the number of channels of the input image, C′ represents the number of channels of the output image, W represents the weight of the convolution kernel in the balanced convolution layer, and I c (x, y) represents the value of the input image at position (x, y), I c’ (x, y) represents the value of the output image at position (x, y) after balanced convolution processing.
5. The crop pest classification method based on progressive generative adversarial network according to claim 1, characterized in that: The steps for implementing the linear right channel in S3 are as follows: Downsample the image using max pooling: Among them, m and n represent the height and width of the pooling window, s is the step size of the pooling, and I c (x, y) represents the value of the input image at position (x, y) and channel c, O c (x, y) represents the value of the output image at position (x, y) and channel c; The obtained image is mapped through a 1×1 convolution layer to obtain the feature image of the right channel.
6. The crop pest classification method based on progressive generative adversarial network according to claim 1, characterized in that: The steps for implementing the fade-in operation in S3 are as follows: Assign weights to the feature images obtained from the left and right channels, and then perform a dot-add operation on the two feature images: F=α×I left +(1-α)×Downsample(I right ) (7) Among them, α is the weight assigned to the current left channel feature image, I left represents the feature image of the left channel, Downsample represents the downsampling operation, (1-α) represents the weight assigned to the current feature image of the right channel, and I right Represents the feature image of the right channel.
7. The crop pest classification method based on progressive generative adversarial network according to claim 1, characterized in that: The steps for implementing the feature enhancement module are as follows: After the feature image is input into the balanced convolution block, it first undergoes a 1×1 balanced convolution to increase the dimension of the feature image in the channel. Then the LeakyRelu activation function is used to perform nonlinear transformation on the feature image; Normalize the feature vector of the feature image at the pixel level, and divide the pixel value by the standard deviation of the color channel where the pixel value is located, so that the pixel value is within the normal distribution with a mean of 0 and a standard deviation of 1. The formula is as follows: Among them, μ c represents the mean of channel c, σ c represents the standard deviation of channel c, I c (x, y) represents the value of the input image at position (x, y) and channel c, I norm,c (x, y) represents the value of the output image at position (x, y) and channel c; Finally, the feature image is input into a 3×3 balanced convolution layer. Local features are extracted in different areas of the feature image through convolution operations. In addition, a built-in edge area is added to the edge of the feature image before the convolution operation. After that, the LeakyRelu activation function and pixel-level feature vector normalization are performed, and finally the feature image is output.
8. The crop pest classification method based on progressive generative adversarial networks according to claim 1, characterized in that: The convolutional attention module includes a channel attention part and a spatial attention part; the channel attention part is implemented as follows: The feature image is input into the channel attention module and weighted according to the importance of the feature map of each channel of the feature image; The channel attention module uses two fully connected layers to compress and expand the features of each channel to obtain the importance score of each channel; Use the Sigmoid function to normalize the importance score to [0, 1] to obtain the attention weight of each channel; Apply the attention weights to the input feature map to obtain weighted features; The steps for implementing the spatial attention part are as follows: The feature image obtained by the channel attention module is input into the spatial attention module to weight the features at different spatial positions; Spatial features are weighted through a bilinear pooling operation and two fully connected layers, where the bilinear pooling captures the spatial relationship in the feature map, and the fully connected layer compresses and expands the features of each spatial position to obtain the importance score of each position; The importance score of each position is normalized and applied to the input feature map to obtain a weighted feature map.
9. The crop pest classification method based on progressive generative adversarial network according to claim 1, characterized in that: The implementation steps of the classifier are as follows: Perform small batch standard deviation processing on the feature image output by the feature extraction network, calculate the standard deviation of each channel in the current batch of samples, and obtain a four-dimensional tensor; The tensor is averaged along the channel dimension to obtain a four-dimensional tensor with one channel, and then the obtained tensor is concatenated to the feature image to obtain a new feature image with one channel added. Pass the new feature image through a 3×3 convolution layer and a 1×1 convolution layer, and reduce the number of channels of the new feature image by 1; The feature image is converted into a feature vector and input into the three-layer fully connected layer. The fully connected layer then performs a nonlinear transformation on the feature vector according to the weights and finally outputs the probability distribution of the category to obtain the classification result.