Small-sample crop disease recognition method based on feature extraction, storage medium
By using small sample learning technology and feature attention module in the crop disease recognition model, combined with the pre-trained ResNet-18 model and Transformer structure, the problem of overfitting traditional models under small sample data is solved, and crop disease recognition with high accuracy and good generalization performance is achieved.
Patent Information
- Application Number
- CN202210242480.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-03-11
AI Technical Summary
Traditional crop disease recognition models based on convolutional neural networks require a large number of labeled instances and are prone to overfitting, making it difficult to effectively identify crop diseases under small sample data.
A small sample crop disease recognition method based on feature extraction is adopted, and a small sample learning technology and feature attention module are used, combined with the pre-trained ResNet-18 model and Transformer structure, embedded functions and distance calculation functions are constructed to improve the recognition accuracy and generalization performance of the model.
Through the combination of small sample learning technology and feature attention module, crop diseases can be effectively identified under small sample data, improving identification accuracy and generalization performance, and reducing the risk of overfitting.
Smart Images

Figure CN114693990B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to a method for recognizing small-sample crop diseases based on feature extraction and a storage medium. Background Art
[0002] The diagnosis and recognition of crop diseases play a crucial role in ensuring the high quality and quantity of food production. Therefore, accurately and timely detecting crop diseases is crucial for ensuring maximum agricultural yields, especially beneficial for farmlands in remote areas. In recent years, with the development of computer vision technology and convolutional neural network technology, the automatic recognition and diagnosis technology of crop diseases has gradually replaced the manual diagnosis method.
[0003] Traditional deep learning models based on convolutional neural networks require thousands of labeled instances for each category, which is a prerequisite for ensuring the performance of disease recognition models. However, in actual situations, agricultural disease image data is difficult to obtain. At this time, the disease recognition model based on convolutional neural networks is very susceptible to the problem of overfitting, thus affecting the implementation of the next protection measures.
[0004] Few-Shot Learning (FSL) generally refers to the method and scenario of learning from a small amount of labeled data. Ideally, a model capable of few-shot learning can also be quickly applied to new fields. Few-shot learning is an idea, not specifically referring to a particular algorithm or model, nor is there a general template and solution. Generally, it is necessary to focus on specific problems and application scenarios. Summary of the Invention
[0005] A method for recognizing small-sample crop diseases based on feature extraction proposed by the present invention uses few-shot learning technology to automatically classify different pests, plants, and their diseases.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for recognizing small-sample crop diseases based on feature extraction includes obtaining crop disease image data and inputting it into a pre-constructed small-sample crop disease recognition model for disease recognition. The construction steps of the small-sample crop disease recognition model are as follows:
[0008] Construct an experimental dataset according to PlantVillage;
[0009] Preprocess the experimental dataset;
[0010] Build a few-shot learning model and start training;
[0011] After training is completed, input the sample images of the test set to verify the performance of the model;
[0012] Among them, the embedding function of the small-sample crop disease recognition model includes a feature extraction module and a feature attention module. The role of the feature extraction module is to extract the features of the sample data and map them into a d-dimensional Euclidean space. The mapping result is a d-dimensional embedding vector, that is, a feature vector; this feature extraction module uses the ResNet-18 model pre-trained on the ImageNet dataset;
[0013] The feature attention module is a feature attention module based on the Transformer structure. The attention module learns the correlation between different classification tasks. Through the set adaptive method, it adapts the feature extraction model, learns the features related to the target task, and makes it adapt to different categories of classification tasks.
[0014] Furthermore, the distance calculation function of the small-sample crop disease recognition model is the Mahalanobis distance. The Mahalanobis distance calculation function measures the similarity between the d-dimensional embedding vectors obtained by calculating two samples through the embedding function respectively. Its calculation formula is as follows: Among them, is the covariance matrix between n∈N categories in the one-dimensional learning task t∈T.
[0015] Furthermore, the small-sample crop disease recognition model is all implemented using the PyTorch deep learning framework. Stochastic gradient descent is used to train the model. The initial learning rate is set to 0.0002 and accompanied by a weight decay strategy. The convolutional neural network models used include the RestNet-18 model and the Transformer model. Among them, the RestNet-18 model is used for the feature extraction module, and the Transformer model is used for the feature attention module.
[0016] Furthermore, before inputting the sample images into the feature extraction module, all images are scaled to a pixel size of 84×84×3. This dataset is divided into 3 different and independent parts, and each part contains a meta-training set and a meta-test set.
[0017] Furthermore, the initial parameters of the ResNet18 model use the pre-trained parameters of this model on the ImageNet public dataset. ResNet18 maps the image into a d-dimensional embedding vector φ x , which corresponds to the feature vector of the image; the Transformer module is based on the obtained embedding vector, and the feature vectors φ x of all images in the support set are calculated through the Transfomer module to obtain a new feature vector ψ with attention informationx ;
[0018]
[0019] Among them,
[0020] The values of the calculation parameters Q, K, and V in the Transformer module are
[0021]
[0022] Q, K, and V are the same and are the set composed of all support samples in the training set;
[0023]
[0024]
[0025]
[0026] W Q , W K , W V are weight matrices, and |Q|, |K|, and |V| represent the number of elements in the set;
[0027] φ x : The input image is a d-dimensional vector obtained by calculating through the Resnet-18 model;
[0028] ψ x : A d-dimensional vector after adding auxiliary information using the Transformer structure on the basis of φ x .
[0029] Furthermore, after the few-shot crop disease recognition model is trained once on the support set, the loss is obtained on the query set, and the Mahalanobis distance is used to measure the similarity between feature vectors:
[0030]
[0031] Among them, is the covariance matrix for n ∈ N categories involved in a single task t ∈ T, and the covariance matrix can be estimated by the regular estimator method;
[0032] The total loss function of the entire model is as follows:
[0033]
[0034] Among them, is the mean of the feature vectors ψ x of all samples in each category, and y qis the true class label corresponding to the query sample in the test dataset, is the predicted output class of the model of the present invention, λ is a constant weight value set during model training, l is the cross-entropy loss function, and the main function of the second part of this formula is to train the Transformer structure in the model.
[0035] Furthermore, the experimental dataset constructed by PlantVillage contains a total of 38 classes, which will be divided into 3 groups using different methods. For each group, 10 classes are selected as the meta-test set, and the remaining 28 classes are used as the meta-training set. The 10 classes included in the meta-test set of each group are non-repetitive. During the meta-training stage, for each meta-task, only 5 samples from 5 classes are selected as the support sample set, and 1 additional sample is selected as the query sample.
[0036] Furthermore, in each ResNet-18 structure of the residual network, its activation function uses a linear activation unit, and batch normalization technology is used to suppress the overfitting phenomenon of the model. A global average pooling layer is added to the last layer of the residual network to generate the required feature vector for calculation.
[0037] Furthermore, the probability of dropout used in the Transformer structure is set to 0.5.
[0038] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.
[0039] The small-sample crop disease recognition method based on feature extraction of the present invention constructs a small-sample crop disease dataset by using the labeled crop disease image dataset open-sourced on some professional agricultural websites or collecting relevant crop disease image data through web crawlers. In actual situations, agricultural disease categories have the characteristics of multi-species diversity, so it is extremely difficult to obtain a large amount of data for a single species and a single disease. Therefore, a small-sample learning paradigm is adopted to construct a meta-training set (Meta Training Set) and a meta-test set (Meta Testing Set). Among them, M sample data are randomly selected from each of the N classes in the meta-training set as the support set (Support Set), and then one remaining sample from each class is selected as the query set (Query Set). During the meta-training (Meta Training) stage, the support set and the query set constitute a meta-learning task (Meta Task). Similarly, the same settings are made during the meta-testing (Meta Testing) stage.
[0040] In each meta-learning task of the model, independent random sampling is performed on the training data. Since different meta-tasks are sampled in each training, overall, the training contains different class combinations. This mechanism enables the model to learn the common parts in different meta-learning tasks, such as how to extract important features and compare the similarity of samples. The model learned through this learning mechanism can also perform classification well when facing new unseen meta-learning tasks. The entire classifier model has two improvements compared to existing few-shot learning techniques:
[0041] a) First, the embedding function consists of two parts: a conventional feature extraction module and a feature attention module. The role of the feature extraction module is to extract the features of the sample data and map them into a d-dimensional Euclidean space. The mapping result is a d-dimensional embedding vector (embeddings), which is also the feature vector. The feature extraction module uses the ResNet-18 model pre-trained on the ImageNet dataset. However, the feature vector calculated by ResNet-18 does not achieve the ideal effect, that is, the feature vector it extracts cannot express some key features hidden in the image samples, and it is not specifically designed for the target task, so it may perform poorly in the process of applying to different category tasks. The present invention improves this deficiency by adding a feature attention module based on the Transformer structure, which can use the self-attention mechanism to make the model pay more attention to the regions in the image that are more important for the embedding vector. This attention module mainly learns the correlation between different classification tasks, adapts the feature extraction model through the set adaptive method, and learns the features related to the target task to make it adapt to different category classification tasks.
[0042] b) Second, the distance calculation function is changed from the conventional Euclidean distance to the Mahalanobis Distance. The role of the distance calculation function is to measure the similarity between the d-dimensional embedding vectors calculated by two samples through the embedding function respectively. Compared with the Euclidean distance, the Mahalanobis distance can be regarded as a correction of the Euclidean distance, which improves the problem that the calculation scales of each dimension are inconsistent and correlated. Its calculation formula is as follows: where, is the covariance matrix between n ∈ N classes in a meta-learning task t ∈ T.
[0043] Based on the publicly available crop disease dataset PlantVillage containing 38 categories, the present invention divides it into three different and independent datasets. Independent means that there are no duplicate category labels in these three datasets. The few-shot learning model of the present invention has a relatively high improvement in recognition accuracy compared to the excellent few-shot models previously applied to crop disease classification.
[0044] As can be seen from the above technical solution, the few-shot crop disease recognition method based on feature extraction of the present invention is an improved few-shot recognition method for automatically classifying and recognizing agricultural diseases based on feature extraction. This model has good recognition accuracy and generalization performance. Brief Description of the Drawings
[0045] Figure 1 It is a schematic diagram of feature vector extraction of an embodiment of this application;
[0046] Figure 2 It is a flowchart of the few-shot disease recognition method of this application. Detailed Description of the Embodiment
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0048] The few-shot crop disease recognition method based on feature extraction described in this embodiment includes the following steps: obtaining crop disease image data and inputting it into a pre-constructed few-shot crop disease recognition model for disease recognition. The steps for constructing the few-shot crop disease recognition model are as follows:
[0049] Construct an experimental dataset according to PlantVillage;
[0050] Preprocess the experimental dataset;
[0051] Build a few-shot learning model and start training;
[0052] After training is completed, input the test set sample images to verify the model performance;
[0053] Among them, the embedding function of the few-shot crop disease recognition model includes a feature extraction module and a feature attention module. The role of the feature extraction module is to extract the features of the sample data and map them into a d-dimensional Euclidean space. The mapping result is a d-dimensional embedding vector, that is, a feature vector. This feature extraction module uses the ResNet-18 model pre-trained on the ImageNet dataset;
[0054] Such asFigure 1 The figure shows a schematic diagram of feature vector extraction for an embodiment of the present application; among them, Input: input sample x; Embeddings: φ x ; TransformEmbeddings: ψ x ; Embeddings and TransformEmbeddings are input into the calculation of the total loss function formula of the entire model;
[0055] The feature attention module is a feature attention module based on the Transformer structure. The attention module learns the correlation between different classification tasks, adapts the feature extraction model through a set adaptive method, learns features related to the target task, and makes it adapt to different categories of classification tasks.
[0056] The following is a specific description respectively:
[0057] All neural network models proposed by the present invention are implemented using the PyTorch deep learning framework, and the model is trained using stochastic gradient descent. The initial learning rate is set to 0.0002 and accompanied by a weight decay strategy. The convolutional neural network models involved are mainly the RestNet-18 model and the Transformer model. Among them, the RestNet-18 model is used for the feature extraction module, and the Transformer model is used for the feature attention module.
[0058] Before inputting the sample image into the feature extraction module, all images are scaled to a pixel size of 84×84×3. The PlantVillage public dataset contains data of 38 categories of different crop disease leaf images and healthy leaf images, with a total of 61,486 images. The present invention divides this dataset into 3 different and independent parts. Each group contains a meta-training set and a meta-test set. Further, the meta-test set of each group contains 10 different categories, and the training set of this group consists of the remaining 28 categories. In addition, the categories included in these three meta-test sets do not repeat.
[0059] The initial parameters of the ResNet18 model described in step 1 adopt the pre-trained parameters of this model on the ImageNet public dataset. The role of ResNet18 is to map the image into a d-dimensional embedding vector φ x , that is, the feature vector corresponding to the image. The Transformer module is based on the embedding vector obtained in step 3, and the feature vectors φ of all images in the support set x After being calculated by the Transfomer module, a new feature vector ψ with attention information is obtained x。
[0060]
[0061] Among them,
[0062] Among them, the values of the calculation parameters Q, K, and V in the Transformer module are
[0063]
[0064] Among them, Q, K, and V are the same, and they are the set composed of all support samples in the training set.
[0065]
[0066]
[0067]
[0068] W Q , W K , W V are weight matrices, and |Q|, |K|, and |V| represent the number of elements in the set.
[0069] φ x : The input image is a d-dimensional vector obtained by calculating through the Resnet-18 model;
[0070] ψ x : A d-dimensional vector after adding auxiliary information using the Transformer structure on the basis of φ x ;
[0071] Among them, ψ x is calculated on the basis of φ x . First, φ x is calculated, and then ψ x is calculated through the Transformer structure;
[0072] After the model is trained once on the support set, the loss is calculated on the query set. The present invention uses the Mahalanobis distance to replace the Euclidean distance to measure the similarity degree between feature vectors.
[0073]
[0074] Among them, is the covariance matrix for n ∈ N categories involved in a single task t ∈ T. The covariance matrix can be estimated by the regular estimator method. To ensure that the feature vector ψ x, and expanding the feature vector ψ calculated from different categories of samples x The distance between them.
[0075] The total loss function of the entire model is as follows:
[0076]
[0077] Where is the mean of the feature vectors ψ of all samples in each category x y q is the true class label corresponding to the query sample in the test dataset, is the predicted output class of the model of the present invention, and λ is a constant weight value set when training the model. l is the cross-entropy loss function, and the main function of the second part of this formula is to train the Transformer structure in the model.
[0078] Dataset division: The PlantVillage dataset contains a total of 38 categories, and there will be 3 groups of different division methods. Each group selects 10 categories as the meta-test set, and the remaining 28 categories as the meta-training set part, where the 10 categories included in the meta-test set of each group are not repeated. In the meta-training stage, for each meta-task, only 5 samples from 5 categories are selected as the support sample set, and another 1 sample is selected as the query sample.
[0079] Specific model details: In each ResNet-18 structure, its activation function uses the rectified linear unit (ReLU), and the batch normalization technique (BN) is used to suppress the overfitting phenomenon of the model. A global average pooling layer is added to the last layer of the residual network to generate the required feature vector. The probability of dropout used in the Transformer structure is set to 0.5.
[0080] Experimental results: As shown in the following table, compared with the classical few-shot learning paradigm, specifically referring to the model without introducing the Transformer structure. After introducing the Transformer feature attention module in the existing technical solution of the present invention, the recognition accuracy rate has been effectively improved.
[0081] Experimental model Group 1 Group 2 Group 3 Classical method 0.53 0.77 0.69 Method of the present invention 0.64 0.82 0.89
[0082] As can be seen from the above technical solution, the few-shot crop disease recognition method based on feature extraction of the present invention is an improved few-shot recognition method for automatically classifying and recognizing agricultural diseases based on feature extraction. This model has good recognition accuracy and generalization performance.
[0083] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of any of the above methods.
[0084] In yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of any of the above methods.
[0085] In yet another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer causes the computer to execute the steps of any of the methods in the above embodiments.
[0086] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples and beneficial effects of the relevant content, reference can be made to the corresponding parts in the above methods.
[0087] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided by the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0088] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0089] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A small-sample crop disease recognition method based on feature extraction, which obtains crop disease image data and inputs it into a pre-constructed small-sample crop disease recognition model for disease recognition. It is characterized in that: The steps for constructing the small-sample crop disease recognition model are as follows: Construct an experimental dataset according to PlantVillage; Preprocess the experimental dataset; Build a small-sample learning model and start training; After training is completed, input the test set sample images to verify the model performance; Among them, the embedding function of the small-sample crop disease recognition model includes a feature extraction module and a feature attention module. The role of the feature extraction module is to extract the features of the sample data and map them into a d-dimensional Euclidean space. The mapping result is a d-dimensional embedding vector, that is, a feature vector. This feature extraction module uses the ResNet-18 model pre-trained on the ImageNet dataset; The feature attention module is a feature attention module based on the Transformer structure. The attention module learns the correlation between different classification tasks. Through the set adaptive method, it adapts the feature extraction model, learns the features related to the target task, and makes it adapt to different classification tasks of different categories; Before inputting the sample images into the feature extraction module, all the images are scaled to a pixel size of 84×84×3. The dataset is divided into 3 different and independent parts, and each part contains a meta-training set and a meta-test set; After the small-sample crop disease recognition model is trained once on the support set, the loss is calculated on the query set. The Mahalanobis distance is used to measure the similarity between feature vectors: Among them, is the covariance matrix for n ∈ N categories involved in a single task t ∈ T, and the covariance matrix can be estimated by the regular estimator method; The total loss function of the entire model is as follows: Among them, is the mean of the feature vectors ψ of all samples in each category x of y q is the true class label corresponding to the query sample in the test dataset, is the predicted output class of the model of the present invention, λ is a constant weight value set during model training, l is the cross-entropy loss function, and the main function of the second part of this formula is to train the Transformer structure in the model.
2. The small-sample crop disease recognition method based on feature extraction according to claim 1, It is characterized in that: The distance calculation function of the small-sample crop disease recognition model is the Mahalanobis distance. The Mahalanobis distance calculation function measures the similarity between the d-dimensional embedding vectors obtained by calculating two samples through the embedding function respectively. The calculation formula is as follows: where, is the covariance matrix between n∈N categories in the one-dimensional learning task t∈T.
3. The small-sample crop disease recognition method based on feature extraction according to claim 2, It is characterized in that: The small-sample crop disease recognition model is all implemented using the PyTorch deep learning framework. Stochastic gradient descent is used to train the model. The initial learning rate is set to 0.0002 and accompanied by a weight decay strategy. The convolutional neural network models used include the RestNet-18 model and the Transformer model. Among them, the feature extraction module uses the RestNet-18 model, and the feature attention module uses the Transformer model.
4. The small-sample crop disease recognition method based on feature extraction according to claim 3, It is characterized in that: The initial parameters of the ResNet-18 model are the pre-trained parameters of this model on the ImageNet public dataset. The ResNet-18 model maps an image into a d-dimensional embedding vector φ x , which corresponds to the feature vector of the image. The Transformer model, on the basis of obtaining the embedding vector, takes the feature vectors φ x of all images in the support set and obtains a new feature vector ψ with attention information after being calculated by the Transfomer model x ; Among them, The values of the calculation parameters Q, K, and V in the Transformer module are Q, K, and V are the same, and they are the set composed of all support samples in the training set; W Q ,W K ,W V is the weight matrix, and |Q|, |K|, |V| represent the number of elements in the set; φ x is a d-dimensional vector obtained by calculating the input image through the Resnet-18 model; ψ x is a d-dimensional vector obtained by adding auxiliary information to φ x using the Transformer structure.
5. The small-sample crop disease recognition method based on feature extraction according to claim 1, It is characterized in that: The experimental dataset constructed by PlantVillage contains a total of 38 categories, which will be divided into 3 groups. For each group, 10 categories are selected as the meta-test set, and the remaining 28 categories are used as the meta-training set. The 10 categories included in the meta-test set of each group are non-repetitive. During the meta-training stage, for each meta-task, only 5 samples from 5 categories are selected as the support sample set, and another 1 sample is selected as the query sample.
6. The few-shot crop disease recognition method based on feature extraction according to claim 3, characterized in that: In each ResNet-18 structure, the activation function uses a linear activation unit, and batch normalization technology is used to suppress the overfitting phenomenon of the model. A global average pooling layer is added to the last layer of the residual network to generate the required feature vector for calculation.
7. The few-shot crop disease recognition method based on feature extraction according to claim 3, characterized in that: The probability of dropout used in the Transformer structure is set to 0.
5.
8. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Small sample and zero sample image classification method based on metric learning and meta-learning
CN109961089A
Plant disease identification method and device, electronic equipment and storage medium
CN113869098A