Small sample crop disease and pest identification method and device based on efficient parameter transfer learning
By using efficient parameter migration paradigm, channel space adapter and scaling offset module in the identification of small sample crop diseases and pests, the problem of insufficient identification accuracy and generalization capabilities in the existing technology small sample scenarios is solved, and higher identification accuracy and stronger adaptability are achieved.
Patent Information
- Application Number
- CN202510238656.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has problems of insufficient accuracy and generalization ability in crop pest identification in small sample scenarios, especially in scenarios where data is scarce and spatial characteristics are significant.
The efficient parameter migration paradigm is adopted to adapt to the spatial channel dimension and feature offset problems of the pretrained model by freezing the backbone network parameters of the pretrained model and inserting the channel space adapter and scaling offset module in the model backbone layer.
It significantly improves the model's identification accuracy and generalization ability on small-sample crop pest data, reduces its dependence on large-scale annotation data, and enhances the model's adaptability to spatial feature differences.
Smart Images

Figure CN120107752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of crop pest and disease identification and deep learning technology, and in particular to a small sample crop pest and disease identification method and device based on efficient parameter transfer learning. Background Art
[0002] Timely and accurate identification of crop pests and diseases is crucial to reduce yield losses and ensure sustainable production. Pests and diseases seriously affect agricultural production efficiency. Traditional pest and disease identification methods, such as clustering and manual feature extraction techniques, have been widely used in the past. However, the limitations of these methods in accuracy and efficiency restrict their applicability in modern agricultural scenarios. With the rapid development of camera hardware and computer vision technology, deep learning-based models have become the mainstream method for crop pest and disease identification, significantly surpassing the performance of traditional methods in accuracy and generalization ability.
[0003] Currently, most crop pest and disease identification models rely on transfer learning, which uses model parameters pre-trained on large-scale datasets as initialization and adapts them to downstream tasks through full tuning or linear probes. Transfer learning accelerates model convergence and improves performance, making deep learning technology more feasible in practical applications. However, a key challenge remains: the training of these models usually requires large-scale annotated datasets, which are expensive and time-consuming to collect in real scenarios.
[0004] In the few-sample scenario, traditional transfer learning methods show obvious shortcomings. In traditional transfer learning methods, the full tuning method updates all model parameters, which is prone to overfitting and has limited generalization ability; the linear probe method freezes the backbone network and only updates the classification head. Although it alleviates the overfitting problem, it performs poorly when facing new data sets with large distribution differences from the pre-training data.
[0005] It can be seen that the existing technology still has many deficiencies in crop disease and pest identification, especially in scenarios where data is scarce and spatial characteristics vary significantly. Therefore, a new method is urgently needed to solve these problems in order to improve the accuracy and efficiency of crop disease and pest identification. Summary of the invention
[0006] In order to solve the problems of insufficient adaptation to spatial channel dimensions and insufficient correction of feature offsets in the prior art, the primary purpose of the present invention is to provide a small sample crop pest and disease identification method based on efficient parameter transfer learning that can enhance the model's adaptability to spatial feature differences and effectively adapt to the spatial channel dimensions and feature offset problems of the pre-trained model.
[0007] To achieve the above object, the present invention adopts the following technical solution: a method for identifying small sample crop pests and diseases based on efficient parameter transfer learning, the method comprising the following steps in order:
[0008] (1) Collect and preprocess crop disease and insect pest image data. The preprocessed crop disease and insect pest image data form a data set, which is divided into a training set, a validation set, and a test set;
[0009] (2) Construct a pre-trained model, which includes an image block embedding layer, a model backbone layer, and a model classification layer. The target dimension of the model classification layer is set to the number of categories of crop pest and disease data;
[0010] (3) Using an efficient parameter transfer paradigm, the training set is input into the pre-trained model, the pre-trained model is transferred, and the trained model is obtained;
[0011] (4) Input the image data of crop pests and diseases to be identified into the trained model to obtain the identification results.
[0012] In step (1), the pretreatment specifically includes the following steps:
[0013] (1a) Label the acquired crop pest and disease image data by disease type;
[0014] (1b) Scaling the annotated image data, all images are scaled to 224 × 224 pixel resolution by bilinear interpolation;
[0015] (1c) Perform color standardization and standardize the three RGB channels separately.
[0016] Step (2) specifically includes the following steps:
[0017] (2a) Constructing a pre-trained model. The pre-trained model uses the Vision Transformer model, namely the ViT model. The ViT model includes an image block embedding layer, a model backbone layer, and a model classification layer.
[0018] (2b) Obtaining image data pre-training parameters, and importing the image data pre-training parameters into the image block embedding layer and the model backbone layer of the ViT model;
[0019] (2c) According to the disease types of the crop disease and insect pest data to be identified, the definition of the model classification layer of the pre-trained model is modified, and the channel dimensionality reduction parameter of the model classification layer is modified to the number of disease types of the crop disease and insect pest data to be identified.
[0020] Step (3) specifically includes the following steps:
[0021] (3a) Freeze the parameters of the image block embedding layer and the model backbone layer of the pre-trained model, that is, do not update these parameters during the transfer training process;
[0022] (3b) Insert the first efficient parameter module, namely the channel space adapter, into all the model backbone layers of the pre-trained model. The channel space adapter is composed of the channel dimension reduction matrix , channel dimension matrix , an intermediate layer 3×3 convolution module, an image token change matrix and nonlinear activation layers Composition; set the intermediate feature map of the input model backbone layer ; The length of the spatial dimension of , N is an image token of length N , 1 is a category token of length 1 ; The channel dimension is ; Channel space adapter pair The processing formula is as follows:
[0023] ;
[0024] in, express After channel dimension reduction matrix , nonlinear activation layer Features after processing;
[0025] Will Decompose into image tokens along the sequence dimension and category tokens , using a 3×3 convolutional module and an image token change matrix , for image tokens and category tokens Perform spatial dimension adaptation:
[0026] ;
[0027] ;
[0028] in, represents global average pooling, represents the image token after spatial dimension adaptation, Represents the category token after spatial dimension adaptation; ; Represents the convolution operation;
[0029] Finally and Splice along the spatial dimension and use the channel dimension-raising matrix Map feature channels to channel dimension C:
[0030] ;
[0031] In the formula, represents the output of the channel space adapter;
[0032] (3c) Insert the second efficient parameter module, namely the scaling offset module, into all the model backbone layers of the pre-trained model. The scaling offset module consists of the scaling tensor and the offset tensor Composition, scaling and offset module The processing formula is as follows:
[0033] ;
[0034] In the formula, Represents the output of the scaling and offset module:
[0035] (3d) The processing of the model backbone layer after adding the channel space adapter and the scaling offset module is as follows:
[0036] ; (1)
[0037] The output of formula (1) is used as the input of formula (2):
[0038] ; (2)
[0039] In the formula, For channel space adapter, For the scaling offset module, is the multi-head attention layer, is the layer normalization layer, is a multi-layer perception layer. To scale the hyperparameters, is the first output feature of the model backbone layer, It is the second output feature of the backbone layer of the model;
[0040] (3e) Use the channel space adapter, the scaling and offset module, and the parameters of the model classification layer to transfer the pre-trained model on the crop pest and disease data.
[0041] Another object of the present invention is to provide an electronic device, comprising:
[0042] Processor; and
[0043] A memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the small sample crop pest identification method based on efficient parameter transfer learning as described above.
[0044] The present invention also provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the small sample crop pest and disease identification method based on efficient parameter transfer learning as described above.
[0045] It can be seen from the above technical scheme that the beneficial effects of the present invention are as follows: First, the present invention improves the recognition accuracy and generalization ability under small sample data: by introducing an efficient parameter transfer paradigm, the spatial channel dimension and feature offset problem of the pre-trained model can be effectively adapted in the case of a small amount of labeled data, significantly improving the recognition accuracy and generalization ability of the model on small sample crop pest and disease data, and overcoming the problem of full tuning overfitting or insufficient linear probe performance in traditional transfer learning methods; Second, the present invention reduces the dependence of model training on large-scale labeled data: by freezing the backbone network parameters of the pre-trained model and only updating the efficient parameter module and the classification layer parameters, the dependence on large-scale labeled data is greatly reduced, which not only reduces the cost of data collection and annotation, but also improves the applicability of the model in actual agricultural scenarios, especially in data-scarce environments; Third, the present invention enhances the adaptability of the model to spatial feature differences: through the design of the channel space adapter and the scaling offset module, the model can dynamically adjust the degree of adaptation of the image spatial features and the correction ability of the feature offset, so that the model can still maintain a high recognition performance when facing a new data set with a large difference in distribution from the pre-training data, thereby better adapting to the spatial feature differences of different crop pests and diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of the method of the present invention;
[0047] Figure 2 Schematic diagram of transfer training for a pre-trained model. DETAILED DESCRIPTION
[0048] like Figure 1 As shown, a small sample crop pest and disease identification method based on efficient parameter transfer learning includes the following steps in sequence:
[0049] (1) Collect and preprocess crop disease and insect pest image data. The preprocessed crop disease and insect pest image data form a data set, which is divided into a training set, a validation set, and a test set;
[0050] (2) Construct a pre-trained model, which includes an image block embedding layer, a model backbone layer, and a model classification layer. The target dimension of the model classification layer is set to the number of categories of crop pest and disease data;
[0051] (3) Using an efficient parameter transfer paradigm, the training set is input into the pre-trained model, the pre-trained model is transferred, and the trained model is obtained;
[0052] (4) Input the image data of crop pests and diseases to be identified into the trained model to obtain the identification results.
[0053] In step (1), the pretreatment specifically includes the following steps:
[0054] (1a) Label the acquired crop pest and disease image data by disease type;
[0055] (1b) Scaling the annotated image data, all images are scaled to 224 × 224 pixel resolution by bilinear interpolation;
[0056] (1c) Perform color standardization and standardize the RGB channels separately:
[0057] ;
[0058] in, is the channel mean of ImageNet, is the standard deviation of ImageNet, .
[0059] Step (2) specifically includes the following steps:
[0060] (2a) Constructing a pre-trained model. The pre-trained model uses the Vision Transformer model, namely the ViT model. The ViT model includes an image block embedding layer, a model backbone layer, and a model classification layer.
[0061] The image block embedding layer is used to divide the input image into blocks of 16×16 pixels, a total of 14×14 blocks, and generate a 768-dimensional embedding vector through linear projection;
[0062] The model backbone layer consists of 12 Transformer encoder layers, each of which contains a multi-head self-attention layer (MHSA) and a multi-layer perceptron (MLP);
[0063] Model classification layer: The output dimension of the ViT model is 1,000 (corresponding to the number of categories in the ImageNet pre-training data).
[0064] (2b) Obtaining image data pre-training parameters, and importing the image data pre-training parameters into the image block embedding layer and the model backbone layer of the ViT model;
[0065] (2c) According to the disease types of the crop disease and insect pest data to be identified, the definition of the model classification layer of the pre-trained model is modified, and the channel dimensionality reduction parameter of the model classification layer is modified to the number of disease types of the crop disease and insect pest data to be identified.
[0066] like Figure 2 As shown, step (3) specifically includes the following steps:
[0067] (3a) Freeze the parameters of the image block embedding layer and the model backbone layer of the pre-trained model, that is, do not update these parameters during the transfer training process;
[0068] (3b) Insert the first efficient parameter module, namely the channel space adapter, into all the model backbone layers of the pre-trained model. The channel space adapter is inserted into the multi-head self-attention layer and the multi-layer perception layer in the model backbone layer in parallel, and a scaling hyperparameter is used. The proportion of the channel space adapter output in the model backbone layer output is adjusted according to the different crop data categories.
[0069] Scaling Hyperparameters .
[0070] like Figure 2 As shown, the channel space adapter consists of a channel dimension reduction matrix , channel dimension matrix , nonlinear activation layer and a spatial adaptation module; the spatial adaptation module includes an intermediate layer 3×3 convolution module, an image token change matrix ; Set the intermediate feature map of the input model backbone layer ; The length of the spatial dimension of , N is an image token of length N , 1 is a category token of length 1 ; The channel dimension is ; Channel space adapter pair The processing formula is as follows:
[0071] ;
[0072] in, express After channel dimension reduction matrix , nonlinear activation layer Features after processing;
[0073] Will Decompose into image tokens along the sequence dimension and category tokens , using a 3×3 convolutional module and an image token change matrix , for image tokens and category tokens Perform spatial dimension adaptation:
[0074] ;
[0075] ;
[0076] in, represents global average pooling, represents the image token after spatial dimension adaptation, Represents the category token after spatial dimension adaptation; ; Represents the convolution operation;
[0077] Finally and Splice along the spatial dimension and use the channel dimension-raising matrix Map feature channels to channel dimension C:
[0078] ;
[0079] In the formula, represents the output of the channel space adapter;
[0080] (3c) Insert the second efficient parameter module, namely the scaling offset module, into all the model backbone layers of the pre-trained model. The scaling offset module is inserted in series into the output part of the multi-head self-attention layer and the multi-layer perception layer to ensure timely correction of the offset effect of the feature on the downstream data. The scaling offset module consists of the scaling tensor and the offset tensor Composition, scaling and offset module The processing formula is as follows:
[0081] ;
[0082] In the formula, Represents the output of the scaling and offset module:
[0083] (3d) The processing of the model backbone layer after adding the channel space adapter and the scaling offset module is as follows:
[0084] ; (1)
[0085] The output of formula (1) is used as the input of formula (2):
[0086] ; (2)
[0087] In the formula, For channel space adapter, For the scaling offset module, is a multi-head self-attention layer, is the layer normalization layer, is a multi-layer perception layer. To scale the hyperparameters, is the first output feature of the model backbone layer, It is the second output feature of the backbone layer of the model;
[0088] (3e) Use the channel space adapter, the scaling and offset module, and the parameters of the model classification layer to transfer the pre-trained model on the crop pest and disease data.
[0089] Embodiment 1
[0090] This embodiment uses an apple leaf disease dataset as an example. The apple leaf disease dataset is the AppleLeaf9 dataset, which contains 9 types of apple leaf disease images and a total of 14,582 image samples. According to the type label of each image disease, a structured small sample conditional constraint dataset is formed.
[0091] 1000 samples were uniformly sampled from the 14582 image samples as the training set for the final training, and the remaining 13582 samples were used as the test set. The selected 1000 samples were further divided into 800 and 200 samples uniformly divided as the training set and validation set used for the model hyperparameter search in the fine-tuning stage.
[0092] Test set accuracy: The Top-1 accuracy of the final model on the test set is shown in Table 1. The comparison methods include full-tuning of the traditional transfer learning paradigm, linear classification head (Linear-probe), adapter (Adapter) of the efficient parameter transfer paradigm, low-rank adaptation (LoRA) and deep visual prompt fine-tuning (VPT-deep). All experiments were repeated three times on three different random seeds, and the results reported are the mean and standard deviation of the three experimental results.
[0093] Table 1 Experimental results under 1k sample setting
[0094] Method Param. (M) AppleLeaf9 Full-tuning 85.82 76.62 ± 2.72 Linear-probe 0.02 83.89 ± 0.07 <![CDATA[Adapter 8 ]]> 0.34 <![CDATA[ 94.38 ± 0.54]]> <![CDATA[LoRA 16 ]]> 0.32 93.95 ± 0.11 <![CDATA[VPT-deep 32 ]]> 0.32 93.58 ± 0.41 CSAdapter+SSM of the present invention 0.35 94.85 ± 0.13
[0095] The experimental results under the extremely small sample setting are shown in Table 2:
[0096] Table 2
[0097] Method\shot 1-shot 2-shot 4-shot 8-shot Full-tuning 28.18 ± 2.70 34.46 ± 5.30 31.74 ± 6.84 40.12 ± 0.81 Linear-probe 34.60 ± 0.87 38.38 ± 0.51 66.45 ± 0.27 64.70 ± 0.22 <![CDATA[Adapter 8 ]]> <![CDATA[ 34.95 ± 0.44]]> <![CDATA[ 46.80 ± 2.07]]> 77.77 ± 0.57 <![CDATA[ 77.26 ± 0.56]]> <![CDATA[LoRA 16 ]]> 34.05 ± 1.49 44.19 ± 1.51 72.42 ± 1.31 76.24 ± 1.32 <![CDATA[VPT-deep 32 ]]> 34.79 ± 1.51 44.53 ± 1.85 61.29 ± 2.80 74.21 ± 0.86 CSAdapter+SSM 36.94 ± 1.75 47.78 ± 2.59 <![CDATA[ 75.92 ± 1.31]]> 78.97 ± 0.49
[0098] Parameter efficiency: The channel space adapter and scaling offset module only introduce 0.35M trainable parameters (full fine-tuning requires 85.82M original ViT-B / 16 parameters), significantly reducing the computational cost.
[0099] In summary, the present invention improves the recognition accuracy and generalization ability under small sample data: by introducing an efficient parameter transfer paradigm, the spatial channel dimension and feature offset problems of the pre-trained model can be effectively adapted in the case of a small amount of labeled data, significantly improving the recognition accuracy and generalization ability of the model on small sample crop pest and disease data, and overcoming the problems of full tuning overfitting or insufficient linear probe performance in traditional transfer learning methods; the present invention reduces the dependence of model training on large-scale labeled data: by freezing the backbone network parameters of the pre-trained model and only updating the efficient parameter module and the classification layer parameters, the dependence on large-scale labeled data is greatly reduced, which not only reduces the cost of data collection and annotation, but also improves the applicability of the model in actual agricultural scenarios, especially in data-scarce environments; the present invention enhances the adaptability of the model to spatial feature differences: through the design of the channel space adapter and the scaling offset module, the model can dynamically adjust the degree of adaptation of the image spatial features and the correction ability of the feature offset, so that the model can still maintain a high recognition performance when facing a new data set with a large difference in distribution from the pre-training data, thereby better adapting to the spatial feature differences of different crop pests and diseases.
Claims
1. A small sample crop pest identification method based on efficient parameter transfer learning, characterized by: The method comprises the following steps in order: (1) Collect and preprocess crop disease and insect pest image data. The preprocessed crop disease and insect pest image data form a data set, which is divided into a training set, a validation set, and a test set; (2) Constructing a pre-trained model, which includes an image block embedding layer, a model backbone layer, and a model classification layer. The target dimension of the model classification layer is set to the number of categories of crop pest and disease data; (3) Using an efficient parameter transfer paradigm, the training set is input into the pre-trained model, the pre-trained model is transferred, and the trained model is obtained; (4) Input the image data of crop pests and diseases to be identified into the trained model to obtain the identification results.
2. The small sample crop pest identification method based on efficient parameter transfer learning according to claim 1 is characterized by: In step (1), the pretreatment specifically includes the following steps: (1a) Label the acquired crop pest and disease image data by disease type; (1b) Scaling the annotated image data, all images are scaled to 224 × 224 pixel resolution by bilinear interpolation; (1c) Perform color standardization and standardize the three RGB channels separately.
3. The small sample crop pest identification method based on efficient parameter transfer learning according to claim 1 is characterized in that: Step (2) specifically includes the following steps: (2a) Constructing a pre-trained model. The pre-trained model uses the Vision Transformer model, namely the ViT model. The ViT model includes an image block embedding layer, a model backbone layer, and a model classification layer. (2b) Obtaining image data pre-training parameters, and importing the image data pre-training parameters into the image block embedding layer and the model backbone layer of the ViT model; (2c) According to the disease types of the crop disease and insect pest data to be identified, the definition of the model classification layer of the pre-trained model is modified, and the channel dimensionality reduction parameter of the model classification layer is modified to the number of disease types of the crop disease and insect pest data to be identified.
4. The small sample crop pest identification method based on efficient parameter transfer learning according to claim 1 is characterized in that: Step (3) specifically includes the following steps: (3a) Freeze the parameters of the image block embedding layer and the model backbone layer of the pre-trained model, that is, do not update these parameters during the transfer training process; (3b) Insert the first efficient parameter module, namely the channel space adapter, into all the model backbone layers of the pre-trained model. The channel space adapter is composed of the channel dimension reduction matrix , channel dimension matrix , an intermediate layer 3×3 convolution module, an image token change matrix and nonlinear activation layers Composition; set the intermediate feature map of the input model backbone layer ; The length of the spatial dimension of , N is an image token of length N , 1 is a category token of length 1 ; The channel dimension is ; Channel space adapter pair The processing formula is as follows: ; in, express After channel dimension reduction matrix , nonlinear activation layer Features after processing; Will Decompose into image tokens along the sequence dimension and category tokens , using a 3×3 convolutional module and an image token change matrix , for image tokens and category tokens Perform spatial dimension adaptation: ; ; in, represents global average pooling, represents the image token after spatial dimension adaptation, Represents the category token after spatial dimension adaptation; ; Represents the convolution operation; Finally and Splice along the spatial dimension and use the channel dimension-raising matrix Map feature channels to channel dimension C: ; In the formula, represents the output of the channel space adapter; (3c) Insert the second efficient parameter module, namely the scaling offset module, into all the model backbone layers of the pre-trained model. The scaling offset module consists of the scaling tensor and the offset tensor Composition, scaling offset module The processing formula is as follows: ; In the formula, Represents the output of the scaling and offset module: (3d) The processing of the model backbone layer after adding the channel space adapter and the scaling offset module is as follows: ;(1) The output of formula (1) is used as the input of formula (2): ;(2) In the formula, For channel space adapter, For the scaling offset module, is the multi-head attention layer, is the layer normalization layer, is a multi-layer perception layer. To scale the hyperparameters, is the first output feature of the model backbone layer, It is the second output feature of the backbone layer of the model; (3e) Using the parameters of the channel space adapter, the scaling offset module, and the model classification layer to migrate the pre-trained model on the crop pest data. An electronic device comprises: a processor; and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the small sample crop pest identification method based on efficient parameter transfer learning as described in any one of claims 1 to 4.
5. A computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the processor executes the small sample crop pest and disease identification method based on efficient parameter transfer learning as described in any one of claims 1 to 4.