Method for identifying plant diseases and insect pests of tomato leaves
Through deep learning technology, the multi-integrated deep model network is constructed, which solves the problems of low accuracy and low efficiency of tomato pest recognition in the existing technology, and achieves efficient and accurate pest recognition, which improves the yield and quality of tomato crops.
Patent Information
- Application Number
- CN202510042647.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems of low accuracy, low efficiency and high cost in tomato pest identification, and cannot be widely used in large areas of tomato crops. Human error identification may lead to wrong prevention and control strategies, affecting the yield and quality of tomato crops.
Deep learning technology is adopted to build a deep learning model through the Pytorch framework, use preprocessing modules, classification modules and fusion modules, combine ResNet and ResNeXt models, add residual structures and batch normalization to build a multi-integrated deep model network to identify pests and diseases of tomato leaf.
It improves the yield and quality of tomato crops, reduces the economic losses caused by tomato pests, and achieves efficient and accurate pest identification.
Smart Images

Figure CN119964148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a method for identifying tomato leaf diseases and insect pests. Background Art
[0002] The outbreak of tomato pests and diseases will lead to a serious decline in tomato yield and quality, thus having a great impact on the entire agricultural economy. Therefore, the detection and identification of tomato pests and diseases is particularly important for the prevention and control of tomato pests and diseases. At present, the manual identification of tomato pests and diseases has problems such as low accuracy, low efficiency, and high cost, and cannot be widely used in large areas of tomato crops. In addition, due to manual misidentification, the wrong prevention and control strategy may be used, which will affect the yield and quality of tomato crops, thereby causing losses to the local agricultural economy. Summary of the invention
[0003] In view of this, the present invention provides a method for identifying tomato leaf diseases and pests, so as to improve the yield and quality of tomato crops and reduce the economic losses caused by tomato diseases and pests.
[0004] In a first aspect, the present invention provides a method for identifying tomato leaf diseases and insect pests, the method comprising:
[0005] Step 1: Select and preprocess sample images, 70% of which are used as training samples and 30% as verification samples;
[0006] Step 2: Build a deep learning model based on the Pytorch framework;
[0007] Step 3: Train and optimize the constructed deep learning model through training samples, and verify the trained and optimized model through verification samples to identify tomato leaf diseases and pests.
[0008] Optionally, step 1 includes:
[0009] The sample images come from experimental materials, from which image data of 8 common diseases including tomato health are selected; the sample images are preprocessed and cropped to a size of 256×256.
[0010] Optionally, the step 2 includes:
[0011] The deep learning model includes a preprocessing module, a classification module, and a fusion module;
[0012] The addition of the residual structure can effectively improve the network degradation problem. By making the output of each layer contain not only the feature representation of the current layer but also the information of the previous layer, it helps the network to better transfer gradients, alleviate the gradient disappearance and gradient explosion problems, and improve model performance.
[0013] Batch Normalization is used to adjust numerical data to a common scale without distorting its shape. The overall steps include: normalization, rescaling, and offset; Normalization is the process of converting data to a mean of 0 and a standard deviation of 1; In this step, there is a batch input from layer h. First, the mean of this hidden activation needs to be calculated, and its expression is:
[0014]
[0015] Where m is the number of neurons in layer h; the next step is to calculate the standard deviation of the hidden activations, which is expressed as:
[0016]
[0017] Then subtract the mean from each input and divide it by the sum of the standard deviation and the smoothing term ε. The smoothing term ε is a very small constant that prevents the denominator from being zero to ensure numerical stability in the operation.
[0018]
[0019] Rescaling and Shifting In the final operation, the input is rescaled and shifted; the rescaling parameter γ and the shift parameter β are expressed as:
[0020] h i =γ×h i(norm) +β;
[0021] After batch normalization, the output of each neuron follows the standard normal distribution of the entire batch, making the optimization smoother and the optimizer work easier. It will not be affected by gradient vanishing or gradient exploding when the learning rate becomes larger.
[0022] The ResNeXt model is proposed based on ResNet, which improves the accuracy while reducing the model complexity and the number of hyperparameters; and adds group convolution;
[0023] ResNeXt model architecture: The 3×3 convolution in ResNet is replaced by group convolution. A concept of Cardinality is added to the model, which is actually a convolution group. Compared with increasing the width and depth of the network, adding groups is better. The ResNeXt residual block first groups the 1×1 convolution into Groups=32.
[0024] In order to improve the recognition accuracy, a multi-integrated deep model network is obtained by combining the ResNet and ResNeXt models.
[0025] In the technical solution provided by the present invention, the method includes selecting and preprocessing sample images; constructing a deep learning model based on the Pytorch framework; training and optimizing the constructed deep learning model through training samples, and verifying the trained and optimized model through verification samples to identify tomato leaf diseases and pests. This method improves the yield and quality of tomato crops and reduces the economic losses caused by tomato diseases and pests. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 A flow chart of a method for identifying tomato leaf diseases and insect pests provided by an embodiment of the present invention;
[0028] Figure 2 A schematic diagram of a sample image provided by an embodiment of the present invention;
[0029] Figure 3 A schematic diagram of a ResNet residual structure provided by an embodiment of the present invention;
[0030] Figure 4 A schematic diagram of a ResNeXt residual block provided by an embodiment of the present invention;
[0031] Figure 5 A schematic diagram of an integrated depth model provided by an embodiment of the present invention;
[0032] Figure 6 The sample image processing results provided by the embodiment of the present invention;
[0033] Figure 7 A schematic diagram of training loss and accuracy provided by an embodiment of the present invention;
[0034] Figure 8 A schematic diagram of the loss and accuracy of a validation sample input model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.
[0038] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0039] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0040] With the rapid development of computer vision and deep learning technology and the continuous advancement of artificial intelligence technology, in the field of agricultural pest control, it provides strong technical and theoretical support for the use of computer vision image recognition technology to strengthen the identification of crop pests and diseases and the timely warning and prevention of pests and diseases. The pest and disease recognition method based on image processing has been applied to other agricultural pest and disease recognition fields, such as rice pest and disease recognition, corn pest and disease recognition, etc. Currently, the commonly used methods include DenseNet, YOLO, and ResNet convolutional neural network models. Among them, the ResNet network has higher accuracy and wider application due to its ultra-deep network structure, residual connection module and batch normalization to accelerate training.
[0041] The deepest ResNet network has more than 1,000 layers. Although the deepening of the network can reduce the error to a certain extent, it will bring a series of network degradation problems: 1. Gradient vanishing and gradient exploding. When the number of network layers increases, the gradient may gradually become smaller or larger, leading to the problem of gradient vanishing or gradient exploding. This problem makes it impossible for the network to update parameters effectively, resulting in a decline in model performance. 2. Overfitting. Deep networks have strong expressive power and are prone to overfitting on training data. When the number of network layers increases, the complexity of the model also increases, increasing the risk of overfitting and causing the model to perform poorly on the test set. 3. Lack of effective feature representation. As the number of network layers increases, the network pays more attention to the learning and expression of high-level features, while ignoring the importance of low-level features. This may cause the network to lose some effective feature representation capabilities, leading to degradation of model performance.
[0042] Figure 1 A flow chart of a method for identifying tomato leaf pests and diseases provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes:
[0043] Step 1: Select and preprocess sample images, 70% of which are used as training samples and 30% as verification samples.
[0044] In the embodiment of the present invention, step 1 includes:
[0045] The sample images come from the test materials, from which 8 common disease image data including tomato health are selected, as shown in Table 1; and the sample images are preprocessed and cropped to a size of 256×256.
[0046] Table 1. Number distribution of tomato dataset
[0047]
[0048] The experimental materials are taken from the dataset used in the 2018 Artificial Intelligence Challenge; Figure 2 As shown, the sample pictures mainly include common tomato disease categories, such as general tomato early blight, severe tomato early blight, healthy tomato, general tomato yellow leaf curl virus disease, severe tomato yellow leaf curl virus disease, general tomato red spider mites damage, general tomato leaf spot, and severe tomato leaf spot.
[0049] Step 2: Build a deep learning model based on the Pytorch framework.
[0050] In the embodiment of the present invention, step 2 includes:
[0051] The deep learning model includes a preprocessing module, a classification module, and a fusion module;
[0052] Since the ResNet network relies on increasing the depth or width of the network during optimization, such as Figure 3 As shown in the figure, however, as the number of network parameters increases, problems such as adjusting parameters, designing networks, and computing costs also increase; when there are too many hyperparameters, it is often difficult to ensure that all parameters are optimized to the best state. In addition, the network trained by the traditional model is highly dependent on hyperparameters, which requires frequent parameter adjustments on different data sets, which affects the scalability of the model. For this reason, the ResNeXt model is proposed based on ResNet, which does not increase or even reduces the complexity of the model while improving the accuracy, and reduces the number of hyperparameters. The method of adding group convolution provides a new idea for optimization.
[0053] The addition of the residual structure can effectively improve the network degradation problem. By making the output of each layer contain not only the feature representation of the current layer but also the information of the previous layer, it helps the network to better transfer gradients, alleviate the gradient disappearance and gradient explosion problems, and improve model performance.
[0054] Batch Normalization is used to adjust numerical data to a common scale without distorting its shape. The overall steps include: normalization, rescaling, and offset; Normalization is the process of converting data to a mean of 0 and a standard deviation of 1; In this step, there is a batch input from layer h. First, the mean of this hidden activation needs to be calculated, and its expression is:
[0055]
[0056] Where m is the number of neurons in layer h; the next step is to calculate the standard deviation of the hidden activations, which is expressed as:
[0057]
[0058] Then subtract the mean from each input and divide it by the sum of the standard deviation and the smoothing term ε. The smoothing term ε is a very small constant that prevents the denominator from being zero to ensure numerical stability in the operation.
[0059]
[0060] Rescaling and Shifting In the final operation, the input is rescaled and shifted; the rescaling parameter γ and the shift parameter β are expressed as:
[0061] h i =γ×h i(norm) +β;
[0062] After batch normalization, the output of each neuron follows the standard normal distribution of the entire batch, making the optimization smoother and the optimizer work easier. It will not be affected by gradient vanishing or gradient exploding when the learning rate becomes larger.
[0063] The ResNeXt model is proposed based on ResNet, which improves the accuracy while reducing the model complexity and the number of hyperparameters; and adds group convolution;
[0064] ResNeXt model architecture: The 3×3 convolution in ResNet is replaced by group convolution. A concept of Cardinality is added to the model, which is actually a convolution group. Compared with increasing the width and depth of the network, adding a group is better; Figure 4 As shown in the figure, the ResNeXt residual block first performs a grouping of 1×1 convolutions to Groups=32; although the dimension of the ResNeXt residual block after the first 1×1 convolution is 32×4=128, it is increased compared to the original residual block 64. Since grouped convolution can reduce the number of parameters, even if the dimension increases, the total number of parameters is reduced.
[0065] In order to improve the recognition accuracy, the ResNet and ResNeXt models are combined to obtain a multi-integrated deep model network. Figure 5 As shown, the recognition accuracy can be greatly improved.
[0066] Step 3: Train and optimize the constructed deep learning model through training samples, and verify the trained and optimized model through verification samples to identify tomato leaf diseases and pests.
[0067] In the embodiment of the present invention, since the data has been screened and is relatively balanced, in order to ensure efficient use of the data set, the image needs to be scaled and enhanced. The image is adjusted to 224 pixels × 224 pixels to meet the model requirements, and then the data is enhanced by horizontal flipping, vertical flipping, multi-angle rotation, etc.
[0068] Then, experiments and parameter adjustments were carried out in the built environment. In this application, the experiments were run in a physical machine with a CPU of i7-13700KF, a GPU of RTX4060Ti, and a memory of 16G, using the Windows 11 operating system. In terms of the deep learning environment, Pytorch 2.1.1, CUDA 12.1, and torchvision 0.16.1 were used.
[0069] The present invention pre-processes tomato sample images and then inputs them into a deep learning model for deep learning, obtains the best model by adjusting parameters, and then completes artificial intelligence recognition of pests and diseases in tomato images by calling the model.
[0070] The specific steps are as follows:
[0071] (1) Preprocessing of sample images:
[0072] You need to crop the sample to 256×256, traverse the images in the original folder, and save the resized images to a new folder, such as Figure 6 As shown in the figure, the sample image processing results. (2) Transformation of input data:
[0073] According to the Pytorch framework, the sample data is input in a standard form, and pandas is used to read the csv file, remove the header part, calculate the length, and put the header into a list; read the first column, which contains the name of the image file, such as images / 0.jpg; the second column is the leaf type label of the image; the validation set can prevent overfitting and observe the fitting results; image augmentation, random horizontal flipping, random cropping, modification of brightness, contrast, saturation, random rotation of a certain angle, and channel standardization, that is, first subtract the mean and then divide by the standard deviation; the test set only needs to return the image, get the string label of the image, consult the dictionary, convert the type to a number, return the image data and the corresponding label corresponding to each index, and return the number of training / validation / test / images.
[0074] (3) Model training:
[0075] Perform deep learning training and parameter adjustment on the models ResNeXt18 and ResNeXt50, calculate custom cropping areas, normalize, train functions, and use cross entropy as a performance metric for classification tasks; define learning rate decay and initialize the optimizer. You can fine-tune some hyperparameters by yourself, such as learning rate. Adam is used here; ensure that the model is in training mode before training, record training information, iterate the training set in batches, and the batch consists of image data and corresponding labels. Move the data to the GPU, crop the image CUTMIX training code, generate random cropping weights, shuffle the samples to generate spliced samples, generate cropping areas, change the bbx1:bbx2,bby1:bby2 areas in the original samples to the areas corresponding to the random sample labels, recalculate lambda to accurately match the pixel ratio, so It is possible to crop beyond the boundary; bring the graphic data into the model to calculate the prediction, calculate the cross entropy loss, and note that the loss of the two samples is weighted summed according to the segmentation ratio; clear the gradient stored in the parameters in the previous step, calculate the gradient of the parameters, update the parameters with the calculated gradient, calculate the accuracy of the current batch, record the loss and accuracy, and the average loss and accuracy of the training set are the average of the recorded values; update the learning rate; during verification, performing corresponding transformations on the image may be able to extract more features, ensure that the model is in eval mode so that modules such as dropout can be disabled and work properly, iterate the validation set batch by batch, calculate the loss, calculate the accuracy of the current batch, record the loss and accuracy, and the average loss and accuracy of the entire validation set are the average of the recorded values; record the model, and if it is improved, save a checkpoint at this point in time, such as Figure 7 As shown, a diagram of training loss and accuracy.
[0076] (4) Adjustment and determination of hyperparameters:
[0077] The final parameters are determined by adjusting different hyperparameters, including batch size, learning rate, weight decay, and the number of iterations during training.
[0078] (5) Model accuracy verification:
[0079] like Figure 8 As shown, the loss and accuracy of the validation sample input model.
[0080] (6) Image recognition:
[0081] Load the model, add tomato images for recognition, inherit the pytorch dataset, create Data, get the file name corresponding to the index from image_arr, read the image file, image processing, the test set only needs to return the image, data transformation; model prediction, load model 1, load model 2, load TTA, initialize the list for storing predictions, identify probabilities, iterate the test set batch by batch, take the largest logit class as the prediction, and record it, and restore the predicted numerical category to the category name.
[0082] Prediction results and probabilities:
[0083] Prediction category: 'Tomato leaf spot - general',
[0084] Prediction probability: 98.63%.
[0085] The present invention can not only detect tomato diseases and pests in a timely and accurate manner and reduce agricultural losses through fine-grained identification of tomato diseases and pests, but also facilitate accurate use of drugs during treatment, improve food safety and reduce environmental pollution. The multiple models of deep learning can be used in combination, which can not only increase recognition accuracy, but also greatly improve the generalization ability of the model, and can be directly applied to the recognition of other images, which can provide a direction for the research of deep learning.
[0086] In the technical solution provided by the present invention, the method includes selecting and preprocessing sample images; constructing a deep learning model based on the Pytorch framework; training and optimizing the constructed deep learning model through training samples, and verifying the trained and optimized model through verification samples to identify tomato leaf diseases and pests. This method improves the yield and quality of tomato crops and reduces the economic losses caused by tomato diseases and pests.
[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying tomato leaf pests and diseases, characterized in that: The method comprises: Step 1: Select and preprocess sample images, 70% of which are used as training samples and 30% as verification samples; Step 2: Build a deep learning model based on the Pytorch framework; Step 3: Train and optimize the constructed deep learning model through training samples, and verify the trained and optimized model through verification samples to identify tomato leaf diseases and pests.
2. The method according to claim 1, characterized in that The step 1 comprises: The sample images come from experimental materials, from which image data of 8 common diseases including tomato health are selected; the sample images are preprocessed and cropped to a size of 256×256.
3. The method according to claim 1, characterized in that The step 2 comprises: The deep learning model includes a preprocessing module, a classification module, and a fusion module; The addition of the residual structure can effectively improve the network degradation problem. By making the output of each layer contain not only the feature representation of the current layer but also the information of the previous layer, it helps the network to better transfer gradients, alleviate the gradient disappearance and gradient explosion problems, and improve model performance. Batch Normalization is used to adjust numerical data to a common scale without distorting its shape. The overall steps include: normalization, rescaling, and offset; Normalization is the process of converting data to a mean of 0 and a standard deviation of 1; In this step, there is a batch input from layer h. First, the mean of this hidden activation needs to be calculated, and its expression is: Where m is the number of neurons in layer h; the next step is to calculate the standard deviation of the hidden activations, which is expressed as: Then subtract the mean from each input and divide it by the sum of the standard deviation and the smoothing term ε. The smoothing term ε is a very small constant that prevents the denominator from being zero to ensure numerical stability in the operation. Rescaling and Shifting In the final operation, the input is rescaled and shifted; the rescaling parameter γ and the shift parameter β are expressed as: h i =γ×h i(norm) +b; After batch normalization, the output of each neuron follows the standard normal distribution of the entire batch, making the optimization smoother and the optimizer work easier. It will not be affected by gradient vanishing or gradient exploding when the learning rate becomes larger. The ResNeXt model is proposed based on ResNet, which improves the accuracy while reducing the model complexity and the number of hyperparameters; and adds group convolution; ResNeXt model architecture: The 3×3 convolution in ResNet is replaced by group convolution. A concept of Cardinality is added to the model, which is actually a convolution group. Compared with increasing the width and depth of the network, adding groups is better. The ResNeXt residual block first groups the 1×1 convolution into Groups=32. In order to improve the recognition accuracy, a multi-integrated deep model network is obtained by combining the ResNet and ResNeXt models.
Citation Information
Patent Citations
Tomato disease image recognition method
CN113158754A
Leaf disease identification method and device, medium and product
CN117975273A
Method and system for detecting scene text
US20220207890A1