Plant disease identification method, equipment and medium
By optimizing ResNet50 model and data augmentation technology, the problems of insufficient training data and insufficient computing resources in plant disease recognition are solved, and efficient and accurate disease identification and real-time prevention and control guidance are achieved, which is suitable for edge devices.
Patent Information
- Application Number
- CN202510506841.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
The existing plant disease recognition methods have problems such as insufficient training data, limited recognition accuracy, and large model calculation volume and difficulty in running efficiently on mobile or edge devices, resulting in low recognition accuracy and poor practicality and real-time performance.
The optimized ResNet50 model is adopted, combined with transfer learning, automatic mixed accuracy strategy, momentum stochastic gradient descent optimization algorithm and weight decay, and trained through the StepLR learning rate scheduling algorithm, and image preprocessing and data enhancement are added to construct plant disease heat maps to improve recognition accuracy and model adaptability.
The accuracy of plant disease identification is improved to 98.5%, the use of computing resources is reduced, the applicability and interpretability of the model on edge devices is enhanced, and the rapid and accurate disease identification and prevention and control guidance is achieved.
Smart Images

Figure CN120375084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular, to a method, device, and medium for plant disease recognition. Background Art
[0002] In the modern agricultural production system, the accurate recognition and scientific management of plant diseases have become the key links determining crop yield and quality. With the continuous improvement of the scale and intensification of agriculture, the expansion of the planting area of single crops and the frequent cross-regional agricultural trade, the occurrence of plant diseases has shown significant characteristics of complex types, rapid spread, and increased harm. As the primary link in the prevention and control system, disease recognition is undergoing a technological upgrade from traditional manual inspections to intelligent monitoring.
[0003] Traditional disease recognition methods mainly rely on the empirical judgment of agronomy experts. This method is not only time-consuming and laborious but also easily affected by subjective factors, resulting in inaccurate recognition results. In recent years, the application of computer vision technology in the agricultural field has gradually increased, and among them, image recognition technology based on deep learning has been widely used in plant disease detection.
[0004] However, the existing plant disease recognition methods still have the following deficiencies: Insufficient training data leads to poor generalization ability of the model. When the training sample size is less than 1000 cases, the test accuracy of mainstream models such as ResNet-50 will decrease by 15%-20% compared with the scenario of sufficient data, and the recognition accuracy of early diseases with lesion spots less than 5mm is less than 40%; The recognition accuracy is limited and is easily affected by external environmental factors such as light and background. Uneven light increases the blurriness of lesion edges by 30%-50%, the spectral reflectance of leaves fluctuates by more than 20% in low-light environments, and complex backgrounds lead to a decrease in the accuracy of target crop area segmentation by more than 40%; In addition, some models have a large amount of calculation and are difficult to operate efficiently on mobile devices or edge devices, resulting in the difficulty of popularizing the existing deep learning-based plant disease and pest recognition methods, with low practicality and real-time performance. Summary of the Invention
[0005] To solve the above problems, the present invention provides a plant disease recognition method applied to the server side, including the following steps: Obtain an image uploaded by the client; Input the image into a pre-trained optimized ResNet50 model to obtain a prediction result, where the prediction result includes various possible plant diseases analyzed by the model and the probability of each plant disease; Perform visual analysis on the prediction result to obtain a plant disease heat map, increasing interpretability, and enabling users to quickly lock the disease area through the heat map; Transmit the plant disease heatmap and prediction results to the client.
[0006] Based on the above solution, optimizing the training process of the ResNet50 model includes: Collect images of diseased and healthy plant leaves, preprocess the collected images, and construct a disease recognition dataset. Normalization is to ensure the consistent distribution of all input data and avoid the impact of scale differences in different input data on the model training process. Construct an optimized ResNet50 model based on transfer learning. Input the disease recognition dataset into the optimized ResNet50 model for pre-training, and process it in combination with the automatic mixed precision strategy, momentum stochastic gradient descent optimization algorithm, and weight decay. The automatic mixed precision combines 16-bit and 32-bit floating-point precisions to dynamically adjust the gradient, making the calculation more efficient while reducing GPU memory occupancy. At the same time, the momentum stochastic gradient descent optimization algorithm and weight decay are also used jointly for parameter optimization. During the training process, adjust the learning rate of the ResNet50 model through the StepLR learning rate scheduling algorithm to improve the generalization ability of the model.
[0007] When constructing the optimized ResNet50 model based on transfer learning, perform operations on the ResNet50 model including freezing the first 4 convolutional blocks and replacing the fully connected layer structure, and adjust the model to suit the output structure for plant disease recognition.
[0008] Based on the above solution, the StepLR learning rate scheduling algorithm is also added during the training process of optimizing the ResNet50 model. After every 10 rounds of training, the learning rate is decayed to 0.1 times the original. This strategy can converge quickly in the initial stage of training, optimize stably in the later stage of training, and avoid gradient oscillation. Compared with the fixed learning rate model, the StepLR dynamic learning rate model has a faster convergence speed.
[0009] Based on the above solution, the preprocessing includes resizing, normalizing, and data augmentation of the images. Among them, resizing and normalizing the images are the standardization processes of the images to ensure the consistent distribution of all input image data and avoid the impact of scale differences in different input data on the training process.
[0010] Based on the above solution, data augmentation includes random horizontal flipping, random cropping, random scaling, and random rotation. Data augmentation is to increase the diversity of data and the robustness of the model. Among them, random horizontal flipping can enhance the model's adaptability to direction changes, random cropping can enhance the model's learning of local features, and random scaling and rotation can enhance the model's recognition ability for different scales and angle changes. These data augmentation methods can effectively increase the diversity of training data, thereby improving the generalization ability of the model and reducing overfitting.
[0011] The second aspect of the present invention provides a plant disease identification device, which is applied to the server side and includes: A device for acquiring the images uploaded by the client; A device for inputting the images into a pre-trained optimized ResNet50 model to obtain a prediction result; A device for visually analyzing the prediction distribution result to obtain a plant disease heat map; A device for transmitting the plant disease heat map and the prediction result to the client.
[0012] The third aspect of the present invention provides a computer-readable storage medium, in which a runnable computer program is stored. When the program runs, it can execute the above-mentioned plant disease identification method.
[0013] The fourth aspect of the present invention provides a plant disease identification method, which is applied to the client side and includes the following steps: Taking a plant image and uploading the plant image to the server side; Obtaining the plant disease heat map and the prediction result returned by the server to highlight the disease area concerned by the model. A control is set in the plant disease area of the plant disease heat map; When it is detected that the control is triggered, jumping to the disease prevention and control interface corresponding to the prediction result, which increases the usability and practicality of the present invention and can provide a solution on-site after shooting.
[0014] The fifth aspect of the present invention provides a plant disease identification device, which is applied to the client side and includes: A device for taking a plant image and uploading it to the server side; A device for obtaining the plant disease heat map and the prediction result returned by the server. A control is set in the plant disease area of the plant disease heat map; A device for jumping to the disease prevention and control page corresponding to the prediction result item when it is detected that the control is triggered.
[0015] The advantages of the present invention are as follows: 1. Breaking through the limitations of traditional manual inspections and existing deep learning-based methods, solving problems such as insufficient training data, limited recognition accuracy, and large model computational complexity that are difficult to operate efficiently on mobile devices or edge devices; 2. Combining transfer learning with the ResNet50 model, by loading existing pre-trained weights, only the top classification layer needs to be fine-tuned to achieve a good recognition and analysis effect, reducing the parameter training requirements, and at the same time greatly shortening the training cycle and total training time; 3. During model training, automatic mixed precision training is combined to significantly reduce the use of computing resources while maintaining floating-point precision, and further accelerate the model training speed under the same hardware configuration; 4. By controlling the growth of model parameters through weight decay to prevent overfitting, and introducing "inertia" in the gradient direction update by the momentum mechanism to avoid oscillations. The combination of the two improves the stability of training, makes the model have stronger generalization ability, and improves the accuracy of the final test set; 5. The model network is lightweight. The residual structure based on ResNet50 is adopted to avoid the problem of gradient disappearance and improve the efficiency of deep network training. The overall number of parameters is small, which is suitable for deployment in medium and small-sized GPUs or edge devices and has low requirements for the hardware environment; 6. The visualization heat map display is added. Different from the existing "black box" models, it can more intuitively provide the key points of model analysis for users, increase interpretability, enable users to quickly lock the disease area, and increase the practicability of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of an implementation environment related to an embodiment of the present invention.
[0017] Figure 2 It is a schematic diagram of the overall interaction of the present invention.
[0018] Figure 3 It is a schematic diagram of the training process of the optimized ResNet50 model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following further describes the solution in conjunction with the specific embodiments and the accompanying drawings of the specification.
[0020] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment related to an embodiment of the present invention. The implementation environment includes: an edge device 100, a camera device 101, and an identification object 102.
[0021] The edge device 100 is used to pre-train the optimized ResNet50 model of the present invention and run the pre-trained optimized ResNet50 model. The edge device 100 has low requirements for hardware computing power. This embodiment is carried out on a device installed with RTX 3060. The whole process from input image preprocessing to GPU loading is optimized, and the average inference time per unit image is less than 40 ms / image, meeting the requirements of real-time recognition. The edge device 100 is also used to receive the pictures sent by the camera device 101, and the receiving methods include wireless and wired connections.
[0022] The imaging device 101 is used to capture pictures of the recognition object 102 and send the captured pictures to the edge device 100 for analysis and processing. There is no clear division of labor between the imaging device 101 and the edge device 100, that is, the function of displaying the analysis and processing results can be flexibly allocated between the two.
[0023] In this embodiment, pictures are captured by the imaging device 101 and the analysis and processing results sent by the edge device 100 are received. Visual display and disease control method query are performed on the imaging device 101. This method is more in line with modern people's usage habits: using a mobile phone to capture pictures and search for results. This solution supports users to customize data sets and train models that meet their own needs. Compared with existing general recognition models, the recognition method and recognition model provided by this solution are more professional; and the deployment, training, and usage processes are all completed locally, with higher data security.
[0024] In other embodiments, the imaging device 101 is only responsible for image capture, sends the captured images to the edge device 100, and the edge device 100 analyzes and processes them, and directly outputs and displays the analysis and processing results and disease control queries on the edge device 100.
[0025] In other embodiments, the imaging device 101 is integrated on the edge device 100, and image capture, image processing, output of processing results, and query of disease control methods are all concentrated on the edge device 100.
[0026] The recognition object 102 is the plant object recognized by this solution, including various crops and fruit trees. This solution supports the analysis and recognition of 38 types of plant diseases, with a recognition accuracy rate as high as 98.5%.
[0027] Please refer to Figure 2 , which shows the overall interaction schematic diagram of the present invention. The server side and the client side here are for facilitating the understanding of the distinction between the pre-trained model data processing process and the display process of the data processing results of this solution. The server side represents the processing process of the optimized ResNet50 model pre-trained by this solution, and this process is invisible to users. The client side represents the operations visible to users of this solution, such as capturing images, receiving processing results, and jumping to the disease prevention page. The client side and the server side can be located on the same physical device.
[0028] The client captures an image of the plant to be recognized and sends the captured image to the server. The server is equipped with the optimized ResNet50 model pre-trained by this solution. The captured image is analyzed and recognized by the ResNet50 model to obtain a prediction result. The prediction result includes the key recognition area and the recognition result. A heat map is generated by analyzing the prediction result to show the recognition basis to the user, improve the credibility of the model, and facilitate the user's understanding and use. Finally, the server sends the heat map and the recognition result to the client.
[0029] In this embodiment, Grad-CAM is used to perform visual analysis on the prediction result, and the formula is as follows:
[0030] Where is the score of the predicted class, is the gradient weight of the th feature map, is the activation value of the th feature map. The heat map calculated by Grad-CAM is superimposed on the original image, which can highlight the disease area that the model focuses on and help the user accurately locate the disease position.
[0031] While the client displays the recognition result, it provides a query interface for the user. By clicking on the corresponding area of the output image, the user can automatically jump to the relevant disease prevention and control knowledge page. The disease prevention and control knowledge page can also be the user-defined prevention and control strategy. If the user does not have a user-defined prevention and control strategy, it will jump to the browser to automatically search and display the corresponding disease knowledge. The user can quickly consult the disease characteristics and prevention methods. Providing a user-defined prevention and control strategy facilitates the expansion of functions and user use, and enhances the application of this solution in multiple scenarios.
[0032] Please refer to Figure 3 , which shows a schematic diagram of the training process of the optimized ResNet50 model of the present invention, including the following steps: Step S201, collect plant disease leaf images and healthy leaf images, and preprocess the collected images to construct a disease recognition data set; The preprocessing includes resizing, normalizing, and data augmentation of the images. The resizing and normalizing of the images are the standardization of the images, and the formula is as follows:
[0033] Where and are the mean and standard deviation of the training data respectively, is the standardized image. This process ensures that the distributions of all input data are consistent and avoids affecting the overall training process due to the scale differences of different input data. Data augmentation is to improve the robustness of the model. In this embodiment, a multi-level data augmentation strategy is applied to the dataset, including: random horizontal flipping, random cropping, and random scaling and rotation.
[0034] Random horizontal flipping increases the diversity of data through random horizontal flipping, enhancing the model's adaptability to direction changes.
[0035] Random cropping enhances the model's learning of local features by randomly cropping the image. The formula is as follows:
[0036] where, is the cropped image, is the image cropping method, is the original image, is the random cropping size.
[0037] Random scaling and rotation enhance the model's recognition ability for different scales and angle changes by scaling and rotating the image. The random scaling formula is as follows:
[0038] where, is the scaled image, is the image scaling method, is the random scaling ratio.
[0039] The random rotation formula is as follows:
[0040] where, is the rotated image, is the image rotation method, is the random rotation angle.
[0041] Finally, the images are uniformly scaled to the RGB format of 224×224×3 using bicubic interpolation.
[0042] The above data augmentation method can effectively increase the diversity of training data, thereby enhancing the model's generalization ability and reducing overfitting.
[0043] In other embodiments, the brightness and saturation of the pictures are also randomly adjusted to improve the model's robustness to illumination changes.
[0044] In other embodiments, Gaussian noise is added to the pictures to improve the model's anti-interference ability.
[0045] Step S202: Construct an optimized ResNet50 model based on transfer learning, and input the disease recognition dataset into the optimized ResNet50 model for training, combined with automatic mixed precision (AMP) training, momentum stochastic gradient descent optimization algorithm, and weight decay processing.
[0046] After preprocessing and data augmentation, the disease recognition dataset is input into the model to start the training process. In the process, the automatic mixed precision (AMP) technology is first used. This technology combines 16-bit and 32-bit floating-point precisions, and by dynamically adjusting the gradient, it reduces the GPU memory occupancy. This is equivalent to increasing the data throughput of the model per unit time, making the calculation more efficient. The formula for dynamically adjusting the gradient is as follows:
[0047] where, is the learning rate, is the current model parameter, and the loss function is , represents the gradient of the loss function, is the scaling factor used to dynamically adjust the gradient. Through the above AMP technology, while maintaining the accuracy during the model training process, the calculation efficiency is significantly improved.
[0048] To achieve more efficient gradient updates and stronger model generalization ability, this solution adopts the momentum stochastic gradient descent optimization algorithm in the loss function optimization stage, integrating the weight decay mechanism to prevent the model from overfitting during the training process. In this embodiment, the weight decay mechanism is implemented through L2 regularization, and the formula is as follows:
[0049] where, is the loss function, is the weight decay coefficient, is the squared norm of the parameter vector; Then calculate the gradient of the integrated regularization term, and the formula is as follows:
[0050] This gradient is applied to the momentum gradient update rule and can be combined with the historical gradient accumulation term to form a smooth direction update. The momentum gradient update formula is as follows:
[0051]
[0052] where, is the current momentum vector, is the momentum factor, which is set to 0.9 in this embodiment, is the learning rate, is the gradient of the loss function, is the L2 weight decay term. Weight decay and momentum optimization are additively fused in the above gradient update, that is, the weight decay term directly participates in the gradient derivation process and is executed synchronously with the momentum formula, rather than being independent operations before and after. This strategy has the following advantages: Control the growth of model parameters through weight decay to prevent overfitting; The momentum mechanism introduces "inertia" in the gradient direction update to avoid oscillations; The combination of the two can improve the stability of training while maintaining the generalization ability.
[0053] At the same time, since Automatic Mixed Precision (AMP) can automatically manage gradient scaling and precision switching, in the above gradient calculation, scaling processing will be synchronously performed through GradScaler to ensure numerical stability. After gradient scaling, the regularization term still participates, and it is ensured to cooperate with the momentum mechanism.
[0054] Step S203, during the training process, update the learning rate through the StepLR learning rate scheduling algorithm. In this embodiment, every 10 rounds of training, the learning rate is decayed to 0.1 times the original. The formula is as follows:
[0055] where, is the initial learning rate, is the decay factor, is the current training step, is the learning rate of the current step, is the number of training rounds, represents rounding down.
[0056] Step S204, train until the model converges and the amplitude is small to obtain the optimized ResNet50 model with pre-training completed. In this embodiment, the model training can be completed through 30 - 50 rounds.
[0057] The comparison between this solution and the traditional convolutional neural network is as follows in the table:
[0058] Compared with the traditional convolutional neural network, the optimized ResNet50 model provided by this solution has fewer training rounds, shorter training time, and faster convergence speed. Existing convolutional neural networks require high-precision floating-point operations during training. Insufficient floating-point operation precision will affect the model training speed, precision, and generalization ability, while too high floating-point precision also requires a higher GPU. This solution adopts automatic mixed precision to reasonably allocate the floating-point operation precision of 16 bits and 32 bits, obtaining a model with less GPU memory occupancy and more efficient calculation.
[0059] Existing large-scale convolutional neural networks (such as ResNet) usually require more than 16G of high memory. When applied to edge devices, the model needs to be further compressed by methods such as pruning and quantization. Processing the model through pruning will significantly reduce the accuracy and generalization ability of the model. At present, some pruning strategies have less impact on the accuracy and generalization ability of the model, but they require higher requirements for operators and it is also difficult to transplant the model to edge devices. Processing the model through quantization will reduce the accuracy and stability of the model. The model provided by this application has high accuracy and good stability, and can be adapted to edge devices at the same time.
[0060] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
[0061] Although the specific implementation manners of the present invention have been described above, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.
Claims
1. A method for identifying plant diseases, characterized in that, Applied to the server side, it includes the following steps: Obtain the images uploaded by the client; Input the images into the pre-trained optimized ResNet50 model to obtain prediction results; Conduct visual analysis on the prediction results to obtain a plant disease heat map; Transmit the plant disease heat map and the prediction results to the client.
2. The plant disease recognition method according to claim 1, characterized in that, The training process of the optimized ResNet50 model includes: Collect plant disease leaf images and healthy leaf images, preprocess the collected images, and construct a disease recognition dataset; Construct an optimized ResNet50 model based on transfer learning, input the disease recognition dataset into the optimized ResNet50 model for pre-training, and process it in combination with the automatic mixed precision strategy, momentum stochastic gradient descent optimization algorithm, and weight decay; During the training process, adjust the learning rate of the optimized ResNet50 model through the StepLR learning rate scheduling algorithm; Train until the optimized ResNet50 model converges and stabilizes to obtain the pre-trained optimized ResNet50 model.
3. The plant disease recognition method according to claim 2, characterized in that, During the training process, the StepLR learning rate scheduling algorithm is also added. After every 10 rounds of training, the learning rate is decayed to 0.1 times the original.
4. The method for identifying plant diseases according to claim 2, characterized in that The preprocessing includes resizing, normalizing, and data augmentation of the images.
5. The plant disease recognition method according to claim 4, characterized in that, The data augmentation includes random horizontal flipping, random cropping, random scaling, and random rotation.
6. A plant disease identification device, characterized in that, Applied to the server side, it includes: A device for obtaining the images uploaded by the client; A device for inputting the images into the pre-trained optimized ResNet50 model to obtain prediction results; A device for conducting visual analysis on the prediction distribution results to obtain a plant disease heat map; A device for transmitting the plant disease heat map and the prediction results to the client.
7. The plant disease identification device according to claim 6, wherein, The training process of the optimized ResNet50 model includes: Collect plant disease leaf images and healthy leaf images, preprocess the collected images, and construct a disease recognition dataset; Construct an optimized ResNet50 model based on transfer learning, input the disease recognition dataset into the optimized ResNet50 model for pre-training, and process it in combination with the automatic mixed precision strategy, momentum stochastic gradient descent optimization algorithm, and weight decay; During the training process, adjust the learning rate of the optimized ResNet50 model through the StepLR learning rate scheduling algorithm; Train until the optimized ResNet50 model converges and stabilizes to obtain the pre-trained optimized ResNet50 model.
8. A computer-readable storage medium storing a computer program executable thereon, characterized in that, When the computer program runs, it executes a plant disease recognition method as described in any one of claims 1 to 5.
9. A method for identifying plant diseases, characterized in that, Applied to the client side, it includes the following steps: Take a plant image and upload it to the server side; Obtain the plant disease heat map and prediction results returned by the server side, and a control is set in the plant disease area of the plant disease heat map; When it is detected that the control is triggered, jump to the disease prevention and control page corresponding to the prediction result.
10. A plant disease identification device, characterized in that, Applied to the client side, it includes: A device for taking a plant image and uploading it to the server side; A device for obtaining the plant disease heat map and prediction results returned by the server side, and a control is set in the plant disease area of the plant disease heat map; A device for jumping to the disease prevention and control page corresponding to the prediction result when the control is detected to be triggered.