Plant leaf disease and pest detection method and system suitable for edge calculation

By constructing a lightweight Mamba model and combining data augmentation and optimization techniques, the difficulties in deploying on edge computing platforms caused by the time-consuming and large model size of traditional methods are solved, enabling fast and accurate detection of plant leaf diseases and pests, which is suitable for edge computing platforms.

CN121600320APending Publication Date: 2026-03-03ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511804851.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing methods for detecting plant leaf diseases and pests rely on manual inspection, which is time-consuming and susceptible to human error. Furthermore, the models based on YOLOv8 and convolutional neural networks are too large to be suitable for deployment on edge computing platforms.

Method used

A lightweight plant leaf disease and pest classification model based on Mamba was constructed. The image dataset was processed through data augmentation and normalization. The model was optimized by combining spatial attention mechanism and Mamba algorithm, including operator fusion, selective low-rank decomposition and pruning, to adapt to edge computing platforms.

Benefits of technology

It enables fast and accurate pest and disease detection on edge computing platforms, reduces model file size, improves detection speed and resource utilization efficiency, and is suitable for resource-constrained embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600320A_ABST
    Figure CN121600320A_ABST
Patent Text Reader

Abstract

The invention discloses a plant leaf disease and insect pest detection method and system suitable for edge calculation, and relates to the technical field of plant leaf disease and insect pest detection. The method comprises the following steps: collecting and sorting plant leaf pest and disease damage and stressed leaf images, and constructing a basic image data set for subsequent model training and evaluation; preprocessing the basic image data set through data enhancement and standardization, and constructing an image for model training; constructing a Mama-based plant leaf disease and insect pest classification model; by optimizing an edge computing platform, lightweight design is carried out on the model; performing model training to obtain a pre-training model; and detecting plant leaf diseases and insect pests by using the trained plant leaf disease and insect pest classification model. The method can effectively reduce the complexity of the model and improve the detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of plant leaf disease and pest detection technology, and in particular to a method and system for detecting plant leaf diseases and pests suitable for edge computing. Background Technology

[0002] Traditional plant leaf disease and pest detection largely relies on manual inspection. Agricultural experts visually examine leaf surfaces or collect field samples, identifying and controlling diseases based on morphological characteristics. This method is labor-intensive, time-consuming, and susceptible to human error and bias, making it unsuitable for the needs of modern precision agriculture. Furthermore, visual assessment is inherently limited by subjective perception, potentially leading to inconsistent diagnoses. Therefore, there is an urgent need for large-scale, objective, automated, and non-destructive detection technologies for reliable disease management. Currently, neural network-based screening is an effective and efficient detection method. Data acquisition equipment can be carried out on unmanned vehicles or drones to inspect planting areas, reducing intensive labor consumption and improving the level of modern precision agriculture technology.

[0003] Existing technologies 1, "A method for detecting grape leaf diseases based on YOLOv8", and 2, "A lightweight method and device for classifying grape leaf diseases based on convolutional neural networks", both use traditional methods to build models. One is YOLOv8, and the other is a convolutional neural network. However, the above two models have large structures, which will generate model files that occupy more memory, making them unsuitable for deployment on edge computing platforms. Summary of the Invention

[0004] The technical problem to be solved by the present invention is how to provide a plant leaf disease and pest detection method that is suitable for edge computing and can effectively improve the detection speed.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for detecting plant leaf diseases and pests suitable for edge computing, comprising the following steps: Collect and organize leaf images of plant diseases, pests, and stresses to construct a basic image dataset for subsequent model training and evaluation; The basic image dataset is preprocessed through data augmentation and normalization to construct images for model training; Construct a Mamba-based classification model for plant leaf diseases and pests; The model is designed to be lightweight by optimizing it for edge computing platforms; Perform model training to obtain a pre-trained model; The trained plant leaf disease and pest classification model was used to detect plant leaf diseases and pests.

[0006] The present invention also discloses a computer system, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned method for detecting plant leaf diseases and pests suitable for edge computing.

[0007] The beneficial effects of adopting the above technical solution are as follows: The method identifies diseased leaves through a Mamba-based pest and disease classification model, and then optimizes the model for edge computing platforms to facilitate deployment in the real world. It combines the spatial attention module and the Mamba model to enhance attention to different disease characteristics on the leaves through the attention mechanism, and then classifies diseases through the Mamba network with faster detection speed. The operator optimization fusion, pruning and selective low-rank decomposition techniques are arranged in an orderly manner to reduce the model file size and improve the detection speed without losing or with a small loss of model accuracy. Attached Figure Description

[0008] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0009] Figure 1 This is the main flowchart of the method described in Embodiment 1 of the present invention; Figure 2 This is a sample size distribution diagram of the grape leaf disease and pest dataset in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of data augmentation in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the principle of the plant leaf disease and pest classification model in the method described in Embodiment 1 of the present invention; Figure 5 This is an optimization flowchart of the model described in the method of Embodiment 1 of the present invention; Figure 6 This is a schematic diagram of the computer system described in Embodiment 2 of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0011] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0012] Example 1: Overall, such as Figure 1 As shown in the figure, this invention discloses a method for detecting plant leaf diseases and pests suitable for edge computing, the method comprising the following steps: S1: Collect and organize leaf images of plant diseases, pests, and stresses to build a basic image dataset for subsequent model training and evaluation; S2: Preprocess the basic image dataset through data augmentation and normalization to construct images for model training; S3: Construct a Mamba-based classification model for plant leaf diseases and pests; S4: The model is designed to be lightweight by optimizing it for edge computing platforms; S5: Perform model training to obtain a pre-trained model; S6: Use the trained plant leaf disease and pest classification model to detect plant leaf diseases and pests.

[0013] The above steps will be explained in detail below with reference to specific content: Constructing a Pest and Disease Dataset (Taking Wine Grape Pests and Diseases as an Example): With the assistance of plant pathology experts, images were collected and a grape leaf pest and disease dataset was constructed, containing 13,525 images across 11 categories. Besides samples of black rot, leaf blight, and Escafé auxiliaries from the Plant Village dataset, over two-thirds of the remaining images were collected using a mobile phone camera and a Sony A6000 digital camera in various vineyards. Images of three types of diseased leaves—nutrient deficiency, downy mildew, and brown spot—were collected from table grape vineyards in Hangzhou, China. Three diseases—downy mildew, powdery mildew, and viral diseases—and two pests—grape gall mite infestation and grape leafhopper—were collected from wine grape vineyards in Yinchuan, China, along with leaf samples from one type of healthy wine grape. Some images contain only one leaf target, some two or three, and some even more than ten. To standardize size and resolution, each leaf was individually labeled and extracted, and then each sample was resized to 640x640 pixels. Apart from black rot, leaf blight, and Escafé salina from the Plant Village dataset, the other categories of photographs were taken in the field. These field images, with their diverse lighting conditions and complex backgrounds, enhance their practicality and robustness. Figure 2This dataset contains the number of samples of various grape diseases and pests. It can be observed that, except for the nutrient deficiency category, the number of samples in other categories is relatively even, making it a high-quality dataset.

[0014] Data Augmentation and Standardization: To overcome the problem of insufficient nutrient-deficient samples in the dataset, data augmentation (random horizontal flip, random vertical flip, random rotation, random cropping with padding, and color jitter) is prioritized for the nutrient-deficient classes before overall data augmentation. This supplements these samples and addresses the impact of data imbalance on poor training accuracy. Then, a series of data augmentations are performed on the images in the training set, using the same five methods: random horizontal flip, random vertical flip, random rotation, random cropping with padding, and color jitter. These augmentation methods primarily simulate image distortion caused by camera movement and lighting changes, allowing the training process to generate new training samples through random transformations without increasing the original dataset size. This improves the model's generalization ability and prevents overfitting. Finally, the data is standardized by normalizing the pixel values ​​of each channel of the image according to the mean and standard deviation of the ImageNet dataset, improving the stability and convergence speed of model training. Figure 3 The image shown is an illustration of partial image enhancement.

[0015] A leaf disease and pest classification model based on Mamba: Its core logic is to enhance the image representation features of diseases and pests through a spatial attention mechanism, and then combine the selective state space of the Mamba algorithm with the synergistic effect of a hardware-aware parallel algorithm to achieve fast and accurate disease and pest classification. This module consists of four main parts: a convolutional feature extraction module, a spatial attention module, a Mamba core module, and a classifier, such as... Figure 4 As shown.

[0016] First, convolutional feature extraction is performed: the augmented and standardized image data is resampled to 64x64 pixels and then input into a module with four stacked convolutional feature extraction layers. Each feature extraction operation is identical, and the layer order for each operation is: 1) 2D fully connected convolutional layer; 2) 2D batch normalization layer; 3) SiLU activation function layer; 4) 2D fully connected convolutional layer; 5) 2D batch normalization layer; 6) SiLU activation function layer. This process progressively extracts and combines discriminative features from the input image, ranging from low-level details to high-level semantics. Then, a spatial attention module is used. This module takes the high-order feature maps obtained after multi-layer convolutional feature extraction, performs max pooling and average pooling on them, and then concatenates them. The concatenated feature map is then convolved and activated sequentially, and finally concatenated with the input of the spatial attention module before output. This allows the network to adaptively learn and strengthen important spatial regions in the feature maps while suppressing unimportant background or noise regions.

[0017] Mamba Core Module: Since Mamba is a neural network designed for sequence data, it needs to convert feature map data into data usable by the network. First, a feature transformation operation is performed, flattening the spatial feature maps into a sequence, making the convolutional feature map data processable by Mamba. The transformed sequence data is first input into the first layer of the Mamba block for initial feature learning and sequence dependency modeling. Then, layer normalization is used to stabilize and standardize the feature distribution. The normalized features are then input into the second layer of the Mamba block for deeper sequence feature extraction and secondary training. This cascaded design achieves hierarchical sequence modeling, enhancing the model's ability to capture data features.

[0018] The classifier is responsible for mapping high-level features to the class space, and its design balances representational power and generalization performance. This module performs the following operations sequentially: Input features are first linearly transformed through a fully connected layer, then stabilized by a one-dimensional batch normalization layer, followed by the introduction of non-linearity through the SiLU activation function. Finally, a regularization layer (such as Dropout) is applied to prevent overfitting. This structure is repeated twice to achieve deep feature extraction. Output features are mapped to the dimension of the number of classes through the last fully connected layer and transformed into a probability distribution by the Softmax function. The model selects the class with the highest probability as the prediction result. This design ensures training stability through batch normalization and improves generalization ability through regularization, ultimately achieving a robust mapping from features to classification results.

[0019] Optimization for edge computing platforms: Designed for practical applications of leaf disease and pest detection methods. Its core logic involves a series of model compression techniques to significantly reduce model size and improve inference speed while maintaining model performance, which is crucial for deploying the model on resource-constrained embedded or mobile devices. This module mainly consists of three parts: operator optimization and fusion, low-rank decomposition, and pruning. Experiments show that operator optimization and fusion have the least impact on overall model performance and size, while pruning has the greatest impact. Therefore, the order of these three operations is as follows: Figure 5 As shown: Full-process operator optimization and fusion: First, scan each layer of the model to analyze the complete computational graph. Then, mark layers with continuous convolution, batch normalization, and activation function operations as layers suitable for operator fusion. Finally, fuse these layers into a single operator to reduce computational and memory access overhead. This step has minimal impact on model accuracy but lays the foundation for subsequent optimizations.

[0020] Selective low-rank decomposition: First, the weight matrix is ​​selected, choosing a matrix with a generally low weight value. This is because such matrices perform the same calculations as matrices with large weights but with less useful work, thus being relatively redundant. Next, the rank is selected. A higher rank has less impact on accuracy and model size, while a lower rank has a greater impact on both. After matrix decomposition, the model is reconstructed based on the selected weight matrix and rank. The reconstructed model is then fine-tuned, evaluated, and iterated. The results of low-rank decomposition at different levels are dynamically analyzed, and accuracy loss thresholds and model size thresholds are set to find the most suitable low-rank decomposition structure for the current model.

[0021] Pruning: First, the weight contribution of each layer of the model is calculated. Then, based on the importance score of the weights, unimportant weights or structural units (such as neurons) are removed to generate a sparse model and perform unstructured pruning. Finally, fine-tuning is performed based on the model results. Since pruning has the greatest impact on the model structure and may cause a significant decrease in accuracy, it is placed as the last step so that it can be performed on the basis of the first two optimization steps and leave room for fine-tuning to restore accuracy.

[0022] Eleven types of grape leaf diseases and pests were trained using a classification model trained on a Mamba-based grape disease and pest classification module. The resulting model achieved a classification accuracy of 97.38% for different grape leaf diseases and pests, with a model size of 13.4796 MB and a computational cost of 0.3496 FLOPs / G, demonstrating good performance. After optimization for an edge computing platform, the model achieved a classification accuracy of 96.86% for different grape leaf diseases and pests, with a model size of 6.8825 MB and a computational cost of 0.1996 FLOPs / G. This result indicates that the system maintains high detection accuracy while achieving extremely high processing efficiency, providing technical support for the grape disease detection industry.

[0023] The leaf pest and disease classification model based on Mamba, as described in this application, enhances feature representation through spatial attention and leverages the efficient sequence modeling capabilities of the Mamba algorithm to achieve fast and accurate classification. It fully utilizes the advantages of the Mamba model, achieving linear computational efficiency far exceeding traditional neural network architectures and a smaller number of model parameters while maintaining high classification accuracy. This design enables the model to have faster inference speeds and lower resource consumption on edge computing devices, effectively solving the computational latency and storage space bottlenecks faced by complex models during deployment, and providing feasibility for real-time field pest and disease diagnosis.

[0024] The optimization for edge computing platforms effectively addresses the challenges of insufficient storage space and slow computation speed faced by lightweight models deployed on embedded devices through a progressive three-stage compression strategy. This module, by sequentially performing operator fusion, selective low-rank decomposition, and pruning operations, significantly reduces the number of model parameters and computational complexity while ensuring model accuracy. This provides crucial technical support for real-time diagnosis of leaf pest and disease classification systems on resource-constrained field edge devices.

[0025] Example 2 In one exemplary embodiment, the present invention also provides a computer system, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, the computer system includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the plant leaf disease and pest detection method suitable for edge computing described in Embodiment 1.

[0026] Those skilled in the art will understand that Figure 6 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer system to which the present application is applied. A specific computer system may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer system is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0027] In one exemplary embodiment, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0028] In one exemplary embodiment, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0030] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0031] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units, etc., and are not limited to these.

[0032] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0033] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for detecting plant leaf diseases and pests suitable for edge computing, characterized in that... Includes the following steps: Collect and organize leaf images of plant diseases, pests, and stresses to construct a basic image dataset for subsequent model training and evaluation; The basic image dataset is preprocessed through data augmentation and normalization to construct images for model training; Construct a Mamba-based classification model for plant leaf diseases and pests; The model is designed to be lightweight by optimizing it for edge computing platforms; Perform model training to obtain a pre-trained model; The trained plant leaf disease and pest classification model was used to detect plant leaf diseases and pests.

2. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 1, characterized in that, The method for constructing a basic image dataset includes the following steps: Collect images of diseases and pests affecting wine grapes and construct a dataset of grape leaf diseases and pests, which includes a set of images for each category.

3. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 1, characterized in that, The data augmentation and standardization include: Before performing overall data augmentation, data augmentation is prioritized for images with insufficient nutrients to supplement these samples and address the impact of poor training accuracy caused by data imbalance. Images in the training set are randomly flipped horizontally, vertically, rotated, randomly cropped with padding, and have their colors jittered to simulate image distortion caused by camera movement and lighting changes. This allows the training process to generate new training samples through random transformations without increasing the amount of original data. The data is then standardized by normalizing the pixel values ​​of each channel of the image according to the mean and standard deviation of the ImageNet dataset.

4. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 1, characterized in that: The plant leaf disease and pest classification model includes a convolutional feature extraction module, a spatial attention module, a Mamba core module, and a classifier. The convolutional feature extraction module is used to extract discriminative features from the data-enhanced and standardized image data. The spatial attention module is used to adaptively learn and strengthen important regions in the feature map while suppressing unimportant background or noise regions. The Mamba core module is used to perform deep modeling on the feature sequences extracted from the image to capture complex dependencies and high-level features. A classifier is used to map high-level features to a category space and output the classification result.

5. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 4, characterized in that: The convolutional feature extraction module includes a first two-dimensional fully connected convolutional layer. The output of the first two-dimensional fully connected convolutional layer is sequentially connected to the input of a first two-dimensional batch normalization layer, a first SiLU activation function layer, a second two-dimensional fully connected convolutional layer, and the second two-dimensional batch normalization layer and the second SiLU activation function layer. The image data after data augmentation and normalization is resampled to 64 x 64 pixels and then input into a module that performs four stacked convolutional feature extraction operations to gradually extract and combine discriminative features from low-level details to high-level semantics from the input image.

6. The plant leaf disease and pest detection method applicable to edge computing as described in claim 4, characterized in that: The spatial attention module is used to perform max pooling and average pooling on the high-order feature maps obtained after multi-layer convolution feature extraction, and then concatenate them. The concatenated feature maps are then convolved and activated in sequence, and then concatenated with the input of the spatial attention module before output. This enables the network to adaptively learn and strengthen important regions in the spatial location of the feature maps, while suppressing unimportant background or noise regions.

7. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 4, characterized in that: The Mamba core module first performs a feature transformation operation, flattening the spatial dimension feature map into a sequence, thus converting the convolutional feature map data into data that Mamba can process. The transformed sequence data is first input into the first layer of the Mamba block for initial feature learning and sequence dependency modeling, and then the feature distribution is stabilized and standardized through layer normalization. The normalized features are then input into the second layer of the Mamba block for in-depth sequence feature extraction, secondary training, and deep modeling to capture complex dependencies and high-level features in the features.

8. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 4, characterized in that: The classifier's processing method includes the following steps: The input features are first linearly transformed through a fully connected layer, then the data distribution is stabilized by a one-dimensional batch normalization layer, nonlinearity is introduced through the SiLU activation function, and finally a regularization layer is applied to prevent overfitting. This structure is repeated twice to achieve deep feature extraction. The output features are mapped to the dimension of the number of categories through the last fully connected layer, and then transformed into a probability distribution through the Softmax function. The category with the highest probability is selected as the prediction result.

9. The method for detecting plant leaf diseases and pests suitable for edge computing as described in claim 1, characterized in that, The model optimization includes operator optimization and fusion, low-rank decomposition, and pruning: Full-process operator optimization and fusion: First, scan each layer of the model and analyze the complete computation graph of the model. Then, mark the layers with continuous convolution, batch normalization and activation function operations as layers that can be fused into a single operator to reduce computation and memory access overhead. Low-rank decomposition: First, select the weight matrix, choosing a matrix with relatively low overall weight values. Then, select the rank. After completing the matrix decomposition, reconstruct the model based on the selected weight matrix and the size of the rank. Fine-tune, evaluate, and iterate the reconstructed model. Then, dynamically analyze the results of low-rank decomposition at different levels, and set accuracy loss threshold and model size threshold to find the low-rank decomposition structure that is most suitable for this model. Pruning involves first calculating the weight contribution of each layer of the model, then scoring the importance of the weights, removing unimportant weights or structural units, generating a sparse model, and performing unstructured pruning. Finally, fine-tuning is performed based on the model results.

10. A computer system, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the plant leaf disease and pest detection method suitable for edge computing as described in any one of claims 1-9.