Plant disease detection method based on improved YOLOv8n network

By improving the YOLOv8n network, combined with DRGhostConv, multi-scale spatial attention and efficient channel attention module, the problem of difficulty in taking into account both accuracy and efficiency in the existing technology is solved, and high-precision and lightweight plant disease detection is achieved, which is suitable for real-time applications of resource-constrained devices.

CN120164113APending Publication Date: 2025-06-17CHONGQING JIAOTONG UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510366168.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing plant disease detection methods are difficult to balance between accuracy and efficiency. Complex deep learning models require high hardware resources while lightweight design sacrifices detection accuracy, making it difficult to accurately identify subtle disease characteristics.

Method used

The improved YOLOv8n network is adopted to reduce the amount of parameters and improve the calculation efficiency through the DRGhostConv module, and combine the multi-scale spatial attention module and the efficient channel attention module to improve the multi-scale feature expression and detection accuracy.

Benefits of technology

While ensuring detection accuracy, the lightweight design of the model is achieved, which improves real-time detection capabilities on resource-constrained devices, especially in identifying small disease targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164113A_ABST
    Figure CN120164113A_ABST
Patent Text Reader

Abstract

The invention discloses a plant disease detection method based on an improved YOLOv8n network, and the method comprises the steps: optimizing a backbone network and a detection head network of the YOLOv8n network based on a DRGhostConv module and an efficient channel attention module for a to-be-detected plant disease image, and employing a multi-scale space attention module as a neck network of the optimized YOLOv8n network, a plant disease detection model is constructed and obtained, the public data set is used as input of the plant disease detection model for training and testing, parameters of the plant disease detection model are updated with minimization of a CIoU loss function as a target, and the trained plant disease detection model is obtained; and identifying and positioning the to-be-detected plant disease image based on the trained plant disease detection model to obtain a disease detection result of the to-be-detected plant disease image. According to the method, the lightweight design of the model is realized while the detection precision is ensured, and the real-time detection capability on resource-constrained equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection processing, and in particular to a plant disease detection method based on an improved YOLOv8n network. Background Art

[0002] With the improvement of people's requirements for the quality of life, the types and quantities of domestic plants have increased significantly. However, domestic plants are prone to be invaded by various diseases during the growth process, such as leaf spot disease, calcium deficiency, leaf burn disease, leaf blight disease, mosaic disease, and leaf curl virus disease, etc. These diseases not only affect the health and normal growth of plants, but may even cause the yellowing and withering of plant leaves, and in severe cases, lead to the death of the whole plant, bringing great troubles to plant lovers. The existing domestic plant disease detection methods mainly rely on manual observation and diagnosis. However, due to the large variety of domestic plant diseases and the complex symptom manifestations, different diseases may present similar symptoms, which are easy to be confused. This makes manual diagnosis not only require rich plant pathology knowledge and experience, but also be inefficient. Especially when faced with a large number of plants, it is difficult for manual diagnosis to be timely and efficient, and it is easy to miss or misdetect.

[0003] In recent years, with the rapid development of computer vision and deep learning technologies, target detection technology has made remarkable progress in the agricultural field. Currently, disease detection methods are roughly divided into two categories: One method uses complex deep learning models. Although it can provide high detection accuracy, the model is huge, with numerous parameters, and has high requirements for hardware resources, making it difficult to achieve real-time application on resource-constrained devices; Another method focuses on lightweight design. Although it reduces the requirements for hardware performance, it usually sacrifices detection accuracy and is difficult to accurately identify subtle disease features. This trade-off between accuracy and efficiency makes the existing detection methods face challenges in practical applications. Summary of the Invention

[0004] Aiming at the deficiencies of the above-mentioned existing technologies, the present invention provides a plant disease detection method based on an improved YOLOv8n network. By using the DRGhostConv module, the number of parameters is reduced and the calculation efficiency is improved. By using the multi-scale spatial attention module, the expression and fusion ability of multi-scale features are enhanced. And by using the efficient channel attention module, the detection accuracy of disease targets is further improved, thus achieving lightweight design of the model while ensuring the detection accuracy, and improving the real-time detection ability on resource-constrained devices.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A plant disease detection method based on an improved YOLOv8n network, comprising the following steps:

[0007] Obtain the plant disease image to be detected;

[0008] Use the DRGhostConv module and the efficient channel attention module to optimize the backbone network and the detection head network of the YOLOv8n network, and use the multi-scale spatial attention module as the neck network of the optimized YOLOv8n network to construct a plant disease detection model;

[0009] Use the public dataset as the input of the plant disease detection model for training and testing, and use the CIoU loss function to calculate the training loss of the plant disease detection model. Update the parameters of the plant disease detection model with the goal of minimizing the CIoU loss function to obtain the trained plant disease detection model;

[0010] Use the plant disease image to be detected as the input of the trained plant disease detection model, identify and locate the plant disease image to be detected, and obtain the disease detection result of the plant disease image to be detected.

[0011] As a preferred solution, before using the public dataset as the input of the plant disease detection model for training, it includes: annotating the six types of diseases including leaf spot, calcium deficiency, leaf burn, leaf blight, mosaic disease and leaf roll virus in the public dataset, performing size adjustment, normalization processing and data augmentation operation processing on the annotated plant disease image dataset, and dividing the processed plant disease image dataset into a training set, a validation set and a test set according to a ratio.

[0012] As a preferred solution, the plant disease detection model includes a backbone network, a neck network and a detection head network;

[0013] The backbone network includes a DRGhostConv module, 4 multi-scale feature fusion units and a spatial pyramid pooling fast module connected in sequence, where the multi-scale feature fusion unit includes a cascaded DRGhostConv module and a C2f module;

[0014] The neck network includes 2 feature fusion upsampling units, 2 multi-scale connection feature fusion units connected in sequence, and multi-scale spatial attention modules inserted in the 2 feature fusion upsampling units and the 2 multi-scale connection feature fusion units respectively. The feature fusion upsampling unit includes a cascaded upsampling module, a connection layer and a C2f module, and the multi-scale connection feature fusion unit includes a DRGhostConv module, a connection layer and a C2f module;

[0015] The detection head network includes three detection modules connected in sequence and efficient channel attention modules respectively inserted in the first detection module and the third detection module. The three detection modules are in parallel. The detection module includes a bounding box prediction unit and a class prediction unit in parallel. The bounding box prediction unit includes a cascaded convolution module, a convolution layer, and a bounding box loss function. The class prediction unit includes a convolution module, a convolution layer, and a class loss function. The convolution module includes a cascaded convolution layer, a batch normalization layer, and a SiLU activation function layer;

[0016] Among them, the input of the plant disease detection model serves as the input of the backbone network; the inputs of the second multi-scale feature fusion unit and the third multi-scale feature fusion unit in the backbone network also serve as the inputs of the connection layers in the second feature fusion upsampling unit and the first feature fusion upsampling unit in the neck network respectively. The output of the spatial pyramid pooling fast module in the backbone network serves as the input of the neck network and the input of the connection layer in the second multi-scale connected feature fusion unit in the neck network respectively; the outputs of the second feature fusion upsampling unit and the second multi-scale connected feature fusion unit in the neck network pass through the efficient channel attention modules in the detection head network respectively and then serve as the inputs of the first detection module and the third detection module respectively. The output of the first multi-scale connected feature fusion unit in the neck network passes through the second multi-scale spatial attention module and then serves as the input of the second detection module.

[0017] As a preferred solution, the processing process of the DRGhostConv module includes:

[0018] Performing preliminary feature extraction on the input feature map using the main path of a standard convolution with a convolution kernel size of 3×3; extracting features of the preliminary features obtained by the main path through two parallel channel paths respectively. One channel uses a convolution with a convolution kernel size of 3×3 to extract fine-grained features, and the other channel uses a convolution with a convolution kernel size of 5×5 to capture features with a larger receptive field; the two channel paths are processed in parallel with the main path, weighting the output features of the three paths through a channel attention mechanism, and concatenating the weighted features and the output of the main path in the channel dimension to obtain the feature map as the output of the DRGhostConv module.

[0019] As a preferred solution, the output of the DRGhostConv module is shown by the following formula:

[0020] Output=Concat(Conv 3×3 (path1),Conv 5×5 (path2),Attention);

[0021] Attention=σ(Conv1×1 (Input));

[0022] Wherein, Output represents the output of the DRGhostConv module, path1 represents the channel path using 3×3 convolution, path2 represents the channel path using 5×5 convolution, Concat(·) represents the concatenation operation along the channel dimension, Conv 3×3 represents 3×3 convolution, Conv 5×5 represents 5×5 convolution, Attention represents the channel attention mechanism, σ represents the sigmoid activation function, Conv 1×1 represents 1×1 convolution, Input represents the input to the channel attention mechanism.

[0023] As a preferred solution, the processing process of the multi-scale spatial attention module includes:

[0024] Performing preliminary feature extraction on the input feature map using the first convolution block and increasing the non-linear expression ability. The first convolution block includes a cascaded 1×1 convolution, a batch normalization layer, and a ReLU activation function. Taking the output of the first convolution block as the input of the second convolution block to extract the feature map. The second convolution block includes a cascaded 1×1 convolution, a batch normalization layer, and a Sigmoid function; respectively performing feature extraction on the feature map extracted by the second convolution block through two parallel channel paths. One channel path uses a 3×3 convolution layer to extract detailed disease features, and the other channel uses a 5×5 convolution layer to capture the position features of the disease area; concatenating the outputs of the two channel paths with the output of the second convolution block along the channel dimension to obtain the feature map as the output of the multi-scale spatial attention module.

[0025] As a preferred solution, the output of the multi-scale spatial attention module is shown in the following formula:

[0026] C2(F) = σ(BN(Conv 1×1 (RELU(BN(Conv 1×1 (F))))));

[0027] MSSAM(F) = Concat(C2(F), Conv 3×3 (C2(F)), Conv 5×5 (C2(F)));

[0028] Wherein, MSSAM(F) represents the output after being processed by the multi-scale spatial attention module, BN(·) represents batch normalization, RELU(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, C2(F) represents the output feature map of the second convolution block, and F represents the input feature map.

[0029] As a preferred solution, the processing process of the efficient channel attention module includes:

[0030] Perform global average pooling on the input feature map to generate a channel feature map; then, model the local dependence relationship between channels through a one-dimensional convolution operation to generate the attention weight of each channel; finally, multiply the generated channel attention weight by the input feature to obtain an enhanced channel feature map.

[0031] As a preferred solution, the output of the efficient channel attention module is shown in the following formula:

[0032] A c = σ(Conv1D(GAP(F)));

[0033] In the formula, A c represents the enhanced channel feature map output after being processed by the efficient channel attention module, Conv1D represents the one-dimensional convolution operation, and GAP represents the global average pooling operation.

[0034] As a preferred solution, the CIoU loss function is:

[0035]

[0036] In the formula, L CIoU represents the position loss of the prediction box, IoU represents the intersection over union of the prediction box and the ground truth box, ρ 2 (p, p qt ) represents the Euclidean distance between the center points of the prediction box and the ground truth box, p represents the center point of the prediction box, and p qt represents the center point of the ground truth box, c represents the diagonal length of the smallest rectangle containing the prediction box and the ground truth box, α represents the balance coefficient, and v represents the square of the difference in the aspect ratios of the prediction box and the ground truth box.

[0037] Compared with the prior art, the present invention has the following technical effects:

[0038] (1) By introducing the DRGhostConv module to replace the traditional convolution operation, the present invention optimizes the YOLOv8n network. Compared with the traditional convolution operation, the DRGhostConv module combines multiple parallel convolution paths and an attention mechanism, significantly reducing the number of model parameters, improving the computational efficiency, and at the same time enhancing the diversity of feature extraction and the model performance, especially showing remarkable effects when dealing with fine-grained disease features.

[0039] (2) The present invention replaces the feature fusion module of the original network with a multi-scale spatial attention (MSSAM) module, thereby improving the expression and fusion ability of multi-scale features. The MSSAM module combines multi-scale convolution and spatial attention mechanism, enabling the plant disease detection model to better capture disease features at different scales and improving the model's recognition ability for disease targets in complex backgrounds.

[0040] (3) The present invention introduces an efficient channel attention (ECA) module, which enhances the feature correlation between channels and the model's focusing ability on important features in the detection head network, thereby improving the detection accuracy of disease targets, especially performing excellently in the detection of small disease targets. Description of the Drawings

[0041] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:

[0042] Figure 1 is a flowchart of a plant disease detection method based on an improved YOLOv8n network disclosed by the present invention;

[0043] Figure 2 is the overall structure diagram of the plant disease detection model network in this embodiment;

[0044] Figure 3 is the ECA network structure used by the plant disease detection model in this embodiment;

[0045] Figure 4 is the effect diagram of this embodiment. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0047] The present invention will be further described in detail below with reference to the accompanying drawings.

[0048] Figure 1The present invention discloses a plant disease detection method based on an improved YOLOv8n network. Aiming at the problem that existing object detection technologies are difficult to accurately identify subtle disease features with limited parameters, the method includes:

[0049] S1. Obtain plant disease images to be detected;

[0050] S2. Use the DRGhostConv module and the efficient channel attention module to optimize the backbone network and the detection head network of the YOLOv8n network, and use the multi-scale spatial attention module as the neck network of the optimized YOLOv8n network to construct a plant disease detection model;

[0051] S3. Use the public dataset as the input of the plant disease detection model for training and testing, and use the CIoU loss function to calculate the training loss of the plant disease detection model. Update the parameters of the plant disease detection model with the goal of minimizing the CIoU loss function to obtain the trained plant disease detection model;

[0052] S4. Use the plant disease image to be detected as the input of the trained plant disease detection model, identify and locate the plant disease image to be detected, and obtain the disease detection result of the plant disease image to be detected.

[0053] Existing technologies for disease detection methods can be roughly divided into two categories: One method uses complex deep learning models. Although it can provide high detection accuracy, the model is huge, with numerous parameters, and has high requirements for hardware resources, making it difficult to achieve real-time applications on resource-constrained devices; Another method focuses on lightweight design. Although it reduces the requirements for hardware performance, it usually sacrifices detection accuracy and is difficult to accurately identify subtle disease features.

[0054] In response to this, the present invention proposes a plant disease detection method for an improved YOLOv8n-PLPD network. In this method, the DRGhostConv module, the efficient channel attention module, and the multi-scale spatial attention module are used to optimize the basic network structure, and the CIoU loss function is introduced to obtain a plant disease detection model.

[0055] The following further details the plant disease detection method based on the improved YOLOv8n network of the invention.

[0056] 1. Construct a plant disease image dataset

[0057] The plant disease image dataset is formed using plant disease pictures from the Plant Disease Computer Vision Project dataset of Roboflow. 6,155 original images containing six types of diseases, namely leaf spot disease, calcium deficiency, leaf scorch, leaf blight, mosaic disease, and curly leaf virus disease, are selected from the Plant Disease Computer Vision Project dataset of Roboflow. Different disease types are labeled through the open-source software LabelImg, where leaf spot disease is labeled as "0", calcium deficiency is labeled as "1", leaf scorch is labeled as "2", leaf blight is labeled as "3", mosaic disease is labeled as "4", and curly leaf virus disease is labeled as "5". The labeling operation includes manually bounding the disease areas in each image to generate a labeling file in the YOLO format. The labeling file contains the disease category and its corresponding bounding box coordinate information. After the labeling is completed, the images are resized, normalized, and data augmentation operations are performed. After data augmentation, a total of 14,771 images are finally generated, including 12,924 images in the training set, 1,230 images in the validation set, and 617 images in the test set, forming the final plant disease dataset for training and testing.

[0058] 2. Plant Disease Detection Model

[0059] As Figure 2 shown, specifically, in this embodiment, the plant disease detection model includes a backbone network, a neck network, and a detection head network;

[0060] The backbone network includes a DRGhostConv module, 4 multi-scale feature fusion units, and a spatial pyramid pooling fast module connected in sequence, where the multi-scale feature fusion unit includes a cascaded DRGhostConv module and a C2f module;

[0061] The neck network includes 2 feature fusion upsampling units, 2 multi-scale connection feature fusion units, and multi-scale spatial attention modules inserted respectively in the 2 feature fusion upsampling units and 2 multi-scale connection feature fusion units. The feature fusion upsampling unit includes a cascaded upsampling module, a connection layer, and a C2f module, and the multi-scale connection feature fusion unit includes a DRGhostConv module, a connection layer, and a C2f module;

[0062] The detection head network includes three detection modules connected in sequence and efficient channel attention modules respectively inserted into the first detection module and the third detection module. The three detection modules are in parallel. The detection module includes a bounding box prediction unit and a class prediction unit in parallel. The bounding box prediction unit includes a cascaded convolution module, a convolution layer, and a bounding box loss function. The class prediction unit includes a convolution module, a convolution layer, and a class loss function. The convolution module includes a cascaded convolution layer, a batch normalization layer, and a SiLU activation function layer;

[0063] Among them, the input of the plant disease detection model is used as the input of the backbone network; the inputs of the second multi-scale feature fusion unit and the third multi-scale feature fusion unit in the backbone network are also used as the inputs of the connection layer in the second feature fusion upsampling unit and the first feature fusion upsampling unit in the neck network respectively. The output of the spatial pyramid pooling fast module in the backbone network is used as the input of the neck network and the input of the connection layer in the second multi-scale connected feature fusion unit in the neck network respectively; the outputs of the second feature fusion upsampling unit and the second multi-scale connected feature fusion unit in the neck network are respectively used as the inputs of the first detection module and the third detection module after passing through the efficient channel attention modules in the detection head network, and the output of the first multi-scale connected feature fusion unit in the neck network is used as the input of the second detection module after passing through the second multi-scale spatial attention module.

[0064] In this embodiment, YOLOv8n is used as the basic model. The DRGhostConv module is introduced into the original network to replace the traditional convolution operation to reduce the number of parameters and improve the calculation efficiency. At the same time, the attention mechanism is introduced into the backbone network part to enhance the extraction of important features; in the neck network part, the MSSAM (multi-scale spatial attention mechanism) module is used to replace the feature fusion module of the original network to improve the expression and fusion ability of multi-scale features; in the detection head network part, the ECA (efficient channel attention mechanism) and multiple DRGhostConv modules are introduced to further improve the detection accuracy and position regression ability of fine-grained disease targets; finally, the model is trained through loss functions such as Bbox Loss and Cls Loss, and the model weights are obtained after convergence.

[0065] The DRGhostConv module, the multi-scale spatial attention module, and the efficient channel attention module will be introduced in detail below.

[0066] 2.1. DRGhostConv module

[0067] Introduce the DRGhostConv module into the backbone network part of the YOLOv8n model. This module reduces the number of parameters through a multi-path convolution design while maintaining the diversity of convolutions. Specifically, in this embodiment, the DRGhostConv module consists of three parallel paths:

[0068] Main path (Conv2d): Use a 3×3 convolutional kernel to perform standard convolution operations on the input features and extract basic features;

[0069] Path1: Adopt a 3×3 convolutional kernel to specifically extract fine-grained features and enhance the detection ability for small-scale diseases;

[0070] Path2: Adopt a 5×5 convolutional kernel to capture features with a larger receptive field, which is especially suitable for detecting large-area diseases;

[0071] Attention mechanism: After fusing the three paths, introduce an attention mechanism for feature weighting. The attention mechanism can focus on important disease regions according to the weights of the input features, while reducing the attention to the background or irrelevant regions, ensuring that the model can accurately capture disease details;

[0072] Output feature fusion: Finally, concatenate the feature maps of all paths in the channel dimension through the Concat operation to output a diverse and weighted feature map.

[0073] Thus, the processing process of the DRGhostConv module includes: performing preliminary feature extraction on the input feature map using the main path with a 3×3 standard convolution kernel; respectively extracting features of the preliminary features obtained by the main path through two parallel channel paths, where one channel uses a 3×3 convolutional kernel to extract fine-grained features, and the other channel uses a 5×5 convolutional kernel to capture features with a larger receptive field; the two channel paths are processed in parallel with the main path, weight the output features of the three paths through the channel attention mechanism, and concatenate the weighted features with the output of the main path in the channel dimension to obtain the feature map as the output of the DRGhostConv module.

[0074] The output of the DRGhostConv module is shown in the following formula:

[0075] Output = Concat(Conv 3×3 (path1), Conv 5×5 (path2), Attention);

[0076] Attention = σ(Conv 1×1 (Input));

[0077] In the formula, Output represents the output of the DRGhostConv module, path1 represents the channel path using 3×3 convolution, path2 represents the channel path using 5×5 convolution, Concat(·) represents the concatenation operation along the channel dimension, Conv 3×3 represents 3×3 convolution, Conv 5×5 represents 5×5 convolution, Attention represents the channel attention mechanism, σ represents the sigmoid activation function, Conv 1×1 represents 1×1 convolution, Input represents the input to the channel attention mechanism.

[0078] In this embodiment, by introducing the DRGhostConv module to replace the traditional convolution operation, the YOLOv8n network is optimized. Compared with the traditional convolution operation, the DRGhostConv module combines multiple parallel convolution paths and the attention mechanism, significantly reducing the number of model parameters, improving the computational efficiency, and at the same time enhancing the diversity of feature extraction and the model performance, especially showing remarkable effects when dealing with fine-grained disease features.

[0079] 2.2. Multi-scale Spatial Attention Module

[0080] In this embodiment, a multi-scale spatial attention (MSSAM) module is introduced in the neck network part to fuse features of different scales extracted by the backbone network, thereby improving the model's detection ability for multi-scale diseases. Specifically, the processing process of the multi-scale spatial attention (MSSAM) module includes: performing preliminary feature extraction on the input feature map using the first convolutional block and increasing the non-linear expression ability. The first convolutional block includes a cascaded 1×1 convolution, a batch normalization layer, and a ReLU activation function. Taking the output of the first convolutional block as the input of the second convolutional block to extract the feature map. The second convolutional block includes a cascaded 1×1 convolution, a batch normalization layer, and a Sigmoid function; respectively performing feature extraction on the feature map extracted by the second convolutional block through two parallel channel paths. One channel path uses a 3×3 convolutional layer to extract detailed disease features, and the other channel uses a 5×5 convolutional layer to capture the location features of the disease area; concatenating the outputs of the two channel paths with the output of the second convolutional block along the channel dimension to obtain the feature map as the output of the multi-scale spatial attention module.

[0081] The output of the multi-scale spatial attention module is shown in the following formula:

[0082] C2(F) = σ(BN(Conv 1×1 (RELU(BN(Conv 1×1 (F))))));

[0083] MSSAM(F) = Concat(C2(F), Conv 3×3 (C2(F)), Conv 5×5 (C2(F)));

[0084] Wherein, MSSAM(F) represents the output after being processed by the multi-scale spatial attention module, BN(·) represents batch normalization, RELU(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, C2(F) represents the output feature map of the second convolutional block, and F represents the input feature map.

[0085] In this embodiment, the multi-scale spatial attention module is used to replace the feature fusion module of the original network, thereby improving the expression and fusion ability of multi-scale features. The MSSAM module combines multi-scale convolution and spatial attention mechanisms, enabling the plant disease detection model to better capture disease features at different scales and improving the model's recognition ability for disease targets in complex backgrounds.

[0086] 2.3. Efficient Channel Attention Module

[0087] As Figure 3 shown, in this embodiment, an efficient channel attention (ECA) module is introduced into the detection head network to enhance the feature correlation between channels and improve the model's accuracy and detail processing ability in disease detection. Specifically, the processing process of the ECA module in this embodiment includes: First, global average pooling is performed on the input feature map to compress the features of each channel to obtain global information and generate a channel feature map; then, one-dimensional convolution operation is used to model the local dependence relationship between channels to generate the attention weights of each channel; finally, the generated channel attention weights are multiplied by the input features to obtain the enhanced feature map.

[0088] The output of the efficient channel attention module is shown in the following formula:

[0089] A c = σ(Conv1D(GAP(F)));

[0090] Wherein, A c represents the enhanced channel feature map output after being processed by the efficient channel attention module, Conv1D represents one-dimensional convolution operation, and GAP represents global average pooling operation.

[0091] The efficient channel attention module introduced in this embodiment enhances the feature correlation between channels, enhances the model's focusing ability on important features in the detection head network, thereby improving the detection accuracy of disease targets, especially showing excellent performance in the detection of small disease targets.

[0092] 3. Training of the Plant Disease Detection Model

[0093] In the process of model training in this embodiment, by setting model parameters, including input image size, prior box size, target category, initial learning rate, and learning rate adjustment strategy, etc., the pre-annotated and processed plant disease images and the divided plant disease training set are input into the plant disease detection model for training. During the training process, a validation set is used to verify the model, and the CIoU loss function is used to calculate the training loss of the plant disease detection model. The parameters of the plant disease detection model are updated with the goal of minimizing the CIoU loss function, and then the plant disease detection model is trained.

[0094] The CIoU loss function is as follows:

[0095]

[0096] In the formula, L CIoU represents the location loss of the predicted box, IoU represents the intersection over union of the predicted box and the ground truth box, ρ 2 (p, p qt ) represents the Euclidean distance between the center points of the predicted box and the ground truth box, p represents the center point of the predicted box, p qt represents the center point of the ground truth box, c represents the diagonal length of the smallest rectangle containing the predicted box and the ground truth box, α represents the balance coefficient, v represents the square of the difference in aspect ratio between the predicted box and the ground truth box, w gt , h gt respectively represent the width and height of the ground truth box, and w and h respectively represent the width and height of the predicted box.

[0097] When specifically applying this embodiment, first use the ArgumentParser library to define and parse input parameters, then load the pre-trained improved YOLOv8n model through the torch library, and initialize the model parameters. Then preprocess the input image, including operations such as adjusting the image size and normalization, and convert the processed image into a Tensor that conforms to the model input format. Subsequently, input the Tensor into the model for forward propagation to obtain the prediction results. Finally, process the detection results through post-processing steps, including parsing the bounding box coordinates, calculating the confidence, etc. Finally, visualize the detected plant disease targets on the output image and save the final results.

[0098] 4. Model Deployment and Verification

[0099] Deploy the trained model to the computer, load the optimal model weights obtained from training, input the test set in the plant disease image dataset into the model, and through feature extraction and fusion, generate the bounding box coordinates, categories, and confidence levels of the detected target diseases in the Detect layer. Subsequently, remove redundant detection boxes through non-maximum suppression (NMS), retain the optimal detection results, and output and save the final detection results in the form of images; input the processed test set images of domestic plant diseases into the deployed improved YOLOv8n model, detect and output the bounding box positions, categories, and confidence levels of each detected disease. The detection results can be output through a display terminal, facilitating users to view and analyze the disease distribution in real time.

[0100] By evaluating indicators such as the mean average precision (mAP), recall rate, and detection speed (FPS) of the model, the improved YOLOv8n model performs excellently in the disease detection task, especially showing remarkable effects in small target detection.

[0101] Through experimental verification, the improved model can effectively detect domestic plant diseases and generate accurate detection results, including disease positions, categories, and confidence levels. The terminal output results of the model are as Figure 4 shown, and users can monitor domestic plant diseases in real time through the terminal.

[0102] 5. Summary

[0103] This embodiment discloses a plant disease detection method based on an improved YOLOv8n network, which improves and optimizes different task levels of plant disease detection. First, by introducing the DRGhostConv module to replace the traditional convolution operation, the YOLOv8n network is optimized. Compared with the traditional convolution operation, the DRGhostConv module combines multiple parallel convolution paths and an attention mechanism, significantly reducing the number of model parameters, improving the computational efficiency, and at the same time enhancing the diversity of feature extraction and model performance, especially showing remarkable effects when dealing with fine-grained disease features. Then, the multi-scale spatial attention module is used to replace the feature fusion module of the original network, thereby improving the expression and fusion ability of multi-scale features. The MSSAM module combines multi-scale convolution and spatial attention mechanisms, enabling the plant disease detection model to better capture disease features at different scales and improving the model's recognition ability for disease targets in complex backgrounds. Secondly, the efficient channel attention module is introduced to enhance the feature correlation between channels and the focusing ability of the model on important features in the detection head network, thereby improving the detection accuracy of disease targets, especially showing excellent performance in the detection of small disease targets. This method realizes the lightweight design of the model while ensuring the detection accuracy and improves the real-time detection ability on resource-constrained devices. Finally, non-maximum suppression (NMS) is used to optimize the detection results, improving the accuracy and stability of plant disease detection.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described by referring to the preferred embodiments of the present invention, those of ordinary skill in the art should understand that various changes can be made in form and details without departing from the spirit and scope of the present invention defined by the appended claims.

Claims

1. A plant disease detection method based on an improved YOLOv8n network, characterized in that: The steps include: Acquire an image of a plant disease to be detected; Use the DRGhostConv module and efficient channel attention module to optimize the backbone network and detection head network of the YOLOv8n network, and use the multi-scale spatial attention module as the neck network of the optimized YOLOv8n network to build a plant disease detection model. The public dataset is used as the input of the plant disease detection model for training and testing, and the CIoU loss function is used to calculate the training loss of the plant disease detection model. The parameters of the plant disease detection model are updated with the goal of minimizing the CIoU loss function to obtain the trained plant disease detection model. The plant disease image to be detected is used as the input of the trained plant disease detection model, the plant disease image to be detected is identified and located, and the disease detection result of the plant disease image to be detected is obtained.

2. The plant disease detection method based on the improved YOLOv8n network according to claim 1, characterized in that: Before using the public data set as the input of the plant disease detection model for training, the method includes: labeling six types of diseases in the public data set, including leaf spot, calcium deficiency, leaf burn, leaf blight, mosaic disease and leaf curl virus, resizing, normalizing and data enhancement operations on the labeled plant disease image data set, and dividing the processed plant disease image data set into a training set, a validation set and a test set in proportion.

3. The plant disease detection method based on the improved YOLOv8n network according to claim 1, characterized in that: The plant disease detection model includes a backbone network, a neck network and a detection head network; The backbone network includes a DRGhostConv module, four multi-scale feature fusion units and a spatial pyramid pooling fast module connected in sequence, wherein the multi-scale feature fusion unit includes a cascaded DRGhostConv module and a C2f module; The neck network includes two feature fusion upsampling units, two multi-scale connection feature fusion units, and multi-scale spatial attention modules respectively inserted into the two feature fusion upsampling units and the two multi-scale connection feature fusion units, wherein the feature fusion upsampling unit includes a cascaded upsampling module, a connection layer, and a C2f module, and the multi-scale connection feature fusion unit includes a DRGhostConv module, a connection layer, and a C2f module; The detection head network includes three detection modules connected in sequence and efficient channel attention modules respectively inserted into the first detection module and the third detection module, wherein the three detection modules are connected in parallel, wherein the detection module includes a bounding box prediction unit and a category prediction unit connected in parallel, the bounding box prediction unit includes a cascaded convolution module, a convolution layer and a bounding box loss function, the category prediction unit includes a convolution module, a convolution layer and a category loss function, and the convolution module includes a cascaded convolution layer, a batch normalization layer and a SiLU activation function layer; Among them, the input of the plant disease detection model is used as the input of the backbone network; the input of the second multi-scale feature fusion unit and the third multi-scale feature fusion unit in the backbone network are also used as the input of the connection layer in the second feature fusion upsampling unit and the input of the connection layer in the first feature fusion upsampling unit in the neck network, respectively; the output of the spatial pyramid pooling fast module in the backbone network is used as the input of the neck network and the input of the connection layer in the second multi-scale connection feature fusion unit in the neck network, respectively; the output of the second feature fusion upsampling unit and the second multi-scale connection feature fusion unit in the neck network, respectively, passes through the efficient channel attention module in the detection head network, and then serves as the input of the first detection module and the third detection module, respectively; the output of the first multi-scale connection feature fusion unit in the neck network passes through the second multi-scale spatial attention module as the input of the second detection module.

4. The plant disease detection method based on the improved YOLOv8n network according to claim 3, characterized in that: The processing process of the DRGhostConv module includes: The input feature map is subjected to preliminary feature extraction using the main path of standard convolution with a kernel size of 3×3. The preliminary features extracted by the main path are extracted through two parallel channel paths, one of which uses convolution with a kernel size of 3×3 to extract fine-grained features, and the other uses convolution with a kernel size of 5×5 to capture features with a larger receptive field. The two channel paths are processed in parallel with the main path, and the output features of the three paths are weighted through the channel attention mechanism. The weighted features are concatenated with the output of the main path in the channel dimension to obtain the feature map output by the DRGhostConv module.

5. The plant disease detection method based on the improved YOLOv8n network according to claim 4, characterized in that: The output of the DRGhostConv module is shown in the following formula: Output=Concat(Conv 3×3 (path1),Conv 5×5 (path2),Attention); Attention=σ(Conv 1×1 (Input)); Where Output represents the output of the DRGhostConv module, path1 represents the channel path using 3×3 convolution, path2 represents the channel path using 5×5 convolution, Concat(·) represents the concatenation operation along the channel dimension, and Conv 3×3 Represents 3×3 convolution, Conv 5×5 represents 5×5 convolution, Attention represents the channel attention mechanism, σ represents the sigmoid activation function, Conv 1×1 Represents a 1×1 convolution, and Input represents the input of the channel attention mechanism.

6. The plant disease detection method based on the improved YOLOv8n network according to claim 3, characterized in that: The processing process of the multi-scale spatial attention module includes: The first convolution block is used to perform preliminary feature extraction on the input feature map and increase the nonlinear expression ability. The first convolution block includes a cascaded 1×1 convolution, a batch normalization layer and a ReLU activation function. The output of the first convolution block is used as the input of the second convolution block to extract the feature map. The second convolution block includes a cascaded 1×1 convolution, a batch normalization layer and a Sigmoid function. The feature map extracted by the second convolution block is respectively subjected to feature extraction through two parallel channel paths, one of which uses a 3×3 convolution layer to extract detailed disease features, and the other channel uses a 5×5 convolution layer to capture the location features of the diseased area. The outputs of the two channel paths are spliced ​​with the output of the second convolution block in the channel dimension to obtain a feature map output by the multi-scale spatial attention module.

7. The plant disease detection method based on the improved YOLOv8n network according to claim 6, characterized in that: The output of the multi-scale spatial attention module is shown in the following formula: C2(F)=σ(BN(Conv 1×1 (RELU(BN(Conv 1×1 (F)))))); MSSAM(F)=Concat(C2(F),Conv 3×3 (C2(F)),Conv 5×5 (C2(F))); Where MSSAM(F) represents the output after processing by the multi-scale spatial attention module, BN(·) represents batch normalization, RELU(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, C2(F) represents the output feature map of the second convolutional block, and F represents the input feature map.

8. The plant disease detection method based on the improved YOLOv8n network according to claim 3, characterized in that: The processing process of the efficient channel attention module includes: The input feature map is globally averaged pooled to generate a channel feature map. Then, the local dependency between channels is modeled through a one-dimensional convolution operation to generate the attention weights of each channel. Finally, the generated channel attention weights are multiplied with the input features to obtain the enhanced channel feature map.

9. The plant disease detection method based on the improved YOLOv8n network according to claim 8, characterized in that: The output of the efficient channel attention module is shown in the following formula: A c Nσ(Conv1D(GAP(F))) In the formula, A c It represents the enhanced channel feature map output after being processed by the efficient channel attention module, Conv1D represents a one-dimensional convolution operation, and GAP represents a global average pooling operation.

10. The plant disease detection method based on the improved YOLOv8n network according to claim 1, characterized in that: The CIoU loss function is: Where, L CIoU represents the position loss of the predicted box, IoU represents the intersection-over-union ratio between the predicted box and the real box, ρ 2 (p,p qt ) represents the Euclidean distance between the center point of the predicted box and the true box, p represents the center point of the predicted box, and p qt represents the center point of the real box, c represents the diagonal length of the minimum rectangle containing the predicted box and the real box, α represents the balance coefficient, and v represents the square of the difference in aspect ratio between the predicted box and the real box.

Citation Information

Cited By

  • Underground disease detection system and detection method based on YOLOv8n

    CN120495282A

  • Underground disease detection system and detection method based on YOLOv8n

    CN120495282B

  • Water supply pipeline disease identification method fusing lightweight trunk and multi-scale features, computer system, readable storage medium and program product

    CN120747904A

  • Oral panoramic film tooth disease detection system based on multi-scale feature extraction mechanism

    CN121169796A

  • Improved StarNet-YOLOv13-based unmanned aerial vehicle field tobacco virus disease lightweight detection method

    CN121767895A