Tea disease detection method based on improved YOLOv8

By constructing a data set of tea diseases in complex scenarios and improving the YOLOv8 model, the problem of multi-scale disease identification of tea disease detection in complex natural scenarios is solved, accurate detection and early warning of tea diseases are achieved, and the level of intelligence of the tea industry is improved.

CN120356010APending Publication Date: 2025-07-22KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510528384.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing tea disease detection technology is difficult to deal with multi-scale and multi-type diseases in complex natural scenarios, and the image resolution is insufficient, resulting in low detection accuracy and efficiency, making it difficult to meet the actual needs of large-scale tea gardens.

Method used

A complex scene tea disease data set was constructed, and the improved SSM-YOLO model was adopted. By introducing the SPPFCSPC structure and SS-Conv module, combined with the improved MPDIoU loss function, the model's fusion ability and target positioning accuracy of multi-scale features were improved.

Benefits of technology

It has achieved accurate detection of tea diseases in complex backgrounds, improved detection efficiency and reliability, supported intelligent management of the tea industry, reduced the use of pesticides, and promoted sustainable development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356010A_ABST
    Figure CN120356010A_ABST
Patent Text Reader

Abstract

The invention discloses a tea disease detection method based on improved YOLOv8. A data set containing 6560 high-quality RGB images is constructed, eight common tea diseases and health control classes are covered, and a rich data basis is provided for model training. In a model design level, an SPPFCSPC structure is introduced to replace part of an SPPF layer, the advantages of spatial pyramid pooling and a full convolutional network are fused, and multi-scale feature efficient fusion is realized; an SS-Conv module is adopted to replace traditional convolution, spatial pyramid decomposition convolution and a self-adaptive attention mechanism are integrated, and the low-resolution image and small target feature extraction capacity is improved. Meanwhile, an improved MPDIOU loss function is provided, and the target positioning precision is optimized by calculating multi-point distance information between a prediction frame and a real frame, so that the model pays more attention to position deviation, and the detection precision of disease targets with different scales is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart agriculture, and particularly to a method for detecting tea diseases based on improved YOLOv8. Background Art

[0002] With the rapid development of agricultural informatization and intelligent agriculture, computer vision and deep learning technologies have shown great application potential in the field of crop disease detection. Computer vision technology, especially deep learning algorithms, can automatically learn image features and perform efficient object detection and image segmentation by simulating the working principle of the human brain neural network. In agricultural disease detection, deep learning models can accurately identify subtle features such as disease spots and insect pests on crop leaves, realizing early warning and precise prevention and control of diseases. The application of this technology not only improves the efficiency of disease detection but also significantly enhances the accuracy and reliability of detection, providing a scientific basis for agricultural production decisions and promoting the development of agriculture towards intelligence and precision.

[0003] However, in complex natural scenarios, tea disease detection still faces many technical challenges. Traditionally, tea disease detection mainly relies on manual inspections, which are not only time-consuming and laborious but also difficult to ensure the comprehensiveness and accuracy of detection. Especially in large-scale plantations, the efficiency of manual detection is low, often resulting in untimely discovery of diseases and delaying the best prevention and control time. In recent years, although computer vision technology based on deep learning has made significant progress in crop disease detection, in complex tea garden environments, factors such as changes in lighting conditions, occlusion by branches and leaves, and diversity of soil backgrounds seriously interfere with the accurate segmentation and recognition of disease areas. In addition, at the initial stage of the disease, it usually shows small-area damage, and the images collected in the field are easily restricted by equipment, resulting in insufficient image resolution, further increasing the detection difficulty. Existing models often show problems of insufficient adaptability when dealing with multi-scale and multi-type diseases, and it is difficult to cope with the diverse disease types and variable morphologies in actual production, restricting the generalization ability and practical application effect of the models. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for detecting tea diseases based on improved YOLOv8 to solve the above technical problems.

[0005] The above technical purpose of the present invention is achieved through the following technical solutions: A method for detecting tea diseases based on improved YOLOv8 includes:

[0006] Constructing a tea disease dataset for complex scenarios:

[0007] Collect a number of high-quality RGB images through a combination of web crawlers and manual screening, covering 8 disease categories including tea anthracnose, tea blister blight, tea zonate leaf spot, tea phyllosticta leaf blight, tea red rust alga disease, tea white star disease, tea shoot blight, and tea sooty mold, as well as a healthy leaf control category;

[0008] Use quantitative indicators such as color histogram entropy, color proportion distribution, and edge density to evaluate the complexity of the image background, ensuring that the resolution of each image is not less than 2000×2000 pixels and the proportion of the disease area exceeds 15%;

[0009] Design an SSM-YOLO detection model:

[0010] Feature extraction network: Use the SS-Conv module to replace the traditional convolutional layer. The SS-Conv integrates spatial pyramid decomposition convolution and an adaptive attention module, enhancing the ability to extract features of low-resolution images and small targets while maintaining computational efficiency;

[0011] Feature fusion structure: Introduce the SPPFCSPC module to replace part of the spatial pyramid pooling (SPPF) layer. This module realizes the efficient fusion and enhancement of multi-scale features through cascading spatial pyramid pooling, a fully convolutional network, and a channel attention mechanism;

[0012] Loss function optimization: Use an improved multi-point distance intersection over union loss function. Calculate the multi-point Euclidean distance, overlapping area, and union area between the predicted box and the ground truth box through the formula, and at the same time consider the impact of target scale changes on the positioning accuracy, improving the detection accuracy of the model for disease targets of different scales. The calculation formula is as follows:

[0013]

[0014]

[0015] L MPDIoU = 1 - MPDIoU(10)

[0016] In the formula are the position coordinates of the predicted box, are the position coordinates of the ground truth box, and w, h are the width and height of the input image.

[0017] Further preferably, the specific implementation method of the SS-Conv module includes:

[0018] SPDConv component: Replace the stride convolution and pooling layer through a space-to-depth conversion method, retaining the detailed information of the image;

[0019] SimAM attention mechanism: Generate the weight of the feature map by optimizing the energy function, enhancing the feature expression ability;

[0020] Module integration: Take the output of SPDConv as the input of SimAM to form an end-to-end feature enhancement channel.

[0021] Further preferably, the calculation process of the MPDIoU loss function is as follows:

[0022] Distance calculation: Calculate the Euclidean distances d1 and d2 between the upper left corner and the lower right corner of the predicted box and the ground truth box according to formulas (1) and (2) respectively.

[0023] Boundary value determination: Take the maximum values of the predicted box and the ground truth box in the x-axis and y-axis directions as the coordinates of the upper left corner of the intersection area through formulas (5) and (6), and take the minimum values as the coordinates of the lower right corner through formulas (7) and (8).

[0024] Overlap area calculation: Calculate the area of the intersection area using formula (9), and calculate IoU in combination with the union area of the predicted box and the ground truth box.

[0025] Loss value calculation: Calculate the MPDIoU loss value according to formula (10), that is, 1 - IoU, and introduce multi-point distance information to correct the calculation result of IoU.

[0026] Further preferably, the construction method of the complex scene tea disease dataset further includes:

[0027] Data cleaning: Eliminate images with a resolution lower than 2000×2000 or unclear disease characteristics.

[0028] Semi-automated annotation: Use the labelImg tool to annotate the filtered images to generate XML files in PascalVOC format.

[0029] Data augmentation: Perform operations such as rotation, flipping, and brightness adjustment on the annotated dataset to improve the generalization ability of the model.

[0030] In summary, the present invention has the following beneficial effects:

[0031] Firstly, by introducing the SPPFCSPC structure and the SS-Conv module, the model's ability to process multi-scale features in complex scenes is enhanced, making disease detection more accurate. At the same time, by adopting the improved MPDIoU loss function, the target localization accuracy is optimized, further improving the detection performance.

[0032] Secondly, the improved YOLOv8 model has stronger robustness to factors such as complex backgrounds, lighting changes, and foliage occlusion, can work stably in various natural scenes, and improves the reliability of detection.

[0033] Thirdly, this method realizes the real-time detection of tea diseases, greatly improves the detection efficiency, and helps to achieve early warning and timely prevention and control of diseases.

[0034] Fourthly, it provides an efficient technical support for disease detection in the tea industry, promotes the development of agricultural informatization and intelligent agriculture, helps to improve agricultural production efficiency, and reduces labor costs.

[0035] Fifthly, through accurate disease detection, it helps to reduce the overuse of pesticides, protect the ecological environment, and promote the sustainable development of the tea industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is an example diagram for judging the complexity of the picture background;

[0037] Figure 2 It is an example diagram of a self-made dataset;

[0038] Figure 3 It is the network structure diagram of YOLOv8;

[0039] Figure 4 It is the network structure diagram of SSM-YOLO;

[0040] Figure 5 It is the network structure diagram of SPPFCSPC;

[0041] Figure 6 It is the calculation process diagram of SS-Conv;

[0042] Figure 7 It is the attention mechanism structure of SS-Conv;

[0043] Figure 8 It is the performance index curve of SSM-YOLO;

[0044] Figure 9 It is the training time consumption and inference speed of the model;

[0045] Figure 10 It is the comparison diagram of the actual detection effects of different algorithms;

[0046] Figure 11 It is the visualization result of Grad-CAM gradual change of the same algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with embodiments. Those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention.

[0048] Embodiment

[0049] 1. Materials and Methods

[0050] 1.1 Dataset

[0051] 1.1.1 Data Collection

[0052] The transfer and generalization ability of deep learning algorithm models has limitations, making it possible that the content that can be recognized in the laboratory environment may not be correctly recognized or the performance may decline in the natural complex environment. Therefore, in this study, tea leaf disease images with complex backgrounds under natural conditions were collected to construct data to ensure the feasibility and effectiveness of the model's inference and recognition in actual scenarios. The data was sourced from web collection. The study combined manual web search and Python automatic web crawler technology to collect suitable tea leaf disease image data. To ensure data quality, a series of search keywords were formulated, such as tea leaves, diseases, tea diseases, leaf diseases, tea leaf diseases, etc., and relevant images were widely collected on major search engines and professional websites. In addition, well-known public dataset platforms such as Kaggle, Roboflow, and Baidu PaddlePaddle platform were used to further obtain some tea leaf disease image data. Finally, the collected image data was initially screened to remove images with a resolution lower than 2000*2000, and PNG and JPG format images related to the diseases were carefully selected to ensure their clarity and the accuracy of the disease types.

[0053] During this process, to enrich the diversity of the dataset, ensure its authenticity, make the image data more closely resemble the disease images with complex backgrounds under natural conditions, and the disease characteristics are consistent with the disease manifestations in the real natural environment, tea leaf disease images with different light intensities and different disease types were selected, which helps the subsequent model to better learn and generalize.

[0054] In the present invention, the complex environment is defined as a complex background image under quasi-natural conditions in which the image background contains various color, texture, shape, and edge distribution information. A color region segmentation and edge detection algorithm was written in Python language to achieve the quantitative statistics of the pixels in each color region and the edge pixels. Through experimental verification, based on the statistical information and setting appropriate thresholds, the complexity of the image background can be judged. Examples are as Figure 1 shown

[0055] Therefore, all the images preprocessed by web collection were processed and screened by the above algorithm to ensure that the collected image data all meet the complex background. A tea leaf disease dataset of its own was constructed, including 8 disease categories of tea leaves bitten by mosquitoes, tea leaves infested with spider mites, black rot, leaf rust, white spot disease, algal leaf spot, gray blight, and brown blight, and 1 healthy leaf control category, with a total of 6560 images. Example sample images are as Figure 2 .

[0056] 1.1.2 Data Annotation

[0057] Subsequently, in combination with the model-assisted bounding box drawing and the annotation tool labelImg, a semi-automated annotation process was adopted to complete data annotation. In the semi-automated annotation process, first, an advanced object detection model was used to assist in drawing bounding boxes, and then the annotation tool labelImg was used to correct, refine, and supplement the automatically generated bounding boxes. The annotation formats supported by labelImg include PascalVOC, YOLO, and CreateML. In the present invention, it was chosen to annotate image files into XML format files, that is, PascalVOC format, which is convenient for subsequent label file processing and data format conversion. This semi-automated process not only improves the annotation speed but also greatly reduces the huge workload of manually drawing bounding boxes one by one, thereby improving the efficiency of data annotation. The number of pictures in each category is shown in Table 1.

[0058] Table 1 Number of pictures in each category

[0059]

[0060] 1.2 Method

[0061] 1.2.1 YOLOv8 Network Architecture

[0062] The YOLO series of algorithms are representative algorithms in one-stage object detection algorithms. Its core idea is to regard the object detection task as a regression problem and simultaneously predict the location and category of the object through a single neural network. Compared with traditional object detection methods, it has the advantages of real-time performance and simplicity and rapidity. YOLOv8 is a relatively new and stable detection model at present, and its detection performance is good. Its network structure mainly includes an input image, a backbone feature extraction part, a neck network, and a detection head. As Figure 3 shown. It uses the Conv structure as the backbone network to extract features. After the backbone feature extraction network is the SPPF structure, also known as spatial pyramid pooling, which can convert feature maps of any size into fixed-size feature vectors to achieve the feature map fusion of local features and global features. The neck network adopts a specific FPN plus PANet structure to achieve feature fusion at different scales. The detection head is the mainstream decoupled head structure.

[0063] 1.2.2 SSM-YOLO Leaf Disease Detection Model

[0064] Based on the advantages of the YOLOv8 algorithm, the present invention proposes an improved tea leaf disease detection and recognition algorithm, which improves the accuracy of the network in identifying tea leaf diseases in complex backgrounds while ensuring real-time performance and without increasing the model complexity. Figure 4 The overall framework of the SSM-YOLO leaf disease detection model proposed by the present invention is shown as follows.

[0065] (1) Replace the SPPF structure with the SPPFCSPC structure

[0066] SPPFCSPC

[0067] (Spatial Pyramid Pooling for Convolutional Neural Networks for Semantic Segmentation) is a deep learning model for semantic segmentation that combines the advantages of spatial pyramid pooling (SPP) and fully convolutional network (FCN).

[0068] SPPFCSPC has four main features. First is spatial pyramid pooling. By pooling the feature maps at different scales, SPP can capture information at multiple scales, enhancing the model's adaptability to different object sizes. Second is the fully convolutional network. SPPFCSPC applies the fully convolutional network to the segmentation task, avoiding the fully connected layers in traditional convolutional neural networks (CNNs), enabling the model to accept input images of any size. Third is feature fusion. The model performs feature fusion between different levels to improve the accuracy of segmentation. By combining low-level detailed features and high-level semantic features, the quality of the segmentation results is improved. In contrast, the SPPF (Spatial Pyramid Pooling Fast) structure has some drawbacks: Although SPPF aims to preserve spatial information, detail information may still be lost during the pooling process; when dealing with small objects, SPPF may not be able to fully capture the detailed features, resulting in its performance not meeting expectations in specific application scenarios. Therefore, the present invention replaces the SPPF structure.

[0069] While maintaining high accuracy, the computational efficiency of SPPFCSPC is also optimized, making it suitable for real-time application scenarios. By using this model, the ability to recognize and segment objects in complex scenes can be effectively improved. Next Figure 5 The network structure diagram of SPPFCSPC is as follows:

[0070] (2) Introduce the SS-Conv module

[0071] The present invention proposes the SS-Conv module, as Figure 6As shown, SS-Conv is an advanced convolutional neural network (CNN) architecture designed to enhance the processing ability for low-resolution images and small objects. It combines the advantages of the SPDConv module and the SimAM attention module to form a more efficient and flexible feature extraction and representation mechanism. The design of the SPDConv module takes into account the information loss problem often faced by traditional convolutional neural networks when processing low-resolution images. By replacing stride convolution and pooling layers with a spatial-to-depth conversion method, SPDConv can effectively preserve the detailed information in the image. This conversion not only increases the depth of the feature map but also enables the model to capture the features of small objects more precisely, thus achieving better results in complex scenarios. In addition, by retaining spatial information, SPDConv reduces the feature loss caused by traditional operations, thereby enhancing the robustness and adaptability of the model. The SimAM module focuses on optimizing the attention mechanism of the model, as Figure 7 shown. It infers the importance of each neuron through a simple and effective method without adding extra parameters. This not only reduces the complexity of the model but also improves the computational efficiency, making SimAM particularly suitable for use in resource-constrained environments. SimAM further enhances the feature expression ability by optimizing the energy function and assigning weights to the feature map.

[0072] The combination of the SS-Conv module makes the model perform even better when processing complex images. Through the depth feature extraction of SPDConv and the efficient attention mechanism of SimAM, SS-Conv can effectively improve the recognition ability for small objects and low-resolution images while keeping the computation lightweight. The design of this module not only improves the overall performance of the model but also demonstrates excellent adaptability in practical applications. In contrast, ordinary convolution (Conv) has some defects: when using stride convolution and pooling layers, fine-grained information is easily lost; traditional convolutional layers are not sensitive enough to the spatial transformation of the input image and are difficult to effectively capture features at different scales; traditional convolutional layers lack the ability to dynamically adjust the importance of features and cannot effectively focus on key regions, thus affecting the performance of the model in complex scenarios. Therefore, in the present invention, the Conv structure except for the 0th layer is replaced to improve the performance and robustness of the model in various tasks.

[0073] (3) Loss function design: Replace CioU with MPDioU

[0074] MPDioU (Multi-Path Distance Intersection over Union) is an evaluation metric for object detection and segmentation tasks, aiming to more accurately measure the performance of the model when dealing with different objects or regions. By introducing the distance information of multiple points, MPDIoU can comprehensively evaluate the relationship between the predicted bounding box and the ground truth box, improving the sensitivity to the object position. Due to considering the distances of multiple points, MPDIoU usually performs more robustly when dealing with complex scenarios (such as occlusions or partially visible objects) and can better handle the shape changes of objects. MPDIoU performs well in small object detection because it reduces the dependence on a single box through the evaluation of multiple points, enabling the model to better capture the details of small objects. MPDIoU is more likely to converge to a better solution during the training process, especially when there is noise in the dataset. In contrast, the calculation of Ciou loss is relatively complex, requiring consideration of multiple factors such as the distance between the centers of the boxes and the aspect ratio, increasing the computational overhead, especially when dealing with a large number of objects. Although CIoU considers the aspect ratio, in some cases, this sensitivity may lead to unstable loss, especially in scenarios where the object shapes vary greatly. CIoU may easily fall into a local optimum during the training process, especially when the data annotation is inaccurate or there is noise, resulting in the model performance not reaching the best. When dealing with small objects, CIoU may not be able to capture sufficient detail information, affecting the detection effect. Therefore, the present invention replaces the loss function with MPDiou, and its calculation formula is as follows:

[0075]

[0076]

[0077] L MPDIoU = 1 - MPDIoU(10)

[0078] In the formula are the position coordinates of the predicted bounding box, are the position coordinates of the ground truth box, and w, h are the width and height of the input image.

[0079] 2 Results and Analysis

[0080] The present invention introduces the experimental settings and specific details, and conducts in-depth analysis and discussion of the experimental results.

[0081] 2.1 Experimental Details

[0082] The experimental environment is a 64-bit Ubuntu operating system, an AMD R55600x CPU, 16G memory, and an AMD RX6700 10G GPU. The version of PyTorch is 2.3.0. In terms of training strategy, the hyperparameters are determined to be 200 epochs, 16 batch sizes, 1e-2 initial learning rate, 0.0005 weight decay, 0.937 momentum, and SGD optimizer. The dataset is divided into training, validation, and test sets in a ratio of 7:2:1. This study uses precision, recall, average precision (mAP), computational effort, parameter quantity, and model size to evaluate model performance. Precision, recall, and average precision are calculated as follows:

[0083]

[0084] In the formula, TP (True Positive) is the true positive, which is predicted to be 1 and is actually 1, indicating the number of diseases correctly identified by the algorithm. FP (False Positive) is the false positive, which is predicted to be 1 and is actually 0, indicating the number of diseases incorrectly identified by the algorithm. FN (False Negative) is the false negative, which is predicted to be 0 and is actually 1, indicating the number of unidentified diseases. AP stands for average precision, mAP is the mean of AP of different categories, and C is the number of categories, and C=9 in the present invention.

[0085] 2.2 Ablation Experiment

[0086] In this study, the YOLOv8s model was improved and optimized to improve its target detection performance. In order to verify the effectiveness of the improved module in improving network performance, an ablation test was conducted, and the results are shown in Table 2. The performance indicator data changes during the model training process are shown in Table 2. Figure 8 As shown, it indicates that the model has converged.

[0087] Table 2 Ablation study on the dataset

[0088]

[0089] As can be seen from Table 2, after replacing SPPFCPSPC, the model detection accuracy is improved by 1.8%. The SPPFCSPC structure brings rich gradient flow information to the model, which has better computational efficiency while maintaining high performance, but increases the number of model parameters and computation. In addition, SPDConv also rarely increases the number of model computations and parameters, but the model performance is improved, and the model detection accuracy is improved by 0.7%.

[0090] MPDIoU considers the mutual relationship between objects during the regression process of the target bounding box, can more accurately measure the matching degree of the target box, improve the prediction accuracy and precision of the model for the target bounding box, and further enhance the performance of the model. To sum up, this structure is carefully designed to improve detection accuracy. Compared with the baseline model, the accuracy is improved by 3.9%, and the number of parameters and computational volume are increased by 58% and 13% respectively. Through the design and comparative analysis of ablation experiments, the effectiveness of the improved method proposed in the present invention is verified.

[0091] 2.3 Comparative Experiments

[0092] To further verify the superiority of the SSM-YOLO model, the present invention compares it with classic mainstream object detection models such as YOLOv5s, YOLOv6s, and YOLOv7. The results are shown in Table 3.

[0093] Table 3 Verification Results of Comparative Experiments of Different Algorithms

[0094]

[0095] Compared with YOLOv6 and YOLOv7, the accuracy is improved by 5.7% and 4.6% respectively, and the model complexity is also lower than theirs. Compared with YOLOv8s, SSM-YOLO has only 42.7% of its number of parameters and 33.5% of its computational volume, and the accuracy is 2.7% higher than it. SSD has the lowest GFLOPs, but performs poorly in other metrics. While Faster-RCNN has higher complexity than other models, and the lower Recall value also reflects the limitations of its detection performance. In addition, SSM-YOLO has a lower training time-consuming and excellent inference speed, achieving a higher FPS performance, as Figure 9 shown. The comparison with these advanced models further verifies the superiority of the SSM-YOLO model, and the lower model complexity also makes it more suitable for deployment and application in terminal devices.

[0096] 2.4 Visualization Comparative Verification

[0097] Visualization provides a more intuitive view of the detection effectiveness of the algorithm, and the position, size, and category information of the detected objects can be observed to help better evaluate the accuracy and efficiency of the algorithm and debug and optimize it. Figure 10 The visualization results of the model detection are shown. As can be seen from the figure, the algorithm proposed in the present invention has better detection effect. In Figure 10 -b brown spot disease detection, the baseline model YOLOv8n missed detections and did not completely detect the lesions, while SSM-YOLO detected the missed targets. In Figure 10-c, SSM-YOLO can detect a slightly blurred small disease spot in the lower right corner, while the baseline model YOLOv8n fails to recognize it. In Figure 10 -a and 10-d, SSM-YOLO shows a higher confidence level in detection performance.

[0098] In addition, to deeply explore the interpretability of the model, the present invention uses the Grad-CAM method to visualize its class activation map, as Figure 11 shown. Through this method, we can intuitively observe the attention distribution of the model during object detection and the degree of attention the model pays to different regions.

[0099] By comparing and analyzing the Grad-CAM heatmaps, the SSM-YOLO model shows more concentrated and relevant activations in the key regions of the task, and can accurately identify and locate the disease regions. In contrast, the baseline model shows more scattered or inaccurate attention points, indicating that the SSM-YOLO model demonstrates better performance in understanding the task.

[0100] 3 Conclusion

[0101] The present invention collects and constructs a tea leaf disease dataset in a complex scenario, which is more challenging than the dataset with a single disease in a simple background. In terms of the algorithm, through targeted improvement and optimization of YOLOv8s, an SSM-YOLO model is proposed: replacing part of the SPPF structure with SPPFCSPC, effectively improving the recognition and segmentation ability of objects in complex scenarios; replacing most of the Conv with the SS-Conv structure to enhance the performance of processing low-resolution images and small objects, improving the performance and robustness of the model in complex scenarios, and helping the model better understand and process the key information of the input data; adopting the MPDIoU loss function to improve the localization and recognition ability of small disease spot regions. After sufficient experimental verification and visualization analysis, and comprehensive comparison with mainstream models and baseline models, the proposed SSM-YOLO model significantly improves the object detection performance in complex scenarios, including the detection of occluded targets and small targets. The experimental results show that the proposed model has achieved significant improvements in multiple key indicators. It also has obvious advantages compared with the baseline model and other models. This study proposes an accurate and real-time tea leaf disease detection method, which provides an efficient disease detection technology support for the tea industry. Subsequently, it can be further improved according to actual application requirements and embedded in the terminal device system for application.

[0102] The above embodiments are only explanations of the present invention and are not limitations thereof. Those skilled in the art can make modifications without creative contributions to the embodiments after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A tea disease detection method based on improved YOLOv8, characterized in that, Including: Constructing a tea disease dataset for complex scenarios: By combining web crawling and manual screening, a number of high-quality RGB images are collected, covering 8 disease categories including tea anthracnose, tea blister blight, tea zonate leaf spot, tea phyllosticta leaf blight, tea red rust alga disease, tea white star disease, tea shoot blight and tea sooty mold, as well as a control category of healthy leaves; Using quantitative indicators such as color histogram entropy, color proportion distribution, and edge density to evaluate the complexity of the image background, ensuring that the resolution of each image is not less than 2000×2000 pixels, and the proportion of the disease area exceeds 15%; Designing the SSM-YOLO detection model: Feature extraction network: The SS-Conv module is used to replace the traditional convolutional layer, where the SS-Conv integrates spatial pyramid decomposition convolution and an adaptive attention module, enhancing the ability to extract features of low-resolution images and small targets while maintaining computational efficiency; Feature fusion structure: The SPPFCSPC module is introduced to replace part of the spatial pyramid pooling (SPPF) layer. This module realizes the efficient fusion and enhancement of multi-scale features through cascading spatial pyramid pooling, a fully convolutional network, and a channel attention mechanism; Loss function optimization: An improved multi-point distance intersection over union (MPDIoU) loss function is adopted. The multi-point Euclidean distance, overlapping area, and union area between the predicted box and the ground truth box are calculated through formulas, and the influence of target scale changes on the positioning accuracy is considered to improve the detection accuracy of the model for disease targets of different scales. The calculation formula is as follows: L MPDIoU = 1 - MPDIoU(10) In the formula are the position coordinates of the predicted bounding box, are the position coordinates of the ground truth bounding box, and w is the width and height of the input image.

2. The tea disease detection method based on improved YOLOv8 according to claim 1, characterized in that: The specific implementation method of the SS-Conv module includes: SPDConv component: The spatial-to-depth conversion method is used to replace the stride convolution and pooling layers, retaining the detailed information of the image; SimAM attention mechanism: The feature map weights are generated by optimizing the energy function to enhance the feature expression ability; Module integration: The output of the SPDConv is used as the input of the SimAM to form an end-to-end feature enhancement channel.

3. The tea disease detection method based on improved YOLOv8 according to claim 1 is characterized in that: The calculation process of the MPDIoU loss function is: Distance calculation: The Euclidean distances d1 at the upper left corner and d2 at the lower right corner of the predicted box and the ground truth box are calculated according to formulas (1) and (2) respectively; Boundary value determination: The maximum values of the predicted box and the ground truth box in the x-axis and y-axis directions are taken as the upper left corner coordinates of the intersection area through formulas (5) and (6), and the minimum values are taken as the lower right corner coordinates through formulas (7) and (8); Overlapping area calculation: The area of the intersection area is calculated using formula (9), and the IoU is calculated in combination with the union area of the predicted box and the ground truth box; Loss value calculation: The MPDIoU loss value is calculated according to formula (10), that is, 1 - IoU, and the calculation result of IoU is corrected by introducing multi-point distance information.

4. A tea disease detection method based on improved YOLOv8 according to claim 1, characterized in that: The construction method of the complex scenario tea disease dataset also includes: Data cleaning: Images with a resolution lower than 2000×2000 or unclear disease characteristics are removed; Semi-automatic annotation: The labelImg tool is used to annotate the selected images to generate XML files in PascalVOC format; Data augmentation: Operations such as rotation, flipping, and brightness adjustment are performed on the annotated dataset to improve the generalization ability of the model.

Citation Information

Cited By

  • Multi-scale image segmentation and damage assessment method for surface cracks of bridge structure

    CN121392622A