Method for detecting plant diseases and insect pests of tomato leaves based on improved YOLOv8
By using GhostNet and RFA attention mechanisms in the detection of pests and diseases of tomato leaves, combined with the Wise IOU loss function, the problem of high computing burden and poor real-time performance of traditional models is solved, and efficient and accurate pest detection is achieved.
Patent Information
- Application Number
- CN202510074439.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In the prior art, in the detection of tomato leaf pests and diseases, traditional object detection models rely on deep convolutional neural networks, resulting in increased computing burden and reduced real-time performance, making it difficult to perform well in resource-constrained environments and agricultural scenarios that require rapid response.
The GhostNet lightweight network is used to combine the RFA attention mechanism to reduce the number of parameters and calculation amounts, and to introduce attention mechanisms in the feature fusion process of the feature pyramid network to enhance the model's ability to pay attention to features. At the same time, the replacement loss function is Wise IOU to improve detection accuracy and optimize the recognition ability of small objects.
It achieves the performance and real-time performance of YOLOv8 in small object detection while maintaining high accuracy, and can perform target detection more quickly, improving the accuracy and recall of pest and disease detection.
Smart Images

Figure CN120070960A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pest and disease detection, and particularly relates to a method for detecting tomato leaf pests and diseases based on improved YOLOv8. Background Art
[0002] Tomato is a widely cultivated cash crop, and its growth is vulnerable to various pests and diseases, which affect yield, quality and safety. Although traditional algorithms based on feature extraction and classifiers are easy to implement and perform well on small-scale data sets, they have deficiencies in dealing with small targets and complex background features and are difficult to mine potential information in the data. Convolutional neural networks (CNNs), such as YOLO and Faster R-CNN, have become the mainstream detection means, significantly improving the detection accuracy and the ability to process diverse images. However, in the detection of tomato leaf pests and diseases, traditional object detection models rely on relatively deep convolutional neural networks (such as VGG, ResNet) to improve accuracy, but increase the computational burden due to the increase in the number of layers and parameters, reducing real-time performance, and performing poorly in resource-constrained environments and agricultural scenarios that require quick responses. YOLOv8 is an improvement over YOLOv5 and mainly consists of a backbone network, a neck and a head. The backbone network uses CSPNet to extract features in stages to achieve efficient expression; the neck combines the advantages of FPN and PAN to effectively fuse multi-scale features and improve the detection effect of small targets; the head uses decoupling technology to separate classification and detection, and the loss function uses positive and negative sample matching. The features extracted by its backbone network are fused by the neck to form a multi-scale feature representation, and finally the head completes object detection, which has more advantages in tomato pest and disease detection.
[0003] This design has significantly improved both the speed and accuracy of YOLOv8, but there are still certain limitations in small target detection. At the same time, due to the increase in model complexity, its real-time performance may be affected in some application scenarios, resulting in a processing speed lower than expected. Therefore, in order to improve the performance and real-time performance of YOLOv8 in small target detection, it needs to be improved.
[0004] Therefore, the present invention proposes a method for detecting tomato leaf pests and diseases based on improved YOLOv8. Summary of the Invention
[0005] The present invention provides a method for detecting tomato leaf diseases and pests based on improved YOLOv8. By using the GhostNet lightweight network through its efficient feature extraction mechanism, the number of parameters and the amount of computation are reduced, enabling YOLOv8 to perform object detection more quickly while maintaining high accuracy. The attention mechanism RFA is introduced to enhance the model's ability to focus on features; it is introduced during the feature fusion process of the feature pyramid network, enabling the model to adaptively focus on important feature regions and enhancing the detection ability for small targets and complex backgrounds. And the loss function is replaced with Wise IOU to improve the detection accuracy and optimize the recognition ability for small objects, better handle the bounding box regression problem in object detection, and improve the accuracy and recall rate of object detection. Thus, more accurate detection of diseases and pests is achieved.
[0006] The present invention provides a method for detecting tomato leaf diseases and pests based on improved YOLOv8, including:
[0007] S1: Construct a tomato leaf image dataset containing each type of disease and pest;
[0008] S2: Integrate GhostNet and the RFA attention mechanism, modify the loss function to Wise IOU, and construct an improved YOLOv8 model;
[0009] S3: Configure training hyperparameters and update the model weights based on the tomato leaf image dataset to obtain a trained improved YOLOv8 model;
[0010] S4: Perform disease and pest detection on the tomato leaf image to be detected based on the trained improved YOLOv8 model to obtain the disease and pest detection results of the tomato leaf image to be detected.
[0011] Preferably, for the method for detecting tomato leaf diseases and pests based on improved YOLOv8, S2: Integrate GhostNet and the RFA attention mechanism, modify the loss function to Wise IOU, and construct an improved YOLOv8 model, including:
[0012] Load the pre-trained weights of YOLOv8, integrate the RFA attention mechanism into the feature extraction network of YOLOv8, replace the traditional convolution in the neck structure with GhostConv, and replace the feature fusion module in the neck structure with C2fGhost. Then replace the loss function with Wise IOU to construct an improved YOLOv8 model.
[0013] Preferably, for the method for detecting tomato leaf diseases and pests based on improved YOLOv8, integrating the RFA attention mechanism into the feature extraction network of YOLOv8 includes:
[0014] The RFA attention mechanism is introduced during the feature fusion process of the Feature Pyramid Network. The RFAConv has a 3×3 convolutional kernel and uses Group Conv to extract receptive field features. It aggregates information through AvgPool, conducts information interaction through 1×1 group convolution operations, and emphasizes feature importance through softmax. The C2f_RFAConv contains two convolutional layers and multiple Bottleneck_RFAConv modules. The Bottleneck_RFAConv module contains two convolutional layers, where the first convolutional layer reduces the number of input channels, and the second convolutional layer processes using RFAConv and supports shortcut connections.
[0015] Preferably, for the tomato leaf pest and disease detection method based on the improved YOLOv8, the formula for Wise IOU is:
[0016]
[0017] L WIoU = εR WIoU
[0018] In the formula, R WIoU is the degree of overlap between the anchor box and the target box, exp() is the exponential function with the natural constant e as the base, and the value of the natural constant e is 2.71828. x is the abscissa of the center point of the anchor box, y is the ordinate of the center point of the anchor box, x gt is the abscissa of the center point of the target box, y gt is the ordinate of the center point of the target box, W g is the width of the target box, H g is the height of the target box, L WIoU is the output loss value of Wise IOU, and ε is the weighting factor.
[0019] Preferably, for the tomato leaf pest and disease detection method based on the improved YOLOv8, S3: Configure training hyperparameters and update the model weights based on the tomato leaf image dataset to obtain a trained improved YOLOv8 model, including:
[0020] Divide the tomato leaf image dataset into a training set, a validation set, and a test set based on the division ratio of 8:1:1;
[0021] Perform data augmentation on the training set to obtain an optimized training set, and perform standardization or normalization operations on the test set and the validation set to obtain a standard test set and a standard validation set;
[0022] Configure the training hyperparameters, set the batch size to 32 and the number of iterations to 300 rounds, and use the optimized training set to train the improved YOLOv8 model. In each training iteration, calculate the loss and update the model weights, and monitor the performance metrics on the standard test set until the latest improved YOLOv8 model is obtained;
[0023] Calculate the recognition results of the currently obtained latest improved YOLOv8 model on the test set and the true labels of the test set, draw a confusion matrix, and calculate the accuracy and loss values of the currently obtained latest improved YOLOv8 model on the test set based on the confusion matrix as the model evaluation results of the currently obtained latest improved YOLOv8 model;
[0024] Based on whether the model evaluation results of the currently obtained latest improved YOLOv8 model meet the requirements, if so, regard the improved YOLOv8 model as the trained improved YOLOv8 model, otherwise, continue to train and test verify the currently obtained latest improved YOLOv8 model until a trained improved YOLOv8 model that meets the requirements is obtained.
[0025] Preferably, for the tomato leaf pest and disease detection method based on improved YOLOv8, S4: Perform pest and disease detection on the tomato leaf image to be detected based on the trained improved YOLOv8 model, and obtain the pest and disease detection results of the tomato leaf image to be detected, including:
[0026] Perform first-level super-resolution reconstruction processing on the tomato leaf image to be detected and divide it into S×S grids. Input each grid into the trained improved YOLOv8 model, predict B bounding boxes in each grid, and calculate the center coordinates, width, and height of each bounding box;
[0027] x = σ(t x ) + c x
[0028] y = σ(t y ) + c y
[0029]
[0030] In the formula, x is the abscissa value of the center coordinate of the bounding box, y is the ordinate value of the center coordinate of the bounding box, σ() is the sigmoid function, t x is the horizontal offset of the bounding box center relative to the grid it is in, t y is the vertical offset of the bounding box center relative to the grid it is in, c x is the abscissa value of the center point of the grid corresponding to the currently calculated bounding box, c yis the vertical coordinate value of the center point of the grid corresponding to the currently calculated bounding box, w is the width of the bounding box, p w is the width of the prior box, e is the natural constant with a value of 2.71828, t w is the relative change in the width of the bounding box, h is the height of the bounding box, p h is the height of the prior box, t h is the relative adjustment of the height of the bounding box;
[0031] Determine the intersection over union (IoU) between the predicted box and the ground truth box based on the center coordinates, width, and height of each bounding box, and calculate the confidence of each bounding box based on the IoU between the predicted box and the ground truth box:
[0032] C = P(object)·IoU truth
[0033] In the formula, C is the confidence of the bounding box, IoU truth is the intersection over union between the predicted box and the ground truth box;
[0034] Select the final target bounding boxes based on the non-maximum suppression algorithm;
[0035] Predict the probability distribution of each target bounding box for each type of pest and disease based on the confidence:
[0036]
[0037] In the formula, P(i|box) is the probability of the target bounding box for the i-th type of pest and disease, e is the natural constant with a value of 2.71828, z i is the original predicted score of the target bounding box for the i-th type of pest and disease, z j is the original predicted score of the target bounding box for the j-th type of pest and disease, C is the total number of pest and disease types;
[0038] Obtain the pest and disease detection results of the tomato leaf image to be detected based on the probability distribution of all target bounding boxes in all grids in the tomato leaf image to be detected for each type of pest and disease.
[0039] Preferably, for the tomato leaf pest and disease detection method based on the improved YOLOv8, obtain the pest and disease detection results of the tomato leaf image to be detected based on the probability distribution of all target bounding boxes in all grids in the tomato leaf image to be detected for each type of pest and disease, including:
[0040] Perform pest and disease detection annotation on the tomato leaf image to be detected based on the probability distribution of all target bounding boxes in all grids of the tomato leaf image to be detected for each type of pest and disease, and obtain the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection for the tomato leaf image to be detected;
[0041] Based on the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection of the tomato leaf image to be detected, obtain the pest and disease detection result of the tomato leaf image to be detected.
[0042] Preferably, for the tomato leaf pest and disease detection method based on improved YOLOv8, based on the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection of the tomato leaf image to be detected, obtain the pest and disease detection result of the tomato leaf image to be detected, including:
[0043] Based on the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection of the tomato leaf image to be detected, calculate the pest and disease detection evaluation value of the tomato leaf image to be detected;
[0044] Based on the pest and disease detection evaluation value of the tomato leaf image to be detected, obtain the pest and disease detection result of the tomato leaf image to be detected.
[0045] Preferably, for the tomato leaf pest and disease detection method based on improved YOLOv8, based on the coordinate information of all minimum target pest and disease detection boxes in the super-resolution image of the tomato leaf to be detected and all non-minimum target pest and disease detection boxes in the corresponding non-minimum target pest and disease detection pure annotation result, calculate the pest and disease detection evaluation value of the tomato leaf image to be detected, including:
[0046] Based on the coordinate information of all minimum target pest and disease detection boxes in the super-resolution image of the tomato leaf to be detected and all non-minimum target pest and disease detection boxes in the corresponding non-minimum target pest and disease detection pure annotation result, determine the shortest distance between each minimum target pest and disease detection box and the corresponding non-minimum target pest and disease detection box in the super-resolution image of the tomato leaf to be detected;
[0047] Based on the shortest distance between each minimum target pest and disease detection box and the corresponding non-minimum target pest and disease detection box in the super-resolution image of the tomato leaf to be detected, calculate the pest and disease detection evaluation value of the tomato leaf image to be detected:
[0048]
[0049] Wherein, P is the pest and disease detection and evaluation value of the tomato leaf image to be detected, n is the total number of extremely small target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected, m is the total number of non-extremely small target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected, max(n, m) is the maximum value between the total number of extremely small target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected and the total number of non-extremely small target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected, and P m is the standard similarity between the total number of extremely small target pest and disease detection boxes and the total number of non-extremely small target pest and disease detection boxes, and l ij is the shortest distance between the i-th extremely small target pest and disease detection box and the j-th non-extremely small target pest and disease detection box, and P l is the standard distribution similarity between all extremely small target pest and disease detection boxes and all non-extremely small target pest and disease detection boxes.
[0050] Preferably, for the tomato leaf pest and disease detection method based on the improved YOLOv8, based on the pest and disease detection and evaluation value of the tomato leaf image to be detected, the pest and disease detection result of the tomato leaf image to be detected is obtained, including:
[0051] If the pest and disease detection and evaluation value of the tomato leaf image to be detected is not less than the preset evaluation threshold, then all the extremely small target pest and disease detection boxes and all the non-extremely small target pest and disease detection boxes in the super-resolution image of the tomato leaf to be detected are labeled to obtain the pest and disease detection result of the tomato leaf image to be detected;
[0052] If the pest and disease detection and evaluation value of the tomato leaf image to be detected is less than the preset evaluation threshold, then the tomato leaf image to be detected is subjected to second-level super-resolution reconstruction processing and pest and disease detection until the newly obtained extremely small target pest and disease detection evaluation value is not less than the preset evaluation threshold, and then all the extremely small target pest and disease detection boxes and all the non-extremely small target pest and disease detection boxes in the newly obtained super-resolution image of the tomato leaf to be detected are labeled to obtain the pest and disease detection result of the tomato leaf image to be detected.
[0053] The beneficial effects of the present invention compared with the prior art are as follows: This method uses the GhostNet lightweight network, through its efficient feature extraction mechanism, to reduce the number of parameters and the amount of computation, enabling YOLOv8 to perform object detection more quickly while maintaining high accuracy. The attention mechanism RFA is introduced to enhance the model's ability to focus on features; it is introduced during the feature fusion process of the feature pyramid network, enabling the model to adaptively focus on important feature regions and enhancing the detection ability for small objects and complex backgrounds. And the loss function is replaced with Wise IOU to improve the detection accuracy and optimize the recognition ability for small objects, better handle the bounding box regression problem in object detection, and improve the accuracy and recall rate of object detection. Thus, more accurate pest and disease detection is achieved.
[0054] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0055] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0056] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0057] Figure 1 is the flow chart of the tomato leaf pest and disease detection method based on the improved YOLOv8 in the embodiment of the present invention;
[0058] Figure 2 is the implementation flow chart of the tomato leaf pest and disease detection based on the improved YOLOv8 algorithm in the embodiment of the present invention;
[0059] Figure 3 is the network structure diagram of a YOLOv8 algorithm in the embodiment of the present invention;
[0060] Figure 4 is the schematic diagram of GhostNet in the embodiment of the present invention;
[0061] Figure 5 is the structural schematic diagram of Ghostconv in the embodiment of the present invention;
[0062] Figure 6 is the structural schematic diagram of GhostBottleneck in the embodiment of the present invention;
[0063] Figure 7Schematic diagram of the structure of C2fGhost in the embodiments of the present invention;
[0064] Figure 8 Schematic diagram of obtaining the received field spatial features by transforming the spatial features in the embodiments of the present invention;
[0065] Figure 9 Schematic diagram of the structure of RFAConv in the embodiments of the present invention;
[0066] Figure 10 Schematic diagram of the network structure of the improved YOLOv8 algorithm in the embodiments of the present invention. Detailed implementation manners
[0067] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0068] Embodiment 1:
[0069] The present invention provides a method for detecting tomato leaf diseases and pests based on improved YOLOv8. Referring to Figure 1 and Figure 2 , it includes:
[0070] S1: Construct a tomato leaf image dataset containing each type of disease and pest;
[0071] S2: Integrate GhostNet and RFA attention mechanism, modify the loss function to Wise IOU, and construct an improved YOLOv8 model;
[0072] S3: Configure training hyperparameters and update the model weights based on the tomato leaf image dataset to obtain a trained improved YOLOv8 model;
[0073] S4: Perform disease and pest detection on the tomato leaf image to be detected based on the trained improved YOLOv8 model, and obtain the disease and pest detection results of the tomato leaf image to be detected.
[0074] In this embodiment, the network structure diagram of the YOLOv8 algorithm is as shown in Figure 3 . The features extracted by the backbone network are fused through the neck to form a rich multi-scale feature representation, and finally object detection is performed by the head. This design significantly improves the speed and accuracy of YOLOv8, but there are still certain limitations in small object detection. At the same time, due to the increase in model complexity, its real-time performance may be affected in some application scenarios, resulting in a processing speed lower than expected. Therefore, in order to improve the performance and real-time performance of YOLOv8 in small object detection, it needs to be improved.
[0075] In the network structure of YOLOv8, two convolutional kernel sizes of 1×1 and 3×3 are adopted. In the Backbone main network, in order to enhance the network's feature learning ability while maintaining the original gradient path, YOLOv8 introduces a module called the CSP (Cross Stage Partial) module, which improves the network's expressiveness through cross-stage feature fusion. In addition, YOLOv8 also combines other modules (such as convolutional and activation functions) to further enhance the feature extraction ability. After entering the Head part, YOLOv8 adopts the ideas of FPN and PAN, performs multiple top-down and bottom-up fusions and re-extractions on the generated multiple feature layers, and finally generates feature maps of different sizes, namely 20×20, 40×40, and 80×80, for detecting large, medium, and small targets in the image.
[0076] For the generated feature maps, they are first divided into S×S grids. Each grid is responsible for detecting objects within its area. For each grid, YOLOv8 predicts B bounding boxes, and each box contains five parameters: the center coordinates (x,y), width w, height h, and confidence C. Specifically, the center coordinates (x,y) of the box are the offsets relative to the grid position. In addition, YOLOv8 also generates a probability distribution related to the target class for each predicted box.
[0077] Finally, YOLOv8 uses the non-maximum suppression (NMS) algorithm to filter the final target bounding boxes. The process of NMS is to remove those boxes with too high overlap with the boxes with higher confidence by setting a threshold, ensuring that the finally output bounding boxes are unique and have the highest confidence. Through these steps, YOLOv8 can efficiently process feature maps, perform object classification and localization, and thus achieve accurate object detection in various complex scenarios.
[0078] In this embodiment, the system of the YOLOv8 tomato leaf pest and disease detection method improved based on the GhostNet network and the RFA module includes the following modules:
[0079] Image preprocessing module: This module is responsible for capturing tomato leaf images and performing necessary preprocessing operations to reduce the impact of external factors on the image quality.
[0080] Image recognition module, which contains 4 sub-modules, including the following modules:
[0081] Dataset construction module: Construct a tomato leaf image dataset containing different pest and disease types to ensure the comprehensiveness and representativeness of the data.
[0082] Improved YOLOv8 Detection Model Module: Integrate GhostNet and RFA (Residual Feature Attention) attention mechanism, modify the loss function to Wise IOU, and construct an efficient YOLOv8 model to enhance the model's ability to focus on features, especially in the detection of small objects.
[0083] Model Training Module: Train the improved YOLOv8 model to improve the model's detection performance by optimizing parameters.
[0084] Image Detection Module: Use the trained model to detect pests and diseases in tomato leaf images and identify different types of pests and diseases.
[0085] The specific implementation process is as follows:
[0086] Step 1: Data Preparation:
[0087] Use the publicly available tomato leaf pest and disease dataset, divide the dataset into training set, validation set and test set, and the division ratio is 8:1:1. The configuration uses an i7-12700H CPU, a 64-bit Windows 11 operating system, and the GPU configuration is an NVIDIA RTX 3050Ti. At the same time, the experimental environment includes a GPU server configured with an NVIDIA Tesla A40 48GB 300W GPU card.
[0088] Step 2: Data Preprocessing:
[0089] Perform data augmentation on the training set to expand the training data, including operations such as random rotation, scaling, flipping, and brightness adjustment to increase the generalization ability of the model. Perform standardization or normalization operations on the validation set and test set to ensure that the pixel values of the images are within the same scale range.
[0090] Step 3: Model Construction:
[0091] Load the pre-trained weights of YOLOv8, integrate the RFA attention mechanism into the feature extraction network of YOLOv8, use GhostNet in the feature fusion network, and then replace the loss function with Wise IOU to construct an improved YOLOv8 model to enhance the ability to identify key features.
[0092] Step 4: Model Training:
[0093] Configure the training hyperparameters, set the batch size to 32 and the number of iterations to 300 rounds, and start training the model using the training set data. In each training iteration, calculate the loss and update the model weights, and at the same time monitor the performance metrics on the validation set to prevent overfitting.
[0094] Step 5: Model Evaluation:
[0095] Use the test set to conduct a final evaluation of the model, calculate metrics such as the accuracy and F1-score of the model on the test set, and plot a confusion matrix for visualization to evaluate the performance of the model in practical applications.
[0096] The specific flowchart is as Figure 2 .
[0097] In this embodiment, the tomato leaf image dataset: Create a set that contains various tomato leaf images and clearly labels the type of pests and diseases corresponding to each image. It is used for training and testing the improved YOLOv8 model to achieve the detection and analysis of tomato leaf pests and diseases. Example: For instance, it consists of 1000 tomato leaf images in different pest and disease states, covering the annotation of various pest and disease types such as leaf mold and fusarium wilt.
[0098] In this embodiment, the trained improved YOLOv8 model: An improved YOLOv8 model with good performance obtained through a series of training and optimization processes. It performs excellently on the training set, validation set, and test set, can accurately detect tomato leaf pests and diseases, and has high accuracy and low loss values. Example: An improved YOLOv8 model that after multiple parameter adjustments and trainings, has a pest and disease detection accuracy of 92% on the test set and a loss value of only 0.15.
[0099] In this embodiment, the pest and disease detection result of the tomato leaf image to be detected: The conclusion obtained after using the trained improved YOLOv8 model to process and analyze the tomato leaf image to be detected. It includes detailed information such as the detected type of pests and diseases, location (annotated by a bounding box), and the probability of each type of pest and disease. Example: The detection result shows that the leaf to be detected has anthracnose with a probability of 0.75, and the bounding box accurately marks the range of the disease on the leaf.
[0100] The beneficial effects of the above technologies are as follows: Refer to Figure 10 , this method uses the GhostNet lightweight network. Through its efficient feature extraction mechanism, it reduces the number of parameters and the amount of calculation, enabling YOLOv8 to perform target detection more quickly while maintaining high accuracy. The attention mechanism RFA is introduced to enhance the model's ability to focus on features; it is introduced during the feature fusion process of the feature pyramid network, enabling the model to adaptively focus on important feature regions and enhancing the detection ability for small targets and complex backgrounds. And the loss function is replaced with Wise IOU to improve the detection accuracy and optimize the recognition ability for small objects, and can better handle the bounding box regression problem in target detection, improving the accuracy and recall rate of target detection. Thus, more accurate pest and disease detection is achieved.
[0101] Example 2:
[0102] Based on the method for detecting tomato leaf diseases and pests improved from YOLOv8 in Example 1, S2: Incorporate the GhostNet and RFA attention mechanisms, modify the loss function to Wise IOU, and construct an improved YOLOv8 model, including:
[0103] Load the pre-trained weights of YOLOv8, integrate the RFA attention mechanism into the feature extraction network of YOLOv8, replace the traditional convolution in the neck structure with GhostConv, and replace the feature fusion module in the neck structure with C2fGhost. Then, replace the loss function with Wise IOU to construct an improved YOLOv8 model.
[0104] In this embodiment, GhostNet is a lightweight convolutional neural network architecture designed to reduce computational complexity and model parameters while maintaining high classification accuracy. As Figure 4 shown, the process of GhostNet mainly includes two key steps: feature extraction and feature generation. First, the network extracts the initial features of the input image through a standard convolutional layer. Then, the Ghost module generates additional feature maps using linear operations. Finally, the final feature maps are generated through a Concat operation, which enhances the network's expressive power while maintaining computational efficiency.
[0105] In Figure 5 Ghostconv is a convolutional module in the GhostNet network, specifically designed as a phased convolutional computing module. By parallelly executing feature extraction and cost-effective linear operations, it generates a large number of feature maps, thus demonstrating its computational efficiency. This process first generates half of the feature maps by GhostConv, using a convolutional kernel size that is half of the original convolution. Then, through a 5×5 convolutional kernel and a linear operation with a stride of 1, the remaining half of the feature maps are obtained. Finally, the two feature maps are merged through a concatenation operation to generate the complete feature maps.
[0106] Figure 6 Demonstrates the GhostBottleneck operation. First, the initial GhostConv layer is used as an expansion layer to increase the number of channels. Subsequently, regularization and the Sigmoid linear unit (SiLU) are applied to the feature maps. Next, the second GhostConv layer is used to reduce the number of channels of the output feature maps to match the input channels. Finally, the feature maps obtained in the previous stage are added for feature fusion.
[0107] Figure 7The C2fGhost module is shown, which replaces all bottleneck components in the original network's C2f module with GhostBottleneck units. This structure combines a cross-stage feature fusion strategy and a truncated gradient flow technique to enhance the diversity of learned features between different network layers, reduce the impact of redundant gradient information, and thus improve the learning ability.
[0108] The beneficial effects of the above technologies are as follows: Loading the YOLOv8 pre-trained weights can accelerate the training speed and improve the performance. Integrating the RFA attention mechanism enables the model to focus on key features and improve the detection accuracy. Replacing the traditional convolution in the neck structure with GhostConv reduces the computational complexity and parameters, and improves the computational efficiency; replacing the feature fusion module with C2fGhost enhances the feature diversity and improves the learning ability. Changing the loss function to Wise IOU optimizes the training and makes the prediction more accurate. These improvements make the model more accurate, efficient, and lightweight, which is beneficial for practical detection applications.
[0109] Example 3:
[0110] Based on the tomato leaf pest and disease detection method improved from YOLOv8 in Example 2, the RFA attention mechanism is integrated into the feature extraction network of YOLOv8, including:
[0111] The RFA attention mechanism is introduced during the feature fusion process of the feature pyramid network. And RFAConv has a 3×3 convolutional kernel and uses Group Conv to extract receptive field features, aggregates information through AvgPool, performs information interaction through 1×1 group convolution operations, and emphasizes the feature importance through softmax. C2f_RFAConv contains two convolutional layers and multiple Bottleneck_RFAConv modules. The Bottleneck_RFAConv module contains two convolutional layers, and the first convolutional layer reduces the number of input channels, and the second convolutional layer is processed using RFAConv and supports shortcut connections.
[0112] In this embodiment, in the detection of tomato leaf diseases and pests, RFA (Receptive Field Attention) addresses issues such as insufficient receptive fields, multi-scale feature fusion, and context information capture. By dynamically adjusting the receptive field, it enhances the ability to focus on features of different scales and adaptively extracts context information related to the target, thereby improving the accuracy and robustness of detection. RFA can be regarded as a lightweight plug-and-play module. The designed convolution operation (RFAConv) combines a spatial attention mechanism with convolution operations, optimizing the working mode of the convolution kernel, especially when dealing with spatial features within the receptive field. RFAConv focuses on the spatial features within the receptive field, enabling the network to more effectively understand and process local regions of the image, thus improving the accuracy of feature extraction. The designed C2f_RFAConv module can effectively integrate feature information from different levels through multi-scale feature fusion, enabling the model to accurately identify diseases and pests even in the presence of complex backgrounds or occlusions.
[0113] Receptive-field spatial features refer to the features in the region of the input image that each convolution operation can perceive. These features can include basic visual elements such as color, shape, and texture. Receptive-field spatial features are specifically designed for the convolution kernel and are dynamically generated according to the size of the kernel. As Figure 8 shown, taking a 3×3 convolution kernel as an example. "Spatial Feature" in the following figure refers to the original feature map. "Receptive-Field Spatial Feature" is the feature map transformed from the spatial feature, consisting of non-overlapping sliding windows. Each 3×3-sized window in the receptive-field spatial feature represents a receptive-field slider.
[0114] The receptive-field attention convolution (RFAConv) has a 3×3-sized convolution kernel, and the overall structure is as Figure 7 shown. In RFAConv, a fast method is adopted to extract receptive-field spatial features, namely Group Conv. As mentioned earlier, when using a 3×3 convolution kernel to extract features, each 3×3-sized window represents a receptive-field slider in the receptive-field spatial feature. After using Group Conv to extract the receptive-field features, the original features are mapped to new features. This method is fast and efficient.
[0115] For RFAConv, learning the attention map by interacting with receptive field feature information can enhance network performance. However, interacting with each receptive field feature may lead to additional computational overhead. Therefore, to minimize the computational overhead and the number of parameters, AvgPool is used to aggregate the global information of each receptive field feature. Then, a 1×1 grouped convolution operation is used for information interaction. Finally, softmax is used to emphasize the importance of each feature in the receptive field features. Generally, the calculation of RFA can be expressed as:
[0116] F = Softmax(g l×l (AvgPool(X))) × ReLU(Norm(g k×k (X))) = A rf × F rf
[0117] where g l×l and g k×k represent grouped convolutions of size l×l and k×k respectively, l and k represent the size of the convolutional kernel, Norm() represents normalization, X represents the input feature map, and F is obtained by multiplying the attention map A rf with the transformed receptive field spatial feature F rf ;
[0118] C2f_RFAConv contains two convolutional layers and multiple Bottleneck_RFAConv modules. First, the input is passed through the first convolutional layer for channel expansion, and the output is split into two parts. Then, multiple Bottleneck_RFAConv modules process one part, and finally all the features are concatenated and passed through the second convolutional layer to obtain the final output. This design improves the efficiency and accuracy of feature extraction. Among them, the Bottleneck_RFAConv module is a bottleneck structure that contains two convolutional layers. The first convolutional layer reduces the number of input channels to the hidden channels, and the second convolutional layer uses RFAConv for processing. This module supports shortcut connections. When the number of input and output channels is equal, the input can be directly added to the output, thereby enhancing the feature expression ability of the network.
[0119] To better extract target features, this paper introduces RFA into the target detection network, and specifically integrates RFAConv and C2f_RFAConv into the Backbone structure of YOLOv8. In this way, RFA can effectively enhance the network's feature extraction ability for targets of different scales and positions, and improve the expression performance of the feature map.
[0120] The beneficial effects of the above technologies include: introducing the RFA attention mechanism in the feature fusion process of the Feature Pyramid Network, enhancing the attention and fusion ability for features of different scales. RFAConv uses a 3×3 convolutional kernel and GroupConv to extract receptive field features, improving the speed and efficiency of feature extraction. By aggregating information through AvgPool, performing information interaction through 1×1 group convolution operations, and emphasizing feature importance through softmax, the feature processing and attention allocation are optimized. C2f_RFAConv contains multiple modules and convolutional layers, improving the efficiency and accuracy of feature extraction and enhancing the feature expression ability of the network. The design of the Bottleneck_RFAConv module reduces the computational overhead and the number of parameters during the process of reducing the number of channels and using RFAConv for processing. Introducing RFA into the Backbone structure of YOLOv8 improves the feature extraction ability for targets of different scales and positions, thereby improving the detection accuracy. For problems such as insufficient receptive fields, RFA can dynamically adjust the receptive field, adaptively extract context information, and enhance the detection accuracy and robustness. When dealing with complex backgrounds or occlusion situations, the C2f_RFAConv module can effectively integrate feature information to achieve accurate identification of pests and diseases. Overall, these improvements make the tomato leaf pest and disease detection model more efficient, accurate, and reliable, contributing to improving the effect and practicality of pest and disease detection.
[0121] Example 4:
[0122] Based on the tomato leaf pest and disease detection method that improves YOLOv8 on the basis of Example 2, the formula for Wise IOU is:
[0123]
[0124] L WIoU = εR WIoU
[0125] In the formula, R WIoU is the degree of overlap between the anchor box and the target box (R WIoU ∈[1, e)), exp() is the exponential function with the natural constant e as the base and the value of the natural constant e is 2.71828, x is the abscissa of the center point of the anchor box, y is the ordinate of the center point of the anchor box, x gt is the abscissa of the center point of the target box, y gt is the ordinate of the center point of the target box, W g is the width of the target box, H g is the height of the target box, L WIoU is the output loss value of Wise IOU, and ε is the weighting factor (ε∈[0, 1], which is used to significantly reduce the attention to the distance of the center point when the anchor box and the target box overlap well).
[0126] In this embodiment, to prevent R WIoU from generating gradients that hinder convergence, W g and H g are detached from the computational graph. Since it effectively eliminates the factors that hinder convergence, no new metrics such as aspect ratio are introduced.
[0127] In this embodiment, IOU (Intersection over Union) is a commonly used evaluation metric for object detection models, which is used to measure the overlap degree between the predicted bounding box and the ground truth bounding box. However, when dealing with objects of different sizes, the IOU evaluation results may be biased. For example, when dealing with small objects, even if the predicted bounding box has a high overlap with the ground truth bounding box, the IOU value may be low, which may lead to insufficient attention to small objects during the model training process.
[0128] Wise IOU (WIoU) is an improved IOU calculation method that takes into account the shape and size of the bounding box and introduces a weighting factor to balance the attention to small and large objects in the evaluation metric. It aims to improve the generalization ability of the object detection model, especially in the case where the training data inevitably contains low-quality examples. Traditional geometric metrics (such as distance, aspect ratio) tend to exacerbate the penalty for these examples when dealing with low-quality anchor boxes, resulting in a decline in the generalization performance of the model. By introducing a distance attention mechanism, Wise IOU can significantly weaken the penalty of geometric metrics when the anchor box and the target box overlap well, thereby reducing the interference with low-quality examples and enhancing the overall generalization ability of the model.
[0129] The beneficial effects of the above technologies are as follows: Introducing Wise IOU in YOLOv8 can better handle the relationship between the anchor box and the target box, thereby improving the accuracy and robustness of object detection. Wise IOU combines weighted IOU and traditional IOU, taking into account the distance between the anchor box and the target box, enabling the model to reduce the penalty when facing low-quality anchor boxes and avoiding negative impacts on training. This design not only improves the detection ability for small objects and complex scenes but also enhances the overall generalization performance of the model.
[0130] Embodiment 5:
[0131] Based on the tomato leaf disease and pest detection method that improves YOLOv8 on the basis of Embodiment 1, S3: Configure training hyperparameters and update the model weights based on the tomato leaf image dataset to obtain a trained improved YOLOv8 model, including:
[0132] Divide the tomato leaf image dataset into a training set, a validation set, and a test set based on the division ratio of 8:1:1;
[0133] Perform data augmentation on the training set to obtain an optimized training set, and perform standardization or normalization operations on the test set and the validation set to obtain a standard test set and a standard validation set;
[0134] Configure the training hyperparameters, set the batch size to 32 and the number of iterations to 300 rounds, and use the optimized training set to train the improved YOLOv8 model. In each training iteration, calculate the loss and update the model weights, and at the same time monitor the performance metrics on the standard test set until the latest improved YOLOv8 model is obtained;
[0135] Calculate the recognition results of the currently obtained latest improved YOLOv8 model on the test set and the true labels of the test set, draw a confusion matrix, and calculate the accuracy and loss value of the currently obtained latest improved YOLOv8 model on the test set based on the confusion matrix as the model evaluation results of the currently obtained latest improved YOLOv8 model;
[0136] Based on whether the model evaluation results of the currently obtained latest improved YOLOv8 model meet the requirements, if so, regard the improved YOLOv8 model as the trained improved YOLOv8 model, otherwise, continue to train and test the currently obtained latest improved YOLOv8 model until a trained improved YOLOv8 model that meets the requirements is obtained.
[0137] In this embodiment, the confusion matrix: a matrix table used to evaluate the performance of a classification model. In the detection of tomato leaf diseases and pests, it shows the matching situation between the disease and pest categories predicted by the model and the actual disease and pest categories. The rows of the matrix represent the actual disease and pest categories, the columns represent the disease and pest categories predicted by the model, and each element in the matrix represents the number of samples whose actual category is a certain category and is predicted by the model as another category.
[0138] In this embodiment, evaluate the accuracy Train_Accuracy, loss value Train_Loss of the model on the training set, and the accuracy Val_Accuracy, loss value Val_Loss of the test set.
[0139] Among them, the accuracy rate is:
[0140]
[0141] Among them, TP represents the number of instances correctly classified as positive examples, that is, the number of instances that are actually positive examples and are classified as positive examples by the classifier; TN represents the number of instances correctly classified as negative examples, that is, the number of instances that are actually negative examples and are classified as negative examples by the classifier; FP represents the number of instances misclassified as positive examples, that is, the number of instances that are actually negative examples but are classified as positive examples by the classifier; FN represents the number of instances misclassified as negative examples, that is, the number of instances that are actually positive examples but are classified as negative examples by the classifier.
[0142] The loss value is as follows:
[0143]
[0144] In the formula, N is the total number of samples, i represents one of the output samples, y i is the actual value, x i is the predicted value, and ln is the logarithmic function with the natural constant e as the base.
[0145] In this embodiment, for the technical solution of the present invention, the network structure of YOLOv8 can also be replaced with a weighted bidirectional feature pyramid network (WBFPN) to achieve the extraction and fusion of multi-scale features, thereby achieving the effect of improving the detection accuracy of the model.
[0146] In this embodiment, data augmentation is performed on the training set to obtain an optimized training set, and standardization or normalization operations are performed on the test set and the validation set to obtain a standard test set and a standard validation set: By applying data augmentation techniques such as flipping, rotation, and cropping to the training set, the diversity and richness of the data are increased, thereby obtaining an optimized training set. At the same time, standardization or normalization processing is performed on the test set and the validation set to map the feature values of the data to a specific range or distribution to obtain a standard test set and a standard validation set, so that the data is comparable and consistent in different dimensions. Example: For the tomato leaf images in the training set, randomly perform horizontal flipping, rotation at a certain angle, and randomly crop some areas to expand the training data. Standardize the pixel values of the images in the test set and the validation set so that their mean is 0 and the variance is 1.
[0147] In this embodiment, training hyperparameters are configured: Set the parameters used to control the training process and affect the model performance during model training. These parameters include but are not limited to batch size, number of iterations, learning rate, etc. Example: Set the batch size to 32, the number of iterations to 300 rounds, the initial value of the learning rate to 0.01, and adjust it according to a certain strategy during training.
[0148] In this embodiment, the loss is calculated and the model weights are updated while monitoring the performance metrics on the standard test set until the latest improved YOLOv8 model is obtained: In each training iteration, the loss value is calculated based on the prediction results and the ground truth labels, and then the weight parameters of the model are updated according to the loss value through a specific optimization algorithm. At the same time, the performance metrics on the standard test set, such as accuracy, recall, etc., are continuously monitored, and the model is continuously optimized until the currently trained improved YOLOv8 model is obtained. For example, the cross-entropy loss is calculated in each iteration, and the weights are updated using the stochastic gradient descent algorithm. The accuracy is calculated on the standard test set every 10 epochs. If the accuracy continues to increase, the training continues; otherwise, the training stops and the current model is saved.
[0149] In this embodiment, the ground truth labels of the test set: In the test set, the actual pest and disease categories corresponding to each tomato leaf image are labeled. For example, for a certain tomato leaf image in the test set, its ground truth label is "infected with Alternaria solani".
[0150] The beneficial effects of the above technologies are as follows: The tomato leaf image dataset is divided proportionally, providing a clear and reasonable dataset allocation for model training, validation, and testing. Data augmentation is performed on the training set, increasing the data diversity and helping to improve the generalization ability of the model. Standardization or normalization operations are performed on the test set and the validation set to ensure the consistency and comparability of the data. Appropriate training hyperparameters, such as batch size and number of iterations, are configured to improve the training efficiency and effect. The loss is calculated and the model weights are updated during the training process, enabling continuous optimization of the model performance. The model is evaluated by calculating the accuracy and loss value through plotting the confusion matrix, providing comprehensive and intuitive model performance measurement metrics. Judging whether the requirements are met based on the model evaluation results ensures that the finally obtained optimal model has good performance.
[0151] Embodiment 6:
[0152] Based on the tomato leaf pest and disease detection method of the improved YOLOv8 on the basis of Embodiment 1, S4: The pest and disease detection of the tomato leaf image to be detected is performed based on the trained improved YOLOv8 model, and the pest and disease detection results of the tomato leaf image to be detected are obtained, including:
[0153] The tomato leaf image to be detected is subjected to first-level super-resolution reconstruction processing and divided into S×S grids, each grid is input into the trained improved YOLOv8 model, B bounding boxes in each grid are predicted, and the center coordinates, width, and height of each bounding box are calculated;
[0154] x = σ(t x ) + c x
[0155] y = σ(t y) + c y
[0156]
[0157] In the formula, x is the abscissa value of the center coordinate of the bounding box, y is the ordinate value of the center coordinate of the bounding box, σ() is the sigmoid function, t x is the horizontal offset of the center of the bounding box relative to the grid where it is located, t y is the vertical offset of the center of the bounding box relative to the grid where it is located, c x is the abscissa value of the center point of the grid corresponding to the currently calculated bounding box, c y is the ordinate value of the center point of the grid corresponding to the currently calculated bounding box, w is the width of the bounding box, p w is the width of the prior box, e is the natural constant with a value of 2.71828, t w is the relative change in the width of the bounding box, h is the height of the bounding box, p h is the height of the prior box, t h is the relative adjustment of the height of the bounding box;
[0158] Determine the intersection over union (IoU) between the predicted box and the ground truth box based on the center coordinates, width, and height of each bounding box, and calculate the confidence of each bounding box based on the IoU between the predicted box and the ground truth box:
[0159] C = P(object)·IoU truth
[0160] In the formula, C is the confidence of the bounding box, IoU truth is the intersection over union between the predicted box and the ground truth box;
[0161] Select the final target bounding box based on the non-maximum suppression (NMS) algorithm;
[0162] Predict the probability distribution of each target bounding box for each type of pest and disease based on the confidence:
[0163]
[0164] In the formula, P(i|box) is the probability of the target bounding box for the i-th type of pest and disease, e is the natural constant with a value of 2.71828, z i is the original prediction score of the target bounding box for the i-th type of pest and disease, z j is the original prediction score of the target bounding box for the j-th type of pest and disease, C is the total number of pest and disease types;
[0165] Based on the probability distribution of all target bounding boxes in all grids of the tomato leaf image to be detected for each type of pest and disease, obtain the pest and disease detection result of the tomato leaf image to be detected.
[0166] In this embodiment, YOLOv8 uses the non-maximum suppression (NMS) algorithm to screen the final target bounding boxes. The process of NMS is to set a threshold to remove those boxes with too high overlap with the boxes with higher confidence, ensuring that the finally output bounding boxes are unique and have the highest confidence. Through these steps, YOLOv8 can efficiently process the feature map, perform target classification and localization, so as to achieve accurate target detection in various complex scenarios.
[0167] In this embodiment, perform first-level super-resolution reconstruction processing on the tomato leaf image to be detected: use a specific algorithm or technology to perform the first super-resolution reconstruction operation on the tomato leaf image with pests and diseases to be detected, aiming to improve the clarity and details of the image for more accurate pest and disease detection in the follow-up. Example: Through the super-resolution reconstruction algorithm of deep learning, the clarity of the tomato leaf image to be detected with an original resolution of 640×480 is increased to 1280×960.
[0168] In this embodiment, determine the intersection over union (IoU) between the predicted box and the ground truth box based on the center coordinates, width, and height of each bounding box: According to the calculated center coordinates, width, and height of each predicted bounding box, and the corresponding parameters of the known ground truth box (i.e., the actually annotated box), obtain the overlapping degree ratio between the predicted box and the ground truth box through a specific mathematical formula or calculation method. Example: Suppose the center coordinates of the predicted box are (100, 100), the width is 50, and the height is 50, and the center coordinates of the ground truth box are (80, 80), the width is 60, and the height is 60. The IoU between the two is obtained as 0.6 through the IoU calculation formula.
[0169] The beneficial effects of the above technology include: performing first-level super-resolution reconstruction processing on the tomato leaf image to be detected, improving the quality and details of the image, and helping the model to more accurately detect pests and diseases. Dividing the image into grids and inputting them into the trained improved YOLOv8 model for prediction can more finely analyze different regions of the image. Calculating the center coordinates, width, and height of the bounding box through clear mathematical formulas makes the determination of the bounding box more accurate and scientific. Calculating the confidence of the bounding box based on the intersection over union (IoU) can effectively measure the reliability of the bounding box. Using the non-maximum suppression (NMS) algorithm to filter out the final target bounding boxes reduces duplicate and redundant detection boxes and improves the accuracy of detection. Predicting the probability distribution of the target bounding box for each type of pest and disease based on the confidence provides a quantitative basis for the classification of pests and diseases. Finally, obtaining the detection result based on the probability distribution of the target bounding boxes in all grids makes the detection result more comprehensive and reliable. These steps and calculation methods improve the accuracy, fineness, and reliability of tomato leaf pest and disease detection, and help to discover and handle pest and disease problems in a timely manner.
[0170] Example 7:
[0171] Based on the method for detecting tomato leaf pests and diseases based on the improved YOLOv8 in Example 6, and based on the probability distribution of all target bounding boxes in all grids in the tomato leaf image to be detected for each type of pest and disease, the pest and disease detection result of the tomato leaf image to be detected is obtained, including:
[0172] Performing pest and disease detection annotation on the tomato leaf to be detected based on the probability distribution of all target bounding boxes in all grids in the tomato leaf image to be detected for each type of pest and disease, and obtaining the non-minimal target pest and disease detection annotation result and the fine annotation result of the minimal target pest and disease detection of the tomato leaf image to be detected;
[0173] Based on the non-minimal target pest and disease detection annotation result and the fine annotation result of the minimal target pest and disease detection of the tomato leaf image to be detected, the pest and disease detection result of the tomato leaf image to be detected is obtained.
[0174] In this embodiment, pest and disease detection annotation of the tomato leaf image to be detected is performed based on the probability distribution of all target bounding boxes in all grids of the tomato leaf image to be detected for each type of pest and disease, and the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection of the tomato leaf image to be detected are obtained: According to the probability distribution of each target bounding box corresponding to each type of pest and disease in all grids into which the tomato leaf image to be detected is divided, the annotation work of pest and disease detection of this leaf is carried out. In this process, the detection annotation result of non-minimum target pests and diseases will be obtained respectively, that is, the annotation of more obvious or larger area pest and disease regions; and the fine annotation result of minimum target pest and disease detection, that is, a more detailed and accurate annotation for smaller area or less obvious pest and disease regions. Example: For a tomato leaf image to be detected, after analyzing the pest and disease probability distribution of the target bounding boxes in its grid, a large area of leaf spot disease region on the leaf is marked (non-minimum target pest and disease detection annotation result), and at the same time, a very small leaf spot disease region at the edge of the leaf is finely marked (fine annotation result of minimum target pest and disease detection).
[0175] The beneficial effects of the above technology are as follows: By performing pest and disease detection annotation on the tomato leaf image to be detected through probability distribution, the types and locations of pests and diseases can be intuitively displayed. Distinguishing the detection annotation results of non-minimum target and minimum target pests and diseases makes the detection results more detailed and comprehensive. Combining the annotation results of non-minimum target and minimum target to obtain the final pest and disease detection result improves the accuracy and integrity of the detection. It can more accurately identify pests and diseases of different sizes, providing more targeted guidance for subsequent prevention and control measures. This step-by-step annotation and comprehensive method effectively improves the effect and practical value of tomato leaf pest and disease detection.
[0176] Embodiment 8:
[0177] Based on the tomato leaf pest and disease detection method improved from YOLOv8 in Embodiment 7, and based on the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection of the tomato leaf image to be detected, the pest and disease detection result of the tomato leaf image to be detected is obtained, including:
[0178] Based on the non-minimum target pest and disease detection annotation result and the fine annotation result of minimum target pest and disease detection of the tomato leaf image to be detected, the pest and disease detection evaluation value of the tomato leaf image to be detected is calculated;
[0179] Based on the pest and disease detection evaluation value of the tomato leaf image to be detected, the pest and disease detection result of the tomato leaf image to be detected is obtained.
[0180] In this embodiment, the pest and disease detection evaluation value of the tomato leaf image to be detected is calculated: through a specific calculation method and based on the relevant coordinate information of the extremely small target pest and disease detection frames and non-extremely small target pest and disease detection frames in the tomato leaf image to be detected, a value is obtained to measure the pest and disease detection effect and accuracy of the tomato leaf image to be detected. For example, comprehensively considering the distance, distribution and other coordinate information between each extremely small target pest and disease detection frame and non-extremely small target pest and disease detection frame in the tomato leaf image to be detected, after a series of operations, an evaluation value between 0 and 1 is obtained, such as 0.8, indicating that the pest and disease detection of this image has high accuracy and reliability.
[0181] The beneficial effects of the above technology are as follows: By calculating the pest and disease detection evaluation value based on the detection and annotation results of non-extremely small targets and extremely small targets, the detection results can be quantitatively evaluated. Obtaining the pest and disease detection results based on the evaluation value makes the results more objective and reliable. This way of quantitative evaluation helps to accurately judge the quality of the detection results. It can provide a clear basis for the improvement and optimization of subsequent detection methods. It improves the scientificity and credibility of the tomato leaf pest and disease detection results.
[0182] Embodiment 9:
[0183] Based on the tomato leaf pest and disease detection method improved by YOLOv8 in Embodiment 8, based on the coordinate information of all extremely small target pest and disease detection frames in the super-resolution image of the tomato leaf to be detected and all non-extremely small target pest and disease detection frames in the corresponding non-extremely small target pest and disease detection pure annotation results, the pest and disease detection evaluation value of the tomato leaf image to be detected is calculated, including:
[0184] Based on the coordinate information of all extremely small target pest and disease detection frames in the super-resolution image of the tomato leaf to be detected and all non-extremely small target pest and disease detection frames in the corresponding non-extremely small target pest and disease detection pure annotation results, determine the shortest distance between each extremely small target pest and disease detection frame and the corresponding non-extremely small target pest and disease detection frame in the super-resolution image of the tomato leaf to be detected;
[0185] Based on the shortest distance between each extremely small target pest and disease detection frame and the corresponding non-extremely small target pest and disease detection frame in the super-resolution image of the tomato leaf to be detected, calculate the pest and disease detection evaluation value of the tomato leaf image to be detected:
[0186]
[0187] Wherein, P is the pest and disease detection and evaluation value of the tomato leaf image to be detected, n is the total number of minimum target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected, m is the total number of non-minimum target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected, max(n, m) is the maximum value between the total number of minimum target pest and disease detection boxes and the total number of non-minimum target pest and disease detection boxes contained in the super-resolution image of the tomato leaf to be detected, P m is the standard proximity between the total number of minimum target pest and disease detection boxes and the total number of non-minimum target pest and disease detection boxes, l ij is the shortest distance between the i-th minimum target pest and disease detection box and the j-th non-minimum target pest and disease detection box, P l is the standard distribution proximity between all minimum target pest and disease detection boxes and all non-minimum target pest and disease detection boxes.
[0188] In this embodiment, the standard proximity between the total number of minimum target pest and disease detection boxes and the total number of non-minimum target pest and disease detection boxes: A standard measure for measuring the closeness between the total number of minimum target pest and disease detection boxes and the total number of non-minimum target pest and disease detection boxes. For example, the standard proximity may be calculated as 0.71 through a specific formula.
[0189] In this embodiment, the standard distribution proximity between all minimum target pest and disease detection boxes and all non-minimum target pest and disease detection boxes: A standard index for evaluating the similarity or closeness in spatial distribution between all minimum target pest and disease detection boxes and all non-minimum target pest and disease detection boxes. For example, the standard distribution proximity may be calculated as 0.73 through a specific formula.
[0190] The beneficial effects of the above technologies include: By determining the shortest distance between the minimum target pest and disease detection box and the non-minimum target pest and disease detection box, the positional relationship between the two can be finely measured. Calculating the pest and disease detection and evaluation value based on the shortest distance provides a quantitative and accurate evaluation index for the detection effect. Considering the total number of minimum target and non-minimum target detection boxes and their distribution proximity makes the evaluation more comprehensive and comprehensive. The calculation method of this evaluation value is scientific and reasonable, and can objectively reflect the accuracy and consistency of the detection method for detecting pest and diseases of different sizes. It helps to discover possible problems in the detection process and provides a strong basis for further optimizing the detection model and method. It improves the accuracy and reliability of the tomato leaf pest and disease detection and evaluation, and helps to enhance the performance and application value of the detection technology.
[0191] Example 10:
[0192] Based on Example 8, a tomato leaf pest and disease detection method based on improved YOLOv8 obtains the pest and disease detection results of the tomato leaf image to be detected based on the pest and disease detection evaluation value of the tomato leaf image to be detected, including:
[0193] If the pest and disease detection evaluation value of the tomato leaf image to be detected is not less than the preset evaluation threshold, then all the extremely small target pest and disease detection frames and all non-extremely small target pest and disease detection frames in the super-resolution image of the tomato leaf to be detected are marked to obtain the pest and disease detection results of the tomato leaf image to be detected;
[0194] If the pest and disease detection evaluation value of the tomato leaf image to be detected is less than the preset evaluation threshold, then the tomato leaf image to be detected is subjected to second-level super-resolution reconstruction processing and pest and disease detection until the obtained evaluation value of the extremely small target pest and disease is not less than the preset evaluation threshold, and then all the extremely small target pest and disease detection frames and all non-extremely small target pest and disease detection frames in the latest obtained super-resolution image of the tomato leaf to be detected are marked to obtain the pest and disease detection results of the tomato leaf image to be detected.
[0195] In this embodiment, the preset evaluation threshold is a numerically defined boundary set in advance for determining whether the pest and disease detection evaluation value of the tomato leaf image to be detected is qualified or meets specific requirements. Example: The preset evaluation threshold is set to 0.8. If the calculated pest and disease detection evaluation value of the tomato leaf image to be detected is greater than or equal to 0.8, the detection result is considered qualified; if it is less than 0.8, it is unqualified.
[0196] In this embodiment, the second-level super-resolution reconstruction processing of the tomato leaf image to be detected is to perform a higher-degree or different-way super-resolution reconstruction operation on the tomato leaf image with pests and diseases to be detected to further improve the image quality, thereby possibly obtaining more accurate detection results. Example: The first-level super-resolution reconstruction may double the image resolution, while the second-level super-resolution reconstruction may quadruple it, or use a more complex algorithm for processing.
[0197] The beneficial effects of the above technologies are as follows: When the evaluation value is not less than the preset threshold, the detection results are directly obtained by marking the detection frames, ensuring the reliability and effectiveness of the detection results. For the case where the evaluation value is less than the threshold, second-level super-resolution reconstruction processing and re-detection are carried out, improving the accuracy and integrity of the detection. By presetting the evaluation threshold to judge the usability of the detection results, a clear standard is provided for the detection quality. This flexible processing method based on the evaluation value maximally ensures the accuracy and credibility of the finally obtained pest and disease detection results. It effectively improves the accuracy and reliability of tomato leaf pest and disease detection and can provide strong support for pest and disease control in tomato planting.
[0198] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A tomato leaf disease and insect pest detection method based on improved YOLOv8, characterized in that: include: S1: Build a tomato leaf image dataset containing each pest type; S2: Integrate GhostNet and RFA attention mechanism, modify the loss function to WiseIOU, and build an improved YOLOv8 model; S3: Configure training hyperparameters and update model weights based on the tomato leaf image dataset to obtain a trained improved YOLOv8 model; S4: Based on the trained improved YOLOv8 model, pest and disease detection is performed on the tomato leaf image to obtain the pest and disease detection result of the tomato leaf image to be detected.
2. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 1, characterized in that: S2: Integrate GhostNet and RFA attention mechanisms, modify the loss function to WiseIOU, and build an improved YOLOv8 model, including: Load the pre-trained weights of YOLOv8, integrate the RFA attention mechanism into the feature extraction network of YOLOv8, replace the traditional convolution of the neck structure with GhostConv, replace the feature fusion module in the neck structure with C2fGhost, and then replace the loss function with WiseIOU to build an improved YOLOv8 model.
3. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 2, characterized in that: Integrate the RFA attention mechanism into the feature extraction network of YOLOv8, including: The RFA attention mechanism is introduced in the feature fusion process of the feature pyramid network, and RFAConv has a 3×3 convolution kernel and uses Group Conv to extract receptive field features. It aggregates information through AvgPool, interacts with 1×1 group convolution operations, and emphasizes the importance of features through softmax. C2f_RFAConv contains two convolutional layers and multiple Bottleneck_RFAConv modules. The Bottleneck_RFAConv module contains two convolutional layers and the first convolutional layer reduces the number of input channels. The second convolutional layer is processed by RFAConv and supports shortcut connections.
4. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 2, characterized in that: The formula for WiseIOU is: L WIoU =εR WIoU In the formula, R WIoU is the overlap between the anchor box and the target box, exp() is an exponential function with the natural constant e as the base and the value of the natural constant e is 2.71828, x is the horizontal coordinate of the center point of the anchor box, y is the vertical coordinate of the center point of the anchor box, and x gt is the horizontal coordinate of the center point of the target frame, y gt is the ordinate of the center point of the target box, W g is the width of the target box, H g is the height of the target box, L WIoU is the output loss value of Wise IOU, and ε is the weighting factor.
5. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 1, characterized in that: S3: Configure training hyperparameters and update model weights based on the tomato leaf image dataset to obtain a trained improved YOLOv8 model, including: The tomato leaf image dataset is divided into training set, validation set and test set based on the division ratio of 8:1:1; Perform data augmentation on the training set to obtain an optimized training set, and perform standardization or normalization operations on the test set and validation set to obtain a standard test set and a standard validation set; Configure the training hyperparameters, set the batch size to 32 and the number of iterations to 300, and use the optimized training set to train the improved YOLOv8 model. In each training iteration, calculate the loss and update the model weights, while monitoring the performance indicators on the standard test set until the latest improved YOLOv8 model is obtained. Calculate the recognition results of the latest improved YOLOv8 model on the test set and the true labels of the test set to draw a confusion matrix, and calculate the accuracy and loss value of the latest improved YOLOv8 model on the test set based on the confusion matrix as the model evaluation result of the latest improved YOLOv8 model; Based on judging whether the model evaluation result of the latest improved YOLOv8 model currently obtained meets the requirements, if so, the improved YOLOv8 model is regarded as a trained improved YOLOv8 model, otherwise, the latest improved YOLOv8 model currently obtained is continuously trained and tested and verified until a trained improved YOLOv8 model that meets the requirements is obtained.
6. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 1, characterized in that: S4: Based on the trained improved YOLOv8 model, pest and disease detection is performed on the tomato leaf image to be detected, and the pest and disease detection results of the tomato leaf image to be detected are obtained, including: Perform the first-level super-resolution reconstruction on the tomato leaf image to be detected and divide it into S×S grids. Input each grid into the trained improved YOLOv8 model, predict B bounding boxes in each grid, and calculate the center coordinates, width, and height of each bounding box. x=σ(t x )+c x y=σ(t y )+c y Where x is the horizontal coordinate value of the center coordinate of the bounding box, y is the vertical coordinate value of the center coordinate of the bounding box, σ() is the sigmoid function, and t x is the lateral offset of the center of the bounding box relative to the grid, t y is the vertical offset of the center of the bounding box relative to the grid, c x is the horizontal coordinate value of the center point of the grid corresponding to the currently calculated bounding box, c y is the ordinate value of the center point of the grid corresponding to the currently calculated bounding box, w is the width of the bounding box, and p w is the width of the prior box, e is a natural constant and its value is 2.71828, t w is the relative change in the width of the bounding box, h is the height of the bounding box, and p h is the height of the prior box, t h is the relative adjustment of the bounding box height; Based on the center coordinates, width, and height of each bounding box, the intersection-and-union ratio between the predicted box and the true box is determined, and the confidence of each bounding box is calculated based on the intersection-and-union ratio between the predicted box and the true box: C=P(object)·IoU truth Where C is the confidence of the bounding box, IoU truth is the intersection-over-union ratio between the predicted box and the true box; Filter out the final target bounding box based on the non-maximum suppression algorithm; Based on the confidence, the probability distribution of each target bounding box for each type of pest is predicted: Where P(ibox) is the probability of the target bounding box for the i-th pest type, e is a natural constant with a value of 2.71828, and z i is the original prediction score of the target bounding box for the i-th pest type, z j is the original prediction score of the target bounding box for the jth pest type, and C is the total number of pest types; Based on the probability distribution of all target bounding boxes in all grids in the tomato leaf image to be detected for each type of pests and diseases, the pest and disease detection result of the tomato leaf image to be detected is obtained.
7. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 6, characterized in that: Based on the probability distribution of all target bounding boxes in all grids in the tomato leaf image to be detected for each type of pests and diseases, the pest and disease detection result of the tomato leaf image to be detected is obtained, including: Based on the probability distribution of each pest type for all target bounding boxes in all grids in the tomato leaf image to be detected, pest detection and annotation are performed on the tomato leaf to be detected, and a non-minimum target pest detection and annotation result and a minimum target pest detection and annotation result are obtained for the tomato leaf image to be detected; Based on the non-minimum target pest and disease detection annotation results and the minimum target pest and disease detection fine annotation results of the tomato leaf image to be detected, the pest and disease detection results of the tomato leaf image to be detected are obtained.
8. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 7, characterized in that: Based on the non-minimum target pest and disease detection annotation results and the minimum target pest and disease detection fine annotation results of the tomato leaf image to be detected, the pest and disease detection results of the tomato leaf image to be detected are obtained, including: Based on the non-minimum target pest and disease detection labeling results and the minimum target pest and disease detection fine labeling results of the tomato leaf image to be detected, the tomato leaf image to be detected is obtained and the pest and disease detection evaluation value of the tomato leaf image to be detected is calculated; Based on the disease and insect pest detection evaluation value of the tomato leaf image to be detected, the disease and insect pest detection result of the tomato leaf image to be detected is obtained.
9. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 8, characterized in that: Based on the coordinate information of all the extremely small target pest detection frames in the super-resolution image of the tomato leaf to be detected and all the non-extremely small target pest detection frames in the corresponding pure annotation results of non-extremely small target pest detection, the pest detection evaluation value of the tomato leaf image to be detected is calculated, including: Based on the coordinate information of all the minimal target pest and disease detection frames in the super-resolution image of the tomato leaf to be detected and all the non-minimal target pest and disease detection frames in the corresponding pure annotation results of non-minimal target pest and disease detection, the shortest distance between each minimal target pest and disease detection frame in the super-resolution image of the tomato leaf to be detected and each corresponding non-minimal target pest and disease detection frame is determined; Based on the shortest distance between each minimal target pest detection frame and each corresponding non-minimal target pest detection frame in the super-resolution image of the tomato leaf to be detected, the pest detection evaluation value of the tomato leaf image to be detected is calculated: Wherein, P is the pest detection evaluation value of the tomato leaf image to be detected, n is the total number of extremely small target pest detection frames contained in the tomato leaf super-resolution image to be detected, m is the total number of non-extremely small target pest detection frames contained in the tomato leaf super-resolution image to be detected, max(n,m) is the maximum value of the total number of extremely small target pest detection frames contained in the tomato leaf super-resolution image to be detected and the total number of non-extremely small target pest detection frames contained in the tomato leaf super-resolution image to be detected, P m is the standard similarity between the total number of minimal target pest detection frames and the total number of non-minimal target pest detection frames, l ij is the shortest distance between the i-th minimum target pest detection frame and the j-th non-minimum target pest detection frame, P l It is the similarity between the standard distribution of all minimal target pest and disease detection frames and all non-minimal target pest and disease detection frames.
10. The tomato leaf disease and insect pest detection method based on improved YOLOv8 according to claim 8, characterized in that: Based on the pest and disease detection evaluation value of the tomato leaf image to be detected, the pest and disease detection result of the tomato leaf image to be detected is obtained, including: If the pest detection evaluation value of the tomato leaf image to be detected is not less than the preset evaluation threshold, all minimal target pest detection frames and all non-minimal target pest detection frames in the tomato leaf super-resolution image to be detected are marked to obtain the pest detection result of the tomato leaf image to be detected; If the pest and disease detection evaluation value of the tomato leaf image to be detected is less than the preset evaluation threshold, the tomato leaf image to be detected is subjected to second-level super-resolution reconstruction processing and pest and disease detection until the latest obtained minimum target pest and disease detection evaluation value is not less than the preset evaluation threshold. Then, all minimum target pest and disease detection frames and all non-minimum target pest and disease detection frames in the latest obtained super-resolution image of the tomato leaf to be detected are marked to obtain the pest and disease detection result of the tomato leaf image to be detected.
Citation Information
Patent Citations
Industrial counting method based on improved YOLOv5
CN114724038A
Method for detecting plant diseases and insect pests of tomato leaves based on improved YOLOv5s
CN116994056A
Rice disease detection method, product, medium and equipment
CN118072147A
Asphalt pavement disease detection method based on YOLOv8n model
CN118247636A
Tomato leaf disease detection method based on YOLO V8
CN118735858A
Cited By
Target detection model based on deep learning and application thereof
CN121259480A