Lightweight detection method for tea plant diseases and insect pests based on improved YOLOv8s
By improving the YOLOv8s-HMStea model, the lightweight and high-precision problems of tea pest detection are solved, and efficient tea pest detection is achieved on smart mobile terminals, with the detection accuracy reaching 95.36%, and the model size is only 4.7MB.
Patent Information
- Application Number
- CN202510427233.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
AI Technical Summary
The existing tea pest and disease detection methods rely on manual identification to be time-consuming and subjective. Deep learning-based detection models cannot be lightweight on smart mobile terminals, resulting in insufficient detection accuracy to meet real-time detection requirements.
Using the improved YOLOv8s-HMStea model, by replacing the backbone network with MobileNetV3, introducing the Haar wavelet downsampling layer, combining the Slim-neck network structure, a lightweight tea pest detection model is built, and the anti-interference and feature extraction capabilities of the model are improved through data augmentation technology.
It realizes efficient and accurate tea pest detection on smart mobile terminals, with the detection accuracy reaching 95.36%, and the model size is only 4.7MB, suitable for smartphones and other devices.
Smart Images

Figure CN120259885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of plant pest and disease detection, and particularly relates to a lightweight detection method for tea pests and diseases based on improved YOLOv8s. Background Art
[0002] During the growth process, tea is prone to being affected by pests and diseases, which reduces the yield and quality of tea. Traditional tea pest and disease detection mainly relies on manual identification by tea farmers and plant protection experts, which is time-consuming, laborious, and highly subjective.
[0003] In recent years, object detection technology based on deep learning has developed rapidly. In particular, the use of convolutional neural networks has achieved good results in the research of tea pest and disease detection. It can automatically extract abstract features in images, and through training and optimization, finally realize the classification and positioning of pest and disease targets. To improve the accuracy of tea blight monitoring, Bao et al. added a multi-scale module to the backbone network of YOLOv5 and added two-dimensional hybrid attention to the neck to integrate local and non-local information. The mean average precision (mAP) of the resulting DDMA-YOLO model for detecting tea pests and diseases is 76.8%, but such detection accuracy is not sufficient to meet practical applications. Li et al. collected 10 common tea pests and diseases and trained using the ImageNet pre-trained model. The accuracy of the final model for detecting tea pests and diseases is 98.58%. However, this detection model does not consider model lightweighting and cannot be used on intelligent mobile terminals with weak computing performance such as smartphones, which is not conducive to real-time detection of tea pests and diseases in the tea garden. Summary of the Invention
[0004] Aiming at the above-mentioned technical problems, the present invention provides a lightweight detection method for tea pests and diseases based on improved YOLOv8s, which includes:
[0005] Step S1) Construct a sample data set for tea pests and diseases;
[0006] Step S2) Construct an improved YOLOv8s-HMStea model based on the YOLOv8s network. The YOLOv8s-HMStea model includes: a backbone network, a neck network, and a prediction end. The backbone network is formed by replacing the backbone network of YOLOv8s with a MobileNetV3 network and adding a Haar wavelet-based downsampling layer to the backbone network. The neck network is formed by replacing the neck network of YOLOv8s with a Slim-neck network structure. The prediction end uses the Head network structure of the YOLOv8s network;
[0007] Step S3) Train the YOLOv8s-HMStea model using the sample data set;
[0008] Step S4) Input the tea leaf image into the trained YOLOv8s-HMStea model, and output the detection results of tea pests and diseases.
[0009] Among them, the sample data set includes: image data with tea pests and diseases and corresponding annotation information, and the annotation information includes the category of pests and diseases and the position information on the image data.
[0010] Among them, four methods of brightening, darkening, horizontal flipping, and adding Gaussian noise are used to complete the enhancement of image data. Among them, brightening is used to simulate the sunny lighting environment, darkening is used to simulate the cloudy lighting environment, and Gaussian noise is used to simulate the noise effects caused by mud and branches in the natural environment.
[0011] The beneficial effects of the present invention are as follows:
[0012] The lightweight detection method for tea pests and diseases based on the improved YOLOv8s provided by the present invention uses the existing known YOLOv8s as the baseline model, improves the backbone network and the neck network, replaces the backbone network of the YOLOv8s model with MobileNetV3 to complete the model lightweighting, introduces a downsampling module based on Haar wavelet to improve the anti-interference ability and feature extraction ability of the model, replaces the neck network of the YOLOv8s model with the Slim-neck structure, can fuse features of different scales, reduce the false detection and missed detection of the model, its detection accuracy for tea leaf pests and diseases in the natural environment is as high as 95.36%, and the model lightweighting is realized. The size of the YOLOv8s-HMStea model constructed for detecting tea pests and diseases is only 4.7MB, which is convenient to use on intelligent mobile terminals with weak computing performance such as smart phones. Description of the Drawings
[0013] Figure 1 is a flowchart of a lightweight detection method for tea pests and diseases based on the improved YOLOv8s provided by the present invention.
[0014] Figure 2 is a schematic structural diagram of the YOLOv8s-HMStea model provided by the present invention;
[0015] Figure 3 is a schematic structural diagram of the existing known HWD layer;
[0016] Figure 4 is a schematic structural diagram of the existing known bneck layer, GSConv layer, and VoV_GSCSP layer;
[0017] Figure 5 is a schematic structural diagram using the Head network structure of the YOLOv8s network as the prediction end. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] refer to Figure 1 The present invention proposes a lightweight detection method for tea pests and diseases based on improved YOLOv8s, which specifically includes:
[0020] Step S1) constructs a sample data set about tea pests and diseases, the sample data set includes: image data with tea pests and diseases and corresponding annotation information, the annotation information includes the category of the pests and diseases and the location information on the image data.
[0021] Various pests and diseases on tea leaves are photographed to form multiple image data. In this invention, 6 pests and 3 diseases are taken as examples. All images are taken in a natural environment, with soil, sky, tea branches, etc. as the background, and the distance between the shooting equipment and the leaves is 0.2-0.4m. In order to ensure the richness and diversity of the data, multiple shooting methods such as front light, back light, close distance, long distance, depression angle, elevation angle, etc. are adopted. The characteristics of the 6 common pests and 3 diseases used in the present invention are described in Table 1.
[0022] Table 1
[0023]
[0024]
[0025] To avoid overfitting or underfitting of the model caused by unbalanced image quantities and to enhance the robustness and accuracy of the model, four methods, namely brightening, darkening, horizontal flipping, and adding Gaussian noise, are used to complete data augmentation, thereby obtaining more image data. For example, 2,202 original images are taken, and after data augmentation, 11,010 images can be obtained. Among them, brightening simulates the sunny lighting environment, the darkening method simulates the cloudy lighting environment, and Gaussian noise simulates the noise effects caused by mud, branches, etc. in the natural environment. Existing technologies can be used to implement brightening, darkening, horizontal flipping, and adding Gaussian noise, which will not be elaborated here. The LabelImg tool is used to annotate the damage spots of tea pests and diseases. The annotation information includes: the category of pests and diseases and the position information on the image data. The annotation information is stored in the corresponding txt file, and the labels are TAB (aphid), LL (green plant bug), SCD (tea lace bug), SDH (tea yellow thrips), RSJ (persimmon wax scale), CTW (tea leaf miner), GTM (tea anthracnose), CCM (tea zonate leaf spot), CTBH (tea round scab). Finally, the image data and annotation files are randomly divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 to complete the construction of the sample data set.
[0026] In step S2), an improved YOLOv8s-HMStea model is constructed based on the YOLOv8s network. The YOLOv8s-HMStea model includes: a backbone network, a neck network, and a prediction head. Among them, the backbone network is formed by replacing the backbone network of YOLOv8s with the MobileNetV3 network and adding a downsampling layer based on the Haar wavelet in the backbone network. The neck network is formed by replacing the neck network of YOLOv8s with the Slim-neck network structure, and the prediction head uses the Head network structure of the YOLOv8s network.
[0027] Thus, the improved YOLOv8s-HMStea model includes a backbone network composed of a downsampling module based on the Haar wavelet and the MobileNetV3 network, a neck network composed of the Slim-neck network, and a prediction head composed of the Head network of the YOLOv8s network.
[0028] The backbone network is responsible for extracting features from the input image. The neck network is used for feature fusion and enhancement. The Head network is the decision-making part and is responsible for generating the final detection results. The Head network of the YOLOv8s network uses a multi-branch design. Each branch is responsible for object detection at a specific scale. Its main function is to perform object classification and localization. It classifies (predicts the object category) and localizes (predicts the bounding box) each feature point according to the fused feature map provided by the neck network, and finally outputs the detection results.
[0029] Specifically,
[0030] The backbone network of the YOLOv8s - HMStea model includes: 1 Conv_BN layer (a combined layer of convolutional layer and batch normalization layer) with 16 channels, 4 HWD layers with different channels, 4 Concat layers (fully connected layers), and 11 bneck layers with different channels. The 4 HWD layers with different channels include: 1 HWD layer with 16 channels, 1 HWD layer with 24 channels, 1 HWD layer with 48 channels, and 1 HWD layer with 96 channels. The 11 bneck layers with different channels include: 1 bneck layer with 16 channels, 2 serially connected bneck layers with 24 channels, 3 serially connected bneck layers with 40 channels, 2 serially connected bneck layers with 48 channels, and 3 serially connected bneck layers with 96 channels. Among them, the input of the Conv_BN layer receives the input data. The output of the Conv_BN layer is respectively connected to the input of the 16 - channel HWD layer and the input of the 16 - channel bneck layer. The output of the 16 - channel HWD layer and the output of the 16 - channel bneck layer are respectively connected to the input of the first Concat layer. The output of the first Concat layer is connected to the input of 2 serially connected 24 - channel bneck layers. The output of the 16 - channel HWD layer is also connected to the input of the 24 - channel HWD layer. The output of the 24 - channel HWD layer and the output of 2 serially connected 24 - channel bneck layers are respectively connected to the input of the second Concat layer. The output of the second Concat layer is connected to the input of 3 serially connected 40 - channel bneck layers. The output of the 24 - channel HWD layer is also connected to the input of the 48 - channel HWD layer. The output of 3 serially connected 40 - channel bneck layers is connected to the input of 2 serially connected 48 - channel bneck layers. The output of 2 serially connected 48 - channel bneck layers and the output of the 48 - channel HWD layer are respectively connected to the input of the third Concat layer. The output of the 48 - channel HWD layer is also connected to the input of the 96 - channel HWD layer. The output of the third Concat layer is connected to the input of 3 serially connected 96 - channel bneck layers. The output of 3 serially connected 96 - channel bneck layers and the output of the 96 - channel HWD layer are respectively connected to the fourth Concat layer.
[0031] In one instance, sample data with an image size of 640×640, for example, is input into the backbone network. The MobileNetV3 network extracts the depth features of the tea pest and disease images, and the HWD layer extracts the frequency - domain features of the images. The two are fused in 4 levels in sequence, and feature maps with scales of 160×160×32, 80×80×48, 40×40×96, and 20×20×192 are obtained respectively.
[0032] The neck network of the YOLOv8s-HMStea model includes: 2 Upsample layers, 4 Concat layers, 4 VoV_GSCSP layers, and 2 GSConv layers. The input of the first Upsample layer is connected to the output of the fourth Concat layer. The input of the fifth Concat layer is connected to the outputs of the first Upsample layer and the third Concat layer respectively. The output of the fifth Concat layer is connected to the input of the first VoV_GSCSP layer. The output of the first VoV_GSCSP layer is connected to the input of the second Upsample layer. The input of the sixth Concat layer is connected to the outputs of the second Upsample layer and the second Concat layer respectively. The output of the sixth Concat layer is connected to the input of the second VoV_GSCSP layer. The output of the second VoV_GSCSP layer is connected to the input of the first GSConv layer. The input of the seventh Concat layer is connected to the outputs of the first GSConv layer and the first VoV_GSCSP layer respectively. The output of the seventh Concat layer is connected to the input of the third VoV_GSCSP layer. The output of the third VoV_GSCSP layer is connected to the input of the second GSConv layer. The input of the eighth Concat layer is connected to the outputs of the second GSConv layer and the fourth Concat layer respectively. The output of the eighth Concat layer is connected to the input of the fourth VoV_GSCSP layer. Finally, the outputs of the second, third, and fourth VoV_GSCSP layers are connected to the three input ends of the Head network respectively.
[0033] In the present invention, the HWD layer, bneck layer, GSConv layer, and VoV_GSCSP layer all adopt existing known network structures. The structure of the HWD layer is as Figure 3 shown, and the structures of the bneck layer, GSConv layer, and VoV_GSCSP layer are as Figure 4 shown, which will not be elaborated here.
[0034] Continuing with the above example, the backbone network inputs feature maps of sizes 80×80×48, 40×40×96, and 20×20×192 into the neck network of the Slim-neck structure. In the neck network, further feature extraction and fusion are completed through upsampling, concatenation, GSConv, and VoV-GSCSP modules. The fusion of high-level and low-level information is completed, and the fused information is input into the prediction network.
[0035] The prediction end uses the Head network structure of the YOLOv8s network, as Figure 5 shown, which is an existing publicly known network structure and can perform object detection simultaneously at different scales, improving the detection ability for large and small objects.
[0036] Continuing with the above example, the prediction end receives the feature maps from the neck structure. The size and number of channels of the feature maps correspond to different scales: 80×80×128, high resolution, for small object detection; 40×40×256, medium resolution, for medium object detection; 20×20×512, low resolution, for large object detection. Each detection head at the prediction end includes two main branches: a regression branch and a classification branch. Both branches first enter the CBS module, which is a combination of Conv-BatchNorm-SiLU. Among them, the kernel size k of the convolutional layer is 3 for local feature extraction; s = 1 to keep the size of the input feature map; n = 2 means there are two CBS modules, and then the feature expression ability is enhanced through the BatchNorm and SiLU activation functions. The second CBS module is used to further process the features and match the number of channels of the input feature map with subsequent convolutional operations. Secondly, the feature map enters the regression and classification branches. The regression branch structure includes a 3×3 convolution, and the output number of feature channels is 4×reg_max, where reg_max usually refers to the maximum number of attributes that the model can predict for each anchor box or prediction box. In the YOLOv8s-HMStea model, each bounding box usually needs to regress four parameters: the abscissa (x) of the center point, the ordinate (y) of the center point, the width (w), and the height (h), or the offsets of these parameters or their values relative to a certain reference point. Therefore, 4×reg_max means that if reg_max is 1 (i.e., each prediction only considers one anchor box or bounding box), then c = 4, corresponding to these four regression parameters. If reg_max is greater than 1, it means that the model regresses these four parameters for multiple anchor boxes or bounding boxes, so the total output number of channels will increase. The loss function is Bbox Loss.
[0037] In the detection head of YOLOv8, in addition to regressing the bounding box parameters, it is also necessary to predict the confidence of each bounding box belonging to each category. Therefore, the classification branch is used to predict the category of the target, including a 3×3 convolution, and the output number of channels is nc. c = nc means that the output number of channels is equal to the number of categories. Each channel corresponds to the confidence prediction of a category, and finally the detection result with the highest confidence is output.
[0038] Step S3) Train the YOLOv8s-HMStea model using the sample data set.
[0039] The YOLOv8s-HMStea model can be trained using existing known methods. Use the training set divided from the sample data set in step S1 to train the model, use the validation set to verify the training effect of the model, and use the test set to evaluate the trained model. The training method is not the design key point of the present invention and is only briefly described here.
[0040] The training process uses the AdamW optimizer and is based on the Pytorch 1.11.0 deep learning framework. The training and evaluation of the YOLOv8s-HMStea model are carried out under the Ubuntu20.04 system. The CPU of the server is, for example, Intel Xeon E5-2690 V4, and the GPU is, for example, Nvidia GeForce RTX 3080Ti. The installed CUDA and Cudnn versions are 11.4.0 and 8.2.4 respectively, and the programming language is Python3.8.
[0041] The training settings are as follows:
[0042] Initial learning rate: Set to 0.01.
[0043] Batch size: 8 image samples, that is, the model processes 8 image samples each time.
[0044] Number of training epochs: 200 epochs, and an early stopping mechanism is adopted to prevent overfitting.
[0045] Image input size: All tea leaf pest and disease images are uniformly 640×640 pixels.
[0046] The evaluation metrics are used to determine whether the trained YOLOv8s-HMStea model achieves the expected results. Specifically, the validation set is used to evaluate the performance of the YOLOv8s-HMStea model. When the mean average precision (mAP) is greater than 70%, the number of model parameters (Model Parameters) is less than 20M, the accuracy rate is greater than or equal to approximately 92%, and the recall rate is the same as or lower than that of the YOLOv8s model, the performance of the trained YOLOv8s-HMStea model reaches the expected results and the training ends.
[0047] Step S4) Input the tea leaf image into the trained YOLOv8s-HMStea model, and output the detection results regarding tea leaf pests and diseases.
[0048] Input the tea leaf image with pests and diseases into the trained YOLOv8s-HMStea model, and it outputs the detection results regarding tea leaf pests and diseases. The detection results include: whether there are pests and diseases, and when there are pests and diseases, give the conclusion of which specific pests and diseases they are.
[0049] The results obtained through experiments are as follows: The overall precision rate and mean average precision of the YOLOv8s-HMStea model provided by the present invention on 9 types of pests and diseases are 95.36% and 94.17% respectively. The detection results for each type of pest and disease are represented by the mean average precision. The specific experimental results are as follows: aphids (99.5%), Apolygus lucorum (91.22%), Stephanitis chinensis Drake (95.45%), Scirtothrips dorsalis Hood (85.14%), Ricania sublimbata Jacobi (99.23%), Caloptilia theivora Walsingham (97.04%), Colletotrichum camelliae Mass. (94.12%), Pestalotiopsis theae Sawada (98.32%), and Exobasidium vexans Massee (87.55%). Therefore, the YOLOv8s-HMStea model provided by the present invention has higher recognition accuracy. The present invention is based on the YOLOv8s network and replaces the original backbone network with MobileNetV3, which has the advantages of high efficiency, low latency, and low computational cost, and achieves model lightweighting.
[0050] The above content further elaborates on the present invention in combination with specific implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as falling within the protection scope determined by the claims submitted for the present invention.
Claims
1. A lightweight detection method for tea pests and diseases based on improved YOLOv8s, comprising: Step S1) Construct a sample data set for tea pests and diseases; Step S2) Based on the YOLOv8s network, construct an improved YOLOv8s-HMStea model, which includes a backbone network, a neck network, and a prediction head. The backbone network is formed by replacing the backbone network of YOLOv8s with a MobileNetV3 network and adding a downsampling layer based on Haar wavelets to the backbone network. The neck network is formed by replacing the neck network of YOLOv8s with a Slim-neck network structure. The prediction head uses the Head network structure of the YOLOv8s network; Step S3) Use the sample data set to train the YOLOv8s-HMStea model; Step S4) Input the tea leaf image into the trained YOLOv8s-HMStea model and output the detection result for tea pests and diseases.
2. The lightweight detection method for tea pests and diseases based on the improved YOLOv8s according to claim 1, wherein the sample data set includes: Image data with tea pests and diseases and corresponding annotation information, where the annotation information includes the category of pests and diseases and the position information on the image data.
3. The lightweight detection method for tea pests and diseases based on improved YOLOv8s according to claim 2, wherein four methods, namely brightening, darkening, horizontal flipping, and adding Gaussian noise, are used to complete the enhancement of the image data. Among them, brightening is used to simulate the sunny lighting environment, darkening is used to simulate the cloudy lighting environment, and Gaussian noise is used to simulate the noise effects caused by mud and branches in the natural environment.
4. The lightweight detection method for tea pests and diseases based on the improved YOLOv8s according to claim 1, wherein the backbone network of the YOLOv8s-HMStea model includes: 1 16-channel Conv_BN layer, 4 HWD layers with different numbers of channels, 4 Concat layers, and 11 bneck layers with different numbers of channels. The 4 HWD layers with different numbers of channels include: 1 16-channel HWD layer, 1 24-channel HWD layer, 1 48-channel HWD layer, and 1 96-channel HWD layer. The 11 bneck layers with different numbers of channels include: 1 16-channel bneck layer, 2 cascaded 24-channel bneck layers, 3 cascaded 40-channel bneck layers, 2 cascaded 48-channel bneck layers, and 3 cascaded 96-channel bneck layers. Among them, the input of the Conv_BN layer receives the input data. The output of the Conv_BN layer is respectively connected to the input of the 16-channel HWD layer and the input of the 16-channel bneck layer. The output of the 16-channel HWD layer and the output of the 16-channel bneck layer are respectively connected to the input of the first Concat layer. The output of the first Concat layer is connected to the input of the 2 cascaded 24-channel bneck layers. The output of the 16-channel HWD layer is also connected to the input of the 24-channel HWD layer. The output of the 24-channel HWD layer and the output of the 2 cascaded 24-channel bneck layers are respectively connected to the input of the second Concat layer. The output of the second Concat layer is connected to the input of the 3 cascaded 40-channel bneck layers. The output of the 24-channel HWD layer is also connected to the input of the 48-channel HWD layer. The output of the 3 cascaded 40-channel bneck layers is connected to the input of the 2 cascaded 48-channel bneck layers. The output of the 2 cascaded 48-channel bneck layers and the output of the 48-channel HWD layer are respectively connected to the input of the third Concat layer. The output of the 48-channel HWD layer is also connected to the input of the 96-channel HWD layer. The output of the third Concat layer is connected to the input of the 3 cascaded 96-channel bneck layers. The output of the 3 cascaded 96-channel bneck layers and the output of the 96-channel HWD layer are respectively connected to the fourth Concat layer.
5. The lightweight detection method for tea pests and diseases based on the improved YOLOv8s according to claim 4, wherein the neck network of the YOLOv8s-HMStea model comprises: 2 Upsample layers, 4 Concat layers, 4 VoV_GSCSP layers, 2 GSConv layers, where the input of the first Upsample layer is connected to the output of the fourth Concat layer, the inputs of the fifth Concat layer are respectively connected to the output of the first Upsample layer and the output of the third Concat layer, the output of the fifth Concat layer is connected to the input of the first VoV_GSCSP layer, the output of the first VoV_GSCSP layer is connected to the input of the second Upsample layer, the inputs of the sixth Concat layer are respectively connected to the output of the second Upsample layer and the output of the second Concat layer, the output of the sixth Concat layer is connected to the input of the second VoV_GSCSP layer, the output of the second VoV_GSCSP layer is connected to the input of the first GSConv layer, the inputs of the seventh Concat layer are respectively connected to the output of the first GSConv layer and the output of the first VoV_GSCSP layer, the output of the seventh Concat layer is connected to the input of the third VoV_GSCSP layer, the output of the third VoV_GSCSP layer is connected to the input of the second GSConv layer, the inputs of the eighth Concat layer are respectively connected to the output of the second GSConv layer and the output of the fourth Concat layer, and the output of the eighth Concat layer is connected to the input of the fourth VoV_GSCSP layer.