Method and device for detecting tomato diseases and insect pests based on improved YOLOv8

By introducing the multi-scale upsampling enhancement module and the multi-head attention module into the YOLOv8 model, the problems of decreased accuracy and insufficient target differentiation ability of the YOLOv8 model in tomato pest and disease detection are solved, achieving more efficient pest and disease detection, which is suitable for smart agriculture and environmental monitoring.

CN120708213APending Publication Date: 2025-09-26UNIV OF JINAN

Patent Information

Application Number
CN202510608918.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing YOLOv8 model has problems with decreased accuracy and insufficient ability to distinguish between targets and backgrounds in tomato pest and disease detection, and is unable to effectively mine key feature information.

Method used

The multi-scale upsampling enhancement module MEUM is used to replace the original upsampling module, and the multi-scale feature extraction and multi-head attention module MSEPA is added to the neck network of the YOLOv8 model to enhance feature extraction and target detection capabilities.

Benefits of technology

The model's feature expression capability has been improved, the accuracy and classification effect of tomato disease and pest detection have been enhanced, and traditional detection methods have been optimized. It is suitable for fields such as smart agriculture and environmental monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708213A_ABST
    Figure CN120708213A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision target detection, and particularly relates to a tomato disease and insect pest detection method and device based on improved YOLOv8, and the method comprises the steps: obtaining a tomato disease leaf image, and carrying out the preprocessing of the image, so as to construct an image sample set; a disease and pest detection model is constructed based on an improved YOLOv8 model, and the improved YOLOv8 model comprises the steps that a multi-scale up-sampling enhancement module MEUM is used for replacing an original up-sampling module of the YOLOv8 model, and a multi-scale feature extraction and multi-head attention module MSEPA is added into a neck network of the YOLOv8 model; and dividing the image sample set into a training set, a verification set and a test set, and carrying out training optimization on the disease and pest detection model to obtain a target disease and pest detection model. The invention aims to comprehensively enhance the feature extraction and target detection capability of YOLOv8 through the combination of multi-scale feature extraction and an attention mechanism, so as to achieve the improvement of the precision of the YOLOv8 model in the detection of tomato diseases and pests and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision target detection, and particularly relates to a tomato disease and insect pest detection method and device based on improved YOLOv8. Background Art

[0002] With the rapid development of deep learning technology, object detection algorithms have made significant progress in accuracy and real-time performance, and are widely used in fields such as security monitoring, autonomous driving, and medical image analysis. As an advanced object detection algorithm, YOLOv8 has been widely used in many scenarios due to its high accuracy. However, in practical applications, YOLOv8 is prone to missed detections or decreased accuracy. In addition, it does not fully exploit key feature information of the target, making it unable to effectively distinguish the target from the background, which affects the accuracy of the detection results.

[0003] Among existing improvement methods, obtaining multi-scale spatial information and adding attention mechanisms are two important technical directions. By adding additional feature pyramids, such as YOLOv8-P2, and fusing feature maps at different levels, the model can more accurately locate and classify small objects by combining the strong semantic information of deep features with the strong positional information of shallow features.

[0004] For example, Chinese patent document CN119229261A discloses a corn pest identification method based on improved YOLOv8, comprising obtaining a corn pest image; constructing an improved YOLOv8 model; the improved YOLOv8 model comprising a backbone network, a neck network, and a head network; utilizing the backbone network to extract features from the corn pest image and fusing the feature information based on deformable convolution to generate a multi-scale feature map containing visual information at different levels; utilizing the neck network to fuse the multi-scale feature maps containing visual information at different levels to obtain a feature map fusing the multi-scale information; and utilizing the head network to predict the corn pest identification result based on the feature map fusing the multi-scale information.

[0005] Chinese patent document CN118365936A discloses a YOLOv8 pest detection method and device based on image preprocessing, including: labeling and dividing a self-built pest dataset; constructing an improved image preprocessing network to preprocess the pest images; constructing an improved YOLOv8 pest detection network, introducing an SSFF module that combines global semantic information at multiple scales of the image and a TFE module that captures fine local details of small targets into the YOLOv8 neck; inputting the preprocessed pest images into the improved YOLOv8 pest detection network to enhance feature extraction of small target pests; then using a synthetic dataset to evaluate the model and testing it on the self-built pest dataset.

[0006] However, existing technologies that simply add additional feature pyramids fail to fully leverage the complementarity of features at each level and require more computing resources, slowing down the model in practical applications. Adding an attention mechanism, by introducing an attention module, enhances the network's focus on key features and improves their expressiveness. However, the attention mechanism can lead to excessive focus on certain areas, causing others to be neglected. Excessive local focus can weaken global context.

[0007] In summary, the existing YOLOv8 model still has a series of limitations, such as insufficient feature extraction and weak ability to distinguish between targets and backgrounds. This may lead to inaccurate detection results when performing image detection such as crop pest and disease detection or other target image defect detection. Summary of the Invention

[0008] The present invention aims to overcome at least one of the above-mentioned shortcomings of the prior art and provide a tomato disease and pest detection method based on an improved YOLOv8, so as to solve the problems that the existing YOLOv8 model has limited effect in processing tomato disease and pest detection, is prone to accuracy degradation, and does not fully mine the key feature information of the target, making it impossible to effectively distinguish the target from the background.

[0009] The present invention also discloses a device loaded with a tomato disease and insect pest detection method based on improved YOLOv8.

[0010] The detailed technical solutions of the present invention are as follows: A tomato pest and disease detection method based on improved YOLOv8, the method comprising: S1. Acquire diseased tomato leaf images and preprocess them to construct an image sample set; wherein the tomato disease types in the diseased tomato leaf images include early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold; S2. Build a pest and disease detection model based on an improved YOLOv8 model, wherein the improved YOLOv8 model includes: replacing the original upsampling module of the YOLOv8 model with a multi-scale upsampling enhancement module MEUM, and adding a multi-scale feature extraction and multi-head attention module MSEPA to the neck network of the YOLOv8 model; S3. Divide the image sample set into a training set, a validation set, and a test set and use them to train and optimize the pest and disease detection model to obtain a target pest and disease detection model for tomato pest and disease detection.

[0011] Preferably, according to the present invention, in S1, obtaining images of diseased tomato leaves includes: selecting images of diseased tomato leaves in a natural background from the Tomato-Village dataset, the PlantDoc dataset, and the FieldPlant dataset, and combining images of diseased tomato leaves in a laboratory background from the PlantVillage dataset.

[0012] Preferably, according to the present invention, in S1, the tomato diseased leaf image is preprocessed, including: performing operations such as inverting, rotating or adding noise to the tomato diseased leaf image to enhance the image sample set; and using the labelImg tool to label the tomato diseased leaf image with pests and diseases.

[0013] According to the preferred embodiment of the present invention, in S2, the multi-scale upsampling enhancement module MEUM is used to input feature maps Perform upsampling to obtain feature maps ,Right now: (1); In formula (1): Represents the input feature map, the input size is ; Represents the input feature map Upsampling The feature map after , its size is ; and, through Convolution The obtained feature map Perform dimensionality reduction to obtain feature maps ,Right now: (2); In formula (2): Feature map The size is , that is, the number of channels is halved; Represents the real number domain, indicating that the dimensions of the feature map are: the number of channels is Gao Wei 、Width , whose elements are all real numbers; And, for the obtained feature map Perform average pooling to obtain a smooth feature map ,Right now: (3); In formula (3): represents average pooling, Represents the pooling kernel as , Indicates that the fill is 1, Indicates that the step size is 1; the feature map obtained by this operation The size remains unchanged.

[0014] According to the preferred embodiment of the present invention, in S2, the multi-scale upsampling enhancement module MEUM is also used to calculate the feature map after dimensionality reduction And the feature map after pooling The difference, and through Convolution Perform edge enhancement to obtain edge-enhanced feature maps ,Right now: (4); And use the The feature map is then The pooling convolution operation is performed to gradually extract feature maps of different scales, namely: (5); And, the extracted feature maps of different scales are spliced ​​to obtain a fused feature map ,Right now: (6); Among them, the fusion feature map As the output of the multi-scale upsampling enhancement module MEUM, it enters the next module of the YOLOv8 model.

[0015] Preferably, according to the present invention, in S2, the multi-scale feature extraction and multi-head attention module MSEPA is used to obtain a fusion feature map containing both original input features and multi-scale features, specifically including: Initial input After normalization layer , after a layer of convolution kernel for Convolution After that, it passes through a layer of convolution kernel for ,filling Convolution of 2 After that, we get the feature map ,Right now: (7); And, the feature map obtained After three different convolution operations to obtain the corresponding feature maps, a splicing operation is performed to form a spliced ​​feature map ,Right now: (8); In formula (8): Represents the feature map The convolution kernel is performed as , a convolution operation with a padding of 9 and a dilation rate of 3, Represents the feature map The convolution kernel is performed as , a convolution operation with a padding of 6 and a dilation rate of 3, Represents the feature map The convolution kernel is performed as , convolution operation with padding of 3 and dilation rate of 3; Represents a splicing operation; And, the obtained spliced ​​feature map Multilayer Perceptron After fusion and initial input Add together to get the first fusion feature map ,Right now: (9); Wherein, the multi-layer perceptron Use the GELU activation function.

[0016] According to the preferred embodiment of the present invention, in S2, the multi-scale feature extraction and multi-head attention module MSEPA is also used to obtain the first fusion feature map Perform enhancement operations, including: The first fusion feature map After normalization layer , after a layer Convolution After that, it passes through a layer of convolution kernel for ,filling Convolution of 3 After that, a feature map is formed; at the same time, the first fusion feature map After the normalization layer After that, we first use global average pooling Get channel-level features and then pass a layer of convolution Capture the dependencies between channels and then pass through the Sigmoid activation function Output global information; finally, multiply the feature map by the global information to obtain the intermediate output feature map , which is simple pixel attention: (10); And, the first fusion feature map After normalization layer , after global average pooling Get channel-level features, pass through GELU activation function, pass through Convolution , after Sigmoid activation function After that, the output result is combined with the first fusion feature map Only after normalization layer Multiply the output result after , and get the intermediate output feature map , that is, channel attention: (11); And, the first fusion feature map After normalization layer , after two layers of convolution Get the spatial level features and the attention coefficient of each pixel respectively, and then pass it through a Sigmoid activation function After that, the output result is combined with the first fusion feature map Only after normalization layer Multiply the output result after , and get the intermediate output feature map , that is, pixel attention: (12); And, the intermediate output feature map 、 、 Perform splicing operation, and the splicing result is passed through the multi-layer perceptron After fusion and the first fusion feature map Add together to get the second fusion feature map : (13); Among them, the second fusion feature map That is, the feature map after multi-scale feature extraction and multi-head attention module MSEPA enhancement.

[0017] In another aspect of the present invention, a device for implementing a tomato pest and disease detection method based on an improved YOLOv8 is provided, the device comprising: a data acquisition module for acquiring diseased tomato leaf images and preprocessing them to construct an image sample set; wherein the tomato disease types in the diseased tomato leaf images include early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold; A model construction module is used to build a pest and disease detection model based on an improved YOLOv8 model. The improved YOLOv8 model includes: replacing the original upsampling module of the YOLOv8 model with a multi-scale upsampling enhancement module MEUM, and adding a multi-scale feature extraction and multi-head attention module MSEPA to the neck network of the YOLOv8 model; The model training module is used to divide the image sample set into a training set, a validation set and a test set and to train and optimize the pest and disease detection model to obtain a target pest and disease detection model for tomato pest and disease detection.

[0018] In another aspect of the present invention, an electronic device is provided, comprising: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the tomato disease and pest detection method based on the improved YOLOv8 as described above.

[0019] In another aspect of the present invention, a machine-readable storage medium is provided, which stores executable instructions. When the instructions are executed, the machine executes the tomato pest and disease detection method based on the improved YOLOv8 as described above.

[0020] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention provides a tomato pest and disease detection method based on improved YOLOv8, which replaces the original upsampling module with a multi-scale upsampling enhancement module and adds a multi-scale feature extraction and multi-head attention module to the neck network of the YOLOv8 model to achieve multi-scale feature extraction, long-distance dependency modeling and key feature enhancement, thereby improving the feature expression ability of the model.

[0021] (2) The present invention constructs a pest and disease detection model based on the improved YOLOv8 model, which can efficiently detect and classify tomato pest and disease targets. It not only provides an innovative idea for agricultural disease detection, but can also be widely used in smart agriculture, environmental monitoring and target detection and other fields. It is expected to significantly optimize traditional detection methods and improve recognition accuracy and decision-making efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a flow chart of the tomato disease and pest detection method based on improved YOLOv8 described in the present invention.

[0023] Figure 2 This is a network structure diagram of the pest and disease detection model constructed based on the improved YOLOv8 model in Example 1 of the present invention.

[0024] Figure 3 4 is a network structure diagram of the multi-scale upsampling enhancement module MEUM in Example 1 of the present invention.

[0025] Figure 4 This is a network structure diagram of the multi-scale feature extraction and multi-head attention module MSEPA in Example 1 of the present invention.

[0026] Figure 5 This is a schematic diagram comparing the average precision of the method of the present invention and the conventional YOLOv8 model.

[0027] Figure 6 This is a schematic diagram comparing the accuracy of the method of the present invention and the conventional YOLOv8 model. DETAILED DESCRIPTION

[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0029] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0031] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0032] In response to the shortcomings of the existing technology, the present invention proposes a target detection method based on YOLOv8's multi-scale feature extraction and multi-head attention mechanism combined with multi-scale upsampling enhancement. The method aims to comprehensively enhance the feature extraction and target detection capabilities of the YOLOv8 model through the combination of multi-scale feature extraction and attention mechanism, so as to improve the accuracy of the YOLOv8 model in tomato disease and insect detection.

[0033] The tomato pest and disease detection method and device based on improved YOLOv8 of the present invention are further described below with reference to specific embodiments.

[0034] Example 1 Ginseng Figure 1 This embodiment provides a method for detecting tomato pests and diseases based on an improved YOLOv8, the method comprising: S1. Acquire diseased tomato leaf images and preprocess them to construct an image sample set; wherein the tomato disease types in the diseased tomato leaf images include early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold.

[0035] In this example, we selected diseased tomato leaf images from natural backgrounds (Tomato-Village, PlantDoc, and FieldPlant) and combined them with laboratory background images from PlantVillage. We selected tomato early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold as the research subjects.

[0036] The tomato diseased leaf image is preprocessed, including: performing operations such as inverting, rotating, and adding noise to the tomato diseased leaf image to enhance the image sample set, and using the labelImg tool to label the tomato diseased leaf image with pests and diseases.

[0037] S2. Build a pest and disease detection model based on the improved YOLOv8 model. The improved YOLOv8 model includes: using the multi-scale upsampling enhancement module MEUM to replace the original upsampling module of the YOLOv8 model, and adding a multi-scale feature extraction and multi-head attention module MSEPA to the neck network of the YOLOv8 model.

[0038] In this embodiment, the multi-scale upsampling enhancement module MEUM is used to replace the original upsampling module of the YOLOv8 model to solve the problem of edge information loss that is easily caused in the original YOLOv8 model.

[0039] Specifically, Figure 3 As shown, the multi-scale upsampling enhancement module MEUM performs on the input feature map Perform the following operations: First, the input feature map Perform upsampling operation to obtain feature map : (1); In formula (1): Represents the input feature map, the input size is ; Represents the input feature map Upsampling The feature map after , its size is .

[0040] The above operation uses bilinear interpolation upsampling as the original YOLOv8 model, and the input feature map The adjacent pixel values ​​of produce smoother results, but at the same time may cause blurred image edges, so further processing is required for edge enhancement.

[0041] That is, through Convolution For the feature map obtained above Perform dimensionality reduction to obtain feature maps : (2); In formula (2): Feature map The size is , that is, the number of channels is halved; Represents the real number domain, indicating that the dimensions of the feature map are: the number of channels is Gao Wei 、Width , whose elements are all real numbers.

[0042] Then, the obtained feature map Perform average pooling to obtain a smooth feature map : (3); In formula (3): represents average pooling, Represents the pooling kernel as , Indicates that the fill is 1, Indicates that the step size is 1; the feature map obtained by this operation The size remains unchanged.

[0043] Then, the feature map after dimensionality reduction is calculated And the feature map after pooling The difference between the two is used to highlight the edge part and enhance the edge features; Convolution Perform edge enhancement to obtain the feature map after edge enhancement : (4).

[0044] Then, use the The feature map is then The pooling convolution operation is repeated to gradually extract feature maps of different scales: (5).

[0045] Finally, the feature maps output above are subjected to channel splicing operation, that is, the feature maps extracted at different scales are merged to obtain a fused feature map containing information at different scales. . Instead of the original upsampling output entering the subsequent modules of the YOLOv8 model for processing, it plays a connecting role in the method of the present invention: (6).

[0046] Based on the above, the MEUM module network is able to simultaneously utilize information at multiple scales to enhance image understanding. Through these steps, the MEUM module not only enhances edge information but also provides higher-quality input features for the subsequent attention mechanism, thereby improving model performance.

[0047] Furthermore, in this embodiment, a multi-scale feature extraction and multi-head attention module MSEPA is added to the neck network of the YOLOv8 model to enhance the model's feature extraction capability at different scales.

[0048] Specifically, Figure 4 As shown, the multi-scale feature extraction and multi-head attention module MSEPA are used to extract the initial input Do the following: First, the initial input After normalization layer (Batch Norm), after a layer of convolution kernel for Convolution After that, it passes through a layer of convolution kernel for ,filling Convolution of 2 After that, we get the feature map : (7).

[0049] Then, the feature map obtained above After three different convolution operations, the corresponding feature maps are obtained and then spliced ​​to form a spliced ​​feature map : (8); In formula (8): Represents the feature map The convolution kernel is performed as , a convolution operation with a padding of 9 and a dilation rate of 3, Represents the feature map The convolution kernel is performed as , a convolution operation with a padding of 6 and a dilation rate of 3, Represents the feature map The convolution kernel is performed as , convolution operation with padding of 3 and dilation rate of 3; Represents a splicing operation.

[0050] The above multi-scale feature extraction captures large-scale and small-scale features in the image by using convolution kernels of different sizes in parallel, and uses dilated convolution kernels to obtain a large receptive field.

[0051] Then, the spliced ​​feature map obtained above is Multilayer Perceptron After fusion and initial input Add together to get the first fusion feature map : (9); Among them, the multi-layer perceptron The GELU activation function is used in .

[0052] Through the residual connection, the multi-layer perceptron The fused feature map is added to the original feature map to obtain the first fused feature map that contains both the original input features and the multi-scale features. .

[0053] Then, the first fusion feature map After normalization layer , after a layer Convolution After that, it passes through a layer of convolution kernel for ,filling Convolution of 3 After that, a feature map is formed; at the same time, the first fusion feature map After the normalization layer After that, we first use global average pooling Get channel-level features and then pass a layer of convolution Capture the dependencies between channels and then pass through the Sigmoid activation function Output global information; finally, multiply the feature map by the global information to obtain the intermediate output feature map , which is simple pixel attention: (10).

[0054] And, the first fusion feature map After normalization layer , after global average pooling Get channel-level features, pass through GELU activation function, pass through Convolution , after Sigmoid activation function After that, the output result is combined with the first fusion feature map Only after normalization layer Multiply the output result after , and get the intermediate output feature map , that is, channel attention: (11); In the above operation, two layers of convolution and activation functions are used to capture the first fusion feature map. Dependencies between channels.

[0055] And, the first fusion feature map After normalization layer , after two layers of convolution Get the spatial level features and the attention coefficient of each pixel respectively, and then pass it through a Sigmoid activation function After that, the output result is combined with the first fusion feature map Only after normalization layer Multiply the output result after , and get the intermediate output feature map , that is, pixel attention: (12).

[0056] Finally, the intermediate output feature map 、 、 Perform splicing operation, and the splicing result is passed through the multi-layer perceptron After fusion and the first fusion feature map Add together to get the second fusion feature map : (13).

[0057] The second fusion feature map obtained above This is the feature map after multi-scale feature extraction and multi-head attention module MSEPA enhancement. It will continue to enter the subsequent convolution processing or enter the detection head to perform the detection task.

[0058] Based on the above model improvements, the network structure of the pest detection model constructed in this embodiment is as follows: Figure 2 shown.

[0059] S3. Divide the image sample set into a training set, a validation set, and a test set and use them to train and optimize the pest and disease detection model to obtain a target pest and disease detection model.

[0060] Specifically, the preprocessed image sample set was randomly divided into training, validation, and test sets in a ratio of 7:2:1. These sets were then fed into the improved YOLOv8 model and trained with parameter adjustments. The YOLOv8 model was adjusted and the corresponding training parameters were configured, including setting the number of epoches to 300, the number of batches to 32, and the number of imgsz to 640. Training was terminated after reaching 300 rounds, at which point the loss function or mAP accuracy had stabilized, indicating that the trained target pest and disease detection model was complete.

[0061] Finally, the trained target pest and disease detection model is deployed in practical applications to realize pest and disease detection on diseased tomato leaf images.

[0062] Figure 5 The results show a comparison of the mAP50-95 (mean Average Precision) metric between the proposed method and the conventional YOLOv8 model. mAP50-95 is a key metric for measuring the overall performance of an object detection algorithm, encompassing detection results at IoU thresholds from 0.5 to 0.95. The figure clearly shows that the proposed method achieves higher mAP values ​​across multiple IoU thresholds, demonstrating superior detection accuracy compared to the conventional YOLOv8. This demonstrates the significant effectiveness of the introduced Multi-Scale Upsampling Enhancement Module (MEUM) and Multi-Scale Feature Extraction and Multi-Head Attention Module (MSEPA) in improving the model's feature extraction capabilities, particularly for small objects and complex backgrounds.

[0063] Figure 6 The comparison shows the performance of the proposed method and the conventional YOLOv8 model in terms of Precision. Precision reflects how many of the targets detected by the model are truly correct results and is an important indicator for measuring the false positive rate of the model. Figure 6 As can be seen from the results, the proposed method maintained a high accuracy throughout the entire testing process, with fewer false positives than the conventional YOLOv8. This demonstrates that the proposed structural optimization effectively suppresses background interference and enhances the ability to discriminate pests and diseases, further improving the robustness and practical application value of the model.

[0064] Example 2 This embodiment provides a device for implementing a tomato pest and disease detection method based on an improved YOLOv8, the device comprising: a data acquisition module for acquiring diseased tomato leaf images and preprocessing them to construct an image sample set; wherein the tomato disease types in the diseased tomato leaf images include early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold; A model construction module is used to build a pest and disease detection model based on the improved YOLOv8 model. The improved YOLOv8 model includes: replacing the original upsampling module of the YOLOv8 model with a multi-scale upsampling enhancement module MEUM, adding a multi-scale feature extraction and multi-head attention module MSEPA to the neck network of the YOLOv8 model, and replacing the CIoU loss function with the SDIoU loss function; The model training module is used to divide the image sample set into a training set, a validation set and a test set and to train and optimize the pest and disease detection model to obtain a target pest and disease detection model for tomato pest and disease detection.

[0065] Example 3 This embodiment further provides an electronic device, including: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the tomato disease and pest detection method based on the improved YOLOv8 as described above.

[0066] In this embodiment, electronic devices may include, but are not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and the like.

[0067] Example 4 This embodiment also provides a machine-readable storage medium storing executable instructions, which, when executed, enable the machine to perform the tomato disease and pest detection method based on the improved YOLOv8 as described above.

[0068] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes that implement the functions of any of the above-mentioned embodiments are stored, and a computer or processor of the system or device can read and execute instructions stored in the readable storage medium.

[0069] In this case, the program code itself read from the machine-readable medium can implement the functions of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.

[0070] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.

[0071] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0072] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0073] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0075] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the technical solutions of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A tomato pest and disease detection method based on improved YOLOv8, characterized in that: The method comprises: S1. Acquire diseased tomato leaf images and preprocess them to construct an image sample set; wherein the tomato disease types in the diseased tomato leaf images include early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold; S2. Build a pest and disease detection model based on an improved YOLOv8 model, wherein the improved YOLOv8 model includes: replacing the original upsampling module of the YOLOv8 model with a multi-scale upsampling enhancement module MEUM, and adding a multi-scale feature extraction and multi-head attention module MSEPA to the neck network of the YOLOv8 model; S3. Divide the image sample set into a training set, a validation set, and a test set and use them to train and optimize the pest and disease detection model to obtain a target pest and disease detection model for tomato pest and disease detection.

2. The tomato pest and disease detection method based on improved YOLOv8 according to claim 1, characterized in that: In the above S1, acquiring diseased tomato leaf images includes: selecting diseased tomato leaf images in natural background from the Tomato-Village dataset, the PlantDoc dataset, and the FieldPlant dataset, and combining diseased tomato leaf images in laboratory background from the PlantVillage dataset.

3. The tomato pest and disease detection method based on improved YOLOv8 according to claim 2, characterized in that: In S1, the tomato diseased leaf image is preprocessed, including: performing operations such as inverting, rotating or adding noise to the tomato diseased leaf image to enhance the image sample set; and using the labelImg tool to label the tomato diseased leaf image with pests and diseases.

4. The tomato pest and disease detection method based on improved YOLOv8 according to claim 1, characterized in that: In S2, the multi-scale upsampling enhancement module MEUM is used to input feature maps Perform upsampling to obtain feature maps ,Right now: (1); In formula (1): Represents the input feature map, the input size is ; Represents the input feature map Upsampling The feature map after , its size is ; and, through Convolution The obtained feature map Perform dimensionality reduction to obtain feature maps ,Right now: (2); In formula (2): Feature map The size is , that is, the number of channels is halved; Represents the real number domain, indicating that the dimensions of the feature map are: the number of channels is Gao Wei 、Width , whose elements are all real numbers; And, for the obtained feature map Perform average pooling to obtain a smooth feature map ,Right now: (3); In formula (3): represents average pooling, Represents the pooling kernel as , Indicates that the fill is 1, Indicates that the step size is 1; the feature map obtained by this operation The size remains unchanged.

5. The tomato pest and disease detection method based on improved YOLOv8 according to claim 4, characterized in that: In S2, the multi-scale upsampling enhancement module MEUM is also used to calculate the feature map after dimensionality reduction And the feature map after pooling The difference, and through Convolution Perform edge enhancement to obtain edge-enhanced feature maps ,Right now: (4); And use the The feature map is then The pooling convolution operation is performed to gradually extract feature maps of different scales, namely: (5); And, the extracted feature maps of different scales are spliced ​​to obtain a fused feature map ,Right now: (6); Among them, the fusion feature map As the output of the multi-scale upsampling enhancement module MEUM, it enters the next module of the YOLOv8 model.

6. The tomato pest and disease detection method based on improved YOLOv8 according to claim 1, characterized in that: In S2, the multi-scale feature extraction and multi-head attention module MSEPA is used to obtain a fusion feature map containing both original input features and multi-scale features, specifically including: Initial input After normalization layer , after a layer of convolution kernel for Convolution After that, it passes through a layer of convolution kernel for ,filling Convolution of 2 After that, we get the feature map ,Right now: (7); And, the feature map obtained After three different convolution operations to obtain the corresponding feature maps, a splicing operation is performed to form a spliced ​​feature map ,Right now: (8); In formula (8): Represents the feature map The convolution kernel is performed as , a convolution operation with a padding of 9 and a dilation rate of 3, Represents the feature map The convolution kernel is performed as , a convolution operation with a padding of 6 and a dilation rate of 3, Represents the feature map The convolution kernel is performed as , convolution operation with padding of 3 and dilation rate of 3; Represents a splicing operation; And, the obtained spliced ​​feature map Multilayer Perceptron After fusion and initial input Add together to get the first fusion feature map ,Right now: (9); Wherein, the multi-layer perceptron Use the GELU activation function.

7. The tomato pest and disease detection method based on improved YOLOv8 according to claim 6, characterized in that: In S2, the multi-scale feature extraction and multi-head attention module MSEPA is also used to obtain the first fusion feature map Perform enhancement operations, including: The first fusion feature map After normalization layer , after a layer Convolution After that, it passes through a layer of convolution kernel for ,filling Convolution of 3 After that, a feature map is formed; at the same time, the first fusion feature map After the normalization layer After that, first pass the global average pooling Get channel-level features and then pass a layer of convolution Capture the dependencies between channels and then pass through the Sigmoid activation function Output global information; finally, multiply the feature map by the global information to obtain the intermediate output feature map , which is simple pixel attention: (10); And, the first fusion feature map After normalization layer , after global average pooling Get channel-level features, pass through GELU activation function, pass through Convolution , after Sigmoid activation function After that, the output result is combined with the first fusion feature map Only after normalization layer Multiply the output result after , and get the intermediate output feature map , that is, channel attention: (11); And, the first fusion feature map After normalization layer , after two layers of convolution Get the spatial level features and the attention coefficient of each pixel respectively, and then pass it through a Sigmoid activation function After that, the output result is combined with the first fusion feature map Only after normalization layer Multiply the output result after , and get the intermediate output feature map , that is, pixel attention: (12); And, the intermediate output feature map 、 、 Perform splicing operation, and the splicing result is passed through the multi-layer perceptron After fusion and the first fusion feature map Add together to get the second fusion feature map : (13); Among them, the second fusion feature map That is, the feature map after multi-scale feature extraction and multi-head attention module MSEPA enhancement.

8. A device for implementing a tomato pest and disease detection method based on improved YOLOv8, characterized in that: The device comprises: a data acquisition module for acquiring diseased tomato leaf images and preprocessing them to construct an image sample set; wherein the tomato disease types in the diseased tomato leaf images include early blight, late blight, bacterial spot, mosaic virus, yellows virus, and leaf mold; A model construction module is used to build a pest and disease detection model based on an improved YOLOv8 model. The improved YOLOv8 model includes: replacing the original upsampling module of the YOLOv8 model with a multi-scale upsampling enhancement module MEUM, and adding a multi-scale feature extraction and multi-head attention module MSEPA to the neck network of the YOLOv8 model; The model training module is used to divide the image sample set into a training set, a validation set and a test set and to train and optimize the pest and disease detection model to obtain a target pest and disease detection model for tomato pest and disease detection.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the tomato disease and pest detection method based on the improved YOLOv8 according to any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that The machine-readable storage medium stores executable instructions, which, when executed, enable the machine to execute the tomato disease and insect pest detection method based on improved YOLOv8 according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • YOLOv8 pest detection method and device based on image preprocessing

    CN118365936A

  • Corn pest identification method based on improved YOLOv8

    CN119229261A

Cited By

  • Crop leaf disease and pest detection method and system

    CN121438115A