Fruit tree pest detection method and device based on improved YOLOv11

Through the improved YOLOv11 model, combined with CSP-PMFEM and LKAP modules, the problems of target size diversity, complex background interference and large differences in pest and disease characteristics in fruit tree leaf disease detection are solved, and efficient and accurate pest and disease detection and model robustness are achieved.

CN120147871APending Publication Date: 2025-06-13NANJING FORESTRY UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510234326.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has problems such as diversity of target size, complex background interference and large differences in pest and disease characteristics in fruit tree leaf disease detection, resulting in insufficient detection accuracy and robustness.

Method used

The improved YOLOv11 model integrates advanced feature extraction, hierarchical feature fusion and enhanced spatial perception functions by introducing CSP-PMFEM module and LKAP module, improving the robustness and generalization capabilities of the detection model.

Benefits of technology

It realizes efficient and accurate detection of pests and diseases of fruit tree leaves, improves the robustness and generalization ability of the detection model, and maintains high performance especially in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147871A_ABST
    Figure CN120147871A_ABST
Patent Text Reader

Abstract

The invention discloses a fruit tree disease and pest detection method and device based on improved YOLOv11, relates to the technical field of automatic machine learning, and aims to solve the problem that YOLOv11 in the prior art is still difficult in apple disease and pest detection. Taking the fruit tree shot image as input, and outputting based on the fruit tree pest detection model to obtain a fruit tree processing graph; and judging whether the fruit tree has diseases and pests and the types of the diseases and pests according to the fruit tree processing graph, wherein the fruit tree disease and pest detection model is obtained based on improvement of a YOLOv11 model. According to the method, advanced feature extraction, layered feature fusion and enhanced spatial perception functions are integrated, efficient and accurate detection of fruit tree diseases and insect pests is realized, the robustness and generalization ability of a detection model are improved, and meanwhile, relatively high performance is kept in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for detecting fruit tree diseases and pests based on an improved YOLOv11, belonging to the technical field of automated machine learning. Background Art

[0002] As one of the important fruit trees widely planted globally, the yield and economic value of fruit trees occupy an important position in the agricultural industry. However, fruit trees are often invaded by various diseases and pests during the growth process, which not only affects the yield and quality of fruits, but also brings huge economic losses to fruit farmers. For example, fruit tree aphid disease, fruit tree cotton aphid disease, fruit tree gray rot disease, and fruit tree scab disease are common diseases and pests in the process of fruit tree planting. These diseases and pests usually cause fruit rot, leaf damage, and even affect the growth of the whole tree. According to research, the actual loss rate of yield per hectare in non-prevention areas (i.e., areas without disease and pest prevention) is as high as 41.26%. These diseases and pests will not only reduce the yield and quality of fruit trees, but also cause growers to invest a large amount of time and cost in prevention and control. Therefore, it is of great significance to detect and control diseases and pests on fruit tree leaves in a timely and accurate manner.

[0003] Traditional methods for detecting diseases and pests mainly rely on methods such as manual observation, unmanned aerial vehicle monitoring, and sensor technology. These methods not only consume a large amount of human and material resources, are limited in application scenarios, but are also easily interfered by environmental factors, resulting in inaccurate detection results. In recent years, with the development of computer vision and deep learning technologies, automated methods for detecting diseases and pests based on image processing have gradually been applied in the agricultural field. Compared with traditional manual detection, computer vision technology can process a large number of images in a short time and efficiently and accurately identify diseases and pests, which provides technical support for the real-time detection of fruit tree diseases and pests.

[0004] Traditional methods for detecting diseases and pests mostly rely on manual visual inspection and laboratory chemical analysis. Although these methods can provide a certain degree of accuracy, they are usually time-consuming and laborious and are easily affected by environmental changes. In the past few years, the emergence of machine learning has shined in object detection.

[0005] Based on the above domestic and international status quo, although YOLOv11 has made significant improvements in performance, it still faces some difficulties and challenges when applied to the detection of fruit tree leaf diseases and pests:

[0006] First, there is a problem of target size diversity in the process of detecting fruit tree leaf diseases and pests. The sizes of diseases and pests on fruit tree leaves vary significantly. In particular, some small targets (such as aphids on leaves) are extremely easy to be ignored in detection. Although YOLOv11 has the ability to extract multi-scale features, it still has limitations in detecting extremely small targets, especially when the colors of diseases and pests are similar to the leaf background, which is prone to missed detection.

[0007] Second, there will be interference from complex backgrounds. In the natural environment, the detection of leaf pests and diseases in fruit trees will be interfered by complex backgrounds, such as sunlight, shadows, overlapping leaves, and the presence of other non-target objects. These background factors are likely to interfere with the discriminative ability of the model. The performance of YOLOv11 under complex backgrounds still needs to be optimized. Especially under high-light and low-contrast conditions, the detection accuracy of the model often decreases.

[0008] Third, there are significant differences in the characteristics of pests and diseases. There are many types of pests and diseases on the leaves of fruit trees, and the visual characteristics of different pests and diseases are significantly different. Some diseases show similar visual characteristics, such as the similar spot shapes and colors of different disease types, which are prone to false detections and missed detections. In addition, pests and diseases will present different appearance characteristics as time goes by, which poses higher requirements for the generalization ability of the model. Summary of the Invention

[0009] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and device for detecting pests and diseases in fruit trees based on an improved YOLOv11. It integrates advanced feature extraction, hierarchical feature fusion, and enhanced spatial perception functions, realizes efficient and accurate detection of pests and diseases on fruit tree leaves, improves the robustness and generalization ability of the detection model, and maintains high performance in complex environments.

[0010] To achieve the above purpose, the present invention is implemented by the following technical solutions:

[0011] On the one hand, the present invention provides a method for detecting pests and diseases in fruit trees based on an improved YOLOv11, including:

[0012] Obtain the captured images of fruit trees;

[0013] Use the captured images of fruit trees as input and obtain the processed images of fruit trees based on the output of the fruit tree pest and disease detection model;

[0014] Judge whether there are pests and diseases on the fruit trees and the types of pests and diseases according to the processed images of fruit trees;

[0015] The fruit tree pest and disease detection model is improved based on the YOLOv11 model. The YOLOv11 model includes a backbone, a neck, and a detection head connected in sequence. Replace the C3k2 module in the backbone with a CSP-PMFEM module, replace the SPPF module with an LKAP module, and set an efficient local attention module in the neck to obtain the fruit tree pest and disease detection model;

[0016] The CSP-PMFEM module is used to segment the feature map and extract features of different scales using multi-scale convolutional kernels. The LKAP module is used to capture context information by introducing decomposed large convolutional kernels. The efficient local attention module is used to perform multi-scale feature fusion on the feature map.

[0017] Further, the input feature map of the CSP-PMFEM module is segmented to obtain a segmented feature map. The segmented feature map passes through a convolutional layer and multiple partial multi-scale feature extraction modules in sequence and is concatenated with the input feature map to obtain the output feature map of the CSP-PMFEM module;

[0018] The partial multi-scale feature extraction module includes a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence. The input feature map of the partial multi-scale feature extraction module is divided into a first partial feature map and a second partial feature map along the channel direction after extracting small target features through the first convolutional layer. The first partial feature map is divided into a third partial feature map and a fourth partial feature map along the channel direction again after capturing medium-scale target features through the second convolutional layer. The third partial feature map is concatenated with the second partial feature map and the fourth partial feature map after extracting large-scale target features through the third convolutional layer, and is connected with the input feature map through a residual connection via a fourth convolutional layer to obtain the output feature map of the partial multi-scale feature extraction module. The input feature map of the CSP-PMFEM module is segmented to obtain an input feature map. The segmented feature map passes through a convolutional layer and multiple partial multi-scale feature extraction modules in sequence and is concatenated with the input feature map to obtain the output feature map of the CSP-PMFEM module;

[0019] The partial multi-scale feature extraction module includes a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence. The input feature map of the partial multi-scale feature extraction module is divided into a first partial feature map and a second partial feature map along the channel direction after extracting small target features through the first convolutional layer. The first partial feature map is divided into a third partial feature map and a fourth partial feature map along the channel direction again after capturing medium-scale target features through the second convolutional layer. The third partial feature map is concatenated with the second partial feature map and the fourth partial feature map after extracting large-scale target features through the third convolutional layer, and is connected with the input feature map through a residual connection via a fourth convolutional layer to obtain the output feature map of the partial multi-scale feature extraction module.

[0020] Further, the data processing expression of the CSP-PMFEM module is as follows:

[0021] ;

[0022] ;

[0023] ;

[0024] ;

[0025] ;

[0026] ;

[0027] ;

[0028] ;

[0029] Among them, represents the input feature map of the module, Split represents the splitting operation, represents a part of the feature map after splitting, represents the other part of the feature map after splitting, represents the input feature map of the partial multi-scale feature extraction module, represents the input of the first branch, represents the output of the first branch, represents the input of the second branch, represents the output of the second branch, represents the input of the third branch, represents the output of the third branch, represents the output after the concatenated feature map is processed by the fourth convolutional layer, represents the output feature map of the partial multi-scale feature extraction module, represents the output feature map of the nth partial multi-scale feature extraction module, represents the output feature map of the module.

[0030] Furthermore, the LKAP module includes a first convolutional layer, which is used to reduce the number of channels of the input feature map by half. Its output passes through multiple max-pooling layers in sequence. The max-pooling layer is used to extract the global information of the image from different scales. The output of each max-pooling layer is concatenated with the output of the convolutional layer in channels to obtain a multi-scale feature combination. The multi-scale feature combination passes through a large separable convolutional attention module and a second convolutional layer in sequence to obtain an output feature map. The large separable convolutional attention module is used to expand the convolutional receptive field in the spatial domain to obtain an output feature map, and the second convolutional layer is used to restore the number of channels of the feature map.

[0031] Furthermore, the number of max-pooling layers is 3, and the data processing expression of the LKAP module is as follows:

[0032] ;

[0033] ;

[0034] ;

[0035] ;

[0036] ;

[0037] Among them, represents the input feature map of the module, represents the output of the first convolutional layer, represents the output of the first max pooling layer, represents the output of the second max pooling layer, represents the output of the third max pooling layer, represents the multi-scale feature combination, represents the output of the large-scale separable convolutional attention module, represents the output feature map of the module.

[0038] Furthermore, the backbone is used to extract multi-scale image features to obtain three feature maps with different scales, which are the first feature map, the second feature map, and the third feature map respectively;

[0039] The neck is used to perform feature fusion on the first feature map, the second feature map, and the third feature map. It includes multiple efficient local attention modules. The first feature map passes through the efficient local attention module and the two-dimensional convolutional layer in sequence to obtain the first high-resolution feature map;

[0040] The second feature map passes through three branches respectively. Among them, the first branch includes a transposed convolutional layer and an efficient local attention module connected in sequence, the second branch includes an efficient local attention module and a two-dimensional convolutional layer connected in sequence, and the third branch includes a transposed convolutional layer and an efficient local attention module connected in sequence;

[0041] The third feature map passes through the efficient local attention module and the two-dimensional convolutional layer in sequence to obtain the third high-resolution feature map;

[0042] The output of the first branch is multiplied element-wise with the first high-resolution feature map and then concatenated to obtain the first fusion feature map;

[0043] The third high-resolution feature map is input into the transposed convolutional layer of the third branch. The outputs of the second branch and the third branch are multiplied element-wise and concatenated with the output of the transposed convolutional layer of the third branch to obtain the second fusion feature map;

[0044] The first fusion feature map, the second fusion feature map, and the third high-resolution feature map are respectively input into the detection head through the CSP-PMFEM module.

[0045] Furthermore, it also includes pre-training the fruit tree pest and disease detection model. The pre-training method includes:

[0046] S1. Obtain the fruit tree pest and disease image dataset;

[0047] S2. Label and augment the image dataset of fruit tree pests and diseases to obtain a preprocessed image dataset of fruit tree pests and diseases;

[0048] S3. Divide the preprocessed image dataset of fruit tree pests and diseases into a training set, a test set, and a validation set;

[0049] S4. Use the training set data as input to train the fruit tree pest and disease detection model, and adjust the model parameters using the loss function during training; use the validation set data as input to validate the fruit tree pest and disease detection model to obtain a validation result, and adjust the hyperparameters according to the validation result;

[0050] S5. Use the test set data as input to test the fruit tree pest and disease detection model to obtain a test result, calculate evaluation metrics based on the test result. If the evaluation metrics are higher than the preset value, the training is completed, and a pre-trained fruit tree pest and disease detection model is obtained. If the evaluation metrics are lower than the preset value, repeat S4 - S5 until the evaluation metrics are higher than the preset value, and a pre-trained fruit tree pest and disease detection model is obtained.

[0051] Further, the annotation information includes the pest and disease area and the type of pest and disease. The data augmentation includes random rotation and flipping, color transformation, random cropping and scaling, blurring, noise addition, and geometric transformation. The evaluation metrics include mean average precision, model parameters, model size, and computational efficiency.

[0052] Further, determining whether there are pests and diseases on the fruit tree and the type of pests and diseases according to the fruit tree processing diagram includes:

[0053] If there are pests and diseases on the fruit tree, the fruit tree processing diagram includes a detection frame, the detection frame is the pest and disease area, and the type and confidence of the pest and disease are marked on the detection frame;

[0054] If there are no pests and diseases on the fruit tree, the fruit tree processing diagram is the fruit tree captured image.

[0055] On the other hand, the present invention also provides a fruit tree pest and disease detection device based on an improved YOLOv11, including:

[0056] A fruit tree image acquisition module configured to acquire a fruit tree captured image;

[0057] A pest and disease detection module configured to use the fruit tree captured image as input and output a fruit tree processing diagram based on the fruit tree pest and disease detection model;

[0058] A pest and disease judgment module configured to determine whether there are pests and diseases on the fruit tree and the type of pests and diseases according to the fruit tree processing diagram.

[0059] Compared with the prior art, the beneficial effects achieved by the present invention:

[0060] The present invention innovatively proposes a CSP-PMFEM module. The Partial Multi-Scale Feature Extraction Module (PMFEM) combines multi-scale convolution operations and the CSP idea to achieve more efficient feature representation. This module uses multi-scale convolution to process the input features in blocks respectively to capture feature information of different scales, and enhances the representation ability of the network through feature aggregation, ultimately effectively improving the detection accuracy of fruit tree pests and diseases. Especially in complex environments and lighting conditions, it demonstrates more excellent detection effects and feature extraction capabilities.

[0061] The present invention improves the backbone of the YPLOv11 model to construct a hierarchical feature pyramid network, realizing the efficient fusion of multi-scale features. The LKAP module effectively expands the receptive field range of the network by introducing a large-scale kernel attention mechanism. While maintaining the computational efficiency, this module enhances the ability to obtain spatial features through position-sensitive attention calculation, and is particularly suitable for processing disease regions with irregular shapes on the surface of fruit trees. In practical applications, the LKAP module significantly improves the precise positioning ability of the model for disease boundaries and enhances the detection accuracy.

[0062] The fruit tree pest and disease detection model provided by the present invention significantly improves the network's detection ability for disease targets of different sizes. Especially for tiny disease spots and initial disease symptoms on fruit trees and their leaves, it demonstrates excellent detection performance and greatly improves the detection recall rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic structural diagram of the existing YOLOv11 model;

[0064] Figure 2 It is a schematic structural diagram of the fruit tree pest and disease detection model in an embodiment of the fruit tree pest and disease detection method based on the improvement of YOLOv11 according to the present invention;

[0065] Figure 3 It is a schematic structural diagram of the CSP-PMFEM module of the fruit tree pest and disease detection model in an embodiment of the fruit tree pest and disease detection method based on the improvement of YOLOv11 according to the present invention;

[0066] Figure 4 It is a schematic structural diagram of the LKAP module of the fruit tree pest and disease detection model in an embodiment of the fruit tree pest and disease detection method based on the improvement of YOLOv11 according to the present invention;

[0067] Figure 5 It is a comparison schematic diagram of map@50 (mean average precision) of different models in Embodiment 2 of the present invention;

[0068] Figure 6Schematic diagram for comparison of effective receptive fields of YOLOv11 and YOLO-PEL in Embodiment 2 of the present invention;

[0069] Figure 7 Schematic diagram for comparison of heat maps generated by YOLOv11 and YOLO-PEL in Embodiment 2 of the present invention. Among them, (A), (D), (G), and (J) are apple aphid disease, apple woolly aphid disease, apple gray mold disease, and apple scab disease labeled by LabeImg respectively; (B), (E), (H), and (K) are the heat maps generated by (A), (D), (G), and (J) under YOLOv11 respectively; (C), (F), (I), and (L) are the heat maps generated by (A), (D), (G), and (J) under YOLO-PEL respectively;

[0070] Figure 8 Visualization diagram of the detection results of different target pests and diseases by YOLOv11 in Embodiment 2 of the present invention. Among them, (A) is apple aphid disease, (B) is apple woolly aphid disease, (C) is apple gray mold disease, and (D) is apple scab disease;

[0071] Figure 9 Visualization diagram of the detection results of different target pests and diseases by YOLO-PEL in Embodiment 2 of the present invention. Among them, (E) is apple aphid disease, (F) is apple woolly aphid disease, (G) is apple gray mold disease, and (H) is apple scab disease. Detailed implementation method

[0072] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0073] Embodiment 1:

[0074] The embodiment of the present invention provides a fruit tree pest and disease detection method improved based on YOLOv11, including the following steps:

[0075] Obtain fruit tree photographed images. In this embodiment, the apple photographed images are three-channel (RGB) color images obtained by a Nikon 7200d camera, and the image resolution is 4000×6000 pixels.

[0076] Take the fruit tree photographed images as input, and output the fruit tree processed images based on the fruit tree pest and disease detection model (YOLO-PEL).

[0077] If there are no pests and diseases on the fruit tree, the fruit tree processed images are the same as the fruit tree photographed images, without any processing marks.

[0078] If there are pests and diseases on the fruit tree, a detection frame is marked on the fruit tree processing diagram. The area within the detection frame is the pest and disease area, and the type and confidence level of the pests and diseases are marked on the detection frame. The confidence level is one of the indicators to measure the accuracy of the model's judgment.

[0079] Embodiment 2:

[0080] Based on Embodiment 1, the structure of the fruit tree pest and disease detection model is introduced in detail in this embodiment.

[0081] The structure of the traditional YOLOv11 model is as Figure 1 shown, which includes a backbone, a neck, and a detection head connected in sequence. The fruit tree pest and disease detection model in this embodiment is improved based on the YOLOv11 model, as Figure 2 shown:

[0082] The backbone is used to extract multi-scale image features and further optimize the efficiency. However, the traditional YOLOv11 backbone network has obvious limitations when dealing with fruit tree pest and disease detection tasks. The C3k2 module used in its backbone network can extract multi-level features, but it performs poorly when dealing with fruit tree pest and disease images with complex backgrounds and multi-scale features. Specifically, the C3k2 module uses a fixed-size convolutional kernel for feature extraction, which leads to insufficient adaptability when dealing with pest and disease targets of different scales. Especially when dealing with small pest and disease targets, due to the limitations of feature extraction ability, missed detections are likely to occur.

[0083] Therefore, in the fruit tree pest and disease detection model of this embodiment, the C3k2 module in the backbone is replaced with the CSP-PMFEM module.

[0084] The CSP-PMFEM module innovatively integrates the CSP idea. By integrating multi-scale feature extraction and channel segmentation strategies, it significantly improves the model's detection ability for targets of different scales. The CSP_PMFEM module mainly includes two parts: the CrossStage Partial part and the Part Multi-Scale Feature Extraction Module part. The CSP part reduces the computational amount by splitting the feature map and promotes the diversity of feature expression, while the PMFEM part uses multi-scale convolutional kernels to extract features of different scales.

[0085] As Figure 3 shown, the structure and data processing process of the CSP-PMFEM module include:

[0086] First, the input feature map is split into two parts:

[0087] ;

[0088] Among them, represents the input feature map of the module, and Split represents the splitting operation. represents a part of the feature map after splitting. represents the other part of the feature map after splitting.

[0089] Then, in the passed branches, first pass through a 1×1 convolution, and the output is :

[0090] ;

[0091] Among them, represents the input feature map of the partial multi-scale feature extraction module.

[0092] Input to the PMFEM module. Inside the PMFEM module:

[0093] First, divide into two parts along the channel dimension. One part is given as , whose number of channels is c. First, pass through a 3×3 convolution kernel to obtain the output for extracting the features of small targets:

[0094] ;

[0095] Next, divide into two parts along the channel direction, and apply a 5×5 convolution to half of them to capture the features of medium scale, obtaining the intermediate output :

[0096] ;

[0097] For divide it into blocks again, and send one part of them into a 7×7 convolution to extract the features of large scale, and the output is :

[0098] ;

[0099] Through convolution kernels of different scales, the PMFEM module extracts rich feature information at different scales, laying a foundation for subsequent feature fusion. After feature extraction, the three parts of features are concatenated and integrated into a new feature map for subsequent processing. At this time, the PMFEM module applies a 1×1 convolution to the concatenated feature map to further fuse the features of each scale and compress the number of channels to ensure that the output feature map maintains the computational efficiency of the model, and the output is :

[0100] ;

[0101] Next, in order to preserve the original information in the input features, the module adopts a residual connection method to add the input feature map and the fused feature map to obtain :

[0102] ;

[0103] Since there are multiple PMFEM modules, the final output of the CSP - PMFEM module can be expressed as:

[0104] ;

[0105] where represents the output feature map of the nth partial multi - scale feature extraction module, represents the module output feature map.

[0106] In addition, in the traditional SPPF layer of the YOLOv11 model, multi - scale pooling operations are used to perform multi - level spatial pyramid pooling on the features to capture context information with different receptive fields. However, the standard SPPF layer shows certain limitations in extracting fine features of apple leaf pests and diseases. Due to the fixed size of its convolutional kernel, the SPPF layer is difficult to effectively capture the smaller and more complex edges and texture features in leaf pests and diseases, thus affecting the performance of the model in detecting tiny pests and diseases. In addition, the traditional SPPF layer lacks adaptability to different feature regions and is difficult to focus on important pest and disease features.

[0107] Therefore, in this embodiment, the SPPF module in the backbone of the apple pest and disease detection model is replaced with the LKAP module.

[0108] As Figure 4 shown, in the LKAP module, a series of pooling operations are performed on the input feature map to generate multi - scale feature maps respectively, so as to ensure that the model can capture multi - level spatial information. The specific data processing process includes:

[0109] First, the input feature map passes through a 1×1 convolutional layer to reduce its number of channels to half of the input. This process effectively reduces the computational amount while preserving the necessary feature information:

[0110] ;

[0111] Next, the feature map is sequentially passed into three max - pooling layers. Through layer - by - layer pooling operations, the receptive field is gradually increased to extract the global information of the image from different scales. After pooling respectively, the outputs are , and :

[0112] ;

[0113] These three pooling operations respectively expand the receptive field of features from local to global, thus capturing the key information of the image at different scales. After that, , , and are concatenated in channels to generate a multi-scale feature combination :

[0114] ;

[0115] After completing the multi-scale concatenation, the combined feature map is input into the LKA (Large Separable Kernel Attention) module. The LKA module uses a large convolutional kernel (k = 11) to generate an attention map, and captures extensive context information by expanding the convolutional receptive field in the spatial domain:

[0116] ;

[0117] Finally, the feature map restores the number of channels through a 1×1 convolutional layer to generate the final output feature map :

[0118] ;

[0119] By applying the LKAP module to fruit tree pest and disease detection, multiple challenges faced by the dataset are effectively solved. In terms of the diversity of target sizes, the large-kernel convolutional attention design of LKA significantly improves the model's detection ability for small target pests and diseases. Especially for the tiny disease spots and pests on the leaves, the detection rate is significantly improved. For the problem of complex background interference, the attention mechanism can adaptively suppress the feature responses in non-critical regions, enhancing the model's robustness under complex lighting and overlapping leaf conditions. In addition, aiming at the problem of large differences in pest and disease characteristics, LKAP improves the model's ability to distinguish different types of pests and diseases through multi-scale feature extraction and attention enhancement. Especially for disease types with similar visual features, the improved model shows stronger discriminative ability, effectively reducing the false detection rate. In small object detection, the design of separable convolution retains more spatial detail information while maintaining a large receptive field, significantly improving the detection accuracy of tiny pest and disease targets.

[0120] After extracting multi-scale image features through the backbone, three feature maps of different scales are obtained, namely the first feature map, the second feature map, and the third feature map.

[0121] Next, the first feature map, the second feature map, and the third feature map are fused through the neck, which includes multiple efficient local attention modules. The first feature map passes through the Efficient Local Attention (ELA) module and the two-dimensional convolutional layer (Conv2d) in sequence to obtain the first high-resolution feature map. The second feature map passes through three branches respectively. Among them, the first branch includes a transposed convolutional layer (ConvTranspose2d) and an efficient local attention module connected in sequence, the second branch includes an efficient local attention module and a two-dimensional convolutional layer connected in sequence, and the third branch includes a transposed convolutional layer and an efficient local attention module connected in sequence. The third feature map passes through the efficient local attention module and the two-dimensional convolutional layer in sequence to obtain the third high-resolution feature map.

[0122] The output of the first branch is multiplied element-wise with the first high-resolution feature map and then concatenated to obtain the first fusion feature map. The third high-resolution feature map is input into the transposed convolutional layer of the third branch. The outputs of the second branch and the third branch are multiplied element-wise and concatenated with the output of the transposed convolutional layer of the third branch to obtain the second fusion feature map. The first fusion feature map, the second fusion feature map, and the third high-resolution feature map are respectively input into the detection head through the CSP-PMFEM module.

[0123] The detection head performs object classification and detection box prediction on the input feature map. Using feature maps of different resolutions as inputs, it accurately predicts the object positions for each feature map, thus achieving multi-scale detection.

[0124] This embodiment also includes pre-training the fruit tree pest and disease detection model. The pre-training method includes:

[0125] S1. Obtain the fruit tree pest and disease image dataset. In this embodiment, the apple pest and disease subset in the Turkey_Plant Dataset is used. This dataset contains unconstrained images, including different scenarios such as soil, trees, leaves, and sky. This dataset covers various pest and disease types common in the apple tree planting process, such as apple woolly aphid disease and apple scab disease. The apple pest and disease image dataset contains 1416 high-resolution apple leaf pest and disease images, which are collected from apples at different times, under different lighting conditions, and in different regions.

[0126] S2. Since there may be problems such as insufficient sample size, uneven distribution, and data noise when directly using the original data for model training, which will affect the learning effect and generalization ability of the model. Therefore, in this study, we adopted a method combining manual annotation and data augmentation to expand and optimize the original dataset.

[0127] In terms of data annotation, the LabelImg tool was used to annotate the pest and disease areas in each image one by one to ensure high-precision and high-consistency annotation. The annotation categories cover multiple pest and disease types, and the labels for each category are strictly defined according to the classification criteria provided by the dataset. In addition, the boundaries of the disease spots and the distribution characteristics of the pests are also accurately marked to ensure that the target detection task fully captures the detailed information. In the data augmentation stage, the advanced Albumentations tool was used to augment the data.

[0128] Albumentations has become a widely used data augmentation library with its rich augmentation operators and efficient implementation. The following augmentation strategies were adopted in this study: random rotation and flipping, color transformation, random cropping and scaling, blurring and noise addition, geometric transformation. Through the above augmentation operations, the dataset was expanded from the original 1416 images to 5664 images, containing more diverse scenes and target distributions, and a preprocessed apple pest and disease image dataset was obtained.

[0129] S3. The preprocessed apple pest and disease image dataset was divided into a training set, a test set, and a validation set according to the ratio of 8:1:1.

[0130] S4. The training set data was used as input to train the fruit tree pest and disease detection model, and the loss function was used to adjust the model parameters during the training process; the validation set data was used as input to verify the fruit tree pest and disease detection model to obtain the verification results, and the hyperparameters were adjusted according to the verification results.

[0131] The experimental environment and training parameters of this embodiment are specifically shown in Table 1:

[0132] Table 1: Experimental environment and training parameters

[0133]

[0134] S5. The test set data was used as input to test the fruit tree pest and disease detection model to obtain the test results, and the evaluation indicators were calculated according to the test results. In this embodiment, the evaluation indicators include mean average precision, model parameters, model size, computational efficiency, etc. If the evaluation indicators are higher than the preset values, the training is completed, and a pre-trained fruit tree pest and disease detection model is obtained. If the evaluation indicators are lower than the preset values, repeat S4~S5 until the evaluation indicators are higher than the preset values, and a pre-trained fruit tree pest and disease detection model is obtained.

[0135] Next, the actual effect of the pre-trained fruit tree pest detection model will be verified in the detection of apple tree pests and diseases.

[0136] In this embodiment, four groups of comparative experiments are designed to compare the performance of different backbone structures, different neck structures, different downsampling modules, and different models. In all experiments, the mean average precision is used as the evaluation index, and GFLOPS and Param are combined for auxiliary analysis to ensure the comprehensiveness and scientificity of the results.

[0137] (1) Compare different backbone modules on the apple pest and disease subset of the Turkey_Plant Dataset

[0138] In the C3k1 module comparison experiment, the performance of the C3k2-iRMB, C3k2-RVB-EMA, C3k2-Star-CAA, C3k2-AdditiveBlock, C3k2-IdentityFormer, and CSP-PMFEM modules was tested. The results are shown in Table 2.

[0139] Table 2: Comparison results of different modules with the C3k1 module

[0140]

[0141] This table shows that the model using the CSP-PMFEM module achieves the best balance between accuracy and computational complexity, with mAP@50 reaching 0.708, which is 4.9% higher than the sub-optimal C3k2-IdentityFormer.

[0142] Although the GFLOPS of CSP-PMFEM is 7.6, slightly higher than that of C3k2-IdentityForme and C3k2-RVB-EMA, it is much lower than that of C3k2-Star-CAA. In addition, its number of parameters is 2.62M, only 0.43M more than that of C3k2-IdentityFormer, still at a relatively low level. It is not difficult to see that CSP-PMFEM maintains good computational efficiency while improving the detection accuracy, meeting the application requirements of actual pest and disease detection tasks.

[0143] (2) Compare different neck structures on the apple pest and disease subset of the Turkey_Plant Dataset

[0144] In the neck comparison experiment, three structures of GhostHGNet, Goldyolo, and GDFPN were compared respectively, and the performance of different attention mechanisms CA-HSFPN, CAA-HSFPN, and EHFPN (in this embodiment) was also compared. The results are shown in Table 3.

[0145] Table 3: Comparison Results of Different Neck Structures

[0146]

[0147] As can be seen from Table 3, EHFPN achieves the best balance in both mAP@50 and computational efficiency. The mAP@50 reaches 71.5%, significantly better than other modules. Specifically, it improves by 6.6% compared to CAA-HSFPN and 7.4% compared to CA-HSFPN, and also surpasses GDFPN with higher computational complexity. In terms of GFLOPS and the number of parameters, the GFLOPS of EHFPN is 5.7, only slightly higher than CA-HSFPN, but significantly lower than Goldyol and GDFPN, demonstrating high computational efficiency. At the same time, its number of parameters is 2.51M. While significantly improving the accuracy, it only increases by 0.51M compared to CAA-HSFPN, far lower than Goldyolo, maintaining a good lightweight design.

[0148] In contrast, CA-HSFPN and CAA-HSFPN have lower computational overhead in terms of GFLOPS and the number of parameters, but perform weakly in detection accuracy. Although GDFPN and Goldyolo have certain accuracy advantages, their computational complexity and the number of parameters increase significantly, making it difficult to meet the lightweight requirements. By introducing the efficient local attention mechanism (ELA), EHFPN is significantly optimized in feature extraction and fusion, taking into account both detection accuracy and computational efficiency, showing broad application value.

[0149] (3) Compare different downsampling modules on the apple pest and disease subset of the Turkey_Plant Dataset

[0150] In the comparison of downsampling modules, the performances of AIFIRepBN, FocalModulation, AIFI, and LKAP were tested. The results are shown in Table 4.

[0151] Table 4: Comparison Results of Different Downsampling Modules

[0152]

[0153] LKAP far exceeds other modules with an mAP@50 accuracy of 0.729, improving by 8.8% compared to the sub-optimal AIFI (0.67) and 17.0% compared to FocalModulation (0.623), showing its significant advantages in multi-scale feature extraction.

[0154] In terms of computational complexity, the GFLOPS of LKAP is 7.3, which is basically equivalent to that of AIFI and AIFIRepBN (both are 7.4), slightly higher than that of FocalModulation (7.2), but it achieves a greater lead in terms of accuracy. At the same time, the number of parameters of LKAP is 2.3M, significantly lower than that of AIFI (2.65M) and AIFIRepBN (2.65M), and it also has advantages in lightweight design. By optimizing feature fusion and introducing a large-scale kernel attention mechanism, LKAP achieves the best balance between accuracy and efficiency, demonstrating strong application potential.

[0155] (4) Comparing different models on the apple pest and disease subset of the Turkey_Plant Dataset

[0156] In the model comparison, experiments respectively tested the overall network models of YOLOv5, YOLOv6, YOLOv8n, YOLOv9t, YOLOv10n, SSD, Faster R-CNN, YOLOv11 and YOLO-PEL for comparison. The experimental data is shown in Table 5, and the experimental data indicators are as Figure 5 shown.

[0157] Table 5: Comparison results of different models

[0158]

[0159] Combined with Table 5 and Figure 5 it can be known that YOLO-PEL shows significant advantages in multiple indicators. Its mAP@50 reaches 72.9%, significantly better than other YOLO versions, SSD and RetinaNet. At the same time, while maintaining a high detection accuracy, YOLO-PEL has a low computational cost (7.3 GFLOPS), and the number of parameters is only 2.30M. In addition, from the convergence curve, it can be observed that YOLO-PEL not only has a faster convergence speed but also shows higher stability during the training process. These results prove that YOLO-PEL achieves a good balance among detection accuracy, efficiency and model compactness, and is suitable for real-time and resource-constrained scenarios.

[0160] Secondly, in order to verify the contribution of the proposed module to the model performance, a series of ablation experiments were conducted on YOLOv11, including separately introducing the PMFEM, EHFPN and LKAP modules, as well as combined experiments of different modules. The performance was evaluated based on the mAP@50 index, and at the same time, the number of parameters (Param / M) and the amount of computation (GFLOPS) were recorded. The results are shown in Table 6. The independent introduction and combination of each module have varying degrees of improvement on the model performance, verifying the effectiveness and rationality of the module improvement and providing a theoretical basis for model optimization.

[0161] Table 6: Test Results of Ablation Experiments on Different Models

[0162]

[0163] According to the experimental results, the design of the ablation experiment aims to verify the impact of each module on the performance of YOLOv11. The original YOLOv11 model (without the improved module) only reaches 68.6% on mAP@50 as a comparison benchmark. When the CSP-PMFEM module is introduced, mAP@50 significantly increases to 70.8%, indicating that this module reduces the computational cost and enhances the key feature expression ability by introducing partially shareable features and a more refined feature fusion structure. After further introducing the EHFPN module, mAP@50 reaches 71.5%, reflecting that the structure combining the local efficient attention (ELA) mechanism and the hierarchical scale feature pyramid (HSFPN) not only enhances the detection ability for multi-scale targets but also improves the recognition effect of small targets. And by adding the LKAP module, the model also shows a good improvement on mAP@50-95, indicating its effectiveness in multi-scale feature fusion and enhancing the spatial receptive field.

[0164] The final model combining the three improved modules has the best performance, with mAP@50 reaching 72.9%, and both the computational cost (GFLOPS) and the number of parameters (Param / M) remaining within a reasonable range, showing the overall effectiveness of the design and verifying the rationality and practical value of the model optimization strategy.

[0165] Finally, in deep learning, the receptive field refers to the area of the input image that a neuron in a certain layer of the network can perceive. A larger receptive field enables the model to capture more context information, which is particularly important for complex scenes and multi-object detection. Therefore, the parameter thresh = 0.99 is selected to measure the coverage range of global features. At this threshold, the area ratio of the original model is 0.779, and the side length of the rectangle is 565, showing the locality of its receptive field and the limitation of capturing global information. While in our model, the area ratio of the model increases to 0.876, and the side length of the rectangle increases to 599, indicating that the model has made significant progress in expanding the receptive field and capturing global context information.

[0166] To more intuitively show this change, the following is a visual comparison of the receptive fields of the two models at the same threshold, as Figure 6As shown, the wider the distribution of the dark area indicates the larger the effective receptive field. It can be seen that in the improved model, the coverage range of the receptive field is significantly expanded, the distribution of high-contribution regions is more extensive, and the captured feature information is more abundant. This result verifies the effectiveness of YOLO-PEL in enhancing the receptive field and further proves its advantages in dealing with complex scenes and multi-scale targets.

[0167] In the visual comparison of the receptive field, we can intuitively observe the degree of attention of the model to key regions. In the heatmap, the depth of color reflects the contribution intensity of different regions, and the darker the color, the greater the contribution to the model decision-making, such as Figure 7 As shown, through the heatmap, we can clearly see the advantages of the YOLO-PEL model after the receptive field is expanded. Compared with the original model, the improved model not only captures more high-contribution information in a wider area, but also shows stronger context awareness in complex background and multi-object scenarios. The heatmap intuitively demonstrates how the model processes details and global features in the image. The improved model can identify and focus on more important regions, further proving the improvement of the receptive field expansion on the target detection performance.

[0168] To more intuitively demonstrate the advantages of the proposed improved module in the detection of apple leaf diseases and pests, Figure 8 and Figure 9 visually display the detection results of different targets, including the detection confidence and the accuracy of the localization box for different types of diseases and pests. For example, Figure 8 the (A)-(D) in Figure 9 show the detection results without introducing the improved module, and there are some cases where the confidence is low or the detection box is inaccurate for some targets, while

[0169] Example 3:

[0170] Based on the above embodiments, this embodiment also provides a fruit tree diseases and pests detection device based on the improved YOLOv11, including:

[0171] A fruit tree image acquisition module, configured to acquire fruit tree captured images.

[0172] A diseases and pests detection module, configured to take the fruit tree captured image as input and output a fruit tree processed image based on the fruit tree diseases and pests detection model.

[0173] A diseases and pests judgment module, configured to judge whether there are diseases and pests on the fruit tree and the types of diseases and pests according to the fruit tree processed image.

[0174] Conclusion:

[0175] The fruit tree pest and disease detection model YOLO-PEL improved based on YOLOv11 in the present invention aims to improve the accuracy and efficiency of pest and disease detection, especially in the detection of complex backgrounds and small targets. By integrating PMFEM, EHFPN, and LKAP modules, YOLO-PEL enhances multi-scale feature extraction, small target detection, and receptive field expansion. Experimental results show that the mAP@50 of YOLO-PEL on the Turkey_Plant dataset reaches 72.9%, significantly superior to existing algorithms such as YOLOv11 and YOLOv8n, with an overall accuracy improvement of 4.3%. At the same time, it maintains a high advantage in terms of computational efficiency and the number of parameters, demonstrating its effectiveness in practical applications.

[0176] The innovation of the present invention lies in combining multiple deep learning technologies to propose a high-precision and low-computation-overhead fruit tree pest and disease detection solution suitable for resource-constrained environments. YOLO-PEL not only provides technical support for precision agriculture but also contributes to the development of intelligent agriculture. In the future, the model will be further optimized to improve real-time processing capabilities, especially the detection accuracy under extreme lighting and large-size images. In addition, future research will also expand the application of this model in the monitoring of pest and diseases of other crops, promoting the further development of intelligent agriculture technology.

[0177] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A fruit tree pest detection method based on improved YOLOv11, characterized in that: include: Acquire images of fruit trees; The fruit tree images are taken as input, and the fruit tree processing map is obtained based on the output of the fruit tree pest and disease detection model; Determine whether the fruit trees have pests and diseases and the types of pests and diseases based on the fruit tree treatment diagram; The fruit tree pest and disease detection model is improved based on the YOLOv11 model. The YOLOv11 model includes a trunk, a neck and a detection head connected in sequence. The C3k2 module in the trunk is replaced by a CSP-PMFEM module, the SPPF module is replaced by an LKAP module, and an efficient local attention module is set in the neck to obtain the fruit tree pest and disease detection model; The CSP-PMFEM module is used to segment the feature map and extract features of different scales using multi-scale convolution kernels. The LKAP module is used to capture contextual information by introducing decomposed large convolution kernels. The efficient local attention module is used to perform multi-scale feature fusion on the feature map.

2. The fruit tree pest and disease detection method based on improved YOLOv11 according to claim 1, characterized in that: The input feature map of the CSP-PMFEM module is segmented to obtain a segmentation feature map, and the segmentation feature map is sequentially passed through a convolution layer and a plurality of partial multi-scale feature extraction modules and then spliced ​​with the input feature map to obtain an output feature map of the CSP-PMFEM module; The partial multi-scale feature extraction module includes a first convolution layer, a second convolution layer and a third convolution layer connected in sequence. The input feature map of the partial multi-scale feature extraction module is divided into a first partial feature map and a second partial feature map along the channel direction after the small target feature is extracted by the first convolution layer. The first partial feature map is divided into a third partial feature map and a fourth partial feature map again along the channel direction after the medium-scale target feature is captured by the second convolution layer. The third partial feature map is spliced ​​with the second partial feature map and the fourth partial feature map after the large-scale target feature is extracted by the third convolution layer, and then residually connected with the input feature map through the fourth convolution layer to obtain the output feature map of the partial multi-scale feature extraction module.

3. The fruit tree pest and disease detection method based on improved YOLOv11 according to claim 2, characterized in that: The data processing expressions of the CSP-PMFEM module include the following: ; ; ; ; ; ; ; ; in, Represents the module input feature map, Split represents the segmentation operation, Represents a part of the feature map after segmentation, Represents another part of the feature map after segmentation, Represents the input feature map of some multi-scale feature extraction modules, represents the first branch input, represents the first branch output, represents the second branch input, represents the second branch output, represents the third branch input, Indicates the third branch output, It indicates that the concatenated feature map is output after being processed by the fourth convolutional layer. Represents the output feature map of some multi-scale feature extraction modules, represents the output feature map of the nth partial multi-scale feature extraction module, Represents the module output feature map.

4. The fruit tree pest detection method based on improved YOLOv11 according to claim 1, characterized in that: The LKAP module includes a first convolutional layer, which is used to reduce the number of channels of the input feature map to half, and its output passes through multiple maximum pooling layers in sequence, and the maximum pooling layer is used to extract global image information from different scales. The output of each maximum pooling layer is channel-joined with the output of the convolutional layer to obtain a multi-scale feature combination, and the multi-scale feature combination is sequentially passed through a large separation convolutional attention module and a second convolutional layer to obtain an output feature map. The large separation convolutional attention module is used to expand the convolution receptive field in the spatial domain to obtain the output feature map, and the second convolutional layer is used to restore the number of feature map channels.

5. The fruit tree pest detection method based on improved YOLOv11 according to claim 4, characterized in that: The number of the maximum pooling layers is 3, and the data processing expression of the LKAP module includes the following: ; ; ; ; ; in, represents the module input feature map, represents the output of the first convolutional layer, represents the output of the first maximum pooling layer, represents the output of the second maximum pooling layer, represents the output of the third maximum pooling layer, represents a combination of multi-scale features, represents the output of a large disjoint convolutional attention module, Represents the module output feature map.

6. The fruit tree pest detection method based on improved YOLOv11 according to claim 1, characterized in that: The backbone is used to extract multi-scale image features to obtain three feature maps of different scales, which are respectively a first feature map, a second feature map and a third feature map; The neck is used to perform feature fusion on the first feature map, the second feature map and the third feature map, and includes a plurality of efficient local attention modules, wherein the first feature map is sequentially passed through the efficient local attention module and the two-dimensional convolution layer to obtain a first high-resolution feature map; The second feature map passes through three branches respectively, wherein the first branch includes a transposed convolution layer and an efficient local attention module connected in sequence, the second branch includes an efficient local attention module and a two-dimensional convolution layer connected in sequence, and the third branch includes a transposed convolution layer and an efficient local attention module connected in sequence; The third feature map is sequentially passed through an efficient local attention module and a two-dimensional convolutional layer to obtain a third high-resolution feature map; The output of the first branch is multiplied element by element with the first high-resolution feature map, and then concatenated to obtain a first fused feature map; The third high-resolution feature map is input into the transposed convolution layer of the third branch, the outputs of the second branch and the third branch are element-wise multiplied, and concatenated with the output of the transposed convolution layer of the third branch to obtain a second fused feature map; The first fused feature map, the second fused feature map and the third high-resolution feature map are respectively input into the detection head via the CSP-PMFEM module.

7. The fruit tree pest detection method based on improved YOLOv11 according to claim 1, characterized in that: The method also includes pre-training the fruit tree pest and disease detection model, wherein the pre-training method includes: S1. Obtain fruit tree pest and disease image dataset; S2, annotating and data enhancement of the fruit tree disease and insect pest image dataset to obtain a preprocessed fruit tree disease and insect pest image dataset; S3, dividing the preprocessed fruit tree pest and disease image dataset into a training set, a test set and a validation set; S4, using the training set data as input to train the fruit tree pest and disease detection model, and using the loss function to adjust the model parameters during the training process; using the validation set data as input to validate the fruit tree pest and disease detection model to obtain validation results, and adjusting the hyperparameters according to the validation results; S5. Use the test set data as input to test the fruit tree pest and disease detection model to obtain test results, and calculate the evaluation index based on the test results. If the evaluation index is higher than the preset value, the training is completed and the pre-trained fruit tree pest and disease detection model is obtained. If the evaluation index is lower than the preset value, repeat S4~S5 until the evaluation index is higher than the preset value and the pre-trained fruit tree pest and disease detection model is obtained.

8. The fruit tree pest and disease detection method based on improved YOLOv11 according to claim 7, characterized in that: The annotation information includes the pest area and the pest type, the data enhancement includes random rotation and flipping, color transformation, random cropping and scaling, blurring, noise addition and geometric transformation, and the evaluation indicators include average precision mean, model parameters, model size and computational efficiency.

9. The fruit tree pest detection method based on improved YOLOv11 according to claim 1, characterized in that: The method of judging whether the fruit trees have pests and diseases and the types of pests and diseases according to the fruit tree treatment diagram includes: If there are pests and diseases on the fruit trees, the fruit tree treatment map includes a detection frame, which is the pest and disease area, and the pest and disease type and confidence level are marked on the detection frame; If the fruit tree is free from pests and diseases, the fruit tree treatment image is the photographed image of the fruit tree.

10. A fruit tree pest detection device based on improved YOLOv11, characterized in that: include: A fruit tree image acquisition module is configured to acquire photographed images of fruit trees; The pest and disease detection module is configured to take the photographed image of the fruit tree as input and obtain a fruit tree processing map based on the output of the fruit tree pest and disease detection model; The pest and disease judgment module is configured to judge whether the fruit tree has pests and diseases and the types of pests and diseases according to the fruit tree treatment map.

Citation Information

Cited By

  • Target detection model and method based on deep learning

    CN121095726A

  • Trachinotus ovatus image processing method and system based on machine vision

    CN121236799A

  • Pest target detection method based on improved YOLOv11

    CN121259591A

  • Camellia oleifera tree disease detection method and system based on unmanned aerial vehicle and YOLO algorithm

    CN121708485A