Crop pest detection method based on improved YOLOv11s model

By introducing the AMdown dual-branch downsampling module, dynamic texture attention module and spatial reconstruction multi-scale expansion attention module in the YOLOv11s model, the problem of poor feature extraction ability in pest detection is solved, and higher detection accuracy and robustness are achieved.

CN120014434APending Publication Date: 2025-05-16DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411963631.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Pest detection based on deep learning models has the problem of poor feature extraction capabilities, especially when the appearance of different types of pests is small and the appearance of large differences in growth stages.

Method used

The improved YOLOv11s model is adopted, and the multi-scale expansion attention module is added by adding the AMdown dual-branch downsampling module, dynamic texture attention module and spatial reconstruction multi-scale expansion attention module, which enhances the model's feature extraction and attention mechanism, and improves the detection performance of small objects and multi-scale feature modeling capabilities.

Benefits of technology

Without increasing the number of parameters and FLOPs, the accuracy and robustness of the model in the pest detection task is significantly improved, meeting the needs of agriculture for pest detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014434A_ABST
    Figure CN120014434A_ABST
Patent Text Reader

Abstract

The invention provides a crop pest detection method based on an improved YOLOv11s model, and the method comprises the steps: employing YOLOv11s as an original model, providing an AMdown double-branch down-sampling module, a dynamic texture attention module and a spatial reconstruction multi-scale expansion attention module in the original YOLOv11s model, forming the improved YOLOv11s model, and employing the improved YOLOv11s model as a pest detection network model; inputting the preprocessed pest data set into an improved YOLOv11s model, and carrying out model training; and detecting an actual pest image by using the trained improved YOLOv11s model. According to the method, the YOLOv11s model is improved, an AMdown double-branch down-sampling module, a dynamic texture attention module and a spatial reconstruction multi-scale expansion attention module are provided, the generalization ability and robustness of the model are improved, and the detection precision is improved while the parameter quantity and FLOPs are not increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection and relates to a pest detection method, in particular to a pest detection method for crops based on an improved YOLOv11s model. Background Art

[0002] The world's population is huge and continues to rise, and people's demand for food crops is increasing. Pest detection related to food crops has always been an important concern in the field of agricultural planting. Adopting effective pest control methods, increasing crop yields, and reducing agricultural economic losses have always been the crop pest control effects generally pursued by the industry.

[0003] In agriculture, the one-stage algorithm represented by the YOLO series of algorithms is a relatively mature technology. Kangshun et al. proposed a passion fruit pest detection algorithm based on the improved YOLOv5 model, using a new point-line distance loss function (PLDIoU Loss) to reduce redundant calculations and shorten detection time, while adding an attention module to improve detection and recognition rates. In addition, they introduced the Mixup online data enhancement algorithm to increase the diversity of the training set, improve the robustness of the model and prevent overfitting. Yang Di et al. proposed a lightweight attention-based TP-YOLO network for micro-pest detection. The network includes a context transformer (CoT) and an omnidirectional dynamic convolution (ODConv) module to enhance feature extraction capabilities, and demonstrated an average precision (AP) of 42.0% and an average precision of 50% (AP50) of 66.8% on the Pest24 dataset. Yunong Tian et al. proposed a multi-scale dense YOLO (MD-YOLO) model, which uses DenseNet blocks to reduce feature information loss, combines adaptive attention modules (AAM) and multi-scale feature fusion networks, and effectively integrates features of different scales. Saman M. Omer et al. proposed a model based on improved YOLOv5 for cucumber leaf disease and pest detection. By replacing the C3 module with the Bottleneck CSP module in the backbone and neck network of YOLOv5l and introducing the convolutional block attention module (CBAM), the average precision (mAP) reached 80.10%, the precision and recall were 73.8% and 73.9% respectively, and the model memory occupancy was only 13.6MB.

[0004] Although the above studies have achieved good results, pest detection based on deep learning models still has the following problems: Since the appearance of different types of pests is very different, while the appearance of the same type of pests at different growth stages is very different, and they need to rely on the texture on the pests to distinguish them, the model has poor feature extraction capabilities for pests, which poses a challenge to the YOLOv11 model in extracting pest features. Summary of the invention

[0005] In order to overcome the shortcomings of the prior art, the present invention provides a crop pest detection method based on an improved YOLOv11s model, which improves the detection accuracy without increasing the number of parameters and FLOPs, and meets the requirements of actual agriculture for pest detection.

[0006] The technical solution adopted by the present invention to solve the technical problem is:

[0007] A crop pest detection method based on an improved YOLOv11s model, the steps of which include:

[0008] Build a pest detection network model: Use YOLOv11s as the original model, add the AMdown dual-branch downsampling module, dynamic texture attention module and spatial reconstruction multi-scale expansion attention module to the original YOLOv11s model to form an improved YOLOv11s model, and use the improved YOLOv11s model as the pest detection network model;

[0009] Preprocessing of images from pest dataset;

[0010] The preprocessed pest dataset is input into the improved YOLO11s model for model training;

[0011] Use the trained improved YOLOv11s model to detect actual pest images.

[0012] Based on the above scheme, YOLOv11s is used as the original model, and the AMdown dual-branch downsampling module is added to better reduce information loss and capture global and local information; a dynamic texture attention module is added to dynamically amplify shallow texture information and filter out redundant background noise to improve the small object detection performance; a spatial reconstruction multi-scale expansion attention module is added to make up for the limitations of the dynamic texture attention module in locality and sparsity modeling, and effectively model the local and global features in pest detection, thereby significantly improving the accuracy of the model in pest detection tasks; by enhancing and preprocessing image data, the robustness and generalization ability of the model can be improved, so that the model can better handle pest images in various practical scenarios.

[0013] Furthermore, the improved YOLOv11s model is specifically improved by:

[0014] Configure the environment and input the pest dataset into the original YOLOv11s model;

[0015] Configure hyperparameters: set the optimal hyperparameter combination of learning rate, batch size, number of iterations, number of image channels, image cropping size, and learning rate momentum;

[0016] The following improvements are made to the original YOLOv11s model: a. The AMdown dual-branch downsampling module is used to replace the CBS module in the original YOLOv11s model network; b. A dynamic texture attention module is added to the neck network of the original YOLOv11s model; c. A spatial reconstruction multi-scale dilation attention module is added to the neck network of the original YOLOv11s model.

[0017] Based on the above scheme, the AMdown dual-branch downsampling module replaces the CBS downsampling module in the original YOLOv11s model, which helps the model capture global information and local details and reduce information loss; the dynamic texture attention module is added to the neck network of the original YOLOv11s model, which helps the model fuse multi-scale context information, dynamically amplify shallow texture information, and filter out background noise to improve the performance of small object detection; the spatial reconstruction multi-scale expansion attention module is added to the neck network of the original YOLOv11s model, which helps to make up for the limitations of the dynamic texture attention module in locality and sparsity modeling, effectively modeling local and global features in pest detection, thereby significantly improving the accuracy of the model in pest detection tasks. These improved methods jointly improve the performance of the improved YOLOv11s model in pest detection tasks, including higher accuracy and better robustness, by optimizing the model's feature extraction, enhancing the texture perception ability, and strengthening the attention mechanism.

[0018] Furthermore, the AMdown dual-branch downsampling module processes the input features as follows: the input features are passed through two parallel branches respectively, the first branch performs average pooling, 3×3 convolution, batch normalization, and SiLU activation function processing in sequence, and the second branch performs maximum pooling, 1×1 convolution, batch normalization, and SiLU activation function processing in sequence.

[0019] Based on the above scheme, the input features are passed through two branches respectively. The first branch samples average pooling to extract global information, which not only helps to extract local features, but also maintains certain global feature information, which helps to enhance the model's ability to recognize objects of different scales. The second branch uses maximum pooling to capture salient features in the image. This setting helps to highlight local salient features, thereby enhancing the robustness of the model in complex scenes.

[0020] Furthermore, the dynamic texture attention module has two input branches, namely the dynamic texture branch and the background noise branch; the dynamic texture branch is a three-branch structure, in which each branch obtains different receptive field features through maximum pooling and 3×3 dilated convolutions with different expansion rates for the input features, and then obtains the attention coefficient through channel attention, and fuses the receptive field features obtained by maximum pooling and 3×3 dilated convolutions with different expansion rates; the background noise branch performs upsampling, average pooling, 1×1 convolution, batch normalization, and Sigmod activation function processing in sequence, and finally the dynamic texture branch minus the background noise branch obtains the output features of the dynamic texture attention module.

[0021] Based on the above scheme, each branch in the dynamic texture branch expands the receptive field through maximum pooling and dilated convolutions of various sizes, and then combines the different receptive field features obtained from these two methods along the channel dimension, and the channel attention considers the importance of the receptive field. The background noise branch is used to suppress distracting low-resolution features in the information-rich high-resolution features.

[0022] Furthermore, the spatial reconstruction multi-scale dilated attention module first obtains the query, key and value through linear projection, and then divides the channel of the feature map into multiple heads, applies different dilation rates on different heads to perform the following operations: spatial reconstruction is used to optimize the input of the features before the multi-scale dilated attention mechanism, and then the multi-scale dilated attention module models long-distance dependencies through a sliding window dilated attention operation.

[0023] Based on the above scheme, the spatial reconstruction multi-scale dilated attention module suppresses redundant information by introducing spatial reconstruction before the multi-scale dilated attention, thereby enhancing the representation ability of features and providing optimized input for the subsequent multi-scale dilated attention. Then, the multi-scale dilated attention modeled long-distance dependencies through the sliding window dilated attention operation, so that the attention mechanism can not only focus on local features, but also take into account long-distance semantic interactions, effectively improving the multi-scale capability of feature extraction, and enabling the dynamic texture attention module to more flexibly handle complex image scenes based on multi-scale shallow features.

[0024] Furthermore, the model training specifically includes the following steps: automatically generating a priori boxes using the K-mean clustering method, obtaining the boundary size through bounding box regression prediction, classifying the bounding boxes using a classifier, obtaining the defect type probability corresponding to each bounding box, and then sorting the classification probability of each bounding box through the non-maximum suppression method to obtain the bounding box prediction value with the highest confidence, the confidence threshold is set to 0.25, the IOU threshold is set to 0.7, and then the loss value between the predicted value and the true value is calculated through the loss function, and back propagation is performed according to the loss value until the preset number of iterations is reached, and the network model training is completed.

[0025] Based on the above scheme, using the K-means clustering method to automatically generate a priori boxes can help the model better adapt to the sizes and shapes of different targets, improve the generalization ability of the model, and enable it to detect a variety of targets; through bounding box regression, the model can more accurately predict the bounding box size of the target, which helps to improve the accuracy of detection, especially for targets of different sizes; the classifier is used to classify the bounding boxes, thereby determining the probability of defect types corresponding to each bounding box, allowing the model to not only detect the existence of the target but also classify the type of the target; the classification probability of each bounding box is sorted using the non-maximum suppression method, and highly overlapping bounding boxes are removed, thereby reducing repeated detection and improving the quality of the detection results; setting the confidence threshold and IOU threshold helps control the output results of the model, ensuring that only bounding boxes with sufficiently high confidence are retained, improving the robustness and accuracy of the model; the difference between the predicted value and the true value is calculated through the loss function, and the model can be back-propagated to update the model parameters to reduce the loss;

[0026] Furthermore, the image data is preprocessed, specifically: when the width or height is proportionally scaled to 640, the remaining part is filled with background grayscale.

[0027] The beneficial effects of the present invention include:

[0028] Improvements to the YOLOv11s model for pest detection include the AMdown dual-branch downsampling module, which enables the model to better balance the relationship between global information and local salient features when downsampling, ensuring that important image details are not lost while reducing the amount of computation. The proposed dynamic texture attention module integrates multi-scale contextual information and dynamically amplifies texture information by reducing background noise to emphasize small objects in shallower layers, thereby improving the performance of detecting small objects. Spatial reconstruction multi-scale dilated attention helps to make up for the limitations of the dynamic texture attention module in locality and sparsity modeling, effectively modeling local and global features in pest detection, thereby significantly improving the accuracy of the model in pest detection tasks. The accuracy of pest detection is improved without increasing the number of parameters and FLOPs, meeting the requirements of agriculture for pest detection, and laying a technical foundation for the ultimate establishment of a pest detection system. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flow chart of the crop pest detection method based on the improved YOLOv11s model of the present invention;

[0030] Figure 2 Schematic diagram of the AMdown dual-branch downsampling module proposed in the present invention;

[0031] Figure 3is a schematic diagram of the dynamic texture attention module proposed in the present invention;

[0032] Figure 4 Schematic diagram of the spatial reconstruction multi-scale dilation attention module proposed in the present invention;

[0033] Figure 5 is a schematic diagram of a space reconstruction module used in the present invention;

[0034] Figure 6 It is a schematic diagram of the improved algorithm structure of the present invention;

[0035] Figure 7 It is a comparison chart of the test results of the improved model of the present invention and the original model. DETAILED DESCRIPTION

[0036] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0038] Example 1

[0039] This embodiment proposes a crop pest detection method based on the YOLOv11s model, which integrates artificial intelligence technology and pest detection technology, aims to strengthen the adaptive feature fusion of the neural network during feature extraction and improve the detection effect. Without increasing the number of parameters and FLOPs, the model's detection accuracy for pests is improved to meet the needs of agriculture for pest detection.

[0040] The specific operations of the crop pest detection method based on the YOLOv11s model in this embodiment are as follows:

[0041] 1. Build a pest detection network model based on the improved YOLOv11s model:

[0042] (1) Environment configuration: Download the official source code of the YOLOv11s model from GitHub, then use Anaconda to create a virtual YOLOv11 environment. According to the requirements file of the YOLOv11s source code, install matplotlib>=3.2.2, opencv-python>=4.6.0, Pillow>=7.1.2, PyYAML>=5.3.1, requests>=2.23.0, scipy>=1.4.1, tor ch>=1.7.0, torchvision>=0.8.1, tqdm>=4.64.0, tensorboard>=2.13.0, dvclive>=2.11.0, clearml, come t, pandas>=1.1.4, seaborn>=0.11.0, coremltools>=6.0, onnx>=1.12.0, onnxsim>=0.4.1, nvidia-pyindex nvidia-tensorrt, scikit-learn==0.19.2, tensorflow>=2.4.1, tflite-support, tensorflowjs>=3.9.0openvino-dev>=2022.3, psutil, thop>=0.1.1, ipython, algorithms>=1.0.3, pycocotools>=2.0.6, roboflow;

[0043] (2) Data input: Input the pest dataset into the YOLOv11s network structure model, including obtaining pest images, pest image labels, YOLOv11s network structure model pre-trained weights, and YOLOv11s network structure model configuration file;

[0044] (3) Configure hyperparameters: Set the optimal hyperparameter combination of learning rate, batch size, iteration number, image channels, image crop size, and learning rate momentum, specifically: learning rate lr = 0.01, batch size Batchsize = 16, iteration number Epoch = 100, image channels Channels = 3, image crop size Cropsize = 640 × 640, learning rate momentum = 0.937;

[0045] (4) Model improvement: The original model used in the present invention is YOLOv11s, and the YOLOv11s network structure model is improved in the following three aspects:

[0046] 1) Design an AMdown dual-branch downsampling module (hereinafter referred to as the AMdown module) to replace the CBS module in the original model YOLOv11s to enhance the feature extraction capability and multi-scale information fusion. The AMdown dual-branch downsampling module optimizes the downsampling process through a dual-branch structure while retaining key feature information, thereby improving the detection accuracy of the model. The structure of the proposed AMdown dual-branch downsampling module is shown in Figure 1. Figure 2 shown.

[0047] The AMdown module contains two parallel branches, which process the input feature map through average pooling (AvgPool) and maximum pooling (MaxPool) operations respectively. The first branch extracts global information through average pooling, and then applies a 3×3 convolution layer (Conv). The convolved feature map is then nonlinearly transformed through batch normalization (BN) and SiLU activation function. This design not only helps to extract local features, but also maintains certain global feature information, which helps to enhance the model's ability to recognize objects of different scales.

[0048] The second branch uses maximum pooling to capture the salient features in the image, further compresses the number of channels of the feature map through a 1×1 convolution layer, and then uses batch normalization and SiLU activation function for processing. This setting helps to highlight local salient features, thereby enhancing the robustness of the model in complex scenes.

[0049] After the feature extraction of the two branches, the AMdown module concatenates all the generated feature maps in the channel dimension (Concatenation, CAT). This multi-scale feature fusion strategy can capture rich contextual information in different receptive fields, making the model perform better when performing multi-scale detection of targets. In addition, by integrating the advantages of average pooling and maximum pooling, the AMdown module can better balance the relationship between global information and local salient features, ensuring that important image details are not lost while reducing the amount of calculation.

[0050] 2) A dynamic texture attention module (DTA) is designed to be added to the neck of YOLOv11s, which integrates multi-scale contextual information and dynamically amplifies texture information to emphasize small objects in lower layers by reducing redundant semantics, thereby improving the performance of detecting small objects. The structure of the proposed DTA is shown in Figure 3 shown.

[0051] Due to excessive background noise, lower-level features have poor discrimination in detecting small objects. Although the standard dilated convolution can obtain a larger receptive field size without increasing the kernel parameters, it is still limited by the fixed receptive field size of the output of the dilated convolution. To solve this problem and better represent lower features, this embodiment adds a small amount of receptive field size to the shallow features P. i A multi-branch structure that fuses multi-scale contextual information is adopted. Each branch expands the receptive field through maximum pooling and dilated convolutions of various sizes, and then combines the different receptive field features obtained from these two methods along the channel dimension. The former extracts features of different levels by using different pooling sizes, and the latter fuses multi-scale information by using dilated convolutions with different expansion rates.

[0052] Specify the dilation rate and pooling size, and set the convolution kernel size to 3×3. Where H, W, and C represent the height, width, and number of channels of the feature map, respectively. The three parallel dilated convolution and maximum pooling branches are represented as follows:

[0053]

[0054] where x i and i They represent the three input features after dilated convolution and maximum pooling. Then, this embodiment adds the output results of different branches one by one to obtain the specific representation of each branch feature map, and the formula is defined as follows:

[0055]

[0056] The input of this module includes two feature layers: shallow feature P i and high-level features P i+2 In this embodiment, DTA is designed to integrate P i and P i+2 In this embodiment, the shallow feature P i To this end, this embodiment removes the large object features that weaken the small object features from the P containing small objects. i Subtract P i+2 Get the output shallow feature map P i dt In this embodiment, the filtered feature map P can be generated as follows: i dt :

[0057]

[0058] Where P d represents shallow features enhanced by dynamically utilizing maximum pooling and dilated convolutions with various dilation rates to more accurately represent small objects, Ps represents the attention map used to suppress distracting low-resolution features from informative high-resolution features. σ(·), ω(·) and They represent the sigmoid function, 1×1 convolutional learnable weights, and upsampling layer respectively.

[0059] Specifically, the shallow feature P d is the cascade feature P that is aggregated by using the attention score α cat The aggregation uses the output F of maximum pooling and dilated convolution with different dilation rates (r = 1, 2, 3) added one by one i The attention weights are applied to P via the Hadamard product cat After that, it is split and combined by element-wise summation. d Formulated as:

[0060] P d =ψ(P cat (F i )⊙α) (5)

[0061] where ψ(·) represents the operation of sequential segmentation and element-wise summation. The α attention score represents the importance of the channel attention mechanism considering the receptive field, as shown below:

[0062] α=σ(W2δ(W1(P avg (P cat (F i ))))) (6)

[0063] in, are the learnable weights of the first 1×1 convolutional layer, are the learnable weights of the second 1×1 convolutional layer. avg (·), δ(·), σ(·) represent the operations of average pooling, ReLU function and sigmoid function. The reduction rate is set to r = 16.

[0064] In order to obtain the semantic attention map P s In this embodiment, the characteristic graph P i+2 Upsampling was performed to achieve the same i Then, the upsampled feature map is passed through an average pooling layer, a 1×1 convolution layer, and a sigmoid function to generate P s , as shown below:

[0065]

[0066] in, are the learnable weights of the 1×1 convolutional layer, P avg(·) and σ(·) represent upsampling, average pooling and sigmoid function operations respectively.

[0067] 3) A spatial reconstruction multi-scale dilation attention module is designed to further improve the multi-scale modeling capability of the model in shallow feature extraction. The structure of the spatial reconstruction multi-scale dilation attention module is as follows: Figure 3 shown.

[0068] The spatial reconstruction multi-scale dilation attention module first performs spatial reconstruction through the spatial reconstruction module to optimize the representation of the input feature map. The structure is as follows Figure 4 Specifically, given the feature map Where N is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively. The standardized feature map X is calculated by Group Normalization (GN) out :

[0069]

[0070] Among them, μ and σ are the mean and standard deviation of the input feature X, ε is a small constant added for stability, and γ and β are trainable affine transformation parameters. By measuring γ for each channel, the weight W can be obtained γ :

[0071]

[0072] Next, the weights are processed using the sigmoid activation function and the gating mechanism to obtain the informative weight W1 and the redundant weight W2. Finally, the reconstructed features are obtained through the following operations:

[0073] W=Gate(Sigmoid(W γ GN(X))) (10)

[0074] The reconstruction operation separates and fuses features by weighting and adding the informative and redundant parts. Specifically, for the input feature X, it is element-wise multiplied with W1 and W2 to obtain the weighted feature and Then perform cross reconstruction operation to finally obtain the spatially reconstructed feature X w :

[0075]

[0076] At this point, the spatial reconstruction operation has been completed, and the representation ability of the features has been enhanced by suppressing redundant information, providing optimized input for the subsequent multi-scale dilated attention (SWDA). Then, the locality and sparsity characteristics of Vision Transformers (ViTs) in shallow global attention are used to model long-distance dependencies through the sliding window dilated attention (SWDA) operation. Specifically, in the SWDA operation, the keys and values ​​are sparsely selected within a sliding window centered on the query image block, so that the attention mechanism can focus on local features while also taking into account long-distance semantic interactions.

[0077] Given a feature map X W The multi-scale dilated attention module first obtains the query, key, and value through linear projection, and then divides the channels of the feature map into multiple heads, and performs SWDA on different heads with different dilation rates to achieve multi-scale feature aggregation. The query features of each head are sparsely selected in a sliding window manner with its center position as the sliding window to control the receptive field of the attention operation. Dilation rate The setting of determines the degree of sparsity, so that the attention mechanism can flexibly model long-distance dependencies in local windows. For each query position (i, j), the output of SWDA can be expressed by the following formula:

[0078]

[0079] Where H and W are the height and width of the feature map, respectively, and K r and V r Represents the keys and values ​​from the feature maps K and V. k represents the scaling factor; q ij The query vector representing the position (i, j) on the feature map.

[0080] In the spatial reconstruction multi-scale dilation attention module, each head is configured with a different dilation rate to achieve the integration of multi-scale features. The feature channel is divided into n heads, and each head performs SWDA operation in turn to extract multi-scale information. The formula is as follows:

[0081] h i =SWDA(Q i ,K i ,V i ,r i ),1≤i≤n (13)

[0082] X=Linear(Concat[h1,...,h n ]) (14)

[0083] where r i is the expansion rate of the ith head, Qi , K i 、V i Respectively represent the feature slices input to the i-th head. After the outputs of multiple heads are spliced, the features are further aggregated through a linear layer. By setting multiple expansion rates on different heads, the spatial reconstruction multi-scale expansion attention module can efficiently integrate multi-scale semantic information, which not only retains the advantages of SWDA in locality and sparsity, but also significantly reduces the redundancy of the shallow attention mechanism. With a slight increase in computational overhead, the spatial reconstruction multi-scale expansion attention effectively improves the multi-scale capability of feature extraction, enabling the dynamic texture attention module to more flexibly handle complex image scenes based on multi-scale shallow features.

[0084] 2. Image preprocessing:

[0085] The specific steps of preprocessing include:

[0086] When the width / height is scaled proportionally to 640, the remaining part is filled with background grayscale.

[0087] 3. Training using the improved YOLOv11s model

[0088] (1) The image data enhancement and preprocessed pest data set are input into the improved YOLOv11s model with set hyperparameters for training. The configuration used in the present invention is CPU: Intel Core i9-11900K; CPU main frequency: 3.50GHz; memory: 32G; GPU: NVIDIA GeForce RTX 3080; video memory: 8G; deep learning framework is PyTorch, and the development environment is Pytoch 2.0.1, Python 3.11, Cuda 11.8;

[0089] (2) Model training: The training set of the pest data set is input into the improved YOLOv11s model for training. The K-mean clustering method is used to automatically generate a priori boxes. The boundary size is obtained through bounding box regression prediction. The bounding boxes are classified by the classifier to obtain the defect type probability corresponding to each bounding box. The classification probability of each bounding box is then sorted by the non-maximum suppression (NMS) method to obtain the bounding box prediction value with the highest confidence. The confidence threshold is set to 0.25 and the IOU threshold is set to 0.7. The loss value between the predicted value and the true value is then calculated using the loss function. Backpropagation is performed based on the loss value until the preset number of iterations is reached and the network model training is completed.

[0090] 4. Test and evaluate the improved YOLOv11s model:

[0091] (1) Model testing: The performance changes caused by changes in the network structure were verified through ablation experiments. Among them, we call the AMdown dual-branch downsampling module that replaces the CBS module of the original model AMdown, the dynamic texture attention module added to the neck network DTA, and the spatial reconstruction multi-scale expansion attention module SEMSDA. A total of four experiments were trained: YOLOv11s, YOLOv11s-AMdown, YOLOv11s-AMdown-DTA, and YOLOv11s-AMdown-DTA-SRMSDA (the present invention). The experimental results are as follows: Figure 7 And as shown in Table 1.

[0092] Table 1 Evaluation table of ablation experiment of the present invention

[0093]

[0094] As can be seen from Table 1, after replacing the downsampling module of YOLOv11s with the AMdown module, all indicators of the model have been improved. After adding DTA and SRMSDA to the neck network, all indicators of the model have been improved, and mAP@0.5 has increased by 3.7%, indicating that SRMSDA has an auxiliary effect on DTA. Figure 7 The performance comparison between the proposed method and YOLOv11s in different scenarios is shown. The image on the left is the detection result of YOLOv11s, and the image on the right is the detection result of the proposed method. In the scenario where the background and pests are clearly distinguished, there is almost no difference between the proposed method and YOLOv11s. In the detection of larvae and adults, the proposed method is superior to the original model YOLOv11s, such as Figure 7 (b) As shown. In scenes with complex backgrounds, the present invention can effectively detect pests, such as Figure 7 (c) is shown. Figure 7 In (d), we can see that the original model misdetects pests. The results show that this method is better than YOLOv11s.

[0095] 5. Use the improved YOLOv11s model to detect pest images:

[0096] The pest image is input into the improved YOLOv11s model to complete the detection of the pest image.

[0097] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. A crop pest detection method based on an improved YOLOv11s model, its characteristic steps include: Build a pest detection network model: Use YOLOv11s as the original model. In the original YOLOv11s model, modify the original downsampling module to the AMdown dual-branch downsampling module, add a dynamic texture attention module and a spatial reconstruction multi-scale expansion attention module to form an improved YOLOv11s model, and use the improved YOLOv11s model as the pest detection network model; Preprocessing of images from pest dataset; The preprocessed pest dataset is input into the improved YOLOv11s model for model training; Use the trained improved YOLOv11s model to detect actual pest images.

2. The crop pest detection method based on the improved YOLOv11s model according to claim 1, characterized in that: The improved YOLOv11s model is specifically improved as follows: Configure the environment and input the constructed pest dataset into the original YOLOv11s model; Configure hyperparameters: set the optimal hyperparameter combination of learning rate, batch size, number of iterations, number of image channels, image cropping size, and learning rate momentum; The following improvements are made to the original YOLOv11s model: a. The AMdown dual-branch downsampling module is used to replace the CBS module in the original YOLOv11s model network; b. A dynamic texture attention module is added to the neck network of the original YOLOv11s model; c. A spatial reconstruction multi-scale dilation attention module is added to the neck network of the original YOLOv11s model.

3. The crop pest detection method based on the improved YOLOv11s model according to claim 2, characterized in that: The AMdown dual-branch downsampling module processes input features as follows: the input features are passed through two parallel branches respectively, the first branch performs average pooling, 3×3 convolution, batch normalization, and SiLU activation function processing in sequence, and the second branch performs maximum pooling, 1×1 convolution, batch normalization, and SiLU activation function processing in sequence.

4. The crop pest detection method based on the improved YOLOv11s model according to claim 2, characterized in that: The dynamic texture attention module has two input branches, namely a dynamic texture branch and a background noise branch; the dynamic texture branch is a three-branch structure, in which each branch obtains different receptive field features through maximum pooling and 3×3 dilated convolutions with different dilation rates, and then obtains the attention coefficient through channel attention, and fuses the receptive field features obtained by maximum pooling and 3×3 dilated convolutions with different dilation rates; The background noise branch is processed by upsampling, average pooling, 1×1 convolution, batch normalization, and Sigmod activation function in sequence. Finally, the dynamic texture branch is subtracted from the background noise branch to obtain the output features of the dynamic texture attention module.

5. The crop pest detection method based on the improved YOLOv11s model according to claim 2, characterized in that: The spatial reconstruction multi-scale dilated attention module first obtains the query, key and value through linear projection, and then divides the channels of the feature map into multiple heads, and applies different dilation rates on different heads to perform the following operations: spatial reconstruction is used to optimize the input of the features before the multi-scale dilated attention mechanism, and then the multi-scale dilated attention module models long-distance dependencies through sliding window dilated attention operations.

6. The crop pest detection method based on the improved YOLOv11s model according to claim 1 or 2, characterized in that: The model training specifically includes the following steps: using the K-mean clustering method to automatically generate a priori boxes, obtaining the boundary size through bounding box regression prediction, using a classifier to classify the bounding boxes to obtain the defect type probability corresponding to each bounding box, and then sorting the classification probability of each bounding box through a non-maximum suppression method to obtain the bounding box prediction value with the highest confidence, and then calculating the loss value between the predicted value and the true value through a loss function, and performing back propagation according to the loss value until a preset number of iterations is reached, and the network model training is completed.

7. The crop pest detection method based on the improved YOLOv11s model according to claim 6, characterized in that: The image data preprocessing is specifically as follows: when the width or height is scaled proportionally to 640, the remaining part is filled with background grayscale.

Citation Information

Cited By

  • Traditional Chinese medicine decoction piece dispensing and weighing electronic scale based on machine vision recognition and intelligent recognition method

    CN120685182A

  • SLAM method based on double-flow feature fusion

    CN120740570A

  • Target detection model and method based on deep learning

    CN121095726A