Accurate pest detection method based on depth vision converter
Through the improved YOLOv8 detection model, combined with the EfficientViT network and AFPN network, the problems of high missed detection rate, difficulty in early identification and poor target accuracy in pest detection are solved, and the accuracy and practicality of the detection are improved, providing more reliable technical support for agriculture.
Patent Information
- Application Number
- CN202510131392.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, pest detection has problems such as high missed detection rate, difficulty in early identification, and poor accuracy of small target detection, resulting in low detection accuracy.
Pest detection is performed through improved YOLOv8 detection models, including EfficientViT network, AFPN network and prediction network, using precise pest detection methods based on deep vision transformers. This model is characterized by the EfficientViT network, feature fusion of AFPN network, and detection is performed in the detection layer of the prediction network at three scales: large, medium and small.
It improves the accuracy and practicality of pest detection, reduces the missed detection rate and early identification difficulties, and provides more reliable technical support for agricultural production.
Smart Images

Figure CN119992336A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of accurate pest detection, and in particular relates to an accurate pest detection method based on a deep vision transformer. Background Art
[0002] High-frequency pest outbreaks pose a serious threat to agricultural production and cause significant economic losses. Therefore, it is of great significance to accurately and timely detect pests and reduce the impact of pests on agriculture, guide farmers to plant scientifically, and promote the continuous development of agriculture towards high quality and high yield. Traditional pest detection tasks require agricultural experts to work on the spot, but the agricultural planting area is large and there are many types of pests. There are problems such as time-consuming, labor-intensive, and insufficient agricultural experts. With the development of machine learning and deep learning, some new technologies have gradually replaced the original manual detection methods. However, machine learning needs to rely on manual extraction of pest feature information, which is easily affected by subjective factors. The extracted features have poor robustness, resulting in low accuracy of detection results. Deep learning technology based on convolutional neural networks automatically extracts feature information of pests through convolutional neural networks, effectively making up for the shortcomings of traditional machine learning, and greatly improving both feature extraction capabilities and detection accuracy. Summary of the invention
[0003] In order to solve the above technical problems, the present invention proposes an accurate pest detection method based on a deep vision transformer, which can solve the problems existing in the prior art such as high missed detection rate, difficulty in early identification, and poor accuracy in detecting small targets, improve the accuracy and practicality of pest detection, and provide more reliable technical support for agricultural production.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is: an accurate pest detection method based on a deep visual transformer, preparing pest images to be detected, dividing them into training sets and verification sets, and inputting them into an improved YOLOv8 detection model for pest detection, wherein the improved YOLOv8 detection model includes an EfficientViT network, an AFPN network and a prediction network, and the model is named EAS-YOLO; the collected images are first input into the EfficientViT network for feature extraction, and then all the extracted features are fused in the AFPN network, and finally the fused features are input into the large, medium and small scale detection layers of the prediction network for detection to obtain the detection results.
[0005] The improved YOLOv8 detection model uses YOLOv8n as the basic model and replaces the backbone network with the EfficientViT network. The EfficientViT network is composed of an efficient convolution module integrated into the Transformer architecture. The efficient convolution module combines the global self-attention mechanism and the local convolution kernel. The global self-attention mechanism adopts a lightweight multi-scale linear attention module. The multi-scale linear attention module aggregates nearby tokens by using small convolution kernels to generate multi-scale tokens and performs ReLU-based global attention on these multi-scale tokens. The local convolution kernel first receives the input feature map and then performs a convolution operation on each local area of the input feature map. For each position on the input feature map, the local convolution kernel generates an output value.
[0006] Furthermore, the improved YOLOv8 detection model replaces the FPN-PAN structure of the neck network with an AFPN (asymptotic feature pyramid network) structure; the AFPN structure first obtains the feature map of the bottom layer and performs an upsampling operation, and then adds the upsampled feature map to the feature map of the previous layer element by element. This process can be iterative, and each time the new fused feature map is used as the input of the next round of fusion until the feature map of the highest layer is reached.
[0007] Furthermore, the improved YOLOv8 detection model replaces the C2f module of the neck network with a self-designed C2fSE module. The input features of the C2fSE module first pass through a CBS module and then perform feature map segmentation. The input feature map is divided into two, one part of the feature map enters the subsequent bottleneck module for processing, and the other part directly participates in the subsequent splicing operation, and then passes through n bottleneck-SE modules. The bottleneck-SE module is composed of a bottleneck module and a SENet module, and the bottleneck module includes multiple convolutional layers; the SENet module first uses global average pooling to calculate the global features of each channel, and then learns the importance weights of the channels through two fully connected layers and a nonlinear activation function, and finally multiplies the channel weights obtained in the previous step by the original feature map channel by channel; then, the output feature maps of all bottleneck-SE modules and the previously segmented feature maps directly involved in the splicing are spliced, and then pass through a CBS module to obtain the output features of the C2fSE module.
[0008] Furthermore, the collected pest images are from the IP102 dataset.
[0009] Furthermore, the EfficientViT network includes a first Conv module, a first DSConv module, a first MBConv module, a second MBConv module, a third MBConv module, a first EfficientViT Module structure, a fourth MBConv module, a second EfficientViT Module structure, a fifth MBConv module and a first SPPF module in sequence. The main function of the EfficientViT network is to extract features from the input image and transmit the extracted features to the AFPN network.
[0010] Further, the processing process of the AFPN network is as follows: feature P2 is obtained from the fourth MBConv module of the EfficientViT network, feature P2 is processed by the second EfficientViT Module module to obtain feature P3, feature P3 is fused with feature P1 extracted by the second MBConv module of the EfficientViT network and feature P2 extracted by the fourth MBConv module, and the fused features are processed by the fifth MBConv module and the first SPPF module to obtain feature P4; the bottom-level features P2 and P3 are first input into the feature pyramid network, and after feature fusion, they are respectively passed through a C2fSE module to obtain features P5 and P6, and then P4 is added, and features P7, P8 and P9 are obtained after feature fusion and C2fSE modules again; feature P7 is input into the large target detection layer for detection, feature P8 is input into the medium target detection layer for detection, and feature P9 is input into the small target detection layer for detection.
[0011] The present invention takes YOLOv8n as the basic model and improves it. The backbone network in YOLOv8n is replaced with the EfficientViT network, so that the model can capture richer image information while maintaining efficient calculation; the FPN-PAN module of YOLOv8n is replaced with the AFPN module to solve the problem of feature loss and degradation in the multi-scale fusion process; the C2f module in YOLOv8n is replaced with the C2fSE module to improve the image capture ability of the algorithm while reducing noise or redundancy. The present invention helps farmers to accurately and timely detect pests, reduce the impact of pests on agricultural production, guide farmers to plant scientifically, and promote the continuous development of agriculture towards high quality and high yield. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a diagram of the improved YOLOv8 detection model framework.
[0013] Figure 2Schematic diagram of the improved YOLOv8 detection model structure.
[0014] Figure 3 Schematic diagram of AFPN structure.
[0015] Figure 4 Schematic diagram of the C2fSE structure. DETAILED DESCRIPTION
[0016] The present invention is further described in detail below with reference to the embodiments and drawings. Figure 1 , an accurate pest detection method based on deep visual transformer, obtains images from IP50 pest dataset and inputs them into improved YOLOv8 detection model for pest detection, the improved YOLOv8 detection model includes EfficientViT network, AFPN network and prediction network; the obtained image is first input into EfficientViT network for feature extraction, then all the extracted features are fused in AFPN network, and finally the fused features are input into large target detection layer, medium target detection layer and small target detection layer of prediction network for detection to obtain detection results.
[0017] This embodiment uses YOLOv8n as the basic model and improves it. Figure 2 As shown in the figure, the improved YOLOv8 detection model includes an EfficientViT network, an AFPN network and a prediction network. The backbone network of YOLOv8n is replaced with an EfficientViT network, which includes the first Conv module, the first DSConv module, the first MBConv module, the second MBConv module, the third MBConv module, the first EfficientViT Module structure, the fourth MBConv module, the second EfficientViT Module structure, the fifth MBConv module and the first SPPF module. The main function of the EfficientViT network is to extract features from the input image and transmit the extracted features to the AFPN network. The EfficientViTModule structure is composed of an MSA module and an MBConv module. The MSA module first obtains Q / K / V tokens through the Linear layer, and then aggregates nearby tokens through convolution to generate multi-scale Q / K / V tokens. Then, ReLU linear attention acts on the multi-scale tokens, and the outputs are connected and transmitted to the final Linear layer for feature fusion.
[0018] like Figure 2As shown, in order to improve the feature extraction capability of the neck network, this embodiment replaces the neck network of the basic model YOLOv8n with the AFPN module. Feature P2 is obtained from the fourth MBConv module of the EfficientViT network, and feature P2 is processed by the second EfficientViT Module module to obtain feature P3. Feature P3 is fused with feature P1 extracted by the second MBConv module of the EfficientViT network and feature P2 extracted by the fourth MBConv module. The fused features are then processed by the fifth MBConv module and the first SPPF module to obtain feature P4. The bottom features P2 and P3 are first input into the feature pyramid network, and after feature fusion, they are respectively processed by a C2fSE module to obtain features P5 and P6, and then P4 is added, and features P7, P8 and P9 are obtained again after feature fusion and C2fSE modules. Feature P5 is input into the large target detection layer for detection, feature P6 is input into the medium target detection layer for detection, and feature P7 is input into the small target detection layer for detection.
[0019] like Figure 3 As shown in the figure, before feature fusion, the last layer of features is extracted from the EfficientViT network, represented as {P2, P3, P4}. The bottom-level features P3 and P4 are first input into the feature pyramid network, and then P2 is added. After feature fusion, a set of multi-scale features {P5, P6, P7} are extracted. In the bottom-up feature extraction process of Backbone, the AFPN structure first fuses two adjacent low-level features, and then gradually fuses high-level features.
[0020] like Figure 4 As shown in the figure, the input features of the C2fSE module first pass through a CBS module, and then perform feature map segmentation. The input feature map is divided into two, one part of the feature map enters the subsequent bottleneck module for processing, and the other part directly participates in the subsequent splicing operation, and then passes through n bottleneck-SE modules, the bottleneck-SE module is composed of a bottleneck module and a SENet module, and the bottleneck module contains multiple convolutional layers; the SENet module first uses global average pooling to calculate the global features of each channel, and then learns the importance weights of the channels through two fully connected layers and a nonlinear activation function, and finally multiplies the channel weights obtained in the previous step by the original feature map channel by channel; then, the output feature maps of all bottleneck-SE modules and the previously segmented feature maps directly involved in the splicing are spliced, and then pass through a CBS module to obtain the output features of the C2fSE module.
[0021] This example uses the IP50 dataset of agricultural pests with 50 categories and relatively uniform numbers for training. 50 categories of pests are selected from the IP102 dataset for target classification tasks. After data cleaning and labeling, the images are subjected to data enhancement processing, including 90-degree rotation (clockwise, counterclockwise, and upside-down), position transformation, and size change, and finally the IP50 dataset is obtained.
[0022] The experimental environment configuration of this embodiment is: NVIDIA GeForce RTX 3090 graphics card with 24GB video memory, Intel(R) Xeon(R) Gold 6330 CPU @ 2.00GHz processor, Windows 10 operating system; Pytorch 2.0 is used as the deep learning model framework, CUDA version is 11.8, and the image size is adjusted to 416 × 416 pixels. During the model training process, the learning rate is set to 0.01, the weight decay is set to 0.0005, the training batch size is 16, the training is 300 rounds, and the SGD optimizer is used.
[0023] The evaluation indicators of this embodiment are Parameters, GFLOPs (Giga floating point operations per second), P (precision), R (recall), mAP 50 (Average precision over all categories at IoU=0.5).
[0024] The ablation experiment gradually deletes or modifies specific parts of the model to observe how these changes affect the performance of the model. The ablation experiment results of this embodiment are shown in Table 1. After improving the backbone network of YOLOv8, the P and mAP of the model are improved. 50 The P and mAP of the model with AFPN introduced on the basis of EfficientViT are 2.9% and 2.8% respectively, which shows that the introduction of EfficientViT enables the model to better focus on small target pests, thereby improving the detection performance of the model. 50 It brings 0.7% and 0.3% improvement respectively, and reduces the amount of calculation by 0.5, which shows that the feature fusion method of AFPN can enhance the network's ability to fuse multi-scale features. Adding the C2fSE module proposed in this paper to EfficientViT improves P and mAP 50 It brings 0.9% and 0.5% improvement respectively, indicating that the C2fSE module can help the model focus on key features more accurately. Based on the premise that EfficientViT is used as the backbone network, the C2fSE module is added to the AFPN structure to increase P and mAP 50 They increased by 1.4% and 1% respectively, reaching 88.5% and 89.8%, which proves the effectiveness of the algorithm improvement.
[0025] Table 1. Ablation experiment results Number +EfficientViT +AFPN +C2F P / % <![CDATA[mAP 50 / %]]> GFLOPs YOLOv8n 83.5 85.7 8.1 1 √ 86.4 88.5 9.5 2 √ 85.4 87.6 7.7 3 √ 83.4 86.1 8.1 4 √ √ 87.1 88.8 9.0 5 √ √ 87.3 89.0 9.5 6 √ √ 86.1 87.6 7.7 7 √ √ √ 88.5 89.8 9.0
[0026] The comparative experimental results of this embodiment are shown in Table 2. It can be seen that the algorithm in this paper is significantly better than other models in terms of mAP value under multiple IoU threshold conditions, and the accuracy P is also the highest. Compared with the YOLOv8n algorithm, the accuracy P is increased by 5%, the recall rate R is increased by 6%, and the mAP is 50 and mAP 50:95 They increased by 4.1% and 7.9% respectively, with the number of parameters only increasing from 3.01M to 4.26M, and GFLOPs only increasing from 8.1 to 9.0. Among them, the number of parameters and GFLOPs of the YOLOv5n algorithm are the lowest, but its precision P, recall R, and mAP are not ideal. Although the recall R of YOLOv5s reaches 85.6%, the other indicators are not as good as the EAS-YOLO algorithm.
[0027] Table 2. Comparative experimental results Algorithm P / % R / % <![CDATA[mAP 50 / %]]> <![CDATA[mAP 50:95 / %]]> Parameters / M GFLOPs YOLOv3-tiny 82.5 75.4 82.6 55.0 12.15 19.0 YOLOv5n 81.1 76.4 83.5 56.7 2.51 7.1 YOLOv5s 86.8 85.6 89.3 58.8 7.14 16.2 YOLOv8s 86.1 84.7 89.3 64.0 11.14 28.5 YOLOv8n 83.5 79.2 85.7 58.8 3.01 8.1 YOLOv10n 85.8 82.1 88.3 62.1 2.71 8.0 MobileNetV4 76.6 73.6 79.6 53.5 3.8 8.4 EAS-YOLO 88.5 85.2 89.8 66.7 4.26 9.0
[0028] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. An accurate pest detection method based on deep visual transformer, characterized in that: Pest images are collected and input into an improved YOLOv8 detection model for pest detection. The improved YOLOv8 detection model includes an EfficientViT network, an AFPN network and a prediction network, and the model is named EAS-YOLO. The pest image is first input into the EfficientViT network for feature extraction, and then all the extracted features are fused in the AFPN network. Finally, the fused features are input into the large, medium and small scale detection layers of the prediction network for detection to obtain the detection results.
2. The method for accurate pest detection based on deep vision transformer according to claim 1, characterized in that: The backbone network is replaced with the EfficientViT network, which is composed of an efficient convolution module integrated into the Transformer architecture. The efficient convolution module combines the global self-attention mechanism with the local convolution kernel. The global self-attention mechanism adopts a lightweight multi-scale linear attention module. The multi-scale linear attention module aggregates nearby tokens using small convolution kernels to generate multi-scale tokens and performs ReLU-based global attention on these multi-scale tokens. The local convolution kernel first receives the input feature map, and then performs a convolution operation on each local area of the input feature map. For each position on the input feature map, the local convolution kernel generates an output value.
3. The accurate pest detection method based on deep vision transformer according to claim 1, characterized in that: The FPN-PAN structure of the neck network is replaced with an AFPN (asymptotic feature pyramid network) structure; the AFPN structure first obtains the feature map of the bottom layer and performs an upsampling operation, and then adds the upsampled feature map to the feature map of the previous layer element by element. This process can be iterated, and each time the new fused feature map is used as the input of the next round of fusion until the feature map of the highest layer is reached.
4. The method for accurate pest detection based on deep vision transformer according to claim 1, characterized in that: The C2f module of the neck network is replaced with a self-designed C2fSE module. The input features of the C2fSE module first pass through a CBS module and then perform feature map segmentation. The input feature map is divided into two. One part of the feature map enters the subsequent bottleneck module for processing, and the other part directly participates in the subsequent splicing operation, and then passes through n bottleneck-SE modules. The bottleneck-SE module consists of a bottleneck module and a SENet module. The bottleneck module contains multiple convolutional layers. The SENet module first uses global average pooling to calculate the global features of each channel, and then learns the importance weights of the channels through two fully connected layers and a nonlinear activation function. Finally, the channel weights obtained in the previous step are multiplied by the original feature map channel by channel. After that, the output feature maps of all bottleneck-SE modules and the previously segmented feature maps directly involved in the splicing are spliced, and then pass through a CBS module to obtain the output features of the C2fSE module.
5. The accurate pest detection method based on deep vision transformer according to claim 1, characterized in that: The collected pest images are from the IP50 dataset.
6. The method for accurate pest detection based on deep vision transformer according to claim 5, characterized in that: Before inputting the IP50 data set into the improved YOLOv8 model, the method further includes: preprocessing the IP50 data set.
Citation Information
Cited By
Improved YOLO11-based water hyacinth target rapid detection method
CN120635395A