Lightweight forestry pest target detection system based on multi-attention mechanism feature processing and generation method
Through the lightweight forestry pest target detection system based on the characteristics of multiple attention mechanisms, the problem of low detection accuracy of small objects under low parameter limits is solved, and efficient pest target detection in forestry environments is achieved.
Patent Information
- Application Number
- CN202510466441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
The existing forestry pest target detection system is difficult to achieve high accuracy of detection of small objects under low parameter limits, and the hardware environment limitations lead to poor detection results.
A lightweight forestry pest target detection system based on feature processing based on multiple attention mechanisms is adopted, including the main backbone network layer, SATC encoder layer, decoder layer and output layer. A lightweight multi-scale attention feature extraction module, space group enhancement feature interaction module and feature pyramid structure are used to reduce the system parameters while improving detection accuracy.
Under low parameter conditions, the accuracy of pest target detection is significantly improved, efficient small object detection is achieved, and hardware needs in forestry environments are met.
Smart Images

Figure HDA0005358622810000011 
Figure HDA0005358622810000012 
Figure HDA0005358622810000013
Abstract
Description
Technical Field
[0001] This invention relates to the field of forest pest target detection. Specifically, it designs a lightweight forest pest target detection system and generation method based on multi-attention feature processing. This system and method aims to deeply mine key feature information from forest pest images under low parameter constraints through innovative feature extraction and fusion strategies, thereby improving the accuracy of pest target detection. Background Art
[0002] Pest management in precision forestry requires efficient and accurate target detection technology, but the diversity of pests and the complexity of their environments pose a significant challenge to this technology. Specifically, pests vary in size, from tiny aphids to large longhorn beetles. This vast difference in scale requires detection algorithms to possess multi-scale processing capabilities, capable of flexibly adapting to targets of varying sizes. Furthermore, the diverse colors, shapes, and textures of pests make comprehensive coverage difficult with traditional methods based on fixed features, further highlighting the advantage of deep learning in automatically learning complex features.
[0003] Compared to traditional target detection methods, deep learning-based solutions demonstrate significant advantages. Instead of relying on manually designed features, they automatically learn high-level, abstract features from massive amounts of data through convolutional neural networks. These features are more robust to factors such as illumination variations, occlusion, and deformation. Furthermore, the parallel computing capabilities of deep learning models enable greater efficiency in processing large-scale datasets, meeting the demands of detection applications. With the continuous advancement of computing power and algorithm optimization, deep learning-based target detection methods will play an even more important role in forest pest detection. By combining more advanced network architectures and optimization algorithms, such as computer vision, it is expected that automatic identification, classification, counting, and behavioral analysis of pests can be achieved, providing strong technical support for the precise management and protection of forest ecosystems.
[0004] However, in real-world applications, hardware limitations mean existing detection systems often struggle to achieve the required model performance. First, existing detection systems haven't undergone specialized refinements, resulting in suboptimal detection results for small-target pests. Second, existing detection systems have a large number of parameters, which doesn't meet the hardware limitations of forestry applications. Low-parameter detection systems often underperform, making them difficult to implement. Summary of the Invention
[0005] In order to solve the requirement of high small object detection accuracy in the field of forestry pest target detection under the limitation of low target detection system parameters, the present invention proposes a lightweight forestry pest target detection system based on multiple attention mechanism feature processing.
[0006] To achieve the above object, the present invention is implemented by the following technical solutions:
[0007] The present invention proposes a lightweight forest pest target detection system based on multi-attention mechanism feature processing. The system includes: a main backbone network layer, a SATC encoder layer, a decoder layer, and an output layer.
[0008] The purpose of the main backbone network layer is to extract rough features of forest pest images. It consists of a convolution of size 3×3, batch normalization, ReLU, and the lightweight multi-scale attention feature extraction module we proposed, which is used to enhance the system's image extraction ability while maintaining low system parameters.
[0009] The encoder layer is used to further extract the rough features output by the initial layer. It consists of SATC including the proposed AIFI-SGE and CTDA, which is used to further extract fine features while maintaining low system parameters.
[0010] The decoder layer can process three features of different sizes input, so as to be used by the output layer.
[0011] The output layer is used to output the probabilities of different classes and the information of the object positions. It consists of a series of fully connected layers.
[0012] The present invention proposes a method for generating a lightweight forest pest target detection system based on multi-attention mechanism feature processing, including the following steps:
[0013] (1) Build a development platform for implementing a lightweight forest pest target detection system based on multi-attention mechanism feature processing. The hardware platform of the present invention is based on an I5-13600KF CPU, an RTX 4060TI GPU (with a video memory of 16GB), and a memory of 128GB. The software platform is the Ubuntu 18.04 operating system, and has an operating environment with CUDA 11.3, Torchvision 0.15.2, and Python3.8.
[0014] (2) Forest pest image data division. The forest pest images contain 7,163 images of 31 different types of pests. To ensure the training effect, the forest pest dataset is randomly divided according to the following ratio: Train:Val = 9:1, that is, the training set contains 6,446 images, and the validation set contains 717 images.
[0015] (3) Build a lightweight forest pest target detection system based on multi-attention mechanism feature processing. The constructed forest pest detection system is as described above.
[0016] (4) Training and testing of the forest pest target detection system. The system in this embodiment adopts 300 rounds of training. During the training process, in order to improve the convergence speed and stability of the system, a validation set is used for verification after each round of training. To ensure fairness, all systems do not use pre-trained weights.
[0017] (5) The performance of the forest pest target detection system was evaluated. AP, AR and parameter quantity were used as experimental evaluation indicators. These three parameters have been widely used in the field of target detection.
[0018] We conducted comparative experiments with several leading and mainstream target detection systems, including YOLOv5 and LWDETR, another real-time detection system in the DETR series, which was developed concurrently with RTDETR. All the experimental systems used were official versions. The experimental results are shown in Table 1. Our proposed lightweight forest pest target detection system based on multi-attention feature processing achieved the best performance, demonstrating the advanced nature of our invention.
[0019] MODEL params <![CDATA[AP 0.5:0.95 > <![CDATA[AP 50 > <![CDATA[AP 75 > <![CDATA[AP S > <![CDATA[AP M > <![CDATA[AP L > mAR Efficient-Det-d7 51.9M 52.7% 71.9% 56.7% 36.2 56.9 66.2 38.5% LWDETR-S 14.6M 83.4% 96.1% 91.4% 45.5% 64.8% 88.2% 66.5% Yolov5-S 7.2M 90.1% 99.3% 97.0% 68.1% 80.8% 93.1% 69.2% Propose 12M 91.5% 99.0% 97.3% 75.2% 80.9% 94.5% 69.7%
[0020] Table 1 Test results BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0022] Figure 1 It is an overall block diagram of the system of the present invention;
[0023] Figure 2 is a schematic diagram of a lightweight multi-scale attention feature extraction module in the system of the present invention;
[0024] Figure 3 is a schematic diagram of a spatial group enhancement feature interaction module in the system of the present invention;
[0025] Figure 4 is a schematic diagram of the CKE module in the system of the present invention;
[0026] Figure 5 is a schematic diagram of the C3KE and BottleneckE modules in the system of the present invention;
[0027] Figure 6 is a schematic diagram of a lightweight multi-scale feature fusion architecture in the system of the present invention;
[0028] Figure 7 is a flow chart of the system generation method of the present invention;
[0029] Figure 8It is a comparison schematic diagram of the real target image and the detection result image obtained by the system of the present invention on the forestry pest dataset; Detailed implementation manners
[0030] Figure 1 It is the overall block diagram of the system of the present invention. The system is composed of a main backbone network layer, a SATC encoder layer, a decoder layer, and an output layer.
[0031] Figure 2 It is a schematic diagram of the lightweight multi-scale attention feature extraction module in the system of the present invention. It is composed of a convolution with a size of 3×3, batch normalization, ReLU, and the lightweight multi-scale attention feature extraction module we proposed, which is used to enhance the system's ability to extract images while maintaining low system parameters. The lightweight multi-scale attention feature extraction module uses the lightweight large-field-of-view convolution (WTPConv) proposed by combining the local convolution (PConv) and the wavelet transform convolution (WTConv) and combines it with the efficient multi-scale attention (EMA) attention mechanism. PConv utilizes the redundant information in the feature map and only performs convolution operations on some input channels while keeping other channels unchanged. This can reduce computational redundancy and the number of memory accesses, reduce the number of parameters, and improve the computational speed. However, this approach may lose some low-frequency information. Since WTConv can obtain a nearly global receptive field without significantly increasing the number of parameters, it can improve the system's ability to capture low-frequency information. Therefore, we use WTConv to replace the traditional convolution in PConv. By combining these two convolutions, the WTPConv we proposed takes into account the high-efficiency parameter calculation of PConv and the low-parameter characteristics of WTConv and enhances the ability of the backbone network to recognize the receptive field, thereby improving the system's ability to detect small objects while significantly reducing the system's number of parameters. Based on the lightweight large-field-of-view convolution WTPConv we proposed, by introducing the EMA module, we proposed the lightweight multi-scale attention feature extraction module WTPConvEMA-Block. EMA (Efficient Multi-scale Attention) is an attention mechanism used in computer vision tasks. Its main purpose is to effectively capture multi-scale features without reducing the channel dimension. Therefore, we introduce this module to avoid losing some important information due to the reduction of the channel dimension caused by multiple convolutions.
[0032] Figure 3It is a schematic diagram of the spatial group enhanced feature interaction module in the system of the present invention. The Transformer in this module is quite efficient in distinguishing the features of different objects, and the SGE module depends on the similarity relationship between the global features and local features within the group to calculate the similarity factor. Therefore, by combining the scale interaction of the Transformer and the spatial group enhancement module SGE, the spatial distribution pattern of the feature map can be optimized, and the discriminability and noise resistance of the feature representation can be improved, ultimately achieving a systematic improvement in the detection accuracy.
[0033] Figure 4 It is a schematic diagram of the CKE module in the system of the present invention. CKE is the core module of CTDA, which is improved from C3K2. The C3K2 module is composed of the classic structure Bottleneck of the YOLO architecture and the C3K module composed of Bottlenecks, and is formed by the residual structure and multiple stacks, aiming to efficiently improve the feature extraction ability of the system. The design of the C3K2 module not only reduces the computational amount, but also significantly improves the feature expression ability through multi-level convolution operations and residual mechanisms, enabling it to efficiently extract complex features in the object detection task while maintaining high training stability and system performance.
[0034] Figure 5 It is a schematic diagram of the C3KE and BottleneckE modules in the system of the present invention. C3KE and BottleneckE are respectively composed of C3K combined with ELA and Bottleneck combined with ELA, and are used to form the basic module units of CKE.
[0035] Figure 6 It is a schematic diagram of the lightweight multi-scale feature fusion architecture in the system of the present invention. It mainly includes four modules: TFAM, CKE, Dy sample, and Adown. CKE is the core module of CTDA, which is improved from C3K2. And we use TFAM to replace the Concat operation. This module uses channel and spatial attention to determine the important parts of the features, and uses temporal information to determine the important parts between the dual-temporal features to further improve the system performance. Dy sample and Adown are lightweight and efficient upsampling and downsampling modules respectively. Using these two modules to replace the convolutional downsampling and upsampling in the original structure can further improve the system performance while reducing the number of parameters of the system. In summary, the CTDA we proposed has significantly improved the average performance and parameter efficiency of the system through the fusion of features at different scales and the combination of the feature pyramid structure.
[0036] Figure 7It is a flowchart of the system generation method of the present invention. The system generation method is mainly divided into five steps: (1) building a system development platform; (2) dividing forestry pest image data; (3) constructing a forestry pest target detection system based on feature processing with an attention mechanism; (4) training and testing the forestry pest target detection system; (5) evaluating the performance of the forestry pest target detection system.
[0037] Figure 8 It is a comparison schematic diagram of the real object image and the detection result image obtained by the system of the present invention on the forestry pest image dataset. In order to verify the performance of the system of the present invention, this embodiment is evaluated on the forestry pest image dataset. On the forestry pest image dataset, the system of the present invention has achieved remarkable detection effects, and the detection result image is very close to the real target image, fully demonstrating the advantages of the present invention.
Claims
1. A lightweight forest pest target detection system and generation method based on multi-attention mechanism feature processing, and the specific implementation includes the following steps: (1) Build a development platform for implementing the lightweight forest pest target detection system based on multi-attention mechanism feature processing. The hardware platform of the present invention is based on an I5-13600KF CPU, an RTX 4060TI GPU (with a video memory of 16GB), and a memory of 128GB. The software platform is the Ubuntu 18.04 operating system, and has an operating environment with CUDA 11.3, Torchvision 0.15.2, and Python 3.
8. (2) Division of forest pest image data. The forest pest images include 7,163 images of 31 different types of pests. To ensure the training effect, the forest pest dataset is randomly divided according to the following ratio: Train:Val = 9:1, that is, the training set contains 6,446 images, and the validation set contains 717 images. (3) Construct a lightweight forest pest target detection system based on multi-attention mechanism feature processing. The constructed forest pest detection system mainly includes: The purpose of the main backbone network layer is to extract the rough features of forest pest images. It consists of a convolution of size 3×3, batch normalization, ReLU, and the lightweight multi-scale attention feature extraction module we proposed, which is used to enhance the system's image extraction ability while maintaining low system parameters. The encoder layer is used to further extract the rough features output by the initial layer. It consists of the SATC containing the proposed AIFI-SGE and CTDA, which is used to further extract fine features while maintaining low system parameters. The decoder layer can process the three features of different sizes input to facilitate the output layer. The output layer is used to output the probabilities of different classes and the information of the object position, and it consists of a series of fully connected layers. (4) Training and testing of the forest pest target detection system. The system in this embodiment is trained for 300 rounds. During the training process, in order to improve the convergence speed and stability of the system, the validation set is used for verification after each round of training. To ensure fairness, no pre-trained weights are used in all systems. (5) Evaluate the performance of the forest pest target detection system.
2. The lightweight forestry pest target detection system based on multi-attention mechanism feature processing according to claim 1, characterized in that, It includes a main backbone network layer, an encoder layer, a decoder layer, and an output layer; the main backbone network layer integrates the WTPConvEMA-Block module, the encoder layer contains the AIFI-SGE module and the CTDA module, and the CTDA module consists of TFAM, CKE, DySample, and Adown.
3. The lightweight forest pest target detection system based on multi-attention mechanism feature processing according to claim 1, characterized in that, The WTPConvEMA-Block realizes efficient feature extraction with low parameters by fusing PConv and WTConv and combining the EMA attention mechanism.
4. The lightweight forest pest target detection system based on multi-attention mechanism feature processing according to claim 1, characterized in that The AIFI-SGE module combines the scale interaction of Transformer and the spatial group enhancement module SGE module to optimize the spatial distribution pattern of the feature map. Improve the discriminability and noise resistance of feature representation, and finally realize a systematic improvement in detection accuracy.
5. The lightweight forest pest target detection system based on multi-attention mechanism feature processing according to claim 1, characterized in that, The CKE module is a residual stacking structure improved based on C3K2, and uses ELA (Efficient Local Attention) to enhance the feature integration ability.