Multi-class pest identification and detection method and system based on Mask-RCNN-CBAM fusion attention mechanism
By introducing the Mask-RCNN-CBAM fusion attention mechanism in pest detection, the detection accuracy problem of small target pests in complex backgrounds is solved, and high-precision multi-type pest recognition and extraction is achieved, false positives and false negatives are reduced, and detection efficiency is improved.
Patent Information
- Application Number
- CN202510498078.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-01
AI Technical Summary
When existing pest detection technology deals with small target pests, especially in complex backgrounds, there are low detection accuracy, false alarms and missed responses, making it difficult to effectively extract multiple categories of small target pests.
Using a method based on Mask-RCNN-CBAM fusion attention mechanism, a Mask-RCNN-CBAM deep learning network is constructed by introducing the CBAM attention mechanism module and a multi-scale feature enhancement module, combining the attention mechanism and multi-level semantic features to improve the recognition accuracy of small-target pests.
Under complex backgrounds and pest-intensive conditions, the accuracy and performance of pest identification are significantly improved, the false alarm rate and missed alarm rate are reduced, and the detection accuracy and efficiency of the network are improved.
Smart Images

Figure CN120412019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pest identification and detection, and more specifically to a multi-class pest identification and detection method and system based on the Mask-RCNN-CBAM fusion attention mechanism. Background Art
[0002] Pests are one of the greatest challenges faced by global agricultural production. The annual crop yield loss caused by pests is estimated to be between 20% and 40%. These pests not only directly threaten food security, but also result in huge economic losses, exacerbate the famine crisis, and pose non-negligible ecological risks. In recent years, the yields of major crops such as rice, wheat, and corn have increased steadily, and the yields of cash crops such as cotton, soybeans, peanuts, fruits, and tea have also increased. However, the impact of climate change has exacerbated the pest and disease pressure, having a negative impact on crop yields in some regions. Since the 1970s, the frequency of pest and disease outbreaks has tripled, and rising temperatures have further accelerated the growth and reproduction of pests, especially at night. During the crop planting process, various crops face their own unique pest and disease challenges. The most common pests include Spodoptera frugiperda, Cnaphalocrocis medinalis, Ostrinia furnacalis, thrips, Helicoverpa armigera, locusts, and Agrotis ypsilon, etc. They not only reduce crop yields and cause economic losses, but if the prevention and control are not timely, the crops may even face the serious consequence of complete failure.
[0003] Therefore, quickly identifying and effectively controlling major pests is crucial for reducing yield losses and ensuring food security. In the early years, agronomists mainly relied on manual counting of pests to predict pest outbreaks. This method not only requires a large amount of manpower, but is also time-consuming and inefficient. With the progress of technology, the emergence of intelligent pest monitoring lights has solved the problem of manual counting. These lights use trapping lights to attract specific pests with light of a specific wavelength. These lights capture pest images at preset time intervals through industrial cameras, and then remotely transmit the images to a designated network port for processing, and use object detection algorithms to classify and count the pests in the images. The core challenge of the intelligent pest monitoring system focuses on the detection and recognition algorithms of pest images. This algorithm generally uses object detection technology in the field of computer vision to achieve pest detection. Currently, the research on pest detection mainly focuses on improving the efficiency of recognition and classification algorithms. In particular, deep learning models have achieved excellent performance in agricultural pest detection. For example, the YOLO series algorithms, SSD (Single Shot MultiBox Detector), and Faster R-CNN (Faster Region-based Convolutional Neural Network) provide a good balance in finding a balance between the speed and accuracy of pest detection. In addition, the introduction of attention mechanisms, multi-modal fusion, and super-resolution technologies has made significant progress in improving pest detection performance.
[0004] However, detecting small or densely distributed small-target pests remains challenging, particularly due to difficulties in extraction, susceptibility to background interference, and poor performance in complex environments. To address these issues, research should focus on improving specialized technologies for small-target pest detection and exploring efficient, lightweight models suitable for agricultural production scenarios. Compared to traditional machine learning methods, image recognition models based on deep learning can achieve higher accuracy. Among various deep learning network models, Mask-RCNN (masked region-based convolutional neural network) offers significant advantages in pest detection and identification accuracy. For example, the Jinye Intelligent Insect Monitoring Light utilizes an improved Mask-RCNN model to quickly and accurately identify and classify a variety of pest species. An apple orchard pest identification method based on the improved Mask R-CNN has also demonstrated its high accuracy and robustness in pest identification tasks. Mask-RCNN is an instance segmentation model that adds a mask branch to FasterRCNN for pixel-by-pixel segmentation. It utilizes anchor boxes for classification and regression, combining pixel-level segmentation and classification for more accurate classification. In recent years, many researchers have focused on the detection and identification of pests. Mask-RCNN has been widely adopted as a foundational recognition algorithm in numerous studies. For example, research by DEEPIKA et al. demonstrated that the algorithm achieved an accuracy exceeding 92% when detecting a single target or single category within each image. Furthermore, research by MENDOZA et al. also demonstrated that Mask-RCNN achieved high accuracy in detecting multiple categories within a single image. Liu et al. achieved an accuracy of 92% for detecting four sparsely distributed categories, but this accuracy did not account for non-target similar pests or high-density conditions, leaving room for improvement. Therefore, while Mask-RCNN excels in single-image multi-target classification, further research and optimization are needed for detecting small, multi-category targets.
[0005] Therefore, due to the high intra-class similarity, large target scale variation, and complex background of small target pest images collected by pest monitoring lights, the detection accuracy is low, and there are false positives and missed positives. How to propose a multi-class pest recognition and detection method and system based on the Mask-RCNN-CBAM fusion attention mechanism to solve the problems of complex background and insufficient multi-scale feature extraction, and combine the attention mechanism with multi-level semantic features to achieve accurate extraction of small target pests is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0006] In view of this, the present invention innovatively proposes a multi-class pest recognition and detection method and system based on the Mask-RCNN-CBAM fusion attention mechanism, aiming to address the problems of complex backgrounds and multi-scale feature extraction. By combining the attention mechanism with multi-level semantic features, high-precision recognition of small target pests is achieved. To achieve the above objectives, the present invention specifically adopts the following technical solutions:
[0007] A multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism, comprising:
[0008] Collect multi-class pest image data and construct an experimental data set;
[0009] Introduce the CBAM attention mechanism module in the feature processing stage of the Mask-RCNN model, and connect the multi-scale feature enhancement module to construct the Mask-RCNN-CBAM deep learning network;
[0010] Train the Mask-RCNN-CBAM deep learning network through the experimental data set;
[0011] Collect real-time data and input it into the trained Mask-RCNN-CBAM deep learning network to obtain multi-class pest recognition and detection results.
[0012] Optionally, the CBAM attention mechanism module includes a channel attention module and a spatial attention module connected in sequence. After the backbone extracts the backbone features, the input image is sent into the CBAM attention mechanism module. In this module, the channel attention module first acts on the input features, fuses the generated channel attention features with the original features, and then generates the final channel attention features. The generated channel attention features are input into the spatial feature map, and the final attention features are obtained through pooling and convolutional connections.
[0013] Optionally, the channel attention module aggregates the spatial information of the channel feature map through average pooling and max pooling operations to generate two different spatial context descriptors, representing the average pooling feature and the max pooling feature respectively; subsequently, these features are transmitted to the shared network, which includes a multi-layer perceptron MLP and a hidden layer. The activation size of the hidden layer is carefully set to C / r, and finally a 1×1×C channel attention map Mc(F) is generated, where C is the number of neurons and r is the attenuation rate. The activation function is ReLu. After applying the shared network to each descriptor, the sum of the combined feature elements is output as a feature vector.
[0014] Optionally, the calculation formula for channel attention is:
[0015]
[0016] Among them: F C max =MaxPool(F) is the global maximum pooling, F C avg =AvgPool(F) is the global average pooling, σ is the S-type function; W0 and W1 are the shared weights of the two input features.
[0017] Optionally, the spatial attention module generates a feature map by exploiting the spatial relationships of features. It then uses two pooling layers to aggregate the channel information of the feature map to generate two 2D features, representing the average pooled feature and the maximum pooled feature. The processed features are concatenated and convolved through a standard convolutional layer and then fed into a sigmoid function to generate the spatial attention feature map.
[0018] Optionally, the calculation formula of the spatial attention is:
[0019] M s (F)=σ(f 7×7 (AvgPool(F)));
[0020] (MaxPool(F)))=σ(f 7×7 (F avg ; F max ));
[0021] Where: M s (F) is the spatial feature map, σ is the S-shaped function; f 7×7 Represents a convolution operation with a filter size of 7×7.
[0022] The multi-scale feature enhancement module consists of two main components: top-down downsampling feature extraction and bottom-up upsampling feature extraction. The downsampling component aims to extract deeper semantic information and stronger image features to obtain high-resolution features. The upsampling feature extraction, on the other hand, focuses on the spatial information of shallow-layer features, transferring them to deeper layers for fusion. Ultimately, the transferred features from the upper layers are combined with the shallow spatial information and deep semantic information. During the feature transfer process, a dual-channel downsampling convolution operation is used to effectively minimize the loss of detailed features during downsampling, ensuring that image details are preserved.
[0023] Optionally, the dual-channel downsampling convolution operation includes: fusing two transition modules, the left branch adopts a 2×2 maximum pooling operation, followed by a 1×1 convolution layer, the right branch adopts a 1×1 convolution layer, followed by a 3×3 convolution layer, with a stride of 2×2, and the two branches superimpose the results and output them.
[0024] Optionally, the loss function of the Mask-RCNN-CBAM deep learning network includes:
[0025] L = L cls + L box + L mask ;
[0026] In the formula, L cls represents the classification loss, which is used to judge the category to which each ROI belongs; L box represents the bounding box offset loss, which is used to regress the bounding box of each ROI; L mask represents the pixel segmentation mask generation loss, which is used to generate a mask for each ROI and each category.
[0027] Optionally, a multi-class pest recognition and detection system based on the Mask-RCNN-CBAM fusion attention mechanism includes:
[0028] An acquisition module: used to acquire multi-class pest image data and construct an experimental data set;
[0029] A model construction module: used to introduce a CBAM attention mechanism module in the feature processing stage of the Mask-RCNN model and connect a multi-scale feature enhancement module to construct a Mask-RCNN-CBAM deep learning network;
[0030] A training module: used to train the Mask-RCNN-CBAM deep learning network through the experimental data set;
[0031] A recognition and detection module: used to collect real-time data and input it into the trained Mask-RCNN-CBAM deep learning network to obtain multi-class pest recognition and detection results.
[0032] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a multi-class pest recognition and detection method and system based on the Mask-RCNN-CBAM fusion attention mechanism, and has the following beneficial effects:
[0033] The present invention proposes a multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism, including: collecting multi-class pest image data and constructing an experimental data set; introducing a CBAM attention mechanism module in the feature processing stage of the Mask-RCNN model and connecting a multi-scale feature enhancement module to construct a Mask-RCNN-CBAM deep learning network; training the Mask-RCNN-CBAM deep learning network through the experimental data set; collecting real-time data and inputting it into the trained Mask-RCNN-CBAM deep learning network to obtain multi-class pest recognition and detection results.
[0034] In the input stage of the present invention, the image passes through the feature extraction layer of the ResNet101 backbone network, and the extracted features are fed into the CBAM attention mechanism module. The attention mechanism helps the model focus on important regions for processing. By sequentially connecting the channel attention module and the spatial attention module, the learning ability of the network for complex object features is improved, and false positives and false negatives in complex backgrounds are avoided. The network also introduces a multi-scale feature fusion pyramid module to perform feature transmission at different scales, alleviating the impact of insufficient multi-scale features on object detection. To further solve the problem of information loss caused by traditional downsampling methods, a dual-channel downsampling module is introduced to retain information and extract features using two channels, thereby improving the model accuracy. The method proposed by the present invention performs better in detecting pest-dense areas and extracting pests. The Mask-RCNN-CBAM network effectively filters out complex backgrounds, improves the recognition rate of pest features, and is more sensitive to the background. In various pest detection tasks, the Mask-RCNN-CBAM network exhibits good recognition performance, with higher accuracy and smoother masks. This is because the network integrates a dual-channel attention mechanism, assigns greater weights to pest features, while reducing the impact of background features, effectively distinguishing pest and background features. In addition, the feature-enhanced FPN (Feature Pyramid Network) fuses the shallow and deep features of the image, further enriching the detailed information and significantly improving the detection accuracy and efficiency of the network. Aiming at the problems of false alarms and missed detections in pest extraction caused by complex backgrounds and dense pest superposition, a Mask-RCNN-CBAM pest extraction network integrating an attention mechanism is proposed in the present invention. By integrating the CBAM attention mechanism into the feature processing process, enhancing multi-scale feature extraction and fusion through the feature pyramid module, the ability of the network to extract context information is enhanced, and the loss of image detail features is reduced. Experimental results show that the Mask-RCNN-CBAM network has excellent target extraction effects on the pest dataset. Especially under complex backgrounds and dense pest conditions, it achieves high accuracy and good performance, achieving the highest performance in terms of Precision, F1score, Recall, etc., and having a low false alarm rate and missed detection rate, indicating that the network model is reliable and applicable. Compared with other pest extraction methods, the Mask-RCNN-CBAM network can better extract pest feature information and optimize detailed information. Brief Description of the Drawings
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0036] Figure 1 Schematic diagram of the Mask-RCNN network framework provided by the present invention.
[0037] Figure 2 Schematic diagram of the process of a multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism provided by the present invention.
[0038] Figure 3 Structural framework diagram of the channel attention module provided by the present invention.
[0039] Figure 4 Structural framework diagram of the spatial attention module provided by the present invention.
[0040] Figure 5 Structural framework diagram of the multi-scale feature enhancement module provided by the present invention.
[0041] Figure 6 Structural framework diagram of the dual-channel downsampling module provided by the present invention.
[0042] Figure 7 Visualization schematic diagram of the comparison results of classical network experiments provided by the present invention. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] The embodiments of the present invention disclose a multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism, including:
[0045] Collect multi-class pest image data and construct an experimental data set;
[0046] Introduce the CBAM attention mechanism module in the feature processing stage of the Mask-RCNN model, and connect the multi-scale feature enhancement module to construct the Mask-RCNN-CBAM deep learning network;
[0047] Train the Mask-RCNN-CBAM deep learning network through the experimental data set;
[0048] Collect real-time data and input it into the trained Mask-RCNN-CBAM deep learning network to obtain multi-class pest recognition and detection results.
[0049] In the specific implementation, image data of multiple types of pests are collected, and the experimental dataset is constructed, including:
[0050] First, the training dataset is filtered to remove image samples with low quality. To ensure the high recognition accuracy of the deep learning algorithm, the image resolution needs to reach a high standard. Therefore, images with a resolution lower than 4096×2160 and images containing incomplete leaves or pests are excluded. Next, images in different seasons are extracted in a certain proportion to ensure that there are sufficient numbers of images for each type of pest to meet the requirements of deep learning training.
[0051] Given the diversity of pest pictures, a dataset containing multiple types of pests is constructed. Before the experiment, all pictures are uniformly processed. The pictures are cropped into nine equal parts, and the resolution of each cropped picture is set to 1824 pixels × 1216 pixels. To ensure the diversity of experimental samples, data augmentation techniques are used to increase the size of the dataset, including horizontal flipping, vertical flipping, 90° rotation, 180° rotation, 270° rotation, adding noise, etc. Through these operations, the size of the dataset is increased by 7 times, totaling 7000 pictures. The LabelMe tool is used for labeling, and the labels of each picture are saved in the json file format in the directory where the picture is located. In the Mask-RCNN dataset training, the dataset format follows the COCO format specification, including picture information, annotation details, and class definitions, etc. Finally, the dataset is split into a training set, a validation set, and a test set according to the ratio of 6:1:3. In this embodiment, only the pests that cause serious damage to crops are labeled, and the pests that cause minor damage, beneficial insects, and non-crop pests are not labeled.
[0052] In the specific implementation, a multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism, the specific implementation steps include:
[0053] The network structure adopted in this embodiment is based on the classic Mask-RCNN architecture. The Mask-RCNN structure is simple and can be used for various tasks such as object detection, semantic segmentation, instance segmentation, and human pose recognition. Its network structure is as Figure 1As shown in the figure. The Mask-RCNN model follows the idea of Faster-RCNN. It adopts the ResNet-FPN architecture for feature extraction and adds a Mask prediction branch for semantic segmentation, with both detection and extraction functions. Different from Faster-RCNN using VGG as the backbone feature extraction network, Mask-RCNN uses ResNet-50 and ResNet-101 as the backbone feature extraction networks and adds a Feature Pyramid Network (FPN) to the backbone structure. Using different combinations of backbone networks will generate feature layers of different sizes. Then, through the Region Proposal Network (RPN), Anchor boxes are generated at each point of the effective feature layer for rough screening to produce local feature layers. After that, ROI alignment is performed on these local feature layers, and they are fed into the classification, regression model, and Mask model for classification and mask generation, and finally, the classification and segmentation results are output.
[0054] As Figure 2 As shown in the figure, the Mask-RCNN-CBAM deep learning network proposed in this embodiment introduces the CBAM (Convolutional Block Attention Module) attention mechanism in the feature processing stage based on the original Mask-RCNN structure. The CBAM module assigns weights to the feature information of different regions by enhancing the channel and spatial attention of the image, enabling the model to effectively learn and highlight the target features. Compared with the original network, the enhanced Feature Pyramid module adds a bottom-up feature transfer path, allowing shallow features to be smoothly transferred to deeper layers for effective fusion. At the same time, it fuses the features passed down from the upper layer with the spatial information of the shallow layer and the semantic information of the deep layer. During the feature transfer process, dual-channel downsampling convolution operations are used to reduce the loss of detailed features during downsampling, retain the image's detailed features, and improve the network's feature extraction and detail optimization capabilities.
[0055] Specifically, in the input stage, the image passes through the ResNet101 backbone network feature extraction layer, and the extracted features are fed into the CBAM attention mechanism module. The attention mechanism helps the model focus on important regions for processing. By connecting the attention module and the spatial attention module in sequence, the network's ability to learn the features of complex objects is improved, avoiding false positives and false negatives in complex backgrounds. The network also introduces a multi-scale feature fusion pyramid module to perform feature transfer at different scales, alleviating the impact of insufficient multi-scale features on object detection. To further solve the problem of information loss caused by traditional downsampling methods, a dual-channel downsampling module is introduced to retain information and extract features using two channels, thereby improving the model's accuracy.
[0056] In the specific implementation manner, the CBAM attention mechanism module includes:
[0057] CBAM is an optimization algorithm that combines a channel attention module and a spatial attention module. After the backbone features are extracted from the image by the backbone and fed into the CBAM module, the features will first pass through the channel attention module to aggregate the features generated by the channel attention with the input features to generate the final channel attention features. The generated features are input into the spatial feature map, and the final attention features are obtained through pooling and convolutional connections. The CBAM module, as an attention mechanism, can sequentially infer the attention maps of the channel and spatial dimensions for adaptive feature optimization. This mechanism is very effective in dealing with multi-scale feature extraction tasks, and can highlight the effective features of the target and reduce redundant information. In this embodiment, the channel attention map is generated using the channel attention relationship of the features. Each channel of the features represents a special feature detector, and the attention of different channel features will be assigned corresponding weight coefficients. The channel attention mechanism can calculate the channel attention more efficiently by focusing on meaningful features and compressing the spatial dimension of the input feature map. This enables the model to provide more accurate feature representations for regions with rich features in the image.
[0058] (1) As Figure 3 shown, the channel attention module utilizes the maximum pooling output and average pooling output of the shared network. The specific process includes: ① Aggregate the spatial information of the channel feature map through average pooling and maximum pooling operations to generate two different spatial context descriptors, representing the average pooling feature and the maximum pooling feature respectively; ② Transmit the generated features into the shared network to generate a 1×1×C channel attention map Mc(F). The shared network consists of a multi-layer perceptron MLP and a hidden layer. To reduce the parameter cost, the activation size of the hidden layer is set to C / r, where C is the number of neurons and r is the decay rate, and the activation function is ReLu. After applying the shared network to each descriptor, the sum of the combined feature elements is output as a feature vector. Briefly, the specific calculation method of the channel attention is as follows:
[0059]
[0060] where: F C max = MaxPool(F) is the global maximum pooling, F C avg = AvgPool(F) is the global average pooling, σ is the sigmoid function; W0 and W1 are the shared weights of the two input features.
[0061] (2) As Figure 4As shown, different from the channel attention mechanism, the spatial attention mechanism focuses on the location of features and mainly uses the spatial relationship of features to generate a spatial attention feature map. Spatial attention and channel attention are complementary to each other. When calculating spatial attention, first, average pooling and max pooling operations of channels are adopted, and then the results of these operations are concatenated to generate an effective feature descriptor. Then, the feature descriptor is used to generate the spatial feature map Ms(F) through a convolutional layer. The specific process includes: ① Using two pooling layers to aggregate the channel information of the feature map to generate two two-dimensional features, representing the average pooling feature and the max pooling feature respectively; ② Connecting and convolving the features through a standard convolutional layer and inputting them into the sigmoid function to generate the spatial attention feature map. In short, the calculation method of spatial attention is as follows:
[0062] M s (F) = σ(f 7×7 (AvgPool(F)));
[0063] (MaxPool(F)) = σ(f 7×7 (F avg ; F max ));
[0064] Where: σ is the sigmoid function; f 7×7 represents a convolutional operation with a filter size of 7×7.
[0065] In the specific implementation manner, the multi-scale feature enhancement module, as Figure 5 shown, includes two parts: downsampling feature extraction from top to bottom and upsampling feature extraction from bottom to top. Downsampling extracts deeper and more semantically rich image features, thus obtaining higher-resolution features. Upsampling shallow features often contain rich spatial information of objects, which is crucial for locating targets in the image. For example, when the input image feature is 1024×1024×3, a convolutional layer with a stride of 2 is used for dimensionality reduction. Since convolutional downsampling will cause certain feature loss, the IdentityBlock module is used to enhance the network during the sampling process, enabling deeper layers to learn more complex features. After being processed by this block, the features are marked as Ci, i ∈ [2, 3, 4, 5], where the C5 feature is fused with the C4 feature of the previous layer through upsampling. After upsampling, a convolutional operation is performed on the fused feature to generate a new feature P4. After the bottom-up upsampling, a pooling operation is performed on the C5 feature to reduce information redundancy and prevent overfitting, obtaining a new feature layer Pj, j ∈ [2, 3, 4, 5, 6], which contains more semantic and spatial information.
[0066] In the specific implementation manner, the dual-channel downsampling convolutional operation specifically includes:
[0067] The downsampling module is a method for reducing the resolution of an image or feature map. However, common downsampling methods often lead to the loss of detailed information. To effectively reduce the possible feature loss during downsampling, this embodiment particularly optimizes and improves the downsampling module. The improved dual-channel downsampling module integrates two transition modules. The left branch uses a 2×2 max pooling operation followed by a 1×1 convolutional layer, and the right branch uses a 1×1 convolutional layer followed by a 3×3 convolutional layer (stride 2×2). The results of the two branches are superimposed and then output. Compared with traditional methods, this module can better capture and process the features of the input image, improve the object detection performance, and reduce the size of the feature map without changing the depth of the feature map, effectively reducing information loss and enabling the network to retain important detailed information. The structure is as Figure 6 shown.
[0068] In the specific implementation, the loss function of the Mask-RCNN-CBAM deep learning network includes: The loss function of Mask-RCNN-CBAM is a multi-task loss function that combines the losses of classification, localization, and segmentation masks, as shown in the following formula:
[0069] L = L cls + L box + L mask ;
[0070] In the formula, L cls represents the classification loss, which is used to determine which category each ROI belongs to; L box represents the bounding box offset loss, which is used to regress the bounding box of each ROI; L mask represents the pixel segmentation mask generation loss, which is used to generate a mask for each ROI and each category.
[0071] In the specific implementation, experiments and result analysis are carried out on the multi-class pest recognition and detection method based on Mask-RCNN-CBAM fusion attention mechanism:
[0072] Experimental Environment
[0073] The programming language used in the experimental environment is Python 3.8.10, the deep learning framework is TensorFlow 2.4.0, the hardware environment configuration is Intel(R) Core(TM) i7-12700K * 20, the operating system is Windows 10, the graphics card is NVIDIA GeForce RTX 3080, the graphics card driver configuration is CUDA 11.6 and cuDNN 8.0.6, the initial learning rate is set to 0.001, and the learning momentum is set to 0.9. The mini-batch size is set to 128, the weight decay is set to 0.0005, and the total number of iterations is also set to 300.
[0074] Evaluation Metrics
[0075] In this embodiment, four accuracy evaluation metrics, namely Intersection over Union (IoU), Precision, Recall, and F1-score, are used to measure the performance of the model in extracting pests. IoU represents the degree of overlap between the predicted region and the ground truth region, and is usually used to represent the accuracy of the detected target location. Precision represents the ratio of the samples predicted as positive by the model to the true positive samples. Recall represents the ratio of the number of pests correctly detected by the model to the number of detected targets. F1 is the weighted average of Precision and Recall, as shown in the following formula:
[0076]
[0077] Among them, Intersection represents the area of overlap between the model's predicted region and the ground truth region; Union represents the area of the union of the model's predicted region and the ground truth region; TP represents the pixels correctly classified as pests; FP represents the positive samples misclassified.
[0078] Analysis of Experimental Results
[0079] To verify the accuracy and effectiveness of the Mask-RCNN-CBAM network in pest detection, three classic networks are introduced in this embodiment for comparative analysis: ResNet, Faster R-CNN, and Mask R-CNN. Mask R-CNN, as the best paper of ICCV2017, was jointly developed by Kaiming He and Ross Girshick of the FAIR team. It adds an instance segmentation function on the basis of Faster R-CNN and realizes object detection and pixel-level segmentation through a parallel branch network. ResNet solves the problem of gradient disappearance in deep networks by introducing residual connections, enabling information and gradients to propagate efficiently in deep networks, thus making the training of very deep neural networks possible and significantly improving the model performance. Faster R-CNN realizes end-to-end training by introducing a Region Proposal Network (RPN) and shared convolutional features, efficiently generating candidate regions, and providing high precision and efficiency in object detection. Mask-RCNN adds a fully connected segmentation network on the basis of Faster R-CNN for semantic segmentation and introduces the ROIAlign module to accurately align pixels and handle semantic segmentation problems. To verify the effectiveness of the proposed Mask-RCNN-CBAM network, a comparative analysis was carried out with these classic networks. To more intuitively display the performance of the improved model, this study input the test set data into the four networks and output the pest detection results of each image in an end-to-end manner. By comparing the performance metrics between different models, such as mAP and F1Score, this system can accurately detect and classify various pest targets on the surface of crops. The comparison of the recognition results of each model is as Figure 7 shown. Figure 7 The pest recognition results of three groups of images in the pest dataset are shown. The methods mentioned above generally perform well in pest recognition. However, in cases where the background is complex and the pest population is dense, there are differences in the recognition ability of the models. From the overall performance, the Mask-RCNN-CBAM model proposed in this embodiment achieved the best recognition effect. Figure 7 The yellow boxes in Figure 7In the second and third rows, the ResNet and Faster R-CNN models perform poorly in extracting from pests with complex structures or dense distributions. In the detection process, all three models - ResNet, Faster R-CNN, and Mask-RCNN - exhibit false negatives and false positives, with the false negative problem being particularly prominent in the Mask-RCNN network. In terms of mask smoothness and segmentation quality, although ResNet is superior to Faster R-CNN in some aspects, Mask-RCNN, especially the version combined with the FPN architecture and mask prediction branch, as well as the further optimized Mask-RCNN-CBAM network, demonstrates better performance.
[0080] The method proposed in this embodiment performs better in detecting dense pest areas and extracting pests. The Mask-RCNN-CBAM network effectively filters and significantly improves the recognition rate of pest features, being more sensitive to the background. In various pest detection tasks, the Mask-RCNN-CBAM network shows good recognition performance, with higher accuracy and smoother masks. This benefits from its integrated dual-channel attention mechanism, which assigns greater weights to pest features while weakening the influence of background features, effectively distinguishing pest and background features. In addition, the feature-enhanced FPN (Feature Pyramid Network) fuses the shallow and deep features of the image, enriching the detailed information and significantly improving the detection accuracy and running efficiency of the network.
[0081] To evaluate the effectiveness of the proposed method, a quantitative analysis of the experimental results was conducted. The ResNet, Faster R-CNN, Mask-RCNN, and Mask-RCNN-CBAM models were tested on the pest dataset, and Table 1 presents the results of the test set. As can be seen from Table 1, the proposed method is superior to the other three methods in terms of Precision, MIoU, and F1 score. Compared with ResNet, Precision, MIoU, and F1 are increased by 0.73%, 1.38%, and 0.11% respectively; compared with Faster R-CNN, these three metrics are increased by 2.65%, 5.14%, and 3.35% respectively; compared with Mask-RCNN, these three metrics are increased by 2.44%, 0.6%, and 1.21% respectively. For example, in the research on improving the accuracy of robot mixed disassembly codes, the Mask-RCNN algorithm has achieved a significant improvement in efficiency and accuracy compared with traditional algorithms. In terms of the recall rate metric, the method proposed in this embodiment is increased by 2.67% and 1.64% compared with Faster R-CNN and Mask-RCNN respectively, while the improvement compared with ResNet is not obvious, which may be due to the insufficient size of the dataset, resulting in less obvious learning effects.
[0082] Comparison Results of Evaluation Metrics for Three Classic Network Models in Table 1
[0083]
[0084] In addition, as shown in Table 1, the size of the proposed network parameters is 63.73 MB, which is 1.39 MB smaller than the Mask-RCNN network. This is due to the optimization of some parameters during the convolution process by the improved dual-channel attention module and the feature-enhanced FPN module. Compared with the original network, the proposed method not only has fewer parameters, but also is more accurate in the extraction and segmentation of pests, and has better detection performance, indicating that the proposed improved method achieves a better balance between segmentation accuracy and efficiency.
[0085] In a specific embodiment, the influence of the ablation of the attention mechanism module on the extraction result is evaluated. The specific steps include:
[0086] To evaluate the influence of the attention mechanism module on the pest extraction result, an ablation experiment was conducted. The basic network of the experiment is Mask-RCNN-CBAM, and the dataset used is the constructed pest dataset. The results of the ablation experiment for quantitatively verifying the influence of the attention mechanism module on the pest extraction effect are shown in Table 2.
[0087] Table 2 Comparison of Metrics before and after Incorporating the Attention Mechanism Module
[0088] Method Training Time (seconds / time) F1(%) AP(50) / (%) AR(small) / (%) With Module 25 58.8 76.3 35.0 Without Module 24 53.3 70.0 32.0
[0089] As shown in Table 2, after adding the attention mechanism module, the F1 score of the network increased by 5.5%, and the extraction accuracy of small targets increased by 3%. After adding the attention mechanism module, the network's ability to extract important information and assign weights to pest features has been significantly improved, thereby enhancing the network's sensitivity to pest features and improving the accuracy and efficiency of feature extraction. In addition, after adding the attention mechanism module, the running time of the network only increased slightly (1 second slower), which has little impact on the overall running efficiency of the network. The experimental results confirm that the attention mechanism can improve the accuracy of target detection, improve the feature extraction performance, help the model focus on important features, reduce the sensitivity to noise or irrelevant information, reduce overfitting, and accelerate the convergence speed.
[0090] In a specific embodiment, the ablation of the multi-scale feature enhancement module is evaluated for its impact on the extraction results. The specific steps include: The FPN network enhances the feature information and makes full use of multi-scale features. In this embodiment, the FPN network is improved by adding a top-down branch to better mine the detailed features of the image. To verify the effectiveness of the proposed feature enhancement pyramid module in agricultural pest detection, this study conducted a quantitative experiment, and the results are shown in Table 3. The experimental results show that this module significantly improves the model's ability to identify pest features. As shown in the table, after adding this module, the F1 score of the network increases by 0.4%, the AP value increases by 3.3%, and the recall rate for small target pests such as thrips and leafhoppers increases by 6%. However, during the experiment, it was found that although the addition of this module slightly prolongs the network running time, when dealing with a huge dataset, this change may have a certain impact on the running efficiency of the model. But the accuracy of the model is significantly improved. Through this ablation experiment, it can be observed that the multi-scale feature fusion pyramid module enables the Mask-RCNN network to increase the receptive field of the model, enrich the information in the feature layer, better understand the object instances in the image, and adapt to targets of different scales, thereby improving the accuracy of detection and segmentation.
[0091] Table 3 Comparison of indicators before and after integrating the multi-scale feature fusion pyramid module
[0092] Method Running Time (seconds) F1(%) AP(50) / (%) AR(small) / (%) With Module 50 55.2 74.1 41.3 Without Module 46 54.8 70.8 35.0
[0093] In a specific implementation manner, to verify the effectiveness of the proposed dual-channel downsampling module, this embodiment compares the performance of the max-pooling downsampling and the dual-channel downsampling module in the network. For example, in the field of text information processing, by combining max-pooling and average-pooling for hybrid pooling, the model's ability to extract text features can be improved. In sonar image target detection, compared with traditional methods, the dual-channel attention mechanism model can significantly improve the detection accuracy. The dual-channel downsampling module uses a 2×2 max-pooling operation, followed by a 1×1 convolutional compression module, and combines it with another 1×1 convolution, followed by a 3×3 convolutional kernel with a stride of 2×2. This method reduces the downsampling stride during the feature transmission process and reduces the number of channels by stacking convolutional layers. The experimental results are shown in Table 4. The dual-channel downsampling module improves the overall F1 score, AP value, and AR(small) of the network by 0.8%, 3.1%, and 0.8% respectively. The experiment shows that the dual-channel downsampling technique balances the class distribution by reducing the number of samples. Although it may cause information loss, it reduces the computational cost and improves the training speed. In addition, through appropriate methods, such as random downsampling or NearMiss, the loss of important information can be minimized to a certain extent, thereby retaining feature information to a certain extent.
[0094] Table 4 Comparison of indicators before and after integrating the dual-channel downsampling module
[0095] Method F1(%) AP(50) / (%) AR (small) / (%) With Module 54.5 73.2 35.1 Without Module 53.7 70.1 34.3
[0096] In the specific implementation, the analysis of model complexity and efficiency includes the following steps:
[0097] After introducing the attention mechanism module, the feature enhancement FPN module, and the improved dual-channel downsampling module, the number of model parameters did not increase. Instead, compared with the original model, the complexity decreased. The ablation experiment shows that adding the CBAM attention mechanism module enhances the network's ability to extract context information from remote sensing images. Adding the feature enhancement pyramid module further improves the fusion of deep and shallow feature information. The design of the dual-channel downsampling module effectively reduces feature loss during the transmission process. These improved modules effectively enhance the network's feature extraction and analysis capabilities when processing remote sensing image target detection tasks, and thus have a positive impact on the efficiency and accuracy of target recognition. In terms of the model operation efficiency, the introduction of the attention mechanism module and the feature enhancement FPN module has little impact on the model processing time. The implementation shows that the proposed Mask-RCNN-CBAM network performs quite well in the operation efficiency of the pest extraction task.
[0098] Aiming at the problems of false positives and missed detections in pest extraction caused by complex backgrounds and dense superposition of pests, this embodiment proposes a Mask-RCNN-CBAM pest extraction network integrating an attention mechanism. By integrating the CBAM attention mechanism into the feature processing process and enhancing multi-scale feature extraction and fusion through the feature pyramid module, the network's ability to extract context information is enhanced, and the loss of image detail features is reduced. The experimental results show that the Mask-RCNN-CBAM network has excellent target extraction effects on the pest dataset. Especially under complex backgrounds and dense pest conditions, it achieves high accuracy and good performance, and obtains the highest performance in terms of Precision, F1score, Recall, etc. At the same time, it has a low false positive rate and a low missed detection rate, indicating that the network model is reliable and applicable. Compared with other pest extraction methods, the Mask-RCNN-CBAM network can better extract pest feature information and optimize detail information.
[0099] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0100] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism, characterized in that, Including: Collecting image data of multiple types of pests to construct an experimental data set; Introducing a CBAM attention mechanism module in the feature processing stage of the Mask-RCNN model and connecting a multi-scale feature enhancement module to construct a Mask-RCNN-CBAM deep learning network; Training the Mask-RCNN-CBAM deep learning network with the experimental data set; Collecting real-time data and inputting it into the trained Mask-RCNN-CBAM deep learning network to obtain the recognition and detection results of multiple types of pests.
2. The multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 1, wherein, The CBAM attention mechanism module includes a channel attention module and a spatial attention module connected in sequence. After the input image is extracted by the backbone to obtain the backbone features, it is sent into the CBAM attention mechanism module. The input features are first processed by the channel attention module, and the generated channel attention features are aggregated with the original input features to form the final channel attention features. Subsequently, the final channel attention features are sent into the spatial feature map and fused through pooling and convolution operations to finally obtain the attention features.
3. A multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 2, characterized in that, The channel attention module aggregates the spatial information of the channel feature map through average pooling and max pooling operations to generate two different spatial context descriptors, representing the average pooling feature and the max pooling feature respectively; the generated features are transmitted to the shared network to generate a 1×1×C channel attention map Mc(F). The shared network consists of a multi-layer perceptron MLP and a hidden layer. The activation size of the hidden layer is set to C / r, where C is the number of neurons and r is the decay rate, and the activation function is ReLu. After applying the shared network to each descriptor, the sum of the combined feature elements is output as a feature vector.
4. The multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 3, characterized in that, The calculation formula for channel attention is: Where: F C max = MaxPool(F) is global max pooling, F C avg = AvgPool(F) is global average pooling, σ is the sigmoid function; W0, W1 are the shared weights of two input features.
5. The multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 2, characterized in that The spatial attention module generates a spatial attention feature map by utilizing the spatial relationship of the features, and aggregates the channel information of the feature map with the help of two pooling layers to obtain two two-dimensional feature sums, representing the average pooling feature and the max pooling feature respectively. Subsequently, these features are connected and convolved through a standard convolutional layer and finally input into the sigmoid function to generate the spatial attention feature map.
6. The multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 5, characterized in that The calculation formula for the spatial attention is: Ms(F) = σ(f 7×7 (AvgPool(F))); (MaxPool(F))) = σ(f 7×7 (F avg ; F max )); Among them, M s (F) is a spatial feature map, and σ is a sigmoid function; f 7×7 represents a convolution operation with a filter size of 7×7.
7. A multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 1, characterized in that The multi-scale feature enhancement module includes two parts: top-down downsampling feature extraction and bottom-up upsampling feature extraction. The downsampling feature extraction aims to extract deep image features containing semantic information to generate high-resolution features; while the upsampling feature extraction focuses on the spatial information of the shallow features, transmits them to the deeper layer for fusion, and finally synthesizes the features transmitted from the upper layer with the shallow spatial information and the deep semantic information. During the feature transmission process, a two-channel downsampling convolution operation is adopted to reduce the loss of detailed features during downsampling and retain the image detailed features.
8. A multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 7, characterized in that, The two-channel downsampling convolution operation includes: fusing two transition modules. The left branch adopts a 2×2 max pooling operation followed by a 1×1 convolutional layer, and the right branch adopts a 1×1 convolutional layer followed by a 3×3 convolutional layer with a stride of 2×2. The results of the two branches are superimposed and output.
9. A multi-class pest recognition and detection method based on the Mask-RCNN-CBAM fusion attention mechanism according to claim 1, characterized in that, The loss function of the Mask-RCNN-CBAM deep learning network includes: L = L cls + L box + L mask ; where L cls represents the classification loss, which is used to determine the category to which each ROI belongs; L box represents the bounding box offset loss, which is used to regress the bounding box of each ROI; L mask represents the pixel segmentation mask generation loss, which is used to generate a mask for each ROI and each category.
10. A multi-class pest recognition and detection system based on the Mask-RCNN-CBAM fusion attention mechanism, characterized in that, including: Acquisition module: used to acquire multi-class pest image data and construct an experimental data set; Model construction module: used to introduce the CBAM attention mechanism module in the feature processing stage of the Mask-RCNN model and connect the multi-scale feature enhancement module to construct the Mask-RCNN-CBAM deep learning network; Training module: used to train the Mask-RCNN-CBAM deep learning network through the experimental data set; Recognition and detection module: used to collect real-time data and input it into the trained Mask-RCNN-CBAM deep learning network to obtain multi-class pest recognition and detection results.
Citation Information
Patent Citations
Multi-scene ship detection and segmentation method based on mixed attention
CN115631427A
Mask-RCNN-based multi-target detection method in indoor complex environment
CN115937659A
Instance-level ship identification method, terminal equipment and storage medium
CN116486180A
Foggy day high-altitude power equipment corrosion detection method and system based on deep learning
CN117523423A
Rail traffic vehicle chassis real-time detection method, system and equipment and storage medium
CN117690104A
Cited By
Crop disease and pest identification method and system based on multi-task learning
CN120747652A
A crop disease and pest identification method and system based on multi-task learning
CN120747652B