A Maize Pest Identification Method Based on Improved YOLOv8

By introducing a global attention mechanism and deformable convolution module in the YOLOv8 model, the problem of different sizes and sizes in corn pest detection is solved, the detection accuracy and recall rate are significantly improved, and a more efficient corn pest recognition effect is achieved.

CN119229261BActive Publication Date: 2025-05-27GUANGDONG OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411687276.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-27
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of different sizes and sizes of corn pests due to different shooting heights in corn pest detection, resulting in low detection accuracy and recall rate.

Method used

The corn pest recognition method based on improved YOLOv8 is adopted. By adding a global attention mechanism (GAM) module to the neck network and replacing part of the C2f module as a deformable convolution (DCN) module in the backbone network, the feature extraction and fusion capabilities are enhanced and the model's sensitivity to small object detection is improved.

Benefits of technology

The improved GD-YOLOv8 model significantly improves detection accuracy, recall, mAP50 and mAP50-95 in corn pest detection, which are improved by 4.7%, 3.7%, 0.9% and 2.4%, respectively, achieving higher end-to-end pest detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229261B_ABST
    Figure CN119229261B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of maize pest detection, and discloses a maize pest recognition method based on improved YOLOv8, including obtaining maize pest images; constructing an improved YOLOv8 model; the improved YOLOv8 model includes a backbone network, a neck network, and a head network; using the backbone network to extract features from the maize pest images and fusing the feature information based on deformable convolutions to generate multi-scale feature maps containing different levels of visual information; using the neck network to fuse the multi-scale feature maps containing different levels of visual information to obtain a feature map integrating multi-scale information; using the head network to predict the maize pest recognition result according to the feature map integrating multi-scale information. The present invention can provide accurate pest detection and recognition for maize crops, and achieve high-precision end-to-end pest detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of corn pest detection, and specifically relates to a method for identifying corn pests based on improved YOLOv8. Background Art

[0002] In the prevention and control of crop pests and diseases, the primary and most important point is how to quickly identify the pests and diseases that damage crops in a short time, and to protect the growth of crops to the greatest extent.

[0003] Spodoptera frugiperda, an insect native to the tropical and subtropical regions of the Americas. It likes to be active at night. This insect has very strong adaptability, migratory ability and polyphagia, and has quickly become a major agricultural invasive pest across national and continental boundaries. The larval stage of Spodoptera frugiperda shows its harmfulness. They mainly feed on corn and rice. The Asian corn borer is a common pest. It belongs to the Pyralidae family of Lepidoptera. This species has strong destructive power to cereal crops, especially corn. The larvae of the Asian corn borer will invade the corn plants and feed on them, resulting in serious impacts on the growth of corn plants and even directly causing the death of young corn plants.

[0004] In order to accurately detect pests and formulate scientific prevention and control measures, it is necessary to accurately identify pests. However, the identification of Spodoptera frugiperda and Asian corn borer mainly relies on manual identification, and the experience and knowledge of humans are highly subjective. If ordinary people lack understanding of corn pests, they may miss the best prevention and control period when corn pests invade corn. Therefore, deep learning technology can be used to accurately identify corn pests.

[0005] In recent years, some researchers have used image processing and machine learning to detect crop pests and diseases. However, compared with traditional machine learning and deep learning, machine learning is difficult to meet the actual needs in complex scenarios, and the effect of detecting crop pests is poor and the accuracy is low. Deep learning has excellent performance in complex scenarios and can meet the actual needs. Among them, deep learning represented by convolutional neural networks has achieved great success in the fields of computer vision, automatic speech recognition, image classification, etc. Such as AlexNet, GoogleNet, VGGNet, ResNet, Transformer, which have been introduced from the field of natural language processing to the field of computer vision in recent years. Pest target detection is one of the main tasks of plant pest detection, and its purpose is to obtain accurate position and category information of pests. Deep learning technology shows excellent performance in image recognition, including the ability to extract occluded and rotated objects, and directly processes image information during the learning process without cumbersome processes.

[0006] Deep learning models can be divided into two-stage and single-stage according to the framework. The two-stage models are mainly represented by the RCNN series. FAROOQ ALI et al. improved Faster-RCNN, used MobileNet as its basic network, collected local crops to make a dataset, and conducted tests. It can solve problems such as blurring and light changes, as well as the problem of pest sample distortion, and the model has a certain generalization ability. Tiewei Wang et al. proposed an improved recognition and counting method MPest-RCNN based on deep learning and data recombination. Taking three typical pests that damage apples as the research object, an apple pest dataset was established, and the accuracy of apple pest detection can reach 99.11%. AMANI ABDULRAHMAN ALBRAIKAN et al. proposed a Mask RCNN recognition method for the red palm weevil pest that damages palm trees. Using MobileNetv2 as the backbone network, the pre-trained network weights of the ResNet50 architecture were used for transfer learning, and the data was augmented through data augmentation techniques. Detection was carried out on the red palm weevil pest dataset, and the accuracy was improved to 99.27%.

[0007] Single-stage networks achieve the entire process of generating candidate bounding boxes in an integrated manner. The YOLO series is a typical example. Lijuan Zhang et al. proposed a deep learning algorithm for YOLOv5, integrating the lightweight convolutional module C3M of the mobileNetV3 network into the YOLOv5 model and adding the GAM attention mechanism to the neck of the model, further improving the detection ability of the model. When detecting pests in the public dataset IP102, the accuracy was increased by 2.4% compared to the original model. Lianpeng et al. selected corn leaf spot as the research object, used the YOLOv5 model combined with MLFF (Multi-Feature Fusion Layer), added the CABM attention module and Detection Transformer (DETE), and constructed a pest and disease detection MCD-Yolov5 model to achieve dynamic adjustment of the weights of the feature layer features of the input image. In addition, a drone platform was established for on-site inspection. The accuracy of the model for pest and disease identification can reach 88.12%. Shuai et al. proposed a Maize-YOLO network based on YOLOv7. Through the CSPResNeXt-50 module and VoVGSCSP module, it has faster detection accuracy and detection speed than the original model, with a mAP of 76.3% on the public dataset IP102. Zhang et al. took corn borer, fall armyworm, and cotton bollworm as the research objects and proposed a method for identifying corn pests by combining the Adan optimizer to replace the optimizer of YOLOv7. The Adan optimizer can perceive the surrounding gradient information in advance and reduce the computing power of the model. A corn pest dataset was constructed through data augmentation and compared with traditional methods and other commonly used network models. After improvement, it was increased by 2.79% - 11.83% compared to the original YOLOv7, and the computing power far exceeds the performance of the original network.

[0008] Two-stage algorithms need to generate candidate regions and then identify the candidate regions. Single-stage algorithms are an integrated process, saving detection time. Summary of the Invention

[0009] In view of the above deficiencies in the prior art, the present invention provides a method for identifying corn pests based on improved YOLOv8 to solve the problem of inconsistent size scales of corn pests due to different shooting heights.

[0010] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0011] A method for identifying corn pests based on improved YOLOv8, comprising the following steps:

[0012] Obtain corn pest images;

[0013] Construct an improved YOLOv8 model; the improved YOLOv8 model includes a backbone network, a neck network, and a head network;

[0014] Use the backbone network to extract features from the maize pest images and fuse the feature information based on deformable convolution to generate a multi-scale feature map containing different levels of visual information;

[0015] Use the neck network to fuse the multi-scale feature maps containing different levels of visual information to obtain a feature map that fuses multi-scale information;

[0016] Use the head network to predict the maize pest recognition result based on the feature map that fuses multi-scale information.

[0017] Furthermore, the obtaining of the maize pest images further includes:

[0018] Perform image enhancement processing on the obtained maize pest images.

[0019] Furthermore, the performing of image enhancement processing on the obtained maize pest images includes:

[0020] Perform flipping and rotation operations on the obtained maize pest images.

[0021] Furthermore, the backbone network includes a first CBS module, a second CBS module, a first C2f module, a third CBS module, a second C2f module, a fourth CBS module, a third C2f module, a fifth CBS module, a deformable convolution module, and an SPPF module arranged in sequence from top to bottom.

[0022] Furthermore, the deformable convolution module first performs a convolution on the input feature map, then generates a compensation field through a convolutional layer for the output feature map, interpolates the compensation field to obtain an offset, and finally obtains the offset feature map based on the output feature map and the offset.

[0023] Furthermore, a first global attention mechanism module is also included between the fifth convolutional module and the deformable convolution module;

[0024] The first global attention mechanism module first obtains channel attention weights through a multi-layer perceptron, then multiplies the channel attention weights by the input feature map to obtain a channel attention feature map; then uses a convolutional layer to perform spatial information fusion on the channel attention feature map to obtain spatial attention weights, and finally multiplies the channel attention feature map by the spatial attention weights to obtain a global attention feature map.

[0025] Furthermore, the neck network includes:

[0026] The first upsampling module, the first splicing module, the fourth C2f module, the second upsampling module, the second splicing module, the fifth C2f module, the third upsampling module, the third splicing module, and the sixth C2f module are arranged from bottom to top, and the second global attention mechanism module, the sixth CBS module, the fourth splicing module, the seventh C2f module, the seventh CBS module, the fifth splicing module, the eighth C2f module, the eighth CBS module, the sixth splicing module, and the ninth C2f module are arranged from top to bottom;

[0027] The input end of the first upsampling module is connected to the output end of the SPPF module, the input end of the first splicing module is connected to the output end of the deformable convolution module, the input end of the second splicing module is connected to the output end of the second C2f module, the input end of the third splicing module is connected to the output end of the first C2f module, the input end of the second global attention mechanism module is connected to the output end of the sixth C2f module, the input end of the fourth splicing module is connected to the output end of the fifth C2f module, the input end of the fifth splicing module is connected to the output end of the fourth C2f module, and the input end of the sixth splicing module is connected to the output end of the SPPF module.

[0028] Furthermore, the head network includes:

[0029] The first detection head with a size of 20×20, the second detection head with a size of 40×40, the third detection head with a size of 80×80, and the fourth detection head with a size of 160×160;

[0030] The input end of the first detection head is connected to the output end of the ninth C2f module, the input end of the second detection head is connected to the output end of the eighth C2f module, the input end of the third detection head is connected to the output end of the seventh C2f module, and the input end of the fourth detection head is connected to the output end of the second global attention mechanism module.

[0031] Furthermore, the loss function of the improved YOLOv8 model is:

[0032] ;

[0033] Wherein, , respectively represent the squares of the distances between the upper left point and the lower right point between the predicted box and the ground truth box, and respectively represent the widths and heights of the predicted box and the ground truth box, represents the overlapping ratio of the predicted box and the ground truth box.

[0034] The present invention has the following beneficial effects:

[0035] In order to improve the detection accuracy of maize pests and solve the problem of different sizes of maize pests, this study proposes a maize pest detection method based on improved YOLOv8. First, a Global attention mechanism (GAM) module is added to the neck network part to enhance the feature map information, retain the mapping features, and improve the model's fusion ability and detection accuracy. Then, before the SPPF layer of the backbone network, some C2f modules are replaced with Deformable ConvNets (DCN) modules to fuse the information of the feature map with different convolution weights, and the detection head with a scale of 160×160 is combined to enhance the small target detection ability. The experimental results show that for maize pests, the improved GD-YOLOv8 model has improved the precision, recall rate, mAP50, and mAP50-95 by 4.7, 3.7, 0.9, and 2.4 percentage points compared with the YOLOv8 model. This method can provide accurate pest detection and identification for maize crops and achieve high-precision end-to-end pest detection. Description of the Drawings

[0036] Figure 1 It is a schematic flow chart of a maize pest identification method based on improved YOLOv8;

[0037] Figure 2 It is a schematic diagram of the improved YOLOv8 model structure;

[0038] Figure 3 It is a schematic diagram of the head network structure;

[0039] Figure 4 It is a schematic diagram of precision;

[0040] Figure 5 It is a schematic diagram of Recall precision. Detailed Embodiments

[0041] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0042] As Figure 1 shown, a maize pest identification method based on improved YOLOv8 provided by an embodiment of the present invention includes the following steps S1 to S5:

[0043] S1. Obtain maize pest images;

[0044] S2. Construct an improved YOLOv8 model; the improved YOLOv8 model includes a backbone network, a neck network, and a head network;

[0045] S3. Use the backbone network to extract features from the maize pest images and fuse the feature information based on deformable convolution to generate multi-scale feature maps containing different levels of visual information;

[0046] S4. Use the neck network to fuse the multi-scale feature maps containing different levels of visual information to obtain a feature map that fuses multi-scale information;

[0047] S5. Use the head network to predict the maize pest recognition result based on the feature map that fuses multi-scale information.

[0048] In an optional embodiment of the present invention, due to various factors such as ecological and climate changes that can affect the growth of pests, it has become quite challenging to collect agricultural pest data. In this embodiment, the pest dataset publicly available on the Kaggle website was sorted out, and a total of 278 original pest images were obtained. These images are stored in JPG format with different sizes, high clarity, and rich colors, fully showing the morphological characteristics of the pests. Among them, the collected images of Spodoptera frugiperda and Ostrinia furnacalis pests cover different life cycles. Some of these pests are similar to the background, and some of the images have complex backgrounds, which increases the difficulty of identifying maize pests to a certain extent.

[0049] Therefore, after obtaining the maize pest images, this embodiment further includes:

[0050] Performing image enhancement processing on the obtained maize pest images.

[0051] Furthermore, the performing image enhancement processing on the obtained maize pest images includes:

[0052] Performing flipping and rotation operations on the obtained maize pest images.

[0053] This embodiment uses data augmentation technology, which can effectively expand the dataset to solve the network fitting problem and the data sample imbalance problem. In this embodiment, the original dataset is expanded through data augmentation methods such as flipping and rotation to obtain 668 pictures. According to the images, they are divided into a training set, a test set, and a validation set according to 8:1:1.

[0054] In an optional embodiment of the present invention, YOLOv8 is a new network framework based on the YOLO series of algorithms. On the basis of the previous YOLO versions, it integrates a number of new technologies to improve the detection speed and accuracy of the model. YOLOv8 is designed based on the concept of being fast, accurate, achieving higher performance, and being easy to use, making it an excellent choice for a wide range of object detection, image segmentation, and image classification tasks.

[0055] YOLOv8 mainly consists of three parts: the backbone network, the neck network, and the head network. The backbone network part is mainly composed of multiple C2f modules, which are optimized based on the CSPDarknet network structure design. The CSPDarknet network structure is characterized by dividing the entire network into two main parts, each consisting of multiple residual blocks. This design not only ensures that the network can fully learn and extract deep features of the input data, but also effectively alleviates the problem of gradient disappearance through residual connections, making the model more stable during training.

[0056] The neck network part adopts the PAN-FPN structure. Compared with the Feature Pyramid Network (FPN), PAN-FPN adopts a bidirectional path architecture. This means that feature information can flow not only from high-level to low-level, but also from low-level to high-level. This bidirectional flow design enhances the context information flow of features, enabling more effective fusion of features between different levels. In YOLOv8, the network structure has been further optimized. By removing the upsampling convolution structure, the transmission of low-level feature information to the high-level of the network becomes more direct and efficient. This simplification not only reduces the complexity of the network, but also promotes the effective fusion of feature information between different levels.

[0057] The head network part adopts an Anchor-Free detection method, which can directly predict the center point and width-height ratio of the target, rather than predicting the position and size of the Anchor box. In this way, the model can reduce the number of Anchor boxes to improve the detection speed and accuracy of the model.

[0058] In an alternative embodiment of the present invention, as Figure 2 shown, the backbone network includes a first CBS module, a second CBS module, a first C2f module, a third CBS module, a second C2f module, a fourth CBS module, a third C2f module, a fifth CBS module, a deformable convolution module, and an SPPF module arranged in sequence from top to bottom.

[0059] The deformable convolution module first performs a convolution on the input feature map, then generates a compensation field for the output feature map through a convolutional layer, interpolates the compensation field to obtain an offset, and finally obtains the offset feature map based on the output feature map and the offset.

[0060] In traditional convolutional neural network architectures, convolutional units, also known as convolutional kernels, are responsible for performing sampling operations at fixed positions on the input feature map. They extract features from local regions and integrate these features into a new feature map through convolutional operations. After the convolutional layer, a pooling layer usually follows. Its main role is to gradually reduce the dimension of the feature map, thereby reducing the number of model parameters and preventing overfitting to a certain extent. The pooling layer reduces the size of the feature map by selecting the maximum, minimum, or average value of a specific region, etc.

[0061] The pooling layer is another important component in CNN. The Region of Interest (RoI) pooling layer is a special type of pooling layer that generates RoIs (Regions of Interest) with spatial position constraints. The RoI pooling layer can map RoIs of different sizes to a feature map of a fixed size, facilitating subsequent operations such as classification or regression. However, despite the significant success of these network layers in image processing, they still have some limitations.

[0062] A major limitation is that due to the fixed nature of the convolutional kernel weights, the receptive field size remains consistent when the network processes different positions in the image. This characteristic means that the network's receptive field cannot be adaptively adjusted according to the size changes of objects in the image, thus limiting the flexibility of the network in encoding position information. In practical applications, different regions of an image may correspond to objects of different scales or that have undergone different deformations, which requires the model to be able to automatically adjust its scale to adapt to these changes. Therefore, in order to more effectively handle geometric transformations in images, a mechanism that can adaptively adjust the scale is needed to better handle this situation.

[0063] Deformable ConvNets (DCN) do not change the way of convolution calculation, but add a learnable parameter to the scope of the convolution operation . Similarly for each output , it is necessary to upsample 9 positions from . These 9 positions are obtained by spreading from the central position ( ) to the surrounding areas. However, there is an additional that allows the sampling points to spread into a non-grid shape, expressed as:

[0064] ;

[0065] Among them, represents the pixel value at position on the output feature map, represents the weight, represents the input feature map Indicates the position of the central sampling point; Indicates the position of the sampling point; Indicates the index number of each block; Indicates the weight; Indicates the position of the sampling point Offset relative to the central position; Indicates the mapping relationship between the original image and the feature map obtained by convolution.

[0066] In this embodiment, by introducing a deformable convolutional network, some C2F modules of the YOLOv8 network model are replaced with deformable convolutional DCNs to extract features of different sizes in the feature map scale, so as to realize feature extraction for pests of different scales.

[0067] In an alternative embodiment of the present invention, a first global attention mechanism module is further included between the fifth convolutional module and the deformable convolutional module.

[0068] The first global attention mechanism module first obtains channel attention weights through a multi-layer perceptron, and then multiplies the channel attention weights by the input feature map to obtain a channel attention feature map; then uses a convolutional layer to perform spatial information fusion on the channel attention feature map to obtain spatial attention weights, and finally multiplies the channel attention feature map by the spatial attention weights to obtain a global attention feature map.

[0069] The global attention mechanism module further includes a global average pooling layer and a global max pooling layer;

[0070] The global average pooling layer performs global average pooling operation on the input feature map, extracts the global features of each channel, and then outputs the global features of each channel to the multi-layer perceptron;

[0071] The global max pooling layer performs global max pooling operation on the input feature map, extracts the global features of each channel, and then outputs the global features of each channel to the multi-layer perceptron.

[0072] The global attention mechanism module further includes a convolutional layer;

[0073] The convolutional layer performs a convolutional operation on the input feature map and then outputs the output feature map to the global average pooling layer and the global max pooling layer respectively.

[0074] The global attention mechanism module further includes a pooling layer;

[0075] The pooling layer performs feature fusion on the global attention feature map to obtain the final output feature map.

[0076] The purpose of integrating the Global Attention Mechanism (GAM) into the YOLOv8 network in this embodiment is to improve the efficiency and quality of image feature extraction. The goal of the global attention mechanism is to enhance the overall performance of deep neural networks by optimizing information flow and strengthening global interactions between features. Based on CBAM, GAM makes sequential adjustments and sub-module reconstructions to the channel-spatial attention mechanism. Especially in the spatial attention sub-module, a double-layer convolutional structure is adopted. Through multi-level information fusion, this structure enables more effective integration of spatial information, thereby increasing the attention to spatial features. To avoid information loss that may be caused by the max pooling operation, the pooling operation is removed in this module. This is because while reducing the feature dimension, the max pooling operation may also lead to the loss of some important information. By omitting this step, feature maps can be better retained, thus improving the accuracy of feature extraction. However, the spatial attention module may lead to an increase in the number of parameters. To effectively control the number of parameters, the module adopts grouped convolution with channel shuffle. This technique can not only reduce the number of parameters to a certain extent but also improve the generalization ability of the model, thereby further enhancing the performance of the model.

[0077] The Global Attention Mechanism (GAM) is an improvement based on the Convolutional Block Attention Module (CBAM). CBAM applies attention mechanisms separately in the channel and spatial dimensions, while GAM further introduces a global attention mechanism, comprehensively considering the attention weights in the three dimensions of channels, spatial width, and spatial height, significantly enhancing the performance of the model in object detection tasks.

[0078] GAM adopts a unique approach to channel attention. First, it uses global average pooling and global max pooling operations to extract the global features of each channel respectively. These two global features contain information from all positions within the channel and can fully reflect the importance of the channel.

[0079] Next, GAM feeds these two global features into a shared multi-layer perceptron (MLP). Through multi-layer non-linear transformations, the MLP converts and reduces the dimension of the global features, further extracting the dependencies and high-level features between channels. The output of the MLP is used as the basis for calculating the attention weights of each channel, and these weights can reflect the importance of different channels in object detection tasks.

[0080] After calculating the attention weights of each channel, GAM multiplies these weights with the original feature map to achieve channel attention weighting. In this way, the model will pay more attention to the channels with higher weights in the subsequent feature extraction and object detection processes, thereby enhancing the performance of the model.

[0081] In this embodiment, the global attention mechanism GAM is used to fuse deformable convolution and GAM attention with the YOLOv8 network, further enhancing the ability of the improved YOLOv8 model to extract feature image information.

[0082] In this embodiment, the GAM global attention mechanism is introduced in the Backbone part, and some C2f modules in the YOLOv8 network are replaced with C2f-DCN. First, the data-augmented images are input and convolved, mainly to expand the channels, reduce the feature map, and extract deep features. Then the image enters the C2f module to obtain richer gradient information. Before entering the SPPF layer, we add the GAM attention mechanism to the network. The image is convolved, and the convolutional layer uses a convolutional kernel with a stride of 2 for downsampling operation to obtain a feature map of 1024*20*20. Then, max pooling and average pooling are performed on the input feature map, and after being processed by the MLP respectively, the max pooling and average pooling of the feature map are superimposed, and then convolved and activated by Sigmoid. Then, the feature map is input. First, ordinary convolution is performed on the input image, and then a convolutional layer is added to the feature map to obtain the offset of the deformable convolution. The offset uses the difference algorithm and is learned through backpropagation to more easily focus on the pest features. Finally, the feature map is input to the pooling layer for feature fusion to achieve a richer gradient combination.

[0083] In an alternative embodiment of the present invention, the neck network includes:

[0084] A first upsampling module, a first splicing module, a fourth C2f module, a second upsampling module, a second splicing module, a fifth C2f module, a third upsampling module, a third splicing module, and a sixth C2f module arranged in sequence from bottom to top, and a second global attention mechanism module, a sixth CBS module, a fourth splicing module, a seventh C2f module, a seventh CBS module, a fifth splicing module, an eighth C2f module, an eighth CBS module, a sixth splicing module, and a ninth C2f module arranged in sequence from top to bottom;

[0085] The input end of the first upsampling module is connected to the output end of the SPPF module, the input end of the first splicing module is connected to the output end of the deformable convolution module, the input end of the second splicing module is connected to the output end of the second C2f module, the input end of the third splicing module is connected to the output end of the first C2f module, the input end of the second global attention mechanism module is connected to the output end of the sixth C2f module, the input end of the fourth splicing module is connected to the output end of the fifth C2f module, the input end of the fifth splicing module is connected to the output end of the fourth C2f module, and the input end of the sixth splicing module is connected to the output end of the SPPF module.

[0086] The head network includes:

[0087] A first detection head with a size of 20×20, a second detection head with a size of 40×40, a third detection head with a size of 80×80, and a fourth detection head with a size of 160×160;

[0088] The input end of the first detection head is connected to the output end of the ninth C2f module, the input end of the second detection head is connected to the output end of the eighth C2f module, the input end of the third detection head is connected to the output end of the seventh C2f module, and the input end of the fourth detection head is connected to the output end of the second global attention mechanism module.

[0089] In this embodiment, on the basis of YOLOv8n, the C2f-DCN convolutional network structure is used to replace the C2f module in YOLOv8n, the GAM attention mechanism is introduced, and a small target detection head with a scale of 160×160 is additionally added. Using MPDIoU as the loss function, the problem that the YOLOv8 model has inaccurate recognition of complex backgrounds, different sizes of corn pests, and crop occlusion of pests is solved.

[0090] In the self-built dataset, due to different shooting angles, there will be a problem of inconsistent sizes of the pests captured in the pictures. The large target pests are more obvious in morphology, size, and color characteristics, while the characteristics of small target pests may become blurred and difficult to identify. The detection layer of the original YOLOv8 outputs three different sizes of feature maps, namely 20×20, 40×40, and 80×80. When the pest target is too small, it cannot be detected. This embodiment expands on the YOLOv8n model architecture. Specifically, a small target detection branch with a resolution of 160×160 is added to the model to enhance the detection sensitivity to small target pests. GD-YOLOv8 extracts feature information at the sixth layer of the backbone network and initially captures key feature information. The feature fusion strategy is realized by using the Concat operation, and the features extracted in the shallow stage of the neck network are fused with the GAM attention mechanism. By obtaining more small target feature information, the detection ability of the GD-YOLOv8 model for small target pests is greatly improved, and the misdetection and missed detection of pests at different scales are effectively reduced. The improved detection layer is as Figure 3 shown.

[0091] During the shooting process of corn pests, the corn stalks and leaves will cause a certain degree of occlusion to the corn pests. To solve this problem, consider the overlapping or non-overlapping regions, the distance between the center points, and the deviation of the width and height as the loss function, and calculate the IoU by minimizing the point distance between the predicted bounding box and the ground truth bounding box, thus simplifying the calculation process. The calculation formula is as follows:

[0092] ;

[0093] Among them, the bounding box and the ground truth box represent the predicted bounding box and the ground truth bounding box, ( ), ( ), ( ), ( ) are the coordinates of the upper left point and the lower right point of the bounding box and the ground truth box respectively. , respectively represent the squares of the distances between the upper left points and the lower right points between the predicted box and the ground truth box, and respectively represent the widths and heights of the predicted box and the ground truth box, represents the overlapping ratio of the predicted box and the ground truth box.

[0094] Through the improvement of the above loss function in this embodiment, GD-YOLOv8 is more sensitive to occluded small targets in corn pest detection compared with the YOLOv8 model.

[0095] In this embodiment, deformable convolution and GAM attention are fused with the YOLOv8 network, and the precision, recall rate, mAP50, and mAP50-95 are improved by 4%, 2.4%, 0.8%, and 1.8% respectively compared with the original model.

[0096] After the model is constructed in this embodiment, Precision, Recall, and mAP are used as the main indicators for model verification. For the object detection task, Precision is the best indicator to measure the detection accuracy of the model, mAP is calculated based on the IOU between the predicted box and the ground truth box to calculate the accuracy of the model, and Recall represents the proportion of truly detected corn pests. The higher the proportion, the more corn pests the model detects. The calculation formulas of the verification indicators are as follows:

[0097] ;

[0098] ;

[0099] ;

[0100] Among them, TP (True Positive) represents the number of positive classes predicted as positive classes, that is, the number of correctly predicted corn pest samples; FP (False Positives) represents the number of negative classes predicted as positive classes, that is, the number of samples that identify non-pests as pests; FN (False Negative) represents the number of positive classes predicted as negative classes, that is, the number of incorrectly predicted corn pest samples; represents the average accuracy of the i-th target, represents the number of target categories.

[0101] The present invention will now analyze and illustrate the performance of the method of the present invention in conjunction with specific experiments.

[0102] This experiment was conducted on servers equipped with RTX 2080 Ti (11GB) and Intel(R) Xeon(R) Platinum 8255C CPU. These GPUs were deployed based on Python 3.8 under the Linux Ubuntu 20.04 LTS operating system, and the model was built under the Pytorch 2.0.0 deep learning framework.

[0103] During the training of the corn crop pest detection model, the YOLOv8n weight parameters provided by the ultralytics framework were used as model learning initialization parameters and hyperparameter tuning to achieve the best detection performance of the entire network.

[0104] In order to distinguish the characteristics of the GAM module and the DCN module, this experiment conducts a comparative experiment, adding the GAM and DCN modules to the YOLOv8 network respectively, and analyzes the impact of each module on YOLOv8, such as Figure 4 , Figure 5 As shown in Table 1 and Table 2.

[0105] Table 1. Comparative test analysis table

[0106]

[0107] Table 2 Ablation experiment analysis table

[0108]

[0109] It can be seen that after adding the DCN module, Precision increased by 7 percentage points. The deformable convolution assigns weights to feature images with complex backgrounds, and the effect of improving precision is more significant. This shows that the deformable convolution DCN performs well in balancing model performance and precision. By adding only the GAM module, the recall rate of the model is 6.3 percentage points higher than that of the YOLOv8 model, making it pay more attention to the global information of the features and strengthening the correct identification of pests in pest samples. By integrating the two into the YOLOv8 model, the improved GD-YOLOv8 model has improved the precision, recall rate, mAP50, and mAP50-95 by 4.7, 3.7, 0.9, and 2.4 percentage points respectively compared with the YOLOv8 model, effectively improving the performance of the model and indicating the effectiveness of the proposed improved algorithm.

[0110] In complex backgrounds, the YOLOv8 model may fail to detect small targets. The improved YOLOv8 model has obvious advantages. By adding a small target detection head and using the MPDIoU loss function, it can be more sensitive to small targets.

[0111] This invention aims at the rapid detection of corn pests in complex natural field environments. Based on YOLOv8n, it replaces the C2f module in the model with a deformable convolutional network structure and introduces a global attention mechanism to enhance the model's feature fusion ability, enabling the model to better learn the position and size of pest features. Finally, a corn pest detection model with relatively high detection accuracy is developed.

[0112] The model established in this invention has only achieved good results in the recognition tasks of two common corn pests. However, there are numerous types of corn pests. The next research work will enrich the pest sample information for further recognition and real-time detection research. At the same time, the model will be transplanted and verified on edge mobile platforms and improved to make the model lighter and easier to deploy.

[0113] This invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0116] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0117] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed by the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A corn pest identification method based on improved YOLOv8, characterized in that: The following steps are involved: Obtain corn pest images; Constructing an improved YOLOv8 model; the improved YOLOv8 model includes a backbone network, a neck network and a head network; The backbone network includes a first CBS module, a second CBS module, a first C2f module, a third CBS module, a second C2f module, a fourth CBS module, a third C2f module, a fifth CBS module, a deformable convolution module and an SPPF module, which are arranged in sequence from top to bottom; the deformable convolution module first performs a convolution on the input feature map, then generates a compensation field for the output feature map through a convolution layer, and interpolates the compensation field to obtain an offset, and finally obtains the offset feature map according to the output feature map and the offset; the fifth CBS module and the deformable convolution module also include a first global attention mechanism module; the first global attention mechanism module first obtains the channel attention weight through a multi-layer perceptron, and then multiplies the channel attention weight with the input feature map to obtain a channel attention feature map; then uses the convolution layer to perform spatial information fusion on the channel attention feature map to obtain the spatial attention weight, and finally multiplies the channel attention feature map with the spatial attention weight to obtain the global attention feature map; The neck network includes a first upsampling module, a first splicing module, a fourth C2f module, a second upsampling module, a second splicing module, a fifth C2f module, a third upsampling module, a third splicing module and a sixth C2f module, and a second global attention mechanism module, a sixth CBS module, a fourth splicing module, a seventh C2f module, a seventh CBS module, a fifth splicing module, an eighth C2f module, an eighth CBS module, a sixth splicing module and a ninth C2f module, which are sequentially arranged from bottom to bottom; the input end of the first upsampling module is connected to the input end of the SPPF module. The output ends of the first splicing module are connected, the input end of the first splicing module is connected to the output end of the deformable convolution module, the input end of the second splicing module is connected to the output end of the second C2f module, the input end of the third splicing module is connected to the output end of the first C2f module, the input end of the second global attention mechanism module is connected to the output end of the sixth C2f module, the input end of the fourth splicing module is connected to the output end of the fifth C2f module, the input end of the fifth splicing module is connected to the output end of the fourth C2f module, and the input end of the sixth splicing module is connected to the output end of the SPPF module; The head network includes a first detection head with a size of 20×20, a second detection head with a size of 40×40, a third detection head with a size of 80×80, and a fourth detection head with a size of 160×160; the input end of the first detection head is connected to the output end of the ninth C2f module, the input end of the second detection head is connected to the output end of the eighth C2f module, the input end of the third detection head is connected to the output end of the seventh C2f module, and the input end of the fourth detection head is connected to the output end of the second global attention mechanism module; the backbone network is used to extract features of corn pest images and the feature information is fused based on deformable convolution to generate a multi-scale feature map containing visual information at different levels; The neck network is used to fuse multi-scale feature maps containing visual information at different levels to obtain a feature map that fuses multi-scale information; The head network is used to predict corn pest recognition results based on feature maps that fuse multi-scale information.

2. A corn pest identification method based on improved YOLOv8 according to claim 1, characterized in that: The obtaining of corn pest images also includes: The acquired corn pest images are subjected to image enhancement processing.

3. A corn pest identification method based on improved YOLOv8 according to claim 2, characterized in that: The image enhancement processing of the acquired corn pest image comprises: The acquired corn pest image is flipped and rotated.

4. A corn pest identification method based on improved YOLOv8 according to claim 1, characterized in that: The loss function of the improved YOLOv8 model is: in, Respectively represent the square of the distance between the upper left point and the lower right point between the predicted box and the real box, ω and h represent the width and height of the predicted box and the real box, respectively. Represents the overlap ratio between the predicted box and the true box.

Citation Information

Patent Citations

  • Method for improving YOLOv8 network and application of method in strip steel surface defect detection

    CN118365599A

  • YOLOv8n-based lightweight insect small target detection method and device

    CN118470402A