Multi-target pest detection method and system in unstructured environment based on improved YOLOv5s
Patent Information
- Application Number
- CN202411227881.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-09-03
AI Technical Summary
但少数的害虫检测则因害虫目标在图像中尺寸相对较小、自然环境复杂、新兴技术使用门槛过高等问题效果不佳,难以投入实际农业生产的工程应用中
[0019]1、去除YOLOv5s网络模型中的四个检测头中大检测头,通过采用三层探测头减少模型参数量,提高检测速度,使改进后的模型更适用于实际农业场景中的小目标害虫检测任务。
Smart Images

Figure CN118968312B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a method and system for detecting multi-target pests in an unstructured environment based on the improved YOLOv5s. Background Technology
[0002] Wheat, rice, and sugarcane, as major crops grown in many of my country's main agricultural regions, have always faced significant challenges in pest control, a problem that requires close attention and resolution in agricultural production. In the past, pest identification and observation for these crops was largely conducted manually. However, large-scale pest identification is hampered by the sheer number of pests and the vast planting areas. Firstly, manual identification and observation require prior skills training for technicians. Secondly, relying on agricultural technicians to conduct on-site observations and labor-intensive pest counting to predict the severity of pests is time-consuming, labor-intensive, subjective, prone to errors, and prone to prediction delays, making it difficult to implement effectively. Delayed pest control information leads to large-scale, indiscriminate application of pesticides, and the overuse of pesticides raises environmental concerns. These situations underscore the necessity of innovative and environmentally friendly technological solutions that enable real-time and accurate pest detection. In recent years, with the introduction of new agricultural IoT devices and the rapid development of machine vision and deep learning in agriculture, the time-consuming and labor-intensive problems of manual methods have been effectively alleviated. However, some pest detection methods are ineffective due to issues such as the relatively small size of pest targets in images, complex natural environments, and high barriers to entry for emerging technologies, making them difficult to apply in actual agricultural production.
[0003] With the continuous advancement of modern machine vision technology, deep learning, as an emerging approach, has demonstrated outstanding performance and remarkable results in numerous tasks such as image recognition, target detection, and instance segmentation for pests. Simultaneously, driven by the increasing computing power of mobile devices, deep learning technology has been effectively integrated with hardware devices, potentially further expanding the application scenarios and lowering the barrier to entry for pest monitoring, thus providing more possibilities for pest detection, counting, and engineering applications under machine vision. Therefore, there is an urgent need to provide an effective and feasible method and application for pest detection.
[0004] The above-mentioned problems urgently need to be solved. Summary of the Invention
[0005] To address existing problems, this invention designs a multi-target pest detection method and system based on an improved YOLOv5s unstructured environment. The aim is to provide an effective and feasible pest detection method and application.
[0006] To overcome at least one of the aforementioned drawbacks of existing technologies, this invention provides, in one aspect, a method for multi-target pest detection in unstructured environments based on an improved YOLOv5s architecture. The method includes: S1: acquiring images of crops infested by sugarcane aphids and preprocessing the images to create a first target dataset; S2: randomly selecting and recombining the first target dataset with the optimized and labeled AgriPest dataset to create a second target dataset; S3: dividing the second target dataset into a training set, a validation set, and a test set; S4: constructing an improved YOLOv5s network model as a multi-target pest detection network model, using the YOLOv5s network model as a baseline model, including: removing four... The detection head, a large detection head, reconstructs the backbone network using a GhostNet module for feature extraction. A BiFPN module is introduced into the FPN part for feature fusion. The ordinary convolutional module in the neck network is replaced with a DWConv module. The output features are fed into a detector with an added CBAM module for feature point judgment. S5: The training set is input into the multi-target pest detection network model for model training. S6: The parameters of the multi-target pest detection network model are improved based on the training results to obtain the trained multi-target pest detection network model. S7: The validation set is input into the trained multi-target pest detection network model for model evaluation. S8: The test set is input into the trained multi-target pest detection network model to obtain the multi-target pest detection results.
[0007] Furthermore, in step S4, the reconstruction of the backbone network for feature extraction using the GhostNet module includes: S41: replacing the C3 module in the original YOLOv5s network model with a Ghost Bottleneck module with a stride of 1, and replacing the ordinary convolutional Conv in the original YOLOv5s network model with a Ghost Bottleneck module with a stride of 2, thus forming a feature extraction network with the GhostNet module as the backbone.
[0008] Furthermore, the convolution process in the Ghost Bottleneck module includes: S411: obtaining an intrinsic feature map through regular convolution, used to extract basic features of the input data; S412: applying a linear transformation to the feature map of each channel to generate a Ghost feature map, including performing depthwise convolution on each channel using depthwise separable convolution to generate a Ghost feature map; S413: concatenating the intrinsic feature map and the Ghost feature map to obtain the final output feature map.
[0009] Furthermore, the introduction of the BiFPN module into the FPN part for feature fusion in step S4 includes: S42: deleting the feature pyramid network to reduce its impact on the overall network and simplify the bidirectional network; S43: adding connection edges between the original input nodes and output nodes of the same feature layer to reduce the increase in computational cost.
[0010] Furthermore, in step S4, replacing the ordinary convolutional modules in the neck network with DWConv modules includes: S44: The two Ghost Bottleneck modules are connected by a depth-separable convolution DWConv with a stride of 2, and DWConv replaces all ordinary convolutional blocks Conv in the neck convolution BiFPN module.
[0011] Furthermore, in step S4, the output features are fed into a detector with an added CBAM module for feature point determination, including: S45: Integrating the CBAM module into the downsampling step of the neck network, and the output of the C3 module is transmitted to the detection head through the CBAM module.
[0012] Furthermore, the CBAM module includes two independent convolutional steps that sequentially infer the input feature map, namely a channel attention module and a spatial attention module. The workflow includes: S451: Given an intermediate feature map F∈R C×H×W As input, the input image F is subjected to global max pooling and global average pooling by channel in the channel attention module, resulting in two one-dimensional vectors; S452: The vectors are processed through a fully connected layer, performing element-wise summation and sigmoid activation to generate a one-dimensional channel attention mechanism feature map M. c ∈R C×1×1 S453: As shown in the formula M c Element-wise multiplication with the input feature map F yields the channel attention-adjusted feature map F. ′ S454: Perform global max pooling and average pooling spatially; S455: Concatenate the two 2D vectors generated by pooling and then perform convolution; S456: The dimensionality-reduced feature map is activated by a Sigmoid operation to generate a spatial attention feature map M. S ∈R 1×H×W S457: Spatial attention feature map M S Multiplying the input feature map F' by the input feature map yields the final generated feature. In the formula, This represents element-wise multiplication, where M represents the multiplication operation. c This indicates that attention extraction is performed along the channel dimension, M. c (F) represents the weight coefficients obtained using the channel attention mechanism, M sThis indicates that attention extraction is performed in the spatial dimension, M S (F ′ F represents the weight coefficients obtained using the spatial attention mechanism. ′ F' represents the output feature map after channel attention adjustment of the input feature map F, and F''' represents the final accurate output.
[0013] Furthermore, the removal of the large detection head from the four detection heads of the YOLOv5s network in step S4 includes: S46: In the neck network, a new P2 small-scale prediction head is added, including a 160×160×128 upsampling step, and the large target prediction head is removed, including a 20×20×1024 feature layer output.
[0014] Furthermore, step S7, which involves inputting the validation set into the trained multi-target pest detection network model for model evaluation, includes: inputting the validation set into the trained multi-target pest detection network model and evaluating it based on the precision P, recall R, and mean precision mAP of the multi-target pest detection network model. In the formula, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, and FP is the number of incorrectly predicted positive samples.
[0015] Secondly, this invention also provides a multi-target pest detection system based on an improved YOLOv5s unstructured environment. The system includes: a first target dataset generation unit, suitable for collecting images of crops infested by sugarcane aphids and preprocessing the images to create a first target dataset; a second target dataset generation unit, suitable for randomly selecting and recombining the first target dataset with the optimized and labeled AgriPest dataset to create a second target dataset; a dataset partitioning unit, suitable for partitioning the second target dataset into a training set, a validation set, and a test set; and a model improvement unit, suitable for constructing an improved YOLOv5s network model as a multi-target pest detection network model based on the YOLOv5s network model, including: removing four... The detection head consists of a large detection head, which reconstructs the backbone network using a GhostNet module for feature extraction, introduces a BiFPN module into the FPN part for feature fusion, replaces the ordinary convolutional module in the neck network with a DWConv module, and sends the output features to a detector with an added CBAM module for feature point judgment; a model training unit, which is suitable for inputting the training set into the multi-target pest detection network model for model training; a model improvement unit, which is suitable for improving the parameters of the multi-target pest detection network model based on the training results to obtain the trained multi-target pest detection network model; a model evaluation unit, which is suitable for inputting the validation set into the multi-target pest detection network model for model evaluation; and a result generation unit, which is suitable for inputting the test set into the multi-target pest detection network model to obtain the multi-target pest detection results.
[0016] In another aspect, the present invention also provides a computer-readable storage medium storing one or more instructions for causing the computer to execute the above-described multi-target pest detection method based on the improved YOLOv5s unstructured environment.
[0017] In another aspect, the present invention provides an electronic device, comprising: a memory and a processor; the memory storing at least one program instruction; the processor loading and executing the at least one program instruction to implement the above-mentioned multi-target pest detection method based on the improved YOLOv5s unstructured environment.
[0018] The beneficial effects of this invention are:
[0019] 1. The large detector head among the four detector heads in the YOLOv5s network model was removed. By adopting a three-layer detector head, the number of model parameters was reduced, the detection speed was improved, and the improved model was made more suitable for small target pest detection tasks in actual agricultural scenarios.
[0020] 2. The GhostNet module is used to replace the original backbone module for feature extraction. At the same time, in order to better combine feature information at different scales, the BiFPN model is introduced into the FPN part for feature fusion, thereby improving the feature extraction accuracy of the GhostNet lightweight network.
[0021] 3. Replace all ordinary convolutional modules in the entire neck module with DWConv (included in the GhostBottleNeck module), and feed the output features into a detector with added CBAM for feature point judgment to ensure the uniformity of the overall convolutional modules.
[0022] 4. The improved YOLOv5s network model enhances the ability to identify small-target, high-density pests in unstructured environments. Attached Figure Description
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0024] Figure 1 This is a flowchart of a multi-target pest detection method based on an improved YOLOv5s unstructured environment, provided in Embodiment 1 of the present invention.
[0025] Figure 2 This is a structural diagram of a multi-target pest detection network model based on YOLOv5s improvement provided in Embodiment 1 of the present invention.
[0026] Figure 3 This is a schematic diagram illustrating the application of a Ghost module and a Ghost Bottleneck module provided in Embodiment 1 of the present invention in an improved YOLOv5s network.
[0027] Figure 4 This is a schematic diagram of a basic FPN module (PANet) and a feature fusion BiFPN module structure provided in Embodiment 1 of the present invention.
[0028] Figure 5 This is a schematic diagram of an attention mechanism module (CBAM) provided in Embodiment 1 of the present invention.
[0029] Figure 6 This is a comparison chart of feature levels before and after model improvement provided in Embodiment 1 of the present invention.
[0030] Figure 7 This is a comparison chart of the detection effects of aphid datasets before and after the improvement of the benchmark algorithm provided in Embodiment 1 of the present invention.
[0031] Figure 8 This is a schematic diagram of a multi-target pest detection system based on an improved YOLOv5s unstructured environment, provided in Embodiment 2 of the present invention.
[0032] Figure 9 This is a partial block diagram of the electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0033] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0034] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0035] For ease of understanding, the technical terms appearing in the following embodiments are explained here:
[0036] DWConv module: depthwise convolution.
[0037] BiFPN module: Bi-directional Feature Pyramid Network, a weighted bidirectional feature pyramid network.
[0038] FPN: Feature Pyramid Networks, used for object detection.
[0039] CBAM: Convolutional Block Attention Module, is an attention mechanism that combines channel attention and spatial attention, designed to improve the feature representation capabilities of convolutional neural networks.
[0040] The present invention will now be described in detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0041] Example 1
[0042] like Figure 1-2 The diagram shows a flowchart and a structural diagram of an improved YOLOv5s model for detecting multi-target pests in an unstructured environment, provided by the present invention.
[0043] As an example, the method includes:
[0044] S1: Collect images of crops infested by sugarcane aphids and preprocess the images to serve as the first target dataset.
[0045] S2: Randomly select and recombine the first target dataset and the optimized and labeled AgriPest dataset to create the second target dataset.
[0046] Preferably, steps S1 and S2 include: constructing a novel pest dataset for micro-targets by selecting four types of pests from three crops, including 860 self-built images of sugarcane aphids (sorghum crop) migrating to sorghum leaves, 1400 images of wheat spiders (wheat crop), 1000 images of wheat aphids (wheat crop), and 700 images of rice planthoppers (rice crop) from the AgriPest public dataset. The self-built sugarcane aphid dataset images were collected from the Texas A&M University Agricultural Life Sciences Research Institute farm, which covers approximately 40,000 m2 and where approximately 70-80% of the sorghum plants are infested with sugarcane aphids. In the dataset, the original images are processed using a method of first cutting and then labeling, taking into account the characteristics of the sugarcane aphid dataset images. The original images are divided into 16 sub-images of equal 4×4 dimensions, with an image resolution of 816×612. Other pest species in the dataset retain the original image format of the AgriPest public dataset, and the target annotations are optimized so that the size of the pest targets does not occupy more than 3% of the entire image area.
[0047] S3: Divide the second target dataset into a training set, a validation set, and a test set.
[0048] Preferably, the images are randomly divided into a training set, a validation set, and a test set in a 3:1:1 ratio.
[0049] S4: An improved YOLOv5s network model is constructed based on the YOLOv5s network model as a multi-target pest detection network model. This includes: removing the large detector head from the four detector heads of the YOLOv5s network, reconstructing the backbone network through the GhostNet module for feature extraction, introducing the BiFPN module into the FPN part for feature fusion, replacing the ordinary convolutional module in the neck network with the DWConv module, and feeding the output features into a detector with an added CBAM module for feature point judgment.
[0050] Preferably, after the mainstream deep learning models were built, they were trained and tested using the Tpest-3960 dataset, and the results are shown in Table 1. The comparison results of the YOLOv5s model with other object detection networks show that it performs excellently in terms of model complexity and detection performance parameters. Among them, Faster R-CNN is one of the classic two-stage object detection algorithms (ResNet50 was used as the feature extraction layer in the experiments); RetinaNet, YOLOX, YOLOv7-tiny, YOLOv7, and YOLOv8s are single-stage object detection algorithms.
[0051] Table 1 Performance comparison of different algorithms
[0052]
[0053] In terms of detection accuracy, Faster R-CNN, EfficientDet-D0, and RetinaNet all have an accuracy of less than 50%, making it difficult to accurately detect small-sized pest targets. YOLOX, YOLOv7-tiny, and YOLOv7 all have an accuracy of around 70%, but less than YOLOv5s' 80%. In terms of detection speed, YOLOv5s has an FPS of 222.22, which is better than other target detection algorithms. In terms of model size, Faster R-CNN and RetinaNet have model sizes of over 100MB, which greatly increases the deployment cost of the algorithms on hardware devices, while YOLOv5s has a model size of 13.7MB, which is only 64% of YOLOv8s. In terms of model parameter count and floating-point operations, YOLOv5s is second only to YOLOv7-tiny and the difference is not significant, but the detection accuracy of YOLOv7-tiny is far inferior to that of YOLOv5s. In summary, YOLOv5s exhibits the best overall performance in terms of detection accuracy, detection rate, and model complexity, achieving a certain balance. It reduces the storage and computing resource requirements of mobile hardware devices, lowers the difficulty of algorithm deployment, and maintains good detection accuracy and performance. Therefore, this embodiment uses the YOLOv5s network model as a baseline model for improvement.
[0054] Preferred, such as Figure 2As shown, the existing YOLOv5s model backbone network consists of a Focus module, a Conv module, a C3 module, and a Spatial Pyramid Pooling (SPP) module. It extracts features from the input image and transmits them to the neck network. The neck module employs a combination of a feature pyramid network and a path aggregation network, enhancing the detection capability for targets at different scales through bidirectional integration of low-level spatial features and high-level semantic features. The head network generates detection boxes, which are then applied to the multi-scale feature maps of the neck network to display the target's location, category, and confidence level. YOLOv5 allows users to adjust the model size by changing the network depth and width control factors, resulting in four models ranging from small to large: v5s, v5m, v5l, and v5x.
[0055] The network improvements based on YOLOv5s mainly include model lightweighting and performance optimization. Detailed structure is as follows: Figure 2 As shown, the specific steps are as follows:
[0056] Step S4, which involves reconstructing the backbone network for feature extraction using the GhostNet module, includes: S41: Replacing the C3 module in the original YOLOv5s network model with a Ghost Bottleneck module with a stride of 1, and replacing the ordinary convolutional Conv module in the original YOLOv5s network model with a Ghost Bottleneck module with a stride of 2, thus forming a feature extraction network with the GhostNet module as the backbone. In the figure, G-bneck represents the Ghost Bottleneck module.
[0057] Preferably, the convolution process in the Ghost Bottleneck module includes:
[0058] S411: The intrinsic feature map is obtained through conventional convolution, which is used to extract the basic features of the input data. Specifically, for ordinary convolution, assuming the number of channels in the input feature map is n and the convolution stride is s, the number of channels after the first part of convolution is m1, as shown in the formula m1 = m = n / s.
[0059] S412: Apply a linear transformation to the feature map of each channel to generate a Ghost feature map, including performing depthwise convolution on each channel using depthwise separable convolution to generate the Ghost feature map. Specifically, group convolution is performed, and for each feature map in m1 channels (obtained from the first step of convolution), it is mapped s times (including 1 identity mapping and (s-1) low-cost linear transformation mappings), as shown in the formula m2=m1×(s-1)=(n / s)(s-1). The output channel obtained through the second part is m2.
[0060] S413: Concatenate the intrinsic feature map and the Ghost feature map to obtain the final output feature map. Specifically, this is the Identity concatenation. Add and concatenate the feature layer with channel number m1 obtained in step S411 and the feature layer with channel number m2 obtained in step S412 to obtain the final output feature layer, as shown in the formula m3 = m1 + m2 = m1 + m1 × (s-1) = m × s, where the number of channels is m3. Because m = n / s, the number of output channels after convolution remains n. This allows the module to perform a convolution operation with fewer parameters while maintaining the same number of output channels as the input, making it plug-and-play for any convolutional neural network model. Using this module can significantly reduce model parameters while performing feature extraction.
[0061] Preferably, the network structure of the Ghost module and the Ghost Bottleneck modules stacked therefrom with step sizes of 1 and 2 are as follows: Figure 3 As shown in (a), the Ghost Bottleneck is mainly composed of two stacked Ghost modules. Ghost module 1 is an extension layer that increases the feature dimension by increasing the number of channels. Ghost module 2 is responsible for reducing the number of channels, which helps to reduce the feature dimension and maintain feature consistency through shortcut connections at the input. The connection between Ghost modules 1 and 2 introduces BN and ReLU layers. Only BN layers are applied in the post-connection of Ghost module 2, where BN and ReLU represent batch normalized rectified linear units, and Conv represents convolution.
[0062] Preferably, replacing the ordinary convolutional modules in the neck network with DWConv modules in step S4 includes:
[0063] S44: The two Ghost Bottleneck modules are connected by a depthwise separable convolution DWConv with a stride of 2, replacing all ordinary convolutional blocks (Conv) in the neck convolution BiFPN module. Specifically, DWConv, as a lightweight convolution operation, produces feature maps with the same number of channels as the input layer, reducing parameter computation. This is completely different from conventional convolutions where each kernel operates on every channel of the input image simultaneously; in the DWConv module, each channel is operated on by only one kernel, which is beneficial for extracting pest and disease features of almost equal size. To more effectively utilize model parameters, DWConv has been successfully applied in lightweight networks Xception and MobileNets. In the Ghost Bottleneck module, for the stride of 2, as shown... Figure 3 As shown in (b), the two Ghost modules are connected via a DWConv with a step size of 2.
[0064] Preferably, step S4, which involves introducing the BiFPN module into the FPN part for feature fusion, includes:
[0065] S42: Remove the feature pyramid network to reduce its impact on the overall network and simplify the bidirectional network. S43: Add connection edges between the original input and output nodes of the same feature layer to reduce the incremental computational cost. Specifically, the recently proposed BiFPN can achieve effective bidirectional cross-scale connections and weighted feature fusion. Compared to the feature pyramid network (FPN for object detection), its superior performance is mainly reflected in... Figure 4 The two steps shown are: S42: Deleting the feature pyramid network reduces its impact on the overall network (because a single input contributes very little to feature fusion) and simplifies the bidirectional network; S43: Adding connection edges between the original input nodes and output nodes of the same feature layer reduces the incremental computational cost. In the figure, P1-P2 represent level 1-2 feature extraction, and P3-P7 represent level 3-7 multi-scale features fused from top to bottom. By improving the original FPN+PAN structure to a BiFPN structure, the aim is to enhance feature extraction capabilities by fusing more feature layers through simple residual operations without increasing resource consumption.
[0066] Preferably, in step S4, the output features are fed into a detector with an added CBAM module for feature point determination, including: S45: Integrating the CBAM module into the downsampling step of the neck network, the output of the C3 module is transmitted to the detection head through the CBAM module. Specifically, the CBAM attention mechanism module is a plug-and-play lightweight module that can be seamlessly integrated into the convolutional neural network architecture and assists each feature layer in capturing useful information based on the layer depth. This module mainly involves two independent convolutional steps that sequentially infer the input feature map, namely the channel attention module and the spatial attention module.
[0067] Preferred, such as Figure 5 As shown, the CBAM module includes two independent convolutional steps that sequentially infer the input feature map: a channel attention module and a spatial attention module. The workflow includes:
[0068] S451: Given an intermediate feature map F∈R C×H×W As input, the input image F is subjected to global max pooling and global average pooling by channel in the channel attention module to obtain two one-dimensional vectors;
[0069] S452: The vector is processed through a fully connected layer, performing element-wise summation and sigmoid activation to generate a one-dimensional channel attention mechanism feature map M. c ∈RC×1×1 ;
[0070] S453: As shown in the formula M c Element-wise multiplication with the input feature map F yields the channel attention-adjusted feature map F. ′ ;
[0071] S454: Perform global max pooling and average pooling based on space;
[0072] S455: Concatenate the two two-dimensional vectors generated by pooling and then perform a convolution operation;
[0073] S456: The dimensionality-reduced feature map is activated by a Sigmoid operation to generate a spatial attention feature map M. S ∈R 1×H×W ;
[0074] S457: Spatial attention feature map M S and input feature mapping F ′ Perform multiplication to obtain the final generated features.
[0075] In the formula, This represents element-wise multiplication, where M represents the multiplication operation. c This indicates that attention extraction is performed along the channel dimension, M. c (F) represents the weight coefficients obtained using the channel attention mechanism, M s This indicates that attention extraction is performed in the spatial dimension, M S (F ′ F represents the weight coefficients obtained using the spatial attention mechanism. ′ This represents the output feature map after the input feature map F has undergone channel attention adjustment. ′′ This indicates the final, precise output.
[0076] Preferably, step S4, removing the large detector head from the four detector heads of the YOLOv5s network, includes: S46: In the neck network, a new P2 small-scale prediction head is added, including a 160×160×128 upsampling step, removing the large target prediction head, including a 20×20×1024 feature layer output. Specifically, the feature level comparison before and after model improvement is as follows: Figure 6 As shown, a new P2 small-scale prediction head (160×160×128 upsampling step) was added to the neck network, while the large target prediction head (20×20×1024 feature layer output) was removed. Specifically, as... Figure 6 As shown in (b), unlike Figure 6In (a), the original model's P3, P4, and P5 feature levels are improved by omitting the 20×20×1024 downsampling step in the backbone and downsampling the feature map to a 40×40×512 feature layer. The output of this feature layer is 1 / 16 of the original image size, and this output will be fed into the neck network.
[0077] Preferably, step S5: inputting the training set into the multi-target pest detection network model for model training; S6: improving the parameters of the multi-target pest detection network model based on the training results to obtain the trained multi-target pest detection network model, including:
[0078] The main tools in the development environment that the improved model depends on are shown in Table 2. These toolkits and related Python libraries can be imported into the corresponding environment through the Visual Studio Code platform terminal. Simply enter pipinstall and the name of the toolkit or Python library to automatically install and import them.
[0079] To accelerate the training speed of the improved model, a transfer learning-based approach was used. Transfer learning is a process in convolutional neural network training where a model from one project is used as initial weights, and the learned knowledge is transferred from that project to find similarities with targets. This knowledge is then reused to train a model for another project, thereby improving the learning performance of that project. Transfer learning allows models pre-trained on large datasets to be transferred to datasets with fewer images. This enables models trained on datasets with limited image samples to achieve good detection performance even with limited image samples, and avoids training from scratch, saving computational resources and shortening training time. In this study, YOLOv5s was used as the baseline network model. The parameters of the improved model were set to the weights of YOLOv5s.pt pre-trained on the COCO (Common Objects in Context) dataset, and then fine-tuned. The main hyperparameters used during training are shown in Table 3.
[0080] Table 2. Development Environment for the Improved YOLOv5s Model
[0081]
[0082]
[0083] Table 3 Hyperparameters for network training
[0084]
[0085] Preferably, step S7, which involves inputting the validation set into the trained multi-target pest detection network model for model evaluation, includes: inputting the validation set into the trained multi-target pest detection network model and evaluating it based on the precision P, recall R, and mean precision mAP of the multi-target pest detection network model.
[0086]
[0087] In the formula, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, and FP is the number of incorrectly predicted positive samples.
[0088] This embodiment uses mAP as a performance reference metric. In balancing lightweight design and performance, the number of model parameters and computational cost must also be considered. Therefore, floating-point operations (GFLOPs), number of parameters, and detection speed (FPS) need to be introduced as detailed parameters of the model architecture.
[0089] The experimental results in Table 4 show that the improved model outperforms other advanced object detection models in terms of model size, number of parameters, and FLOPs—three lightweight metrics. In terms of model size, it is only 20% of the YOLOv5s (the baseline model), compared to 8%, 24%, 4%, 13%, 3%, 19%, and 2% for YOLOX, YOLOv7-tiny, YOLOv7, YOLOv8s, Faster R-CNN, EfficientDet-D0, and RetinaNet, respectively. The improved YOLOv5s model also has the fewest parameters, ranging from 1% to 26% of the other eight models. Furthermore, the improved YOLOv5s model has FLOPs comparable to EfficientDet-D0, but significantly smaller than the other seven models (2-58%).
[0090] Table 4 Performance Comparison of Improved Methods and Different Algorithms
[0091]
[0092] The improved model also demonstrates a good balance between detection accuracy and efficiency. On the Tpest-3960 dataset, it achieves an mAP of 80.1%, which is higher than YOLOv5s, YOLOv7, YOLOX, YOLOv7-tiny, RetinaNet, EfficientDet-D0, and Faster R-CNN by 0.1%, 10.0%, 10.3%, 10.3%, 51.1%, 51.3%, and 53.9%, respectively. Although YOLOv8s has a slightly higher detection accuracy (0.9%) than the improved model, the improved model exhibits a faster detection speed, with an FPS 1.6% higher than YOLOv8s. Furthermore, the improved model has a significantly smaller size, only 13% of that of YOLOv8s, making it more suitable for real-time detection and mobile deployment.
[0093] Preferably, S8: Input the test set into the trained multi-target pest detection network model to obtain multi-target pest detection results.
[0094] like Figure 7 As shown, the Grad-Cam heatmap method was used to analyze and count 50 randomly selected images. Compared with YOLOv5s, the improved model showed better detection performance for densely distributed or scattered pests in complex backgrounds (such as overlapping and closure of leaves). Figure 7 As shown in (a), the improved model accurately detected three wheat aphids (WM) hidden deep within the leaves. This improved model demonstrated accurate detection capabilities even when the WMs were scattered across different leaves, including those farther from the camera. In contrast, YOLOv5s failed to identify these scattered WMs due to overlapping and closed leaves. Unlike YOLOv5s, which showed more than one missed detection for both sugarcane aphids and wheat aphids (e.g., ...), the improved model also detected multiple missed detections. Figure 7 (b) and 7(c) sugarcane aphid (SCA) and wheat aphid (WA), the improved model also showed good recognition accuracy in detecting densely distributed and adherent pests, significantly reducing the problem of missed detection due to adhesion. In addition, such as Figure 7 As shown in (d), the improved model can accurately identify rice planthoppers (RPs) even when some are covered and blocked from view.
[0095] The heatmap further demonstrates that the improved model exhibits higher sensitivity for small target detection compared to the reference baseline method (YOLOv5s). This improvement is achieved by combining BiFPN-enhanced multi-channel feature fusion with the spatial attention matrix of CBAM, enabling the network to automatically select important channels and spatial information during computation. Therefore, the enhanced network attention capability significantly improves the recognition results and effectively reduces the probability of missed and false detections.
[0096] The above embodiments improve detection speed by removing the large detector from the four detectors in the YOLOv5s network model and using a three-layer detector to reduce the number of model parameters, making the improved model more suitable for small-target pest detection tasks in real agricultural scenarios. The GhostNet module replaces the original backbone module for feature extraction. To better combine feature information at different scales, a BiFPN model is introduced into the FPN part for feature fusion, improving the feature extraction accuracy of the lightweight GhostNet network. All ordinary convolutional modules in the entire neck module are replaced with DWConv (included in the GhostBottleNeck module), and the output features are fed into a detector with added CBAM for feature point judgment, ensuring the uniformity of the overall convolutional modules. The improved YOLOv5s network model enhances its ability to identify small-target, high-density pests in unstructured environments.
[0097] Example 2
[0098] Please see Figure 8 This embodiment provides a schematic diagram of a multi-target pest detection system based on an improved YOLOv5s unstructured environment.
[0099] As an example, the system includes:
[0100] The first target dataset unit 810 is used to collect images of crops infested by sugarcane aphids and preprocess the images to serve as the first target dataset.
[0101] The second target dataset unit 820 is used to randomly select and recombine the first target dataset and the optimized labeled AgriPest dataset to create the second target dataset.
[0102] Dataset partitioning unit 830 is suitable for partitioning the second target dataset into a training set, a validation set, and a test set.
[0103] Model Improvement Unit 840 is suitable for building an improved YOLOv5s network model based on the YOLOv5s network model as a multi-target pest detection network model. It includes: removing the large detector head from the four detector heads of the YOLOv5s network, reconstructing the backbone network through the GhostNet module for feature extraction, introducing the BiFPN module into the FPN part for feature fusion, replacing the ordinary convolutional module in the neck network with the DWConv module, and feeding the output features into a detector with the CBAM module for feature point judgment.
[0104] The model training unit 850 is suitable for inputting the training set into the multi-target pest detection network model for model training.
[0105] The model improvement unit 860 is suitable for improving the parameters of a multi-target pest detection network model based on the training results to obtain the trained multi-target pest detection network model.
[0106] The model evaluation unit 870 is suitable for inputting a validation set into the multi-target pest detection network model for model evaluation.
[0107] The result generation unit 880 is suitable for inputting the test set into a multi-target pest detection network model to obtain multi-target pest detection results.
[0108] It is not difficult to see that this embodiment is a system implementation corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0109] It is worth mentioning that all units involved in this embodiment are logical units. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0110] Example 3
[0111] This invention also proposes a storage medium storing a multi-target pest detection method based on an improved YOLOv5s unstructured environment. When executed by a processor, the multi-target pest detection program based on the improved YOLOv5s unstructured environment implements the steps of the multi-target pest detection method based on the improved YOLOv5s unstructured environment described above. Since this storage medium adopts all the technical solutions of all the above embodiments, it possesses at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be elaborated upon here.
[0112] Example 4
[0113] Please see Figure 9 The present invention also provides an electronic device, including: a memory and a processor; the memory stores at least one program instruction; the processor loads and executes the at least one program instruction to implement the multi-target pest detection method based on the improved YOLOv5s unstructured environment provided in Embodiment 1.
[0114] The memory 602 and processor 601 are connected via a bus, which may include any number of interconnecting buses and bridges, connecting various circuits of one or more processors 601 and memory 602 together. The bus may also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 601 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 601.
[0115] Processor 601 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 602 can be used to store data used by processor 601 during operation.
[0116] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, based on the guidance provided in this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A method for detecting multi-target pests in an unstructured environment based on improved YOLOv5s, characterized in that, The method includes: S1: Collect images of crops infested by sugarcane aphids and preprocess the images to serve as the first target dataset; S2: Randomly select and recombine the first target dataset and the optimized and labeled AgriPest dataset to create the second target dataset; S3: Divide the second target dataset into a training set, a validation set, and a test set; S4: An improved YOLOv5s network model was constructed based on the YOLOv5s network model as a multi-target pest detection network model. S5: Input the training set into the multi-target pest detection network model for model training; S6: Improve the parameters of the multi-target pest detection network model based on the training results to obtain the trained multi-target pest detection network model; S7: Input the validation set into the trained multi-target pest detection network model for model evaluation; S8: Input the test set into the trained multi-target pest detection network model to obtain the multi-target pest detection results; S4: The step of constructing an improved YOLOv5s network model as a multi-target pest detection network model based on the YOLOv5s network model includes: S41: The C3 module in the original YOLOv5s network model is replaced with the Ghost Bottleneck module with a stride of 1, and the ordinary convolutional Conv in the original YOLOv5s network model is replaced with the Ghost Bottleneck module with a stride of 2, forming a feature extraction network with the GhostNet module as the backbone. S42: Remove the feature pyramid network to reduce its impact on the overall network and simplify the bidirectional network; S43: Add connecting edges between the original input nodes and output nodes of the same feature layer to reduce the increase in computational cost; S44: The two Ghost Bottleneck modules are connected by a depth-separable convolution DWConv with a stride of 2, replacing all ordinary convolutional blocks Conv in the neck convolution BiFPN module with DWConv. S45: Integrate the CBAM module into the downsampling step of the neck network, and transmit the output of the C3 module to the detection head through the CBAM module; S46: In the neck network, a new P2 small-scale prediction head was added, including an upsampling step of 160×160×128, and the large target prediction head was removed, including a feature layer output of 20×20×1024.
2. The multi-target pest detection method based on the improved YOLOv5s unstructured environment according to claim 1, characterized in that, The convolution process in the Ghost Bottleneck module includes: S411: The intrinsic feature map is obtained through conventional convolution and used to extract the basic features of the input data; S412: Apply a linear transformation to the feature map of each channel to generate a Ghost feature map, including performing a depthwise convolution on each channel using a depthwise separable convolution to generate a Ghost feature map. S413: The intrinsic feature map and the Ghost feature map are concatenated to obtain the final output feature map.
3. The multi-target pest detection method based on the improved YOLOv5s unstructured environment according to claim 1, characterized in that, The CBAM module includes two independent convolutional steps that sequentially infer the input feature map: a channel attention module and a spatial attention module. The workflow includes: S451: Given an intermediate feature map As input, the input image is processed in the channel attention module. Perform global max pooling and global average pooling on each channel to obtain two one-dimensional vectors; S452: The vector is processed through a fully connected layer, performing element-wise summation and sigmoid activation to generate a one-dimensional channel attention mechanism feature map. ; S453: As shown in the formula ,Will Mapping with input features Perform element-wise multiplication to obtain the feature map for channel attention adjustment. ; S454: Perform global max pooling and average pooling based on space; S455: Concatenate the two two-dimensional vectors generated by pooling and then perform a convolution operation; S456: The dimensionality-reduced feature map is then subjected to a Sigmoid activation operation to generate a spatial attention feature map. ; S457: Spatial attention feature map and input feature mapping Perform multiplication to obtain the final generated features. ; In the formula, This represents element-wise multiplication, during the multiplication operation, This indicates that attention extraction is performed along the channel dimension. This represents the weight coefficients obtained using the channel attention mechanism. This indicates that attention extraction is performed in the spatial dimension. This represents the weight coefficients obtained using the spatial attention mechanism. Represents the input feature map The feature map output after channel attention adjustment This indicates the final, precise output.
4. The multi-target pest detection method based on the improved YOLOv5s unstructured environment according to claim 1, characterized in that, Step S7, which involves inputting the validation set into the trained multi-target pest detection network model for model evaluation, includes: S71: Input the validation set into the trained multi-target pest detection network model and evaluate it based on the precision P, recall R and mean precision mAP of the multi-target pest detection network model. ; ; ; ; In the formula, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, and FP is the number of incorrectly predicted positive samples.
5. A multi-target pest detection system based on an improved YOLOv5s unstructured environment, characterized in that, The system includes: The first target dataset unit is generated, which is suitable for collecting images of crops infested by sugarcane aphids and preprocessing the images to serve as the first target dataset. The second target dataset unit is suitable for randomly selecting and recombining the first target dataset and the optimized and labeled AgriPest dataset to create the second target dataset. The dataset partitioning unit is suitable for dividing the second target dataset into a training set, a validation set, and a test set; The model improvement unit is suitable for building an improved YOLOv5s network model based on the YOLOv5s network model as a multi-target pest detection network model. The model training unit is suitable for inputting the training set into the multi-target pest detection network model for model training. The model improvement unit is suitable for improving the parameters of a multi-target pest detection network model based on the training results to obtain a trained multi-target pest detection network model. The model evaluation unit is suitable for inputting the validation set into the multi-target pest detection network model for model evaluation. The result generation unit is suitable for inputting the test set into the multi-target pest detection network model to obtain multi-target pest detection results; The model improvement unit includes replacing the C3 module in the original YOLOv5s network model with a Ghost Bottleneck module with a stride of 1, and replacing the ordinary convolutional Conv in the original YOLOv5s network model with a Ghost Bottleneck module with a stride of 2, forming a feature extraction network with GhostNet module as the backbone. The feature pyramid network is removed to reduce its impact on the overall network and to simplify the bidirectional network. Adding connecting edges between the original input and output nodes of the same feature layer can help reduce the increase in computational cost. The two Ghost Bottleneck modules are connected by a depth-separable convolution DWConv with a stride of 2, replacing all ordinary convolutional blocks Conv in the neck convolution BiFPN module. The CBAM module is integrated into the downsampling step of the neck network, and the output of the C3 module is transmitted to the detection head through the CBAM module; In the neck network, a new P2 small-scale prediction head was added, including an upsampling step of 160×160×128, and the large target prediction head was removed, including a feature layer output of 20×20×1024.