A method for detecting stored grain pests based on the YOLOv5s algorithm
Through the YOLOv5s algorithm-based grain storage pest detection method, the model structure and training process are optimized, and the problems of poor real-time performance and complex model detection in the existing technology are solved, achieving efficient and accurate detection results.
Patent Information
- Application Number
- CN202310758408.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-06-26
AI Technical Summary
In the prior art, the real-time real-time detection process of grain storage pests is poor and the model structure is complex, making it difficult to achieve efficient detection.
The grain storage pest detection method based on the YOLOv5s algorithm is adopted. By obtaining the sample images of grain storage pests, the sample set is constructed and pre-processed, the model structure is optimized, including data enhancement and lightweight processing of the backbone network, and the grain storage pest detection model is trained.
It improves the accuracy and speed of detection and positioning of grain storage pests, reduces the probability of missed detection and error detection, simplifies the model structure, and improves detection efficiency.
Smart Images

Figure CN116797834B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image target detection, and specifically to a method for detecting stored-grain pests based on the YOLOv5s algorithm. Background Art
[0002] Stored-grain pests are one of the main reasons for losses in the grain storage process. Efficiently detecting the types and quantities of stored-grain pests is an important task for ensuring safe grain storage. Traditional image-processing-based methods for detecting stored-grain pests (such as support vector machines and backpropagation neural networks, etc.) often fail to achieve satisfactory results. Image-processing methods based on deep learning avoid the cumbersome steps of manually designing grain pest features, can automatically learn and generalize the features of a large amount of image data, classify the feature vectors of grain pests, and quickly identify different types of stored-grain pests. Therefore, applying deep learning models to the small-target detection of stored-grain pests can achieve efficient detection of stored-grain pests, reduce economic losses, and improve the quality of the grain storage environment. Existing detection methods based on deep learning models are mainly divided into two categories: two-stage object detection methods based on region proposals and one-stage object detection methods based on regression. The most representative two-stage object detection algorithms are mainly the RCNN series, and the one-stage object detection algorithms are the YOLO series, SSD series, etc. Whether it is a two-stage or one-stage object detection algorithm, due to the lack of adaptive improvements for stored-grain pests, there are generally problems such as poor real-time performance in the detection process and complex model structures. Summary of the Invention
[0003] To solve the problems of poor real-time performance in the detection process and complex model structure in the prior art, the present invention provides a method for detecting stored-grain pests based on the YOLOv5s algorithm, including the following steps:
[0004] S1. Obtain sample images of stored-grain pests to construct a sample set and perform preprocessing to obtain a stored-grain pest image data set, and divide the image data set into a training set and a validation set;
[0005] S2. Based on the YOLOv5s algorithm framework, obtain an optimized model through data augmentation processing and backbone network lightweight processing;
[0006] S3. Input the training set into the optimized model for training to obtain a stored-grain pest detection model;
[0007] S4. Obtain an image of a stored-grain pest to be recognized, and input it into the trained stored-grain pest monitoring model for recognition to obtain the type and location of the stored-grain pest.
[0008] As a further optimization of the above method for detecting stored-grain pests based on the YOLOv5s algorithm, in S1, the specific method for constructing the sample set includes:
[0009] S11. Determine multiple stored - grain pests as detection targets, and randomly select multiple detection targets as a group;
[0010] S12. Take pictures of a group of detection targets against a white - board background to obtain stored - grain pest sample images;
[0011] S13. After data augmentation of the stored - grain pest sample images, combine them to obtain a sample set.
[0012] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm: The data augmentation method includes rotation, flipping, mirroring, scaling, and / or translation.
[0013] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm, in S2, the specific process of data enhancement includes: upgrading from stitching n pictures to get one picture as a training sample to stitching m pictures to get one picture as a training sample, and m > n.
[0014] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm, in S2, the specific process of lightweight backbone network includes: removing the fully - connected layer and the Softmax layer.
[0015] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm: The optimized model also introduces label smoothing in classification.
[0016] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm: The optimized model also integrates the BiFPN feature pyramid structure.
[0017] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm: The optimized model also integrates the Swin Transformer block into the detection head Head.
[0018] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm: The specific method of S3 includes: continuously adjusting the weights in the network using the loss function, and then calculating the mean average precision using the validation set.
[0019] As a further optimization of the above - mentioned stored - grain pest detection method based on the YOLOv5s algorithm: If the mean average precision reaches the set threshold, it is considered qualified, and the qualified weight file is loaded into the improved YOLOvs5 algorithm.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention applies the multi-layer convolutional neural network structure in deep learning to solve the problems of low real-time performance and complex model structure existing in the current detection of stored-grain pests, effectively improving the accuracy and speed of the detection and positioning of stored-grain pests, reducing the probability of missed detection and misdetection, enriching the existing detection technology system at the same time, improving the efficiency of the detection of stored-grain pests, and promoting the scientific prevention and control of stored-grain pests. Description of the Drawings
[0021] Figure 1 is a flowchart of a method for detecting stored-grain pests based on the YOLOv5s algorithm proposed by the present invention. Detailed Embodiments
[0022] The following further elaborates on the technical solutions of the present invention in combination with specific embodiments. For parts not detailedly recorded and disclosed in the following embodiments of the present invention, they should all be understood as the prior art known or should be known to those skilled in the art.
[0023] As Figure 1 shown, a method for detecting stored-grain pests based on the YOLOv5s algorithm includes the following steps:
[0024] S1. Obtain stored-grain pest sample images to construct a sample set and perform preprocessing to obtain a stored-grain pest image data set, and divide the image data set into a training set and a validation set;
[0025] S2. Based on the YOLOv5s algorithm framework, enhance the original Mosaic data augmentation to Mosaic-9 data augmentation, enhancing the data volume of small target objects and the complexity of the image background, and reducing the computational amount required for batch normalization; obtain an optimized model through data augmentation processing and backbone network lightweight processing;
[0026] S3. Input the training set into the optimized model for training to obtain a stored-grain pest detection model;
[0027] S4. Obtain the stored-grain pest image to be recognized, and input it into the trained stored-grain pest monitoring model for recognition to obtain the types and positions of the stored-grain pests.
[0028] In S1, the specific method for constructing the sample set includes:
[0029] S11. Determine multiple stored-grain pests as detection targets, and randomly select multiple detection targets as a group. Select Tribolium castaneum, Cryptolestes ferrugineus, Sitophilus zeamais, Rhizopertha dominica, and Plodia interpunctella as the sample set, and randomly combine any three stored-grain pests, with three of each stored-grain pest;
[0030] S12. Take pictures of a group of detection targets against a whiteboard background to obtain stored-grain pest sample images. The original image resolution is set to 720×1280. After shooting, preprocess the images to obtain 1990 images with a resolution of 640×640, and divide them into a training set and a validation set according to the ratio of 8:2.
[0031] S13. After data augmentation of the stored-grain pest sample images, combine them to obtain a sample set, and use image processing methods such as rotation, flipping, mirroring, scaling, and / or translation to augment the stored-grain pest images.
[0032] In S2, the specific processing process of data enhancement includes: splicing n pictures to obtain one picture as a training sample and upgrading it to splicing m pictures to obtain one picture as a training sample, and m > n. In the existing YOLOv5s algorithm, the Mosaic method is to splice four pictures to obtain one picture as a training sample, that is, n = 4, and the Mosaic-9 method is to splice nine pictures to obtain one picture as a training sample, that is, m = 9. And the splicing methods include random scaling, random cropping, and / or random arrangement. The arrangement method can be nine pictures arranged in a column, in a row, or in a matrix arrangement. In this embodiment, a matrix arrangement is adopted, that is, 3×3. This not only enriches the background of the dataset images, but also calculates the data of nine pictures at one time during calculation, so that good results can be achieved even at a small size and the generalization ability is enhanced.
[0033] In S2, the specific processing process of lightweight backbone network includes: replacing the backbone network of YOLOv5s with a lightweight network, such as MobileNetv3, and removing the fully connected layer and the Softmax layer. This not only reduces the volume of the model, but also improves the feature extraction ability.
[0034] The input picture size is 640×640 and contains 3 color channels. After preprocessing, it first passes through the Focus module and uses slicing operations to be converted into a feature map of 320×320×12, where 12 = 4×3 means that the original picture is divided into 4×3 = 12 regions for processing, and the feature maps of each sub-region are concatenated to generate a new feature map. Finally, through the convolution operation with 32 convolutional kernels, it becomes a feature map of 320×320×32, which is used to generate richer and more complex feature representations for detecting target objects.
[0035] After passing through the Focus slicing operation, stack the Conv (convolution module) and C3 modules three times to gradually obtain feature maps of different sizes, realizing the full extraction of fine-grained features of shallow information and deep high-level semantic information.
[0036] To avoid relying too much on manual labels, the optimized model also introduces label smoothing in classification. The label smoothing method converts the original one-hot label y into a new probability distribution y′, where the probability of each label is:
[0037]
[0038] where K is the total number of labels, and ∈ is a small positive number, taking a relatively small value such as 0.1 or 0.05. The meaning of this formula is: for label i, the weight of the original one-hot probability (u i = 1) is reduced by ∈ and distributed to all other labels, so that their probabilities are all increased by thus obtaining a smoother label distribution. During training, the smoothed label y′ is used instead of the original one-hot label y for model training to prevent the model from concentrating its decisions too much on one category and distributing its attention more evenly to all categories. In addition, label smoothing can also introduce noise, thereby improving the generalization ability of the model.
[0039] In the original YOLOv5s, the Neck part uses FPN and PAN for feature extraction and fusion. After fusion processing, it is input into the prediction network to locate and classify the objects in the feature map. Although it can solve the problem of large differences in object scales in images of different scenarios, it is not the optimal solution for feature fusion. Therefore, the optimized model also incorporates a BiFPN feature pyramid structure, adopts a bidirectional feature context propagation mechanism, and simultaneously introduces a multi-way feature fusion and gradient weighting strategy to achieve better feature representation ability and higher computational efficiency. In BiFPN, there are two branches, one is the top-down branch and the other is the bottom-up branch. Among them, the top-down branch increases the resolution of features through upsampling, and conversely, the bottom-up branch reduces the resolution of features through downsampling. These two branches are responsible for performing feature fusion in different directions on the input feature map respectively, thus realizing the modeling of global features and the fusion of multi-scale features. At the same time, during the feature fusion process, BiFPN also adopts a multi-way feature fusion method to fuse feature maps from different depths to enhance the diversity and reliability of features. To ensure the weights of features at different depths, BiFPN also introduces a gradient weighting strategy to adjust the contributions of different paths of features through the supervision signal, so as to make better use of information.
[0040] The optimized model also integrates the Swin Transformer blocks into the detection head (Head) to improve the resolution of the feature map at the end of the network. Specifically, the Swin Transformer is a new self-attention mechanism module that can enhance the global view and semantic information of the feature map in a step-by-step decomposition manner and optimize the computational efficiency using a windowing approach. In the detection head of YOLOv5s, the Swin Transformer is used to complete two tasks: one is to construct the feature pyramid, and the other is to perform the operation of feature upsampling. The construction of the feature pyramid means that the Swin Transformer is used to gradually fuse the feature map at the bottom layer with the feature map at the top layer to obtain a series of feature pyramids with different resolutions. This process involves the cascading of multiple Swin Transformer modules, where each module is responsible for processing its input feature map, generating a higher-level feature representation, and inputting it to the next module for processing. The purpose of this is to retain the detailed information at the bottom layer while enhancing the semantic information at the high layer, enabling the model to better identify targets of different sizes and shapes. The operation of feature upsampling means that the Swin Transformer is used to perform upsampling (also known as transposed convolution or deconvolution) of the feature map to restore the low-resolution feature map to the original size. This process also involves the cascading of multiple Swin Transformer modules, where each module is responsible for spatially expanding and reconstructing the input feature map and outputting it to the next module for processing. The purpose of this is to avoid the information loss and blurring problems caused by traditional interpolation algorithms (such as bilinear interpolation), enabling the model to more accurately locate the position and size of stored grain pests.
[0041] The specific method of S3 includes: continuously adjusting the weights in the network using the loss function, and then calculating the average precision using the validation set. If the average precision reaches the set threshold, it is considered qualified, and the qualified weight file is loaded into the improved YOLOv5s algorithm.
[0042] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting stored - grain pests based on YOLOv5s, characterized in that, The steps include: S1. Obtain the sample images of stored-grain pests to construct a sample set and perform preprocessing to obtain a stored-grain pest image dataset, and divide the image dataset into a training set and a validation set; S2. Based on the YOLOv5s algorithm framework, obtain an optimized model through data augmentation processing and backbone network lightweight processing; the optimized model is also integrated with a BiFPN feature pyramid structure; the optimized model also introduces label smoothing in classification; the optimized model also integrates the Swin Transformer block into the detection head Head; In S2, the specific process of backbone network lightweight processing includes: removing the fully connected layer and the Softmax layer; S3. Input the training set into the optimized model for training to obtain a stored-grain pest detection model; S4. Obtain the stored-grain pest image to be recognized, and input it into the trained stored-grain pest monitoring model for recognition to obtain the type and location of the stored-grain pests.
2. The method for detecting stored grain pests based on YOLOv5s according to claim 1, characterized in that, In S1, the specific method for constructing the sample set includes: S11. Determine multiple stored-grain pests as detection targets, and randomly select multiple detection targets as a group; S12. Take pictures of a group of detection targets against a whiteboard background to obtain stored-grain pest sample images; S13. Combine the stored-grain pest sample images after data augmentation to obtain a sample set.
3. The method for detecting stored grain pests based on YOLOv5s according to claim 2, characterized in that: The data augmentation method includes rotation, flipping, mirroring, scaling, and / or translation.
4. The method for detecting stored grain pests based on YOLOv5s according to claim 1, wherein, In S2, the specific process of data augmentation includes: splicing n pictures to obtain one picture as a training sample and upgrading it to splicing m pictures to obtain one picture as a training sample, and m > n.
5. The method for detecting stored grain pests based on YOLOv5s according to claim 1, characterized in that, The specific method of S3 includes: continuously adjusting the weights in the network using the loss function, and then calculating the average precision using the validation set.
6. The method for detecting stored - grain pests based on YOLOv5s according to claim 5, wherein: If the average precision reaches the set threshold, it is considered qualified, and the qualified weight file is loaded into the improved YOLOvs5 algorithm.
Citation Information
Patent Citations
Germinated potato image recognition method based on improved yolov5 model
CN114120037A
Rapid pest detection method based on improved YOLO V4
CN114220035A