Warehouse grain insect detection method based on improved YOLOv8n network

By improving the YOLOv8n network, a lightweight warehousing grain insect detection model is built, which solves the problems of low detection accuracy and high computing resources, and realizes efficient, accurate detection and timely early warning in complex environments. It is suitable for resource-constrained scenarios such as granaries and farms.

CN120339742APending Publication Date: 2025-07-18HENAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510335841.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology has problems with low detection accuracy and high computing resource requirements in the detection of grain storage insects, especially in complex environments and intensive grain storage scenarios, which are difficult to achieve timely and efficient pest monitoring.

Method used

Using the improved YOLOv8n network, by introducing lightweight feature extraction modules, enhanced YOLO detection head modules and in-scale feature interaction modules, we will build a lightweight warehousing detection model, including data acquisition, preprocessing, model training and evaluation, and optimize the model architecture to improve detection accuracy and reduce computational complexity.

Benefits of technology

It improves the accuracy and robustness of small pest detection, reduces the computing resource requirements, is suitable for resource-constrained granary and farm environments, and achieves accurate detection and timely early warning in complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339742A_ABST
    Figure CN120339742A_ABST
Patent Text Reader

Abstract

The invention discloses a storage grain insect detection method based on an improved YOLOv8n network, and the method comprises the following steps: S1, data collection and preprocessing; s2, model architecture selection; s3, training a storage grain insect identification model; s4, performing model evaluation; s5, acquiring an image to be detected; and S6, inputting the preprocessed to-be-detected storage grain insect image into the trained storage grain insect detection model to obtain a storage grain insect detection result. According to the invention, the problems of low detection precision and high computing resource demand when small storage grain insects are detected are solved. The method can effectively improve the detection precision of small storage grain insects, remarkably reduces the calculation amount and the parameter amount, and is suitable for an intelligent detection scene in granary management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of warehouse grain pest detection, and particularly to a warehouse grain pest detection method based on an improved YOLOv8n network. Background Art

[0002] As an important strategic resource of the country, grain plays a crucial role in ensuring people's livelihood and stabilizing the economy. The grain storage industry not only needs to ensure the safe storage of bulk grain but also prevent losses caused by warehouse grain pests. Data shows that every year, about 10% of the stored grain volume is affected by pest infestation, with the loss amount reaching up to 2 billion yuan. In addition, the secretions and corpses of pests will also contaminate the grain, resulting in a decline in its quality.

[0003] Therefore, developing an efficient and accurate warehouse grain pest identification technology has important theoretical and application values for improving grain storage efficiency and ensuring food safety. Traditional pest monitoring methods mainly include probe sampling and acoustic detection methods, etc. Probe sampling extracts grains (0.5 - 1 kg) from the storage bin through a probe, and then screens the insects in the grains with a sieve. This method is easily affected by human factors and environmental factors, and if the sampling method is not scientific, the sampling points are insufficient, or the sampling quantity is not enough, the sample will not be able to truly reflect the quality of the entire batch of grain, affecting the accuracy of the detection results. Acoustic detection can estimate the types and densities of insects in stored grains by monitoring the sounds emitted by insect movement and feeding. This method requires separating the specific frequency sounds of pests from the background sound signal, and the background sound signal is greatly affected by environmental noise in the actual grain storage scenario. High cost and sensor sensitivity also limit the applicability of acoustic devices.

[0004] The intelligent grain pest detection device with the publication number CN202583073U for image processing provides a new type of remote automatic diagnosis device for grain borers, mainly including a controller and several CCD cameras. Among them, the rear of the CCD camera is connected to a multiplexer and connected to the controller through a band - pass filter. At the same time, an amplification and shaping circuit is also provided between the band - pass filter and the multiplexer; and the controller is connected to an input element in addition to connecting to an external display system.

[0005] The patent application with the publication number CN118887388A discloses a warehouse grain pest detection method based on the YOLOv8m and ShuffleNetv2 networks, including the following steps: data collection and pre - processing; constructing a model for warehouse grain pest detection; training the warehouse grain pest detection model; obtaining the image of the warehouse grain pests to be detected; obtaining the warehouse grain pest detection results.

[0006] Compared with traditional methods, deep learning is efficient, accurate, and automated. In this context, the introduction of deep learning technology provides an efficient solution for pest detection, which is a problem worthy of research. Summary of the Invention

[0007] To address the deficiencies in the above-mentioned existing technologies, the purpose of the present invention is to provide a lightweight storage grain pest detection method based on YOLOv8n, which is used to improve the detection accuracy of small pests in the granary environment and reduce the demand for computing resources; to avoid the situation of untimely discovery of grain pests and difficulty in efficient early warning, especially in complex environments and dense grain storage scenarios, and to provide more accurate and timely pest monitoring capabilities; to improve the detection speed and accuracy, reduce the computational complexity, ensure that the system can still operate efficiently in resource-constrained environments, be able to issue pest early warnings in a timely manner, and provide support for intelligent pest control.

[0008] The purpose of the present invention is achieved as follows:

[0009] A storage grain pest detection method based on an improved YOLOv8n network, comprising the following steps:

[0010] S1: Data collection and preprocessing: Collect 5 common storage grain pest images through a data collection method, and perform annotation, data augmentation, and dataset division operations on the collected storage grain pest images to form a storage grain pest image dataset;

[0011] S2: Model architecture selection: Train current mainstream object detection models, and then evaluate the models on the test set, and select the optimal model based on the comparison of evaluation metrics;

[0012] S3: Training of the storage grain pest detection model: Build a deep learning training environment, based on the YOLOv8n network, introduce a lightweight feature extraction module, a detail-enhanced YOLO detection head module, and a scale-internal feature interaction module, and reconstruct the neck network to build an improved storage grain pest detection model; Input the storage grain pest image dataset into the improved storage grain pest detection model for training to generate the optimal weight file of the model;

[0013] S4: Model evaluation: Based on the optimal weight file obtained in step S3, perform model evaluation and analysis;

[0014] S5: Acquisition of the image to be detected: Obtain the image to be detected from the image captured by the camera or the image uploaded locally;

[0015] S6: Generate detection results: Input the image to be detected into the improved storage grain pest detection model to generate detection results;

[0016] The training of the storage grain pest detection model in step S1 includes the following steps:

[0017] S1.1: Collect 5 common live samples of stored - grain insects, including Rhizopertha dominica, Tribolium castaneum, Plodia interpunctella, Sitophilus zeamais, and Sitotroga cerealella;

[0018] S1.2: Use a smartphone to take 2000 images of stored - grain insects;

[0019] S1.3: Enhance the image data using four methods: brightness transformation, Gaussian noise, horizontal transformation, and vertical transformation;

[0020] S1.4: Use the labelimg tool to annotate the enhanced images, and the annotation format is yolo;

[0021] S1.5: Divide the annotated images and annotation files into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0022] The model - architecture selection in the S2 step includes the following steps:

[0023] S2.1: Select the current mainstream object - detection models;

[0024] S2.2: Configure the training environment. In terms of hardware, use an RTX A4000 GPU and an Intel(R)Xeon(R)Gold 5318Y@2.10GHz CPU. In terms of software, use the Ubuntu20.04 system, and install the Python language, CUDA components, and third - party libraries required for deep learning;

[0025] S2.3: Configure the hyperparameters required for training. Set the number of iterations to 16, set the optimizer to SGD, reshape the input - image size to 640*640, and the rest are default configurations;

[0026] S2.4: Train the selected mainstream object - detection models;

[0027] S2.5: Evaluate each index of the selected mainstream object - detection models;

[0028] S2.6: Select the optimal model according to the obtained indexes.

[0029] The training of the stored - grain insect detection model in the S3 step includes the following steps:

[0030] S3.1: Try various methods to improve the YOLOv8n model and obtain the optimal improvement combination: By introducing an optimized lightweight feature - extraction module, improve the efficiency and accuracy of feature extraction; By enhancing the YOLO detection - head module in details, reduce parameter redundancy and improve the ability to capture small - target details; Combine the scale - internal feature interaction module to enhance the interaction between features, reconstruct the neck part to increase the detection ability for small targets, and reduce the computational overhead;

[0031] S3.2: During the training process, save the checkpoint weights generated in each round of training and perform validation after each round of training to confirm whether the model is fitted.

[0032] S3.3: Test the optimal weight file of the model after training is completed and obtain the optimal improvement combination: by introducing an optimized lightweight feature extraction module, improve the efficiency and accuracy of feature extraction; by using a detail-enhanced YOLO detection head module, reduce parameter redundancy and improve the ability to capture details of small targets; combine the intra-scale feature interaction module to enhance the interaction between features, reconstruct the neck part to increase the detection ability of small targets, and reduce the computational overhead.

[0033] Based on the warehouse grain pest detection model constructed based on the YOLOv8n network in the step S3, the warehouse grain pest detection model includes five parts: an input end, a backbone part, a neck part, a head part, and an output end.

[0034] The first part is the input. The input end is responsible for receiving image data from the front end and performing basic preprocessing on the data. Specifically, the basic image size of the input dataset is 640*640, perform image scaling and Mosaic data augmentation operations, and turn off data augmentation in the final training stage to improve the generalization ability of the model.

[0035] The second part is the backbone. The backbone is responsible for extracting features of the input image. It uses an intra-scale feature interaction module to replace the spatial pyramid module in the original YOLOv8n to enhance the detailed capture ability between different feature scales.

[0036] The third part is the neck. The neck part adopts a path aggregation network structure and adds a small target detection layer to enhance the detection ability of small targets. At the same time, the large target detection layer for large target detection is removed to reduce the computational complexity and improve the detection accuracy of small targets.

[0037] The fourth part is the detection layer. It uses a detail-enhanced YOLO detection head module to replace the detection head of the original YOLOv8n. By sharing the convolutional layer structure, reduce parameter redundancy, and improve the detail capture ability of small target detection through detail-enhanced convolution; and perform classification and regression. The calculation of the loss function includes classification loss and regression loss.

[0038] The fifth part is the output, which is used to output the results of classification and regression, including the category prediction of pests and the bounding box regression. By scaling and adjusting the feature map, ensure that the model can maintain good detection performance for targets of different scales.

[0039] Furthermore, set the improvement method in the step S3.1:

[0040] Furthermore, replace the C2f module in the network with a lightweight feature extraction module, called the FasterEMA_C2f (FE_C2f) module; the FE_C2f module is an improvement on the original C2f module in YOLOv8n, mainly combining the Faster_Block and the efficient multi-scale attention mechanism (Efficient Muti-Scale Attention Module, EMA), aiming to reduce redundant calculations and improve the ability to capture pixel-level details; the Faster_Block reduces the computational complexity through partial convolution (PartialConvolution, PConv), while maintaining the feature representation ability. The computational process of PConv can be expressed as the following formula:

[0041] x1,x2=split(x,[dim_conv3,dim_untouched]) (1)

[0042] x'1=Conv 3×3 (x1) (2)

[0043] y=cat(x'1,x2) (3)

[0044] In the formula, x is the feature map input to PConv; dim_conv3 and dim_untouched respectively represent the number of channels of the two parts of the input feature map; x1 and x2 respectively represent the feature map after convolution processing and the feature map without convolution processing after splitting; x'1 represents the feature map part after the 3×3 convolution operation; y represents the feature map after concatenating x'1 and x2, which is the output of PConv.

[0045] Therefore, PConv effectively reduces the computational cost by reducing the number of channels of the convolution operation; its computational amount can be described by the following formula:

[0046]

[0047] In the formula, h and w are the height and width of the feature map, k is the convolution kernel size, C p is the number of channels of the part to be convolved.

[0048] Since C pIt is only 1 / 4 of that of C. Compared with conventional convolution, the computational cost of PConv is only 1 / 16 of that of conventional convolution. At the same time, by keeping some channels un-convolved, PConv can still make full use of the information of all channels in the subsequent 1×1 convolution operation to achieve efficient spatial feature extraction. PConv enables Faster_Block to obtain a higher number of floating-point operations per second (FLOPS) while ensuring a lower FLOPs, which is beneficial for deployment on resource-constrained devices. The multi-scale attention mechanism enhances the ability to capture details at different scales by combining information of different convolution kernel sizes, enabling the model to more effectively detect small targets in complex backgrounds.

[0049] Furthermore, replace the detection head in the network with a Detail-Enhanced YOLO Head (DEYH) module: The Detail-Enhanced YOLO Head module improves the head part in YOLOv8n. By introducing a shared Detail-Enhanced Convolution (DEConv) module, it significantly reduces parameter redundancy and improves the detection ability for small targets. First, a 1x1 convolution is used for channel reduction in each detection layer. Subsequently, the SiLU activation function is introduced. SiLU can better retain fine-grained information and enhance the robustness of the model when dealing with small targets by retaining negative values and having a smooth activation curve. The SiLU activation function can be expressed as the following formula:

[0050]

[0051] In the formula, x is the input feature value.

[0052] The introduction of the DEConv module and group normalization operations aims to improve the feature expression ability by enhancing the high-frequency information of the image (such as edges and contours), and also enhance the stability and adaptability of the enhanced feature expression. The DEConv module includes central difference convolution, angular difference convolution, horizontal and vertical difference convolution, etc., which can capture detail features in different directions. In addition, through reparameterization technology, DEConv simplifies parallel convolution into standard convolution to ensure that no additional computational overhead is added during the inference stage. Specifically, this process is expressed as the following formula:

[0053]

[0054] In the formula, F in represents the input feature map, and K i are parallel convolution kernels respectively, and K cvt represents the equivalent convolution kernel after reparameterization.

[0055] It should be noted that the detail-enhanced convolution realizes parameter sharing among multiple detection layers; when each detection layer extracts features, the same DEConv module is used for convolution operations, thereby reducing redundant computational amount and ensuring cross-level feature consistency; this design not only improves the efficiency of the model but also ensures accurate detection of targets on multi-scale feature maps; after DEConv processing, the SiLU activation function is used to process the feature map again; then, the number of channels is restored to the original state through 1x1 convolution; finally, the feature map is input into the Classification and Regression Branches (CRB), which contains two paths: the Conv_Cls module is used for class prediction, while the Conv_Reg module is used for bounding box regression; in order to alleviate the negative effects that may be brought by shared parameters and ensure the performance of multi-scale object detection, a scaling operation is introduced after Conv_Reg; the scaling operation ensures that the model can maintain good detection performance for targets of different scales by appropriately scaling and adjusting the feature map.

[0056] Furthermore, the spatial pyramid module is replaced with an intra-scale feature interaction module: the Attention-based Intra-scale Feature Interaction (AIFI) module is an introduced attention mechanism module that combines the Transformer self-attention mechanism and two-dimensional position embedding, enhancing the interaction between features and the ability to retain spatial information; the intra-scale feature interaction module weights the features through the multi-head attention mechanism and combines position embedding, enabling the model to capture the dependencies between different features, thereby enhancing the saliency of small targets in complex backgrounds and reducing the cases of missed detection and false detection.

[0057] Furthermore, the neck network is reconstructed: the neck part adopts a reconstructed path aggregation network structure, adding small target detection layers for the detection of small targets and removing the large target detection layers for the detection of large targets; the small target detection layers extract information from high-resolution feature maps, making the model more sensitive to small targets and improving the detection accuracy and stability; while removing the large target detection layers reduces the redundant calculation for large targets, reducing the computational complexity and amount of the model, making the model more lightweight and efficient, and suitable for applications in resource-constrained scenarios.

[0058] The generation of the detection result in the step S6 includes the following steps:

[0059] S6.1: Preprocess the incoming image, reshape the image size to 640*640, and convert the image into a feature map;

[0060] S6.2: The feature map is fed into the network, and the backbone network extracts basic features, while the neck network fuses the features output by the backbone network;

[0061] S6.3: The detection head receives the output of the neck network, generates detection anchor boxes and class predictions, and finally the detection results are output by the output end.

[0062] Positive and beneficial effects:

[0063] By improving the YOLOv8n network, the present invention proposes a lightweight and high-precision method for detecting grain insects in warehouses, having the following technical effects and advantages:

[0064] High-precision detection: By introducing a lightweight feature extraction module and a detailed enhanced YOLO detection head module, the present invention effectively improves the detection accuracy of the model for small pests in the grain storage environment; experimental results show that the improved model is significantly superior to existing mainstream object detection models in terms of indicators such as accuracy, recall rate, and mean average precision.

[0065] Reducing computational complexity: By using lightweight module designs (such as a lightweight feature extraction module and an intra-scale feature interaction module), the present invention significantly reduces the number of model parameters and computational overhead while ensuring detection accuracy, making the model more suitable for deployment in resource-constrained environments such as edge devices or embedded systems.

[0066] Enhancing the small target detection ability: By adding a small target detection layer and deleting the large target detection layer, the model focuses more on capturing and detecting the features of small targets, improving the detection ability for small pests in complex backgrounds and dense grain storage environments, and reducing the cases of missed detection and false detection.

[0067] Adapting to resource-constrained environments: The design of this model is particularly suitable for running in resource-constrained environments; its number of parameters is only 1.3M, and the computational complexity is 4.7G, having a significant lightweight advantage compared to other complex object detection models, and is suitable for scenarios such as granaries and farms that require real-time detection.

[0068] Robustness and generalization ability: By using a variety of data augmentation techniques and an intra-scale feature interaction module, the present invention enhances the model's detection ability for different grain storage environments and different types of pests; the model shows good detection effects on pest images with complex backgrounds, illumination changes, and different perspectives in the experiment, having strong robustness and generalization ability.

[0069] In summary, the model of the present invention has significant advantages in terms of detection accuracy, computational efficiency, and resource requirements, and is particularly suitable for achieving precise detection and timely warning of pests in complex and dynamic grain storage environments, thus effectively ensuring food security. Description of the Drawings

[0070] Figure 1 This is a schematic diagram of the framework for the grain pest detection steps of the present invention;

[0071] Figure 2 This is the structural diagram of the detection model for stored - grain pests involved in the present invention. Detailed implementation manners

[0072] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments:

[0073] Embodiment 1

[0074] A method for detecting stored - grain pests based on an improved YOLOv8n network, comprising the following steps:

[0075] S1: Data collection and pre - processing: Five common types of stored - grain pest images are collected through a data collection method, and the collected stored - grain pest images are labeled, data - enhanced, and dataset - divided to form a stored - grain pest image dataset;

[0076] S2: Model architecture selection: Train current mainstream object - detection models, then evaluate the models on the test set, and select the optimal model based on the comparison of evaluation metrics;

[0077] S3: Training of the stored - grain pest detection model: Set up a deep - learning training environment, based on the YOLOv8n network, introduce a lightweight feature extraction module, a detail - enhanced YOLO detection head module, and a feature interaction module within the scale, and reconstruct the neck network to construct an improved stored - grain pest detection model. Input the stored - grain pest image dataset into the improved stored - grain pest detection model for training to generate the optimal weight file of the model;

[0078] S4: Model evaluation: Based on the optimal weight file obtained in step S3, conduct model evaluation and analysis;

[0079] S5: Acquisition of the image to be detected: Obtain the image to be detected from the image captured by the camera or the image uploaded locally.

[0080] S6: Generation of detection results: Input the image to be detected into the improved stored - grain pest detection model to generate detection results;

[0081] 1. The method for detecting stored - grain pests according to claim 1, characterized in that, based on YOLOv8n in step S3, a stored - grain pest detection model is constructed, and the stored - grain pest detection model is set to be divided into five parts.

[0082] (1) The input end is responsible for receiving image data from the front end and preprocessing the data. Specifically, it includes: the basic size of the input dataset pictures is 640*640, image scaling is performed, the Mosaic data augmentation operation is applied and turned off in the last 10 training cycles.

[0083] (2) The backbone part uses the optimized CSPDarknet network for feature extraction. In the main part, the original C2f module is replaced with a lightweight feature extraction module, and a scale-internal feature interaction module is combined to enhance the ability to capture fine details of different feature scales, thereby achieving more efficient feature fusion and extraction.

[0084] (3) The neck part is a reconstructed path aggregation network structure for feature hierarchy extraction. The multi-layer feature outputs of the backbone are used for multi-scale feature fusion. By combining the added small object detection layer, the small object detection ability is improved, and the large object detection layer is deleted so that network resources can be more concentrated on processing high-resolution feature maps.

[0085] (4) The head part uses a detail-enhanced YOLO detection head to improve the stability and detail capture ability of the model in small object detection through shared detail-enhanced convolutions. The calculation of the loss function includes classification loss and regression loss.

[0086] (5) The output part is responsible for outputting the model prediction results and identifying and locating the targets in the image.

[0087] Furthermore, the improvement method in step S3.1 is set as follows:

[0088] Furthermore, the C2f module in the network is replaced with a lightweight feature extraction module: The lightweight feature extraction module is an improvement on the original C2f module in YOLOv8n, mainly combining Faster_Block and the Efficient Muti-Scale Attention Module (EMA), aiming to reduce redundant calculations and improve the ability to capture pixel-level details. Faster_Block reduces the computational complexity through Partial Convolution (PConv) while maintaining the feature expression ability. The computational process of PConv can be expressed as the following formula:

[0089] x1,x2=split(x,[dim_conv3,dim_untouched]) (1)

[0090] x'1=Conv 3×3 (x1) (2)

[0091] y=cat(x'1,x2) (3)

[0092] In the formula, \(x\) is the feature map input to PConv; \(dim\_conv3\) and \(dim\_untouched\) respectively represent the number of channels of the two parts of the input feature map; \(x1\) and \(x2\) respectively represent the feature map after splitting for convolution processing and the feature map not undergoing convolution processing; \(x'1\) represents the part of the feature map after the \(3\times3\) convolution operation; \(y\) represents the feature map after concatenating \(x'1\) and \(x2\), which is the output of PConv.

[0093] Therefore, by reducing the number of channels of the convolution operation, PConv effectively reduces the computational cost. Its computational complexity can be described by the following formula:

[0094]

[0095] In the formula, \(h\) and \(w\) are the height and width of the feature map, \(k\) is the convolution kernel size, \(C\) p is the number of channels of the part to be convolved.

[0096] Since \(C\) p is only 1 / 4 of \(C\), compared with the conventional convolution, the computational complexity of PConv is only 1 / 16 of that of the conventional convolution. At the same time, by keeping some channels untouched by convolution, PConv can still make full use of the information of all channels in the subsequent \(1\times1\) convolution operation to achieve efficient spatial feature extraction. PConv enables Faster_Block to obtain a high FLOPS while ensuring a low FLOPs, which is beneficial for deployment on resource-constrained devices. The multi-scale attention mechanism (Efficient Muti-Scale Attention Module, EMA) enhances the ability to capture details at different scales by combining information of different convolution kernel sizes, enabling the model to more effectively detect small targets in complex backgrounds.

[0097] Furthermore, replace the detection head in the network with the Detail-Enhanced YOLO Head (DEYH) module: The Detail-Enhanced YOLO Head module improves the Head part in YOLOv8n. By introducing a shared Detail-Enhanced Convolution (DEConv) module, it significantly reduces parameter redundancy and improves the detection ability for small targets. In each detection layer, a \(1\times1\) convolution is first used for channel dimensionality reduction. Subsequently, the SiLU activation function is introduced. SiLU can better retain fine-grained information and improve the robustness of the model when dealing with small targets by retaining negative values and having a smooth activation curve. The SiLU activation function can be expressed by the following formula:

[0098]

[0099] Where x is the input eigenvalue.

[0100] We introduce the DEConv module and group normalization operation, aiming to improve the feature expression ability by enhancing the high-frequency information (such as edges and contours) of the image, and enhancing the stability and adaptability of the feature expression. The DEConv module includes central difference convolution, angular difference convolution, horizontal and vertical difference convolution, etc., which can capture detailed features in different directions. In addition, through reparameterization technology, DEConv simplifies parallel convolutions into standard convolutions, ensuring no additional computational overhead during the inference stage. Specifically, this process is expressed as the following formula:

[0101]

[0102] Where F in represents the input feature map, and K i are parallel convolution kernels respectively, and K cvt represents the equivalent convolution kernel after reparameterization.

[0103] It should be noted that the detail enhancement convolution realizes parameter sharing among multiple detection layers. When each detection layer extracts features, the same DEConv module is used for convolution operations, thereby reducing redundant computational amounts and ensuring cross-level feature consistency. This design not only improves the efficiency of the model but also ensures the accurate detection of targets on multi-scale feature maps. After DEConv processing, the feature map is processed by the SiLU activation function again. Then, the number of channels is restored to the original state through 1x1 convolution. Finally, the feature map is input into the Classification and Regression Branches (CRB), which contains two paths: the Conv_Cls module is used for class prediction, and the Conv_Reg module is used for bounding box regression. To alleviate the possible negative effects of shared parameters and ensure the performance of multi-scale object detection, a scaling operation is introduced after Conv_Reg. The scaling operation appropriately scales the feature map to ensure that the model can maintain good detection performance for targets of different scales.

[0104] Furthermore, the spatial pyramid module is replaced by the intra-scale feature interaction module: The (intra-scale feature interaction) module is an introduced attention mechanism module that combines the Transformer self-attention mechanism and two-dimensional position embedding, enhancing the interaction between features and the ability to retain spatial information. The intra-scale feature interaction module weights the features through the multi-head attention mechanism and combines the position embedding, enabling the model to capture the dependency relationships between different features, thereby enhancing the saliency of small targets in complex backgrounds and reducing the cases of missed detection and false detection.

[0105] Furthermore, reconstruct the neck network: The neck part adopts a reconstructed path aggregation network structure, adds a small object detection layer for small object detection, and removes the large object detection layer for large object detection. The small object detection layer extracts information from the high-resolution feature map, making the model more sensitive to small objects and improving the detection accuracy and stability. Removing the large object detection layer reduces the redundant calculation for large objects, lowers the computational complexity and amount of calculation of the model, makes the model more lightweight and efficient, and is suitable for applications in resource-constrained scenarios.

[0106] Among them, in order to verify the accuracy of the warehouse grain pest detection and recognition of the present invention, the obtained weight file is used to test the warehouse grain pest images in the test set; at the same time, in order to verify the superiority and effectiveness of the detection algorithm proposed by the present invention compared with the current popular object detection models, evaluate the performance of the detection model in the mainstream networks, and select the Faster R-CNN, SSD, RT-DETR, YOLOv5s, YOLOv8n, YOLOv7-tiny, YOLOv8s, YOLOv9t and YOLOv10n algorithms to conduct experimental comparisons under the same conditions (unified environment configuration and parameter settings and the same data set); to confirm the effectiveness of the improvements in this study, ablation experiments are carried out. The improvement methods of this study are added to the original YOLOv8n model one by one. '√' indicates that the module is added to the model, and '-' indicates that the module is not added to the model. The environment configuration and parameter settings remain consistent in the ablation experiments.

[0107] Table 1 Comparison results of the detection performance of different models

[0108]

[0109] According to the experimental results, the improved model designed by the present invention shows significant advantages in the accuracy index. Specifically, the improved YOLOv8n achieves 97.5% and 60.0% respectively in mAP@0.5 and mAP@0.5:0.95, exceeding other mainstream models and demonstrating its excellent detection ability. At the same time, while maintaining high detection accuracy, the number of parameters is only 1.3M and the FLOPs is 4.7G, showing excellent lightweight and computational efficiency advantages, and is very suitable for resource-constrained application environments.

[0110] Table 2 Results of ablation experiments

[0111]

[0112] First, the impact of introducing a small object detection layer and removing the large object detection layer on performance was explored. The experimental results showed that when the small object detection layer was used for small object detection, the precision increased to 92.5%, while mAP@0.5 and mAP@0.5:0.95 increased to 95.6% and 57.4% respectively, and the recall rate slightly decreased to 91.1%. However, this change significantly increased the computational requirements, increasing the FLOPs to 11.6G, indicating that the introduction of the small object detection layer can improve the small object detection ability, but the computational resources need to be weighed. Next, after adding a lightweight feature extraction module, the precision, recall rate, and mAP@0.5 were further improved, reaching 93.8%, 91.9%, and 96.7% respectively, and mAP@0.5:0.95 increased to 58.9%. At the same time, the number of parameters and FLOPs were reduced to 1.5M and 9.9G respectively. The FE_C2f module improved the detection accuracy and effectively reduced the computational redundancy by optimizing feature extraction. After further introducing the DEYH module, the accuracy was significantly improved to 95.1%, the recall rate, mAP@0.5, and mAP@0.5:0.95 remained at a high level, while the FLOPs were significantly reduced to 4.5G and the number of parameters decreased to 1.3M. This indicates that the detailed enhanced YOLO detection head module significantly reduces the computational overhead and optimizes the computational efficiency of the model by sharing convolutional parameters while maintaining high detection performance. Finally, the improved YOLOv8n model combining the small object detection layer, FE_C2f module, DEYH module, and AIFI module demonstrated the best comprehensive performance. The precision, recall rate, mAP@0.5, and mAP@0.5:0.95 reached 95.0%, 93.6%, 97.5%, and 60.0% respectively, and the FLOPs remained at 4.7G and the number of parameters was 1.3M, fully demonstrating the effectiveness of the collaborative optimization of each module.

[0113] Working principle of the present invention:

[0114] The working principle of the present invention is based on the deep learning object detection framework of the YOLOv8n network, and a series of optimization modules are introduced to improve the recognition ability of small objects and reduce the use of computational resources. The specific working principle is as follows:

[0115] 1. Data preprocessing and input: First, the image data of stored grain insects is collected and preprocessed through step S1, including operations such as image scaling and data augmentation to improve the diversity of data and the generalization ability of the model. The processed image data is input into the input part of the model in a fixed size (640*640).

[0116] 2. Feature Extraction (Backbone Part): After the input part receives the image data, it enters the backbone part for feature extraction. The backbone uses a convolutional module, a lightweight feature extraction module, and an intra-scale feature interaction module to capture the detailed features in the pest image by enhancing different feature scales. The intra-scale feature interaction module combines the self-attention mechanism to enhance the model's ability to capture the internal relationships of features, thereby improving the accuracy of detecting small target pests.

[0117] 3. Feature Fusion and Small Target Detection (Neck Part): Next, the feature map undergoes feature fusion through the path aggregation network structure in the neck part. The path aggregation network structure effectively integrates the information of different feature layers through path aggregation, thereby improving the representation ability of the feature pyramid. In particular, the neck part adds a small target detection layer to enhance the detection of small targets, and at the same time deletes the large target detection layer used for large target detection to concentrate the model's computing resources on small targets, further enhancing the detection ability for small pests.

[0118] 4. Detail-Enhanced YOLO Detection Head Module: After being processed by the neck part, the feature map is passed to the detail-enhanced YOLO detection head module. This module enables the model to better identify and locate pests, improving the detection accuracy and robustness by introducing the SiLU activation function and the detail-enhanced convolutional module. And it performs class prediction and bounding box regression.

[0119] 5. Output Classification and Regression: Finally, the feature map processed by the detail-enhanced YOLO detection head module is passed to the output part for class prediction and bounding box regression. The classification and regression operations in the output part are used to determine the class and location of the pests, and output the final results of pest detection, including the coordinates of its location, the class, and the corresponding confidence score.

[0120] 6. Lightweight and Efficient Detection: By using the lightweight feature extraction module, unnecessary redundant convolutional operations are reduced during the calculation process, thereby reducing the computational cost of the model. At the same time, through the collaborative optimization between modules, the improved model significantly reduces the number of parameters and computational complexity while maintaining high accuracy, making it suitable for deployment on resource-constrained devices to achieve real-time pest detection.

[0121] In summary, the working principle of the present invention is based on the modular improvement of the traditional YOLOv8n network. Through means such as multi-scale feature enhancement, lightweight design, and detail enhancement, it effectively improves the accuracy of warehouse grain pest detection and significantly reduces the demand for computing resources, enabling it to work efficiently in complex and dynamic grain storage environments.

[0122] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting stored grain pests based on an improved YOLOv8n network, characterized in that, It includes the following steps: S1: Data collection and preprocessing: Five common images of stored - grain insects are collected through a data collection method. Labeling, data augmentation, and dataset division operations are performed on the collected images of stored - grain insects to form a dataset of images of stored - grain insects; S2: Model architecture selection: Train current mainstream object - detection models, and then evaluate the models on the test set. Select the optimal model based on the comparison of evaluation metrics; S3: Training of the stored - grain insect detection model: Build a deep - learning training environment. Based on the YOLOv8n network, introduce a lightweight feature extraction module, a detail - enhanced YOLO detection head module, and a feature interaction module within the scale, and reconstruct the neck network to construct an improved stored - grain insect detection model; Input the dataset of images of stored - grain insects into the improved stored - grain insect detection model for training to generate the optimal weight file of the model; S4: Model evaluation: Based on the optimal weight file obtained in step S3, conduct model evaluation and analysis; S5: Acquisition of the image to be detected: Obtain the image to be detected from the image captured by the camera or the image uploaded locally; S6: Generation of the detection result: The image to be detected is input into the improved stored - grain insect detection model to generate the detection result.

2. The method for detecting stored grain pests based on the improved YOLOv8n network according to claim 1, characterized in that, Based on the stored - grain insect detection model constructed based on the YOLOv8n network in step S3, the stored - grain insect detection model includes an input end, a backbone part, a neck part, a head part, and an output end.

3. The method for detecting stored - grain insects based on the improved YOLOv8n network according to claim 2, characterized in that, The input end includes: The input end is responsible for receiving image data from the front end and preprocessing the data; The backbone part uses an optimized CSPDarknet network for feature extraction. The main part replaces the original C2f module with a lightweight feature extraction module and combines a feature interaction module within the scale to enhance the ability to capture fine details of different feature scales, thereby achieving more efficient feature fusion and extraction; The neck part is a reconstructed path aggregation network structure for feature hierarchical structure extraction. The multi - layer feature outputs of the backbone network are used for multi - scale feature fusion. Combining the added small - object detection layer improves the small - object detection ability, and deleting the large - object detection layer enables network resources to be more concentrated on processing high - resolution feature maps; The head part uses a detail - enhanced YOLO detection head to improve the stability and detail - capturing ability of the model in small - object detection through shared detail - enhanced convolutions. The calculation of the loss function includes classification loss and regression loss; The output end is responsible for outputting the model prediction result and identifying and locating the targets in the image.

4. The method for detecting stored - grain insects based on the improved YOLOv8n network according to claim 1, wherein: In step S3.3, the method for network improvement is as follows: By introducing an optimized lightweight feature extraction module, the efficiency and accuracy of feature extraction are improved; By using the detail - enhanced YOLO detection head module, parameter redundancy is reduced and the ability to capture small - object details is improved; Combining the feature interaction module within the scale enhances the interaction between features, reconstructing the neck part to increase the small - object detection ability, and reducing the computational overhead.

5. The method for detecting stored - grain insects based on the improved YOLOv8n network according to claim 4, characterized in that: An optimized lightweight feature extraction module, a detail - enhanced YOLO detection head module, and an intra - scale feature interaction module are introduced, and the neck network is reconstructed: Lightweight feature extraction module: The lightweight feature extraction module is an improvement on the original C2f module in YOLOv8n. It mainly combines the Faster_Block and the efficient multi - scale attention mechanism (EMA), aiming to reduce redundant calculations and improve the ability to capture pixel - level details; Faster_Block reduces the computational complexity through partial convolution (PConv) while maintaining the feature representation ability. The computational process of PConv can be expressed as the following formula: x1,x2=split(x,[dim_conv3,dim_untouched]) (1) x1' = Conv 3×3 (x1) (2) y=cat(x1',x2) (3) In the formula, x is the feature map input to PConv; dim_conv3 and dim_untouched respectively represent the number of channels of the two parts of the input feature map; x1 and x2 respectively represent the feature map after convolution processing and the feature map without convolution processing after splitting; x1' represents the part of the feature map after the 3×3 convolution operation; y represents the feature map after splicing x1' and x2, which is the output of PConv; Therefore, PConv effectively reduces the computational cost by reducing the number of channels of the convolution operation, and its computational volume (Floating - Point Operations, FLOPs) can be described by the following formula: Where h and w are the height and width of the feature map, k is the convolutional kernel size, and C p is the number of partial channels to be convolved; Since C p is only 1 / 4 of C, compared with conventional convolution, the computational cost of PConv is only 1 / 16 of that of conventional convolution; at the same time, by keeping some channels unconvolved, PConv can still make full use of the information of all channels in the subsequent 1×1 convolution operation to achieve efficient spatial feature extraction; PConv enables Faster_Block to obtain a higher floating-point operations per second while ensuring a lower computational cost, which is beneficial for deployment on resource-constrained devices; the multi-scale attention mechanism enhances the ability to capture details at different scales by combining information of different convolution kernel sizes, enabling the model to more effectively detect small targets in complex backgrounds; Detail - enhanced YOLO Head (DEYH) module: The detail - enhanced YOLO detection head module improves the head part in YOLOv8n. By introducing a shared detail - enhanced convolution (DEConv) module, it improves the detection ability for small targets; first, a 1x1 convolution is used for channel reduction in each detection layer; subsequently, the SiLU activation function is used to improve the reliability of the model when processing small targets. SiLU can better retain fine - grained information by retaining negative values and having a smooth activation curve. The SiLU activation function can be expressed as the following formula: In the formula, x is the input feature value; The DEConv module and group normalization operation are introduced, aiming to improve the feature representation ability by enhancing the high - frequency information of the image and enhancing the stability and adaptability of the enhanced feature representation; the DEConv module captures detail features in different directions through central difference convolution, angular difference convolution, horizontal and vertical difference convolution. DEConv simplifies the parallel convolution to a standard convolution through re - parameterization; specifically, this process is expressed as the following formula: where F in represents the input feature map, and K i are parallel convolutional kernels respectively, and K cvt represents the equivalent convolutional kernel after reparameterization; The detail-enhanced convolution realizes parameter sharing among multiple detection layers; when each detection layer extracts features, the same DEConv module is used for convolution operations, thus reducing redundant computational volume and ensuring cross-level feature consistency; this design not only improves the efficiency of the model but also ensures the accurate detection of targets on multi-scale feature maps; after DEConv processing, the feature map is processed again using the SiLU activation function; then, the number of channels is restored to the original state through 1x1 convolution; finally, the feature map is input into the Classification and Regression Branches (CRB), which contains two paths: the Conv_Cls module is used for class prediction, and the Conv_Reg module is used for bounding box regression; to alleviate the possible negative effects brought by shared parameters and ensure the performance of multi-scale object detection, a scaling operation is introduced after Conv_Reg; the scaling operation ensures that the model can maintain good detection performance for objects of different scales by appropriately scaling and adjusting the feature map. Intra-scale feature interaction module: The intra-scale feature interaction module weights features through the multi-head attention mechanism and combines position embeddings, enabling the model to capture the dependencies between different features. Reconstructed neck network: The neck network adopts the structure of the reconstructed path aggregation network, adding small-object detection layers and removing large-object detection layers; the small-object detection layers extract information from high-resolution feature maps, making the model more sensitive to small objects and improving the detection accuracy and stability; while removing the large-object detection layers reduces the redundant computation for large objects, lowering the computational complexity and volume of the model, making the model more lightweight and efficient, and suitable for applications in resource-constrained scenarios.

6. The method for detecting stored grain insects based on the improved YOLOv8n network according to claim 1, characterized in that, The data acquisition and preprocessing in step S1 include the following steps: S1.1: Collect 5 common live samples of stored-grain insects, including Rhizopertha dominica, Tribolium castaneum, Plodia interpunctella, Sitophilus zeamais, and Sitotroga cerealella. S1.2: Use a smartphone to take 2000 images of stored-grain insects. S1.3: Enhance the image data using four methods: brightness transformation, Gaussian noise, horizontal transformation, and vertical transformation. S1.4: Use the labelimg tool to annotate the enhanced images, and the annotation format is yolo. S1.5: Divide the annotated images and annotation files into a training set, a validation set, and a test set according to the ratio of 7:2:

1.

7. The method for detecting stored grain insects based on the improved YOLOv8n network according to claim 1, characterized in that, The model architecture selection in step S2 includes the following steps: S2.1: Select the current mainstream object detection model. S2.2: Configure the training environment. In terms of hardware, it is an RTX A4000 GPU and an Intel(R) Xeon(R) Gold 5318Y @ 2.10GHz CPU. In terms of software, use the Ubuntu20.04 system and install the Python language, CUDA components, and third-party libraries required for deep learning. S2.3: Configure the hyperparameters required for training. Set the number of iteration rounds to 16, set the optimizer to SGD, reshape the input image size to 640*640, and keep the rest as default configurations; S2.4: Train the selected mainstream object detection model; S2.5: Evaluate various metrics of the selected mainstream object detection model; S2.6: Select the optimal model according to the obtained metrics.

8. The method for detecting stored grain insects based on the improved YOLOv8n network according to claim 1, wherein The training of the warehouse grain pest detection model in step S3 includes the following steps: S3.1: Try various methods to improve the YOLOv8n model and obtain the optimal improvement combination: by introducing an optimized lightweight feature extraction module, improve the efficiency and accuracy of feature extraction; reduce parameter redundancy and improve the ability to capture small target details by enhancing the YOLO detection head module in details; combine the scale-in feature interaction module to enhance the interaction between features, reconstruct the neck part to increase the detection ability for small targets, and reduce the computational overhead; S3.2: Save the checkpoint weights generated in each round of training during the training process and perform validation after each round of training to confirm whether the model is fitted; S3.3: Test the optimal weight file of the model after training is completed and obtain the optimal improvement combination: by introducing an optimized lightweight feature extraction module, improve the efficiency and accuracy of feature extraction; reduce parameter redundancy and improve the ability to capture small target details by enhancing the YOLO detection head module in details; combine the scale-in feature interaction module to enhance the interaction between features, reconstruct the neck part to increase the detection ability for small targets, and reduce the computational overhead.

9. The method for detecting stored grain insects based on the improved YOLOv8n network according to claim 1, characterized in that, The generation of detection results in step S6 includes the following steps: S6.1: Preprocess the incoming image, reshape the image size to 640*640, and convert the image into a feature map; S6.2: The feature map is fed into the network, and the backbone network extracts basic features, and the neck network fuses the features output by the backbone network; S6.3: The detection head receives the output of the neck network, generates detection anchor boxes and class predictions, and finally the output end outputs the detection results.

Citation Information

Patent Citations

  • Wheat storage grain insect detection method based on YOLOv8m and ShuffleNetv2 network

    CN118887388A

  • Intelligent grain bristletail detection device based on image processing

    CN202583073U