Insect situation monitoring method and system based on improved YOLOv10n
By improving the network structure and training process of the YOLOv10n model and combining edge computing equipment, high-precision detection and automated monitoring of small target pests and overlapping pests are achieved, solving the problems of low efficiency and difficulty in automation of traditional insect situation monitoring methods.
Patent Information
- Application Number
- CN202510113777.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional insect situation monitoring methods are inefficient and subjective, and the existing deep learning algorithms are insufficient in detecting small-target pests and dense targets, making it difficult to achieve automated and efficient operations in field environments.
Based on the improved YOLOv10n insect situation monitoring method, the SPD-Conv module, iRMB inverted residual block attention mechanism and Inner-SIoU loss function are introduced to improve the pest recognition accuracy, and combined with edge computing equipment to realize real-time detection and automated sticky insect board replacement.
High-precision detection of small target pests and overlapping pests is achieved, and the problems of low detection accuracy and inability to achieve automation are solved, which significantly improves the model training effect and robustness, and provides an intelligent and efficient solution for field insect situation monitoring.
Smart Images

Figure CN120088725A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pest monitoring equipment, and particularly to a pest situation monitoring method and system based on improved YOLOv10n. Background Art
[0002] The sticky pest board is a tool for trapping pests by utilizing the tendency of pests towards specific colors and odors, and is widely used for pest control in the fields. However, the traditional pest situation monitoring method relies on manual identification and recording of pests on the sticky pest board. This method has low efficiency, strong subjectivity and requires technicians to have high professional knowledge. In recent years, object detection technologies based on deep learning, such as the YOLO series algorithms, have been applied to pest identification. However, these algorithms have problems with insufficient accuracy when detecting small target pests and densely and overlappingly distributed targets on the sticky pest board. In addition, traditional systems are difficult to meet the requirements of automated and efficient operation in the field environment. Summary of the Invention
[0003] Aiming at the many problems existing in the above-mentioned prior art, the present invention provides a pest situation monitoring method and system based on improved YOLOv10n. Based on the improved YOLOv10n object detection model, the present invention improves the pest identification accuracy by optimizing the network structure and training process. Combining with edge computing devices, pest images are collected and detected in real time, and the replacement operation of the sticky pest board is automatically controlled to ensure the continuity and accuracy of the monitoring process, realizing intelligent monitoring and data collection of the pest situation in the fields.
[0004] As Figure 1 shown, a pest situation monitoring method based on improved YOLOv10n includes the following steps:
[0005] Obtain pest images on the sticky pest board, and preprocess the obtained images to generate an image dataset containing pest category and location information;
[0006] Obtain pest images on the field sticky pest board, specifically including: placing yellow sticky pest boards at different intervals in the field, and then using an image acquisition device to take pictures of the pests on the sticky pest board every 1 to 2 days. The pictures contain multiple categories of pests on the sticky pest board.
[0007] After preprocessing the acquired images, they are labeled and a dataset of pest images on sticky traps is established, including: preprocessing the acquired pest images. First, images containing other non-pest impurities, damaged sticky traps, and unclear images are removed, and the remaining images are uniformly cropped to the same size. Then, they are labeled, and the labeling information includes the category information and location information of the pests. After that, data augmentation is performed on the labeled pest images, including operations such as vertical flipping, scaling, rotation, shearing, adjusting brightness, adjusting saturation, and adding noise to enhance the generalization and robustness during model training. Finally, the pest image set after data augmentation is divided into a training set, a validation set, and a test set. During the division, randomness is ensured and the distribution of pests of each category in the training set, validation set, and test set is made uniform to avoid the model not being able to learn the characteristics of pests well during training and affecting the training results.
[0008] Improve the network structure of the YOLOv10n model. The improvements include introducing the SPD-Conv module to replace the convolutional stride and pooling operations, adding the iRMB inverted residual block attention mechanism, and using the Inner-SIoU loss function to improve the detection performance for small targets and overlapping targets;
[0009] Improving the network structure of the YOLOv10n model includes:
[0010] Use the SPD-Conv module to replace the convolutional stride and pooling operations of the Yolov10n model, so that the model can avoid information loss in the downsampling method;
[0011] Add the iRMB inverted residual block attention mechanism in the Yolov10n model to balance the model's attention to local detail features and global semantic features;
[0012] Use the Inner-SIoU loss function in the Yolov10n model to solve the problems of missed detection and false detection of small targets and overlapping targets by the model.
[0013] Preferably, improving the network structure of the YOLOv10n model includes using the SPD-Conv module to replace the convolutional stride and pooling operations. The SPD-Conv module performs convolutional operations after dividing the input feature map into multiple small blocks and gradually rearranging them, so as to achieve downsampling without losing detail information.
[0014] The SPD-Conv module consists of a space-to-depth layer (SPD) and a Conv layer (non-strided convolution), and can be applied to most CNN structures. SPD-Conv achieves its goal by replacing the convolutional stride and pooling layer of the original model and introducing a convolutional layer with a stride of 1, completely abandoning the currently widely used convolutional stride and pooling operations, and downsampling the feature map without losing learnable information.
[0015] Preferably, the SPD-Conv module performs a space-to-depth conversion on the input feature map and applies a convolutional kernel to each small block, generating a feature map that reduces the number of channels while maintaining the spatial resolution, so as to more efficiently extract the spatial features of small targets.
[0016] The SPD layer rearranges the pixel blocks into the depth dimension through a space-to-depth conversion operation, increasing the depth of the feature map and retaining all the information in the channel dimension, thus avoiding information loss in traditional sampling methods. Then, the convolutional layer with a stride of 1 processes each region of the feature map, thus maximizing the retention of image information and generating a richer feature representation, reducing the number of channels of the feature map without losing details and effectively processing the features. Through this mechanism, SPD-Conv retains more information in the feature extraction stage, thus improving the model's recognition ability for small target images and being very suitable for the detection of small target pests on yellow sticky boards. The core principle of SPD-Conv includes the following steps: (a) The input feature map has a height, width, and number of channels C 1 ; (b) The SPD layer rearranges the pixel blocks into the depth dimension through a space-to-depth conversion operation, expanding the number of channels to 4C 1 , while reducing the spatial resolution by half; (c) Merging different channel groups in the channel dimension; (d) Adding the merged feature map to other feature maps; (e) Finally, applying a convolutional operation with a stride of 1 to the resulting feature map, reducing the number of channels to C while keeping the spatial resolution unchanged 2 .
[0017] Preferably, the improvement of the network structure also includes introducing the iRMB inverted residual block attention mechanism. The iRMB inverted residual block extracts features by first reducing the number of channels and then expanding the number of channels, and combines an attention weight distribution strategy in the expansion stage to strengthen the feature expression of key regions.
[0018] The iRMB inverted residual block attention mechanism adopts the design concept of the inverted residual block (IRB), extends the IRB of the traditional CNN to an attention-based structure, integrates the efficiency of the CNN architecture in capturing local features and the long-range interaction ability of the Transformer architecture in dynamic modeling, considers the entire input space when extracting features, and is effectively helpful for the detection of global targets. The iRMB combines depthwise separable convolution and self-attention mechanism. The 1x1 convolution is used for channel compression and expansion to optimize the computational efficiency, the depthwise separable convolution is used to capture spatial features, and the attention mechanism is used to capture the global dependencies between features. In addition, the Meta-Mobile Block is abstracted to utilize multiple expansion ratios and efficient operators, realizing a modular design. In a pluggable manner, different operations are integrated into a unified framework, making the model more flexible. The iRMB improves the model's attention to pests in dense areas and is also suitable for deploying the model on edge mobile devices for pest detection.
[0019] Preferably, the attention weight allocation strategy generates two sets of weight tensors by calculating the global average value and the maximum value of each channel, and then performs channel-wise weighted fusion to highlight the features of specific regions, thereby enhancing the model's ability to recognize target regions in complex backgrounds.
[0020] The attention weight allocation strategy aims to enhance the model's ability to express the features of target regions in complex backgrounds. Specifically, this strategy performs global operations on the feature maps of each channel, calculates the global average value and the maximum value of the channels respectively. These two statistical information respectively reflect the overall characteristics of the feature maps and the characteristics of the local most significant regions. According to these statistical information, two sets of weight tensors are generated, respectively representing the importance of the channels. Subsequently, fusion processing is performed on these two sets of weight tensors, and the weights are applied to the original feature maps using the channel-wise weighted method, thereby strengthening the expression of the features of the key regions. This process can highlight the features of the target regions, suppress the interference of irrelevant backgrounds, and thus improve the target detection accuracy of the model in complex backgrounds. This weight allocation strategy not only optimizes the feature extraction effect of the model, but also improves the model's ability to recognize small targets and occluded targets, ensuring the robustness and applicability of the model in different scenarios.
[0021] Preferably, the improvement of the network structure includes using the Inner-SIoU loss function. The Inner-SIoU loss function calculates the ratio of the overlapping area of the predicted bounding box and the ground truth bounding box to the overall area, and adjusts the remaining bounding box parts to reduce the detection error of overlapping targets.
[0022] The Inner-SIoU loss function is the fusion of the Inner-IoU loss function and the SIoU loss function, which can solve the problem of detecting pests in the overlapping part. By focusing on the core internal overlap of the bounding box rather than the whole, Inner-IoU provides a more accurate evaluation of the overlapping area. By introducing an auxiliary bounding box to calculate the loss function, it can be integrated into the existing IoU-based loss functions and has strong generalization ability for different detection tasks.
[0023] Preferably, the Inner-SIoU loss function introduces the center point distance as an additional weight factor during the optimization process to further correct the position offset of the bounding box and ensure the segmentation accuracy of overlapping targets.
[0024] The calculation process of Inner-IoU is as follows:
[0025]
[0026] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter
[0027]
[0028] L Inner-IoU = 1 - IoU inner
[0029] Where, represents the center points of the ground truth bounding box and the inner ground truth bounding box, x c , y c represents the center points of the predicted bounding box and the inner predicted bounding box, w gt and h gt represent the width and height of the ground truth bounding box, while the width and height of the predicted bounding box are represented by w and h. The variable ratio is a scaling factor in the range of 0.5 - 1.5. represents the left deviation of the abscissa of the center point of the ground truth bounding box, represents the right deviation of the abscissa of the center point of the ground truth bounding box, represents the lower deviation of the ordinate of the center point of the ground truth bounding box, represents the upper deviation of the ordinate of the center point of the ground truth bounding box, b l represents the left deviation of the abscissa x c of the center point of the predicted bounding box, b rRepresents the abscissa x of the center point of the predicted bounding box c The right deviation, b t Represents the ordinate y of the center point of the predicted bounding box c The lower deviation, b b Represents the ordinate y of the center point of the predicted bounding box c The upper deviation, inter represents the intersection of the internal ground truth bounding box and the internal predicted bounding box, union represents the union of the ground truth bounding box and the predicted bounding box, IoU inner Represents the internal intersection over union, L Inner-IoU Then represents the definition of the Inner-IoU loss function.
[0030] The SIoU loss function redefines the penalty metric by incorporating angle considerations, considering the angle of the expected regression direction, addressing the limitations of previous loss functions for region evaluation. It consists of four parts: Angle cost, Distance cost, Shape cost, and IoU cost. This improvement effectively reduces the degrees of freedom, enabling the SIoU loss function to be easily incorporated into any object detection and contributing to better results. The formula for the SIoU loss function is as follows:
[0031]
[0032] Among them, Δ represents the distance loss, Ω represents the shape loss, B represents the predicted bounding box, B Gt Represents the ground truth bounding box, IoU represents the intersection over union of the bounding boxes, L SIoU Then represents the definition of the SIoU loss function. Combining the ideas of Inner-IoU and SIoU fusion forms Inner-SIoU, enabling the model to achieve maximum efficiency in small object detection. Finally, the formula for Inner-SIoU is as follows:
[0033] L Inner-SIou = L SIoU + IoU - IoU inner
[0034] Among them, L SIoU Represents the SIoU loss function, IoU represents the intersection over union of the bounding boxes, IoU inner Represents the internal intersection over union, L Inner-SIou Represents the definition of the Inner-SIoU loss function.
[0035] Based on the armyworm board pest image dataset, a computing device is used to train the improved YOLOv10n model, and GPU is used for training acceleration. The training effect of the model is improved through the AdamW optimizer and the early stopping mechanism;
[0036] Training the improved model on a computer using a pest image dataset includes: using a GPU to accelerate the training process, using the AdamW optimizer based on the gradient descent algorithm to avoid unnecessary effects of weight decay on bias parameters, and improving the training effect of the model. Prevent overfitting during training through an early stopping mechanism.
[0037] Preferably, the training process of the model uses the AdamW optimizer, which calculates the second-order moment estimate of the gradient for each parameter and applies weight decay to reduce the risk of model overfitting. At the same time, a method of dynamically adjusting the learning rate is combined to improve the stability of training.
[0038] The training process of the model uses the AdamW optimizer to optimize the weight update method and improve the training effect of the model. When calculating parameter updates, the AdamW optimizer introduces a weight decay mechanism. Different from traditional L2 regularization, its weight decay directly acts on the parameters themselves, rather than through the loss function. This method can effectively prevent the weights from growing too large and reduce the risk of overfitting. Specifically, the AdamW optimizer uses first-order and second-order momentum estimates for the gradient calculation of each parameter, and stabilizes the update process by weighted averaging the squares of the gradients. Compared with the traditional Adam optimizer, AdamW improves the generalization ability of the model while ensuring fast convergence by adjusting the weight update method of gradient descent. In addition, this optimizer also combines a strategy of dynamically adjusting the learning rate, automatically adjusting the size of the learning rate according to the optimization progress during training, thus avoiding oscillations caused by too high a learning rate or slow convergence caused by too low a learning rate. This optimization method significantly improves the stability of training, enables the model to reach the best performance more efficiently, and enhances the adaptability to changes in data distribution, thus ensuring the robustness and reliability of the model in practical applications.
[0039] Deploy the trained improved YOLOv10n model to an edge computing device, detect real-time collected pest images, and drive a mechanical device to complete the replacement operation of the sticky board according to the detection results.
[0040] Testing and evaluating the pest images using the trained model includes: evaluating the detection effect of the model using evaluation metrics such as precision, mean average precision (mAP), floating point operations (GFLOPs), and detection speed (FPS).
[0041] Deploy the optimal weight file of the trained model to the edge computing device for inference, and implement pest detection on the edge computing device, including: convert the PyTorch format weight file obtained by training the model on the computer into an ONNX file through the ONNX library in the YOLO model training configuration environment, and install and run the dependencies and configuration environment of YOLO on the edge computing device. Then compile NCNN on the edge computing device, then transfer the ONNX format file to the edge computing device, and then use the onnx2ncnn tool of NCNN to convert the ONNX file into the NCNN format. Then write the detection code and run the detection code to perform inference on the model. Connect the camera to the edge computing device through the USB interface, control the camera to capture the pest image on the yellow sticky board through OpenCV, and the edge computing device uses the deployed model to perform real-time detection on the input pest image.
[0042] The edge computing device uses a Raspberry Pi 4B development board. The Raspberry Pi 4B development board is powerful and equipped with rich GPIO interfaces, supporting a series of mainstream frameworks and algorithms, making it more convenient to integrate the model into the device, and is widely used in tasks such as image classification and object detection.
[0043] Preferably, the edge computing device drives the automatic replacement device of the sticky board through a stepper motor based on the detection result. The replacement device includes a sticky board reel, a winding component and a fixing component to ensure the continuous operation of the sticky board and the stable collection of monitoring data.
[0044] The edge computing device automatically controls the operation of the stepper motor component in the mechanical module of the device according to the detection result to drive the sticky worm roll to rotate, so as to realize the automatic replacement process of the sticky worm roll, including: adding a counting module to the Raspberry Pi 4B detection code, defining a counting function as a counter in the source code to count the pests in the images captured by the Raspberry Pi 4B camera. Connect the stepper motor to the corresponding GPIO interface of the stepper motor driver and the Raspberry Pi 4B, and then write the stepper motor drive code in Python on the Raspberry Pi 4B, and define the rotation direction and number of steps of the stepper motor in the motor drive code. The code uses the RPi.GPIO library of the Raspberry Pi 4B to operate the GPIO pins to control the rotation of the stepper motor. Then write a Python script. According to the counting result of the pests in the detection code, when the set threshold is reached, run the stepper motor drive script to control the rotation of the stepper motor to realize the automatic unfolding and replacement of the sticky worm roll.
[0045] The automatic control method specifically includes: by editing the local file in Raspberry Pi 4B and adding the Python script to be executed, the Raspberry Pi 4B can be set to automatically start this script when powered on. By editing the crontab in the Raspberry Pi to add scheduled tasks, the script can be run at regular intervals. Thus, the camera can be scheduled to take pictures of the sticky worm roll to monitor pests, and whether to control the rotation of the stepper motor can be determined according to the results, thereby realizing the automation process.
[0046] The sticky worm roll is made by connecting the commercially available yellow sticky worm boards with a specification of 25 cm x 25 cm into a roll shape, and one side for catching worms is covered with a film.
[0047] The mechanical module specifically includes a frame. A stepper motor is provided on one side of the frame. The output end of the stepper motor is provided with a coupling and a shaft 1 connected to the coupling. A buckle 1 and a synchronous belt driving wheel are fixed on the shaft 1. A limit rod 1 is arranged beside the shaft 1. A shaft 2 is provided on the other side of the frame. A buckle 2 and a synchronous belt driven wheel cooperating with the synchronous belt driving wheel are fixed on the shaft 2. A shaft 3 is arranged beside the shaft 2. The shaft 3 is fixed on the frame. A limit rod 2 is arranged beside the shaft 3. A synchronous belt is arranged between the synchronous belt driving wheel and the synchronous belt driven wheel. An edge computing device and a power supply are provided below the frame. A rain shield is provided above the frame. A camera is provided on one side below the rain shield.
[0048] Both the shaft 1 and the shaft 2 are connected by deep groove ball bearings, and the shaft 3 is fixed on the frame and cannot rotate.
[0049] The stepper motor is a 28BYJ-48 stepper motor, and the UN2003 driver is selected as the stepper motor driver.
[0050] The limit rod 1 and the limit rod 2 are coated with a non-stick coating to prevent the sticky worm roll from sticking to the limit rods.
[0051] As Figure 2 shown, a pest situation monitoring system based on improved YOLOv10n is used to implement the pest situation monitoring method based on improved YOLOv10n. This system includes:
[0052] An image acquisition module, which is used to acquire the pest images on the sticky worm board;
[0053] An image preprocessing module, which is used to preprocess the acquired pest images to generate an image data set containing pest categories and location information;
[0054] A model improvement module, which is used to improve the network structure of the YOLOv10n model. The improvements include introducing the SPD-Conv module, adding the iRMB inverted residual block attention mechanism, and adopting the Inner-SIoU loss function;
[0055] A model training module for training the improved YOLOv10n model based on an image dataset, using GPU for training acceleration, and combining the AdamW optimizer and early stopping mechanism to improve the training effect;
[0056] A model deployment module for deploying the trained improved YOLOv10n model to edge computing devices;
[0057] An image detection module for detecting pest images collected in real time through edge computing devices;
[0058] A mechanical control module for driving a mechanical device to complete the replacement operation of the sticky pest board according to the detection results.
[0059] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0060] By optimizing the network structure of the YOLOv10n model, including introducing the SPD-Conv module, adding the iRMB inverted residual block attention mechanism, and using the Inner-SIoU loss function, the present invention realizes high-precision detection of small target pests and overlapping pests.
[0061] Based on this improved model, the present invention combines edge computing devices to perform real-time detection of pest images, and at the same time automatically controls the sticky pest board replacement device, completely solving the problems of low detection accuracy and inability to achieve automation in the prior art.
[0062] By optimizing the training process (such as the AdamW optimizer and early stopping mechanism), the present invention significantly improves the model training effect and robustness, providing an intelligent and efficient solution for field pest situation monitoring. Brief Description of the Drawings
[0063] Figure 1 It is a flowchart of the method of the present invention;
[0064] Figure 2 It is a structural block diagram of the system of the present invention;
[0065] Figure 3 It is a flowchart of the steps in the embodiment of the present invention;
[0066] Figure 4 It is a network structure diagram of the improved YOLOv10n model in the embodiment of the present invention;
[0067] Figure 5 It is a schematic diagram of the mechanical module structure in the embodiment of the present invention;
[0068] Figure 6 It is a detection effect diagram of the improved model on the pest image of the sticky pest board in the embodiment of the present invention;
[0069] Description of main reference numerals: 1, frame; 2, stepper motor; 3, coupling; 4, shaft 1; 5, synchronous belt driving pulley; 6, limit rod 1; 7, shaft 2; 8, synchronous belt driven pulley; 9, shaft 3; 10, limit rod 2; 11, synchronous belt; 12, Raspberry Pi 4B development board and power supply; 13, rain shield; 14, camera. Specific embodiments
[0070] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0071] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0072] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0073] As Figure 3 shown, in a specific embodiment provided by the present invention, the function implementation includes the following steps:
[0074] Step 1, obtaining pest images on a field sticky board;
[0075] Step 2, preprocessing the obtained images, annotating them, and establishing a sticky board pest image data set;
[0076] Step 3, improving the network structure of the YOLOv10n model;
[0077] Step 4, training the improved model on a computer through the pest image data set;
[0078] Step 5, using the trained model to test and evaluate the pest images;
[0079] Step 6, deploying the optimal weight file of the trained model to an edge computing device for inference and realizing pest detection on the edge computing device;
[0080] Step 7: The edge computing device automatically controls the operation of the stepping motor component in the mechanical module of the device to drive the rotation of the sticky worm roll, thereby realizing the automatic replacement process of the sticky worm roll.
[0081] Step 1: First, manually place multiple yellow sticky boards at different positions in the field, including between crop rows or at a certain height from the crops, at intervals. Then, every 1 - 2 days, use image acquisition devices such as smartphones and single-lens reflex cameras to take pictures of the sticky boards with a large number of pests on the surface to obtain photos. After taking the pictures, remove the used sticky boards, replace them with new sticky boards to continue catching pests, and take pictures again after 1 - 2 days. Thus, a total of 466 yellow sticky board pest images are obtained. After manually identifying the types of pests, select the three types of pests with the largest number as the research objects, namely flea beetles, whiteflies, and hoverflies.
[0082] Step 2: Preprocess the 466 pest images obtained. Due to the complex field environment and weather factors, among the 466 pest images obtained, there are a large number of images containing other non-pest impurities, damaged sticky boards, and unclear images, and these images cannot be used for research. Therefore, these images are removed. After removal, the remaining number of images is 327. To standardize the images, the images are uniformly cropped to the same size, such as 640x640 pixels. Then use LabelImg to label the pests in the images, and the annotation information includes the category information and location information of the pests. The label format is saved in the txt format suitable for the YOLO algorithm. Since the sample size is small, to enhance the generalization and robustness during model training, data augmentation is performed on the 327 pest images, including operations such as vertical flipping, scaling, rotation, shearing, adjusting brightness, adjusting saturation, and adding noise to the images. Finally, 2257 pest images are obtained. Then, divide the images into a training set, a validation set, and a test set according to the ratio of 7:2:1.
[0083] Step 3: As Figure 4As shown in the figure, the improvements of the present invention to the YOLOv10n model include: using the SPD-Conv module to replace the convolutional stride and pooling operations of the Yolov10n model. The SPD-Conv module consists of a space-to-depth layer (SPD) and a Conv layer (non-strided convolution). The SPD layer rearranges pixel blocks to the depth dimension through a space-to-depth conversion operation, increasing the depth of the feature map and retaining all information in the channel dimension, thus avoiding information loss in traditional downsampling methods. Then, the convolutional layer with a stride of 1 processes each region of the feature map, thereby maximizing the retention of image information and generating a richer feature representation, reducing the number of channels of the feature map and effectively processing features without losing details. Through this mechanism, SPD-Conv retains more information in the feature extraction stage, thus improving the model's recognition ability for small target images; adding the iRMB inverted residual block attention mechanism to the Yolov10n model to balance the model's attention to local detail features and global semantic features. The iRMB combines depthwise separable convolution and self-attention mechanisms. 1x1 convolution is used for channel compression and expansion to optimize computational efficiency, depthwise separable convolution is used to capture spatial features, and the attention mechanism is used to capture the global dependencies between features, thereby increasing the model's attention to pests in dense areas; using the Inner-SIoU loss function in the Yolov10n model focuses on the core internal overlap part of the bounding box rather than the whole, providing a more accurate evaluation of the overlapping region, thus solving the problems of missed detection and false detection of small targets and overlapping targets by the model.
[0084] Step four, the computer used to train the model is equipped with the Windows 11 system and 64GB of RAM, as well as the 13th Gen Intel(R) Core(TM) i9-13900 CPU@2.00GHz and the NVIDIA RTX 4090 GPU with 24GB of video memory. The software environment used in the experiment is Python 3.11, Cuda 12.4, and the PyTorch 2.4.1 deep learning framework. Preferably, the training parameters are set as follows: the size of the input image is 640x640 pixels, the batch size Batch-size for each training is 32, the initial learning rate is 0.01, the number of training epochs is 200, the early stopping mechanism is 100, and the optimizer is AdamW.
[0085] Step five, use the optimal weight file obtained from the training in step four to evaluate on the validation set and the test set. The evaluation metrics use precision, mean average precision (mAP), floating point operations (GFLOPs), and detection speed (FPS) to evaluate the detection effect of the model.
[0086] Step 6: Convert the optimal weight file in PyTorch format obtained by training the model on the computer into an ONNX file through the ONNX library in the YOLO model training configuration environment. This file is first stored on the computer side. The official Raspbian system is pre-burned on the Raspberry Pi 4B development board, and then the Raspberry Pi 4B is remotely connected to the computer using the SSH service, enabling operations on the Raspberry Pi 4B system from the computer side. Install the dependencies and configure the environment for running YOLO on the Raspberry Pi 4B. Then, compile NCNN on the Raspberry Pi 4B development board. The compilation steps are as follows: First, clone the NCNN library, then create and enter the build directory, and finally configure the compilation options. After compilation, the NCNN library files will be installed in the specified path of the Raspberry Pi system. Subsequently, transfer the ONNX format file to the Raspberry Pi 4B, and then use the onnx2ncnn tool of NCNN to convert the ONNX file into NCNN format. Then, write the detection code and run the detection code to perform inference on the model. Using the NCNN format has a faster inference speed than the ONNX format. Connect the camera to the Raspberry Pi 4B development board through the USB interface, control the camera to capture images of pests on the yellow sticky trap board through the OpenCV library and input them into the Raspberry Pi 4B, and then perform real-time detection on the input pest images, and display the identified pest species on the external screen.
[0087] Step 7: Add a counting module to the detection code on the Raspberry Pi 4B. Use the count method in solutions.ObjectCounter of the Ultralytics library in the code, and define a counting function as a counter in the source code to count the pests in the images captured by the Raspberry Pi 4B camera. Connect the stepper motor to the stepper motor driver and the corresponding GPIO interface of the Raspberry Pi 4B, and then write stepper motor drive code in Python on the Raspberry Pi 4B. Define the rotation direction and number of steps of the stepper motor in the motor drive code. The code uses the RPi.GPIO library of the Raspberry Pi 4B to operate the GPIO pins to control the rotation of the stepper motor. Subsequently, write a Python script. According to the pest counting result in the detection code, when the set threshold is reached, run the stepper motor drive script to control the rotation of the stepper motor to realize the automatic unfolding and replacement of the sticky trap roll.
[0088] To achieve automation, by editing the local file in the Raspberry Pi 4B and adding the Python script to be executed, the Raspberry Pi 4B can be set to automatically start this script when powered on. By editing the crontab in the Raspberry Pi and adding a scheduled task, the script can be run at a scheduled time. Thus, the camera can be scheduled to take pictures of the sticky trap roll to monitor pests, and decide whether to control the rotation of the stepper motor according to the results, thereby realizing the automation process.
[0089] The precision rate, recall rate, mAP50, and mAP50-95 of the improved model of the present invention on the sticky board pest dataset reached 83.2%, 83.2%, 86.8%, and 41.3% respectively. The detection speed reached 139 FPS, and the GFLOPs was only 8.8. The mAP50 was 1.7% higher than that of the original model, achieving high-precision detection of pests.
[0090] As Figure 5 shown, the specific operation process of the mechanical module of the present invention is as follows: When the stepper motor 2 rotates, it drives the shaft one 4 to rotate through the coupling 3. The shaft one 4 unfolds the sticky worm roll fixed on its surface. At the same time, the synchronous belt driving wheel 5 located on the shaft one 4 drives the shaft two 7 to rotate through the synchronous belt 11, so as to unfold the surface film of the sticky worm roll fixed on the shaft two 7, and thus the sticky worm roll on the fixed shaft three 9 is gradually unfolded for catching worms. The limiting rod one 6 and the limiting rod two 10 are used to keep the surface of the sticky worm roll parallel to the camera 14 to ensure the normal image shooting angle.
[0091] As Figure 6 shown, the detection results of the improved model of the present invention for the sticky board pests show that in the case of small targets, dense distribution, and overlapping distribution, the improved model can better detect the categories of pests, indicating that the model has good detection accuracy and robustness.
[0092] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0093] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for insect monitoring based on improved YOLOv10n, characterized in that: The following steps are involved: Acquire pest images on the sticky insect board, and pre-process the acquired images to generate an image dataset containing pest category and location information; Improve the network structure of the YOLOv10n model, including introducing the SPD-Conv module to replace the convolution step and pooling operation, adding the iRMB inverted residual block attention mechanism, and using the Inner-SIoU loss function to improve the detection performance of small and overlapping objects; The improved YOLOv10n model was trained using computing devices based on the sticky insect board pest image dataset, GPU was used to accelerate the training, and the AdamW optimizer and early stopping mechanism were used to improve the model training effect. The trained improved YOLOv10n model is deployed to the edge computing device to detect the pest images collected in real time, and the mechanical device is driven to complete the replacement operation of the sticky insect board based on the detection results.
2. The monitoring method according to claim 1, characterized in that: The improvement of the network structure of the YOLOv10n model includes using an SPD-Conv module to replace the convolution step and pooling operation. The SPD-Conv module divides the input feature map into multiple small blocks and gradually rearranges them before performing a convolution operation, thereby achieving downsampling without losing detail information.
3. The monitoring method according to claim 2, characterized in that: The SPD-Conv module performs space-to-depth conversion on the input feature map and applies a convolution kernel to each small block. The generated feature map reduces the number of channels while maintaining the spatial resolution, so as to more efficiently extract the spatial features of small targets.
4. The monitoring method according to claim 1, characterized in that: The improvement of the network structure also includes the introduction of the iRMB inverted residual block attention mechanism. The iRMB inverted residual block extracts features by first reducing the number of channels and then expanding the number of channels. At the same time, the attention weight allocation strategy is combined in the expansion stage to strengthen the feature expression of key areas.
5. The monitoring method according to claim 4, characterized in that: The attention weight allocation strategy calculates the global average and maximum value of each channel, generates two sets of weight tensors accordingly, and then performs channel-by-channel weighted fusion to highlight the characteristics of specific regions, thereby enhancing the model's ability to recognize target areas in complex backgrounds.
6. The monitoring method according to claim 1, characterized in that: The improvement of the network structure includes using the Inner-SIoU loss function, which calculates the ratio of the overlapping area of the predicted box and the true box to the overall area, and adjusts the remaining bounding box parts to reduce the detection error of overlapping targets.
7. The monitoring method according to claim 6, characterized in that: The Inner-SIoU loss function introduces the center point distance as an additional weight factor in the optimization process to further correct the position offset of the bounding box and ensure the segmentation accuracy of overlapping targets.
8. The monitoring method according to claim 1, characterized in that: The training process of the model adopts the AdamW optimizer, in which the second-order moment estimate of the gradient is calculated for each parameter and weight decay is applied to reduce the risk of model overfitting, while the method of dynamically adjusting the learning rate is combined to improve the stability of training.
9. The monitoring method according to claim 1, characterized in that: The edge computing device drives the automatic replacement device of the sticky insect board through a stepper motor based on the detection results. The replacement device includes a sticky insect board reel, a winding component and a fixing component to ensure the continuous operation of the sticky insect board and the stable collection of monitoring data.
10. An insect monitoring system based on improved YOLOv10n, used to implement the insect monitoring method based on improved YOLOv10n according to any one of claims 1 to 9, characterized in that: include: An image acquisition module is used to acquire images of pests on the sticky insect board; An image preprocessing module is used to preprocess the collected pest images to generate an image data set containing pest category and location information; A model improvement module is used to improve the network structure of the YOLOv10n model, including introducing the SPD-Conv module, adding the iRMB inverted residual block attention mechanism, and adopting the Inner-SIoU loss function; Model training module, used to train the improved YOLOv10n model based on image datasets, using GPU for training acceleration, and combining AdamW optimizer and early stopping mechanism to improve training results; Model deployment module, used to deploy the trained improved YOLOv10n model to edge computing devices; An image detection module, used to detect pest images collected in real time through edge computing devices; The mechanical control module is used to drive the mechanical device to complete the replacement operation of the insect sticky board according to the detection results.