YOLOv5-EKKD-based method for detecting quality of eggs of anallocrocis mediterraneus
By improving the YOLOv5-EKKD model, using the RepVGG module, RFEM module and BiFormer attention mechanism, the accuracy and robustness of Mediterranean borer egg detection in the prior art are solved, and high-precision target recognition and overlapping target detection are achieved.
Patent Information
- Application Number
- CN202510097340.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-08
AI Technical Summary
The existing object detection algorithms have insufficient recognition performance in small object detection and complex backgrounds, especially in the case of light changes or occlusion, making it difficult to accurately identify small and dense targets such as Mediterranean Powder Eggs.
Using the improved YOLOv5-EKKD model, by replacing CSPDarknet53 with RepVGG module, combining RFEM module and TFE module, BiFormer attention mechanism is introduced to build a feature fusion network, and the model's detection ability in small goals and complex contexts is improved.
High-precision detection of Mediterranean Powder Eggs is achieved, able to effectively identify overlapping targets, and improve the accuracy and robustness of the detection.
Smart Images

Figure CN120279370A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and particularly to a method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD. Background Art
[0002] Predatory mites are natural enemies of pests and can control prey at a low level without the application of pesticides, which is of great significance for the integrated control of harmful organisms. The most intractable problem currently faced in enhancing the production capacity of predatory mites is that the artificial diet necessary for large-scale production - Ephestia kuehniella eggs - cannot be obtained. This is mainly because the large-scale artificial cultivation of Ephestia kuehniella eggs has always been a technical problem. The product quality of Ephestia kuehniella eggs must be obtained through quality evaluation, which requires accurate count data of Ephestia kuehniella eggs in the product and accurate weight data of the product. Therefore, to break through the industrial production of Ephestia kuehniella eggs, this paper proposes a method for detecting the quality of Ephestia kuehniella eggs based on object detection in deep learning.
[0003] In recent years, object detection algorithms have undergone significant development, gradually evolving from early traditional methods to advanced models based on deep learning. In particular, the proposal of region convolutional neural networks and their variants has opened a new era of object detection. Subsequently, real-time detection algorithms such as YOLO and SSD have emerged one after another, greatly improving the detection speed and accuracy. Through end-to-end learning methods, these models have enabled object detection to be widely applied in various application scenarios, including autonomous driving, security monitoring, and medical image analysis. However, existing object detection algorithms still have some limitations. First, in the detection of small objects, many models are difficult to effectively capture detailed information, resulting in a low recognition rate of small objects. Second, the object recognition performance in complex backgrounds is often insufficient, especially in the case of light changes or occlusions, and the robustness of the algorithm is challenged. Therefore, directly using existing object detection models to detect tiny and dense objects such as Ephestia kuehniella eggs has problems of decreased detection accuracy and difficulty in recognizing overlapping objects, and has great limitations. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD.
[0005] The technical solution adopted by the present invention is:
[0006] A method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD, which includes the following steps:
[0007] S1. Real-time collect images using a detection platform after the impurity removal process of Ephestia kuehniella eggs;
[0008] S2. Use annotation tools to annotate the collected images of Ephestia kuehniella eggs;
[0009] S3. Use data augmentation algorithms to augment the annotated data to obtain the dataset for model training;
[0010] S4. Build the YOLOv5-EKKD model; the model includes a network backbone (Backbone), a neck (Neck), and a head (Head); the network backbone uses CSPDarknet53 as the main feature extraction network, and the C3 module in CSPDarknet53 is replaced by the RepVGG module; the neck includes the RFEM module and the TFE module. The RFEM module enables the output feature map to contain sufficient non-local background information, and the TFE module can capture the local details of small target objects. The shape features of Ephestia kuehniella eggs are detected and drawn using an elliptical detection box in the head (Head).
[0011] Specifically, the input (Input) is an image of Ephestia kuehniella eggs with a size of 640×640; the Backbone part uses CSPDarknet53 as the main feature extraction network and replaces its C3 module with the RepVGG module. P2, P3, P4, and P5 obtained from the Backbone part enter the Neck part for fusion respectively. The Neck part is mainly composed of two main component networks, namely the RFEM module and the TFE module. The RFEM module enables the output feature map to contain sufficient non-local background information, and the TFE module can capture the local details of small target objects.
[0012] Specifically, the RepVGG module structure consists of a 3×3 convolution branch, a 1×1 convolution residual branch, and an identity connection residual branch. This module uses different network structures in the training and inference stages. The model pays more attention to model accuracy during the training stage. It uses 3×3 convolution as the main branch structure and introduces a 1×1 convolution residual branch and an identity mapping residual branch on the basis of the 3×3 main branch; during the inference stage, the model pays more attention to speed and equivalently transforms all network layers into a network structure with a 3×3 convolution as the main branch through the strategy of reparameterization fusion, ensuring that the model has a powerful feature extraction ability and can achieve efficient inference. The RFEM module first effectively fuses the feature maps of P3, P4, and P5, and then the input feature map will pass through dilated convolutions with three different dilation rates (2, 4, 8) and perform channel concatenation in sequence. The advantage of this module is that it can effectively capture balanced non-local context features and local target features and solve the problem of missing local semantics of small targets.
[0013] S5. Use the YOLOv5-EKKD model for quality detection to detect the qualified and contaminated parts of the eggs in the Ephestia kuehniella egg image.
[0014] Further, in step S1, images are collected in real time using a detection platform, which includes a CMOS industrial camera, a lens, a light source, a vibration motor, and a computer system. The culture dish containing the eggs of the Mediterranean flour moth is adhesively bonded to the vibration motor and placed in the middle of the annular light source. Before collecting images, the vibration motor is powered on to act on the culture dish to make the eggs of the Mediterranean flour moth more evenly distributed. The CMOS industrial camera is installed on the eyepiece of the microscope for shooting, and the computer system is connected to the CMOS industrial camera to collect the captured images in real time.
[0015] Further, in step S2, before annotation, the original image is cropped to a size of 640×640 pixels, and then the cropped image is annotated. The label form is saved in the YOLO format.
[0016] Further, in step S3, the amplification method includes at least two random combinations of scaling, translation, and rotation.
[0017] Further, the YOLOv5-EKKD model includes a BiFormer attention mechanism.
[0018] Further, the RepVGG module includes a 3×3 convolution branch, a 1×1 convolution residual branch, and an identity connection residual branch; the RepVGG module uses different network structures in the training and inference stages; in the training stage, 3×3 convolution is used as the main branch structure, and at the same time, a 1×1 convolution residual branch and an identity mapping residual branch are introduced on the basis of the 3×3 main branch; in the inference stage, by using the strategy of reparameterization fusion, all network layers are equivalently transformed into a network structure with a 3×3 convolution as the main branch, ensuring that the model has a strong feature extraction ability and can achieve efficient inference.
[0019] Further, the RFEM module first effectively fuses multiple feature maps output by the network backbone. The input feature maps pass through dilated convolutions with three different dilation rates (2, 4, 8), and channel concatenation is performed in sequence. The advantage of this module is that it can effectively capture balanced non-local context features and local target features, and solve the problem of missing local semantics of small targets.
[0020] The present invention adopts the above technical solutions, and the beneficial effects are as follows: (1) Introducing the deep learning method in the quality detection of the eggs of the Mediterranean flour moth improves the accuracy of the detection of the eggs of the Mediterranean flour moth. (2) Aiming at the characteristics of small and dense targets of the eggs of the Mediterranean flour moth, a feature fusion network composed of the RFEM module, the TFE module, and the BiFormer attention mechanism is constructed, enabling the model to effectively identify overlapping targets while achieving high-precision detection. Description of the Drawings
[0021] The following further describes the present invention in detail with reference to the drawings and specific embodiments;
[0022] Figure 1 This is the diagram of the detection platform for Ephestia kuehniella eggs of the present invention;
[0023] Figure 2 This is the diagram of the YOLOv5-EKKD model;
[0024] Figure 3 This is the structural diagram of the RepVGG module;
[0025] Figure 4 This is the structural diagram of the RFEM module;
[0026] Figure 5 This is the structural diagram of the TFE module;
[0027] Figure 6 This is the diagram of the training data of the YOLOv5-EKKD model. Specific embodiments
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0029] As Figures 1 to 6 shown in one of them, the present invention discloses a method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD, including the following steps:
[0030] S1. After the impurity removal process of Ephestia kuehniella eggs, use the detection platform to collect images in real time.
[0031] Specifically, as shown in the appendix Figure 1 shown, the detection platform mainly consists of a CMOS industrial camera, a microscope (using a 0.5× lens and a 0.5× auxiliary objective lens), a lens, a light source, a vibration motor, and a computer system and other related components.
[0032] Furthermore, when collecting images, the culture dish containing the Ephestia kuehniella egg sample is bonded to the vibration motor. Before shooting, the motor is powered on to vibrate the culture dish, so as to make the distribution of the eggs more uniform and reduce the overlapping phenomenon to a certain extent. The culture dish is placed in the middle of the ring light source, and the computer system monitors the images taken by the CMOS industrial camera in real time, adjusts the lens to obtain clear images, and saves them in the computer.
[0033] S2. Use the annotation tool to perform annotation processing on the collected Ephestia kuehniella egg images.
[0034] Specifically, since the size of the originally collected image is 2048×1536 pixels and the number of eggs in the image is too large, making annotation extremely difficult, the original image is cropped to 640×640 pixels before annotation, and then the cropped image is annotated, and the label form is saved in YOLO format.
[0035] S3. Use the data augmentation algorithm to augment the annotated data to obtain the dataset for model training.
[0036] Specifically, the original dataset after annotation is augmented. The augmentation methods include random combinations of scaling, translation, and rotation. Finally, 1000 pictures are augmented as the dataset.
[0037] S4. Build the YOLOv5-EKKD model, as shown in the appendix Figure 2 It is composed of three parts: Backbone, Neck, and Head; Input is the Ephestia kuehniella egg image with a size of 640×640; the Backbone part uses CSPDarknet53 as the backbone feature extraction network, and replaces its C3 module with the RepVGG module. P2, P3, P4, and P5 obtained from the Backbone part enter the Neck part for fusion respectively. The Neck part is mainly composed of two main component networks, namely the RFEM module and the TFE module. The RFEM module enables the output feature map to contain sufficient non-local background information, and the TFE module can capture the local details of small target objects.
[0038] Specifically, the structure of the RepVGG module is as shown in the appendix Figure 3 It is mainly composed of a 3×3 convolution branch, a 1×1 convolution residual branch, and an identity connection residual branch. This module uses different network structures in the training and inference stages. The model pays more attention to model accuracy in the training stage. It uses 3×3 convolution as the main branch structure, and at the same time introduces a 1×1 convolution residual branch and an identity mapping residual branch on the basis of the 3×3 main branch; in the inference stage, the model pays more attention to speed, and through the strategy of reparameterization fusion, all network layers are equivalently transformed into a network structure with a 3×3 convolution as the main branch, ensuring that the model has a strong feature extraction ability and can achieve efficient inference.
[0039] Specifically, the structure of the RFEM module is as shown in the appendix Figure 4 The RFEM module first effectively fuses the feature maps of P3, P4, and P5, and then the input feature map will pass through dilated convolutions with three different dilation rates (2, 4, 8), and channel concatenation is performed in turn. The advantage of this module is that it can effectively capture balanced non-local context features and local target features, and solve the local semantic loss of small targets.
[0040] Furthermore, the input feature map will undergo dilated convolutions with three different dilation rates (2, 4, 8), and channel concatenation will be performed sequentially. If the resolution of F is 62×62 and the receptive field is 1, then the receptive fields of R1, R2, R3, and P are 5, 13, 29, and 29 respectively. The receptive field of P reaches half the size of F, which will enable the output feature map to contain sufficient non-local background information.
[0041] Specifically, the structural diagram of the TFE module is as shown in the appendix Figure 5 The TFE module captures the detailed information of small targets by concatenating features of three different sizes, large, medium, and small, in the spatial dimension, thereby enhancing the detection of dense Ephestia kuehniella eggs.
[0042] Furthermore, before feature encoding, the number of feature channels is first adjusted to be consistent with the main-scale features. After processing the large-scale feature map, its number of channels is adjusted to 1C, and then a hybrid structure of max pooling and average pooling is used for downsampling, which helps to retain the effectiveness and diversity of high-resolution features and cell images. For the small-scale feature map, a convolutional module is also used to adjust the number of channels, and then the nearest neighbor interpolation method is used for upsampling. This helps to maintain the richness of the local features of the low-resolution image and prevent the loss of small target feature information. Finally, the three feature maps of the same size, large, medium, and small, are convolved once and then concatenated in the channel dimension as follows:
[0043] F TFE = Concat(F l 、F m 、F s )
[0044] where F TFE represents the feature map output by the TFE module, and F l 、F m 、F s represent the large, medium, and small-scale feature maps respectively. F TFE has the same resolution as F m , and the number of channels is three times that of F m .
[0045] The BiFormer attention mechanism is introduced in the P3 branch. The BiFormer attention mechanism uses the Bi-Level Routing Attention (BRA) of double-layer routing attention as the basic building block, implementing a dynamic, query-aware sparse attention mechanism. On the basis of reducing the computational amount, it enhances the network's learning ability for multi-scale dense target features and improves the detection accuracy of the model in high-density Ephestia kuehniella eggs.
[0046] S5. Use the YOLOv5-EKKD model for quality inspection to detect the qualified and contaminated parts of the eggs in the Ephestia kuehniella Zeller egg image. As shown in Table 1, the comparison data between the present invention and other object detection models are presented.
[0047] Table 1: Comparison data between the present invention and other object detection models
[0048]
[0049] The present invention adopts the above technical solutions, and the beneficial effects are as follows: (1) Introduce the deep learning method in the quality inspection of Ephestia kuehniella Zeller eggs, which improves the accuracy of Ephestia kuehniella Zeller egg detection. (2) In view of the characteristics of small and dense Ephestia kuehniella Zeller eggs, construct a feature fusion network composed of an RFEM module, a TFE module and a BiFormer attention mechanism, enabling the model to effectively identify overlapping targets while achieving high-precision detection.
[0050] Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. Generally, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
Claims
1. A method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD, characterized in that: It includes the following steps: S1. After the impurity removal process of Ephestia kuehniella eggs, use the detection platform to collect images in real time; S2. Use the annotation tool to perform annotation processing on the collected Ephestia kuehniella egg images; S3. Use the data augmentation algorithm to perform augmentation processing on the annotated data to obtain the dataset for model training; S4. Build the YOLOv5-EKKD model; the model includes a network backbone, a neck, and a head; the network backbone uses CSPDarknet53 as the main feature extraction network, and the C3 module in CSPDarknet53 is replaced by the RepVGG module; the neck includes the RFEM module and the TFE module. The RFEM module enables the output feature map to contain sufficient non-local background information, and the TFE module can capture the local details of small target objects; the shape features of Ephestia kuehniella eggs are detected and drawn using an elliptical detection box at the head; S5. Use the YOLOv5-EKKD model for quality detection to detect the qualified and contaminated parts of the eggs in the Ephestia kuehniella egg images.
2. The method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD according to claim 1, wherein: In step S1, use the detection platform to collect images in real time. The detection platform includes a CMOS industrial camera, a lens, a light source, a vibration motor, and a computer system. The culture dish for placing Ephestia kuehniella eggs is bonded to the vibration motor and placed in the middle of the ring light source. Before collecting images, the vibration motor is powered on to act on the culture dish to make the distribution of Ephestia kuehniella eggs more uniform. The CMOS industrial camera is installed on the eyepiece of the microscope for shooting, and the computer system is connected to the CMOS industrial camera to collect the captured images in real time.
3. The method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD according to claim 1, wherein: In step S2, before annotation, the original image is cropped to a size of 640×640 pixels, and then the cropped image is annotated, and the label form is saved in the YOLO format.
4. The method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD according to claim 1, characterized in that: In step S3, the augmentation method includes at least two random combinations of scaling, translation, and rotation.
5. The method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD according to claim 1, characterized in that: The YOLOv5-EKKD model includes the BiFormer attention mechanism.
6. The Mediterranean flour moth egg quality detection method based on YOLOv5-EKKD according to claim 1, characterized in that: The RepVGG module includes a 3×3 convolution branch, a 1×1 convolution residual branch, and an identity connection residual branch; The RepVGG module uses different network structures in the training and inference stages; In the training stage, 3×3 convolution is used as the main branch structure, and at the same time, a 1×1 convolution residual branch and an identity mapping residual branch are introduced on the basis of the 3×3 main branch; In the inference stage, by using the strategy of reparameterization fusion, all network layers are equivalently transformed into a network structure with 3×3 convolution as the main branch.
7. The method for detecting the quality of Ephestia kuehniella eggs based on YOLOv5-EKKD according to claim 1, characterized in that: The RFEM module first effectively fuses multiple feature maps output by the network backbone. The input feature maps pass through dilated convolutions with three different dilation rates and are sequentially concatenated in channels.