Parasite egg detection method based on improved YOLOv8s

By designing the lightweight object detection model LCN-YOLOv8, combining lightweight convolution, attention mechanism and efficient neck network, the missed and misdetection problems in parasite egg detection in the existing technology are solved, and efficient and accurate detection effects are achieved.

CN119942533AActive Publication Date: 2025-05-06JIANGNAN UNIV +1

Patent Information

Application Number
CN202411939328.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-06
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The prior art has problems of missed detection and misdetection in parasite egg microscopic image detection, and lightweight improvements usually lead to a reduction in detection accuracy, making it difficult to balance the model complexity and detection accuracy.

Method used

A lightweight object detection model LCN-YOLOv8 based on improved YOLOv8s was designed, combining lightweight convolution, attention mechanism and efficient neck network, and using PA-C2f, SPDConv and LHSFPN modules to improve feature extraction and feature fusion capabilities.

Benefits of technology

This model greatly reduces the computational complexity and parameter quantity, improves detection accuracy, and performs better especially when detecting dense insect eggs and occlusion eggs, and can operate efficiently under resource constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942533A_ABST
    Figure CN119942533A_ABST
Patent Text Reader

Abstract

The invention discloses a parasitic ovum detection method based on improved YOLOv8s, and belongs to the technical field of image processing. The method comprises the steps that YOLOv8s is used as a baseline network, a lightweight target detection model LCN-YOLOv8 algorithm is designed by fusing lightweight convolution, an attention mechanism and an efficient neck network, and efficient detection of parasitic ova under complex conditions is achieved; the feature extraction capability of the model is improved under the condition of reducing the calculation amount and memory access through the PA-C2f module; the fine-grained feature loss of the eggs in the feature extraction process is reduced through an SPDConv module; in addition, the multi-scale feature fusion capability of the model is enhanced through an improved HSFPN module; by combining the modules, the detection capability of the model under the conditions of background shielding, small worm eggs and low resolution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a parasite egg detection method based on improved YOLOv8s, and belongs to the technical field of image processing. Background Art

[0002] Parasitic infection is one of the major public health issues worldwide. Parasite eggs such as protozoa, worms and ectoparasites live in the human body and continuously absorb nutrients, leading to malnutrition and anemia, which especially hinders the growth and development of children. Traditional optical microscopy is the golden method for diagnosing parasitic diseases. By observing the morphology and structure of parasites and their effects on host tissues with the naked eye and comparing them with known parasite morphologies, the type of parasite can be identified and the severity of infection can be assessed. However, microscopy-based parasite identification and quantitative analysis are very challenging. Not only is it labor-intensive, but it also requires well-trained researchers, which will cause a huge waste of time and labor costs.

[0003] With the rapid development of computing power and deep learning, artificial intelligence technology has been successfully used in natural language processing, face recognition, biomedical analysis and other applications. Using traditional machine learning algorithms to identify parasite egg microscopic images requires manual feature selection, and the processing efficiency is low. The target detection method based on deep learning can automatically extract egg features from a large amount of image data, quickly complete tasks such as classification, detection, and counting, greatly improve the recognition speed and accuracy of parasite egg microscopic images, provide timely and accurate information to inspectors, alleviate the pressure on microscopy personnel to a certain extent, and reduce detection costs.

[0004] Currently, many parasite microscopic image detection methods based on deep learning have emerged. Wan et al. proposed a coupled composite backbone network C2BNet with a dual-path structure backbone in "C2BNet: adeep learning architecture with coupled composite backbone for parasitic egg detection in microscopic images[J].IEEE Journal of Biomedical and Health Informatics, 2023, 15(02): 1-13." for automatic detection of parasite egg microscopic images. The model heterogeneity is used to learn object features from different angles to enhance the feature representation ability between different paths of the backbone. At the same time, a multi-scale weighted box fusion module is proposed to fuse the positions and confidence scores of all bounding boxes. The average detection accuracy on the Chula-Parasite-Egg dataset reaches 94.56%. Nouar et al. proposed the CoAt Net model using convolutional neural networks and attention modules in "Parasitic egg recognition using convolution and attention network[J].Scientific reports,2023,13:14475.", which was used to detect parasite microscopic images in the Chula-Parasite-Egg dataset, with an average accuracy and F1 of 93%. Graciela et al. proposed a FiCRoN model based on a fully convolutional regression network in "FiCRoN,a deep learning-based algorithm for the automatic determination of intracellular parasite burden from fluorescence microscopy images[J].Medical Image Analysis,2024,91:103036." for detecting intracellular parasites, using small-scale feature maps to detect intracellular parasites and large-scale feature maps to detect host cells, with an average accuracy of 96%.

[0005] With the widespread use of target detection models in industrial sites, model deployment has become increasingly critical. The pursuit of a balance between model detection efficiency and accuracy has also become a research hotspot for deep learning models. Some efficient lightweight models and lightweight convolutions have been proposed and proven to be successfully used in target detection models. Di et al. proposed the concept of deep separable convolution in "AMMNet: a multimodal medical image fusion method based on an attention mechanism and MobileNetV3[J]. Biomedical Signal Processing and Control, 2024, 96: 106561." The concept of deep separable convolution was proposed. The model parameters were reduced by decomposing the ordinary convolution into deep convolution and point-by-point convolution, and the inverted residual structure and linear bottleneck were introduced to reduce information loss while making better use of limited computing resources. CP Udeh et al. proposed using grouped convolution to extract features and reduce the number of parameters in "Improved ShuffleNetv2network with attention for speech emotion recognition[J]. Information Sciences, 2025, 689: 121488." The obtained feature map channels are disrupted by channel shuffling operations to promote the mutual exchange of feature information from different groups. Mercedes et al. first used ordinary convolution to generate a small number of original images in "Ghost Net for hyperspectral image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(12): 10378-10393.", and then generated phantom feature maps through simple linear operations, which can efficiently extract features while reducing the amount of calculation and parameters. In order to solve the feature redundancy problem of ordinary convolution, Chen et al. used conventional convolution to extract features from some input channels while keeping other channels unchanged in "Real-time detection of mature table grapes using ESP-YOLO network on embedded platforms[J]. Biosystems Engineering, 2024, 246: 122-134.", and followed it with point-by-point convolution to fully utilize channel information and reduce computational redundancy and memory access.

[0006] Although the above methods use deep learning models to achieve automatic detection of parasite microscopic images, there are still many missed detections and false detections due to the small size of parasite eggs, complex backgrounds blocking eggs, and low resolution of dense parasites, which requires inspectors to spend a lot of time on review. Although most of the improved models can improve detection accuracy, these improvements often lead to larger model parameters and increased computational burden, which is not conducive to deployment on mobile devices. Although some lightweight improvements can reduce the number of model parameters, they will also lead to reduced detection accuracy, and it is impossible to achieve a good balance between model complexity and detection accuracy. Summary of the invention

[0007] In order to solve the problems existing in the above-mentioned prior art, the present invention takes YOLOv8s as the baseline model, combines lightweight convolution, attention mechanism and efficient neck network, and designs a new target detection model LCN-YOLOv8 for efficiently detecting parasite egg microscopic images. It not only greatly reduces the computational complexity and parameter amount of the model, so that it can still run efficiently under resource-constrained conditions, but also improves the detection accuracy of the model, and has better detection results for dense and obscured eggs.

[0008] A parasite egg detection method based on improved YOLOv8s, comprising:

[0009] Step 1: Construct a parasite egg dataset and divide it into training set, validation set and test set;

[0010] Step 2: Combining lightweight convolution, attention mechanism, and efficient neck network, a new target detection model LCN-YOLOv8 is designed, which includes three parts: backbone feature extraction network Backbone, feature fusion network Neck, and detection head Head;

[0011] Step 3: Design a new feature extraction module PA-C2f, use partial convolution (PConv) instead of regular convolution to reduce the amount of calculation and memory access, and introduce the separated and enhanced attention module (SEAM) to ensure the model's efficient feature extraction capability;

[0012] Step 4: Introduce the Space-to-depth Conv (SPDConv) module to replace all conventional convolution modules except PA-C2f in the model to reduce the loss of fine-grained features in the feature extraction process, improve the detection ability of dense eggs and low-resolution eggs in multi-egg images, and reduce the complexity of the model;

[0013] Step 5: Design the LHSFPN (Lightweight High-level Screening-feature Fusion Pyramid Network) module as the neck network for efficient feature fusion, which reduces the number of model parameters while increasing the multi-scale feature fusion capability of the model and improves the detection accuracy of obscured insect eggs;

[0014] Step 6: Use the training set constructed in step 1 to train the LCN-YOLOv8 network, introduce the total loss function L to constrain the training process of the model, and use Mosaic and Mixup technology to perform data enhancement and amplification samples;

[0015] Step 7: Use the validation set constructed in step 1 to evaluate the performance of the model during the training process, adjust the model's hyperparameters, and train for a total of 300 rounds. If the model performance does not improve after 50 consecutive rounds, stop training early to save the best model.

[0016] Step 8: Use the test set images constructed in step 1 to input the best model saved in step 7 for testing, obtain the visual detection results of the parasite egg microscopic images, and automatically count the number of different types of parasite eggs.

[0017] Optionally, the step 1 includes:

[0018] 1a) Collect microscopic images of parasite eggs;

[0019] 2a) Under the guidance of experts, the collected parasite egg images are annotated and divided into single egg and multiple egg datasets according to the number of egg types in the images;

[0020] 3a) The single egg dataset and the multi-egg dataset are divided into training set, validation set and test set in a ratio of 7:2:1.

[0021] Optionally, the PA-C2f module in step 3 uses PA Block to replace the Bottleneck structure in C2f for feature extraction, and splices the features of different branches to enrich the feature expression capability. Compared with the original Bottleneck, PA Block uses PConv for efficient feature extraction and introduces the SEAM attention module, which helps the model distinguish between eggs and backgrounds of similar colors, reducing the number of model parameters and computational complexity, while improving the model's feature extraction capability and enhancing the model's detection effect on dense eggs in complex backgrounds.

[0022] Optionally, the PA Block first uses 3×3 PConv for efficient feature extraction to reduce the extra computation caused by a large amount of feature redundancy. PConv is followed by two PWConv layers with a convolution kernel size of 1×1, which makes full use of the channel information that is not processed in PConv and concentrates the receptive field in the middle. Finally, after the shortcut operation with the input features, the SEAM module is used to make the model pay more attention to the insect features in a complex background, reduce the interference of other features, and improve the model's ability to detect occluded insects.

[0023] Optionally, PConv and two PWConvs are presented as an inverted residual structure, and the middle PWConv can expand the number of channels of the feature map. Since more use of normalization and activation functions will limit the diversity of features, reduce the model calculation speed, and damage the model's computing performance, PA Block only uses the BN layer and RELU activation function after the middle PWConv to ensure feature diversity and low latency.

[0024] Optionally, the SPDConv in step 4 is divided into two parts: the space-to-depth layer and the non-span convolution layer. First, the space-to-depth layer performs interval sampling on the input feature map of size H×W×C to obtain 4 images of size Then concatenate the sub-feature maps in the channel dimension to obtain a sub-feature map of size The non-span convolution layer then uses a 1×1 non-span convolution to perform dimensionality reduction, retaining all discriminant feature information as much as possible, and finally obtains a size of Through multiple SPDConv modules, more abstract and rich feature information can be extracted. Compared with conventional convolution operations, it has a higher information retention rate and is more suitable for complex tasks such as detecting low-resolution insect eggs and small-pixel insect eggs.

[0025] Optionally, the LHSFPN in step 5 consists of a feature selection module and a feature fusion module. First, the feature maps of different scales extracted by the backbone network are screened by the SEAM attention module in the feature selection module to highlight the important features of the eggs. The feature selection module uses SEAM attention for feature screening, so that the model pays more attention to the shape, edges and other feature information of the eggs, weakens the influence of impurities such as background on the eggs, and uses the lighter and more efficient SPDConv to adjust the number of channels. Subsequently, LHSFPN fuses the high-level information and low-level information in the feature map through the Selective Feature Fusion (SFF) module to generate a high-level feature map F containing rich semantic information. high ∈R C×H×W and low-level feature maps It helps to detect small-pixel eggs and eggs blocked by the background, and improves the multi-scale feature fusion capability of the model.

[0026] Optionally, given a high-level feature map F high ∈R C×H×W and low-level feature maps The SFF module first expands the high-level feature map using a transposed convolution with a step size of 2 and a convolution kernel size of 3×3 to obtain a size of F high ∈R C×2H×2W Then, in order to unify the dimensions of the high-level feature map and the low-level feature map, bilinear interpolation is used to upsample or downsample the high-level feature map to obtain a new feature map.

[0027] F att =BL(TConv(F high ))

[0028] Where BL is the bilinear interpolation operation and TConv is the transposed convolution operation.

[0029] Then, the high-level feature map is converted into the corresponding attention weights through the SEAM module. After obtaining the features with the same dimension, the low-level feature map is filtered. Finally, the filtered low-level feature map is fused with the sampled high-level feature map to enhance the feature representation of the model and obtain the final output.

[0030]

[0031] Where SEAM represents the weight of the SEAM module, Represents the dot product operation.

[0032] The beneficial effects of the present invention are:

[0033] (1) The present invention uses YOLOv8s as the baseline network, integrates lightweight convolution, attention mechanism, and efficient neck network to design a lightweight target detection model LCN-YOLOv8 algorithm for efficient detection of parasite egg microscopic images and reducing the cost of manual microscopy; compared with other target detection networks, it has more efficient feature extraction capabilities and fewer parameters, and has better effects when detecting complex situations such as background occlusion, small eggs, and low resolution.

[0034] (2) The present invention designs a new feature extraction module PA-C2f to replace the C2f module for extracting rich features of parasite eggs, uses partial convolution instead of conventional convolution to reduce the amount of calculation and memory access, and introduces a separation enhanced attention module to improve the model's efficient feature extraction capability.

[0035] (3) The present invention introduces SPDConv to replace conventional convolution, which reduces the loss of fine-grained information of small-pixel eggs and low-resolution eggs during feature extraction, especially downsampling, enhances the model's ability to process spatial information, facilitates the distinction of individual eggs in dense eggs and complex backgrounds, and improves the detection accuracy of low-resolution eggs.

[0036] (4) The present invention designs the LHSFPN module to replace PANet as the neck network, which increases the multi-scale feature fusion capability of the model, and more comprehensively captures the features of eggs of different sizes while reducing the complexity of the model. The SEAM attention module is added to improve the detection accuracy of occluded eggs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0038] Figure 1 The overall flow chart of the parasite egg detection method based on the improved YOLOv8s provided in Example 1 of the present invention;

[0039] Figure 2 The overall structure diagram of the LCN-YOLOv8 model constructed based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention;

[0040] Figure 3 A PA-C2f structure diagram in the LCN-YOLOv8 model based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention;

[0041] Figure 4 A structural diagram of the PConv module in the LCN-YOLOv8 model based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention;

[0042] Figure 5 A schematic diagram of the SEAM module in the LCN-YOLOv8 model based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention;

[0043] Figure 6 A schematic diagram of the SPDConv module in the LCN-YOLOv8 model based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention;

[0044] Figure 7A schematic diagram of the LHSFPN module in the LCN-YOLOv8 model based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention;

[0045] Figure 8 This is a visualization diagram of the detection results obtained based on the improved YOLOv8s parasite egg detection method provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0047] Embodiment 1

[0048] This embodiment provides a parasite egg detection method based on improved YOLOv8s, see Figure 1 , the method comprising:

[0049] Step 1: Under the guidance of experts, the collected microscopic images of parasite eggs were manually annotated and divided into two data sets: single egg and multiple egg data sets according to the number of egg types in the image, and the training set, validation set and test set were divided according to 7:2:1.

[0050] Step 2: Select the YOLOv8s model as the baseline network, integrate the lightweight convolution, attention module and efficient neck network, and build the target detection model LCN-YOLOv8 with fewer parameters and stronger feature extraction capabilities for fast and accurate detection of parasite eggs.

[0051] Step 3: Design the PA-C2f module to replace the C2f module for efficient feature extraction; and introduce the SEAM attention module to ensure the efficient feature extraction capability of the model.

[0052] PA-C2f uses PA Block to replace the Bottleneck structure in C2f for feature extraction, and splices the features of different branches to reduce redundant feature maps and additional computing overhead; the SEAM attention module enables the model to pay more attention to the characteristics of the insect body in complex backgrounds, reduce the interference of other features, and help the model distinguish between insect eggs and backgrounds with similar colors; it not only reduces the number of model parameters and calculations, but also improves the model's feature extraction capabilities, enhancing the model's detection effect on dense insect eggs in complex backgrounds.

[0053] Step 4: Use the SPDConv module to replace all conventional convolution modules except PA-C2f in the model to reduce the loss of fine-grained information of low-resolution eggs during feature extraction, especially downsampling, improve the detection ability of dense eggs and low-resolution eggs in multi-egg images, and reduce the complexity of the model.

[0054] The SPDConv module is divided into two parts: the space-to-depth layer and the non-span convolution layer. The space-to-depth layer samples the input feature map at intervals and concatenates the sampled feature maps in the channel dimension. The non-span convolution layer performs dimensionality reduction through non-span convolution, retains the information of the discriminant features, and obtains the final feature map.

[0055] Step 5: Design LHSFPN as the neck network to reduce the complexity of the model while more comprehensively capturing the features of eggs of different sizes and completing feature fusion more efficiently.

[0056] The LHSFPN module includes a feature selection module and a feature fusion module. First, the feature selection module uses SEAM attention to screen features, so that the model pays more attention to the shape, edge and other feature information of the eggs, and weakens the influence of background and other impurities on the eggs. Subsequently, the selective feature fusion mechanism fuses the high-level information and low-level information in the feature map to generate a feature map with rich semantic content, which helps to detect small-pixel eggs and eggs blocked by the background, and improves the multi-scale feature fusion capability of the model.

[0057] Step 6: Train the LCN-YOLOv8 model, introduce the total loss function L to constrain the model training process, and use Mosaic and Mixup technology to perform data enhancement and amplify samples.

[0058] If the model performance does not improve after 50 consecutive rounds, the training will be stopped early to save training time, and Mosaic and Mixup technologies will be used to enhance the data and expand the samples.

[0059] Step 7: Use the validation set to evaluate the performance of the model during training, adjust the model's hyperparameters, and train for a total of 300 rounds. If the model performance does not improve for 50 consecutive rounds, stop training early to save the best model.

[0060] Step 8: Input the parasite egg microscopic images in the test set into the saved best model for testing to obtain the visual detection results of the parasite eggs.

[0061] Embodiment 2

[0062] This embodiment provides a parasite egg detection method based on improved YOLOv8s, including:

[0063] Firstly, YOLOv8s was selected as the baseline network, PA-C2f and LHSFPN were designed for efficient feature extraction and feature fusion, PConv and SPDConv were introduced to reduce the number of parameters and complexity of the model, and a lightweight target detection model LCN-YOLOv8 was constructed for efficient detection of parasite egg microscopic images. Then, a self-made parasite egg microscopic image dataset was created, and the images of the training set were input into LCN-YOLOv8 training, and the network parameters were optimized using the validation set. Finally, the best model weights were saved to obtain the visual detection results of the parasite egg images in the test set.

[0064] Step 1: Create a data set, manually annotate the collected microscopic images of parasite eggs, and divide them into training sets, validation sets, and test sets. The specific contents of constructing the data set are as follows:

[0065] 1a) Collect several microscopic images of human parasite eggs from a medical college, and use LambelImg software to create labels under the guidance of experts to construct a human parasite microscopic image dataset.

[0066] 1b) According to the number of egg types in the image, the dataset is divided into single egg and multiple eggs. The single egg dataset contains 852 images, and the multiple egg dataset contains 465 images. The single egg dataset and the multiple egg dataset are divided into training set, validation set and test set in the ratio of 7:2:1 respectively.

[0067] 1c) The microscope magnification of the single egg dataset is large, the pixels of the eggs in each image are large, the number is small, and the recognition difficulty is low; the number of eggs in the multi-egg dataset is large and dense, the pixels of a single egg are small and the resolution is low. In some images, the eggs are blocked due to the complex background, and some feature information is lost, making detection difficult.

[0068] Step 2: Select the YOLOv8s model as the baseline network and improve it to design a lightweight target detection model LCN-YOLOv8 with efficient feature extraction and feature fusion. Figure 2 The overall structure diagram of the LCN-YOLOv8 model in the present invention is composed of three parts: the backbone feature extraction network Backbone, the feature fusion network Neck, and the detection head Head. Its specific structure is as follows:

[0069] 2a) Backbone uses the lightweight module PA-C2f for efficient feature extraction and splicing of features from different branches to obtain richer gradient information, enrich the expressive power of features, and improve the detection capability of the model in complex situations.

[0070] 2b) The Neck part uses LHSFPN as the neck network to achieve full fusion of feature maps of different sizes, while reducing the complexity of the model and more comprehensively capturing the characteristics of eggs of different sizes.

[0071] 2c) The Head part completes the regression and classification tasks and detects large, medium and small eggs on images of three different sizes: 13×13, 26×26 and 52×52.

[0072] Step 3: Replace the conventional convolution in C2f with PConv, design a lightweight and efficient feature extraction module PA-C2f, and introduce the SEAM attention module, which helps the model distinguish between eggs and backgrounds of similar colors. This not only reduces the number of parameters and calculations of the model, but also improves the model's feature extraction ability and enhances the model's detection effect on dense eggs in complex backgrounds. The structure diagram of the PA-C2f module and its submodule PA Block is shown in the figure. Figure 3 As shown, Figure 3 (a) is the structural diagram of the PA-C2f module. Figure 3 (b) is a structural diagram of the PA Block submodule of the PA-C2f module; the specific method includes:

[0073] 3a) In PA Block, 3×3 PConv is first used for efficient feature extraction to reduce the extra computation caused by a large amount of feature redundancy. Figure 4 (a) in the figure is the PConv structure diagram. Figure 4 (b) is a conventional convolution structure diagram; PConv selectively applies conventional convolution to only some input channels for spatial feature extraction, and does not process the remaining channels. PConv treats the first or last continuous channel c p It is calculated as the entire feature map, so the floating point number of PConv is:

[0074] FLOPs = h × w × k 2 ×c p 2

[0075] Where h and w are the height and width of the feature map, respectively, and k is the size of the convolution kernel. If the number of channels in conventional convolution processing is c, in a typical channel ratio In this case, the FLOPs of PConv is the same as that of regular convolution.

[0076] 3b) PConv is followed by two PWConv layers with a convolution kernel size of 1×1, which fully utilizes the channel information that is not processed in PConv and concentrates the receptive field in the middle. Figure 3As shown in (b), PConv and two PWConvs present an inverted residual structure, and the middle PWConv can expand the number of channels of the feature map. In addition, since the use of more normalization and activation functions will limit the diversity of features, reduce the model calculation speed, and damage the model's computing performance, PA Block only uses the BN layer and RELU activation function after the middle PWConv to ensure feature diversity and low latency.

[0077] 3c) Finally, after the shortcut operation with the input features, the SEAM module is used to make the model pay more attention to the insect features in the complex background, reduce the interference of other features, and improve the model's ability to detect occluded insects. The structure of the SEAM module is as follows: Figure 5 As shown, Figure 5 (a) is the structure diagram of SEAM. Figure 5 (b) is the structural diagram of CSMM; the first part is the Channel and Spatial Mixing Module (CSMM) with residual connection, which uses depthwise separable convolution to perform channel-by-channel convolution, learn the importance of different channels and reduce model parameters. In order to compensate for the information loss between channels, the output of the deep convolution is combined through point-by-point convolution. SEAM then uses a two-layer fully connected network to aggregate the channel information extracted from CSMM modules of different patch sizes, enhance the connectivity between channels, and use the information relationship between occluded and unoccluded eggs to compensate for the information loss caused by background occlusion. Exponential transformation of the output weights learned by the fully connected layer, expanding the value range from [0,1] to [1,e], can reduce the position error of the result and improve the robustness of the model when dealing with occluded eggs.

[0078] Step 4: LCN-YOLOv8 uses the SPDConv module to replace all conventional convolution modules in the model except PA-C2f to reduce the loss of fine-grained information of small pixel eggs and low-resolution eggs during feature extraction, especially downsampling. The SPDConv module can enhance the model's ability to process spatial information and facilitate the distinction of individual eggs from dense eggs and complex backgrounds. Its structure is as follows: Figure 6 The specific methods include:

[0079] 4a) SPDConv is divided into two parts: space-to-depth layer and non-span convolution layer. First, the input feature map of size H×W×C is sampled at intervals to obtain 4 images of size The sub-feature map contains the global spatial feature information of the original image, which increases the model feature extraction capability. Then the sub-feature maps are concatenated in the channel dimension to obtain a size of feature map.

[0080] 4b) Use 1×1 non-span convolution to reduce the dimension of the concatenated feature map, retaining all the discriminant feature information as much as possible, and finally obtain a size of Through multiple SPDConv modules, more abstract and rich feature information can be extracted. Compared with conventional convolution operations, it has a higher information retention rate and is more suitable for complex tasks such as detecting low-resolution insect eggs and small-pixel insect eggs.

[0081] Step 5: LCN-YOLOv8 uses LHSFPN as the neck network, such as Figure 7 As shown in the figure, while reducing the complexity of the model, the features of eggs of different sizes are captured more comprehensively, and the conventional convolution and C2f modules are replaced with SPDConv and PA-C2f modules respectively to complete feature fusion more efficiently. Figure 7 (a) in the figure is the LHSFPN structure diagram. Figure 7 (b) is a structural diagram of the SFF module; the specific method includes:

[0082] 5a) LHSFPN mainly consists of two parts: feature selection module and feature fusion module. First, the feature map is screened by SEAM attention module in the feature selection module, so that the model pays more attention to the shape, edge and other feature information of the eggs, weakens the influence of background and other impurities on the eggs, and uses the lighter and more efficient SPDConv to adjust the number of channels. Then, LHSFPN fuses the high-level information and low-level information in the feature map through the SFF mechanism to generate a high-level feature map F with rich semantic content. high ∈R C×H×W and low-level feature maps Improve the model's multi-scale feature fusion capabilities.

[0083] 5b) Given a high-level feature map F high ∈R C×H×W and low-level feature maps First, the high-level feature map F high ∈R C×H×W The high-level features are expanded using a transposed convolution with a stride of 2 and a kernel size of 3×3 to obtain a size of F high ∈R C×2H×2W Then, in order to unify the dimensions of the high-level feature map and the low-level feature map, bilinear interpolation is used to upsample or downsample the high-level feature map to obtain a new feature map.

[0084] F att =BL(TConv(F high ))

[0085] Where BL is the bilinear interpolation operation and TConv is the transposed convolution operation.

[0086] Then, the high-level feature map is converted into the corresponding attention weights through the SEAM module. After obtaining the features with the same dimension, the low-level feature map is filtered. Finally, the filtered low-level feature map is fused with the sampled high-level feature map to enhance the feature representation of the model and obtain the final output.

[0087]

[0088] Where SEAM represents the weight of the SEAM module, Represents the dot product operation.

[0089] Step 6: Train the LCN-YOLOv8 network.

[0090] All experiments were conducted on the same computer, with an Intel Xeon Silver 4210R CPU and an NVIDIA GeForce RTX 3090 GPU with a memory size of 24G. The training environment was Pytorch-GPU 1.11.0, and the python version was 3.8.13. During the experiment, the epoch was set to 500, the batch size was set to 32, the SGD optimizer and the cosine learning rate descent method were used, and the initial learning rate was set to 0.01. All data were enhanced using Mosaic and Mixup. No pre-trained weights were used during training. If the model performance did not improve for 100 consecutive epochs, the training was stopped to save time.

[0091] The experiment uses mean Average Precision (mAP), average precision (AP), parameter amount (Parameter), giga floating point operations (GFLOPs), etc. to evaluate the experimental results. Average precision reflects the comprehensive performance index of the algorithm. It uses recall (R) as the horizontal axis and precision (P) as the vertical axis to integrate the formed PR curve. The recall rate indicates the proportion of correctly predicted positive samples in the total number of samples, and the precision rate indicates the proportion of correctly predicted positive samples in all detected positive samples. The calculation method of the above indicators is as follows:

[0092]

[0093] In the formula, TP represents the number of correctly classified positive samples, FP represents the number of misclassified negative samples, FN represents the number of misclassified positive samples, TN represents the number of correctly classified negative samples, and C represents the number of target categories.

[0094] Step 7: Adjust the model’s hyperparameters to save the best model.

[0095] Step 8: Input the image to be tested into the saved best model for detection.

[0096] Although the backbone network in YOLOv8s uses the C2f module for feature extraction and can obtain rich feature information, the Bottleneck module inside C2f still uses conventional convolution for stacking, which has room for improvement. Therefore, this study designs the PA-C2f module for efficient feature extraction. At the same time, in order to verify the effectiveness of this module, it is compared with five other different Bottleneck modules. Among them, the C2f model can refer to the introduction in "Liu Jia, Zhang Zengwei, Chen Dapeng, etc. Improving the positioning accuracy of SLAM in AR based on improved YOLOv8 [J]. Journal of System Simulation, DOI: 10.16182 / j.issn1004731x.joss.24-0564.)", and C2f-GhostConv can refer to "Li Mingyu, Lin Jiaquan. Lightweight driver facial target detection algorithm based on YOLOv8-DF [J]. Journal of System Simulation. DOI:10.16182 / j.issn1004731x.joss.24-0320. ", C2f-DualConv can refer to "Cao Yu, Li Jiayang, Wang Fang. Lightweight model for underwater fish target recognition based on improved YOLOv8n[J]. Journal of Shanghai Ocean University. DOI:31.2024.S.20241218.1525.004.html.", C2f-ScConv can refer to Refer to "Zheng Yunshui, Meng Yang. Railway station signal layout information extraction method based on improved YOLOv8s [J]. China Railway Science, 2024, 45(05): 209-220.)", C2f-Fasterblock can refer to "Yang Hongxin, Chen Yue, Pei Guoquan, etc. Research on Yunnan small-grain coffee bean grading method based on lightweight YOLOv8-FasterBlock model [J]. Food Science. DOI: 11.2206.T S.20241127.1153.032.html.)”, C2f-Msblock can refer to “Wen Tao, Wang Tianyi, Huang Shirui, et al. Crop and Quinoa Detection Algorithm Based on Improved YOLOv8: MES-YOLO[J]. Computer Engineering and Science. DOI:43.1258.tp.20241011.1309.004.html.)”; To maintain the consistency of the experiment, all C2f modules in the model were replaced. The comparative experimental results of the C2f module are shown in Table 1. Although C2f-GhostConv makes the model more lightweight, the use of GhostConv also reduces the feature extraction ability of the model, and the detection accuracy on the multi-egg dataset is significantly reduced, which cannot meet the detection requirements. Among several comparative modules, Fasterblock, which also uses PConv, has better performance in terms of model complexity and single egg detection accuracy, but the detection accuracy on the multi-egg dataset is still not improved.The PA-C2f module not only uses PConv for efficient feature extraction, but also introduces the SEAM attention module, which enables the model to weaken the interference of background information and pay more attention to the characteristics of dense insect eggs and occluded insect eggs. It not only achieves the lightweight of the model, but also improves the detection accuracy, especially in the detection of multiple insect egg datasets, which has a great advantage over other modules.

[0097] Table 1 C2f module comparison experiment

[0098]

[0099] In order to verify whether the three improved modules in LCN-YOLOv8 can effectively improve the detection performance, this section conducts ablation experiments on the single egg and multiple egg datasets respectively. The experimental results are shown in Table 2. The mark √ represents the used module.

[0100] Table 2 Ablation experiment results

[0101]

[0102] As shown in Table 2, all improved modules have effectively reduced the number of parameters and computation of the model, and achieved lightweight model. PA-C2f module uses PConv and SEAM attention to achieve efficient feature extraction, and has a good performance on the single egg dataset, with mAP@0.5 and mAP@0.5:0.95 increased by 0.4% and 0.7% respectively. LHSFPN uses SEAM and SFF modules to efficiently fuse the feature maps extracted by the backbone network, enhancing the model's detection ability for multiple egg images, and mAP@0.5 and mAP@0.5:0.95 on the multiple egg dataset increased by 0.6% and 1.0% respectively. SPDConv mainly solves the detection problem of low-resolution eggs and small-pixel eggs in the multiple egg dataset, and mAP@0.5 and mAP@0.5:0.95 on the multiple egg dataset are both increased by 0.4%. Finally, after all the modules were added, the number of model parameters decreased by 6M, the amount of computation decreased by 7.5G, the mAP@0.5 and mAP@0.5:0.95 of the single egg dataset increased by 0.4% and 1.7% respectively, and the mAP@0.5 and mAP@0.5:0.95 of the multiple eggs dataset increased by 1.1% and 1.6% respectively, achieving a good balance between detection accuracy and model complexity.

[0103] Table 3 Comparison experimental results of different models

[0104]

[0105] LCN-YOLOv8 is improved by using YOLOv8s as the baseline network. In order to verify its superior detection performance, this section compares LCN-YOLOv8 with the most advanced single-stage target detection models such as YOLOv5, YOLOv9, and YOLOv10 under the same experimental environment. The experimental results are shown in Table 3. For details about YOLOv3-tiny, please refer to "Jin Xiaofang, Yue Ding, Liu Jinyu. Research on intelligent reconnaissance virtual training system based on YOLOv3-tiny [J]. Journal of Ordnance Equipment Engineering, 2023, 44(08): 186-190.)", and for YOLOv5m, please refer to "Song Yaolian, Wang Can, Li Dayan, et al. UAV small target detection algorithm based on improved YOLOv5s [J]. Journal of Zhejiang University (Engineering Edition), 2024, 58(12): 24 17-2426.)”, YOLOv8s and YOLOv8m can refer to “Zhang Lifeng, Tian Ying. Improved YOLOv8 multi-scale lightweight vehicle target detection algorithm [J]. Computer Engineering and Applications, 2024, 60(03): 129-137.)”, YOLOv9s can refer to “Wu Yibang, Chen Zhe, Li Zhe et al. Remote sensing image detection method for illegal cultivation areas on steep slopes based on improved YOLOv9 [J]. Transactions of the Chinese Society of Agricultural Engineering, 2024, 40(17): 197-204.)”, YOLOv10s and YOLOv10m can refer to “Yang Haitao, Zhao Junyu, Wang Rui et al. UAV small target detection based on DBB-YOLOv10s [J]. Infrared Technology. DOI: 53.1053.TN.20241126.1129.002.html.)”.

[0106] The experimental results show that although YOLOv3-tiny has the smallest model calculation amount, its accuracy in detecting multiple parasite eggs lags far behind other models. Compared with the two latest lightweight models YOLOv9s and YOLOv10s, LCN-YOLOv8 not only has a significant improvement in detection accuracy on both data sets, but also has less calculation and parameter amounts than the two latest models. Compared with the two larger models YOLOv9m and YOLOv10m, LCN-YOLOv8 not only has a greater advantage in parameter and calculation amount, but also has slightly higher detection accuracy than these two models. LCN-YOLOv8 integrates efficient convolution modules, feature extraction modules and feature fusion modules, which better balances detection accuracy and model complexity and can be effectively used in parasite egg detection.

[0107] Figure 8 This is a comparison chart of the detection visualization results of YOLOv8s and LCN-YOLOv8. Figure 8 (a) in the figure is the detection result of YOLOv8s. Figure 8(b) is the detection result of LCN-YOLOv8. A total of eight detection effect pictures under three conditions, background occlusion, small insect eggs, and low resolution, are selected for display and numbered (1)-(8). Figure 8 (a) and Figure 8 The same numbered images in (b) are different detection results of the same image using two models; and the boxes with different colors represent different types of parasite eggs. In order to solve the serious background occlusion problem in the multi-egg dataset, Figure 8 The detection results show that the traditional YOLOv8s will have more missed detections, while LCN-YOLOv8 can enhance the model's ability to distinguish between eggs and backgrounds and reduce missed detections caused by background occlusions due to the use of SEAM attention and PA-C2f for efficient feature extraction. For small individual eggs in the data set, YOLOv8s is prone to misjudge them as background impurities when detecting them, while the LCN-YOLOv8 model uses improved HSFPN for efficient feature fusion to reduce missed detections of small individual eggs. For low-resolution eggs caused by shooting and other problems, LCN-YOLOv8 uses SPDConv to extract fine-grained features, reduce the loss of detail features such as edges, and reduce the probability of misjudging eggs as impurities. In summary, LCN-YOLOv8 has better results than YOLOv8s in detecting complex situations such as background occlusion, small eggs, and low resolution.

[0108] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.

[0109] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A parasite egg detection method based on improved YOLOv8s, characterized in that: The method comprises: Step 1: Collect microscopic images of parasite eggs to construct a parasite egg dataset. According to the number of parasite egg types, the parasite egg dataset is labeled as a single egg dataset and a multi-egg dataset, and divided into a training set, a validation set, and a test set respectively; Step 2: Select the YOLOv8s model as the baseline network, integrate the lightweight convolution, attention module and efficient neck network, and build the target detection model LCN-YOLOv8 for detecting parasite eggs; Step 3: Design the PA-C2f module to replace the C2f module for feature extraction, and introduce the SEAM attention module to ensure the feature extraction capability of the model; Step 4: Use the SPDConv module to replace all conventional convolution modules except PA-C2f in the model to reduce the loss of fine-grained information of insect eggs during feature extraction; Step 5: Design the HSFPN module including the feature selection module and the selective feature fusion module as the neck network for feature fusion; Step 6: Use the training set constructed in step 1 to train the LCN-YOLOv8 model, introduce the total loss function L to constrain the training process of the model, and use Mosaic and Mixup technology to perform data enhancement and amplification samples during the training process; Step 7: Use the validation set constructed in step 1 to evaluate the performance of the model during training and adjust the model's hyperparameters. If the model performance does not improve, stop training early to save the best model. Step 8: Input the parasite egg microscopic images in the test set constructed in step 1 into the saved best model for testing to obtain the visual detection results of the parasite eggs.

2. The method according to claim 1, characterized in that The step 3 comprises: In the PA-C2f module, PA Block is used to replace the Bottleneck structure in C2f for feature extraction, and features of different branches are spliced; In the PA Block, PConv with a convolution kernel size of 3×3 is used for feature extraction, and two PWConv layers with a convolution kernel size of 1×1 are used to utilize the channel information not processed in PConv. After the input features are shortcutted, the SEAM module is introduced to enhance the model's ability to focus on features.

3. The method according to claim 2, characterized in that The PConv and PWConv in the PA Block present an inverted residual structure, and only the BN layer and RELU activation function are used in the middle PWConv layer to expand the number of channels of the feature map.

4. The method according to claim 3, characterized in that The SEAM structure includes a channel-space hybrid module CSMM with residual connections and a two-layer fully connected network; The channel space mixing module CSMM with residual connection learns the importance of different channels, and combines them through point-by-point convolution after the output of deep convolution to compensate for the information loss between channels; the two-layer fully connected layer network aggregates the channel information extracted from CSMM modules of different patch sizes.

5. The method according to claim 4, characterized in that The SPDConv in step 4 includes a space-to-depth layer and a non-span convolutional layer; The space-to-depth layer performs interval sampling on the feature map of input size H×W×C to obtain a size of Then concatenate the sub-feature maps in the channel dimension to obtain a sub-feature map of size The feature map of The non-span convolution layer uses 1×1 non-span convolution to perform dimensionality reduction operation, retaining the information of all discriminant features, and finally obtains a size of feature map.

6. The method according to claim 5, characterized in that In step 5, the feature selection module performs feature screening through the SEAM attention module and uses SPDConv to adjust the number of channels; The selective feature fusion module generates a high-level feature map F containing semantic information by fusing the high-level information and low-level information in the feature map. high ∈R C×H×W and low-level feature maps 7. The method according to claim 6, characterized in that The selective feature fusion module expands the high-level feature map using a transposed convolution with a step size of 2 and a convolution kernel size of 3×3 to obtain a size of F high ∈R C×2H×2W feature map; use bilinear interpolation to upsample or downsample the high-level feature map to obtain a new feature map F att The expression is: F att =BL(TConv(F high )) Among them, BL is a bilinear interpolation operation, and TConv is a transposed convolution operation; The high-level feature map is converted into the corresponding attention weights through the SEAM module. After obtaining the features with the same dimension, the low-level feature map is filtered; the filtered low-level feature map is fused with the high-level feature map to enhance the feature representation of the model and obtain the final output F out The expression is: Among them, SEAM represents the weight of the SEAM module, Represents the dot product operation.

8. The method according to claim 7, characterized in that The expression of the total loss function L in step 6 is: L=λ cls L cls +λ obj L obj +λ box L box Among them, L cls is the classification loss function, L obj is the target confidence loss function, L box is the bounding box loss function, λ cls is the classification loss function L cls The weight coefficient, λ obj is the target confidence loss function L obj The weight coefficient, λ box is the bounding box loss function L box The weight coefficient is used to balance the contribution of different losses to the total loss; The Mosaic technology enhances data by stitching multiple images into a new image; The Mixup technology generates new samples by linearly combining samples from different data sets to expand the data set.

9. The method according to claim 8, characterized in that The criterion for evaluating the model in step 7 is that the total loss function L no longer decreases with the training process, and the performance of the model no longer improves with the training process. At this time, the model is the best model.

10. The method according to claim 9, characterized in that The division ratio of the training set, validation set and test set in step 1 is 7:2:1.

Citation Information

Patent Citations

  • Worm egg detection method, device and equipment based on improved YOLOv5 and medium

    CN116778479A

  • Knowledge distillation-based lightweight SAR image multi-type target detection method and device, equipment and medium

    CN118552846A

  • Lightweight image detection method and system fusing SPD-HSFPN-RVB

    CN118628907A

  • Model training method and intelligent construction site safety equipment detection method

    CN119131349A

Cited By

  • Small target detection device and method based on lightweight backbone network GhostaneNet and feature fusion network OfficientRepGFPN-EEP

    CN120852939A

  • Online monitoring system for automatic operation of distribution network

    CN121253982A