Blood cell image detection method based on YOLOv8 improved model
By introducing the Swin Transformer module, SimSPPF module, SGE attention mechanism and multi-scale fusion module in the YOLOv8 model, the problem of micro blood cell detection and image overlap in the blood cell detection is solved, and efficient and accurate blood cell detection is achieved.
Patent Information
- Application Number
- CN202510038952.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
Existing blood cell detection technology is difficult to accurately detect overlap problems in tiny blood cells and process blood cell images, resulting in missed and misdetection.
The Swin Transformer module and SimSPPF module were introduced into the backbone network of the YOLOv8 model, and the SGE attention mechanism and multi-scale fusion module were introduced into the neck network to improve the model's ability to extract blood cell characteristics.
The lightweight and high-performance balance of the model is achieved, which significantly improves the accuracy and efficiency of blood cell detection and reduces the missed detection of micro blood cells.
Smart Images

Figure CN119942540A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a blood cell image detection method based on an improved YOLOv8 model. Background Art
[0002] In recent years, with the rapid development of medical technology and the widespread application of precision instruments, blood cell testing has played an important role in clinical medical diagnosis. Blood cells are an important indicator of human health, and abnormalities in their number, shape, and function can directly reflect the presence of a variety of diseases. With the deepening of medical research and the updating of detection technology, the complexity of blood cell samples has increased, the cell morphology has become more diverse, and abnormal cells related to certain diseases may be very small. Although these subtle changes are difficult to observe directly under a microscope, they may contain important diagnostic information, which is directly related to the accuracy of disease diagnosis and the formulation of treatment plans. In order to capture and accurately analyze these subtle blood cell changes in a timely manner, accurate and efficient blood cell testing has become an urgent need in medical research and clinical diagnosis.
[0003] In the field of blood cell detection, many scholars have applied and improved the YOLO algorithm to improve the accuracy and precision of detection. Rohazia et al. replaced the backbone network of YOLOv3 with Alexnet to achieve efficient extraction of blood cell features and achieved an average accuracy of 98% on the blood cell dataset LISC. Tarimo et al. combined the Vision Transformer (ViT) based on the YOLOv5 algorithm, effectively integrated convolution and Transformer, and performed blood cell detection on a private dataset, achieving an accuracy of 96.49%. Wang et al. proposed an improved YOLOv7 algorithm for blood cell detection tasks. Based on the YOLOv7 algorithm, the algorithm embeds an attention mechanism in the neck (Neck), and uses the Focal-loss loss function instead of the original cross entropy loss function to improve the ability to detect small targets. The algorithm has an average accuracy of 66.32% on the detection task of the general dataset BCCD. In order to solve the problem of global dependence on long-range features in target detection, Kang et al. integrated the CNN-Swin Transformer module based on YOLOv7, enhanced the receptive field of the model, and better extracted the information of target features. They also achieved an accuracy of 91.1% on the public blood cell dataset BCCD.
[0004] At the same time, many scholars have studied how to improve the YOLO series of target detection algorithms in a lightweight way. Xu et al. proposed a lightweight detector TF-YOLOF for detecting blood cells, which uses EfficientNet as the backbone network for feature extraction, reducing the number of parameters while maintaining accuracy. In order to solve the problem of missed detection and false detection caused by the high density of blood cells, He et al. proposed a blood cell detector based on the improved YOLOv5s, which integrates Transformer in the backbone network and CBAM attention mechanism in the neck network, reducing the amount of calculation of the model to 1 / 6 of the original. Liu et al. made a lightweight algorithm based on the YOLOv8 algorithm and proposed the ADA-YOLO target detector, which combines the attention mechanism and designs an adaptive head. It is not only more accurate than YOLOv8 on the public blood cell dataset BCCD, but also uses more than three times less space than YOLOv8.
[0005] Although research in the field of blood cells has achieved certain results, it still faces many difficulties and challenges. The first problem is the diversity and complexity of blood cell morphology and types. How to accurately extract effective features is still a difficult problem that needs to be overcome. Secondly, since there are many overlaps of blood cells in the image, repeated detection is prone to occur. There is still room for improvement in the efficient and accurate detection of blood cells. Summary of the invention
[0006] The present invention aims to improve the accuracy and efficiency of blood cell detection and to prevent missed detection and misdetection of tiny blood cells.
[0007] In order to solve the above technical problems, the first aspect of the present invention proposes a blood cell image detection method based on the YOLOv8 improved model, the method comprising: obtaining blood cell image data to be detected, inputting the data into the YOLOv8 improved model, and obtaining a blood cell detection result;
[0008] The YOLOv8 improved model includes a backbone network backbone, a neck network neck and a detection head head; the backbone network includes a first convolution module CBS, a second convolution module CBS, a C2f module, a third convolution module CBS, a first Swin Transformer module, a fourth convolution module CBS, a second Swin Transformer module, a fifth convolution module CBS, a third Swin Transformer module and a SimSPPF module connected in sequence;
[0009] The neck network includes a first upsampling module Upsample, a first SGE_C2f module, a second upsampling module Upsample, a second SGE_C2f module, a sixth convolution module convolution CBS, a third SGE_C2f module, a first spatial attention module SGE, a seventh convolution module CBS, a fourth SGE_C2f module, and a second spatial attention module SGE;
[0010] The detection head includes three-scale decoupling heads, and the second SGE_C2f module, the first spatial attention module and the second spatial attention module are respectively connected to the three decoupling heads.
[0011] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method provided in the first aspect of the present invention.
[0012] In a third aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method provided in the first aspect of the present invention.
[0013] Beneficial effects of the present invention: First, the present invention introduces the Swin Transformer module and the SimSPPF module in the backbone network of the traditional YOLOv8 target detection model. Compared with the original C2f module and SPPF module in the traditional YOLOv8 backbone network, the present invention can not only reduce the number of model parameters but also improve the real-time detection speed of the model; secondly, the present invention introduces the SGE attention mechanism and the SGE_C2f module designed based on the SGE attention mechanism in the neck of the traditional YOLOv8 model, and also introduces the designed Multi-scale Fusion Neck (MFN) module, which can improve the feature extraction capability of blood cells while ensuring that the number of model parameters remains unchanged. The present invention has great improvements in the number of model parameters and detection indicators, and achieves a balance between lightweight and high performance compared to the original model, and significantly improves the accuracy of blood cell detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A schematic diagram of the network structure of the YOLOv8 improved model in an embodiment of the present invention;
[0015] Figure 2 This is a schematic diagram of the Swin Transformer module structure in an embodiment of the present invention;
[0016] Figure 3 Schematic diagram of the structure of the SimSPPF module in an embodiment of the present invention;
[0017] Figure 4 Schematic diagram of the workflow of the SGE attention mechanism module in an embodiment of the present invention;
[0018] Figure 5 This is a schematic diagram of the structure of the SGE_C2f module in an embodiment of the present invention;
[0019] Figure 6 Schematic diagram of the working process of the MFN module in an embodiment of the present invention;
[0020] Figure 7 This is a comparison chart of the differences between the existing YOLOv8 model and the actual detection experiment of the present invention;
[0021] Wherein, 1-24: represents the network layer sequence number of the YOLOv8 improved model of the present invention, which is used to distinguish the network layer of the model;
[0022] Input: input; Output: output; head: detection head;
[0023] The CBS module includes: Conv (convolution layer), batch normalization layer (Batch Normalization, BN) and activation function SiLU; it can realize feature extraction, normalization processing and nonlinear activation;
[0024] C2f (CSP Bottleneck with 2Convolutions): used to achieve cross-stage partial aggregation;
[0025] Swin Transformer: Moving window transformer;
[0026] SimSPPF (Simplified Spatial Pyramid Pooling–Fast): Simplified spatial pyramid;
[0027] Upsample: Upsampling; Concat: Concatenation;
[0028] SGE (Spatial Group-wise Enhance): spatial grouping enhancement;
[0029] SGE_C2f (Spatial Group-wise Enhance_CSP Bottleneck with2Convolutions): cross-segment aggregation with spatial grouping enhancement;
[0030] LN (Layer Normlization): layer normalization;
[0031] MLP (Multilayer Perceptron): Multilayer Perceptron;
[0032] W-MSA (Window Multi-head Self Attention): Window multi-head self-attention;
[0033] SW-MSA (Shifted Window Multi-head Self Attention): Shifted Window Multi-head Self Attention;
[0034] Global Average Pooling: Global average pooling;
[0035] Normalization: Normalization;
[0036] Position-wise Dot Product: Position-wise dot product;
[0037] Sigmoid: activation function;
[0038] SGE_Bottleneck (Spatial Group-wise Enhance_Bottleneck): Bottleneck spatial grouping enhancement. DETAILED DESCRIPTION
[0039] The terms "first", "second", "third", "fourth", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances. This is just a way of distinguishing objects with the same attributes when describing the embodiments of this application.
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] Figure 1 Schematic diagram of the network structure of the YOLOv8 improved model in an embodiment of the present invention.
[0042] The embodiment of the present invention proposes a blood cell image detection method based on the improved YOLOv8 model. Figure 1As shown, the method includes: obtaining blood cell image data to be detected, inputting the image data into the trained YOLOv8 improved model, and obtaining blood cell detection results;
[0043] The YOLOv8 improved model includes a backbone network backbone, a neck network neck and a detection head head;
[0044] The backbone network backbone Figure 1 The first 10 layers (i.e., layers 1-10) are mainly used for extracting shallow features, which include the first convolution module CBS, the second convolution module CBS, the C2f module, the third convolution module CBS, the first moving window transformer module Swin Transformer, the fourth convolution module CBS, the second moving window transformer module Swin Transformer, the fifth convolution module CBS, the third moving window transformer module Swin Transformer, and the simplified spatial pyramid module SimSPPF, which are connected in sequence;
[0045] The neck network Neck Figure 1 There are 11 to 24 layers in it, which include the first upsampling module, the first SGE_C2f module, the second upsampling module, the second SGE_C2f module, the sixth convolution module convolution CBS, the third SGE_C2f module, the first spatial attention module SGE, the seventh convolution module CBS, the fourth SGE_C2f module, and the second spatial attention module SGE.
[0046] The first SGE_C2f module and the first aggregation module are connected to the third aggregation module. The first Swin Transformer module and the third convolution module in the backbone network are connected to the second aggregation module, the second Swin Transformer module and the fourth convolution module are connected to the first aggregation module, the SimSPPF module is connected to the first upsampling layer and the fourth aggregation module at the same time, and the third Swin Transformer module is also connected to the fourth aggregation module. The first, second and fourth aggregation modules are used to extract shallow features of three scales in the backbone network for obtaining feature information in the backbone network.
[0047] This patent has made relevant improvements to the connection structure and bottleneck module in Neck, enhanced the ability to extract features of different scales to achieve improved model performance, introduced the SGE attention mechanism and designed the SGE_C2f module based on the SGE attention mechanism, and improved the detection ability of tiny blood cells.
[0048] The detection head Head includes three-scale decoupling heads for predicting the type and position of blood cells, and the second SGE_C2f module, the first spatial attention module SGE and the second spatial attention module SGE are respectively connected to the three decoupling heads.
[0049] The embodiment of the present invention introduces the Swin Transformer module, which can reduce the number of YOLOv8 model parameters due to its self-attention mechanism and hierarchical design, while more effectively capturing contextual information and fine-grained features in the image.
[0050] Figure 2 Schematic diagram of the Swin Transformer module structure in an embodiment of the present invention.
[0051] The Swin Transformer module includes a first Swin Tansformer module, a second Swin Tansformer module and a third Swin Tansformer module.
[0052] Reference Figure 2 As shown, the Swin Transformer module workflow is as follows:
[0053] S101: The input feature map passes through two LN (Layer Normlization) layers, MLP (Multilayer Perceptron) and a W-MSA (Window Multi-head Self Attention) attention mechanism to extract shallow features.
[0054] Specifically, the input feature map is first normalized, then the processed feature map is uniformly divided into 2×2 windows, and then self-attention learning is performed in each window. Input feature z l-1 exist Figure 2 The extraction calculation formula on the left is as follows:
[0055]
[0056]
[0057] in, represents the features after W-MSA attention learning, z l Represents the output features after MLP.
[0058] S102: Input the shallow features into two LN (Layer Normlization) layers, MLP (Multilayer Perceptron) and a SW-MSA (Shifted Window Multi-head Self Attention) attention mechanism to extract deep features.
[0059] Specifically, firstly, the feature z l The input is sent to the normalization layer for normalization, and then the 2×2 window divided in S101 is shifted downward and to the right, the feature information between different windows is fused, and then self-attention learning is performed on the fused window.
[0060] The features calculated by SW-MSA are input into the normalization layer for normalization, and finally input into the SW-MSA attention layer for attention learning. l exist Figure 2 The extraction calculation formula on the right is as follows:
[0061]
[0062] in, represents the features after SW-MSA attention learning, z l+1 Represents the output features of the multi-layer perceptron.
[0063] Figure 3 It is a schematic diagram of the structure of the SimSPPF module in an embodiment of the present invention.
[0064] Reference Figure 3 As shown in the figure, the workflow of the SimSPPF module includes:
[0065] The input feature map first undergoes feature transformation through the first convolutional layer CBR to reduce the number of channels of the feature map, then performs multiple pooling operations through the maximum pooling layer to capture the features, and concatenates the current pooling result with the previous pooling result, and finally obtains the output feature map through the second convolutional layer CBR.
[0066] The SimSPPF module reduces the computational complexity and improves the computational efficiency by serially processing multiple 5×5 scale maximum pooling operations.
[0067] The embodiment of the present invention introduces the SGE attention mechanism into the neck network of the existing YOLOv8. The SGE attention adjusts the importance of each sub-feature by generating an attention factor for each spatial position in each semantic group. This mechanism enables the model to pay more attention to the areas in the image that are most important to the task, thereby enhancing the learning ability of semantic features.
[0068] Figure 4 Schematic diagram of the workflow of the SGE attention mechanism module in an embodiment of the present invention.
[0069] Reference Figure 4 As shown in the figure, the workflow of the SGE attention mechanism module is as follows:
[0070] S201: Grouping the input feature maps along the channel dimension, each group of feature maps contains sub-features representing specific semantic information;
[0071] For a given H×W×C feature map, it is divided into G groups along the channel dimension C, and each group of features is represented by a specified vector in space X = {x 1m}, m=H×W.
[0072] S202: Generate each group of initial attention weights by calculating the global features and each group of feature vectors;
[0073] Through the spatial average function F gp (.) to obtain the global feature. The calculation formula of the global feature g is:
[0074]
[0075] The global feature g and the local feature x i Perform dot multiplication to get the initial attention weight c i , and its calculation formula is:
[0076] c i =g·x i (6)
[0077] S203: Use the activation function to perform linear transformation and then multiply the original feature vector to obtain an enhanced feature vector.
[0078] Normalize the initial attention weights to get The local feature x i Multiply it with the activation function σ(.) to get the enhanced vector m=H×W, and its calculation formula is:
[0079]
[0080] Among them, γ and β represent different parameters introduced for scaling and shift normalization.
[0081] The present invention constructs the SGE_C2f module and the MFN module based on the SGE attention mechanism.
[0082] Figure 5 Schematic diagram of the structure of the SGE_C2f module in an embodiment of the present invention.
[0083] In a preferred embodiment, the SGE_C2f module includes an eighth convolution module CBS, a Split module, n SGE_Bottleneck modules, a Concat module and a ninth convolution module CBS.
[0084] The SGE_Bottleneck module includes two convolution modules CBS and a third SGE module connected in series and then connected in parallel with a fourth SGE module.
[0085] Specifically, refer to Figure 5 As shown, the workflow of the SGE_C2f module is:
[0086] First, the input feature map passes through a common convolution module CBS to double the number of channels of the original one, and then the feature map with the changed number of channels is grouped;
[0087] Each set of feature maps is passed through the SGE_Bottleneck module to gradually extract more detailed features;
[0088] The extracted multiple detail features are concatenated with the original feature map through Concat;
[0089] The concatenated feature map is passed through ordinary convolution CBS to compress the number of channels and output the feature map of the target number of channels.
[0090] The embodiment of the present invention uses the SGE_C2f module to replace the original C2f module in the YOLOv8 neck module, thereby further improving the accuracy of the YOLOv8 model.
[0091] like Figure 1 As shown in the figure, the existing YOLOv8 connection method is: the Swin Transformer module of the fifth layer is connected to the Concat module of the fifteenth layer; the Swin Transformer module of the seventh layer is connected to the Concat module of the twelfth layer; the SimSPPF module of the tenth layer is connected to the Concat module of the 22nd layer; the SGE_C2f module of the 13th layer is connected to the Concat module of the 18th layer.
[0092] The existing YOLOv8 model is connected by connecting two modules, lacking adjacent layer connections and cross-scale connections. Not only is it difficult to fully integrate features of different scales, but it is also difficult to extract the features of tiny blood cells, which is prone to missed detection.
[0093] Therefore, the embodiment of the present invention improves the connection mode and designs an MFN module. The MFN module is not a conventional module, but a connection mode.
[0094] Figure 6 FIG. 4 is a schematic diagram of the working process of the MFN module in an embodiment of the present invention.
[0095] In a preferred embodiment, referring to Figure 6 As shown in the figure, the workflow of the MFN module includes: fusing the Swin Transformer module in the backbone network of the YOLOv8 model with its adjacent convolutional module CBS and adjacent SimSPPF module, and then connecting them to the SGE_C2f module of the neck network, and performing cross-scale fusion on the specified Swin Transformer module.
[0096] Specifically, refer to Figure 6 As shown, the connection formula of the MFN module is:
[0097] SGE_C2f4=concat(CBS2,SwinTransformer2,Upsample4) (8)
[0098] SGE_C2f5=concat(CBS1,SwinTransformer1,Upsample5) (9)
[0099] SGE_C2f6=concat(CBS6,SwinTransformer2,SGE_C2f4) (10)
[0100] SGE_C2f7=concat(CBS7,SwinTransformer3,SimSPPF) (11)
[0101] As shown in formula (8), the CBS2, Swin Transformer2, and Upsample4 modules are fused and then connected to the SGE_C2f4 module; as shown in formula (9), the CBS1, Swin Transformer1, and Upsample5 modules are fused and then connected to the SGE_C2f5 module; as shown in formula (10), the CBS6 and Swin Transformer2 modules are first fused, then connected to the SGE_C2f4, and finally connected to the SGE_C2f6 module; as shown in formula (11), the CBS27, Swin Transformer3, and SimSPPF modules are fused and then connected to the SGE_C2f7.
[0102] The MFN module enhances the semantic information of features through cross-scale connections and adjacent layer connections, which enhances the feature extraction capability of small cells and can reduce the missed detection of tiny blood cells.
[0103] In a preferred embodiment, the training process of the YOLOv8 improved model includes:
[0104] S301: Acquire a blood cell detection image sample.
[0105] The blood cell detection image samples come from the BCCD public dataset, and the dataset labels are converted into txt training format using homemade Python code. The BCCD dataset is divided into training set, test set, and validation set in a ratio of 8:1:1. The dataset has a total of 364 images and has three categories, including red blood cells (RBC), white blood cells (WBC), and platelets.
[0106] S302: Input the blood cell detection image sample into the YOLOv8 improved model to obtain the blood cell detection result.
[0107] S303: Calculate the loss between the bounding box in the blood cell detection result and the bounding box of the true label of the blood cell detection image sample, and optimize the YOLOv8 improved model through back propagation of the loss.
[0108] The CIOU loss function is used to calculate the loss between the bounding box in the blood cell detection result and the bounding box of the true label of the blood cell detection image sample.
[0109] The calculation formula of the CIOU loss function is:
[0110]
[0111] Among them, IoU represents the overlapping area between the blood cell detection box and the real box in the blood cell label, b represents the predicted box, and b gt represents the true box, ρ represents the Euclidean distance from the center of the detection box to the center of the true box, and c represents the diagonal distance of the minimum rectangle of the detection box and the true box. α is the coefficient related to IoU, and its calculation formula is:
[0112]
[0113] Among them, v is a constant, and its calculation formula is:
[0114]
[0115] Among them, w gt Indicates the width of the real box, h gt represents the height of the real box, w represents the width of the detection box, and h represents the height of the detection box.
[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method provided in the first aspect of the present invention are implemented.
[0117] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of the method provided in the first aspect of the present invention when executed by a processor.
[0118] Experimental verification:
[0119] 1. Experimental Configuration
[0120] The parameter configuration table of the model training of the present invention is shown in Table 1, and the development environment of the present invention is shown in Table 2:
[0121] Table 1 Training parameter configuration table
[0122] Optimizer SGD Batch size (Batch_size) 16 Epochs 300 Learning rate (Learning_rate) 0.01 Input image size (Input_size) 640×480
[0123] Table 2 Experimental environment
[0124] System Ubuntu 22.04.1 Python 3.9.18 CPU 12th Gen Intel(R)Core(TM)i5-12400 CPU NVIDIA GeForce RTX 4070ti (16G) Cuda 11.8 Memory 16G
[0125] In order to verify the effectiveness of the optimization method adopted in this experiment, the present invention sets up an ablation experiment and a comparative experiment.
[0126] 2. Performance indicators
[0127] In the YOLO series of models, the main indicators for evaluating network performance are as follows: Precision (P), Recall (R), and Mean Average Precision (mAP). This experiment uses mAP50 and mAP50:95 as performance reference indicators. mAP50 and mAP50:95 represent the mAP value when the intersection-over-union ratio between the predicted box and the true box is 0.5 and the average mAP value when the intersection-over-union ratio is 0.5 to 0.95, respectively. The larger the mean average precision mAP, the higher the overall accuracy of the model. The calculation formulas for each indicator are as follows:
[0128]
[0129] Among them, P represents accuracy, R represents recall, AP represents precision, mAP represents average precision, TP represents the number of correctly predicted positive samples, FN represents the number of incorrectly predicted negative samples, and FP represents the number of incorrectly predicted positive samples.
[0130] When discussing model performance, the number of model parameters must also be considered, so it is also necessary to introduce the parameter quantity (Params) model architecture detail parameters.
[0131] 3. Method comparison
[0132] The present invention adopts the method of ablation experiment and comparative experiment to verify the effectiveness of the improved algorithm step by step. The ablation experiment is shown in Table 4, and the comparative experiment is shown in Table 5.
[0133] First, YOLOv8 is used as the baseline algorithm in the public dataset BCCD (the BCCD dataset is available at https: / / www.kaggle.com / datasets / konstantinazov / bccd-dataset ). After the experiment, it was found that YOLOv8 has a better detection effect on large-scale blood cells, but there is still room for improvement in the detection performance of tiny blood cells. Therefore, the spatial attention mechanism is introduced for improvement.
[0134] 3.1 Ablation Experiment
[0135] Different modules are added to the traditional YOLOv8 network (i.e., YOLOv8 model) to check the effectiveness of different modules. The test results are shown in Table 4.
[0136] Only add the Swin Transformer module to the backbone network: After replacing the C2f module in the Backbone with the Swin Transformer module, as shown in Table 4, not only the Params are reduced by 6.7%, but also the mAP50 and mAP50:95 of the model performance are improved by 0.6% and 1%, which fully proves the effectiveness of the Swin Transformer module on YOLOv8.
[0137] Only the SPPF module is added to the backbone network Backbone: After replacing the SPPF module in the Backbone with the SimSPPF module, as shown in Table 4: the mAP50 and mAP50:95 of the model performance are slightly reduced, indicating that the simple structure of SimSPPF has a slight impact on the detection accuracy, but as shown in Table 3, SimSPPF can improve the real-time detection speed of the model. After adding the SimSPPF module, the FPS of the model is much greater than the traditional YOLOv8 model.
[0138] Table 3 Comparison between SimSPPF and SPPF
[0139]
[0140] Two SGE attention mechanisms are introduced in the neck: As shown in Table 4, the mAP50 of the model is improved by 0.7%, and the mAP50:95 of the model is improved by 1.3%, indicating that after adding attention, the model's ability to detect tiny cells has been significantly improved, and because of the lightweight characteristics of SGE attention, the number of model parameters remains unchanged, indicating the feasibility of SGE attention on YOLOv8.
[0141] Introducing the designed SGE_C2f module in the neck: As shown in Table 4, the SGE_C2f module is used to replace the C2f module in the neck. The model size is slightly increased, but the mAP50 and mAP50:95 of the model are both improved by 0.3%, indicating that the SGE_C2f module increases the model's ability to extract blood cells, indicating the feasibility of the SGE_C2f module on the YOLOv8 model.
[0142] The Swin Transformer module and the SimSPPF module are introduced into the traditional YOLOv8 model at the same time: As shown in Table 4, the model performance mAP50 is improved by 1.8%, and mAP50:95 is improved by 5.1%, which greatly improves the detection performance of the model, fully proving the ability of the Swin Transformer module and the SimSPPF module in enhancing the extraction of blood cell features.
[0143] MFN module introduced in the neck: As shown in Table 4, the mAP50 and mAP50:95 of the model increased by 0.1% and 0.4%, respectively, indicating that the feature extraction capability of the model can be increased by increasing the connection between adjacent layers in the traditional YOLOv8 model.
[0144] The Swin Transformer module, SimSPPF module, SGE module, SGE_C2f module and MFN module are introduced into the traditional YOLOv8 model. The results are shown in Ours (algorithm of the present invention) in Table 4: the mAP50 and mAP50:95 of the model are greatly improved, and the mAP50 and mAP50:95 of the model are improved by 4.4% and 7.8% respectively, indicating that the combined introduction of multiple modules is very effective for the YOLOv8 model, which greatly improves the model's ability to detect blood cells.
[0145] Table 4. Comparison of ablation experiments
[0146]
[0147] 3.2 Comparative Experiment
[0148] Under the premise of keeping the experimental environment unchanged, the algorithm of the present invention is compared with the commonly used target detection algorithms and the latest methods in terms of performance and parameter scale. These algorithms are Fast R-CNN, SSD, RT-DETR, YOLOv5, YOLOv7, YOLOv8 and YOLOv10, ADA-YOLO, and CST-YOLO. The comparison results are shown in Table 5.
[0149] 3.2.1 Comparison between the proposed algorithm and the classical algorithm
[0150] Compared with the classic target detection algorithms Faster-Rcnn, SSD and RT-DETR, the improved method proposed in the present invention has very significant advantages. The algorithm of the present invention is 16.8% higher than Faster-RCNN in mAP50:95, 21.3% higher than the SSD algorithm in mAP50:95, and 6.5% higher than RT-DETR in mAP50, and the number of parameters is much smaller than that of the classic algorithm. The number of parameters of the YOLOv8 improved model of the present invention is even less than one tenth of that of the classic algorithm, indicating that the algorithm of the present invention can obtain efficient detection results at a very low cost in the blood cell detection task.
[0151] 3.2.2 Comparison of the algorithm of this invention with other YOLO algorithms
[0152] As shown in Table 5, compared with the existing YOLOv5 and YOLOv7, the algorithm of the present invention has improved 4% and 8.1% in mAP50, 2.9% and 7.3% in mAP50:95, and the number of parameters is much smaller than that of YOLOv5 and YOLOv7. Although the Precision of YOLOv5 is slightly higher than that of the algorithm of the present invention, the detection ability of the model mainly depends on mAP50 and mAP50:95. Therefore, compared with the old version of the YOLO model, the algorithm of the present invention has greatly improved. Compared with the new version of the YOLOv10 model, although the Recall of the algorithm of the present invention has decreased by 0.3%, the mAP50 has increased by 7.1%, indicating that the improved YOLOv8 model of the present invention is more suitable for blood cell detection tasks than the new version of the YOLO model.
[0153] 3.2.3 Comparison between the proposed algorithm and the latest algorithm
[0154] As shown in Table 5, compared with the latest blood cell detection algorithms ADA-YOLO and CST-YOLO, the algorithm of the present invention has improved mAP50 by 2.6% and 1.1% respectively, indicating that the algorithm of the present invention has significant advantages over the latest algorithms in blood cell detection tasks, and the algorithm of the present invention has only 3M parameters, which is a lightweight network and is easier to deploy in practical applications. Compared with the traditional YOLOv8, the Precision, Recall, mAP50, and mAP50:95 of the model have increased by 5.5%, 5.6%, 4.4%, and 7.8% respectively, and the model parameters have also been reduced by about 4%, which fully demonstrates the effectiveness of the algorithm of the present invention compared to YOLOv8.
[0155] In summary, the present invention can well balance the lightweight of the model and the algorithm performance, and is also superior to some common target detection algorithms.
[0156] Table 5 Comparison of common target detection algorithms (the bold value is the maximum value of each indicator)
[0157]
[0158]
[0159] 4. Detection effect analysis
[0160] As mentioned above, the YOLOv8 model has certain deficiencies in the detection effect of blood cells. After being improved by the present invention, the detection effect of the present invention (Ours) on blood cells has been effectively improved, and its performance comparison is shown in Table 6. The detection ability of tiny platelets has been greatly improved, with an increase of 9.1% in mAP50, and the detection ability of red blood cells (RBC) and white blood cells (WBC) has also been improved.
[0161] Table 6 Comparison of blood cell detection results
[0162]
[0163] Figure 7 This is a comparison chart of the differences between the existing YOLOv8 model and the actual detection experiment of the present invention. Figure 7 middle, Figure 7 -a and Figure 7 -b is a comparison of two groups of pictures. The white box represents the detection box of white blood cells (WBC), the light blue box represents the detection box of red blood cells (RBC), and the blue box represents the detection box of platelets. Figure 7The comparison results show that the YOLOv8 model missed three platelets and one white blood cell, but did not miss any red blood cells. The model of the present invention did not miss any of the three blood cells. Figure 7 From the comparison results, it can be seen that the YOLOv8 model missed one platelet cell, but did not miss any white blood cells (WBC) or red blood cells (RBC). Moreover, white blood cells can be detected. The model of the present invention did not miss any of the three blood cells, and the confidence of the detection results of the present invention was significantly higher than that of YOLOv8.
[0164] In summary, the improved YOLOv8 model in the present invention has a very high detection accuracy for red blood cells, white blood cells and platelets, which solves the problem that YOLOv8 easily misses platelets, and can also detect blood cells that are not manually labeled. The improved YOLOv8 model of the present invention greatly improves the detection accuracy of blood cells, and maintains a good detection speed and a small model size, further illustrating the efficiency of the model of the present invention. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: ROM, RAM, disk or CD, etc.
[0165] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A blood cell image detection method based on the improved YOLOv8 model, characterized in that: include: Obtain the blood cell image data to be detected, input it into the trained YOLOv8 improved model, and obtain the blood cell detection results; The YOLOv8 improved model includes a backbone network, a neck network and a detection head; The backbone network includes a first convolution module, a second convolution module, a C2f module, a third convolution module, a first moving window transformer module, a fourth convolution module, a second moving window transformer module, a fifth convolution module, a third moving window transformer module, and a simplified spatial pyramid module, which are connected in sequence; The neck network includes a first upsampling module, a first SGE_C2f module, a second upsampling module, a second SGE_C2f module, a sixth convolution module, a third SGE_C2f module, a first spatial attention module, a seventh convolution module, a fourth SGE_C2f module, and a second spatial attention module; The detection head includes three-scale decoupling heads, and the second SGE_C2f module, the first spatial attention module and the second spatial attention module are respectively connected to the three decoupling heads.
2. The blood cell image detection method based on the improved YOLOv8 model according to claim 1, characterized in that: The Swin Transformer module workflow includes: The input feature map passes through two layers of normalization LN, multi-layer perceptron MLP and window multi-head self-attention module W-MSA to extract shallow features; The extracted shallow features are input into two layer normalization layers LN, a multi-layer perceptron MLP and a sliding window multi-head self-attention module SW-MSA to extract deep features.
3. The blood cell image detection method based on the improved YOLOv8 model according to claim 1, characterized in that: The workflow of the SimSPPF module is specifically as follows: The input feature map is first transformed through the first convolutional layer CBR to reduce the number of channels of the feature map, and then multiple pooling operations are performed through the maximum pooling layer to capture the features, and the current pooling result is concatenated with the previous pooling result, and finally the output feature map is obtained through the second convolutional layer CBR.
4. The blood cell image detection method based on the improved YOLOv8 model according to claim 1, characterized in that: The SGE_C2f module includes an eighth convolution module CBS, a Split module, n SGE_Bottleneck modules, a Concat module and a ninth convolution module CBS.
5. The blood cell image detection method based on the improved YOLOv8 model according to claim 4, characterized in that: The SGE_Bottleneck module includes two convolution modules CBS and one SGE module connected in series and then connected in parallel with another SGE module.
6. The blood cell image detection method based on the improved YOLOv8 model according to claim 1, characterized in that: The workflow of the MFN module includes: fusing the Swin Transformer module in the backbone network of the YOLOv8 improved model with its adjacent convolutional module CBS and adjacent SimSPPF module, connecting them to the SGE_C2f module of the neck network, and performing cross-scale fusion on the specified Swin Transformer module.
7. The blood cell image detection method based on the improved YOLOv8 model according to claim 1, characterized in that: The training process of the YOLOv8 improved model includes: Acquire a blood cell detection image sample; The blood cell detection image samples are input into the YOLOv8 improved model to obtain the blood cell detection results; The loss between the bounding box in the blood cell detection result and the bounding box of the true label of the blood cell detection image sample is calculated, and the YOLOv8 improved model is optimized by back propagation through the loss.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Imaging flow cytometry cell detection method based on improved model
CN121305149A