Blood cell detection method based on improved YOLOv11
By improving the YOLOv11 model, replacing the C3k2 module with the C3k2_PConv module and introducing the ADown module, optimizing feature extraction and fusion, the parameters redundancy and computational complexity problems of the deep learning model in blood cell detection are solved, and efficient and lightweight blood cell detection is achieved, which is suitable for edge device deployment.
Patent Information
- Application Number
- CN202510438825.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
Existing deep learning models have problems such as parameter redundancy, high computational complexity and difficulty in deploying edge devices in blood cell detection. Traditional methods are inefficient, limited accuracy and difficult to adapt to complex samples.
The improved YOLOv11 model is adopted, and by replacing the traditional C3k2 module with the C3k2_PConv module, and introducing the ADown module to replace the convolution blocks of the traditional downsampling layer, the backbone network, the neck network and the prediction head network are built, and feature extraction and fusion are optimized, the model parameters are reduced and the detection speed is improved.
While maintaining detection accuracy, the number of model parameters and volume are significantly reduced, the detection speed is improved, and the edge device deployment capability is stronger, and efficient and lightweight blood cell detection is achieved.
Smart Images

Figure CN120339234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a blood cell detection method based on an improved YOLOv11, belonging to the field of image recognition technology. Background Art
[0002] In biomedical research, the automatic detection of blood cell images under a microscope is a key step in disease diagnosis, cell classification, and pathological research. Traditional methods rely on manual annotation or algorithms based on handcrafted features, suffering from limitations such as low efficiency, limited accuracy, and difficulty in adapting to complex samples.
[0003] In recent years, the rise of deep learning technology has brought new breakthroughs to the field of blood cell detection. Deep learning models can automatically learn features in data, avoiding the cumbersome process of handcrafted feature extraction in traditional methods, and significantly improving the efficiency and accuracy of cell detection. YOLO is a classic among them. However, existing models often face challenges such as parameter redundancy, high computational complexity, and difficulty in deploying on edge devices. Summary of the Invention
[0004] In order to overcome the above defects in the prior art, the present invention provides a blood cell detection method based on an improved YOLOv11. The specific steps to implement this method include:
[0005] S1: For the collected blood cell images, expand the blood cell set through data augmentation, annotate the blood cells, and divide the augmented data set into a training set and a validation set;
[0006] S2: Construct an object detection model for detecting and recognizing blood cells, including a backbone network, a neck network, and a prediction head network;
[0007] S3: Input the augmented data set into the object detection model for feature extraction;
[0008] S4: Adjust the parameters and methods in the model training process, and train the object detection model with the method in step S3 and the data set obtained in step S1. Stop training when the set epoch is reached;
[0009] S5: Put the data set into the model for testing, and evaluate the model through evaluation metrics;
[0010] Further, in step S1: The blood cell images include red blood cells, white blood cells, and platelets, and the quantity is expanded by methods such as rotation, flipping, and adding noise. Annotate the positions and categories of all blood cells in the images through an annotation tool.
[0011] Furthermore, in step S2: The backbone network includes a convolutional block, four groups of ADown modules and C3k2_PConv modules, an SPPF module, and a C2PSA module. The four groups of ADown modules and C3k2_PConv modules are defined as the first, second, third, and fourth groups of ADown modules and C3k2_PConv modules, and each group includes an ADown module and a C3k2_PConv module. The ADown module uses average pooling and max pooling to downsample the input and processes the features through two convolutional layers, which can significantly reduce the number of parameters while maintaining or even improving the accuracy of object detection. The C3k2_PConv module is based on the traditional C3k2 module, replacing the Conv convolution in the traditional C3k2 module with PConv convolution. It achieves fast and efficient operations by applying filters only on a few input channels while keeping other channels unchanged. The neck network includes ADown modules and C3k2_PConv modules, upsampling modules, and combination layers. There are four C3k2_PConv modules in the neck network, which are respectively the first, second, third, and fourth C3k2_PConv modules. There are two ADown modules in the neck network, which are respectively the first and second ADown modules. There are four combination layers in the neck network, which are respectively the first, second, third, and fourth combination layers. There are two upsampling modules in the neck network, which are respectively the first and second upsampling modules. The neck network is used to achieve deep fusion of features at different levels.
[0012] Further, in step S3: The image features input into the YOLOv11 model architecture first enter the backbone network and are processed as follows: The image features first pass through a convolutional block for feature extraction, and then sequentially pass through the first group, the second group, the third group, and the fourth group of ADown modules and the C3k2_PConv module for feature extraction. After the output features pass through the SPPF module, they finally pass through the C2PSA module for output; The output features of the C2PSA module first pass through the first upsampling module and enter the first combination layer, where they are fused with the output features of the C3k2_PConv module in the third group of ADown modules and the C3k2_PConv module; The output features of the first combination layer sequentially pass through the first C3k2_PConv module and the second upsampling module; The output features of the second upsampling module enter the second combination layer, where they are fused with the output features of the C3k2_PConv module in the second group of ADown modules and the C3k2_PConv module; The output features of the second combination layer sequentially pass through the second C3k2_PConv module and the first ADown module; The output features of the first ADown module enter the third combination layer, where they are fused with the output features of the first combination layer; The output features of the third combination layer sequentially pass through the third C3k2_PConv module and the second ADown module; The output features of the second ADown module enter the fourth combination layer, where they are fused with the output features of the C2PSA module, and the output features of the fourth combination layer enter the fourth C3k2_PConv module.
[0013] Further, in step S4: After setting the parameters and methods in the model training process, set the epoch to 100, and train the object detection model using the method in step S3 and the dataset obtained in step S1.
[0014] Further, in step S5: Put the trained weights into the model, use the validation set for testing, and evaluate the overall detection performance of the model.
[0015] The present invention has achieved remarkable technological breakthroughs in the field of cell detection technology. By innovatively replacing the traditional C3k2 module with the improved C3k2_PConv module and introducing the ADown module to replace the convolutional block of the traditional downsampling layer, it realizes the dual optimization of significantly reducing the model parameter quantity and improving the detection speed on the premise of almost no loss of detection accuracy. The model can efficiently and accurately detect and identify the specific categories of three types of blood cells. This improvement effectively solves the problems of low efficiency and low recognition rate of traditional manual detection. At the same time, due to its fewer model parameters, smaller volume, and faster detection speed, it has stronger edge device deployment capabilities. Description of the Drawings
[0016] Figure 1 This is the flowchart of the blood cell detection method based on the improved YOLOv11 in the embodiments of the present invention;
[0017] Figure 2 This is the flowchart of the blood cell detection method based on the improved YOLOv11n-C3k2_PConv-ADown in the embodiments of the present invention;
[0018] Figure 3 This is the network structure diagram of the ADown module in the embodiments of the present invention;
[0019] Figure 4 This is the network structure diagram of the C3k2_PConv module in the embodiments of the present invention; Detailed implementation manners
[0020] To make the advantages, technical solutions, etc. of the present invention clearer, the following will be described.
[0021] It is not considered that this description is a limitation to the present invention.
[0022] Embodiment 1
[0023] Based on the YOLOv11 (You Only Look Once version 11) algorithm, the present invention proposes an innovative blood cell detection method, aiming to optimize the detection and recognition of blood cells under a microscope. By cleverly replacing the traditional C3k2 module with an improved C3k2_PConv module and introducing an ADown module to replace the convolutional block of the traditional downsampling layer, while maintaining high detection accuracy, the model effectively reduces the number of model parameters and volume, and at the same time improves the detection speed. This technological innovation not only effectively solves the problems of low efficiency and low recognition rate of traditional manual detection, but also shows excellent edge device deployment capabilities due to its small number of model parameters, small volume and fast detection speed. Through this optimized design, the model can complete complex blood cell detection tasks with higher efficiency and lower resource consumption, providing an efficient and lightweight solution for the field of microscopic image analysis.
[0024] As Figure 1 shown, the blood cell detection method based on the improved YOLOv11 includes the following steps:
[0025] S1: First, for the collected blood cell images, the blood cell dataset is augmented through data augmentation techniques to improve the generalization ability and robustness of the model. On this basis, LabelImg (a deep learning image annotation tool) is used to accurately annotate three blood cell categories (platelets, red blood cells, white blood cells), generating annotation files (in txt format) with a unified size of 640×640 and Z-score normalization. The annotation files detail the position information and class labels of each blood cell, where the class definitions are: 0 represents platelets, 1 represents red blood cells (RBC), and 2 represents white blood cells (WBC). Finally, the augmented dataset is roughly divided into a training set and a validation set in a ratio of 9:1 to ensure the stability of model training and the reliability of validation. Then, a color channel conversion is performed to convert the images from RGB format to CHW format, i.e., channel-height-width format.
[0026] S2: Build an object detection model for detecting and identifying blood cells. Its network structure is as Figure 2 shown, including a backbone network, a neck network, and a prediction head network; the backbone network of the present invention includes a convolutional block, four groups of ADown modules and C3k2_PConv modules, an SPPF module, and a C2PSA module. These four groups of ADown modules and C3k2_PConv modules are respectively defined as the first group, the second group, the third group, and the fourth group, and each group contains an ADown module and a C3k2_PConv module. The ADown module realizes downsampling of the input features by combining average pooling and max pooling, and processes the features through two convolutional layers, thus significantly reducing the model parameter quantity while maintaining or even improving the accuracy of object detection. The C3k2_PConv module is an innovative improvement based on the traditional C3k2 module, replacing the traditional Conv convolution with PCon convolution. The PCon convolution realizes fast and efficient operations by applying filters only on a few input channels while keeping other channels unchanged. The neck network design further enhances the feature fusion ability, including ADown modules, C3k2_PConv modules, upsampling modules, and combination layers. Four C3k2_PConv modules (the first, second, third, and fourth C3k2_PConv modules) and two ADown modules (the first and second ADown modules) are defined in the neck network. In addition, the neck network also includes four combination layers (the first, second, third, and fourth combination layers) and two upsampling modules (the first and second upsampling modules). Through the collaborative action of these modules, the neck network can achieve deep fusion of features at different levels, providing rich feature information for subsequent detection. The detection head network contains three detection heads for outputting the final detection results.
[0027] S3: The input image features first enter the backbone network, and are first subjected to preliminary feature extraction through a convolutional block, and then flow through four groups of ADown modules and C3k2_PConv modules in sequence. The ADown module realizes downsampling through average pooling and max pooling, and at the same time refines the features using two convolutional layers to reduce the number of parameters. The image features input to the YOLOv11 model architecture first enter the backbone network and are processed as follows: The image features first pass through a convolutional block for feature extraction, and then pass through the first group, the second group, the third group, and the fourth group of ADown modules and C3k2_PConv modules in sequence for feature extraction. After the output features pass through the SPPF module, they finally pass through the C2PSA module for output; The output features of the C2PSA module first pass through the first upsampling module and enter the first combination layer, and are fused with the output features of the C3k2_PConv module in the C3k2_PConv module of the third group of ADown modules and C3k2_PConv modules in the first combination layer; The output features of the first combination layer pass through the first C3k2_PConv module and the second upsampling module in sequence; The output features of the second upsampling module enter the second combination layer and are fused with the output features of the C3k2_PConv module in the C3k2_PConv module of the second group of ADown modules and C3k2_PConv modules in the second combination layer; The output features of the second combination layer pass through the second C3k2_PConv module and the first ADown module in sequence; The output features of the first ADown module enter the third combination layer and are fused with the output features of the first combination layer in the third combination layer; The output features of the third combination layer pass through the third C3k2_PConv module and the second ADown module in sequence; The output features of the second ADown module enter the fourth combination layer and are fused with the output features of the C2PSA module in the fourth combination layer, and the output features of the fourth combination layer enter the fourth C3k2_PConv module.
[0028] As shown in S3 Figure 3 The described ADown module includes an average pooling layer, a data splitting layer, a max pooling layer, two convolutional layers, and a combination layer. Among them, the processing flow of the ADown module: The image features input to the ADown module first pass through an average pooling layer with a pooling window of 2*2 and a stride of 1, and then pass through a data splitting layer to evenly divide the channel dimension into two features and input them into two branches respectively; The features of one branch pass through an average pooling layer with a pooling window of 3*3 and a stride of 2, and then pass through a 1*1 convolution to maintain information integrity; The other branch passes through a 3*3 convolution with a stride of 2 to further reduce the dimension; Finally, feature fusion is performed through the combination layer.
[0029] As shown in S3 Figure 4As shown, the described C3k2_PConv module includes two convolutional blocks, a data splitting layer, two PConvs, and a combination layer. The C3k2_PConv module can divide the input features into two parts. One part is directly passed through ordinary convolutional operations, and the other part undergoes deep feature extraction through multiple PVConvs. Finally, the two parts of the features are concatenated and fused through 1*1 convolution. This structure can not only maintain light weight but also effectively extract deep features.
[0030] S4: After setting the parameters and methods in the model training process, set the epoch to 100, and train the object detection model using the method in step S3 and the dataset obtained in step S1.
[0031] S5: Put the trained weights into the model, use the validation set for testing, and evaluate the overall detection performance of the model.
[0032] S51: After training, evaluate the performance of the model. Precision, Recall, AP, and mAP are important indicators for evaluating the model. Compare with the original YOLOv11-n model.
[0033] S51: The calculation formulas for precision and recall are as follows:
[0034]
[0035]
[0036] Taking the identification of white blood cells as an example, in the formula, TP represents the number of correctly detected white blood cells, FP represents the number of non-white blood cells that are wrongly detected as white blood cells, and FN represents the number of white blood cells that are not detected.
[0037] The results are shown in Table 1.
[0038] Table 1 Result Comparison
[0039]
[0040] Analyzing the results in the above table, compared with the original model, the overall precision and recall of the improved new model have not significantly decreased.
[0041] S52: AP is the average precision, and mAP is the average precision of all blood cell types. The calculation formulas are as follows:
[0042]
[0043]
[0044] The results are shown in Table 2.
[0045] Table 2 Comparison of AP and mAP between the improved new model and the original model
[0046]
[0047] Analyzing the results in the above table, compared with the original model, the overall mAP of the improved new model did not decrease significantly.
[0048] S53: Comparison of the number of parameters (Parameter) and weight size (Weigh) between the improved new model and the original model
[0049] Table 3 Comparison of the number of parameters and weight size
[0050]
[0051] Analyzing the results in the above table, the number of parameters decreased by 30.2%, and the memory occupied by the obtained weights decreased by 29.1%, effectively reducing the number of model parameters and volume, and enhancing the deployment ability of edge devices.
[0052] S54: Comparison of the speed of preprocessing (Preprocess), inference stage (Inference), and postprocessing (Postprocess) per image between the improved new model and the original model
[0053] Table 4 Comparison of the speed of preprocessing, inference stage, and postprocessing
[0054]
[0055] Analyzing the results in the above table, the detection speed was improved. The new model was 32% faster than the old model in terms of total time consumption, which is helpful for real-time detection.
Claims
1. An improved YOLOv11-based blood cell detection method, characterized in that, It includes the following steps: S1: For the collected blood cell images, augment the blood cell set through data augmentation, label the blood cells, and divide the augmented data set into a training set and a validation set; S2: Construct an object detection model for detecting and recognizing blood cells, including a backbone network, a neck network, and a prediction head network; S3: Input the augmented data set into the object detection model for feature extraction; S4: Adjust the parameters and methods in the model training process, and train the object detection model with the methods in step S3 and the data set obtained in step S1. Stop training when the set epoch is reached; S5: Put the data set into the model for testing, and evaluate the model through evaluation metrics.
2. The blood cell detection method based on the improved YOLOv11 according to claim 1, wherein: In step S1: The blood cell images include red blood cells, white blood cells, and platelets, and the quantity is augmented by methods such as rotation, flipping, and adding noise. All blood cell positions and categories in the images are labeled through a labeling tool.
3. The blood cell detection method based on the improved YOLOv11 according to claim 1, wherein: In step S2: The backbone network includes a convolutional block, four groups of ADown modules and C3k2_PConv modules, an SPPF module, and a C2PSA module; Define the four groups of ADown modules and C3k2_PConv modules as the first, second, third, and fourth groups of ADown modules and C3k2_PConv modules respectively, and each group includes an ADown module and a C3k2_PConv module; The ADown module uses average pooling and max pooling to downsample the input, and processes the features through two convolutional layers, which can significantly reduce the number of parameters while maintaining or even improving the accuracy of object detection; The C3k2_PConv module is based on the traditional C3k2 module, replacing the Conv convolution in the traditional C3k2 module with PConv convolution. It achieves fast and efficient operations by applying filters only on a few input channels while keeping other channels unchanged; The neck network includes ADown modules and C3k2_PConv modules, an upsampling module, and a combination layer; Define that there are four C3k2_PConv modules in the neck network, which are the first, second, third, and fourth C3k2_PConv modules respectively; define that there are two ADown modules in the neck network, which are the first and second ADown modules respectively; Define that there are four combination layers in the neck network, which are the first, second, third, and fourth combination layers respectively, and define that there are two upsampling modules in the neck network, which are the first and second upsampling modules respectively; The neck network is used to achieve deep fusion of features at different levels, and the detection head network contains three detection heads.
4. The blood cell detection method based on the improved YOLOv11 according to claim 1, characterized in that: In step S3: The image features input into the YOLOv11 model architecture first enter the backbone network and are processed as follows: The image features first go through a convolutional block for feature extraction, and then sequentially go through the first, second, third, and fourth groups of ADown modules and C3k2_PConv modules for feature extraction. After the output features pass through the SPPF module, they are finally output through the C2PSA module; The output features of the C2PSA module first pass through the first upsampling module and enter the first combination layer, where they are fused with the output features of the C3k2_PConv module in the third group of ADown modules and the C3k2_PConv module; The output features of the first combination layer sequentially pass through the first C3k2_PConv module and the second upsampling module; The output features of the second upsampling module enter the second combination layer, where they are fused with the output features of the C3k2_PConv module in the second group of ADown modules and the C3k2_PConv module; The output features of the second combination layer sequentially pass through the second C3k2_PConv module and the first ADown module; The output features of the first ADown module enter the third combination layer, where they are fused with the output features of the first combination layer; The output features of the third combination layer sequentially pass through the third C3k2_PConv module and the second ADown module; The output features of the second ADown module enter the fourth combination layer, where they are fused with the output features of the C2PSA module, and the output features of the fourth combination layer enter the fourth C3k2_PConv module.
5. A blood cell detection method based on the improved YOLOv11 according to claim 1, characterized in that: In step S3: The ADown module includes an average pooling layer, a data splitting layer, a max pooling layer, two convolutional layers, and a combination layer; Among them, the processing flow of the ADown module: The image features input to the ADown module first pass through an average pooling layer with a pooling window of 2*2 and a stride of 1, and then pass through a data splitting layer to evenly divide the channel dimension into two features and respectively input them into two branches; One of the branch features passes through an average pooling layer with a pooling window of 3*3 and a stride of 2, and then passes through a 1*1 convolution to maintain information integrity; The other passes through a 3*3 convolution with a stride of 2 to further reduce the dimension; Finally, feature fusion is performed through the combination layer.
6. The substation foreign object intrusion detection method based on the improved YOLOv11 model according to claim 5, characterized in that: Average pooling layer formula, that is: Max pooling formula, that is: X: Input feature map; Y: Output feature map; k: Size of the pooling window; s: Stride of the pooling window; i, j: Index of the output feature map; m, n: Index within the pooling window.
7. A blood cell detection method based on the improved YOLOv11 according to claim 1, characterized in that: In step S3: The C3k2_PConv module includes two convolutional blocks, a data splitting layer, two PConv, and a combination layer; The C3k2_PConv module can divide the input features into two parts. One part is directly passed through ordinary convolution operations, and the other part undergoes deep feature extraction through multiple PVConv. Finally, the two parts of the features are concatenated and fused through a 1*1 convolution. This structure can not only maintain light weight but also effectively extract deep features.
8. A blood cell detection method based on the improved YOLOv11 according to claim 1, characterized in that: In step S4: After setting the parameters and methods in the model training process, set the epoch to 100, and train the object detection model with the method in step S3 and the dataset obtained in step S1.
9. A blood cell detection method based on the improved YOLOv11 according to claim 1, wherein: In step S5: The trained model is tested using the validation set to evaluate the overall detection performance of the model.
Citation Information
Cited By
Blood cell detection method based on deep learning
CN121214087A