Blood cell detection method based on improved YOLOv8

By introducing the CBAM attention mechanism and BiFPN network in YOLOv8, combined with the improved loss function and non-maximum inhibition algorithm, the deficiency of accuracy and robustness in blood cell detection is solved, and the detection effect and efficiency are significantly improved.

CN120198409APending Publication Date: 2025-06-24HARBIN MEDICAL UNIV DAQING BRANCH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510403161.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has strong subjectivity, time-consuming and difficult to meet the large-scale and efficient detection needs in blood cell detection, especially when dealing with complex backgrounds, cell overlap and morphological diversity, and lack of accuracy and robustness.

Method used

By introducing the CBAM attention mechanism into YOLOv8's backbone network and integrating the BiFPN network into the neck network, the blood cell detection model is iteratively trained and optimized.

Benefits of technology

It significantly improves the accuracy and robustness of blood cell detection, reduces false detection and missed detection, especially in the case of overlap and occlusion, and improves the detection performance of small target platelets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198409A_ABST
    Figure CN120198409A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, and discloses a blood cell detection method based on improved YOLOv8. According to the method, a convolutional attention module (CBAM) is introduced, so that the detection performance of the model on a shielded small target is improved; a new loss function S-MPDIOU is set, so that the accuracy and robustness of the loss function are further improved; the multi-scale feature fusion technology is adopted to better capture blood cell information of different sizes, so that the model can more accurately identify small target platelets; the recall rate is improved by improving a non-maximum suppression (NMS) algorithm, and particularly under the condition of dense blood cells or complex forms, the improved NMS algorithm can more effectively solve the overlapping problem and reduce false detection and missing detection. Through testing, the model can accurately identify and position different types of blood cells such as red blood cells, white blood cells and blood platelets in an image, and a new effective method is provided for blood cell detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection, and relates to but is not limited to a blood cell detection method based on improved YOLOv8. Background Art

[0002] Blood cell detection is a crucial diagnostic tool in clinical medicine, which can provide important information about the patient's health status for doctors. By analyzing the quantity and morphology of red blood cells, white blood cells, and platelets in a blood sample, doctors can identify various disease states, including infections, inflammations, anemia, and blood tumors, etc. Traditional blood cell detection usually relies on manual microscopic examination. Although this method is intuitive, it has problems such as strong subjectivity, time-consuming, and easy fatigue, and it is difficult to meet the needs of large-scale and high-efficiency clinical detection.

[0003] In recent years, the rise of deep learning technology has brought revolutionary changes to medical image processing. Especially in the field of object detection, models based on convolutional neural networks (CNNs) such as the YOLO (You Only Look Once) series stand out in various real-time detection tasks with their excellent detection speed and accuracy. As the latest member of the YOLO family, YOLOv8 not only reaches a new height in detection performance, but also further optimizes the model structure, reduces the number of parameters, and improves the inference speed, making it more suitable for deployment on resource-constrained edge devices.

[0004] However, directly applying YOLOv8 to blood cell detection still faces many challenges: First, blood cell images have complex and variable backgrounds, and the overlapping, occlusion, and morphological diversity between cells bring difficulties to accurate detection; Second, blood samples from different patients have differences in color, concentration, etc., requiring the detection algorithm to have high robustness and generalization ability; Finally, the medical field has extremely high requirements for the accuracy of detection results, and any misdetection or missed detection may have an adverse impact on clinical decisions. Summary of the Invention

[0005] In view of the defects in the prior art and the requirements of blood cell detection, the embodiments of the present invention provide a blood cell detection method based on improved YOLOv8, which can quickly and accurately identify and locate different types of blood cells such as red blood cells, white blood cells, and platelets in an image, providing strong support for clinical diagnosis and pathological research.

[0006] The technical solution of the embodiment of the present invention is specifically as follows: The embodiments of the present invention provide a blood cell detection method based on improved YOLOv8, including: Construct a blood cell dataset and divide it into a training set, a validation set, and a test set according to a ratio of 7:1:2; introduce the CBAM attention mechanism into the backbone network of YOLOv8 and incorporate the BiFPN network into the neck network to construct a blood cell detection model; based on the training set, iteratively train the blood cell detection model in combination with the improved S-MPDIoU loss function and use the validation set to adjust the hyperparameters during model training; wherein, the improved S-MPDIoU loss function introduces more blood cell geometric information and context information according to the characteristics of the blood cell detection task to measure the difference between the predicted bounding box and the ground truth bounding box; use the trained blood cell detection model to perform object detection on the test set, and optimize the model output using the improved non-maximum suppression algorithm to obtain the final blood cell detection result; evaluate the performance of the trained blood cell detection model based on the blood cell detection result.

[0007] In some embodiments, the construction of the blood cell dataset includes: adding complex background images to the public dataset BCCD and performing preprocessing to construct a blood cell dataset; wherein, the complex background images include blood cell images with at least the following conditions: cell overlap, irregular shape, and uneven staining.

[0008] In some embodiments, add a first CBAM module before the first convolutional layer and initialize it according to the number of input channels 64; add a second CBAM module after the downsampling at the 5th layer, with the number of channels being 256; add a third CBAM module after the P8 / 16 downsampling, with the number of channels being 512; add a fourth CBAM module after the P11 / 32 downsampling, with the number of channels being 1024.

[0009] In some embodiments, add the BiFPN module between the upsampling layer and the C2f layer in the neck network of YOLOv8, adjust the scale of the feature map through upsampling and downsampling operations, and use a weighted feature fusion mechanism to fuse feature maps of different scales.

[0010] In some embodiments, the improved S-MPDIoU loss function is calculated by the following formula: ; In the formula, is the intersection over union of the predicted bounding box and the ground truth bounding box, is the distance between the centers of the predicted bounding box and the ground truth bounding box; , are the aspect ratios of the predicted bounding box and the ground truth bounding box respectively; S0 is a scaling factor used to adjust the sensitivity of the final loss to different errors; , , , is a weighting factor used to adjust the contribution of each item to the final loss. is a regularization term used to prevent overfitting.

[0011] In some embodiments, the performance evaluation of the trained blood cell detection model based on the blood cell detection results includes: determining the following evaluation metrics based on the blood cell detection results: precision, recall, average precision, mean average precision, and recognition speed; wherein, the recognition speed is characterized by the recognition time of an average single image; and using the evaluation metrics to verify the blood cell recognition effect of the blood cell detection model.

[0012] In some embodiments, before proportionally dividing the blood cell data set, the method further includes: performing at least one of the following data augmentation operations on the original images in the blood cell data set: horizontal flipping, vertical flipping, adding Gaussian noise and salt-and-pepper noise, and enhancing or weakening the brightness.

[0013] In some embodiments, the method further includes: using the labelimg image annotation tool to annotate red blood cells, white blood cells, and platelets in the blood cell images in the training set and the validation set respectively and setting corresponding category labels.

[0014] In some embodiments, the model output includes the bounding box and category of the blood cells, and the optimization process of the model output includes: attenuating the confidence score of each corresponding bounding box according to the overlap degree between each bounding box and the target bounding box; wherein, the target bounding box is the bounding box with the highest score; removing the bounding boxes with confidence scores lower than a specific threshold according to the attenuated confidence scores, then updating the score list of the remaining bounding boxes to be processed and sorting them; re-determining the target bounding box based on the sorting result, and repeating the above steps until all bounding boxes are processed and then stopping.

[0015] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: In the embodiments of the present invention, by introducing an attention mechanism into the YOLOv8 algorithm, the detection performance of the model for occluded small targets is improved; a new loss function, S-MPDIoU, is set up to further improve the accuracy and robustness of the loss function; multi-scale feature fusion technology is adopted to better capture blood cell information of different sizes, enabling the model to more accurately identify small target platelets; by improving the non-maximum suppression (NMS) algorithm, the recall rate is increased. Especially in the case of dense or morphologically complex blood cells, the improved NMS algorithm can more effectively handle overlapping problems, reducing false detections and missed detections. The detection effect of the improved model on blood cells is significantly improved, especially reducing the occurrence of overlapping and mutually occluding red blood cells, as well as missed and misdetected small target platelets, providing a new effective method for blood cell detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where: Figure 1 is a schematic flowchart of the blood cell detection method based on the improved YOLOv8 provided by the embodiments of the present invention; Figure 2 is a schematic diagram of the data augmentation result of the blood cell image provided by the embodiments of the present invention; Figure 3 is a schematic structural diagram of the traditional YOLOv8 model provided by the embodiments of the present invention; Figure 4 is a schematic structural diagram of the improved YOLOv8 model provided by the embodiments of the present invention; Figure 5 is a schematic diagram of the comparison of the detection results before and after the algorithm improvement provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0018] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. It is understood, however, that "some embodiments" may be the same subset or a different subset of all possible embodiments, and may be combined with each other without conflict.

[0019] It should be noted that the terms "first / second / third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific order for the objects. It is understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.

[0020] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.

[0021] Figure 1 The flowchart of a blood cell detection method based on improved YOLOv8 provided for the embodiments of the present invention is as Figure 1 shown, and the method at least includes the following steps: Step S110, constructing a blood cell data set and dividing it into a training set, a validation set, and a test set according to a ratio of 7:1:2.

[0022] Here, the quality of the data set has a crucial impact on the performance of the detection model. When constructing a blood cell data set for input into the model, image acquisition is a crucial step. This process needs to ensure that the collected image data is diverse, representative, and of high quality, so as to be able to train an accurate and robust blood cell detection model. Preferably, the public data set BCCD covers blood cell images of different morphologies, sizes, and distributions, and can be used as the blood cell data set for object detection in the present invention.

[0023] In order to accurately evaluate the performance of the model and avoid overfitting, the training set, the validation set, and the test set are independent of each other and have no overlap. After data augmentation, a total of 5460 blood cell images are obtained. The data set is divided into a training set, a validation set, and a test set according to a ratio of 7:1:2, obtaining 3822 training set images, 546 validation set images, and 1092 test set images.

[0024] Step S120: Introduce the CBAM attention mechanism into the backbone network of YOLOv8 and incorporate the BiFPN network into the neck network to construct a blood cell detection model.

[0025] Here, although the traditional YOLOv8 model has high accuracy in object detection, due to the mutual occlusion between blood cells, it brings certain difficulties to recognition. Directly using the traditional YOLOv8 network model for the blood cell recognition task has low accuracy for small target platelets and mutually occluded blood cells.

[0026] In view of the above problems and combined with the characteristics of blood cells, the present invention introduces the CBAM attention mechanism into the backbone network to improve the detection performance of the model for occluded small targets. CBAM is mainly composed of two independent and cascaded sub-modules: the channel attention module (Channel Attention Module) and the spatial attention module (Spatial Attention Module). CBAM adaptively adjusts the importance of feature channels using the channel attention mechanism, and at the same time adaptively adjusts the importance of each spatial position in the feature map using the spatial attention mechanism. In the blood cell recognition task, this means that the model can better capture different features of blood cells, thereby improving the recognition accuracy. Especially for small target platelets, CBAM can help the model focus more on these subtle features, thus improving the detection effect. CBAM can help the model better separate the features of each blood cell, enabling the model to pay more attention to important channels and spatial positions, thereby enhancing the ability to capture details. This reduces the cases of false detection and missed detection for mutually occluded blood cells. Moreover, the design of the CBAM attention mechanism is relatively simple. When added to an existing convolutional neural network, the increase in computation and the number of parameters is small. For the blood cell recognition task, this means that the computational cost can be reduced while maintaining the model performance, and the recognition speed can be improved.

[0027] Due to the differences among individuals of different types of blood cells, it is often difficult to comprehensively capture the characteristics of blood cells with feature extraction at a single scale. The Bidirectional Feature Pyramid Network (BiFPN) is a feature fusion network for object detection. It significantly improves the accuracy and efficiency of object detection through multi-level feature fusion and dynamic feature weight allocation. This BiFPN bidirectional feature pyramid network structure allows features to be fused in both top-down and bottom-up directions. Specifically, the BiFPN module adjusts the scale of the feature maps through upsampling and downsampling operations and uses feature fusion techniques to fuse feature maps of different scales. This bidirectional fusion mechanism enables the network to more effectively utilize feature information at different levels, thereby improving detection performance. Compared with traditional feature pyramid networks (such as FPN), BiFPN introduces learnable weights to determine the importance of different input features. This means that the network can dynamically adjust their weights in the fusion process according to the different resolutions and contributions of the input features. This weighted fusion mechanism enables the network to pay more attention to features with more information, thus improving the effect of feature fusion. BiFPN has also made significant contributions to structural optimization. It optimizes cross-scale connections by removing nodes with only one input edge, adding additional edges between input and output nodes at the same level, and treating each bidirectional path as a feature network layer and repeating it multiple times. These optimization measures reduce redundant calculations in the network, improve the efficiency of feature fusion, and enable the network to maximize the effect of feature fusion while maintaining computational efficiency.

[0028] By introducing the BiFPN module into the neck network and using multi-scale feature fusion technology to better capture blood cell information of different sizes, the improved YOLOv8 model can more accurately identify blood cell objects. Especially for the detection of small target platelets, it significantly improves the accuracy and robustness of detection. At the same time, due to the efficient feature fusion mechanism and optimized network structure of the BiFPN module, it can also improve the detection speed without increasing the computational cost.

[0029] Step S130: Based on the training set, iteratively train the blood cell detection model in combination with the improved S-MPDIoU loss function and adjust the hyperparameters using the validation set during model training; wherein, the improved S-MPDIoU loss function introduces more blood cell geometric information and context information according to the characteristics of the blood cell detection task to measure the difference between the predicted bounding box and the true bounding box.

[0030] Here, the S-MPDIoU loss function is a specialized improvement of the MPDIoU loss function in combination with the characteristics of the blood cell detection task. The MPDIoU loss function measures the overlap of bounding boxes more accurately by introducing the center point and aspect ratio information of the bounding boxes. On this basis, S-MPDIoU, aiming at the characteristics of the blood cell detection task, can more effectively handle complex situations such as occlusion and overlap by introducing more geometric information and context information. In addition, the S-MPDIoU loss function can more accurately measure the difference between the predicted bounding box and the ground truth bounding box, thereby guiding the model to make more precise adjustments, optimizing the detection performance, and further improving the accuracy and robustness of the loss function.

[0031] The geometric information includes the shape, size, and position of blood cells. In blood cell detection, specifically, the shape of blood cells is described by high-precision bounding boxes, including the subtle changes in their edges. For example, when detecting platelets, more refined bounding boxes are introduced to capture their irregular shapes and subtle concavities and convexities on the edges. For different types of blood cells (such as red blood cells, white blood cells, platelets, etc.), their sizes and aspect ratios are different. By introducing this geometric information, the S-MPDIoU loss function can more accurately measure the overlap between the predicted bounding box and the ground truth bounding box, thereby improving the detection accuracy. When detecting multiple blood cells, the relative position relationship between them is also important geometric information. For example, when two blood cells are closely adjacent or overlapping, introducing the relative position information between them can help the model better distinguish them and avoid misdetection. The context information refers to the additional information related to the blood cell detection task, which can help the model better understand the content and structure in the image. Specifically, it includes background information. For example, by introducing information such as color, texture, and brightness in the background image, it can help the model better distinguish the foreground (blood cells) and the background, thereby improving the detection accuracy; it includes the distribution of surrounding blood cells. When detecting a single blood cell, the distribution of surrounding blood cells can also provide useful context information. For example, when blood cells are densely distributed in a certain area, it can be inferred that this area may be an important detection area, thereby strengthening the detection efforts in this area; it includes the type and quantity of blood cells: In the detection task, knowing the type and quantity of blood cells in the image is also very important context information. For example, when detecting platelets, knowing that the number of platelets in the image is small, then the model can pay more attention to those areas similar to the characteristics of platelets, thereby improving the detection sensitivity and reducing misdetection.

[0032] Model training is carried out on a deep learning cloud GPU server with a 12-core Xeon Platinum 8260M CPU, an NVIDIA RTX A6000 48G GPU, 86G of memory, and a 350G hard drive. The programming language is Python 3.8, the deep learning framework is PyTorch 1.10.0, and the CUDA version is 11.7.

[0033] In step S140, use the trained blood cell detection model to perform object detection on the test set, and optimize the model output by referring to the improved non-maximum suppression algorithm to obtain the final blood cell detection result.

[0034] Here, in the blood cell detection task, the traditional YOLOv8 algorithm often predicts multiple overlapping or similar bounding boxes, and these bounding boxes may all point to the same blood cell. The traditional NMS algorithm calculates the score of each bounding box, retains the bounding box with the highest score, and removes other bounding boxes with an overlap degree exceeding a specific threshold. However, this simple removal strategy may cause some correct but highly overlapping bounding boxes to be wrongly suppressed, thus affecting the detection accuracy. The improved non-maximum suppression (Non-Maximum Suppression, NMS) algorithm provided by the present invention can improve the recall rate. By introducing the Soft-NMS method, the score of the bounding box can be dynamically adjusted according to the overlap degree instead of directly removing it.

[0035] Specifically, the Soft-NMS method does not directly set the bounding box with a high overlap degree with the highest-score bounding box to 0 or delete it, but attenuates the confidence of the bounding box according to the overlap degree. The higher the overlap degree, the more the confidence is attenuated; the lower the overlap degree, the less or no attenuation of the confidence. This method can retain as many correct bounding boxes as possible while removing redundant bounding boxes. This can retain more bounding boxes that may contain correct blood cells, thus improving the detection accuracy. Especially in the case of dense or morphologically complex blood cells, the improved NMS algorithm can more effectively handle the overlap problem and reduce false detections and missed detections.

[0036] In step S150, perform performance evaluation on the trained blood cell detection model based on the blood cell detection result.

[0037] Here, evaluation metrics such as precision, recall, average precision (Average Precision, AP), mean average precision (mean Average Precision, mAP), and recognition speed are used to verify the blood cell recognition effect.

[0038] In some embodiments, constructing the blood cell dataset includes: adding complex background images to the publicly available dataset BCCD and performing preprocessing to construct the blood cell dataset; wherein, the complex background images include blood cell images with at least the following conditions: cell overlap, irregular shape, and uneven staining.

[0039] Here, the image preprocessing at least includes the following operations: reading the image, using the OpenCV or PIL library to read the blood cell image; resizing the image, using the image scaling method to resize the image to 640x640 pixels; normalizing the image, normalizing the image pixel values from [0, 255] to the range of [0, 1]. (1) The role of preprocessing is to make the input image meet the requirements of the detection model, improve the training effect and overall performance of the model, and accelerate convergence.

[0040] In some embodiments, a first CBAM module is added before the first convolutional layer and initialized according to the number of input channels 64; a second CBAM module is added after the 5th layer downsampling, with the number of channels being 256; a third CBAM module is added after the P8 / 16 downsampling, with the number of channels being 512; a fourth CBAM module is added after the P11 / 32 downsampling, with the number of channels being 1024.

[0041] Here, the CBAM attention mechanism can help the model better separate the features of each blood cell, enabling the model to pay more attention to important channels and spatial positions, thereby enhancing the ability to capture details. This enables the model to reduce the cases of false detection and missed detection for mutually occluded blood cells. And the design of the CBAM attention mechanism is relatively simple. When added to an existing convolutional neural network, the increase in computation and the number of parameters is small. For the blood cell recognition task, this means that the computational cost can be reduced while maintaining the model performance, and the recognition speed can be improved. By introducing the CBAM attention mechanism, the detection performance of the model for occluded small targets is improved.

[0042] The specific implementation is as follows: Create a cbam.py file in the ultralytics / nn / modules folder, and define three classes, namely ChannelAttentionModule, SpatialAttentionModule, and CBAM in it; Modify the ultralytics / nn / modules / init.py file to import the CBAM module into the initialization module; In the ultralytics / nn / tasks.py file, configure the use of the CBAM module and add CBAM to the network structure of YOLOv8; Modify the yolov8.yaml file to add the CBAM module before the first Conv layer and initialize it according to the input channel number 64; Add the CBAM module after the 5th downsampling, with the number of channels being 256; Add the CBAM module after the P8 / 16 downsampling, with the number of channels being 512; Add the CBAM module after the P11 / 32 downsampling, with the number of channels being 1024.

[0043] In some embodiments, the BiFPN module is added between the upsampling layer and the C2f layer in the neck network of YOLOv8, and the scale of the feature map is adjusted through upsampling and downsampling operations, and a weighted feature fusion mechanism is used to fuse feature maps of different scales.

[0044] Here, through the BiFPN module, the model can more accurately identify blood cell objects. Especially for the detection of small target platelets, it significantly improves the detection accuracy and robustness. At the same time, due to the BiFPN module having an efficient feature fusion mechanism and an optimized network structure, it can also improve the detection speed without increasing the computational cost.

[0045] The specific implementation is as follows: Create a BiFPN.py file in the ultralytics / nn / modules folder to store the implementation code of BiFPN; In the ultralytics / nn / tasks.py file, import the created BiFPN module and register this new module in the parse_model method; Modify the yolov8.yaml file to add a reference to the BiFPN layer and replace all Concat modules with BiFPN modules.

[0046] Schematic diagrams of the traditional YOLOv8 model structure and the improved YOLOv8 model structure are respectively as Figure 3 and Figure 4 shown. It can be seen that in the present invention, the CBAM module is added at four different places in the backbone network of the traditional YOLOv8 model, and at the same time, the BiFPN module is added between the upsampling layer and the C2f layer in the neck network.

[0047] In some embodiments, the improved S-MPDIoU loss function is calculated by the following formula: ; In the formula, is the intersection over union of the predicted bounding box and the ground truth bounding box, is the distance between the centers of the predicted bounding box and the ground truth bounding box; , are the aspect ratios of the predicted bounding box and the ground truth bounding box respectively; S0 is a scaling factor used to adjust the sensitivity of the final loss to different errors; , , , are weighting factors used to adjust the contributions of each term to the final loss, is a regularization term used to prevent overfitting.

[0048] In some embodiments, the performance evaluation of the trained blood cell detection model based on the blood cell detection results includes: determining the following evaluation metrics based on the blood cell detection results: precision, recall, average precision, mean average precision, and recognition speed; wherein, the recognition speed is characterized by the recognition time of an average single image; the blood cell recognition effect of the blood cell detection model is verified using the evaluation metrics.

[0049] Here, the calculation formulas for each evaluation metric are as follows: In the formula, P represents precision, R represents recall, TP is the number of correctly recognized blood cells, FP is the number of misrecognized blood cells, FN is the number of missed recognized blood cells, AP is the average precision of a certain category detection, mAP is the average value of multiple AP categories, n represents the total number of categories, and the total number of categories in the blood cell detection task is 3.

[0050] In some embodiments, before dividing the blood cell data set proportionally, the method further includes: performing at least one of the following data augmentation operations on the original images in the blood cell data set: horizontal flipping, vertical flipping, adding Gaussian noise and salt-and-pepper noise, enhancing and weakening brightness.

[0051] Here, the BCCD data set includes 364 blood cell images. After data augmentation, a total of 5460 blood cell images are obtained. The data set is divided into a training set, a validation set, and a test set according to the ratio of 7:1:2, obtaining 3822 training set images, 546 validation set images, and 1092 test set images. The results after data augmentation are as Figure 2As shown. Through these data augmentation operations, the dataset can be enriched and expanded, the generalization ability of the model can be improved, and overfitting can be prevented.

[0052] In some embodiments, the method further includes: using the labelimg image annotation tool to annotate red blood cells, white blood cells, and platelets in the blood cell images in the training set and the validation set respectively and setting corresponding category labels.

[0053] Here, when constructing the blood cell dataset, more attention is paid to the annotation quality of the data. Preferably, the label names of red blood cells, white blood cells, and platelets are "RBC", "WBC", and "Platelets" respectively.

[0054] In some embodiments, the model output includes the bounding box and category of blood cells. The transposing and post-processing of the model output includes: attenuating the confidence score of the corresponding bounding box according to the overlap degree between each bounding box and the target bounding box; wherein, the target bounding box is the bounding box with the highest score; after removing the bounding boxes with confidence scores lower than a specific threshold according to the attenuated confidence scores, updating the score list of the remaining bounding boxes to be processed and sorting; re-determining the target bounding box based on the sorting result, and repeating the above steps until all bounding boxes are processed and then stopping.

[0055] Here, first, calculate the IoU value between the currently highest-scoring bounding box M and all bounding boxes N to be processed. Then, according to the IoU value, apply an attenuation function to calculate the new confidence score. In the improved blood cell recognition method of YOLOv8, the Gaussian attenuation function is applied as the attenuation function of the Soft-NMS method. The Gaussian attenuation function performs well in dealing with overlapping targets because it can smoothly reduce the scores of overlapping boxes instead of directly setting them to 0 as in traditional NMS. Then, remove the bounding boxes with attenuated confidence scores lower than a specific threshold, and the specific threshold is set according to experience, for example, it can be 0.1. This method can remove redundant bounding boxes while retaining as many correct bounding boxes as possible.

[0056] This can retain more bounding boxes that may contain correct blood cells, thereby improving the detection accuracy. Especially in the case of dense or morphologically complex blood cells, the improved NMS algorithm can more effectively handle overlapping problems and reduce false detections and missed detections.

[0057] The following describes the above-mentioned blood cell detection method based on the improved YOLOv8 in combination with a specific embodiment. However, it should be noted that this specific embodiment is only for better explaining the present invention and does not constitute an improper limitation to the present invention.

[0058] The blood cell recognition and detection process provided by this specific embodiment is as follows: (1)Input image: The blood cell image to be detected in the test set is input into the model.

[0059] (2)Preprocessing: The blood cell image is resized to 640*640 pixels, and then the pixel values of the image are normalized so that the pixel value range becomes [0,1]. The role of preprocessing is to make the input image meet the requirements of the detection model, improve the training effect and overall performance of the model, and accelerate convergence.

[0060] (3)Feature extraction: The input blood cell image first enters the backbone network for feature extraction. The backbone network of the improved YOLOv8 model consists of multiple convolutional layers, pooling layers, and CBAM attention mechanism modules, which are used to extract low-level and high-level features in the blood cell image. The CBAM attention mechanism weights the blood cell feature map through two dimensions: channel attention and spatial attention, thereby enhancing the detection ability of the model. Channel Attention Module (CAM) processing: The blood cell image first enters the channel attention module, and global average pooling and global maximum pooling operations are performed on the input image to obtain global context information. The two vectors obtained by global average pooling and global maximum pooling are respectively sent into a shared multi-layer perceptron network for processing. The multi-layer perceptron network consists of two fully connected layers, and the ReLU activation function is used for non-linear transformation in the middle. Through the multi-layer perceptron network, the importance weights of each channel can be learned. The two weight vectors output by the multi-layer perceptron network are added element by element and non-linearly transformed through the Sigmoid activation function to obtain the final channel attention map (Channel Attention Map). The channel attention map is multiplied element by element with the input feature map to obtain the weighted feature map F1.

[0061] Spatial Attention Module (SAM) processing: The weighted feature map F1 output by the channel attention module is subjected to maximum pooling and average pooling operations in the channel dimension to obtain spatial context information. The two two-dimensional feature maps obtained by maximum pooling and average pooling are concatenated to generate a feature descriptor containing richer spatial information. A 7x7 convolutional kernel is used to perform convolution processing on the concatenated feature descriptor to generate a spatial attention map (Spatial Attention Map). The spatial attention map is multiplied element by element with the weighted feature map output by the channel attention module to obtain the feature map F2.

[0062] The feature map F2 passes through the first convolutional layer (Conv), and the output feature map F3 has a size of 320*320*64, with the length and width being 1 / 2 of the original image. Then, the feature map F3 is input into the second convolutional layer, and the size of the convolutional feature map F4 is 160*160*128, with the length and width being 1 / 4 of the original image. Next, the feature map F4 is input into the next C2f layer. The C2f module first performs feature transformation on the input feature map F4 through two convolutional layers (cv1 and cv2). After passing through the cv1 convolutional layer, the feature map F4 is split into two parts. One part of the feature map is directly passed to the subsequent BiFPN block without additional processing. The other part of the feature map is passed to multiple Bottleneck blocks for further processing. Each Bottleneck block contains two convolutional layers, which transform the input feature map and extract higher-level feature representations. The feature map processed by the Bottleneck block is concatenated with the directly passed part of the feature map in the BiFPN block to form the fused feature map F5. This feature fusion operation enables the model to comprehensively utilize information from different branches and enhances the feature expression ability. Continuing to be passed down, it enters the third convolutional layer for convolution operation, and the output feature map F6 has a size of 80*80*256. The length and width of the feature map F6 have become 1 / 8 of the input image. Then, after passing through a CBAM module and a C2f module, it is input into the fourth convolutional layer for another convolution operation, and the output feature map F7 has a size of 40*40*512. The length and width of the feature map F7 have become 1 / 16 of the input image. Then, after passing through a CBAM module and a C2f layer, it is input into the fifth convolutional layer for convolution operation, and the output feature map F8 has a size of 20*20*1024. The length and width of the feature map F8 have become 1 / 32 of the input image. Then, after performing a CBAM and C2f operation once again, the feature map F9 is obtained. Finally, the feature map F9 is input into the SPPF layer.

[0063] (4) Feature Fusion: SPPF uses multiple small pooling layers to replace the pooling operation with a single large kernel. After the fast pooling operation, the SPPF layer concatenates the outputs of different pooling layers in the channel dimension to form a feature vector of a fixed length. This feature vector contains both the spatial information of the original feature map and the feature information of different scales, thus improving the richness and representativeness of the features. This feature vector is then input into the upsampling layer (upsample) of the neck network. The upsample layer is used to increase the size of the feature map for more feature extraction in subsequent layers. Through the upsampling operation, the fineness and detail expression ability of the image can be improved, thereby enhancing the model performance. The upsampling layer of the improved YOLOv8 model is implemented by bilinear interpolation. Specifically, for each pixel to be enlarged, the weighted average of the values of its surrounding four pixels is calculated to obtain the enlarged pixel value, which can maintain the smoothness and continuity of the blood cell image to a certain extent. After this layer, the length and width of the feature map F10 become twice the original, and the number of channels remains unchanged, so the final size of the feature map F10 is 40*40*1024. It continues to be passed down to the BiFPN layer. The BiFPN module adjusts the scale of the feature map through upsampling and downsampling operations and uses feature fusion technology to fuse feature maps of different scales. This two-way fusion mechanism enables the network to more effectively utilize feature information at different levels, thereby improving the detection performance. Dynamically adjust their weights during the fusion process according to the different resolutions and contributions of the input features. This weighted fusion mechanism enables the network to pay more attention to features with more information, thereby improving the effect of feature fusion. Especially for the detection of small target platelets, it significantly improves the accuracy and robustness of the detection. At the same time, due to the efficient feature fusion mechanism and optimized network structure of the BiFPN module, it can also improve the detection speed without increasing the computational cost, and then perform a C2f operation to obtain the feature map F11. The feature map F11 goes through another upsampling layer, then through the BiFPN module and the C2f operation, and the obtained feature map F12 has a size of 80*80*256. The length and width of the feature map F12 have become 1 / 8 of the input image. Then, a convolution operation is performed again, and the output feature map F13 has a size of 40*40*256. After passing through the BiFPN module and the C2f operation again, the final feature map F14 has a size of 20*20*1024. The length and width of the feature map F14 have become 1 / 32 of the input image.

[0064] (4) Object Detection: The outputs of the previous 3 layers are respectively fed into the Detect layer in the Head part for object detection to obtain the positions and categories of red blood cells, white blood cells, and platelets in the blood cell image.

[0065] (5)Result Transposition and Post-Processing: Perform a transposition operation on the model output to make its dimensions meet the requirements of subsequent processing. Perform post-processing operations such as NMS on the transposed output to remove redundant bounding boxes and determine the final detection results. In an example of a blood cell detection task, the YOLOv8 model predicted multiple overlapping bounding boxes, all of which pointed to the same red blood cell. Under the traditional NMS algorithm, only the bounding box with the highest score would be retained, and other overlapping bounding boxes would be removed. Then, other bounding boxes that, although overlapping, might capture different features of the blood cells would be ignored, resulting in missed detections. Especially in the case of dense or overlapping blood cells, this problem is particularly serious. After introducing Soft-NMS, first, the model still predicts multiple overlapping bounding boxes. Then, the Soft-NMS method attenuates the confidence of these bounding boxes according to the degree of overlap (IOU value) between them and the bounding box with the highest score. According to the attenuated scores, update the score list of the boxes to be processed and sort them according to the new scores. Repeat the above steps until the score is lower than the 0.1 threshold or all boxes are processed and stop. Thus, more bounding boxes that may contain correct blood cells are retained, improving the detection accuracy. Especially in the case of dense or complex-shaped blood cells, this improvement can more effectively handle the overlapping problem, reducing false positives and missed detections.

[0066] To verify the effectiveness of the improved YOLOv8 method, compare its blood cell recognition results with the method before improvement in the same hardware environment and on the same test set. The specific quantitative recognition results are shown in Table 1.

[0067] Table 1 Comparison of Blood Cell Recognition Results before and after Improvement According to Table 1, the improved YOLOv8 algorithm achieved a recognition accuracy of 95.32% and a recall rate of 94.87% in the blood cell recognition task. Compared with the algorithm before improvement, the recognition accuracy increased by 5.24 percentage points, and the recall rate increased by 14.15 percentage points. The average recognition time for a single image was 17.2 milliseconds (ms), which was 2.9 ms faster than before improvement. This is due to the introduction of the CBAM attention mechanism, which enables the model to pay more attention to important channels and spatial positions, thereby enhancing the ability to capture details. For mutually occluded blood cells, the cases of false detection and missed detection can be reduced. Moreover, the design of the CBAM attention mechanism is relatively simple, and when added to the existing convolutional neural network, the increase in computation and the number of parameters is small. Additionally, by introducing the BiFPN module, the model can more accurately identify blood cell objects, especially for the detection of small target platelets, significantly improving the detection accuracy and robustness. At the same time, due to the efficient feature fusion mechanism and optimized network structure of the BiFPN module, it can also improve the detection speed without significantly increasing the computational cost. By improving the non-maximum suppression algorithm, the recall rate is increased, and false detection and missed detection are reduced. Therefore, the improved YOLOv8 method can significantly improve the recognition accuracy of blood cells while enhancing the recognition efficiency.

[0068] To further verify the effectiveness of this method, it was compared with other blood cell recognition methods under the same test set and experimental platform, and the recognition results are shown in Table 2. It can be seen from Table 2 that the blood cell recognition method based on the improved YOLOv8 is superior to other algorithms in terms of recognition performance.

[0069] Table 2 Comparison of recognition results of different methods on the same test set To better illustrate the recognition effect of the model on small target platelets and stacked red blood cells, the test images were respectively detected and compared using the original YOLOv8 and the improved YOLOv8 models. The recognition and detection results are as Figure 5As shown. From Sample 1, it can be seen that for blood cell images with less mutual occlusion, both the traditional YOLOv8 detection method and the improved method can accurately identify red blood cells and white blood cells. For the blood cell image of Sample 2 containing platelets, the traditional YOLOv8 fails to identify platelets due to its poor recognition effect on small targets, while the improved method can accurately identify platelets due to the introduction of the CBAM attention mechanism and the multi-scale feature fusion module. For Sample 3 where there are a large number of overlaps and mutual occlusions between cells, the traditional YOLOv8 has missed detections for a large number of overlapping red blood cells, while the improved model can still accurately identify the overlapping red blood cells. By comparing the detection results of the three groups of samples, it can be found that the improved YOLOv8 model is more accurate in identifying both overlapping and mutually occluding red blood cells and small target platelets, improving the detection accuracy and robustness of blood cells and having better practical value.

[0070] The present invention provides a blood cell detection method based on improved YOLOv8. By optimizing the model structure, improving the feature extraction strategy, and introducing the attention mechanism, the detection effect of the improved model on blood cells is significantly improved, which is expected to promote the automation process of blood cell detection, improve the detection efficiency, reduce the work burden of medical staff, and provide more reliable data support for clinical diagnosis and treatment.

[0071] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0072] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0073] In several embodiments provided by the present invention, it should be understood that the disclosed method can be implemented in other ways. The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0074] As mentioned above, it is only the implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claimed rights.

Claims

1. A blood cell detection method based on improved YOLOv8, characterized in that: include: Construct a blood cell dataset and divide it into training set, validation set, and test set in a ratio of 7:1:2; The CBAM attention mechanism is introduced into the backbone network of YOLOv8, and the BiFPN network is integrated into the neck network to construct a blood cell detection model; Based on the training set, the blood cell detection model is iteratively trained in combination with an improved S-MPDIoU loss function, and the hyperparameters are adjusted using the validation set during model training; wherein the improved S-MPDIoU loss function introduces more blood cell geometry information and context information based on the characteristics of the blood cell detection task to measure the difference between the predicted bounding box and the true bounding box; Utilizing the trained blood cell detection model to perform target detection on the test set, and optimizing the model output by using an improved non-maximum suppression algorithm to obtain a final blood cell detection result; The performance of the trained blood cell detection model is evaluated based on the blood cell detection results.

2. The method according to claim 1, characterized in that The constructing of the blood cell dataset comprises: A complex background image is added to the public data set BCCD, and preprocessing is performed to construct a blood cell data set; wherein the complex background image includes blood cell images in at least the following situations: cell overlap, irregular shape, and uneven staining.

3. The method according to claim 1, characterized in that The first CBAM module is added before the first convolutional layer and initialized according to the input channel number 64; the second CBAM module is added after downsampling at the 5th layer, with the channel number 256; the third CBAM module is added after downsampling at P8 / 16, with the channel number 512; the fourth CBAM module is added after downsampling at P11 / 32, with the channel number 1024.

4. The method according to claim 1, characterized in that: The BiFPN module is added between the upsampling layer and the C2f layer in the neck network of YOLOv8, the scale of the feature map is adjusted through upsampling and downsampling operations, and the weighted feature fusion mechanism is used to fuse feature maps of different scales.

5. The method according to any one of claims 1 to 4, characterized in that: The improved S-MPDIoU loss function is calculated by the following formula: ; In the formula, is the intersection-over-union ratio of the predicted bounding box and the true bounding box, is the distance between the center points of the predicted bounding box and the true bounding box; , are the aspect ratios of the predicted bounding box and the true bounding box, respectively; S0 is a scaling factor used to adjust the sensitivity of the final loss to different errors; , , , is a weighting factor used to adjust the contribution of each item in the final loss. is a regularization term used to prevent overfitting.

6. The method according to any one of claims 1 to 4, characterized in that: The performing performance evaluation on the trained blood cell detection model based on the blood cell detection result includes: Determine the following evaluation indicators based on the blood cell detection result: precision, recall, average precision, average precision mean and recognition speed; wherein the recognition speed is characterized by the average recognition time of a single image; The evaluation index is used to verify the blood cell recognition effect of the blood cell detection model.

7. The method according to any one of claims 1 to 4, characterized in that: Before dividing the blood cell dataset proportionally, the method further comprises: At least one of the following data enhancement operations is performed on the original image in the blood cell data set: horizontal flipping, vertical flipping, adding Gaussian noise and salt and pepper noise, and enhancing or reducing brightness.

8. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: The labelimg image annotation tool is used to respectively annotate the red blood cells, white blood cells, and platelets in the blood cell images in the training set and the validation set and set corresponding category labels.

9. The method according to any one of claims 1 to 4, characterized in that: The model output includes a bounding box and a category of blood cells, and the optimization processing of the model output includes: Attenuating the confidence score of the corresponding bounding box according to the degree of overlap between each bounding box and the target bounding box; wherein the target bounding box is the bounding box with the highest score; According to the attenuated confidence score, the score list of the remaining bounding boxes to be processed is updated and sorted after removing the bounding boxes with confidence scores lower than a certain threshold; The target bounding box is re-determined based on the sorting result, and the above steps are repeated until all bounding boxes are processed.

Citation Information

Cited By

  • Detection method, system, equipment and medium for blood sampling tube blood layering identification

    CN120411079A

  • A detection method, system, device and medium for blood stratification identification in blood collection tubes

    CN120411079B

  • Automatic lung slice cell nucleus segmentation method combining microscopic hyperspectrum and artificial intelligence

    CN121213586A

  • Bone marrow blood cell target identification and counting method based on improved single stage

    CN121599982A