Lightweight blood cell detection method based on improved YOLO network
By improving the YOLO network, using the ShuffleNet module for multi-scale feature extraction and ASFF module for adaptive feature fusion, the problems of real-time performance and accuracy of existing blood cell detection models on low-end devices are solved, and fast and accurate blood cell detection is achieved.
Patent Information
- Application Number
- CN202510153778.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-24
AI Technical Summary
Due to the large amount of parameters and calculations, existing blood cell detection models are difficult to achieve real-time performance on low-end devices, and at the same time, the accuracy decreases when detecting platelets and dense red blood cells.
The improved YOLO network is adopted to perform multi-scale feature extraction through the backbone network based on the ShuffleNet module, and the ASFF module is used to perform adaptive spatial feature fusion, reducing model parameters and calculation complexity, while enhancing feature extraction and fusion capabilities.
It achieves the ability to maintain rapid detection while deploying on edge devices, improves the accuracy and real-time performance of blood cell detection, especially when detecting small and dense targets.
Smart Images

Figure CN120198356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of blood cell detection, and particularly to a lightweight blood cell detection method based on an improved YOLO network. Background Art
[0002] Detecting blood cells under a microscope is a common disease diagnosis method, which usually includes identifying red blood cells, white blood cells and platelets. Manually counting various blood cells is not only challenging and time-consuming, but also error-prone. Therefore, it is very necessary to explore the use of computer technology for efficient and accurate blood cell detection.
[0003] Object detection technology based on convolutional neural networks is currently widely used in various computer vision tasks, including medical image analysis, autonomous driving and face recognition, to achieve object recognition, localization and tracking. High-precision object detection algorithms have been proven to assist medical diagnosis through blood cell detection, achieving satisfactory results. However, the common problem of these models is that the number of parameters is too large, making it difficult to deploy the models on edge devices. Moreover, the large computational amount of these models also slows down the detection speed of the detector, making it difficult to complete some tasks with high real-time requirements. Deploying such a blood cell object detector in areas with limited medical and computing resources will cause problems. In addition, the rapid detection ability of medical devices has received increasing attention, and detection real-time is a key issue, which is manifested in many aspects of the detection process. In the application of blood cell detection, the detection speed is equally important as the detection accuracy. Therefore, it is crucial to develop a lightweight model with a small number of parameters and a small computational amount to achieve deployment on edge devices while maintaining a fast detection ability.
[0004] The single-stage object detection algorithm YOLO has shown excellent performance in the fields of image classification and object detection. Its key features are fast processing speed and high detection accuracy, making it particularly suitable for real-time object detection scenarios. This algorithm has been successfully applied to multiple fields such as face recognition, autonomous driving, and medical image analysis. Among the many variants of YOLO, YOLO v5 is the most influential and has achieved good performance in many tasks. However, due to the large number of parameters and computations, YOLO v5 cannot achieve good real-time performance on low-end devices. YOLO-Fastest v2 proposed by Dog-qiuqiu is a lightweight neural network based on the YOLO architecture. This model replaces the backbone network of YOLOv5 with ShuffleNet V2 and reduces the feature pyramid structure. Therefore, it can meet the requirements of real-time detection even on devices with limited computing resources. This model is designed to execute efficiently on resource-constrained hardware. However, when this algorithm is used for blood cell detection, the overlapping of images caused by high-density red blood cells leads to a decrease in accuracy. In addition, the platelet targets are very small, making detection challenging in complex tasks.
[0005] How to improve the accuracy and real-time performance of the model in the blood cell detection task is a key issue. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a lightweight blood cell detection method based on an improved YOLO network, which can improve the efficiency and accuracy of blood cell detection while reducing the model parameters.
[0007] The technical solution adopted by the present invention to solve its technical problems is: to provide a lightweight blood cell detection method based on an improved YOLO network, including the following steps:
[0008] Obtain the blood cell image to be detected;
[0009] Construct a lightweight blood cell detection model based on the improved YOLO network, and detect and identify the blood cells in the blood cell image to obtain a detection result; the improved YOLO network includes:
[0010] Use a backbone network based on the ShuffleNet module to perform multi-scale feature extraction on the blood cell image to obtain multi-scale feature maps;
[0011] Construct a neck network based on the ASFF module for multi-scale feature fusion and generate a set number of detection heads; the ASFF module adaptively learns the spatial fusion weights of the feature maps at each scale and performs dynamic multi-scale feature fusion according to the fusion weights.
[0012] Further, adaptively learning the fusion weights of the feature maps at each scale and performing dynamic multi-scale feature fusion according to the fusion weights includes:
[0013] Using at least one of upsampling operation, 0.5x downsampling operation or 0.25x downsampling operation to align the scale of each feature map with the scale of the output feature map, obtaining the corresponding feature maps to be fused;
[0014] Based on the adaptive fusion mechanism, using the backpropagation algorithm to obtain the spatial fusion weights of the feature maps to be fused, and then calculating that the output feature map is the weighted sum of all the feature maps to be fused.
[0015] Further, the adaptive fusion mechanism is expressed as:
[0016]
[0017] and
[0018] wherein, represents the feature vector at the position (i, j) adjusted from scale n to scale l, represents the feature vector at the position (i, j) on the output feature map, represents the spatial fusion weight for fusing the n feature maps to be fused at different scales to scale l, and λ is a network parameter.
[0019] Further, the upsampling operation includes:
[0020] Using a 1×1 convolutional layer to align the channel size of the feature map with the channel size of the output feature map;
[0021] Performing interpolation processing on the feature map after the alignment operation.
[0022] Further, the 0.5x downsampling operation includes:
[0023] Using a 3×3 convolutional layer with a stride of 2 to simultaneously adjust the resolution and channel number of the feature map.
[0024] Further, the 0.25x downsampling operation includes:
[0025] Performing a max pooling operation with a stride of 2 on the feature map;
[0026] Using a 3×3 convolutional layer with a stride of 2 to simultaneously adjust the resolution and channel number of the feature map after the max pooling operation.
[0027] Further, the multi-scale feature extraction of the blood cell image using the backbone network based on the ShuffleNet module to obtain a multi-scale feature map includes:
[0028] Pass the blood cell image through a max pooling layer, a first improved ShuffleNet module, and 3 first ShuffleNet modules in sequence, and output a first-scale feature map;
[0029] Pass the first-scale feature map through a second improved ShuffleNet module and 7 second ShuffleNet modules in sequence, and output a second-scale feature map;
[0030] Pass the second-scale feature map through a third improved ShuffleNet module, 3 third ShuffleNet modules, and an SPPF module in sequence, and output a third-scale feature map.
[0031] Further, the improved ShuffleNet module includes two downsampling branches. The output vectors of the two downsampling branches are concatenated and then channel rearranged. One of the downsampling branches introduces an attention mechanism to enhance the feature expression in the channel dimension and the spatial dimension.
[0032] Further, one of the downsampling branches of the improved ShuffleNet module includes a 3×3 convolutional layer, a 1×1 convolutional layer, and a CBAM module connected in sequence, and the other downsampling branch includes a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence.
[0033] Further, the multi-scale feature fusion includes:
[0034] Perform convolution operations on the first-scale feature map, the second-scale feature map, and the third-scale feature map respectively to obtain a fourth feature map, a fifth feature map, and a sixth feature map;
[0035] Process the fourth feature map and the fifth feature map through a first ASFF module and a first DWConv module in sequence, and output a seventh feature map;
[0036] Process the fourth feature map and the fifth feature map through a second ASFF module and a second DWConv module in sequence, and output an eighth feature map;
[0037] Process the sixth feature map, the seventh feature map, and the eighth feature map through a third ASFF module and a third DWConv module in sequence and then perform a convolution operation to obtain a first detection feature map;
[0038] The sixth feature map, the seventh feature map, and the eighth feature map are sequentially processed by a fourth ASFF module and a fourth DWConv module and then subjected to a convolution operation to obtain a second detection feature map.
[0039] Beneficial effects
[0040] Due to the above technical solutions, compared with the prior art, the present invention has the following advantages and positive effects: The present invention constructs a backbone network by using an improved ShuffleNet V2, reducing model parameters and computational complexity, making it more suitable for deployment on embedded devices; at the same time, the improved ShuffleNet V2 uses the CBAM (Convolutional Block Attention Module) attention mechanism to suppress noise and focus on significant features, enhancing the network's feature extraction ability. The SPPF module at the end of the backbone network achieves multi-scale feature extraction, significantly improving network performance and efficiency; the present invention designs an ASFF (Adaptive Spatial Feature Fusion) module, introducing an adaptive spatial feature fusion mechanism, enabling the network to directly learn how to perform spatial filtering on features from other levels, enhancing scale-invariant feature representation at the lowest computational cost, enhancing feature fusion with negligible inference overhead, solving the main limitations of single-stage detectors of feature pyramids caused by inconsistencies between different feature scales, and further solving the problem of poor feature extraction performance of traditional pyramid structures for large size differences between red blood cells, white blood cells, and platelets in images, improving the multi-scale object recognition ability of the YOLO network. Brief description of the drawings
[0041] Figure 1 is a schematic diagram of the structure of the lightweight blood cell detection model according to an embodiment of the present invention;
[0042] Figure 2 is a performance comparison diagram between the lightweight blood cell detection model according to an embodiment of the present invention and the existing model;
[0043] Figure 3 is a P-R curve diagram of the lightweight blood cell detection model according to an embodiment of the present invention;
[0044] Figure 4 is a confusion matrix diagram of the lightweight blood cell detection model according to an embodiment of the present invention;
[0045] Figure 5 is a detection result comparison diagram between the lightweight blood cell detection model according to an embodiment of the present invention and the YOLO-Fastest v2 model. Detailed implementation manners
[0046] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0047] An embodiment of the present invention relates to a lightweight blood cell detection method based on an improved YOLO network, including the following steps:
[0048] Obtain a blood cell image to be detected;
[0049] Construct a lightweight blood cell detection model based on the improved YOLO network, and detect and identify blood cells in the blood cell image to obtain a detection result; the improved YOLO network includes:
[0050] Use a backbone network based on the ShuffleNet module to perform multi-scale feature extraction on the blood cell image to obtain a multi-scale feature map;
[0051] Construct a neck network based on the ASFF module for multi-scale feature fusion and generate a set number of detection heads; the ASFF module adaptively learns the spatial fusion weights of feature maps at each scale and performs dynamic multi-scale feature fusion according to the fusion weights.
[0052] The network architecture of the lightweight blood cell detection model based on YOLO in this embodiment is as Figure 1 shown, and it is composed of four modules: input, backbone, neck, and prediction. This embodiment adopts a three-layer scale architecture, which can also be adjusted according to actual needs.
[0053] The backbone network uses an improved ShuffleNet to reduce model parameters and computational complexity, making it more suitable for deployment on embedded devices. In some preferred embodiments, the feature expression can be enhanced by introducing an attention mechanism. For example, the CBAM (Convolutional Block Attention Module) attention mechanism can be used to suppress noise and focus on significant features, thereby enhancing the feature extraction ability of the network. The SPPF module at the end of the backbone network realizes multi-scale feature extraction, significantly improving the network performance and efficiency.
[0054] The neck network adopts the ASFF (Adaptive Spatial Feature Fusion) module to dynamically fuse multi-layer feature maps, enhance feature fusion, and improve the multi-scale object recognition ability. Depthwise separable convolution is used to reduce network parameters, improve computational efficiency, and make the network more easily portable.
[0055] One of the main limitations of single-stage detectors based on feature pyramids is the inconsistency between different feature scales. In the blood cell detection task, there are large size differences between red blood cells, white blood cells, and platelets in the image, resulting in poor feature extraction performance of traditional pyramid structures. To enhance the feature extraction ability of the network, the present invention introduces an adaptive feature fusion mechanism. Adaptive Spatial Feature Fusion (ASFF) enables the network to directly learn how to spatially filter features from other levels, enhance scale-invariant feature representations at minimal computational cost, and introduce negligible inference overhead. The key mechanism of ASFF is to adaptively learn the spatial fusion weights of feature maps at each scale, including two steps: equal scaling and adaptive fusion.
[0056] Two main adjustment strategies are adopted during the scaling process: upsampling and downsampling. For upsampling, a 1x1 convolutional layer is used to align the channel size with that of layer l, and then interpolation is performed to increase the spatial resolution. In the case of 0.5x downsampling, a 3x3 convolutional layer with a stride of 2 is applied to adjust both the resolution and the number of channels simultaneously. For 0.25x downsampling, a max pooling operation with a stride of 2 is pre-added to the convolutional layer. This unified representation enables the network to perform adaptive fusion across multiple scales, enhancing its ability to effectively capture multi-scale features. This process ensures that all feature maps, regardless of their original size, can make a meaningful contribution to the final detection output.
[0057] Taking Figure 1 the network architecture in
[0058]
[0059] and
[0060] where represents the feature vector at position (i, j) adjusted from layer n to layer l, is the vector on the output feature map, represents the spatial importance weights for fusing three different layers into layer l. These are parameters automatically learned by the network and are simple scalar variables that can be shared across all channels. λ is obtained by applying a 1x1 convolutional layer to x 1 , x 2,x 3 and calculated, thus allowing them to learn through standard backpropagation.
[0061] The ASFF module adaptively fuses feature maps from different scales, weights the features at multiple scales, and can better identify and separate overlapping targets. Its core idea is to weight and fuse them according to the effectiveness of feature maps at different scales, so as to highlight important spatial information and reduce unnecessary interference.
[0062] Spatial adaptive weighting: ASFF assigns different weights to feature maps at each scale, so that important features receive more attention while irrelevant features are suppressed. This helps to avoid the mixing of features of multiple targets in the overlapping area, thus reducing misclassification or mislocalization.
[0063] Multi-scale information fusion: ASFF extracts information from different scales (such as low-resolution and high-resolution feature maps) and fuses them. By processing images at multiple scales, the detection ability of the model for small targets, overlapping targets, and edge targets can be improved.
[0064] Through adaptive multi-scale feature fusion, ASFF can distinguish the features of different targets in the overlapping area. Even if they overlap significantly in space, the model can better identify each target, thereby improving the accuracy. Especially in dense scenes, ASFF can adjust the feature weights of each scale to reduce the negative impact of overlapping targets on the detection accuracy.
[0065] The backbone network of the model proposed by the present invention adopts an improved ShuffleNet module, namely ShuffleNet V2. ShuffleNet V2 can improve the detection speed of the model while basically not affecting the detection accuracy, and greatly reduce the total number of model parameters and floating-point operation amounts.
[0066] ShuffleNet V2 consists of two main units, Unit1 and Unit2, as Figure 1As shown. In Unit 1, the input feature map is divided into two groups: one group goes through 1x1 convolution, 3x3 depthwise convolution and another 1x1 convolution, and the other group uses residual connections. Then a channel shuffle operation is performed to merge the results. This design reduces complexity, prevents vanishing gradients, and allows feature reuse, keeping the input and output channels unchanged without changing the image resolution. Unit 2 does not split the channels, but concatenates the results of the two paths before performing channel shuffle, similar to Unit 1. The key difference is that Unit 2 uses a 3x3 depthwise convolution with a stride of 2 for downsampling, halving the size of the input feature map and expanding the receptive field. After concatenation, the output channels are doubled, improving feature extraction. Both units use batch normalization and ReLU activation. Channel shuffle enhances inter-channel communication, improving robustness, accuracy, and computational efficiency, making ShuffleNetV2 very suitable for embedded devices, large-scale image recognition, and real-time detection tasks. Although ShuffleNetV2 uses depthwise separable convolutions to reduce computational complexity, this method may limit the model's ability to learn certain complex features. Additionally, the relatively shallow network architecture and downsampling operations are not conducive to the detection of small objects.
[0067] To address the issues of weak feature extraction ability and difficulty in detecting small objects inherent in ShuffleNetV2, the present invention proposes to incorporate the attention mechanism of the Convolutional Block Attention Module (CBAM). Specifically, the attention mechanism is introduced into one branch of the downsampling Unit 2 in the ShuffleNetV2 architecture. Through the inherent channel attention and spatial attention mechanisms of CBAM, this study aims to emphasize salient feature representations and capture a wider range of context information. This method enhances the network's ability to focus on relevant features, especially those related to small objects, while expanding the effective receptive field. In Figure 1 it can be seen.
[0068] The addition of CBAM allows for adaptive feature refinement, potentially alleviating the limitations imposed by the lightweight nature of ShuffleNetV2. By selectively emphasizing informative features and suppressing less relevant features, this study hypothesizes that this improved architecture will exhibit better performance in small object detection tasks without significantly increasing computational overhead. This modification attempts to strike a balance between the efficiency of ShuffleNetV2 and the enhanced representational capabilities provided by the attention mechanism, potentially resulting in a more powerful and versatile object detection framework.
[0069] To effectively detect targets of different sizes, preserve the spatial information of features, and enhance the robustness of the network, the SPPF method is used in the network model proposed in the present invention. Spatial Pyramid Pooling Fast (SPPF) is an improved version of the SPP module, which reduces the computational overhead while improving the feature extraction efficiency. SPPF uses a maximum pooling operation with a 5×5 kernel and a stride of 1, instead of the parallel pooling branches used in the original SPP. This change significantly reduces the computational cost and memory usage without sacrificing multi-scale feature extraction. In SPPF, the maximum pooling is applied progressively, and the resulting feature maps of each scale are concatenated along the channel dimension. Then, 1x1 convolution adjusts the channel depth, allowing the model to effectively aggregate multi-scale features. This structure captures targets of various scales with lower complexity, making it very effective in real-time object detection tasks, especially in resource-constrained environments.
[0070] To reduce the computational complexity and improve the efficiency while maintaining the performance, depthwise separable convolutions are used instead of standard convolutions in the network proposed in the present invention. Depthwise separable convolutions are a key innovation for reducing the computational complexity and model size in object detection networks. It decomposes the standard convolution into two steps: depthwise convolution, which applies a filter separately to each input channel, and pointwise convolution (1x1 convolution), which combines the outputs across channels. This reduces the number of parameters and floating-point operations while maintaining a similar expressive power. For an input of M channels and N channels, the computational amount is reduced by approximately where D is the kernel size. This efficiency makes it very suitable for mobile and edge devices, reducing latency and energy consumption, while the smaller number of parameters can also improve the generalization ability.
[0071] To evaluate the beneficial effects of the present invention, the publicly available open-source BCCD (BloodCell Count and Detection) dataset is used in the present invention. The BCCD dataset is a small-scale, publicly available dataset designed to evaluate neural network models in the field of blood cell detection and can be obtained at https: / / github.com / Shenggan / BCCD_Dataset. It consists of 364 blood cell images with a size of 640×480×3 pixels. This dataset is mainly characterized by a large number of red blood cells (RBCs), a small number of white blood cells (WBCs), and platelets. The abundance and overlap of red blood cells, combined with the small size of platelets, pose significant challenges to the detection and recognition tasks.
[0072] In the experimental setup, the dataset was divided into a training set, a validation set, and a test set with a split ratio of 6:2:2. Specifically, 60% of the data was used for training the model parameters, 20% for hyperparameter tuning, and the remaining 20% for testing. Considering the limited number of training images, data augmentation techniques were employed to enhance the robustness of the model. The specific augmentation strategies included techniques such as image flipping, grayscale transformation, HSV transformation, brightness and contrast adjustment, and mosaic augmentation. During the training phase, the study set the batch size to 16, initialized the learning rate to 0.01, and the momentum to 0.937. The model was trained for 500 iterations to ensure convergence and optimal performance. During the evaluation phase, an Intersection over Union (IoU) threshold of 0.6 was used to determine the effectiveness of the detection results. If the IoU between the predicted bounding box and the ground truth bounding box was greater than or equal to 0.6, the detection was considered correct. The choice of this threshold was to balance precision and recall, taking into account the inherent uncertainty in bounding box localization in the object detection task.
[0073] The training and testing of this experiment were conducted on different devices. The training phase was completed on Windows 11 using the PyTorch deep learning framework and the SGD optimizer. The processor used was an i9-13900KF, and the GPU was an NVIDIA GeForce RTX 4070. During the testing phase, the device used was a CPU i5-8265U, and the GPU did not support CUDA. For more detailed device configurations in the training and testing phases, please refer to Table 1.
[0074] Table 1
[0075]
[0076]
[0077] To effectively evaluate the performance of the model proposed in this invention, several commonly used technical indicators were adopted to assess the model's performance. These technical indicators included Precision (P), Recall (R), AP, and mAP. Precision represents the proportion of actual positive samples among all samples predicted as positive, and Recall represents the proportion of samples predicted as positive to all samples that should be predicted as positive. AP (Average Precision) represents the average precision of a single class, approximately equal to the area under the P / R curve. mAP is the average of multiple class APs, and the higher the mAP, the higher the precision. The specific calculation formulas are shown as follows.
[0078]
[0079] Figure 2A comparative analysis of various YOLO algorithm model variants and EB-YOLO in terms of parameter count, computational complexity (FLOPs), and frames per second (FPS) is presented. YOLO-Fastest v2 has the fewest number of parameters, while the model proposed in the present invention has a slightly higher number of parameters, almost equivalent to only 289,000 parameters. YOLOv8n has the highest number of parameters, at 3.01M. The advantageous parameter efficiency of EB-YOLO is attributed to the use of the lightweight ShuffleNet v2 architecture, enabling depthwise separable convolutions instead of standard convolutions, and using only two detection heads.
[0080] In terms of computational complexity, the computational requirements of the model proposed in the present invention are second only to YOLO-Fastest v2, at 0.9 GFLOPs, while YOLOv10 requires 8.2 GFLOPs. The reduction in model size translates to a reduction in memory requirements, which is particularly beneficial for deployment on embedded systems and portable devices.
[0081] In terms of inference speed, the model proposed in the present invention demonstrates excellent performance on a low-power desktop-class i5-8265U CPU, at 12.15 frames per second, outperforming other models. This emphasizes the improved efficiency of the model proposed in the present invention in terms of lightweight design and computational speed compared to traditional established models.
[0082] To further evaluate the blood cell detection ability of the model proposed in the present invention, comparative experiments were conducted with mainstream YOLO algorithm models such as YOLOv5, YOLOv8, YOLOv10, and YOLO-fastest v2, which features lightweight design. Researchers used accuracy, recall, mAP, parameter count, and computational complexity (FLOPs) as evaluation metrics. The specific results are shown in Table 2. It can be clearly seen from Table 2 that although YOLO-Fastest v2 has the lowest number of parameters and computational complexity, its performance metrics are suboptimal for the blood cell detection task, with an accuracy of only 71.02%, a recall of 83.4%, and the lowest mAP@50% of 78.9%. YOLOv8n and YOLOv10n exhibit comparable performance, but neither exceeds 90% mAP@50% despite higher parameter counts and computational requirements. Among the classic YOLO algorithm variants, YOLOv5n demonstrates superior performance, reaching 90.1% mAP@50%, while maintaining a lower number of parameters than YOLOv8 and YOLOv10, and almost halving the computational requirements.
[0083] The model proposed by the present invention is superior to YOLOv5n in terms of accuracy and recall, achieving the highest mAP@50% of 92.1%. While obtaining high precision, EB-YOLO has also made significant progress in model compression. The number of parameters is nearly an order of magnitude lower than that of YOLOv10, comparable to YOLO-Fastest v2, and the computational complexity is only 0.9 GFLOPs, slightly higher than YOLO-Fastest v2 but much lower than YOLOv5n. These results indicate that the model proposed by the present invention effectively balances precision and model compression, making it suitable for fast blood cell detection tasks on computationally restricted devices. The considerations in model design include precision and optimization for resource-constrained environments, addressing the need for efficient deployment in scenarios with scarce computational resources.
[0084]
[0085] Table 2
[0086] Figure 3 Shows the P-R curve of the model proposed by the present invention when detecting three types of blood cells in the test set.
[0087] Figure 4 Shows the confusion matrix of the model proposed by the present invention when detecting three types of blood cells in the test set. It is worth noting that EB-YOLO shows extraordinary ability in detecting platelets, which are typical small targets. The prediction accuracies of the model for red blood cells, white blood cells, and platelets reach 83%, 97%, and 100% respectively. These results indicate that the model has strong discrimination ability for various blood cell types existing in the test set.
[0088] Figure 5 Shows the detection results of three representative images in the test set, comparing the detection results of YOLO-Fastest v2 and the model proposed by the present invention. The figure is divided into three rows: the first row is the original test image, the second row is the detection result of YOLO-Fastest v2, and the third row is the result obtained by EB-YOLO proposed by the present invention. The comparative analysis of the detection results shows that the model proposed by the present invention shows superior ability in identifying overlapping red blood cells, while YOLO-Fastest v2 fails to detect many overlapping red blood cells. In addition, the model proposed by the present invention shows remarkable ability to accurately detect and depict almost all target objects in the image.
[0089] The experimental results show that compared with the classical YOLO algorithm, the parameters of the model proposed in the present invention are only 0.289M. While the detection accuracy reaches 92.1%, the computational cost is only 0.9 GFLOPs. The improved network model achieves lightweight design while maintaining high quality. This method provides an effective solution for blood cell detection tasks in resource-constrained environments. It helps to develop efficient deep learning models for medical image analysis, especially in cases where real-time processing and edge device deployment are crucial.
Claims
1. A lightweight blood cell detection method based on an improved YOLO network, characterized in that: The following steps are involved: Acquiring an image of blood cells to be detected; Constructing a lightweight blood cell detection model based on an improved YOLO network, and detecting and identifying blood cells in the blood cell image to obtain a detection result; The improved YOLO network includes: Using a backbone network based on the ShuffleNet module to extract multi-scale features from the blood cell image to obtain a multi-scale feature map; A neck network based on the ASFF module is constructed to perform multi-scale feature fusion, and a set number of detection heads are generated; the ASFF module is used to adaptively learn the spatial fusion weights of the feature maps at each scale, and perform dynamic multi-scale feature fusion according to the fusion weights.
2. The method according to claim 1, characterized in that: The adaptive learning of the fusion weights of the feature maps at each scale and the dynamic multi-scale feature fusion according to the fusion weights include: At least one of an upsampling operation, a 0.5-fold downsampling operation, or a 0.25-fold downsampling operation is used to align the scale of each feature map with the scale of the output feature map to obtain a corresponding feature map to be fused; Based on the adaptive fusion mechanism, the spatial fusion weights of the feature maps to be fused are obtained by using the back propagation algorithm, and then the output feature map is calculated to be the weighted sum of all the feature maps to be fused.
3. The method according to claim 2, characterized in that The adaptive fusion mechanism is expressed as: in, represents the feature vector at position (i, j) adjusted from scale n to scale l, represents the feature vector at position (i, j) on the output feature map, It represents the spatial fusion weight of the feature maps to be fused at n different scales to be fused to scale l, and λ is the network parameter.
4. The method according to claim 2, characterized in that: The upsampling operation includes: Using a 1×1 convolutional layer to align the channel size of the feature map with the channel size of the output feature map; Interpolate the feature map after the alignment operation.
5. The method according to claim 2, characterized in that: The 0.5 times downsampling operation includes: A 3×3 convolutional layer with a stride of 2 is used to simultaneously adjust the resolution and number of channels of the feature map.
6. The method according to claim 2, characterized in that The 0.25 times downsampling operation includes: Perform a maximum pooling operation with a stride of 2 on the feature map; A 3×3 convolutional layer with a stride of 2 is used to simultaneously adjust the resolution and number of channels of the feature map after the maximum pooling operation.
7. The method according to claim 1, characterized in that The backbone network based on the ShuffleNet module is used to extract multi-scale features from the blood cell image to obtain a multi-scale feature map, including: Passing the blood cell image through a maximum pooling layer, a first improved ShuffleNet module, and three first ShuffleNet modules in sequence, outputting a first scale feature map; The first scale feature map is sequentially passed through the second improved ShuffleNet module and seven second ShuffleNet modules. Output the second scale feature map; The second scale feature map is sequentially passed through a third improved ShuffleNet module, three third ShuffleNet modules and an SPPF module to output a third scale feature map.
8. The method according to claim 7, characterized in that The improved ShuffleNet module includes two downsampling branches, and the output vectors of the two downsampling branches are concatenated and then rearranged in channels, wherein one of the downsampling branches introduces an attention mechanism to enhance the feature expression in the channel dimension and the spatial dimension.
9. The method according to claim 8, characterized in that One of the downsampling branches of the improved ShuffleNet module includes a 3×3 convolutional layer, a 1×1 convolutional layer, and a CBAM module connected in sequence, and the other downsampling branch includes a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence.
10. The method according to claim 7, characterized in that The multi-scale feature fusion includes: Performing convolution operations on the first scale feature map, the second scale feature map, and the third scale feature map, respectively, to obtain a fourth feature map, a fifth feature map, and a sixth feature map; After the fourth feature map and the fifth feature map are processed by the first ASFF module and the first DWConv module in sequence, a seventh feature map is output; After the fourth feature map and the fifth feature map are processed by the second ASFF module and the second DWConv module in sequence, an eighth feature map is output; The sixth feature map, the seventh feature map, and the eighth feature map are processed by the third ASFF module and the third DWConv module in sequence and then subjected to a convolution operation to obtain a first detection feature map; The sixth feature map, the seventh feature map and the eighth feature map are processed by the fourth ASFF module and the fourth DWConv module in sequence and then subjected to a convolution operation to obtain a second detection feature map.