Method and system for identifying blood cell image based on multi-scale ViT model

By using a multi-scale Transformer model and image enhancement technology, the problem of insufficient adaptability and robustness of single-scale models in blood smear cell identification was solved, realizing high-precision automated identification and counting of blood smear cells, generating structured reports, and improving the accuracy and efficiency of the detection process.

CN120932233APending Publication Date: 2025-11-11INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510981596.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing single-scale ViT models for automatic cell identification and counting in blood smears suffer from poor adaptability to cell morphology diversity, low identification accuracy in complex backgrounds, and insufficient model robustness, and cannot achieve full automation of the process.

Method used

Employing a multi-scale Transformer model, this system utilizes a multi-scale patch embedding mechanism and a multi-head self-attention mechanism, combined with image enhancement and an end-to-end automatic cell identification and counting reporting system, to achieve full expression and fusion of cell features at different scales. It also integrates a field-of-view quality assessment module to improve the accuracy and stability of cell classification and counting.

Benefits of technology

It significantly improves the accuracy and stability of blood smear cell classification and counting, and realizes full automation from raw blood smear images to cell type identification and counting, reducing human intervention and errors, and improving the accuracy and practicality of the detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932233A_ABST
    Figure CN120932233A_ABST
Patent Text Reader

Abstract

The invention provides a blood cell image recognition method based on a multi-scale Vision Transform model, and the method comprises the steps: constructing a multi-scale feature extraction architecture, combining the global modeling capability of a Transform network, carrying out the parallel processing of Patches of different sizes, and fusing the multi-scale information, thereby achieving the recognition of a blood cell image. The cells in the blood smear can be comprehensively perceived from microscopic details to macroscopic contexts, so that not only is the adaptability of the model to cell size difference and morphological diversity enhanced, but also the recognition stability under the conditions of cell overlapping and dyeing difference in a complex visual field is improved; the accuracy and calculation efficiency of cell classification and counting are effectively improved, and full-process automation from visual field quality evaluation to automatic cell recognition and counting is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing and artificial intelligence technology, specifically relating to a method and system for identifying six types of blood cells based on blood smear images, which can be applied to scenarios such as fully automated blood analysis and assisted diagnosis. Background Technology

[0002] Blood cell classification and counting in whole blood images are fundamental for the diagnosis of hematological diseases and various systemic diseases. Traditional manual blood smear analysis relies on manual selection of appropriate fields of view for identification and classification, which is time-consuming and highly subjective. In recent years, deep learning technology has been gradually introduced into the field of blood image processing. Related research focuses on single cell image classification and analysis of a small number of smear images, mainly using convolutional neural networks (CNN) or transfer learning methods. Tamang T et al. used multiple pre-trained CNN models (such as VGG, ResNet, Inception, etc.) to perform transfer learning classification on leukocyte images and achieved good accuracy on multiple public datasets, but still relied on finely cropped single cell images, making it difficult to process original field of view images (Diagnostics.2022;12(12):2903.). Ahmad, R et al. proposed using image information entropy as a feature constraint to guide the CNN model to optimize feature expression. Although this improved the classification accuracy, the model structure was complex, and the generalization ability still depended on the quality of data preprocessing, making it unable to effectively handle high-noise fields of view (Diagnostics(Basel,Switzerland),13(3),352). Muhamad, H et al. systematically evaluated the performance of different CNN architectures on the WBC classification task, pointing out that the quality of single-cell cropping has a significant impact on the results, the model cannot directly process the original full-field microscopic images, and there is no quality control mechanism (In 2022 8th International Engineering Conference on Sustainable Technology and Development (IEC) (pp.205-211). IEEE.).

[0003] The methods mentioned above are mostly based on single-cell images for classification and recognition, and cannot achieve automatic control of smear quality, automatic region selection, and report generation. They are not suitable for full-field images used directly in clinical practice. In addition, existing methods mostly lack structured output capabilities, and cannot automate the entire process of counting, locating, and generating reports for different cell types.

[0004] While significant progress has been made in blood cell classification and recognition using models such as CNNs or transfer learning, existing methods all have their limitations. Specifically, different CNN architectures rely on finely segmented single-cell images for blood cell image classification tasks, and the segmentation quality of these single cells significantly impacts the results. These models cannot directly process raw full-field microscopic images and lack quality control mechanisms. Furthermore, these models cannot automatically adapt to cell images of different sizes, have poor contextual modeling capabilities, and are limited in handling complex real-world situations such as cell overlap, staining differences, and morphological diversity. The model results are also unstructured, failing to automatically count the number of different white blood cell types and generate complete diagnostic reports. Most existing ViT models employ only a single-scale patch segmentation method, making it difficult to account for the diversity of blood cells in size and morphology. This is particularly problematic in whole blood smears with significant cell size differences and complex backgrounds, severely limiting recognition accuracy and robustness. Moreover, the single-scale design makes it difficult for the model to simultaneously capture detailed cell features and overall spatial distribution, affecting the accuracy of classification and counting.

[0005] The purpose of this invention is to address the problems of poor adaptability to cell morphology diversity, low recognition accuracy in complex backgrounds, and insufficient robustness of existing single-scale ViT models in the automatic identification and counting of cells in blood smears. This invention proposes a multi-scale Transformer-based method and system for the automatic identification and counting of cells in blood smears, which fully expresses and integrates cell features at different scales, improving the accuracy and stability of cell classification and counting. Simultaneously, it integrates a field-of-view quality assessment module to ensure the quality of input data, thereby improving the accuracy and practicality of the overall detection process. Summary of the Invention

[0006] To achieve the above objectives, this invention first provides a method for blood cell image recognition based on a multi-scale Vision Transformer model, the method comprising the following steps:

[0007] S1: Obtain a multi-scale blood cell image dataset based on blood smears at different microscope magnifications, and perform image size unification and image enhancement processing on the blood cell image data. The blood smears are obtained from blood samples after flow cytometry sorting, including erythrocytes, lymphocytes, monocytes, neutrophils, eosinophils, and basophils. The dataset is divided into a training set, a validation set, and a test set. The scales include cell level, block level, and region level.

[0008] S2: Input the blood cell image dataset obtained in step S1 into the pre-trained model based on multi-scale Vision Transformer for training and validation. The network model parameters with the highest accuracy on the validation set are retained as the final training result to obtain the image recognition model based on multi-scale Vision Transformer. The pre-trained model of multi-scale Vision Transformer includes multiple parallel Transformer branches, each branch corresponding to an input image scale.

[0009] S3: Obtain the dataset of blood cell images to be detected under different microscope magnifications based on blood smears, and perform image size unification and image enhancement processing on the blood cell image data, wherein the blood smears are from clinical blood samples;

[0010] S4: Input the blood cell image dataset to be detected obtained in step S3 into the image recognition model based on multi-scale Vision Transformer obtained in step S2 to identify cell types including red blood cells, lymphocytes, monocytes, neutrophils, eosinophils and basophils, and count each cell type.

[0011] S5: Output the count of each cell type obtained in step S4.

[0012] In a preferred embodiment, the image enhancement processing in step S1 includes image size unification: scaling the input image to a specified size; and image data enhancement for image rotation, flipping, random cropping and scaling, color perturbation, noise addition and blurring, wherein the training set, validation set and test set are divided in a 6:2:2 ratio.

[0013] In a preferred embodiment, the multi-scale Vision Transformer-based pre-trained model described in step S2 includes a multi-scale image encoder module for extracting semantically rich and structurally stable deep feature representations from blood cell images at different resolutions; a self-attention mechanism encoder module for fusing image features at various scales to generate image embedding representations of cell morphology features at different scales in the blood image; a feature fusion module for merging the embedding representations of all images into a unified global representation for subsequent cell type discrimination; a fully connected layer for converting the global representation into probability distributions of each cell category, for calculating the training loss of the pre-trained model and feeding the calculation results back to the pre-trained model until the model converges; and a parameter optimization module for adaptively adjusting parameters based on the gradient of the loss evaluation module.

[0014] In a preferred embodiment, the multi-scale image encoder module consists of sub-encoders at various scales, each sub-encoder including a multi-scale patch segmenter and an embedding encoder, and its operation is as follows:

[0015] Perform patch segmentation on the image at each scale; divide the entire image into a fixed number of image patches and flatten them into a vector sequence;

[0016] Map each patch vector to a high-dimensional space; add positional encoding and scale encoding; obtain the token sequence at the scale.

[0017] Given the image of the i-th image sample at the s-th scale, represented as I i s Then, the sub-encoder subEs at the s-th scale maps it to a fixed-dimensional image feature embedding vector:

[0018] z i s =E s (I i s (1)

[0019] In a preferred embodiment, the algorithm by which the self-attention mechanism encoder module fuses features at various scales and calculates zi for the unified image embedding representation is as follows:

[0020] z i =∑α s ·z i s (2)

[0021] Where α s Let z be the weighting coefficient for the s-th scale. i s Let be the image feature embedding vector at the s-th scale. After calculation, the outputs for all scales are:

[0022] z1, z2, ..., zi.

[0023] In a preferred embodiment, the algorithm of the feature fusion module is as follows:

[0024]

[0025] Where zi represents the Tranformer encoded feature at the i-th scale.

[0026] In a preferred embodiment, the algorithm for converting the global representation into probability distributions for each cell category is as follows:

[0027] y = Softmax(W × z) final +b) (4)

[0028] Where W and b are the learnable weight matrix and bias vector of the classification layer, respectively, and y represents the model's predicted probability for red blood cells, lymphocytes, monocytes, neutrophils, eosinophils, and basophils.

[0029] In a preferred embodiment, the loss evaluation module calculates the training loss by minimizing a multi-class cross-entropy loss function, and the algorithm for minimizing the multi-class cross-entropy loss function is as follows:

[0030]

[0031] Where N represents the number of samples, C represents the number of cell types, and y ic Indicates the real label, p ic The probability predicted by the model, where ∈ is a small constant introduced to prevent numerical instability, and is set to 1×10 by default. -6 .

[0032] In a preferred embodiment, the parameter optimization module includes an AdamW optimizer and a cosine annealing scheduling module, wherein the initial learning rate is set to 1e-4, the batch size is 32, the maximum training time is 50 epochs, and training is stopped early when there is no improvement in the performance on the validation set. The stability of the trained model is evaluated using a five-fold cross-validation strategy.

[0033] In a preferred embodiment, the image size unification and image enhancement processing in step S3 includes: image normalization, color space adjustment and noise reduction, and dividing each image into 500 to 2000 local regions.

[0034] In a preferred embodiment, the cell type identification in step S4 includes: cell location identification using a multi-scale ViT model, cell type identification, and DBSCAN clustering of cell locations to identify duplicate markers in dense regions. The neighborhood distance radius is set as follows: erythrocytes: 10, platelets: 8, neutrophils: 15, eosinophils: 13, monocytes: 16, lymphocytes: 12. Subsequently, a non-maximum suppression algorithm is used to remove overlapping candidate boxes. Based on the deduplicated cell identification results, the model calculates the number of each cell type and its proportion in the total number of cells. The IoU value for removing overlapping candidate boxes using the non-maximum suppression algorithm is set as follows: erythrocytes: 0.20, platelets: 0.15, neutrophils: 0.30, eosinophils: 0.30, monocytes: 0.35, lymphocytes: 0.40.

[0035] Secondly, this invention provides a system for blood cell image recognition using the aforementioned method based on the multi-scale Vision Transformer model. The system for blood cell image recognition includes the following modules:

[0036] (1) Blood cell image dataset acquisition and data augmentation module;

[0037] (2) Based on the multi-scale Vision Transformer pre-training module;

[0038] (3) Data acquisition and data augmentation module for blood cell images to be detected;

[0039] (4) Image recognition and counting module based on multi-scale Vision Transformer;

[0040] (5) Image recognition and counting result output module.

[0041] To address the various technical problems existing in current image recognition models for blood cell images, this invention provides the following technical means that differ from existing image recognition models:

[0042] 1. Multi-scale Patch Embedding Mechanism: This invention proposes a multi-scale patch embedding mechanism, designing multiple scale-adjustable patch segmenters and corresponding embedding encoders to segment and extract features from input blood smear images at different granularities. Each scale pathway can focus on the local morphological features of cells (such as nucleoplasmic structure and edge morphology) and global distribution features (such as cell arrangement and density). Subsequently, the multi-scale information is fused and input into the Transformer encoder for unified modeling. This mechanism significantly improves the model's ability to adapt to and recognize blood cells of different sizes and shapes (such as erythrocytes and macrophages), enhances the model's detail expression and semantic resolution, thereby improving the accuracy of cell classification and detection.

[0043] 2. Multi-scale Transformer Encoding Architecture; This invention constructs a multi-scale Transformer encoding architecture suitable for blood cell image recognition. This structure integrates multi-scale patch features during the encoding process and models long-distance dependencies between different regions in the image through a multi-layer self-attention mechanism. It mines the contextual structure and category discrimination information between cells, effectively improving the model's discrimination ability in complex cell mixed distribution scenarios in blood smears. This ensures stable recognition performance even under interference conditions such as multi-cell overlap, blurred cell edges, and complex staining.

[0044] 3. End-to-End Automated Cell Recognition and Counting Reporting System; This invention integrates an end-to-end blood cell recognition and statistical analysis system, including modules for blood smear image input, multi-scale Transformer feature extraction, cell detection and classification, cell counting, and structured report output. This system can accurately identify and automatically count six types of blood cells (erythrocytes, lymphocytes, monocytes, neutrophils, eosinophils, and basophils) in images, ultimately generating a cell distribution analysis report for doctors' reference. It achieves fully automated analysis from raw blood smear images to cell type identification and counting, significantly reducing manual intervention and errors, improving system efficiency and clinical diagnostic usability, and possessing significant engineering deployment and practical value.

[0045] In summary, this invention, based on a multi-scale Transformer structure, overcomes the shortcomings of existing single-scale ViT models in cell recognition, particularly their poor adaptability to cell morphological diversity and complex backgrounds, achieving efficient extraction and fusion of multi-scale cell features. This invention, through multi-scale patch embedding and multi-head self-attention fusion mechanisms, effectively improves the model's accuracy and robustness in recognizing cells of different sizes and shapes, significantly enhancing the precision and stability of blood smear cell classification and counting. Furthermore, the streaming cytometry sorting data annotation and preprocessing module introduced in this invention automatically annotates different cell types as image training sets, ensuring the reliability and consistency of downstream cell recognition analysis and improving the overall efficiency and accuracy of the detection process. This invention integrates an end-to-end automated cell counting and classification report generation system, reducing manual operation costs, realizing intelligent and automated blood smear analysis, and improving diagnostic speed and clinical application value. Compared with existing technologies, this invention significantly improves model performance, application stability, and ease of use, possessing broad application prospects. Attached Figure Description

[0046] Figure 1 Flowchart of a method for recognizing blood cells in images based on a multi-scale ViT model;

[0047] Figure 2 .Structural diagram of the multi-scale ViT cell type recognition module. Detailed Implementation

[0048] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as a result of the description. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of protection defined by the claims of the present invention.

[0049] Example 1. Construction of a multi-scale Vision Transformer (ViT) model

[0050] This invention proposes a method for blood cell image recognition based on a multi-scale Vision Transformer (ViT) model. Figure 1 A flowchart illustrating the overall workflow of the model is provided. The multi-scale ViT prediction model improves the accuracy of cell image recognition by constructing a multi-scale feature extraction architecture and combining it with the global modeling capabilities of the Transformer network. The technical implementation steps of this invention are described in detail below.

[0051] 1. Construction of the training data collection and preprocessing module

[0052] The blood cell image data used in the training of this invention mainly comes from peripheral blood samples that have undergone flow cytometry sorting. Specific steps include:

[0053] 1.1 Flow cytometry was used to perform multi-parameter detection and sorting of cells in blood samples. Based on cell size, morphology, and surface markers, high-purity sorting of major blood cell types (including neutrophils, eosinophils, basophils, monocytes, and lymphocytes) was achieved. Each cell subpopulation was sorted and corresponding to separate cell image data, ensuring accurate labeling and classification of image samples.

[0054] 1.2 Standardized Blood Smear Preparation and Multi-Scale High-Resolution Imaging: In the experiment, target cells (such as T cells, B cells, and NK cells) sorted by flow cytometry were first prepared into standard blood smears. After resuspending the cells in buffer solution, they were evenly spread on the surface of a glass slide at an appropriate concentration. An automated slide spreading device or manual smearing method was used to form a monolayer cell membrane, avoiding overlap and blank areas to ensure intact cell morphology and uniform distribution. After drying, the smears were stained with Wright-Gymsa stain. This staining method has good staining contrast and cell component differentiation ability, clearly displaying the nuclear structure, cytoplasmic chromatin texture, and granule distribution, facilitating subsequent subtype identification and differentiation of various leukocytes and erythrocytes. After staining, the smears were imaged using a multi-magnification automated microscope. Commonly used magnification levels include: low magnification (10×): for acquiring the overall structure, quality, and preliminary cell distribution of the smear; medium magnification (40×): for acquiring the overall morphology of cells, which is the main analytical field of view; and high magnification (100×): for observing subcellular structures such as the nucleus and granules, which is a key scale for identifying cell subtypes. During imaging, the same cell region will be acquired at multiple scales at the above different magnifications. The corresponding image sizes can be set as follows: 256×256 pixels (10×); 512×512 pixels (40×); and 1024×1024 pixels (100×). This "multi-scale image" not only refers to the image acquisition results at different magnifications, but also reflects the differences in spatial resolution and structural information levels of the image, providing rich input features for subsequent deep learning models. To ensure consistent and comparable image quality, all images undergo a standardized image preprocessing workflow, including: color correction: color shift correction of the RGB channels based on a standard whiteboard image to ensure consistent staining; contrast enhancement: Adaptive histogram equalization (CLAHE) to enhance cell boundaries and particle clarity; denoising: Non-local means (NLM) or bilateral filtering to remove imaging noise while preserving edge structure; size normalization: Images at different magnifications are uniformly scaled to a model-compatible standard input size (e.g., 512×512) to form a unified image data format. Ultimately, the system will establish a high-quality image dataset containing multiple magnifications, scales, and consistent staining, serving as the input basis for deep model training and blood cell classification analysis. All images retain corresponding cell type labels (neutrophils, eosinophils, basophils, monocytes, lymphocytes), ultimately forming a training dataset with realistic cell labels and spatial distribution.

[0055] The blood cell image dataset contains Giemsa staining images of six different cell types. Each image is H×W×C in size, where H represents the height of the image, W represents the width of the image, and C is the number of channels of the image (C=3 for RGB images).

[0056] 1.3 Image Preprocessing

[0057] To ensure that the image data input to the cell recognition model has a uniform size specification and strong robustness, the system performs a preprocessing step on the original blood smear image before it is sent to the model. The preprocessing step includes two parts: size uniformization processing and data augmentation processing.

[0058] 1.3.1 Size standardization

[0059] To address the issue of inconsistent image resolution caused by different microscope magnifications (e.g., 10×, 40×, 100×) and different image acquisition devices, this invention performs size normalization on all image data. The size normalization process includes the following steps:

[0060] 1) Obtain the original image size information;

[0061] 2) Use proportional scaling to adjust the long or short side of the image to the target size range while keeping the aspect ratio of the image unchanged;

[0062] 3) For any portion that is smaller than the target size, fill it using symmetrical filling (such as zero filling or reflective filling);

[0063] 4) The target image size is preferably set to 512×512 pixels to balance model recognition performance and computational efficiency.

[0064] By following the steps above, we can ensure that all images have a uniform size specification when inputting into the model, thereby improving the model's training convergence speed and inference stability.

[0065] 1.3.2 Data Augmentation Processing

[0066] This invention further performs data augmentation processing on the image after it has been resized. The image augmentation processing in the training module is characterized by performing various, random data augmentations (such as rotation, flipping, cropping, color transformation, noise addition, blurring, etc.) to increase the diversity of training data, improve the model's generalization ability and robustness, and enable it to adapt to various complex real-world imaging conditions. The data augmentation parameters during training (such as rotation angle range, flipping probability, color perturbation amplitude, etc.) are typically set to a large range of random values. The execution conditions and evaluation criteria for the augmentation processing include:

[0067] 1) Data augmentation should be performed without altering image labels and key information about cell structure;

[0068] 2) The structural similarity index (such as SSIM) between the generated enhanced image and the original image is not less than 0.85 to ensure that cell morphology information is not destroyed;

[0069] 3) Evaluate the performance of the validation set before and after enhancement. Preferably, the accuracy or F1 score of the enhanced model on the validation set should be improved by ≥2%.

[0070] The data augmentation process includes, but is not limited to, one or more of the following image transformation operations:

[0071] 1) Rotation transformation: Randomly rotate the image within the range of 0° to ±30° to simulate cell morphology from different viewpoints;

[0072] 2) Flipping operations: including horizontal and vertical flipping, suitable for cell images with symmetrical structures, enhancing the orientation invariance of the model;

[0073] 3) Random cropping and scaling: Without damaging the main structure of the cells, the image is cropped from the center (the area retained is no less than 70% of the original image area) to simulate the situation of some cells shifting or the edge of the field of view in actual shooting.

[0074] 4) Color perturbation: Make minor adjustments to the brightness, contrast, and saturation of the image to adapt to different dyeing batches or lighting conditions;

[0075] 5) Noise addition and blurring: Introduce Gaussian noise or Gaussian blur to simulate noise or focal length shift in image acquisition.

[0076] The data augmentation operations are performed in random combinations using image augmentation libraries (e.g., Albumentations, imgaug, or torchvision.transforms). Each original image can generate 1 to 3 augmented images, forming an expanded training sample set.

[0077] Through the above preprocessing methods, this invention can construct a high-quality blood smear image dataset that is standardized, robust, and suitable for deep learning model training, which helps to improve the accuracy and stability of blood cell identification systems.

[0078] 2. Construction of a multi-scale image encoder

[0079] 2.1 Multi-scale image encoder structure design:

[0080] The design goals of the multi-scale image encoder provided by this invention include: capturing the morphological structural differences of blood cells at different image scales; extracting deep feature representations with discriminative capabilities to enhance the ability to identify heterogeneous cells. By constructing image pyramids from multiple scale versions of the input blood cell image, such as at different resolutions (original magnification, 40×, 100×, etc.), the multi-scale image encoder extracts and fuses features from different scales in the blood cell image. It extracts semantically rich and structurally stable deep feature representations from blood cell images at different resolutions, effectively capturing the morphological differences of different cell types at multiple granular scales and enhancing the model's ability to identify heterogeneous cells.

[0081] The deep feature representations described in this invention include, but are not limited to, the following types:

[0082] Marginal features: such as the smooth boundary of lymphocytes and the multi-segmented nuclear margin of neutrophil segmented nuclei;

[0083] Texture features: such as lightly stained granules in the cytoplasm of monocytes, and red eosinophilic granules in the cytoplasm of eosinophils;

[0084] Structural features include: nucleus size, cytoplasm-to-nucleus ratio, and distribution of granules in the cytoplasm.

[0085] Color characteristics: Under Wright-Gymsa staining, there are chromatographic differences such as lymphocytes showing blue cytoplasm and erythrocytes showing pink.

[0086] Examples of morphological differences among different types of blood cells at multiple scales: In low-magnification (10×) images, erythrocytes appear as uniformly distributed discs with no visible nucleus; while leukocytes have a larger area and a clear outline; in medium-magnification (40×) images, neutrophils show a segmented nuclear structure, and lymphocytes appear as round cells with a high nucleocytoplasmic ratio; in high-magnification (100×) images, dense eosinophilic granules can be observed in the cytoplasm of eosinophils, while basophils show larger coarse granular chromatin distributed in a purplish-black color.

[0087] The encoder can be built using a pre-trained multi-scale vision Transformer (ViT) architecture.

[0088] The multi-scale ViT encoder design includes multiple parallel Transformer branches (ES), each corresponding to one input image scale. Each branch contains an image embedding module, a self-attention module, and an inter-layer fusion module.

[0089] 2.1.1 Image Embedding Module: Patches images of different resolutions and maps them to tokens;

[0090] Given the image of the i-th image sample at the s-th scale, represented as I i s Then through encoder E s Map it to a fixed-dimensional image feature embedding vector:

[0091] z i s =E s (I i s (1), Where Es is the ViT sub-encoder of the s-th scale, which can be a parameter-shared or non-shared structure.

[0092] Parameter sharing structure (shared weights): Inputs at all scales share the same encoder, reducing model complexity and improving generalization ability;

[0093] Non-shared parameter structure (independent weights): Each scale uses an independent encoder, allowing for the learning of specific features for different resolutions, suitable for situations with significant scale differences or large task differences. In this invention, Es's design strategy is flexible and iterative. Initially, we tend to share the same encoder for all images of different scales, which allows for faster establishment of a model baseline, simplifies the model, and improves generalization ability. During the optimization phase, if we find that image processing at a specific scale is not performing well, we will set up a separate encoder for that scale without sharing parameters to more accurately capture its unique information. Ultimately, our goal is to find an optimal configuration that achieves a perfect balance between model performance, complexity, and training efficiency, which is usually a hybrid strategy combining shared and non-shared encoders.

[0094] In this invention, the ViT sub-encoder Es mainly comprises a multi-scale patch segmenter and an embedded encoder.

[0095] 2.1.1.1 Multi-scale Patch Segmenter

[0096] Perform patch segmentation on each scale image Iis (e.g., 16×16 pixels per patch);

[0097] The entire image is divided into a fixed number of image patches and flattened into a vector sequence;

[0098] This operation corresponds to the image preprocessing module in ViT, and its output is a patch vector xpatchs∈RNs×p2·cx, where: Ns is the number of patches at this scale, p is the patch size (e.g., 16), and c is the number of channels (usually 3).

[0099] 2.1.1.2 Patch Embedding

[0100] Each patch vector is mapped to a high-dimensional space Rd using a linear mapping Linear(·) or convolutional embedding module; position embedding and scale embedding are added; the token sequence at scale s is obtained and used as the input to the ViT encoder Es.

[0101] Therefore, the patch segmenter and the embedding encoder are embodied in the internal implementation details of formula (1), namely: Segmentation and embedding.

[0102] These modules determine whether the model can fully extract fine-grained morphological features at different scales, directly affecting the classification performance after fusion.

[0103] 2.1.2 Self-attention encoding module: Extracts long-range contextual dependencies within the image, capturing local structure and global information;

[0104] Subsequently, attention-weighted methods are used to fuse features at different scales, i.e., multi-scale patch embedding, to calculate the unified image embedding zi:

[0105] z i =∑α s ·z i s (2)

[0106] Where α s The weighting coefficients for the s-th scale can be learned or set according to the image resolution. Ultimately, this image embedding represents the integration of cell morphology features at different scales in the blood image, enhancing the model's ability to distinguish between large cells (such as monocytes) and small cells (such as platelets).

[0107] Because images at different resolutions have a consistent structure and are easy to expand, this multi-scale image encoder has good versatility and scalability, and can be applied to blood cell image recognition tasks under different magnifications and imaging devices.

[0108] After completing multi-scale feature extraction and self-attention encoding, the outputs at all scales are:

[0109] z1, z2, ..., zi

[0110] Where zi represents the Tranformer encoded feature at the i-th scale.

[0111] 2.1.3 Inter-layer fusion module:

[0112] After multi-scale feature extraction, the following image scales are mainly used for subsequent training modules to learn cell type features, thereby capturing various information of cells from macroscopic to microscopic levels and improving the comprehensive recognition ability of blood cells of different types and sizes.

[0113] Low-resolution images (e.g., 10×) are beneficial for observing the relative size and distribution density of cells, and are suitable for identifying the size ratio between monocytes and lymphocytes. Medium-resolution images (e.g., 40×) can clearly distinguish between erythrocytes and neutrophils, facilitating contour recognition. High-resolution images (e.g., 100×) are suitable for observing nuclear lobulation, granule staining, and edge texture details, and are crucial for identifying eosinophils and basophils. The multi-scale ViT structure extracts semantic features at each scale through independent encoders, and then uses attention-weighted fusion to integrate multi-scale information in a unified embedding vector. It merges feature vectors from all scales into a unified global representation, simultaneously expressing: global structure (cell contours, density distribution); local texture (cytoplasmic granules, nuclear features); and color information (stained cytoplasmic features). This information is used for subsequent cell type discrimination, thereby improving the comprehensive recognition ability of different types and sizes of blood cells.

[0114] The inter-layer fusion mentioned refers to the fusion of features at different scales through a cross-scale token interaction module or attention gating mechanism. For example, based on ViT, a multi-scale feature aggregation module is introduced. This module receives the outputs of 40× and 100× image paths and uses feature weighted fusion or attention mechanisms (such as SE, CBAM) to dynamically weight the importance of features at each scale, thereby improving the model's sensitivity to cellular structural details.

[0115] Specifically, the weighted summation feature fusion strategy used in this invention is defined as follows:

[0116]

[0117] z l : Represents the global image embedding representation at the l-th scale.

[0118] This fusion strategy effectively integrates local details and overall structural information from different scales while maintaining computational efficiency, thereby enhancing the model's expressive power in heterogeneous cell recognition tasks.

[0119] Through the above design, the multi-scale image encoder of the present invention can simultaneously capture overall cell distribution information from low resolution and extract fine-grained nuclear structure and particle features from high resolution, effectively improving the model's ability to identify and classify blood cells with various morphological heterogeneity.

[0120] 2.2 Implementation of Three-Level Morphological Feature Extraction

[0121] To better understand the process of blood cell image recognition using the multi-scale ViT model described in this invention... Figure 2 A specific implementation of the multi-scale ViT cell type identification module is shown, which includes the extraction of the following three levels of morphological features:

[0122] 2.2.1 Cell-level (16×16 pixels): Fine-grained morphological feature extraction ( Figure 2 A)

[0123] Input processing: The 256×256 pixel blood image is divided into 256 non-overlapping 16×16 cells (tokens), each token corresponding to a single cell or a local cellular structure. At 20× magnification, the 16×16 pixel area covers approximately 8μm. 2 The region precisely captures the core morphological characteristics of a single blood cell, such as:

[0124] Red blood cells have a biconcave disc-shaped, nucleus-free structure;

[0125] Lymphocytes have round nuclei and narrow cytoplasm;

[0126] The segmented nuclei and pale purple granules of neutrophils.

[0127] Feature encoding: Each 16×16 token is linearly embedded using the ViT_{256}-16 module to generate a 384-dimensional feature vector, which is then combined with positional encoding to preserve spatial information. This module contains 8 Transformer layers and 6 attention heads, each focusing on different features.

[0128] First 1-2: Differentiate the stromal background from erythrocytes (anucleate, homogeneous);

[0129] Head 3-4: Identify cell nucleus morphology (e.g., the kidney-shaped nucleus of monocytes, the binucleate nucleus of eosinophils);

[0130] First 5-6: Capture cytoplasmic granule characteristics (such as large, purplish-black granules of basophils).

[0131] Output: Each 256×256 image aggregates 256 cell-level features through the [CLS] token to form a block-level (256×256) basic representation, integrating the complete morphological information of a single cell.

[0132] 2.2.2 Block-level (256×256 pixels): Integration of spatial relationships within cell communities ( Figure 2 (B)

[0133] Input processing: The 4096×4096 pixel blood image region is divided into 256 256×256 patches. Each patch contains aggregated features of 256 16×16 cell-level tokens. Each 256×256 pixel patch covers an area of ​​approximately 128μm×128μm, capturing spatial distribution patterns between cells, such as:

[0134] The rouleaux arrangement of red blood cells (distinguishing them from their normal dispersed state);

[0135] Aggregation or isolated distribution of white blood cells (such as aggregation of lymphocytes in the focus of infection).

[0136] Feature aggregation: The 256 patch features are re-encoded using the ViT_{4096}-256 module, which includes 4 Transformer layers and 3 attention heads, focusing on learning mesoscale features.

[0137] First 1-3: Focus on spatial interactions between cells, such as the chemotactic distribution of neutrophils around bacteria;

[0138] Head 4-6: Identify areas of high cell density (such as localized clusters in eosinophilia).

[0139] Output: Each 4096×4096 region aggregates patch-level features through [CLS] tokens to form a mid-level representation of the region (4096×4096), integrating the spatial relationship information of the cell community.

[0140] 2.2.3 Region-level (4096×4096 pixels): Global cell proportion and distribution pattern recognition ( Figure 2 (C)

[0141] Input processing: The entire blood smear is divided into multiple 4096×4096 regions, each containing aggregated features of 256 256×256 patches. Each 4096×4096 pixel area covers approximately 2mm×2mm, capturing global cellular composition and distribution, such as:

[0142] The proportions of different cell types (e.g., the percentage of red blood cells, white blood cells, and platelets);

[0143] Global distribution trends of abnormal cells (e.g., diffuse infiltration of leukemia cells vs. staged distribution of normal bone marrow).

[0144] Feature fusion: Multiple region features are finally aggregated using the ViT_{WSI}-4096 module, which includes two Transformer layers and three attention heads. Positional encoding is weakened to accommodate tissue segmentation differences. The focus is on learning:

[0145] The proportion of cell types in different regions (e.g., the global increase of basophils in an allergic state);

[0146] The correlation between cell distribution and pathological patterns (e.g., the global consistency of microcytic hypochromic changes in erythrocytes in iron deficiency anemia).

[0147] Output: Finally, a global representation of the entire smear is generated using the [CLS] token, which is used for the classification of six major cell types (such as distinguishing between neutrophils and eosinophils) and their number statistics.

[0148] 3. Training and Testing Module

[0149] 3.1 Cell type classification

[0150] The fused feature z final The data is fed into a fully connected layer for cell type classification. The output of this fully connected layer is converted into a probability distribution for each cell type using a softmax function.

[0151] y = Softmax(W × z) final +b) (4)

[0152] Where W and b are the learnable weight matrix and bias vector of the classification layer, respectively, and y represents the model's predicted probability for the six major blood cell types (red blood cells, lymphocytes, monocytes, neutrophils, eosinophils, and basophils). The predicted probabilities are used to calculate the loss function (such as cross-entropy loss), and the model parameters are continuously optimized and updated through training by comparing the difference between the predicted probabilities and the true labels.

[0153] 3.2 Loss Function

[0154] The training objective of the model is to minimize the multi-class cross-entropy loss, defined as follows:

[0155]

[0156] Where N represents the number of samples, C represents the number of cell types, and y ic Indicates the real label, p ic The probability predicted by the model, where ∈ is a small constant introduced to prevent numerical instability, and is set to 1×10 by default. -6 .

[0157] 3.3 Composition of the test dataset

[0158] In this embodiment of the invention, the Peripheral Blood Smear (PBS) Dataset disclosed by 10×Genomics was used as the test dataset. This dataset, provided by multiple medical institutions, contains high-resolution blood smear images acquired after Wright-Gymsa staining, totaling over 15,000 images and covering the following cell types: erythrocytes, neutrophils, eosinophils, basophils, lymphocytes, and monocytes. The image acquisition equipment and staining conditions are suitable for multi-center model generalization performance testing. The dataset was partitioned in a 6:2:2 ratio, i.e., 60% for the training set, 20% for the validation set, and 20% for the test set. To prevent data leakage and improve model generalization ability, all image data were strictly partitioned according to patient ID and acquisition batch, ensuring that images from the same patient or the same acquisition batch would not appear simultaneously in the training, validation, and test sets, with no overlap between them.

[0159] 3.4 The AdamW optimizer combined with cosine annealing scheduling improves the model's generalization ability.

[0160] AdamW decouples weight decay and adaptive learning rate, applying weight decay separately to the parameter update step instead of mixing it into gradient calculation. This achieves more stable regularization. Cosine annealing scheduling is a learning rate scheduling strategy that dynamically adjusts the learning rate using a cosine function, achieving periodic changes in the learning rate during optimization, thereby helping the model escape local optima and improve generalization ability.

[0161] During training, the AdamW optimizer was used with an initial learning rate of 1e-4, dynamically adjusted using a cosine annealing scheduling strategy. The batch size was 32, with a maximum of 50 training epochs. The input image size was a uniformly processed 512×512 pixels, and training was stopped early when there was no improvement in validation set performance.

[0162] 3.5 Data Augmentation Strategies

[0163] To enhance the model's adaptability to different microscopic imaging conditions (such as different magnifications, staining batches, or lighting conditions) and improve its robustness to different microscopic imaging conditions, the following data augmentation strategies were adopted during training: random rotation: angle range ±30°; horizontal / vertical flip: probability of 0.5; color jitter: adjusting brightness, contrast, and saturation, with a perturbation amplitude not exceeding ±20%; blur simulation: using a Gaussian blur kernel (1-3 pixels) to simulate focus deviation.

[0164] 3.6 Stability Assessment

[0165] The model employs a five-fold cross-validation strategy for stability evaluation. Five-fold cross-validation is a classic method for evaluating the generalization ability of a model. It involves dividing the dataset into five subsets, using four of them as the training set and one as the validation set in a loop, and finally combining the evaluation results from all rounds.

[0166] 4. Model Performance Evaluation

[0167] 4.1 Comparison Model Structure Description

[0168] To verify the performance advantages of the model proposed in this invention, the following two control models were designed for comparative experiments:

[0169] Model A: Traditional Convolutional Neural Network (CNN) structure, with ResNet-50 as the backbone, extracting image features and then connecting a fully connected classification layer, without using a multi-scale input mechanism;

[0170] Model B: Single-scale Vision Transformer (ViT) structure, the input image is an image at a single resolution (40×), the standard ViT module is used for feature extraction and classification, without a multi-scale feature fusion module.

[0171] Compared with the above-mentioned comparative model, the model of this invention has the following technical improvements:

[0172] Multi-scale input mechanism: Simultaneously introduce 10×, 40×, and 100× images to simulate the overall structure and microscopic morphology of cells;

[0173] Cross-scale feature fusion module: Introduces an attention gating mechanism to achieve weighted integration of semantic features at different scales;

[0174] High-resolution feature preservation design: The network structure introduces skip connection and multi-level semantic preservation mechanism between multiple Transformer layers to enhance the expression of small features.

[0175] 4.2 Model Testing

[0176] After testing on publicly available datasets, the model results are compared as follows:

[0177] Table 1. Comparison of test results for the three models

[0178] Model Name Precision Recall F1 score Accuracy (AUC) Traditional CNN methods 0.925 0.92 0.922 0.972 Ordinary ViT model 0.901 0.89 0.895 0.891 This invention model 0.945 0.95 0.947 0.987

[0179] Test results show that, compared with existing technologies, the multi-scale Transformer structure proposed in this invention has achieved significant improvements in key evaluation indicators such as accuracy, F1 score and AUC, demonstrating the technical advantages of this invention in blood cell recognition tasks.

[0180] 5. Automatic blood cell identification, counting, and reporting module

[0181] This module, also known as the model's inference module, is based on the multi-scale Vision Transformer (ViT) model. It designs a complete automated workflow for blood cell identification and statistics. Based on high-throughput flow cytometry image data, and combined with cell morphology and background context information, the model completes the type determination of each cell in all input image regions. After that, the model will count the number of each type of cell, achieving accurate identification, location, and counting of six major types of blood cells (erythrocytes, neutrophils, eosinophils, basophils, monocytes, and lymphocytes). This enables automatic identification and classification of blood cells, forming cell distribution vectors for generating visual structured reports to assist in subsequent clinical diagnosis or drug screening analysis.

[0182] In the inference module, random data augmentation is typically not performed. During inference, the model is allowed to directly process the raw, as realistic as possible input images to obtain the most accurate predictions. Therefore, only necessary preprocessing is performed, such as:

[0183] Size standardization: Scaling the input image to a fixed size used during model training.

[0184] Normalization, color space adjustment, and noise reduction: these are standardized preprocessing steps to ensure the consistency of the input data format, keeping it consistent with the training process.

[0185] Image segmentation and local region extraction: The inference module divides the image into local view patches to accommodate the model's ability to process large images, but this process does not involve random augmentation.

[0186] The specific steps are as follows:

[0187] Step 1: Image segmentation and local region extraction

[0188] The raw flow cytometry images are first input into the system for preprocessing, including normalization, color space adjustment, and denoising. Subsequently, the images are divided into multiple local regions; for example, the images are divided into local patches of approximately 64×64 to 128×128 pixels in size. Each image is divided into approximately 500 to 2000 local regions, the exact number depending on the image resolution and cell density. This division is achieved using a sliding window (stride 32 or 64) to ensure that no cell is missed and that each cell falls completely within at least one patch.

[0189] Step 2: Multi-scale feature extraction and cell classification

[0190] The segmented image patches are fed into a multi-scale ViT model. This model combines representation capabilities at different resolutions to learn cell features from both global morphology and local texture. Each local region outputs a predicted label (one of six categories) and its corresponding spatial location information (center point coordinates or bounding box coordinates). The model's output is uniformly represented as a recognition result set, resulting in recognition result C, where (xi,yi) represents the position of the i-th identified cell, and yi belongs to {erythrocyte, neutrophil, eosinophil, basophil, monocyte, lymphocyte} to represent the cell type.

[0191] Step 3: Spatial Clustering and Overlap Removal

[0192] Because of the overlapping regions in the sliding window, the same cell may be detected repeatedly. Therefore, based on the recognition result C, the system first performs DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering on the cell locations in the image to identify duplicate markers in dense regions. Subsequently, the Non-Maximum Suppression (NMS) algorithm is used to remove overlapping candidate boxes, outputting the deduplicated cell recognition result D.

[0193] 1. Spatial clustering based on DBSCAN

[0194] (1) The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm is used to spatially cluster the center points of all candidate detection boxes to automatically identify clusters of redundant boxes that are close together. The parameters involved in spatial clustering are set as follows:

[0195] eps (neighborhood distance radius): This defines whether two candidate box centers are considered to belong to the same cluster if the distance between them is less than this value. This value is set based on the average diameter of different blood cells (usually 6-18 μm) and the image resolution (e.g., 0.5 μm / pixel). See Table 2 for specific parameter settings.

[0196] Table 2. EPS settings for six types of blood cells

[0197]

[0198] min_samples (minimum number of points): Set to 2 uniformly, which means that each cluster needs at least 2 candidate boxes to avoid misclustering of isolated points.

[0199] Feature weight adaptive adjustment: The system assigns different eps values ​​to each candidate box according to the predicted category of cell type to achieve category-aware spatial clustering.

[0200] (2) Clustering output processing:

[0201] Within each cluster, the system retains the candidate box with the highest prediction confidence as the representative detection box for that cluster. All representative boxes proceed to the next step of non-maximum suppression (NMS) to further eliminate spatially overlapping boxes within the same category, resulting in the final accurate cell localization results.

[0202] 2. Non-maximum suppression (NMS) based on class confidence

[0203] Even after spatial clustering, some candidate boxes may still exhibit high overlap. To address this, we perform non-maximum suppression (NMS) based on confidence ranking within each cell class to further remove redundant labels. The processing flow is as follows:

[0204] The clustering results are categorized by cell type;

[0205] Candidate boxes for each cell type are sorted from highest to lowest prediction confidence.

[0206] Starting with the highest confidence box, retain them sequentially;

[0207] After each retention, remove other bounding boxes whose IoU (Intersection over Union) is greater than the threshold (see Table 3 for parameter settings);

[0208] Repeat this process until all candidate boxes have been processed.

[0209] Table 3. IoU Threshold Settings for Six Cell Types

[0210]

[0211] For cells of the same type, they are retained in descending order of prediction confidence scores. After retaining one candidate box at a time, other boxes with an IoU (Intersection over Union) greater than a set threshold (e.g., 0.3) are removed to obtain a unique set of cell localization results.

[0212] Step 4: Cell Count and Type Analysis

[0213] Based on the deduplicated cell identification result D, the system automatically calculates the number of cells of each type and their proportion in the total number of cells. The output statistical results E can be used for subsequent clinical indicator comparison, disease screening, or dynamic monitoring.

[0214] Step 5: Automatic Counting Report Generation and Display

[0215] The final statistical results E are integrated into a standardized blood cell count report F. The report includes:

[0216] (1) Absolute number and relative proportion of each type of cell

[0217] (2) Bar charts / pie charts show cell proportions

[0218] (3) The heatmap shows the spatial distribution of cells in the image (based on coordinates ${(xj,yj)}$).

[0219] (4) Visual overlay: The original image is overlaid with the cell classification results and cell types are labeled with different colors.

[0220] (5) Abnormal prompts: If the proportion of a certain type of cell is abnormal (such as eosinophils > 6%, or red blood cell proportion is low), the system will automatically mark it as abnormal and add a prompt text.

[0221] Example 2. Application of multi-scale Vision Transformer model for blood cell image recognition. Taking a clinical application sample as an example, a full-field flow cytometry image is systematically analyzed.

[0222] 1. Case Background:

[0223] The patient, a 42-year-old female, presented to the dermatology and immunology department of the hospital with a chief complaint of "pruritus accompanied by mild fatigue for one week." Preliminary blood tests showed:

[0224] The total white blood cell count was slightly elevated (WBC = 11.5 × 10⁻⁶). 9 / L)

[0225] The percentage of basophils was 0.6%, slightly higher than the reference range (normal value <0.5%).

[0226] The doctor recommended further evaluation of the blood smear cells to confirm possible signs of allergic or parasitic infection.

[0227] 2. Image Analysis Process

[0228] The system loads a full-field flow cytometry image of the patient (2048×2048 pixels, acquisition resolution 40×) and completes the following steps in sequence:

[0229] Image preprocessing and segmentation: Divide the large image into multiple 1024×1024 image blocks;

[0230] Multi-scale modeling: Generate three different scale (1.0×, 0.75×, 0.5×) image inputs for each image patch;

[0231] ViT Feature Extraction and Fusion: Each scale inputs a multi-scale Vision Transformer encoder to extract local and global morphological features, which are then combined through an attention fusion module to obtain the cell embedding vector.

[0232] Cell detection and classification prediction: Based on the detection head, candidate boxes are output along with the corresponding cell types and confidence levels;

[0233] Spatial clustering and deduplication: Redundant candidate boxes are removed by DBSCAN spatial clustering and NMS non-maximum suppression, generating the final deduplicated result set;

[0234] Results statistics and analysis report generation: Statistically analyzes the number and proportion of various cell types, automatically generates structured diagnostic suggestions, and exports the report.

[0235] 3. The model recognition results are as follows:

[0236] Table 4. Results of whole blood cell identification

[0237]

[0238] 4. Intelligent report output:

[0239] The system automatically analyzes and identifies the results, and provides the following clinical assistance suggestions:

[0240] Abnormal findings: The proportion of basophils is high (0.7% > 0.5%). The system suggests considering the possibility of allergic reaction or parasitic infection, and recommends further examination based on medical history.

[0241] The red blood cell and white blood cell counts were stable, and the cell differential ratio was basically consistent with the blood routine report.

[0242] No abnormal morphological variations were detected in the cells.

[0243] The final report is exported in PDF format and automatically uploaded to the Hospital Information System (HIS) or Laboratory Information System (LIS) for doctors to quickly review and archive in electronic medical records.

Claims

1. A method for blood cell image recognition based on a multi-scale Vision Transformer model, characterized in that, The method includes the following steps: S1: Obtain a multi-scale blood cell image dataset based on blood smears at different microscope magnifications, and perform image size unification and image enhancement processing on the blood cell image data. The blood smears are obtained from blood samples after flow cytometry sorting, including erythrocytes, lymphocytes, monocytes, neutrophils, eosinophils, and basophils. The dataset is divided into a training set, a validation set, and a test set. The scales include cell level, block level, and region level. S2: Input the blood cell image dataset obtained in step S1 into the pre-trained model based on multi-scale Vision Transformer for training and validation. The network model parameters with the highest accuracy on the validation set are retained as the final training result to obtain the image recognition model based on multi-scale Vision Transformer. The pre-trained model of multi-scale Vision Transformer includes multiple parallel Transformer branches, each branch corresponding to an input image scale. S3: Obtain the dataset of blood cell images to be detected under different microscope magnifications based on blood smears, and perform image size unification and image enhancement processing on the blood cell image data, wherein the blood smears are from clinical blood samples; S4: Input the blood cell image dataset to be detected obtained in step S3 into the image recognition model based on multi-scale VisionTransformer obtained in step S2 to identify cell types including red blood cells, lymphocytes, monocytes, neutrophils, eosinophils and basophils, and count each cell type. S5: Output the count of each cell type obtained in step S4.

2. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 1, characterized in that, The image enhancement process described in step S1 includes image size unification: scaling the input image to a specified size; and image data enhancement for image rotation, flipping, random cropping and scaling, color perturbation, noise addition and blurring, wherein the training set, validation set and test set are divided in a 6:2:2 ratio.

3. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 1, characterized in that, The multi-scale Vision Transformer pre-trained model described in step S2 includes a multi-scale image encoder module for extracting semantically rich and structurally stable deep feature representations from blood cell images at different resolutions; a self-attention mechanism encoder module for fusing image features at various scales to generate image embedding representations that integrate cell morphology features at different scales in blood images; a feature fusion module for merging the embedding representations of all images into a unified global representation for subsequent cell type discrimination; and a fully connected layer for converting the global representation into probability distributions for each cell category, calculating the training loss of the pre-trained model and feeding the calculation results back to the pre-trained model until the model converges. A parameter optimization module that adaptively adjusts parameters based on the gradient of the loss assessment module.

4. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 3, characterized in that, The multi-scale image encoder module consists of sub-encoders at various scales. Each sub-encoder includes a multi-scale patch segmenter and an embedding encoder, and its operating procedure is as follows: Perform patch segmentation on the image at each scale; divide the entire image into a fixed number of image patches and flatten them into a vector sequence; Map each patch vector to a high-dimensional space; add positional encoding and scale encoding; Obtain the token sequence at the scale. Given the image of the i-th image sample at the s-th scale, represented as I i s Then, through the sub-encoder of the s-th scale, sub-E s Map it to a fixed-dimensional image feature embedding vector: z i s =E s (I i s ) (1)。 5. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 4, characterized in that, The self-attention mechanism encoder module fuses features at various scales and calculates the z-axis of the unified image embedding representation. i The algorithm is as follows: With i =∑α s ·With i s (2) Where α s Let z be the weighting coefficient for the s-th scale. i s Let be the image feature embedding vector at the s-th scale. After calculation, the outputs for all scales are: z1, z2, ..., zi.

6. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 5, characterized in that, The algorithm for the feature fusion module is as follows: Where zi represents the Tranformer encoded feature at the i-th scale.

7. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 6, characterized in that, The algorithm for converting the global representation into probability distributions for each cell category is as follows: y=Softmax(W×z final +b) (4) Where W and b are the learnable weight matrix and bias vector of the classification layer, respectively, and y represents the model's predicted probability for red blood cells, lymphocytes, monocytes, neutrophils, eosinophils, and basophils.

8. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 7, characterized in that, The loss evaluation module calculates the training loss by minimizing the multi-class cross-entropy loss function. The algorithm for minimizing the multi-class cross-entropy loss function is as follows: Where N represents the number of samples, C represents the number of cell types, and y ic Indicates the real label, p ic The probability predicted by the model, where ∈ is a small constant introduced to prevent numerical instability, and is set to 1×10 by default. -6 .

9. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 8, characterized in that, The parameter optimization module includes the AdamW optimizer and the cosine annealing scheduling module. The initial learning rate is set to 1e-4, the batch size is 32, the maximum training time is 50 rounds, and training is stopped early when there is no improvement in the performance on the validation set. The stability of the trained model is evaluated by a five-fold cross-validation strategy.

10. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 1, characterized in that, The image size unification and image enhancement processing described in step S3 includes: image normalization, color space adjustment and noise reduction, and dividing each image into 500 to 2000 local regions.

11. The method for blood cell image recognition based on a multi-scale Vision Transformer model according to claim 1, characterized in that, The cell type identification in step S4 includes: cell location identification using a multi-scale ViT model, cell type identification, and DBSCAN clustering of cell locations to identify duplicate markers in dense regions. The neighborhood distance radius is set as follows: erythrocytes: 10, platelets: 8, neutrophils: 15, eosinophils: 13, monocytes: 16, lymphocytes:

12. Subsequently, a non-maximum suppression algorithm is used to remove overlapping candidate boxes. Based on the deduplicated cell identification results, the model calculates the number of each cell type and its proportion in the total number of cells. The IoU value for removing overlapping candidate boxes using the non-maximum suppression algorithm is set as follows: erythrocytes: 0.20, platelets: 0.15, neutrophils: 0.30, eosinophils: 0.30, monocytes: 0.35, lymphocytes: 0.

40.

12. A system for recognizing blood cells using the method for blood cell image recognition based on a multi-scale Vision Transformer model as described in any one of claims 1-11, characterized in that, The blood cell image recognition system includes the following modules: (1) Blood cell image dataset acquisition and data augmentation module; (2) Based on the multi-scale Vision Transformer pre-training module; (3) Data acquisition and data augmentation module for blood cell images to be detected; (4) Image recognition and counting module based on multi-scale Vision Transformer; (5) Image recognition and counting result output module.

Citation Information

Cited By

  • Automatic classification and counting method for blood cell microscopic images

    CN121259820A

  • Deep learning-based blood smear whole red blood cell abnormal morphology evaluation system

    CN121544632A

  • Optical visual layering system and method for blood separation

    CN121564242A