System for automatically distinguishing leukemia type based on flow cytometry

Through an intelligent diagnostic system based on flow cytometry, using multi-color laser flow cytometers and deep learning technology, leukemia cell data is automatically processed, solving the problems of low recognition accuracy and difficulty in detecting tiny residual lesions in existing technologies, and achieving efficient and accurate leukemia diagnosis.

CN120611262APending Publication Date: 2025-09-09周宏伟
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510562010.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies in leukemia diagnosis have problems such as low recognition accuracy, test results significantly affected by operator experience, limited quantitative analysis accuracy, and difficulty in detecting minimal residual lesions.

Method used

An intelligent diagnostic system based on flow cytometry is adopted, including a data acquisition module, a feature analysis module and a predictive diagnosis module. It uses technologies such as multi-color laser flow cytometer, autoencoder, tensor fusion network, gated neural network and three-dimensional convolutional network to automatically process cell data, generate high-resolution multimodal input, and combine knowledge graphs for pathological analysis to achieve automated diagnosis.

Benefits of technology

It significantly improves diagnostic consistency and sensitivity, shortens analysis time to within 20 minutes, achieves detection sensitivity at the 0.001% level, reduces missed detection rate by over 30%, and can automatically generate diagnostic recommendations that meet WHO standards, providing treatment plan support.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention belongs to the field of intelligent equipment, and particularly relates to a system for automatically distinguishing leukemia types based on flow cytometry. The invention aims to solve the problem of low leukemia identification accuracy in the prior art. The invention provides a system for automatically distinguishing leukemia types based on flow cytometry. The system comprises a data acquisition module, a feature analysis module and a prediction and diagnosis module, the data acquisition module is used for processing bone marrow puncture fluid of a to-be-detected person to obtain a cell event data set of the to-be-detected person; the feature analysis module is used for processing a cell event data set of a to-be-detected person to obtain extraction features; the prediction and diagnosis module is used for processing extracted feature prediction to obtain a diagnosis conclusion; according to the system, through a standardized data preprocessing process, the analysis duration can be compressed to be within 20 minutes, the diagnosis consistency is improved to 98% or above, and the detection sensitivity reaches the 0.001% level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent diagnosis, and in particular relates to a system for automatically distinguishing leukemia types based on flow cytometry. Background Art

[0002] With the deepening development of precision medicine and the deep integration of artificial intelligence technology in the biomedical field, the diagnostic model for hematological diseases is undergoing a revolutionary transformation. In current clinical practice, the diagnosis of leukemia still relies on bone marrow puncture, cell morphology observation, and the comprehensive judgment of experienced hematopathologists. This is plagued by pain points such as long testing cycles, high subjectivity, and the easy misdiagnosis of minimal residual lesions. Especially in primary care institutions, the shortage of high-level testing personnel and the surge in complex cases have created a sharp contradiction. The establishment of an intelligent and standardized diagnostic system has become a key breakthrough in the diagnosis and treatment of hematological diseases.

[0003] The intelligent leukemia diagnosis system, based on the deep integration of flow cytometry and deep learning technology, is an innovative solution to this clinical challenge. Through core technologies such as high-dimensional flow cytometry data acquisition, multi-parameter feature engineering construction, and intelligent clustering of cell populations, the system has achieved full automation of the entire process from raw data input to diagnostic report output. The system can simultaneously process dozens of immune phenotypic markers, use convolutional neural networks to intelligently cluster millions of cell events, and combine knowledge graphs to analyze the pathological significance of abnormal cell populations. In particular, for clinical difficulties such as acute leukemia typing and minimal residual lesion monitoring, the system continuously optimizes the diagnostic model through a transfer learning mechanism, and its multidimensional data analysis capabilities have surpassed the dimensional limits of traditional manual interpretation.

[0004] Even with flow cytometry, the current diagnostic system still suffers from drawbacks such as long data interpretation times (average 3-5 hours per case), significant influence of operator experience (consistency rates between different laboratories are approximately 75-85%), and limited quantitative analysis accuracy (especially for residual disease <0.01%). These limitations result in low accuracy in identifying leukemia. Summary of the Invention

[0005] The present invention aims to solve the problem of low accuracy in identifying leukemia in the prior art. A system for automatically distinguishing leukemia types based on flow cytometry is provided, comprising:

[0006] Data acquisition module, feature analysis module and prediction diagnosis module;

[0007] The data acquisition module is used to process the bone marrow puncture fluid of the person to be tested to obtain a cell event data set of the person to be tested;

[0008] The feature analysis module is used to process the cell event data set of the person to be tested to obtain extracted features;

[0009] The predictive diagnosis module is used to process the extracted features and predict the diagnosis conclusion;

[0010] The acquisition of bone marrow aspirate is well known to hematologists and oncologists. Since the present invention is a detection system, there is no need to describe the specific method used by medical personnel to obtain it. It is sufficient to describe the input of the detection system of the present invention.

[0011] The beneficial effects of the present invention are:

[0012] This system, through standardized data preprocessing, can reduce analysis time to under 20 minutes, improve diagnostic consistency to over 98%, and achieve a detection sensitivity of 0.001%. Clinical validation has shown that the system not only automatically generates diagnostic recommendations that meet WHO classification standards, but also intuitively presents cell evolution trajectories through a visual interface, providing dynamic monitoring support for treatment planning.

[0013] The system can stably detect minimal residual lesions (MRDs) as low as 0.001%, reducing missed detection rates by over 30% compared to traditional methods. These innovations give the system both high sensitivity and strong anti-interference capabilities.

[0014] The development of this intelligent diagnostic system marks a paradigm shift in hematologic malignancy diagnosis from experience-based to data-driven. By deploying a cloud-based collaborative platform, primary care hospitals can obtain diagnostic support comparable to that of a tertiary hospital simply by completing standardized sample preparation, effectively addressing the challenge of uneven distribution of medical resources. By integrating multi-omics data such as single-cell sequencing and digital pathology, this system will serve as a comprehensive intelligent diagnostic hub for hematologic diseases, providing a new technological infrastructure for the era of precision medicine. DETAILED DESCRIPTION

[0015] Specific implementation method 1: Combined with the description of the present invention,

[0016] Data acquisition module, feature analysis module and prediction diagnosis module;

[0017] The data acquisition module is used to process the bone marrow puncture fluid of the person to be tested to obtain a cell event data set of the person to be tested;

[0018] The feature analysis module is used to process the cell event data set of the person to be tested to obtain extracted features;

[0019] The predictive diagnosis module is used to process the extracted features and predict the diagnosis conclusion;

[0020] The acquisition of bone marrow puncture fluid is well known to hematologists and oncologists. The present invention is a detection system, and there is no need to introduce the specific method for medical personnel to obtain it. It is only necessary to introduce what the input of the detection system of the present invention is.

[0021] Specific embodiment 2: The difference between this embodiment and specific embodiment 1 is that:

[0022] The data acquisition module includes: an acquisition unit and a data fusion unit;

[0023] The acquisition unit is a multi-color laser flow cytometer;

[0024] The model requirements of the multi-color laser flow cytometer are: a flow cytometer that supports more than 10 fluorescent channels;

[0025] In flow cytometry, "color" refers to the number of fluorescence detection channels, specifically the types of fluorescent signals that can be detected simultaneously. Each fluorescent signal typically corresponds to a specific fluorescent dye, which emits fluorescence at different wavelengths when excited by laser light of a specific wavelength. Flow cytometry analyzes cell characteristics by detecting these fluorescent signals.

[0026] Here are some examples of fluorescent models, such as:

[0027] FITC (green fluorescence), PE (yellow fluorescence), PerCP (cyan fluorescence), APC (red fluorescence), APC-Cy7 (far-red fluorescence), PE-Cy7 (orange fluorescence), BV421 (blue fluorescence), BV510 (cyan fluorescence), BV605 (orange-red fluorescence), BV711 (deep red fluorescence)

[0028] If a flow cytometer has 10 fluorescence channels, it can simultaneously detect the fluorescence signals emitted by the above 10 fluorescent dyes. Each channel is equipped with a specific filter to separate and detect fluorescence signals in a specific wavelength range.

[0029] The present invention requires multi-parameter analysis based on multi-color fluorescence channels, which allow the simultaneous detection of multiple cell surface or intracellular markers, thereby performing more complex cell phenotypic analysis. Multiple cell surface markers can be detected simultaneously to distinguish different immune cell subpopulations.

[0030] Using a flow cytometer with 10 or more fluorescent channels can reduce the number of experiments and sample size, improving experimental efficiency. Simultaneous detection of multiple markers can reduce errors caused by fractionated detection.

[0031] The data fusion unit includes an autoencoder and a tensor fusion network in sequence;

[0032] Among them, the output of the autoencoder is the input of the tensor fusion network, and the tensor fusion network outputs the cell event dataset (fusion data) of the person to be tested

[0033] The model of the autoencoder is Deep Canonically-Correlated Auto-Encoder (DCCAE) and the model of the tensor fusion network is Tensor Fusion Network (TFN);

[0034] The structures and processing procedures of these two networks are well known to those skilled in the art.

[0035] The specific process of the data acquisition module processing the bone marrow puncture fluid of the person to be tested to obtain the cell event data set of the person to be tested is as follows:

[0036] S1: Use a multi-color laser flow cytometer to collect and process the bone marrow puncture fluid of the person to be tested, and obtain the physical data of the immune markers, the fluorescence data of the immune markers, and the dynamic data of the immune markers;

[0037] Deconvolution processing is performed on the fluorescence data of the immune marker using a deconvolution algorithm to obtain fluorescence data of the immune marker after deconvolution processing;

[0038] The output of a flow cytometer is an FCS file, which records the signal intensity of each cell across multiple fluorescence channels. These channels typically correspond to fluorescence detectors at different wavelengths on the flow cytometer, such as 488nm, 561nm, and 640nm. The signal intensity of each channel reflects the fluorescence intensity of the cell at that wavelength.

[0039] Fluorescence data is represented by a spectral matrix, a mathematical tool used to describe the emission spectrum characteristics of fluorescent dyes at different wavelengths. It is a two-dimensional matrix in which each row represents a fluorescent dye, each column represents a wavelength, and the matrix element value represents the emission intensity of the dye at the corresponding wavelength.

[0040] Processing the spectral matrix using a deconvolution algorithm can automatically correct for fluorescence signal overlap, which is a processing solution well known to those skilled in the art;

[0041] The physical data include: cell size and granularity;

[0042] Cell size refers to the physical dimensions of a cell, such as its volume or diameter, and is usually measured in micrometers (μm). Cell size is a fundamental physical characteristic of cells, and different cell types have different sizes.

[0043] Cell granularity refers to the amount and distribution of granular substances inside cells, reflecting the complexity of cells or the characteristics of their internal structure.

[0044] The immune markers include immune markers of different immune cells such as CD34 and CD45, which are known to those skilled in the art.

[0045] CD34 and CD45 are two common immune cell surface markers. In addition to CD34 and CD45, there are many other markers used to identify different types of immune cells. The following are some common markers and their corresponding cell types:

[0046] These markers are widely used in techniques such as flow cytometry and immunohistochemistry and are well known in immunology and hematology research;

[0047] S2: The physical data of the immune marker, the fluorescence data of the immune marker after deconvolution processing, and the dynamic data of the immune marker are input into the data fusion unit for processing to obtain a cell event data set of the person to be tested;

[0048] The present invention uses a data fusion unit to map the physical, chemical and dynamic signals of cells into the same feature space through correlation alignment and high-order tensor decomposition, significantly improving the separability of weak abnormal phenotypes and providing high-resolution multimodal input submission for downstream convolutional networks.

[0049] Other steps and parameters are the same as those in the first embodiment.

[0050] Specific embodiment three: This embodiment differs from specific embodiment one in that:

[0051] The feature analysis module includes: a gated neural network and a three-dimensional convolutional network;

[0052] The feature analysis module processes the cell event data set of the person to be tested to obtain extracted features; the specific process is:

[0053] A1: Input the cell event dataset of the person to be tested into the gated neural network for processing; obtain the first feature;

[0054] A2: The first feature and the cell event dataset of the person to be tested are input into a three-dimensional convolutional network for processing to obtain the second feature X, which is used as the extracted feature.

[0055] The second feature is a cross-marker combination feature (such as the CD13+CD33+CD117+ phenotype specific for acute myeloid leukemia). In flow cytometry, each cell can be detected by the expression level of multiple different immune markers.

[0056] The cross-marker combination features in the present invention refer to the joint expression features extracted by the feature parsing module (gated neural network and three-dimensional convolutional network) based on the original immune marker channel data collected by multi-color flow cytometry, which can effectively characterize the complex interactive relationship between immune phenotypes.

[0057] CD13, CD33, and CD117 are all immune markers.

[0058] Cross-marker combined features are not directly equivalent to raw data. Instead, they are high-order combined features extracted from the raw data of multiple immune markers through neural network processing (such as gated weighting and convolutional feature extraction).

[0059] The first module collects raw fluorescence data of individual markers;

[0060] The second module (gating + 3D convolution) processes these raw data interactively and extracts the combined features across markers.

[0061] The other steps and parameters are the same as those in the first and second embodiments.

[0062] Specific embodiment 4: This embodiment differs from specific embodiments 1 to 4 in that:

[0063] The gated neural network in A1 includes: a gated network and a shallow fully connected network;

[0064] The cell event data set of the person to be tested is input into the gated neural network for processing; the first feature is obtained; the specific process is:

[0065] A1.1: First, represent the cell event dataset of the person to be tested as a multidimensional feature matrix;

[0066] A1.2: Input the multidimensional feature matrix obtained in A1.1 into the gated network for processing to obtain the importance weights of each channel feature. Use the sigmoid function to generate the importance weights.

[0067] A1.3: Weight the original feature matrix according to the importance weights to obtain the weighted feature matrix;

[0068] A1.4: Input the weighted feature matrix into the shallow fully connected network for processing to obtain the first feature;

[0069] The other steps and parameters are the same as those in the first to third embodiments.

[0070] Specific embodiment 5: This embodiment differs from specific embodiments 1 to 4 in that:

[0071] The A2 three-dimensional convolutional network includes, in sequence: a first convolutional layer, a second convolutional layer, and a third convolutional layer;

[0072] The convolution kernel sizes of the first convolution layer, the second convolution layer, and the third convolution layer are 3×3×3;

[0073] The specific process of inputting the first feature and the cell event dataset of the person to be tested into the three-dimensional convolutional network for processing to obtain the second feature X is as follows:

[0074] A2.1: Concatenate the first feature with the cell event dataset of the person to be tested in the channel dimension to form a fused feature matrix, where the number of channels after fusion is the sum of the number of original channels and the number of channels of the first feature;

[0075] A2.2, input the fused feature matrix into the 3D convolutional network for feature extraction to obtain the second feature X;

[0076] The three-dimensional convolutional network includes a first convolutional layer, a second convolutional layer and a third convolutional layer in sequence, and the convolution kernel size of each convolutional layer is 3×3×3;

[0077] After each convolutional layer, a ReLU activation function is applied to introduce nonlinear expression capabilities, and feature normalization operations are performed as needed to improve training stability and feature expression consistency. Through three-dimensional convolution operations, cell spatial structure features and cross-marker combination features are extracted from the fusion feature matrix, and the second feature X is output for subsequent prediction and diagnosis module processing.

[0078] The other steps and parameters are the same as those in the first to fourth embodiments.

[0079] Specific embodiment 6: This embodiment differs from specific embodiments 1 to 5 in that:

[0080] The prediction and diagnosis module is used to process the extracted features and predict the diagnosis of the person to be tested; the specific process is as follows:

[0081] B1: Perform PCA preprocessing on the second feature X to obtain the PCA preprocessed data;

[0082] The PCA preprocessing process retains 95% cumulative variance,

[0083] The PCA preprocessing, i.e., principal component analysis (PCA),

[0084] "Retaining 95% cumulative variance" means selecting a portion of principal components so that the sum of their variances can explain 95% of the information in the original data (i.e., the total variance). In other words, PCA reduces the dimensionality of the original data while retaining as much information as possible.

[0085] Cumulative variance refers to the sum of the variances of the first few principal components, which is used to indicate how much "information" the first few principal components retain in the original data. Suppose there is a data set, and after PCA dimensionality reduction, 5 principal components are obtained. Their corresponding variances (eigenvalues) are: Variance of principal component 1: 50% Variance of principal component 2: 20% Variance of principal component 3: 15% Variance of principal component 4: 10% Variance of principal component 5: 5%

[0086] If we select the first two principal components, their cumulative variance is 50% + 20% = 70%, which clearly does not reach 95%. If we select the first three principal components, the cumulative variance is 50% + 20% + 15% = 85%, which is still less than 95%. If we select the first four principal components, the cumulative variance is 50% + 20% + 15% + 10% = 95%, which just retains 95% of the variance. This is well known in the art.

[0087] Discarding the remaining 5% of the variance usually does not significantly affect the performance of the model, which can greatly reduce the computational complexity and memory consumption.

[0088] Then perform UMAP dimensionality reduction on the data after PCA processing to obtain the feature X1 after UMAP dimensionality reduction processing;

[0089] UMAP (Uniform Manifold Approximation and Projection) is a commonly used nonlinear dimensionality reduction technique, primarily used for data visualization and high-dimensional data dimensionality reduction. Based on the principles of topology and manifold learning, it has a better ability to preserve local structure and is computationally more efficient than other dimensionality reduction techniques (such as PCA and t-SNE). The following is the UMAP dimensionality reduction process, which is well known in the art.

[0090] B2: Use the HNSW algorithm to process the feature X1 after UMAP dimensionality reduction to obtain the approximate k-NN index.

[0091] The Hierarchical Navigable Small World (HNSW) algorithm is an efficient approximate nearest neighbor search (ANN) algorithm used to find the closest point to a given query point in a high-dimensional space. It is a graph-based search algorithm and a highly efficient approach to approximate search.

[0092] k-NN (k-Nearest Neighbors) is a common supervised learning algorithm widely used in classification and regression problems. Its core idea is to calculate the distance between an unknown sample and all samples in the training set and select the k closest samples. The present invention uses the HNSW algorithm to obtain an approximate k-NN index, i.e., the k closest cells to each cell.

[0093] B3: Construct the local adjacency graph A based on the obtained approximate k-NN index. The specific process is as follows:

[0094] B4: Calculate the local reachability density (LRD) of each cell in the local adjacency graph A;

[0095] Local Reachability Density (LRD) is a method used to measure the "density" of data points in cluster analysis. LRD is commonly used in local outlier detection, especially in the LOF (Local Outlier Factor) algorithm. It measures the density of a data point in its neighborhood, taking into account the meaning of the reachable distance LRD of the point. A high LRD value indicates that the data point is in a high-density area, which is usually a normal point in the data set. A low LRD value indicates that the data point is in a low-density area and may be an anomaly or outlier. Calculating the local reachability density LRD is known to those in the field;

[0096] The exclusive neighborhood radius of each cell is adaptively calculated based on the local reachability density LRD of each cell in the local adjacency graph A;

[0097] The exclusive neighborhood radius of the i-th cell is denoted as εi;

[0098] Adaptively adjust the size of the neighbor range around each cell based on its local density (LRD) in the local adjacency graph A. A small radius in dense areas (with many cells and large LRD) can define enough neighbors.

[0099] In sparse areas (few cells and small LRD), the radius is automatically expanded to capture enough neighbors.

[0100] B5: Based on the local reachability density LRD of each cell in the local adjacency graph A and the exclusive neighborhood radius εi of each cell, the local adjacency graph A is processed using the density peak seeding algorithm and the variable radius depth-first search algorithm to obtain N1 initial clusters;

[0101] B6: Merge the N1 initial clusters to obtain M merged clusters;

[0102] B7: Calculate the cosine similarity value O between the merged clusters and the known leukemia clone template; mark the clusters that meet the marking conditions in the merged clusters as minimal residual disease candidates;

[0103] B8: Diagnosis based on minimal residual disease candidate;

[0104] ; Other steps and parameters are the same as those in one of the specific implementation methods one to five.

[0105] Specific embodiment 7: This embodiment differs from specific embodiments 1 to 6 in that:

[0106] In B3, the local adjacency graph A is constructed based on the obtained approximate k-NN index. The specific process is as follows:

[0107] B3.1: Take the cell data in feature X1 as a node set, where each node represents a cell data.

[0108] B3.2: Based on the approximate k-NN index, connect adjacent cells to form an edge set, where each edge represents the connection between two adjacent nodes.

[0109] B3.3: Construct a local adjacency graph A based on a set of nodes and a set of edges;

[0110] The approximate k-NN index of the present invention represents the K nearest nodes of each node, and each node is connected to the K nearest nodes to obtain K edges;

[0111] Calculate the local reachability density LRD of each cell in the local adjacency graph A;

[0112] Local Reachability Density (LRD) is a method used to measure the "density" of data points in cluster analysis. LRD is commonly used in local outlier detection, especially in the LOF (Local Outlier Factor) algorithm. It measures the density of a data point in its neighborhood, taking into account the meaning of the reachable distance LRD of the point. A high LRD value indicates that the data point is in a high-density area, which is usually a normal point in the data set. A low LRD value indicates that the data point is in a low-density area and may be an anomaly or outlier. Calculating the local reachability density LRD is known to those in the field;

[0113] The other steps and parameters are the same as those in the first to sixth embodiments.

[0114] Specific embodiment eight: This embodiment differs from specific embodiments one to seven in that:

[0115] In B5, based on the local reachable density LRD of each cell in the local adjacency graph A and the exclusive neighborhood radius εi of each cell, the local adjacency graph A is processed using the density peak seeding algorithm and the variable radius depth-first search algorithm to obtain N1 initial clusters. The specific process is as follows:

[0116] B5.1: Use the density peak seeding algorithm to process the local adjacency graph A and obtain N1 cluster seeds;

[0117] Find nodes with relatively high local density (that is, local density peaks) in the local adjacency graph A as cluster seeds;

[0118] B5.2: Based on the N1 cluster seeds and the exclusive neighborhood radius of each cell in the local adjacency graph A, use a variable-radius depth-first search algorithm to process the local adjacency graph A to obtain N1 initial clusters;

[0119] Interpretation of initial clusters: A density peak seeding algorithm and a variable radius depth-first search algorithm are used to adaptively generate multiple initial clusters in the local adjacency graph A. Each initial cluster corresponds to a local density peak cell and a collection of cells in its neighborhood, forming a preliminary cell clustering result.

[0120] These two algorithms are used to complete the following three things based on the local adjacency graph A + the exclusive neighborhood radius εi of each cell:

[0121] Find points with relatively high local density (that is, local density peaks) as clustering seeds; starting from the seeds, search for neighboring nodes that can belong to the same cluster according to each cell's own neighborhood radius εi; and group the entire group of searched cell nodes into an initial cluster.

[0122] The other steps and parameters are the same as those in the first to seventh embodiments.

[0123] Specific embodiment 9: This embodiment differs from specific embodiments 1 to 8 in that:

[0124] In B6, N1 initial clusters are merged to obtain M merged clusters. The specific process is as follows:

[0125] B6.1: Calculate the multi-scale cluster persistence for each initial cluster;

[0126] B6.2: Calculate the cosine similarity between the N1 initial clusters;

[0127] The degree of stability of a cluster at different scales (different cluster density thresholds, neighborhood radius, similarity requirements, etc.) to determine whether it can continue to exist without obvious collapse or fusion.

[0128] It is necessary to first calculate the persistence value of each initial cluster; observe whether the cluster still exists at each scale (such as different ε radius, or different similarity thresholds), and count the number of scales at which a cluster is still detected as a complete cluster for calculation.

[0129] For each initial cluster, multiple scale parameter change conditions (including changes in neighborhood radius and similarity threshold) are set, and cluster stability tests are performed at each scale. The proportion of clusters that remain intact under different scale conditions is statistically analyzed, which is defined as the multi-scale cluster persistence. The process of calculating the multi-scale cluster persistence for each initial cluster is well known to those skilled in the art. The multi-scale cluster persistence is used to measure the stability of the cluster and serves as one of the criteria for subsequent merging of fragmented clusters.

[0130] B6.3: Merge the N1 initial clusters based on the calculated multi-scale cluster persistence of each initial cluster and the cosine similarity between the N1 initial clusters to obtain M merged clusters;

[0131] Among them, the multi-scale cluster persistence and cosine similarity (similarity between feature vectors or centroids of different initial clusters) thresholds are used. Automatically merge the fragmented clusters (clusters in the initial cluster that are very small, few in number, and have poor stability) to obtain M merged clusters;

[0132] In simple terms, multi-scale cluster persistence calculates the state of the initial cluster itself; cosine similarity calculates the state between different initial clusters; based on the two states, clusters with very small volume, small number, and poor stability in the initial cluster are merged. The thresholds of volume, number, and stability are set by the experimenter.

[0133] The other steps and parameters are the same as those in the first to eighth embodiments.

[0134] Specific embodiment 10: This embodiment differs from specific embodiments 1 to 9 in that:

[0135] The marking condition in B7 is expressed as follows:

[0136] S≤0.001×N_total

[0137] O≥0.80

[0138] Where S represents the cluster size; N_total represents the total number of cells; O represents the cosine similarity value between the merged cluster and the known leukemia clone template;

[0139] Clusters with a cosine similarity of ≥0.80 to known leukemia clone templates were marked as minimal residual disease candidates;

[0140] Calculate the cosine similarity between each merged cluster and the leukemia template. Select clusters that meet the size + similarity criteria and label them as minimal residual disease candidate clusters. Minimal residual disease candidate clusters are not leukemia types per se. They represent a small population of cells whose characteristic expression patterns are highly similar to a leukemia clone template. Lesion candidate clusters suggest the presence of minimal residual leukemia, but the specific leukemia type is not directly determined by the cluster itself.

[0141] The other steps and parameters are the same as those in the first to ninth embodiments.

[0142] Combined with the simulation analysis of specific implementation methods one to ten

[0143] The above only describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific implementation methods. Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent replacements and improvements made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A system for automatically distinguishing leukemia types based on flow cytometry, characterized in that: include: Data acquisition module, feature analysis module and prediction diagnosis module; The data acquisition module is used to process the bone marrow puncture fluid of the person to be tested to obtain a cell event data set of the person to be tested; The feature analysis module is used to process the cell event data set of the person to be tested to obtain extracted features; The predictive diagnosis module is used to process the extracted features and predict to obtain a diagnosis conclusion.

2. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 1, characterized in that: The data acquisition module includes: an acquisition unit and a data fusion unit; The acquisition unit is a multi-color laser flow cytometer; The data fusion unit includes an autoencoder and a tensor fusion network in sequence; The specific process of the data acquisition module processing the bone marrow puncture fluid of the person to be tested to obtain the cell event data set of the person to be tested is as follows: S1: Use a multi-color laser flow cytometer to collect and process the bone marrow puncture fluid of the person to be tested, and obtain the physical data of the immune markers, the fluorescence data of the immune markers, and the dynamic data of the immune markers; Deconvolution processing is performed on the fluorescence data of the immune marker using a deconvolution algorithm to obtain fluorescence data of the immune marker after deconvolution processing; S2: The physical data of the immune marker, the fluorescence data of the immune marker after deconvolution processing, and the dynamic data of the immune marker are input into the data fusion unit for processing to obtain a cell event data set of the person to be tested.

3. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 2, characterized in that: The feature analysis module includes: a gated neural network and a three-dimensional convolutional network; The feature analysis module processes the cell event data set of the person to be tested to obtain extracted features; the specific process is: A1: Input the cell event dataset of the person to be tested into the gated neural network for processing; obtain the first feature; A2: The first feature and the cell event dataset of the person to be tested are input into a three-dimensional convolutional network for processing to obtain the second feature X, which is used as the extracted feature.

4. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 3, characterized in that: The gated neural network in A1 includes: a gated network and a shallow fully connected network; The cell event data set of the person to be tested is input into the gated neural network for processing; the first feature is obtained; the specific process is: A1.1: First, represent the cell event dataset of the person to be tested as a multidimensional feature matrix; A1.2: Input the multidimensional feature matrix obtained in A1.1 into the gating network to obtain the importance weights; A1.3: Weight the original feature matrix according to the importance weights to obtain the weighted feature matrix; A1.4: Input the weighted feature matrix into a shallow fully connected network to obtain the first feature.

5. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 4, characterized in that: The three-dimensional convolutional network in A2 includes, in sequence: a first convolutional layer, a first ReLU activation function layer, a second convolutional layer ReLU activation function layer, a third convolutional layer, and a third ReLU activation function layer; The convolution kernel sizes of the first convolution layer, the second convolution layer, and the third convolution layer are all 3×3×3; The specific process of inputting the first feature and the cell event dataset of the person to be tested into the three-dimensional convolutional network for processing to obtain the second feature X is as follows: A2.1: Combine the first feature with the cell event dataset of the person to be tested in the channel dimension to form a fusion feature matrix. A2.2, input the fused feature matrix into the 3D convolutional network for feature extraction to obtain the second feature X.

6. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 5, characterized in that: The prediction and diagnosis module is used to process the extracted features and predict the diagnosis of the person to be tested; the specific process is as follows: B1: Perform PCA preprocessing on the second feature X to obtain the PCA preprocessed data; The PCA preprocessing process retains 95% cumulative variance, Then perform UMAP dimensionality reduction on the data after PCA processing to obtain the feature X1 after UMAP dimensionality reduction processing; B2: Use the HNSW algorithm to process the feature X1 after UMAP dimensionality reduction to obtain the approximate k-NN index. B3: Construct the local adjacency graph A based on the obtained approximate k-NN index. The specific process is as follows: B4: Calculate the local reachability density (LRD) of each cell in the local adjacency graph A; The exclusive neighborhood radius of each cell is calculated based on the local reachability density LRD of each cell in the local adjacency graph A; B5: Based on the local reachability density LRD of each cell in the local adjacency graph A and the exclusive neighborhood radius of each cell, the density peak seeding algorithm and the depth-first search algorithm with a variable radius are used to process the local adjacency graph A to obtain N1 initial clusters; B6: Merge the N1 initial clusters to obtain M merged clusters; B7: Calculate the cosine similarity value O between the merged cluster and the known leukemia clone template; The clusters that meet the marking conditions in the merged clusters are marked as minimal residual disease candidates; B8: Diagnosis based on minimal residual disease candidate.

7. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 6, characterized in that: In B3, the local adjacency graph A is constructed based on the obtained approximate k-NN index. The specific process is as follows: B3.1: Take the cell data in feature X1 as a node set, where each node represents a cell data. B3.2: Based on the approximate k-NN index, connect adjacent cells to form an edge set, where each edge represents the connection between two adjacent nodes. B3.3: Construct a local adjacency graph A based on a set of nodes and a set of edges; Calculate the local reachability density LRD of each cell in the local adjacency graph A.

8. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 7, characterized in that: In B5, based on the local reachable density LRD of each cell in the local adjacency graph A and the exclusive neighborhood radius of each cell, the local adjacency graph A is processed using a density peak seeding algorithm and a depth-first search algorithm with a variable radius to obtain N1 initial clusters. The specific process is as follows: B5.1: Use the density peak seeding algorithm to process the local adjacency graph A and obtain N1 cluster seeds; B5.2: Based on the N1 cluster seeds and the exclusive neighborhood radius of each cell in the local adjacency graph A, a variable-radius depth-first search algorithm is used to process the local adjacency graph A to obtain N1 initial clusters.

9. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 8, characterized in that: In B6, N1 initial clusters are merged to obtain M merged clusters. The specific process is as follows: B6.1: Calculate the multi-scale cluster persistence for each initial cluster; B6.2: Calculate the cosine similarity between the N1 initial clusters; B6.3: Merge the N1 initial clusters based on the calculated multi-scale cluster persistence of each initial cluster and the cosine similarity between the N1 initial clusters to obtain M merged clusters.

10. The system for automatically distinguishing leukemia types based on flow cytometry according to claim 9, characterized in that: The marking condition in B7 is expressed as follows: S≤0.001×N_total O≥0.80 Where S represents the cluster size; N_total represents the total number of cells; and O represents the cosine similarity value between the merged cluster and the known leukemia clone template.

Citation Information

Cited By

  • Disease diagnosis method and device based on expert knowledge optimization, equipment and medium

    CN121260426A