Method and system of processing cytometric data from multiple tests

The method employs supervised machine learning to classify flow cytometry data, addressing the challenges of labor-intensive and inconsistent analysis by enhancing reproducibility and inter-laboratory consistency across different panels and instruments.

WO2025122805A1PCT designated stage expired Publication Date: 2025-06-12AHEAD MEDICINE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058765
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current flow cytometry analysis methods are labor-intensive, lack consistency, and require skilled professionals, leading to interpretational variations and challenges in ensuring laboratory compliance with standardized guidelines.

Method used

A method and system using supervised machine learning to automatically classify flow cytometry data from multiple tests by determining overlapping parameters, preprocessing them, and developing classifiers, enabling reproducible analysis across different panels and instruments.

Benefits of technology

The approach significantly improves the reproducibility of flow cytometry analyses by enabling consistent interpretation and enumeration of cytometric data from various sources, reducing the need for manual expert analysis and enhancing inter-laboratory consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000013_0001
    Figure IMGF000013_0001
  • Figure IMGF000014_0001
    Figure IMGF000014_0001
  • Figure IMGF000016_0001
    Figure IMGF000016_0001
Patent Text Reader

Abstract

A method of processing cytometric data from multiple tests is provided. The method of processing cytometric data from multiple tests includes the steps of: determining overlapping parameters among the multiple tests using flow cytometry; preprocessing the overlapping parameters; and developing classifiers via supervised machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM OF PROCESSING CYTOMETRIC DATA FROM MULTIPLE TESTS CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims a benefit and priority to U.S. Provisional Patent Application No. 63 / 607,305 filed on Dec. 07, 2023, the entire contents of which are hereby incorporated by reference. BACKGROUND OF THE INVENTION 1. Field of the Invention

[0002] The present disclosure relates to a method of processing cytometric data, and in particular to a method of processing cytometric data from multiple tests. 2. Description of the Related Art

[0003] Flow cytometry (FC) is essential in diagnosing and monitoring hematological diseases. To make flow cytometry work, samples of blood, bone marrow or tissue cells are placed in a suspension and injected into the flow cytometer machine. The cells are arranged in a single file line, and then passed in front of a laser beam, scattered light and fluorescent light. Next, the cells are counted and categorized. The data is stored in a computer and reported via a histogram or dot plot for further analysis.

[0004] Current analytical methods and software predominantly depend on expert analysts who manually inspect and interpret a series of 2-dimensional plots through sophisticated, sequential gating processes. This labor-intensive method presents challenges due to a shortage of skilled professionals and the growing demand for rapid, precise, and reproducible FC analysis results. Also, current manual and subjective analysis of clinical FC data often lacks consistency, leading to interpretational variations. Furthermore, since each medical device manufacturer develops its specific kit for a specific use, data processing, analysis, and results reporting must be customized for each kit for each use.

[0005] Numerous efforts have been made to harmonize antibody panels, sample processing, analysis, and results reporting for various disease diagnoses. However, the main challenge lies in ensuring that laboratories fully comply with these guidelines.

[0006] The concept of standardization exists on a spectrum that can evolve over time and across different laboratories. Harmonization serves as a practical initial step, which may eventually lead to the development of a more comprehensive set of standardized methods.

[0007] In the field of flow cytometry, there continues to be significant variation among different centers, despite the consensus that interlaboratory reproducibility of flow cytometry measurements and studies could be enhanced through the adoption of standardized methodologies. Numerous efforts have been made to harmonize antibody panels, sample processing, analysis, and results reporting for various disease diagnoses.

[0008] In order to minimize heterogeneity and improve the precision of clinical flow cytometry testing, numerous guidelines have been created and published by medical societies, as shown in the citations below.

[0009] Borowitz MJ, Craig FE, Digiuseppe JA, Illingworth AJ, Rosse W, Sutherland DR, Wittwer CT, Richards SJ; Clinical Cytometry Society. Guidelines for the diagnosis and monitoring of paroxysmal nocturnal hemoglobinuria and related disorders by flow cytometry. Cytometry B Clin Cytom.2010 Jul;78(4):211-30. doi: 10.1002 / cyto.b.20525. PMID: 20533382.

[0010] Davis BH, Holden JT, Bene MC, Borowitz MJ, Braylan RC, Cornfield D, Gorczyca W, Lee R, Maiese R, Orfao A, Wells D, Wood BL, Stetler-Stevenson M. 2006 Bethesda International Consensus recommendations on the flow cytometric immunophenotypic analysis of hematolymphoid neoplasia: medical indications. Cytometry B Clin Cytom. 2007;72 Suppl 1:S5-13. doi: 10.1002 / cyto.b.20365. PMID: 17803188.

[0011] Borowitz MJ, Wood BL, Keeney M, Hedley BD. Measurable Residual Disease Detection in B-Acute Lymphoblastic Leukemia: The Children's Oncology Group (COG) Method. Curr Protoc.2022 Mar;2(3):e383. doi: 10.1002 / cpz1.383. PMID: 35263042.

[0012] Van Dongen JJ, Lhermitte L, Böttcher S, Almeida J, van der Velden VH, Flores- Montero J, Rawstron A, Asnafi V, Lécrevisse Q, Lucio P, Mejstrikova E, Szczepański T, Kalina T, de Tute R, Brüggemann M, Sedek L, Cullen M, Langerak AW, Mendonça A, Macintyre E, Martin-Ayuso M, Hrusak O, Vidriales MB, Orfao A; EuroFlow Consortium (EU-FP6, LSHB- CT-2006-018708). EuroFlow antibody panels for standardized n-dimensional flow cytometric immunophenotyping of normal, reactive and malignant leukocytes. Leukemia. 2012 Sep;26(9):1908-75. doi: 10.1038 / leu.2012.120. Epub 2012 May 3. PMID: 22552007; PMCID: PMC3437410.).

[0013] These guidelines aim to establish standardized practices for antibody panels, sample processing, analysis, and results reporting. BRIEF SUMMARY OF THE INVENTION

[0014] However, the main challenge lies in ensuring that laboratories fully comply with theseguidelines.

[0015] While the goal of standardization is to achieve reproducibility, harmonization can provide consistent interpretation, enumeration, and even patterns in certain cases. These concepts exist on a spectrum that can evolve over time and across different laboratories. Harmonization serves as a practical initial step, which may eventually lead to the development of a more comprehensive set of standardized methods.

[0016] In our research, we have shown that supervised machine learning (ML) approaches can effectively detect immunophenotype abnormalities and identify different diseases at the specimen level using flow cytometry data measured with the same panel and acquired on the same or two instrument models from the same manufacturer.

[0017] Accordingly, in order to improve the reproducibility of analyses on inter-laboratory data acquired on different panels but with harmonized parameter subsets. This disclosure encompasses exemplary embodiments related to methods and associated system for automatically classifying flow cytometry data. The data is measured with various reagent panels and acquired by different flow instrument models, all for the same purpose, while utilizing common markers measured across reagent panels.

[0018] It is an objective of the present disclosure to provide a method of processing cytometric data from multiple tests, comprising: (a) determining overlapping parameters among the multiple tests using flow cytometry; (b) preprocessing the overlapping parameters; and (c) developing classifiers via supervised machine learning.

[0019] In some embodiments, the determining of the overlapping parameters is performed by recognizing the overlapping parameters in different tests that describe same biological indicators.

[0020] In some embodiments, the preprocessing of the overlapping parameters comprises rearranging the overlapping parameters and performing normalization.

[0021] In some embodiments, rearranging the overlapping parameters comprises performing parameter alignment.

[0022] In some embodiments, the normalization comprises z-score normalization, min-max normalization, quantile normalization, or rank normalization, and can be applied either by site or by sample.

[0023] In some embodiments, the preprocessing of the overlapping parameters further comprises capturing distribution and embedding phenotype characteristics into a phenotype representation.

[0024] In some embodiments, the distribution is a high-dimensional cellular distribution.

[0025] In some embodiments, the high-dimensional cellular distribution is performed by a Gaussian Mixture Model (GMM), Principal Component Analysis (PCA), or deep neural network approaches such as Convolutional Neural Networks (CNN) and Autoencoders.

[0026] In some embodiments, the representation is a high-dimensional phenotype representation.

[0027] In some embodiments, the embedding phenotype characteristics into the distribution is performed by Fisher vectorization to embed phenotype characteristics into a high-dimensional phenotype representation vector at a specimen level.

[0028] In some embodiments, the preprocessing of the overlapping parameters further comprises inputting vectors along with corresponding diagnoses into a support vector machine.

[0029] In some embodiments, the developing of the classifiers comprises data training and sample classification.

[0030] In some embodiments, the method is used to count cells, to sort cells, to determine cell function, to determine cell characteristics, to detect microorganisms such as bacteria, fungus or yeast, to find biomarkers, or to assist diagnosis and potential treatment of blood and bone marrow cancers.

[0031] In some embodiments, the method further comprises performing visualization after the developing of the classifiers.

[0032] In some embodiments, the overlapping parameters are selected from a group comprising light scatter property and fluorescent marker.

[0033] It is another objective of the present disclosure to provide a system of processing cytometric data from multiple tests, which is suitable for signally connecting with one or more cytometric data providing devices in order to receive the cytometric data from the cytometric data providing devices, the system comprising: a storage module; and an automated classification module, configured to be signally connected with the storage module; wherein a plurality of codes is stored in the storage module, and wherein the automated classification module performs the steps of the method of processing cytometric data from multiple tests after the automated classification module executes the plurality of codes stored in the storage module. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] FIG.1 illustrates a schematic diagram showing a system of processing cytometric data from multiple tests according to an embodiment of the present disclosure.

[0035] FIG.2 is a flowchart illustrating the method of processing cytometric data from multiple tests according to an embodiment of the present disclosure.

[0036] FIG.3 is a diagram showing the experiment results of performance of single parameter feature selection according to an embodiment of the present disclosure

[0037] FIG.4 shows an example scheme diagram of developing and visualizing panel agnostic sample classification using flow cytometry data according to an embodiment of the present disclosure.

[0038] FIG 5. illustrated an exemplary framework for developing and visualizing panel agnostic sample classification using flow cytometry data in detail according to an embodiment of the present disclosure.

[0039] FIG.6A and 6B show the visualization results of samples in an exemplary cross-panel AML versus non-neoplastic classification according to an embodiment of the present disclosure.

[0040] FIG 7. shows the exemplary results of identifying minimal parameters required for developing panel agnostic sample classification using flow cytometry data according to an embodiment of the present disclosure.

[0041] FIG.8 illustrates an example framework for developing and visualizing panel agnostic sample classification using flow cytometry data according to an embodiment of the present disclosure.

[0042] FIG. 9A and 9B show the visualization result of samples in an example cross-panel AML versus non-neoplastic classification according to an embodiment of the present disclosure.

[0043] FIG. 10 illustrates an example framework for developing and visualizing Panel Agnostic Sample Classification Using Flow Cytometry Data according to an embodiment of the present disclosure.

[0044] FIG.11 shows the visualization result of samples in an example cross-panel Lymphoid abnormality versus no abnormality classification according to an embodiment of the present disclosure.

[0045] FIG. 12 is a block diagram that illustrates an example of an artificial intelligent (AI) system in which at least some operations described herein can be implemented according to an embodiment of the present disclosure.

[0046] FIG. 13 shows an example scheme diagram of z-score normalization according to an embodiment of the present disclosure.

[0047] FIG. 14 shows an example scheme diagram of density contour according to an embodiment of the present disclosure.

[0048] FIG.15 shows an example scheme diagram of classifier according to an embodiment of the present disclosure.

[0049] FIG. 16 shows an example scheme diagram of evaluation metrics according to an embodiment of the present disclosure.

[0050] FIG. 17 shows an example scheme diagram of evaluation metrics according to an embodiment of the present disclosure.

[0051] FIG.18 shows an example scheme diagram of GMM and Fisher vectorization according to an embodiment of the present disclosure.

[0052] FIG. 19 shows an example scheme diagram of cross validation according to an embodiment of the present disclosure.

[0053] FIG. 20 shows an example scheme diagram of compensation according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0054] To facilitate understanding of the objects, characteristics and effects of the present disclosure, embodiments together with the attached drawings for the detailed description of the present disclosure are provided.

[0055] It should be noted that the steps described herein may be performed sequentially, in reverse order, or by appropriately changing or skipping the order during the control process. It should be noted that the phrase “the first step may be performed after the second step” described in the present disclosure can be expressed as “the first step is followed directly after the second step” and / or “the second step is followed by the other steps (e.g., the third step) and then the first step”.

[0056] In addition, in the context of the present disclosure, it should be noted that terms such as “first”, “second” and “third” are used to distinguish differences between elements, and not to limit the elements themselves or to represent a particular order of elements. It should be noted that the same element or step may be indicated by the same reference numeral in the following description.

[0057] In addition, the term “coupling” as described in the present disclosure may be represented as “directly connected” and / or “indirectly connected”. Specifically, “the first element is configured to be coupled to the second element” can be expressed as “the first element is configured to be directly connected to the second element” and / or “the first element is configured to be indirectly connected to the second element”. Also, the term “signally connect”as described in the present disclosure may be represented as transmitting signal in any means.

[0058] The present disclosure provides a methodology capable of creating an automated classification model across different panels using commonly measured parameters designed for identical diagnostic or monitoring purposes. The present disclosure could significantly improve the reproducibility of analyses on inter-laboratory data acquired on different panels but with harmonized parameter subsets.

[0059] The flow cytometry data used in the disclosure were from the flow cytometry done as part of the clinical evaluation on bone marrow specimens of patients that have been evaluated for potential hematological diseases from different institutions.

[0060] Cross-tests data or data from multiple tests refers to data from two or more tests. In particular, cross-tests data or data from multiple tests refers to data from two or more tests which are processed via different models of flow cytometry instruments, different panels, different kits, different processing methods, different preprocessing methods, or two or more tests having any differences among which.

[0061] FIG.1 illustrates a schematic diagram showing a system of processing cytometric data from multiple tests according to an embodiment of the present disclosure.

[0062] In some embodiments, the system 100 of processing cytometric data from multiple tests is suitable for signally connecting with one or more cytometric data providing devices in order to receive the cytometric data from the cytometric data providing devices. In some embodiments, the system 100 may be an electronic device such as computers, server computers, client computers, personal computers, desktop computers, notebook computers, tablets, workstations, servers, cloud servers, smartphones, personal digital assistants, and / or computing device. The system 100 includes automated classification module 110, storage module 120, I / O interface 130 and communication interface 140.

[0063] In some embodiments, the automated classification module 110 is configured to be signally connected with the storage module 120, in some embodiments, also the I / O interface 130 and the communication interface 140. In some embodiments, the automated classification module 110 may be processor, CPU, GPU or other computing unit which is capable of performing program or executing codes stored in the storage module 120. The automated classification module 110 is configured to perform the method of processing cytometric data from multiple tests after the automated classification module 110 executes the plurality of codes stored in the storage module 120.

[0064] In some embodiments, the storage module 120 is configured to store the codes,programs, procedures and / or operations of the method of the present disclosure, which can be executed by the automated classification module 110. In some embodiments, the storage module 120 may include one or more non-volatile memory and one or more volatile memory. In some embodiments, volatile memory may be a finished product known to a person having ordinarily knowledge in the art to which the present disclosure belongs, such as, but not limited to, various types of dynamic random access memory or static random access memory. In some embodiments, non-volatile memory may be a finished product known to a person having ordinarily knowledge in the art to which the present disclosure belongs, such as, but not limited to, various types of read-only memory or flash memory. In some embodiments, the cytometric data from multiple tests could also be stored in the storage module 120.

[0065] In some embodiments, the I / O interface 130 is configured to allow the user to manipulate the system 100 to perform the methods or procedures, programs, operations of the present disclosure. In some embodiments, the I / O interface 130 may provide visual and / or auditory user interfaces, such as display units, touchscreens, projectors, loudspeakers, phone voices, keyboards, mouses, dynamic detection, and speech recognition, for functioning as mediums of manipulating the system 100.

[0066] In some embodiments, the communication interface 140 is configured to communicate with data outside the system 100. In some embodiments, the communication interface 140 signally connect to one or more flow cytometry instruments, different panels or other cytometric data providing devices outside the system 100, to receive the cytometric data from multiple tests via a physical signal line and / or a virtual signal line. In some embodiments, the communication interface 140 signally connect to database, server, cloud server or other storing devices to receive the cytometric data from multiple tests via a physical signal line and / or a virtual signal line. In some embodiments, the physical signal line, for example, may be a network signal line conforming to the Internet protocol, USB cable, RJ45, RS232, but is not limited thereto. In some embodiments, the virtual signal line, for example, may be Wi-Fi, 4G / 5G / 6G, Bluetooth, short-range communication, which conforms to wireless communication protocols, but is not limited thereto.

[0067] FIG.2 is a flowchart illustrating the method of processing cytometric data from multiple tests according to an embodiment of the present disclosure. The method of processing cytometric data from multiple tests shown in FIG. 2 can be executed by the automated classification module. The method of processing cytometric data from multiple tests shown in FIG. 2 include the steps of: (S100) determining overlapping parameters among the multipletests using flow cytometry; (S200) preprocessing the overlapping parameters; and (S300) developing classifiers via supervised machine learning.

[0068] In step S100, determining overlapping parameters among the multiple tests using flow cytometry. In some embodiments, the method utilizes common markers extracted from various sources as comparable features.

[0069] In some embodiments, the data is sourced from various manufacturers and instruments, employing different panels, which contain reagents that measure multiple markers / parameters, for flow cytometry analysis. In some embodiments, several distinct datasets and different panels are integrated, for some example, five distinct datasets and four different panels, which together encompass ten common parameters are integrated to provide the cytometric data from multiple tests.

[0070] In some embodiments, the data format of related files is managed and processed by following specific protocols, to ensure seamless integration into existing workflows. In some embodiments, the method provides extensive compatibility, supporting various flow cytometer data file formats such as the Flow Cytometry Standard (FCS) formats 2.0, 3.0, 3.1, and LMD listmode files, ensuring versatile application across various scientific disciplines.

[0071] In some embodiments, the parameters among the multiple tests may be collected from one flow cytometry instrument or panel. In some embodiments, the parameters among the multiple tests may be collected from multiple flow cytometry instruments or panels.

[0072] In some embodiments, the naming of the markers related to the cytometric data often varies, such as when markers have different aliases or the same marker is tagged with different fluorophores. In some embodiments, variations necessitate additional alignment efforts to ensure consistency and accuracy in data comparison, such as automatically searching basic information of the markers from a database to check the information and correlation of each marker, determining the spectrum range to identify the markers, etc.

[0073] In step S200, preprocessing the overlapping parameters. Specifically, to correct biases between different sources ensuring the method of processing cytometric data can be consistent and reliable across various datasets.

[0074] In some embodiments, the method employs one or more preprocessing steps including noise removal, data normalization techniques such as sample-wise z-score normalization, site- wise z-score normalization and rank normalization. In some embodiments, calibration techniques leveraging reference material could be employed as well.

[0075] Thus, the method can achieve high accuracy by using all overlapping parameters, andenable a robust comparison and analysis across diverse data sources. Additionally, through feature selection, the method has verified that utilizing one or more of the parameters can also yield comparably fair results. This flexibility allows for effective performance tuning based on specific analytical needs.

[0076] In some embodiments, two feature selection experiments results of cross five sites AML classification are shown below. The data were collected AML and non-neoplastic samples from five laboratories. Four panels and three instrument models were used in these laboratories for measuring. A total of ten common parameters were found across different panels, including FSC-A, FSC-H, SSC-A, CD7, CD19, CD34, CD45, CD56, CD117, and HLA-DR.

[0077] For the single parameter feature selection, Table 1 is a table showing the experiment results of performance of single parameter feature selection according to an embodiment of the present disclosure, and FIG.3 is a diagram showing the experiment results of performance of single parameter feature selection according to an embodiment of the present disclosure.

[0078] [Table 1] Single Parameter Used AUC ACC CD117 97.08% 94.41% SSC-A 95.16% 89.78% CD34 93.23% 87.42% HLA-DR 92.04% 86.50% FSC-H 83.96% 77.67% CD7 75.31% 71.13% CD56 73.32% 67.91% FSC-A 72.53% 66.50% CD19 67.85% 63.26% CD45 67.47% 60.97%

[0079] Referring to Table 1 and FIG.3, this experiment aims at evaluating the importance of each parameter by training and validating our approach with only one of the common parameters. As the table and figure shown in FIG.3, our approach could still achieve high and comparable performance with only one parameter, like CD117 and SSC-A.

[0080] For the sequential feature selection, Table 2 is a table showing the experiment results of performance of sequential feature selection according to an embodiment of the present disclosure.

[0081] [Table 2] Number of D parameterParameters includediscarded parameterAUC10FSC-A, FSC-H, SSC-A, CD7, CD19, CD34,CD45, CD56, CD117, HLA-DR - 99.94%9FSC-A, SSC-A, CD7, CD19, CD34, CD45,CD56, CD117, HLA-DR FSC-H 99.90%8FSC-A, SSC-A, CD19, CD34, CD45, CD56,CD117, HLA-DR CD7 99.90%7FSC-A, SSC-A, CD19, CD45, CD56, CD117,HLA-DR CD34 99.79%6FSC-A, SSC-A, CD19, CD45, CD117, HLA-DR CD56 99.92%5 SSC-A, CD19, CD45, CD117, HLA-DR FSC-A 99.95% 4 SSC-A, CD19, CD117, HLA-DR CD45 99.95% 3 SSC-A, CD19, CD117 HLA-DR 99.40% 2 SSC-A, CD117 CD19 97.93% 1 CD117 SSC-A 97.60%

[0082] Referring to Table 2, this experiment drops parameters one by a time to assess the least amount of parameters the method needs. The discarded parameter for each parameter number was decided by comparing which parameter being dropped produced the best result. The table below presents the best parameter sets for each parameter number. The performance remained or dropped a little as parameters were discarded. The method obtained the best AUC with only five and four parameters included. SSC-A and CD117 were the last two parameters being retained, indicating its importance for our approach to classify AML and non-neoplastic. This result was similar to the finding in the previous single parameter feature selection experiment.

[0083] In step S300, developing classifiers via supervised machine learning. The details of developing classifiers via supervised machine learning are described below.

[0084] In some embodiments, the determining of the overlapping parameters is performed by recognizing the overlapping parameters in different tests that describe same biological indicators. For example, identifying the information labeled with the overlapping parameters to recognize the information such as name, type, applied target, spectrum or biochemical feature, and to recognize the overlapping parameters which describe same biological indicators.

[0085] In some embodiments, the preprocessing of the overlapping parameters comprises rearranging the overlapping parameters and performing normalization. In some embodiments, rearranging the overlapping parameters comprises performing parameter alignment. In some embodiments, as some common markers may not be placed in the same tube or measured within the same fluorescent channel, the alignment process involved in rearranging the data frame so that each data frame has the parameters arranged in the same order. For example, the channel used to measure the CD45 in each panel is in different column positions before the alignment. After the alignment, these two parameters are at the same column order so that the data from two panels can be combined together for following model training.

[0086] In some embodiments, the normalization comprises z-score normalization. In someembodiments, the z-score normalization is defined as

[0087] The mean and standard deviation are calculated per sample over the fluorescence intensity in all cells and channels. ^ is a small value added to the denominator to maintain numerical stability.

[0088] In some embodiments, the preprocessing of the overlapping parameters further comprises capturing distribution and embedding phenotype characteristics into a phenotype representation.

[0089] In some embodiments, the distribution is a high-dimensional cellular distribution.

[0090] In some embodiments, the high-dimensional cellular distribution is performed by a Gaussian mixture model (GMM). In some embodiments, deep neural network approach is applied to generate sample representation.

[0091] In some embodiments, the representation is a high-dimensional phenotype representation.

[0092] In some embodiments, the embedding phenotype characteristics into the distribution is performed by Fisher vectorization to embed phenotype characteristics into a high-dimensional phenotype representation vector at a specimen level.

[0093] In some embodiments, the preprocessing of the overlapping parameters further comprises inputting vectors along with corresponding diagnoses into a support vector machine.

[0094] In some embodiments, the developing of the classifiers comprises data training and sample classification.

[0095] In some embodiments, the method is used to count cells, to sort cells, to determine cell function, to determine cell characteristics, to detect microorganisms such as bacteria, fungus or yeast, to find biomarkers, or to assist diagnosis and potential treatment of blood and bone marrow cancers. In some embodiments, the method is also used to evaluate treatment response, monitoring residual disease, identify biomarkers, cell quality controls, etc.

[0096] In some embodiments, the method further comprises performing visualization after the developing of the classifiers.

[0097] In some embodiments, the overlapping parameters are selected from a group comprising light scatter property and fluorescent marker.

[0098] FIG.4 shows an example scheme diagram of developing and visualizing panel agnosticsample classification using flow cytometry data according to an embodiment of the present disclosure. The flow cytometry data used in this disclosure were from the flow cytometry done as part of the clinical evaluation on bone marrow specimens of patients that have been evaluated for potential hematological diseases from different institutions.

[0099] Cross-tests data or data from multiple tests refers to data from two or more tests. In particular, cross-tests data or data from multiple tests refers to data from two or more tests which are processed via different models of flow cytometry instruments, different panels, different kits, different processing methods, different preprocessing methods, or two or more tests having any differences among them.

[0100]

[0101] Example 1: Cross-Panel AML versus non-neoplastic classification across two panels.

[0102] In some embodiments, 103 Flow Cytometry Standard (FCS) records from bone marrow samples is used, which are obtained from two different sources shown in Table 3.

[0103] [Table 3]

[0104] The Roswell Park Comprehensive Cancer Center (RPCCC) data were acquired using the Beckman Coulter Navios EX Cytometer (Navios EX) and measured with the ClearLLab10C panel, as detailed in Table 4a; and the University of Pittsburgh Medical Center (UPMC) data were acquired using a BD FACSCantoII and measured with a UPMC-developed diagnostic panel, as shown in Table 4b.

[0105] [Table 4]

[0106] FIG 5. illustrated an exemplary framework for developing and visualizing panel agnostic sample classification using flow cytometry data in detail according to an embodiment of the present disclosure. In some embodiments, the classification model development process of the present disclosure was illustrated in FIG. 5. In some embodiments, we utilized flow cytometry list mode data, treating each light scatter property and fluorescent marker as a distinct parameter. In some embodiments, data preprocessing involved parameter alignment and z-score normalization. In some embodiments, a Gaussian mixture model (GMM) was employed to capture high-dimensional cellular distribution, followed by Fisher vectorization to embed phenotype characteristics into a high-dimensional phenotype representation vector at the specimen level. In some embodiments, these vectors, along with their corresponding diagnoses, were input into a support vector machine (SVM) for AML versus non-neoplastic classification.

[0107] In some embodiments, we identified 25 fluorescent parameters that were commonly measured in both RPCCC and UPMC Panels for further data processing including parameter alignment and z-score normalization to develop the cross-panel classification model (Model D) and use three different scenarios as control: Model A- trained with all ClearLLab10C panel parameters on the RPCCC cohort, Model B- trained with these 25 overlapping parameters between ClearLLab10C and UPMC panels on the RPCCC cohort, Model C- trained with these 25 overlapping parameters between ClearLLab10C and UPMC panels on the UPMC cohort, and Model D-trained with these 25 overlapping parameters between ClearLLab10C and UPMC panels on the combined RPCCC and UPMC cohort.

[0108] The classification performance of these models is shown in Table 5. The Model A trained with all ClearLLab10C panel parameters on the RPCCC cohort achieved AUC of 100%. However, when only using overlapping parameters between ClearLLab10C and UPMC panels, the model's AUC marginally decreased to 99.86% (Model B). The cross-panel classifier (ModelD), trained on a combined RPCCC and UPMC dataset using 25 overlapping parameters, achieved an AUC of 99.06%.

[0109] [Table 5]

[0110] FIG.6A and 6B show the visualization results of samples in an exemplary cross-panel AML versus non-neoplastic classification according to an embodiment of the present disclosure. FIG 7. shows the exemplary results of identifying minimal parameters required for developing panel agnostic sample classification using flow cytometry data according to an embodiment of the present disclosure.

[0111] In some embodiments, dimension reduction was applied to each sample-level phenotype vector to visualize samples from a trained model. Each sample is represented as a dot on a 3D plot, with colors indicating the actual diagnosis. As depicted in FIG. 6A and 6B, Model D sample visualization showed good separation of AML from non-neoplastic samples in both the RPCCC and UPMC dataset.

[0112] In some embodiments, to evaluate the importance of each parameter and its interaction with other parameters, and to determine if the number of parameters could be reduced whilemaintaining performance, we conducted feature selection experiments. In these experiments, we paired the parameter that provided the highest accuracy with each of the remaining parameters to identify the combination that yielded the highest accuracy. Subsequently, we added the remaining parameters individually to identify the combination with the next-highest accuracy. As depicted in FIG.7, a model developed using a combination of parameters - CD117, CD123, CD15, CD34, and optical parameters achieved an equivalent performance with an accuracy of 97.01%, compared to the 96.01% accuracy of Model D (Table 5).

[0113]

[0114] Example 2: Cross-Panel AML versus non-neoplastic classification across three panels.

[0115] In some embodiments, the disclosures are discussed below in connection with an example of Cross-Panel AML versus non-neoplastic classification across three panels.

[0116] In this example, one hundred and thirty-nine Flow Cytometry Standard (FCS) records from peripheral blood and bone marrow samples obtained from three different sources shown in Table 6.

[0117] [Table 6]

[0118] Table 7 shows the sample level data panel composition; Table 8 shows the cell level data panel composition; Table 9 shows the cell level classification training set.

[0119] In some embodiments, the Roswell Park Comprehensive Cancer Center (RPCCC) data were acquired using the Beckman Coulter Navios EX Cytometer (Navios EX) and measured with the ClearLLab10C panel, as detailed in first row of Table 7; and the University ofPittsburgh Medical Center (UPMC) data were acquired using a BD FACSCantoII and measured with a UPMC-developed diagnostic panel, as shown in second row of Table 7; The National Taiwan University Cancer Center (NTUCC) data were acquired using the Beckman Coulter FACSLyric Flow Cytometer and measured with the Euroflow AML / MDS panel, as detailed in third row of Table 7.

[0120] [Table 7]

[0121] [Table 8]

[0122] [Table 9]

[0123] In some embodiments, to prevents bias towards majority class, we will balance the training set sizes for each class across 3 datasets.

[0124] In some embodiments, we identified 7 fluorescent parameters that were commonly measured in all RPCCC, UPMC, and NTUCC Panels for further data processing including parameter alignment and z-score normalization to develop the cross-panel classification model (Model D). In some embodiments, as the common markers may not be placed in the same tube or measuring using the same fluorescent channel, the alignment process involved in rearranging the data frame so that each data frame has the parameters arranged in the same order. For example, the CD34 measured in the ClearLLab10C panel, the CD34 measured in the Tube 1 of the UPMC panel, and the CD34 measured in Euroflow AML / MDS panel may be in different column positions before the alignment. After the alignment, make sure these three parameters are at the same column order so that the data from three panels can be combined together for model training.

[0125] In some embodiments, z-score normalization is defined as.

[0126] In some embodiments, the mean and standard deviation are calculated per sample over the fluorescence intensity in all cells and channels. ^ is a small value added to the denominator to maintain numerical stability.

[0127] In some embodiments, the model (Model D) is trained by aligning the optical parameters (FSC-A, FSC-H, SSC-A) together with parameters that were measured in all three panels: CD34, CD45, CD117, HLA-DR, for each sample to make sure each data frame has the same order of parameters and these aligned data were subjected to normalization, sample representation and classifier training steps.

[0128] In some embodiments, using gaussian mixture model (GMM) to capture the complex cellular distribution. Then, in some embodiments, a Fisher gradient vectorization approach was applied to embed phenotype characteristics in terms of the learned probability distribution in the derived specimen level high-dimensional phenotype representation. Each vector is coupled with classification labels such as AML or non-neoplastics. In some embodiments, then a supervised machine learning algorithm such as SVM was used to train AML versus non- neoplastic classifier.

[0129] In some embodiments, using three different scenarios as control: Model A- trained with all ClearLLab10C panel parameters on the RPCCC cohort; Model B- trained with all UPMC panel parameters on the UPMC cohort; Model C-trained with all Euroflow AML / MDS panel parameters on the NTUCC cohort; Model D-trained with there 7 overlapping parameters among the ClearLab10C, UPMC, and Euroflow AML / MDS panels on the RPCCC cohort; Model E- trained with there 7 overlapping parameters on UPMC cohort; Model F-trained with there 7 overlapping parameters on the NTUCC cohort; Model G-rained with there 7 overlapping parameters on the combined RPCCC, UPMC, and NTUCC cohort.

[0130] In some embodiments, the classification performance of these models is shown in Table 10. The Model A trained with all ClearLLab10C panel parameters on the RPCCC cohort achieved AUC of 100%. However, when only using overlapping parameters among all three panels, the model's AUC marginally decreased to 92.88% (Model D). The cross-panel classifier (Model G), trained on a combined RPCCC, UPMC, and NTUCC dataset using 7 overlapping parameters, achieved an AUC of 99.68%.

[0131] [Table 10]

[0132]

[0133] Example 3: Cross-Panel AML versus non-neoplastic classification across five panels.

[0134] In some embodiments, the disclosures are discussed below in connection with an example of Cross-Panel AML versus non-neoplastic classification across five panels.

[0135] In this example, two hundred and twenty Flow Cytometry Standard (FCS) records from bone marrow samples obtained from five different sources shown in Table 11.

[0136] [Table 11]

[0137] Data from each source is measured and acquired with different panel and instrument sets, as listed in Table 11. The detailed composition of panels was shown in Table 12.

[0138] [Table 12]

[0139] FIG.8 illustrates an example framework for developing and visualizing panel agnostic sample classification using flow cytometry data according to an embodiment of the present disclosure. The classification model development process of this example was illustrated in FIG. 8. In some embodiments, utilized flow cytometry list mode data, treating each light scatter property and fluorescent marker as a distinct parameter. In some embodiments, data preprocessing involved parameter alignment, compensation, downsampling, and z-score normalization. In some embodiments, a Gaussian mixture model (GMM) was employed to capture high-dimensional cellular distribution, followed by Fisher vectorization to embed phenotype characteristics into a high-dimensional phenotype.

[0140] In some embodiments, the framework pipeline shown in FIG. 8 are listed below: (1)Select and rearrange common parameters across different panels; (2) Pre-process data with compensation, downsampling (ratio=0.2), and sample-wise z-score normalization; (3) Utilize Gaussian Mixture Model (GMM) and Fisher Vector to transform each sample’s data into a 1D representation vector; and (4) Apply 3-fold cross validation to train and validate a Support Vecotr Macine (SVM) for binary classification.

[0141] In some embodiments, identified 10 optical and fluorescent parameters that were commonly measured across all panels for further data processing to develop the cross-panel classification model (Model K). In some embodiments, as the common markers may not be placed in the same tube or measured within the same fluorescent channel, the alignment process involved in rearranging the data frame so that each data frame has the parameters arranged in the same order. For example, the channel used to measure the CD34 in the ClearLLab10C panel and in the UPMC panel is in different column positions before the alignment. After the alignment, make sure these two parameters are at the same column order so that the data from two panels can be combined together for model training.

[0142] In some embodiments, compensation is a standard process in flow cytometry to resolve the influence of spillover. Spillover is a common situation in regular cytometers that one channel could collect light emitted both by the paired marker and markers of the other channels, resulting in an inaccurate measurement of light intensity. To resolve the issue, compensation is conducted with a compensation matrix, retrieved from measuring standard beads, and is defined as (x) ̂=x〖c〗^(-1), where c is the compensation matrix, x is the channel intensity matrix with spillover, and x is the channel intensity matrix without spillover.

[0143] In some embodiments, in the downsampling step, the amount of cells is decreased to the targeted ratio of the original amount. It is done by randomly sampling cells from the data.

[0144] In some embodiments, z-score normalization is defined as.

[0145] The mean and standard deviation are calculated per sample over the fluorescence intensity in all cells and channels. ^ is a small value added to the denominator to maintain numerical stability.

[0146] The model (Model K) is trained by aligning the optical and fluorescent parameters measured in all panels, including FSC-A, FSC-H, SSC-A, CD7, CD19, CD34, CD45, CD56,CD117, and HLA-DR. It is conducted on each sample to make sure each data frame has the same order of parameters and these aligned data were subjected to compensation, downsampling, normalization, sample representation, and classifier training steps.

[0147] In some embodiments, used a gaussian mixture model (GMM) to capture the complex cellular distribution. Then, in some embodiments, a Fisher gradient vectorization approach was applied to embed phenotype characteristics in terms of the learned probability distribution in the derived specimen level high-dimensional phenotype representation. Each vector is coupled with classification labels such as AML or non-neoplastics. Then a supervised machine learning algorithm such as SVM was used to train an AML versus non-neoplastic classifier.

[0148] In some embodiments, used ten different scenarios as control: Model A to E are trained with individual dataset and site dependent parameters. Model F to J are also trained with individual dataset, but only using the 10 commonly measured parameters across panels. Model K is trained with all data from five datasets and using the 10 common parameters.

[0149] In some embodiments, the classification performance of these models is shown in Table 3. By comparing the results of models trained on the same dataset with different parameter sets, a marginal performance drops or even score is found in all models. For example, Model D and Model I are both trained with UPMC dataset individually, but Model D slightly outperforms Model I in all metrics; Model E and Model J are also trained with the same individual dataset, NTUH, but both models reach 100% in each aspect. Model K is trained on all dataset together with commonly measured parameters across four different panels. It achieves great performance and outperforms most of the other models, with only two samples being mis- classified in two hundred and twenty cases.

[0150] [Table 13]

[0151] FIG. 9A and 9B show the visualization result of samples in an example cross-panel AML versus non-neoplastic classification according to an embodiment of the present disclosure. Dimension reduction was applied to each sample-level phenotype vector to visualize samples from a trained model. Each sample is represented as a dot on a 3D plot, with colors indicating the actual diagnosis and shapes representing the belonging source. FIG.9A and 9B shows the visualization result of samples in an example cross-panel AML versus non-neoplastic classification. As depicted in FIG. 9A and 9B, Model D sample visualization showed good separation of AML from non-neoplastic samples in all datasets.

[0152]

[0153] Example 4: Cross-Panel lymphoid abnormality versus no abnormality classification across two panels.

[0154] In some embodiments, the disclosures are discussed below in connection with an example of Cross-Panel lymphoid abnormality versus no abnormality classification across two panels.

[0155] In this example, three hundred and ninety-seven Flow Cytometry Standard (FCS) records from bone marrow samples obtained from two different sources shown in Table 14.

[0156] [Table 14]

[0157] Data from each source is measured and acquired with different panel and instrument sets, as listed in Table 14. The detailed composition of panels was shown in Table 15.

[0158] [Table 15]

[0159] FIG. 10 illustrates an example framework for developing and visualizing Panel Agnostic Sample Classification Using Flow Cytometry Data according to an embodiment of the present disclosure. The classification model development process of this example was illustrated in FIG. 10. In some embodiments, utilized flow cytometry list mode data, treating each light scatter property and fluorescent marker as a distinct parameter. In some embodiments, data preprocessing involved parameter alignment, compensation, downsampling, and z-score normalization. In some embodiments, a Gaussian mixture model (GMM) was employed to capture high-dimensional cellular distribution, followed by Fisher vectorization to embed phenotype characteristics into a high-dimensional phenotype.

[0160] In some embodiments, identified 5 optical and fluorescent parameters that were commonly measured across both panels for further data processing to develop the cross-panel classification model. In some embodiments, as the common markers may not be placed in the same tube or measured within the same fluorescent channel, the alignment process involved in rearranging the data frame so that each data frame has the parameters arranged in the same order. For example, the channel used to measure the CD45 in each panel is in different columnpositions before the alignment. After the alignment, make sure these two parameters are at the same column order so that the data from two panels can be combined together for model training.

[0161] In some embodiments, compensation is a standard process in flow cytometry to resolve the influence of spillover. Spillover is a common situation in regular cytometers that one channel could collect light emitted both by the paired marker and markers of the other channels, resulting in an inaccurate measurement of light intensity. To resolve the issue, compensation is conducted with a compensation matrix, retrieved from measuring standard beads, and is defined as x ̂=xc^(-1), where c is the compensation matrix, x is the channel intensity matrix with spillover, and x ̂ is the channel intensity matrix without spillover.

[0162] In some embodiments, in the downsampling step, the amount of cells is decreased to the targeted ratio of the original amount. It is done by randomly sampling cells from the data.

[0163] In some embodiments, z-score normalization is defined as.

[0164] The mean and standard deviation are calculated per sample over the fluorescence intensity in all cells and channels. ^ is a small value added to the denominator to maintain numerical stability.

[0165] In some embodiments, the model is trained by aligning the optical and fluorescent parameters measured in both panels, including FSC-H, SSC-H, CD3, CD5, CD38, and CD45. It is conducted on each sample to make sure each data frame has the same order of parameters and these aligned data were subjected to compensation, downsampling, normalization, sample representation, and classifier training steps.

[0166] In some embodiments, used a gaussian mixture model (GMM) to capture the complex cellular distribution. Then, in some embodiments, a Fisher gradient vectorization approach was applied to embed phenotype characteristics in terms of the learned probability distribution in the derived specimen level high-dimensional phenotype representation. Each vector is coupled with classification labels such as lymphoid abnormality or no abnormality. Then a supervised machine learning algorithm such as SVM was used to train a lymphoid abnormality versus no abnormality classifier.

[0167] In some embodiments, the classification performance of these models is shown in Table 16. The model achieves good performance with only five parameters, which is a lot less than both of the original panel designs.

[0168] [Table 16]

[0169] FIG.11 shows the visualization result of samples in an example cross-panel Lymphoid abnormality versus no abnormality classification according to an embodiment of the present disclosure.

[0170] In some embodiments, dimension reduction was applied to each sample-level phenotype vector to visualize samples from a trained model. Each sample is represented as a dot on a 3D plot, with colors indicating the actual diagnosis and shapes representing the belonging source. FIG. 11 shows the visualization result of samples in an example cross-panel lymphoid abnormality versus no abnormality classification. As depicted in FIG.11, sample visualization showed good separation of lymphoid abnormality from no abnormality samples in all datasets.

[0171] FIG. 12 is a block diagram that illustrates an example of an artificial intelligent (AI) system 500 in which at least some operations described herein can be implemented according to an embodiment of the present disclosure. As shown, the AI system 500 can include a set of layers, which conceptually organize elements within an example network topology for the AI system’s architecture to implement a particular AI model 530. Generally, an AI model 530 is a computer-executable program implemented by the AI system 500 that analyzes data to make predictions. Information can pass through each layer of the AI system 500 to generate outputs for the AI model 530. The layers can include a data layer 502, a structure layer 504, a model layer 506, and an application layer 508. The algorithm 516 of the structure layer 504 and the model structure 520 and model parameters 522 of the model layer 506 together form the example AI model 530. In some embodiments, the optimizer 526, loss function engine 524, and regularization engine 528 work to refine and optimize the AI model 530, and the data layer 502 provides resources and support for application of the AI model 530 by the application layer 508.

[0172] The data layer 502 acts as the foundation of the AI system 500 by preparing data for the AI model 530. As shown, the data layer 502 can include two sub-layers: a hardware platform 510 and one or more software libraries 512. The hardware platform 510 can be designed to perform operations for the AI model 530 and include computing resources for storage, memory, logic and networking, such as the resources described in relation to FIG. 10. The hardwareplatform 510 can process amounts of data using one or more servers. The servers can perform backend operations such as matrix calculations, parallel calculations, machine learning (ML) training, and the like. Examples of servers used by the hardware platform 510 include central processing units (CPUs) and graphics processing units (GPUs). CPUs are electronic circuitry designed to execute instructions for computer programs, such as arithmetic, logic, controlling, and input / output (I / O) operations, and can be implemented on integrated circuit (IC) microprocessors. GPUs are electric circuits that were originally designed for graphics manipulation and output but may be used for AI applications due to their vast computing and memory resources. GPUs use a parallel structure that generally makes their processing more efficient than that of CPUs. In some instances, the hardware platform 510 can include Infrastructure as a Service (IaaS) resources, which are computing resources, (e.g., servers, memory, etc.) offered by a cloud services provider. The hardware platform 510 can also include computer memory for storing data about the AI model 530, application of the AI model 530, and training data for the AI model 530. The computer memory can be a form of random-access memory (RAM), such as dynamic RAM, static RAM, and non-volatile RAM.

[0173] The software libraries 512 can be thought of as suites of data and programming code, including executables, used to control the computing resources of the hardware platform 510. The programming code can include low-level primitives that form the foundation of one or more low-level programming languages, such that servers of the hardware platform 510 can use the low-level primitives to carry out specific operations. The low-level programming languages do not require much, if any, abstraction from a computing resource’s instruction set architecture, allowing them to run quickly with a small memory footprint. Examples of software libraries 512 that can be included in the AI system 500 include Intel Math Kernel Library, Nvidia cuDNN, Eigen, and Open BLAS.

[0174] The structure layer 504 can include an ML framework 514 and an algorithm 516. The ML framework 514 can be thought of as an interface, library, or tool that allows users to build and deploy the AI model 530. The ML framework 514 can include an open-source library, an application programming interface (API), a gradient-boosting library, an ensemble method, and / or a deep learning toolkit that work with the layers of the AI system facilitate development of the AI model 530. For example, the ML framework 514 can distribute processes for application or training of the AI model 530 across multiple resources in the hardware platform 510. The ML framework 514 can also include a set of pre-built components that have the functionality to implement and train the AI model 530 and allow users to use pre-built functionsand classes to construct and train the AI model 530. Thus, the ML framework 514 can be used to facilitate data engineering, development, hyperparameter tuning, testing, and training for the AI model 530. Examples of ML frameworks 514 that can be used in the AI system 500 include TensorFlow, PyTorch, Scikit-Learn, Keras, Cafffe, LightGBM, Random Forest, and Amazon Web Services.

[0175] The algorithm 516 can be an organized set of computer-executable operations used to generate output data from a set of input data and can be described using pseudocode. The algorithm 516 can include complex code that allows the computing resources to learn from new input data and create new / modified outputs based on what was learned. In some implementations, the algorithm 516 can build the AI model 530 through being trained while running computing resources of the hardware platform 510. This training allows the algorithm 516 to make predictions or decisions without being explicitly programmed to do so. Once trained, the algorithm 516 can run at the computing resources as part of the AI model 530 to make predictions or decisions, improve computing resource performance, or perform tasks. The algorithm 516 can be trained using supervised learning, unsupervised learning, semi-supervised learning, and / or reinforcement learning.

[0176] Using supervised learning, the algorithm 516 can be trained to learn patterns (e.g., map input data to output data) based on labeled training data. The training data may be labeled by an external user or operator. For instance, a user may collect a set of training data, such as by capturing data from sensors, images from a camera, outputs from a model, and the like. In an example implementation, training data can include the parameters described above. The user may label the training data based on one or more classes and trains the AI model 530 by inputting the training data to the algorithm 516. The algorithm determines how to label the new data based on the labeled training data. The user can facilitate collection, labeling, and / or input via the ML framework 514. In some instances, the user may convert the training data to a set of feature vectors for input to the algorithm 516. Once trained, the user can test the algorithm 516 on new data to determine if the algorithm 516 is predicting accurate labels for the new data. For example, the user can use cross-validation methods to test the accuracy of the algorithm 516 and retrain the algorithm 516 on new training data if the results of the cross-validation are below an accuracy threshold.

[0177] Supervised learning can involve classification and / or regression. Classification techniques involve teaching the algorithm 516 to identify a category of new observations based on training data and are used when input data for the algorithm 516 is discrete. Said differently,when learning through classification techniques, the algorithm 516 receives training data labeled with categories (e.g., classes) and determines how features observed in the training data relate to the categories. Once trained, the algorithm 516 can categorize new data by analyzing the new data for features that map to the categories. Examples of classification techniques include boosting, decision tree learning, genetic programming, learning vector quantization, k- nearest neighbor (k-NN) algorithm, and statistical classification.

[0178] Regression techniques involve estimating relationships between independent and dependent variables and are used when input data to the algorithm 516 is continuous. Regression techniques can be used to train the algorithm 516 to predict or forecast relationships between variables. To train the algorithm 516 using regression techniques, a user can select a regression method for estimating the parameters of the model. The user collects and labels training data that is input to the algorithm 516 such that the algorithm 516 is trained to understand the relationship between data features and the dependent variable(s). Once trained, the algorithm 516 can predict missing historic data or future outcomes based on input data. Examples of regression methods include linear regression, multiple linear regression, logistic regression, regression tree analysis, least squares method, and gradient descent. In an example implementation, regression techniques can be used, for example, to estimate and fill-in missing data for machine-learning based pre-processing operations.

[0179] Under unsupervised learning, the algorithm 516 learns patterns from unlabeled training data. In particular, the algorithm 516 is trained to learn hidden patterns and insights of input data, which can be used for data exploration or for generating new data. Here, the algorithm 516 does not have a predefined output, unlike the labels output when the algorithm 516 is trained using supervised learning. Said another way, unsupervised learning is used to train the algorithm 516 to find an underlying structure of a set of data, group the data according to similarities, and represent that set of data in a compressed format.

[0180] A few techniques can be used in supervised learning: clustering, anomaly detection, and techniques for learning latent variable models. Clustering techniques involve grouping data into different clusters that include similar data, such that other clusters contain dissimilar data. For example, during clustering, data with possible similarities remain in a group that has less or no similarities to another group. Examples of clustering techniques include density-based methods, hierarchical based methods, partitioning methods, and grid-based methods. In one example, the algorithm 516 may be trained to be a k-means clustering algorithm, which partitions n observations in k clusters such that each observation belongs to the cluster with the nearestmean serving as a prototype of the cluster. Anomaly detection techniques are used to detect previously unseen rare objects or events represented in data without prior knowledge of these objects or events. Anomalies can include data that occur rarely in a set, a deviation from other observations, outliers that are inconsistent with the rest of the data, patterns that do not conform to well-defined normal behavior, and the like. When using anomaly detection techniques, the algorithm 516 may be trained to be an Isolation Forest, local outlier factor (LOF) algorithm, or K-nearest neighbor (k-NN) algorithm. Latent variable techniques involve relating observable variables to a set of latent variables. These techniques assume that the observable variables are the result of an individual’s position on the latent variables and that the observable variables have nothing in common after controlling for the latent variables. Examples of latent variable techniques that may be used by the algorithm 516 include factor analysis, item response theory, latent profile analysis, and latent class analysis.

[0181] The model structure 520 describes the architecture of the AI model 530 of the AI system 500. The model structure 520 defines the complexity of the pattern / relationship that the AI model 530 expresses. Examples of structures that can be used as the model structure 520 include decision trees, support vector machines, regression analyses, Bayesian networks, Gaussian processes, genetic algorithms, and artificial neural networks (or, simply, neural networks). The model structure 520 can include a number of structure layers, a number of nodes (or neurons) at each structure layer, and activation functions of each node. Each node’s activation function defines how a node converts data received to data output. The structure layers may include an input layer of nodes that receive input data, an output layer of nodes that produce output data. The model structure 520 may include one or more hidden layers of nodes between the input and output layers. The model structure 520 can be an Artificial Neural Network (or, simply, neural network) that connects the nodes in the structured layers such that the nodes are interconnected. Examples of neural networks include Feedforward Neural Networks, convolutional neural networks (CNNs), Recurrent Neural Networks (RNNs), Autoencoder, and Generative Adversarial Networks (GANs).

[0182] The model parameters 522 represent the relationships learned during training and can be used to make predictions and decisions based on input data. The model parameters 522 can weight and bias the nodes and connections of the model structure 520. For instance, when the model structure 520 is a neural network, the model parameters 522 can weight and bias the nodes in each layer of the neural networks, such that the weights determine the strength of the nodes and the biases determine the thresholds for the activation functions of each node. Themodel parameters 522, in conjunction with the activation functions of the nodes, determine how input data is transformed into desired outputs. The model parameters 522 can be determined and / or altered during training of the algorithm 516.

[0183] The loss function engine 524 can determine a loss function, which is a metric used to evaluate the AI model’s 530 performance during training. For instance, the loss function engine 524 can measure the difference between a predicted output of the AI model 530 and the actual output of the AI model 530 and is used to guide optimization of the AI model 530 during training to minimize the loss function. The loss function may be presented via the ML framework 514, such that a user can determine whether to retrain or otherwise alter the algorithm 516 if the loss function is over a threshold. In some instances, the algorithm 516 can be retrained automatically if the loss function is over the threshold. Examples of loss functions include a binary-cross entropy function, hinge loss function, regression loss function (e.g., mean square error, quadratic loss, etc.), mean absolute error function, smooth mean absolute error function, log-cosh loss function, and quantile loss function.

[0184] The optimizer 526 adjusts the model parameters 522 to minimize the loss function during training of the algorithm 516. In other words, the optimizer 526 uses the loss function generated by the loss function engine 524 as a guide to determine what model parameters lead to the most accurate AI model 530. Examples of optimizers include Gradient Descent (GD), Adaptive Gradient Algorithm (AdaGrad), Adaptive Moment Estimation (Adam), Root Mean Square Propagation (RMSprop), Radial Base Function (RBF) and Limited-memory BFGS (L- BFGS). The type of optimizer 526 used may be determined based on the type of model structure 520 and the size of data and the computing resources available in the data layer 502.

[0185] The regularization engine 528 executes regularization operations. Regularization is a technique that prevents over- and under-fitting of the AI model 530. Overfitting occurs when the algorithm 516 is overly complex and too adapted to the training data, which can result in poor performance of the AI model 530. Underfitting occurs when the algorithm 516 is unable to recognize even basic patterns from the training data such that it cannot perform well on training data or on validation data. The regularization engine 528 can apply one or more regularization techniques to fit the algorithm 516 to the training data properly, which helps constrain the resulting AI model 530 and improves its ability for generalized application. Examples of regularization techniques include lasso (L1) regularization, ridge (L2) regularization, and elastic (L1 and L2 regularization).

[0186] FIG. 13 shows an example scheme diagram of z-score normalization according to anembodiment of the present disclosure.

[0187] In some embodiments, z-score normalization is a strategy of normalizing data. This technique scales the values of a feature to have a mean of 0 and a standard deviation of 1. The formula for z-score normalization is below:

[0188] z = (x - μ) / σ

[0189] Where x is the original value, μ is the mean value of the feature and σ is the standard deviation of the feature.

[0190] Referring to FIG.13, FIG.13 shows an example with x=60, μ=50, σ=10. After z-score normalization, z=(60-50) / 10=1, then μ=0, σ=1.

[0191] In some embodiments, for sample z-score normalization, perform z-score normalization on each channel of the sample individually. In some embodiments, for site z-score normalization, perform z-score normalization on each channel of the dataset individually.

[0192] FIG. 14 shows an example scheme diagram of density contour according to an embodiment of the present disclosure.

[0193] In some embodiments, probability density function (PDF) likes a density map, where higher values indicate a higher probability of the data falling within that region. Density contour represent the probability density function in a two-dimensional space.

[0194] FIG.15 shows an example scheme diagram of classifier according to an embodiment of the present disclosure.

[0195] In some embodiments, for the Support Vector Machine (SVM) Classifier, SVM work by finding an optimal decision boundary, called a hyperplane, that best separates different classes of data points. This hyperplane is positioned in a way that maximizes the margin between the data points closest to it. In some embodiments, for Extreme Gradient Boosting (XGBoost) Classifier, the XGBoost classifier is a ensemble learning algorithm that uses multiple decision trees for prediction. It employs a strategy called gradient boosting. Each tree focuses on correcting the errors from the previous tree, resulting in a more accurate ensemble model. The final prediction is often obtained by summing the individual predictions from each decision tree.

[0196] FIG.16 and 17 show the example scheme diagrams of evaluation metrics according to an embodiment of the present disclosure.

[0197] In some embodiments, the formula of ACC(Accuracy) is show below:

[0198]

[0199] The overall ability of a model to correctly predict both actual positive and actual negative.

[0200] In some embodiments, the formula of Sensitivity = Recall (True Positive Rate) is show below:

[0201]

[0202] The ability of a model to correctly predict actual positive.

[0203] In some embodiments, the formula of Specificity (True Negative Rate) is show below:

[0205] The ability of a model to correctly predict actual negative.

[0206] In some embodiments, the formula of AUC (Area Under the ROC Curve) is show below:

[0207] For True Positive Rate

[0208] For False Positive Rate (FPR):

[0209] In some embodiments, the formula of UAR (Unweighted Average Recall) is show below:

[0211] The sum of class-wise recall divided by number of classes.

[0212] In some embodiments, the formula of Weighted F1 is show below:

[0213]

[0214] The F1 score combines precision and recall using their harmonic mean. Weighted F1 is calculated by weighting each class's F1 score by its number of true instances and then averaging them.

[0215] In some embodiments, in cell level classification, we use macro-average to calculate the average performance across multiple classes in a classification task for both ACC and AUC. It gives equal weight to each class, regardless of class imbalance.

[0216] FIG.18 shows an example scheme diagram of GMM and Fisher vectorization according to an embodiment of the present disclosure.

[0217] In some embodiments, GMM (Gaussian Mixture Model) is a probabilistic model that assumes data is generated from a mixture of several Gaussian distributions, where each gaussiandistribution represents a cluster within the data. GMM offers greater flexibility by allowing for more diverse cluster shapes and providing probabilities for each data point's belonging to each cluster.

[0218] In some embodiments, the Fisher vector utilizes GMM to capture statistical properties of features. It calculates gradients for each original feature with respect to the means and variances of each Gaussian distribution. Through this process, it helps us understand the influence of features on the fisher vector. Features with larger gradients have a more significant impact on the Fisher Vector, suggesting their potential importance in characterizing the data distribution.

[0219] FIG. 19 shows an example scheme diagram of cross validation according to an embodiment of the present disclosure.

[0220] In some embodiments, the cross validation steps are listed as below:

[0221] 1. Divide the Data:

[0222] The dataset is divided into k subsets of equal size. These are often called "folds". (We set k = 5 in our experiment.)

[0223] 2. Training and Testing:

[0224] The model is trained k times, each time using k-1 folds for training and the remaining fold for testing. So, in each iteration, a different fold is used for testing while the other k-1 folds are used for training.

[0225] 3. Aggregation & Evaluation:

[0226] After training and testing the model k times, the testing results from each iteration are combined, and the performance of the model is evaluated.

[0227] FIG. 20 shows an example scheme diagram of compensation according to an embodiment of the present disclosure.

[0228] In some embodiments, fluorophores emit light at specific wavelengths, but there's often overlap in their emission spectra, like with PE and FITC. This overlap means that fluorescence from one fluorophore might be detected in the channel meant for another, creating false positive signals. This is especially problematic in multicolor experiments. To address this, compensation is necessary to ensure that the fluorescence detected accurately reflects the fluorophore being measured.

[0229] In some embodiments, the compensation steps are listed as below:

[0230] 1. Use flowutils to obtain the spill matrix from the FCS files.

[0231] 2. Then, employing the spill matrix as the coefficient matrix, we perform matrixoperations on FCS data to compensate for spectral overlap effects.

[0232] In some embodiments, the individual steps in the method of processing cytometric data from multiple tests described in the present disclosure may be further combined, replaced, repeatedly performed and / or modified, so as to generate new embodiments within the scope disclosed in the present disclosure.

[0233] In some embodiments, the individual steps of the method of processing cytometric data from multiple tests described in the present disclosure may be stored in a computer-readable recording medium that may be a non-transitory computer-readable recording medium such as a hard disk, optical disc, magnetic disk, flash drive, or a database accessible by a network, but is not limited thereto. After the computer-readable recording medium loads the computer program product stored in the computer-readable recording medium through the computing device, and executes the computer program product, it can realize any one of the methods of processing cytometric data from multiple tests described in the present disclosure.

[0234] In some embodiments, the computer program product of processing cytometric data from multiple tests described in the present disclosure may include individual steps of the method of processing cytometric data from multiple tests described in the present disclosure, so that the computing device can realize any one of the methods of processing cytometric data from multiple tests described in the present disclosure after loading the computer program product and executing the computer program product.

[0235] While the present disclosure contains many specifics, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in the present disclosure in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0236] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results.Moreover, the separation of various system components in the embodiments described in the present disclosure should not be understood as requiring such separation in all embodiments.

[0237] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in the present disclosure.

Claims

WHAT IS CLAIMED IS:

1. A method of processing cytometric data from multiple tests, comprising: determining overlapping parameters among the multiple tests using flow cytometry; preprocessing the overlapping parameters; and developing classifiers via supervised machine learning.

2. The method of claim 1, wherein the determining of the overlapping parameters is performed by recognizing the overlapping parameters in different tests that describe same biological indicators.

3. The method of claim 1, wherein the preprocessing of the overlapping parameters comprises rearranging the overlapping parameters and performing normalization.

4. The method of claim 3, wherein rearranging the overlapping parameters comprises performing parameter alignment.

5. The method of claim 3, wherein the normalization comprises z-score normalization.

6. The method of claim 3, wherein the preprocessing of the overlapping parameters further comprises capturing distribution and embedding phenotype characteristics into a phenotype representation.

7. The method of claim 6, wherein the distribution is a high-dimensional cellular distribution.

8. The method of claim 7, wherein the high-dimensional cellular distribution is performed by a Gaussian mixture model (GMM).

9. The method of claim 6, wherein the representation is a high-dimensional phenotype representation.

10. The method of claim 6, wherein the embedding phenotype characteristics into the distribution is performed by Fisher vectorization to embed phenotype characteristics into a high-dimensional phenotype representation vector at a specimen level.

11. The method of claim 6, wherein the preprocessing of the overlapping parameters further comprises inputting vectors along with corresponding diagnoses into a support vector machine.

12. The method of claim 1, wherein the developing of the classifiers comprises data training and sample classification.

13. The method of claim 1, wherein the method is used to count cells, to sort cells, to determine cell function, to determine cell characteristics, to detect microorganisms such as bacteria, fungus or yeast, to find biomarkers, or to assist diagnosis and potential treatment of bloodand bone marrow cancers.

14. The method of claim 1, wherein further comprising performing visualization after the developing of the classifiers.

15. The method of claim 1, wherein the overlapping parameters are selected from a group comprising light scatter property and fluorescent marker.

16. A system of processing cytometric data from multiple tests, which is suitable for signally connecting with one or more cytometric data providing devices in order to receive the cytometric data from the cytometric data providing devices, the system comprising: a storage module; and an automated classification module, configured to be signally connected with the storage module; wherein a plurality of codes is stored in the storage module, and wherein the automated classification module performs the steps of the method of processing cytometric data from multiple tests according to any one of claims 1 to 15 after the automated classification module executes the plurality of codes stored in the storage module.

Citation Information

Patent Citations

  • Systems and Methods for Analyzing Mixed Cell Populations

    US20200176080A1

  • Automated classification of immunophenotypes represented in flow cytometry data

    WO2022056478A2