Detection of minimal residual disease associated with cancer using flow cytometry data

Flow ARC addresses the inefficiencies of existing MRD detection methods by using a multi-level classification approach with ResNet architectures, achieving over 97% accuracy in MRD detection, thereby enhancing automation and reducing resource intensity.

WO2026072067A1PCT designated stage Publication Date: 2026-04-02MEMORIAL SLOAN KETTERING CANCER CENT +2
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Current flow cytometry methods for detecting minimal residual disease (MRD) in leukemia are labor-intensive, time-consuming, and reliant on experienced personnel, with existing machine learning models falling short of clinically acceptable accuracy for MRD detection.

Method used

The Flow ARC model employs a multi-level classification strategy combining cell-level and sample-level classification using ResNet architectures, leveraging both manual and automated gating techniques to accurately detect MRD by truncating flow cytometry data into a subset of cells with high tumor likelihood, followed by sample-level analysis.

Benefits of technology

Flow ARC achieves high accuracy and efficiency in detecting MRD, reducing turnaround time and resource intensity, with accuracy rates exceeding 97% on real-world patient cases, outperforming previous models like CellCNN.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024049298_02042026_PF_FP_ABST
    Figure US2024049298_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are systems and methods of classifying samples based on flow cytometry data. A computing system may receive a dataset comprising a plurality of event sequences for a plurality of flow channels. The flow channels may correspond to white blood cell types in a sample obtained from a subject at risk of or diagnosed with cancer. The computing system may apply the dataset to a first machine learning (ML) model to determine a plurality of first values for the plurality of event sequences. The computing system may sort the dataset by the plurality of first values. The computing system may apply the sorted dataset to a second ML model to determine a second value indicating a probability of tumor associated with cancer in the sample of the subject. The computing system may generate a classification in accordance with the second value.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt. No.: 115872-3009DETECTION OF MINIMAL RESIDUAL DISEASE ASSOCIATED WITH CANCER USING FLOW CYTOMETRY DATABACKGROUND

[0001] A computing system may receive input data. Using a machine learning model, the computing system may process the input data to generate output data.SUMMARY

[0002] Aspects of the present disclosure are directed to systems and methods of classifying samples based on flow cytometry data. One or more processors coupled with memory may receive, from a flow cytometer, a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels. Each of the plurality of event sequences may identify a respective set of events in a corresponding flow channel of the plurality of flow channels arranged in a first order. The set of events may correspond to at least one of a plurality of white blood cell types in a sample obtained from a subject at risk of or diagnosed with cancer. The one or more processors may apply the first dataset to a first machine learning (ML) model to determine a plurality of first values corresponding to the plurality of event sequences. Each first value of the plurality of first values may indicate a respective probability of tumor associated with cancer. The one or more processors may sort the set of events in the plurality of event sequences arranged in the first order in the first dataset by the plurality of first values to generate a second dataset comprising the plurality of event sequences. Each of the plurality of event sequences in the second dataset may include the respective set of events in a second order.

[0003] The one or more processors may apply the second dataset to a second ML model. The second ML model may be established using a plurality of training samples. Each training sample of the plurality of training samples may include a respective dataset including a corresponding plurality of event sequences labeled with presence or absence of tumor associated with cancer. The one or more processors may determine, based on applying the second dataset to the second ML model, a second value indicating a probability of tumor associated with cancer in the sample of the subject. The one or more processors-1-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 may generate a classification indicating a presence or absence of tumor associated with cancer in the sample in accordance with the second value. The one or more processors may store, using one or more data structures, an association between the subject and the classification.

[0004] In some embodiments, the one or more processors may generate the classification to indicate the presence of the tumor associated with leukemia in the sample, responsive to the second value satisfying a threshold indicative of a presence of a minimum residual disease (MRD) in the subject. The one or more processors may provide an output identifying at least one of (i) the classification to indicate the presence of the tumor, (ii) the presence of the MRD in the subject, or (iii) the subject as to be administered with an antileukemia therapy for leukemia. In some embodiments, the anti-leukemia therapy may include at least one of a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor.Examples of FLT inhibitors include, but are not limited to, sunitinib, sorafenib, midostaurin, lestaurtinib, quizartinib, gilteritinib, crenolanib, and gilteritini. Examples of IDH inhibitors include, but are not limited to, ivosidenib, olutasidenib, vorasidenib, and enasidenib.Examples of BTK inhibitors include, but are not limited to, ibrutinib, acalabrutinib, Zanubrutinib, and pirtobrutinib. Examples of BCL2 inhibitors include, but are not limited to, venetoclax, navitoclax, obatoclax, oblimersen sodium (G3139), Palcitoclax, AT-101, and LP-118. Examples of hedgehog pathway inhibitors include, but are not limited to, glasdegib, vismodegib, sonidegib, LEQ-506 (NVP-LEQ506), itraconazole, saridegib, BMS- 833923 / XL139, Arsenic Trioxide (ATO) and Taladegib.

[0005] In some embodiments, the one or more processors may generate the classification to indicate the absence of the tumor associated with cancer in the sample, responsive to the second value not satisfying a threshold indicative of a MRD in the subject. The one or more processors may provide an output identifying at least one of (i) the classification to indicate the absence of the tumor or (ii) the absence of the MRD in the subject. In some embodiments, the one or more processors may transform the plurality of event sequences in the second dataset from a time-series domain to a frequency domain.-2-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009The one or more processors may apply the second dataset in the frequency domain to second ML model to determine the second value for the sample.

[0006] In some embodiments, the one or more processors may apply each event in each of the set of events in the plurality of event sequences to the first ML model to generate a corresponding first value of the first plurality of first values. The one or more processors may apply the plurality of event sequences of the second dataset in entirety to the second ML model to determine the second value for the sample. In some embodiments, the one or more processors may truncate the plurality of event sequences based on the plurality of first values to select a subset of the plurality of event sequences to include in the second dataset

[0007] In some embodiments, the first ML model may be established using a plurality of second training samples. Each second training sample of the plurality of second training samples may include a respective dataset including a corresponding plurality of event sequences for a respective plurality of flow channels. The corresponding plurality of event sequences may include a set of events each labeled as one of a normal cell or a tumorous cell. In some embodiments, at least one of the plurality of training samples used to train the respective dataset generated from a combination of a plurality of portions from a plurality of datasets. The plurality of datasets may correspond at least one of a normal sample or tumor sample.

[0008] In some embodiments, the one or more processors may receive the first dataset comprising the plurality of event sequences for a corresponding plurality of flow channels including at least one of a forward scatter channel, a side scatter channel, or a fluorescence channel. In some embodiments, the leukemia may include at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL). In some embodiments, the plurality of white cell blood types may include at least one of CD10+, CD19+, CD20+, CD34+, CD38+, or CD45+.

[0009] In some embodiments, the cancer or tumor is a carcinoma, sarcoma, a melanoma, or a hematopoietic cancer. In some embodiments, the cancer or tumor is selected-3-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 from among adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin's disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, non-Hodgkin's lymphoma, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] FIGs. 1A-D: Block diagram of FlowARC, and multi-module deep-learningbased model, follows a four-step approach for detecting tumor presence in flow cytometry data. The initial step (A) is a pipeline to preprocess the data. Second, in the Audit stage (B), a subset of data is meticulously reviewed by pathologists, assigning each cell as either normal or tumor. These labeled cells are then used to train the Cell-level module, a deeplearning model that takes a single cell measurement as input and assigns a probability of it being a tumor cell. Next, in the Reorder stage (C), all cells within a flow sample are rearranged based on their tumor probabilities as determined by the Cell-level module. This reordering places cells with a higher likelihood of being tumors towards the top of the sample, allowing for the truncation of hundreds of thousands of cells into a few thousand relevant cells for tumor detection in each sample. Finally, in the Classify stage (D), the reordered and truncated flow samples, labeled as either normal or tumor, are used to train the Sample-level module. This module determines if a given flow sample is tumorous.

[0011] FIGs. 2A-E: (A) The results of FlowARC -Manual on the Synthetic test dataset of the annotated cohort as a confusion matrix. (B) The results of FlowARC -Manual on the patient test cases of the annotated cohort as a confusion matrix. (C) The results of FlowARC-Auto on the Synthetic test dataset of the clinical archive as a confusion matrix. (D) The results of FlowARC -Auto on the patient test cases of the clinical archive as a confusion matrix. E The comparison between FlowARC and CellCNN on the patient test -4- 4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 cases of the annotated cohort (left) and the clinical archive (right) is shown using the area under the ROC curve metric with 95% confidence interval.

[0012] FIGs. 3A-C: The visualization chart using the FlowARC-Auto cell-level module on six patient test cases of the clinical archive is shown. In this chart, each cells are color-coded according to tumor likelihood obtained from the cell-level module. (A) shows two cases that were correctly predicted by FlowARC-Auto. (B) shows two cases that were incorrectly predicted by FlowARC-Auto. (C) shows two cases that could not be injected onto the FlowARC-Auto sample-level module due to their B-cell population being less than 7500.

[0013] FIGs. 4 A and 4B: The performance of Flow ARC -Manual on predicting the tumor cell population on the B- ALL positive (tumor) synthetic test samples (A) and patient test cases (B) is shown. On each panel, left graph show the results from the FlowARC- Manual quantification module. For cases with quantification module prediction of 7200 and above, right graph shows the results from the averaged Flow ARC -Manual cell-level module predictions. The perfect prediction for each graph is denoted by a dashed line.

[0014] FIGs. 5A-5C: A shows the FlowARC -Manual cell-level module results as a set of bi-parametric plots on a random representative subset of test cells. Each cell is represented by a dot whose color and size reflect the classification outcomes, as shown in the legend at the bottom- right. B shows the SHAP summary plot for cell predictions, tumor (top) and normal (bottom). The cell markers are arranged in rows according to their significance for each category. The overlapping SHAP values are jittered along the y-axis. The color bar encodes the input value of the feature. C shows the t-SNE projection of the penultimate layer (features) of the Flow ARC -Manual sample-level module predictions on the synthetic test dataset. Each synthetic test sample is represented by a dot whose color and size reflect the classification outcomes, as shown in the legend at the bottom.

[0015] FIG. 6: The distribution of patient cases and the process of obtaining labeled cells for datasets used in this study, annotated cohort and the clinical archive, is shown.

[0016] FIG. 7: The four-step approach on auto-gating the CD 19+ B-cells from patient cases using flowDensity is shown. Step 1 removes doublets and margin events.-5-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009Step 2 and 3 selects the viable mononuclear population. Step 4 selects the B-cell population.

[0017] FIGs. 8A-C: A shows the distribution of cases within a synthetic dataset and their respective tumor population ranges as a pie-chart. B illustrates the process of generating a normal synthetic case containing N normal B-cells. C illustrates the process of generating a tumor synthetic case containing a total of N cells with NN normal B-cells and NA tumor cells.

[0018] FIGs. 9A-D: (A) shows the heatmap of the cell markers and cell diagnosis (Dx) in a random synthetic flow sample with 65,000 tumor cells before and after reordering using the Flow ARC- Manual cell-level module. (B) is a histogram showing the percentage of actual tumor population within the top 7,500 cells of synthetic tumor test set. (C) is a histogram showing the false negative rate for different minimum tumor populations within a synthetic tumor test set. (D) shows the cell distributions of the labeled cells of the annotated cohort (left) and the clinical archive (right) as pie-charts.

[0019] FIGs. 10A and 10B: (A) B cell-count distribution among the annotated cases after preprocessing and the train / test split. (B) The cell-level classifier prediction is shown on selected bi-parametric flow cytometry plots and a tSNE plot. Each cell is color-coded according to tumor likelihood obtained from the cell-level module. The upper panel represents a normal case, and the lower panel represents an abnormal case.

[0020] FIGs. 11 A-C: (A) A process is depicted. The manual annotations for all cases were obtained. The cases were then subjected to automated preprocessing for compensation, transformation, and B-cell gating using flowDensity. The cases were split into training and testing sets in a 70 / 30 split for normal and tumor groups (Subpanel 1). The cells from the training set were used to train a cell-level classifier (Subpanel 2). The process of generating a Normal Synthetic sample (Subpanel 3) and a Tumor Synthetic sample (Subpanel 4) is shown. 7500 synthetic training sets were generated using the training cells. The ratio of normal to tumor samples is 1 : 1. Each cell on those synthetic samples was passed through the Cell-level classifier and sorted in the descending order of tumor probability (Subpanel 5). The cells on the testing set were also passed through the-6-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009Cell-level module and sorted on the descending order of tumor probability (Subpanel 6). The synthetic training set was then used to train a sample-level classifier, which was then tested on the testing set. Panel B) The result from the Sample-level module, applied to the testing set, is shown as a Confusion matrix. Panel C) The ROC curve of the testing set predictions shows an AUC of 0.985[0021J FIG. 12 depicts a block diagram of a system for classifying samples based on flow cytometry data in accordance with an illustrative embodiment.[00221 FIG. 13 depicts a block diagram of a process to receive event datasets from flow cytometers in the system for classifying samples in accordance with an illustrative embodiment.

[0023] FIG. 14 depicts a block diagram of a process to train machine learning models in the system for classifying samples in accordance with an illustrative embodiment.

[0024] FIG. 15 depicts a block diagram of a process to determine probabilities of tumor in samples using the machine learning models in the system for classifying samples in accordance with an illustrative embodiment.[0025 J FIG. 16 depicts a block diagram of a process to generate outputs in the system for classifying samples in accordance with an illustrative embodiment.

[0026] FIG. 17 depicts a flow diagram of a method of classifying samples based on flow cytometry data in accordance with an illustrative embodiment.

[0027] FIG. 18 depicts a block diagram of a server system and a client computer system, in accordance with one or more implementations.DETAILED DESCRIPTION

[0028] Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for classifying samples based on flow cytometry data. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways, as the-7-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.

[0029] Section A describes deep learning based automatic detection of minimal residual disease in acute lymphoblastic leukemia / lymphoma.

[0030] Section B describes deep learning-based automatic detection of b-cell neoplasms in clinical flow cytometry.

[0031] Section C describes systems and methods of classifying samples based on flow cytometry data.

[0032] Section D describes a network environment and computing environment which may be useful for practicing various computing related embodiments described herein.A. Deep Learning Based Automatic Detection of Minimal Residual Disease in Acute Lymphoblastic Leukemia / Lymphoma

[0033] The evaluation of minimal residual disease (MRD) in B-lymphoblastic leukemia (B-ALL) using flow cytometry methods has been widely used in clinical practice. However, this analysis is resource-intensive, time-consuming, and relies on extensively trained personnel. To address these challenges, a stepwise deep learning model may be used (referred herein as Flow ARC) to automatically detect MRD in B-ALL. First, the celllevel module, a cell classifier, is used to extract suspicious cells from the CD 19+ B-cell population of a case. From these cells, the sample-level module, a sample classifier, then accurately detects the presence of MRD. Retrospectively from a clinical archive of 1681 cases from 333 patients, cell annotations were obtained on 137 B-ALL positive and 81 B- ALL negative cases by an experienced practicing pathologist.

[0034] Two Flow ARC models, Flow ARC -Manual and FlowARC-Auto, were trained based on manual or automated extraction of CD 19+ B-cells pre-analysis.FlowARC -Manual achieved 100% accuracy in identifying all 24 B-ALL positive and 14 B- ALL negative cases out of 38 clinical test cases. FlowARC-Auto correctly classified 239-8-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 out of 248 B-ALL positive and 229 out of 232 B-ALL negative cases out of 480 clinical test cases, achieving an accuracy of 97.5%, sensitivity of 96.4%, and specificity of 98.7%. In addition, Flow ARC includes an automated preprocessing pipeline, a visualization chart for tumor cell assessment and verification, and a quantification module for estimating the size of the tumor population. This groundbreaking MRD detection model for B-ALL surpasses existing models and has the potential to revolutionize MRD assessments by serving as an assistant tool for hematopathologists, resulting in significant cost reduction and faster turnaround time.1. Introduction

[0035] Minimal / measurable residual disease (MRD) serves as an independent prognostic metric, indicating the residual cancer cells that may persist post-treatment, even when the patient appears to be in remission. Traditionally, MRD is defined as any detectable disease below a 5% abnormal cell threshold that could not be identified by only morphological assessment. MRD detection is crucial for assessing the efficacy of treatment and guiding subsequent therapeutic strategies, especially for patients with hematopoietic malignancies such as acute lymphoblastic leukemia, acute myeloid leukemia, and multiple myeloma. Consequently, an accurate MRD detection workflow is vital for any clinical hematopathology laboratory.

[0036] MRD detection can be carried out using either multi -parameter flowcytometry (MFC) or PCR-based molecular techniques that provide assay detection sensitivities as low as 10'6. MFC allows for the simultaneous evaluation of multiple markers on individual cells, generating data on thousands to millions of cells from a single case1. Compared to molecular MRD diagnostics, MFC is faster, more cost-effective, and does not require access to the patient's diagnostic sample. However, its specificity and sensitivity were not as good as what can be achieved through molecular techniques historically2'4. In recent years, the performance of MFC analysis has been enhanced by incorporating additional concurrent markers, particularly in B-lymphoblastic leukemia (B- ALL), making it a preferred technique.-9-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0037] Regrettably, these advancements in MFC analysis have made it more labor- intensive and reliant on the experience of the analysts. Therefore, development of automated solutions for accurate MRD analysis using flow cytometry data is a critical clinical need to broaden the accessibility and availability of this technology.

[0038] Machine learning methods can automatically extract relevant patterns directly from the data and provide objective interpretation. Therefore, automation tools are widely utilized. In MFC analysis, ML-based tools can be used to automatically extract complex patterns from the vast datasets. This, in turn, enables the detection of subtle shifts in cell populations and the identification of unknown or rare cell populations. Unsupervised ML methods like flowSOM5, FlowMeans6, FLOCK7, and SWIFT8have demonstrated their ability to enhance cellular heterogeneity resolution and group similar cells by projecting data in high-dimensional space. However, these methods primarily focus on grouping cells with similar characteristics and have not yet demonstrated the capability to adequately account for rare cell populations, which are crucial for MRD detection.

[0039] Supervised ML methods that utilize the available diagnostic sample labels have been shown to improve the detection of the rare populations using MFC data. For example, CellCNN, a deep learning model adapted for cellular data, can identify disease- associated cell subsets at very low frequencies, sometimes as rare as 0.01% within small flow cytometry datasets9. Deep CNN10and Cell Scoring Neural Network (CSNN)11, like CellCNN, use the diagnostic sample labels to extract scores at the cell- level that are then aggregated for a sample-level prediction. These supervised models represent a significant advancement, compared to fully unsupervised models, in automated MRD detection for leukemias. However, their reported accuracies still fall short of the clinically acceptable level required for MRD detection.

[0040] Flow ARC may be used for accurate and automated MRD detection using flow cytometry data. In contrast to previous methods, Flow ARC uses both cell annotations and sample labels. It utilizes a multi-level classification strategy that combines cell classification with subsequent sample classification to improve MRD detection accuracy. Flow ARC was implemented and evaluated on the in-house B-ALL clinical archive containing 1681 flow cytometry cases.-10-4891 -3791 -9722.1Atty. Dkt. No.: 115872-30092. Results2.1 Study design and overview

[0041] A retrospective cohort of patient cases tested for B-ALL by an 8-color flow cytometry assay at the institution between 2015 and 2019 was gathered. Cases with multiple diagnosis were then filtered out resulting in a total of 1681 cases from 333 patients (Male: 184, Female: 149). The average patient age was 32.8 with standard deviation of 22.2. Flow ARC was applied to this clinical archive with 1150 normal cases and 531 B- ALL positive cases.

[0042] As Flow ARC requires at least a subset of initial dataset to have explicit cell level annotations, a sub-cohort of 218 cases from the clinical archive were manually annotated by an experienced pathologist to obtain their CD 19+ B-cell population and abnormal / tumor B-cell population. This manually annotated cohort includes 81 normal cases and 137 positive cases for BALL.

[0043] The workflow of Flow ARC includes four stages. In the first stage, an initial data preprocessing pipeline performs standardization of the flow data, followed by gating to obtain the CD19+ B cell population (FIG. 1 A). In the second stage, a cell-level module is trained to classify cells as either tumor or normal (FIG. IB). In the third stage, the cells within the preprocessed flow cases are ranked based on their probability of being tumorous and the input data was truncated to a set number of cells which are most likely tumorous (FIG.1C). Finally, in the fourth stage, a sample-level module is trained to classify these reordered and truncated cases for MRD diagnosis (FIG. ID).

[0044] In the preprocessing stage, two approaches are employed for CD 19+ B-cell gating: manual gating by pathologists and automated gating using flowDensity12, a published algorithm for density-based cell population identification. The manual B-cell gates are only available for the annotated cohort cases. Hence, for all clinical archive cases, automated B-cell gating is performed. Recognizing that different gating techniques may introduce systematic differences, two models were built: Flow ARC -Manual for manually gated patient cases from the annotated cohort, and FlowARC-Auto for auto- gated patient cases from the clinical archive.-11-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0045] The normal and B-ALL positive (tumor) cases in the annotated cohort are split (75% / 25%) into training / testing for Flow ARC -Manual. All cells within each of these sets have manual annotations and become the training / testing set for the cell-level module of Flow ARC -Manual. For Flow ARC- Auto, the normal cases in the clinical archive are split (65% / 35%) into normal training / testing cases, whereas all B-ALL positive cases lacking cell annotations (75%) are set as the B-ALL positive testing cases. For the cell-level module of Flow ARC- Auto, all cells within the auto-gated normal training / testing cases are annotated as normal training / testing, and the B-ALL positive cases with manual cell annotations are split (75% / 25%) into tumor training / testing cells.

[0046] The annotated cells for both Flow ARC models train and validate their respective cell- level module. Further, these annotated cells also generate synthetic datasets to train and validate their sample-level module. The synthetic datasets encompass thousands of cases (training / validation / testing: 15,000 / 6,000 / 9,000) with vast range of MRD levels that are generated by randomly mixing annotated (training / training / testing) cells in varying proportions to mimic the composition of real-world B-cell population. Half of the cases in each synthetic dataset are B-ALL positive which are further divided into four equal subsets with their total tumor cells in the ranges: 25 to 500 (low MRD), 500 to 25K (high MRD), 25K to 50K (low tumor), 50K to 500K (high tumor).

[0047] Additionally, it is noted that only a set number of cells per case are input to the sample- level module rather than the whole parent cell population. This number is considered a hyperparameter and has been set at 7500, as it yielded the optimal validation area under the ROC curve. Consequently, a minimum of 7500 B-cells per flow case is required for obtaining the sample-level module prediction.

[0048] The cell-level module is based on the ID ResNet-18 architecture, as implemented by Lang et al.13, without modification. The sample-level module is based on the ResNet- 10114 architecture without modification. Both module output two probabilities: normal and B-ALL positive / tumor, and are optimized using the area under the ROC curve of the validation set. To evaluate the performance of each module, metrics including the area under the ROC curve and the R2are used. Each module is generated five times using different random initializations. These iterations are then bootstrapped to obtain -12-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 a Confidence Interval (CI). The results obtained from both Flow ARC models are presented for comprehensive analysis.2.2 Flow ARC makes accurate automated prediction of MRD.[0049 J On the synthetic testing set of 9000 cases, Flow ARC -Manual achieves an area under the ROC curve, AUCROC of 0.988 (95% CI, 0.980-0.996). Using the model with the best validation AUCROC, an exemplary 99.0% accuracy, 99.0% sensitivity, and 99.0% specificity were achieved. The associated confusion matrix with 4457 true normal, 4457 true tumor, 43 false normal, and 43 false tumor cases is displayed in FIG. 2A. This Flow ARC -manual model is then tested on the 38 manually gated patient cases of the annotated cohort test set with at least 7500 B-cells. All cases were correctly identified, as shown by the confusion matrix in FIG. 2B.

[0050] On the synthetic testing set of 9000 cases, FlowARC-Auto achieves AUCROC of 0.994 (95% CI, 0.992-0.996). Using the model with the best validation AUCROC, a 98.8% accuracy, 98.5% sensitivity, and 99.1% specificity are achieved with the associated confusion matrix with 4459 true normal, 4431 true tumor, 41 false normal, and 69 false tumor cases displayed in FIG. 2C. This FlowARC-Auto model is then subsequently tested on the auto-gated original patient cases from the test set of the clinical archive. The test set encompasses a total of 480 cases — 232 normal and 248 as tumorcontaining a minimum of 7.5K cells within their B-cell gate. The model achieves an AUCROC of 0.990, with an accuracy of 97.5%, a sensitivity of 96.4%, and specificity of 98.7%. The corresponding confusion matrix with 229 true normal, 239 true tumor, 3 false normal, and 9 false tumor cases is presented in the FIG. 2D.

[0051] Both Flow ARC -Manual and FlowARC-Auto exhibit excellent performance in predicting MRD. To further evaluate the performance of Flow ARC, its prediction results with those of CellCNN were compared, the most sensitive previously reported model for detection rare abnormal cell detection in the literatures. For a fair comparison, use the exact B-cell population of the training and testing patient cases that were employed for FlowARC -Manual and FlowARC-Auto to train two CellCNN models: one for manually gated cases and another for auto-gated cases.-13-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0052] For the manually gated test cases, CellCNN achieves an AUCROC of 0.942 (95% CI, 0.922-0.963). For the auto-gated test cases, CellCNN achieves an AUCROC of 0.869 (95% CI, 0.860-0.879). A comparison of the AUCROC values between CellCNN and Flow ARC is presented in FIG. 2E. Clearly, Flow ARC shows a significant improvement compared to CellCNN.2.3 FlowARC shows capability to assist pathologists to visualize and verify suspicious cell population.

[0053] Flow cytometry data traditionally is visualized on many two-dimensional plots by demonstrating two makers on each plot. Pathologists diagnose MRD in B-ALL by visualizing a tumor cell cluster separated from the normal B-cells population on these plots. Here, similar plots were used to visualize and verify the most suspicious B-cell population within flow cases as predicted by the cell-level module. Dimension reduction is performed using the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm15to visualize the total high-dimensional information of all cells on a two dimensional plot. These CD 10 vs CD20, CD45 vs CD20, CD34 vs CD38, CD 10 vs CD38, CD34 vs CD45, and t-SNE comp- 1 vs comp-2 plots are grouped together to make a visualization chart.

[0054] The prediction performance of the cell-level module is important for this chart. The Flow ARC -Manual cell-level module achieves a test area under the ROC curve of 0.978 (95% CI, 0.978-0.978) on the test cells. Similarly, the FlowARC-Auto cell-level module achieves a test area under the ROC curve of 0.948 (95% CI, 0.948-0.948). These results show that both of the cell-level modules can separate normal and tumor cells well. As FlowARC-Auto can be applied to any patient case, the capability of the visualization chart is demonstrated using its cell-level module.

[0055] The FlowARC-Auto cell-level module with best area under the ROC curve on validation cells (random subset of training cells) is used on patient test cases of the clinical archive to generate their visualization chart. For true tumor cases, there is a distinct clustering of cells with very high tumor likelihood (darker cells shown in FIG. 3 A) in most of the plots. In true normal cases, even though there can be cells with higher tumor likelihood, their arrangement is relatively diffuse in most plots (FIG. 3B). With this-14-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 expected pattern, the visualization chart can be used as a prediction verification tool. Furthermore, the cases mentioned that with less than 7500 cells in B cell gate cannot be injected to the sample-level module. However, these cases can be evaluated by visualizing the clustering pattern for the true tumor cases and scattered pattern for the true normal cases. (FIG. 3C). This method allows us to be able to assign a prediction to the cases with fewer B cell which is not uncommon post therapy.2.4 Flow ARC shows capability to predict the tumor content.[0056J The enumeration of tumor cells within a positive flow cytometry case is an essential clinical parameter for disease management. To enhance the capability of Flow ARC in providing precise estimations of tumor cell prevalence, a quantification module may be designed focused on population estimation. The architecture of the quantification module is the exact ResNet-10114model used in the sample-level module, with the output changed to a number between 0 and 7500. This module is trained and validated on the B-ALL positive cases of the synthetic dataset, and is optimized using the validation coefficient of determination, R2.

[0057] To achieve the best performance, an integrated approach may be designed for cases with different tumor levels. Every B-ALL positive case is first analyzed by the quantification module. If the subsequent prediction is below 7200 cells, the total tumor burden is accepted. For the cases predicted to have more than 7200 cells, the number of cells predicted to be positive were aggregated by cell level model to provide an estimate of the total tumor burden

[0058] For the Flow ARC -Manual synthetic testing tumor set, R2is 0.998 (CFO.998 - 0.999), indicating an exceptional alignment between the true and predicted tumor population. Additionally, the mean absolute error (MAE) is found to be 0.012 (CFO.010 - 0.014), further attesting to the accuracy and precision of the quantification module. Using the model with best validation R2, the comprehensive evaluation of the true and predicted tumor population is plotted in FIG. 4A (left). For cases with predicted tumor population above 7200 cells, the estimation from the cell-level module achieves R2of 0.998 as shown in FIG. 4 A (right).-15-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0059] For the manually gated real-world test tumor cases, the quantification module with the best validation R2achieves, coefficient of determination, R2, of 0.922 and a mean absolute error (MAE) of 0.048, signifying a good estimation of tumor burden. The correlation between the true and predicted tumor population is graphically depicted in FIG.6 (left). For cases with predicted tumor population above 7200 cells, the estimation from the cell-level module achieves R2of 0.997 as shown in FIG. 4B (right).2.5 Flow ARC explainability analysis

[0060] Here, the inner workings of the cell-level and sample-level module were explored, and analyze their prediction pattern in more detail.

[0061] First, the classifications of the cell-level module (Flow ARC -Manual) are visualized in the set of marker vs marker plots using a random representative subset of the test cells in FIG. 5A. The visual representation employs color and dot size to reflect the classification outcomes from the Cell-level module, as shown in the legend at the bottomright of the figure. It becomes apparent upon examination that most misclassifications are localized around the boundaries separating the two cell populations.

[0062] For a more nuanced understanding of the operational mechanisms of the Cell-level module, the Shapley Additive exPlanations (SHAP) algorithm16a widely used method for elucidating the decision-making processes of deep-learning models were applied. The SHAP algorithm dissects the model's predictions by pinpointing the contributing input features and appraising their relative significance in shaping the confidence of the classification. By utilizing SHAP values that are rooted in game theory, the algorithm imparts a quantifiable importance to each feature within the model. Features with positive SHAP values positively influence tumor predictions, while those with negative values have a negative influence.

[0063] The SHAP summary plot offers a synthesis of feature importance and their respective effects. On this plot, each data point signifies a SHAP value for a feature in relation to a specific instance (cell). The summary plot, presented in FIG. 7, showcases these SHAP values for two distinct scenarios: Tumor prediction and Normal prediction. The features, or cell markers, are arrayed in rows according to their significance, with -16-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009SHAP values that overlap being jittered along the y-axis. The color bar encodes the input value of the feature. From the plot, a lower forward scatter area (FSC-A) coupled with a higher forward scatter height (FSC-H) is more indicative of a normal cell prediction, while lower expression of CD38 is more suggestive of a tumor cell prediction may be discerned.

[0064] The sample-level module is a robust neural network capable of extracting rich and valuable features from its input. These extracted features are then used to classify each case in the model's final layer. Understanding how these features differ for each class may provide us insight into the model itself. To this end, t-SNE is applied to project the penultimate 2048-dimensional feature layer into a lower-dimensional space and plotted in FIG. 7. The normal and tumor cases are well separated, signifying the Sample-level module's capacity to accurately differentiate and predict the two distinct classes.3. Discussion

[0065] Flow ARC represents a significant advancement in the automatic and accurate detection of minimal residual disease (MRD) using flow cytometry data. The distinguishing feature of this model lies in its multi-step approach, wherein the challenge posed by a large number of randomly ordered cells within each flow case is first addressed by a ID cell- level module. This module truncates each case into an ordered table, consisting of the most tumor-like cell subset, which is then thoroughly analyzed by a 2D sample-level module to provide an accurate prediction. This approach overcomes the limitations of previous works, which either rely on unsupervised clustering methods that may miss crucial rare cell populations or use neural networks to extract single-cell information without effectively utilizing the 2D information within the cases. In contrast, Flow ARC leverages the 2D information among the relevant subset of cells to accurately predict the presence of minimal residual disease.

[0066] FlowARC comprises a ID cell-level module trained on manually annotated cells and a 2D sample-level module trained on annotated cases. Both modules make use of the powerful ResNet architecture, a state-of-the-art deep learning model. Additionally, FlowARC incorporates a clinically relevant estimator for the tumor population. As a supervised deep learning model, FlowARC requires a substantial amount of annotated data-17-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 for training, including millions of annotated cells and thousands of annotated cases. However, the process of obtaining these annotations can be time-consuming and may present a bottleneck. To address this issue, Flow ARC can be trained using synthetic cases generated from a smaller set of annotated cases. These synthetic cases closely mimic real- world cases and cover a wide range of tumor population sizes, including low levels of MRD and advanced cases of the disease. This approach significantly reduces the training requirements and enhances the accessibility of Flow ARC.

[0067] Using a clinical archival dataset of B-ALL cases comprising 1681 cases, a Flow ARC model was developed for the automatic detection of MRD. The model achieved highly accurate results on real-world patient cases, demonstrating an accuracy of 97% on a test set of 480 cases. The model was trained on manual annotations of 103 B-ALL cases, while the remaining training set was auto-gated. These results highlight that even a small set of annotated flow cases for a specific disease panel is sufficient to train a Flow ARC model with clinical-grade performance. This has the potential to significantly reduce the turnaround time for flow MRD analysis, ultimately leading to improved patient outcomes.

[0068] It is important to note that MRD detection using flow cytometry panels requires a minimum prevalence of tumor cells, known as the sensitivity of the flow test. In the synthetic case generation process, a minimum tumor population of 25 cells was set, which corresponds to a test sensitivity of up to 10'5depending on the number of cells tested. It should be acknowledged that in Flow ARC, the false negative rate tends to increase as the tumor population decreases. Therefore, this minimum tumor population can be adjusted to favor either higher test sensitivity or a lower false negative rate in Flow ARC. Additionally, the model can be fine-tuned during training to achieve a lower false negative rate by assigning greater weight to the tumor cases. This flexibility to adapt to different use cases is a significant strength of Flow ARC.

[0069] Furthermore, Flow ARC is a generic model that can be trained to detect and estimate MRD for any leukemia flow cytometry test. It can also be applied to modem flow cytometer panels with a larger number of features, including mass flow cytometers.Consequently, Flow ARC has the potential to automate MRD detection workflows and find utility in various clinical settings.-18-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0070] The proficiency of the Flow ARC system was successfully highlighted in pinpointing aberrant precursor B cells within the confines of the CD 19+ B cell population. A pivotal need for the identification of the infrequent CD 19- B-cell acute lymphoblastic leukemia (B-ALL) phenotype was recognized, which has become increasingly relevant due to the escalated adoption of anti-CD19 immunotherapies in the treatment of patients with relapsed and refractory B-cell acute leukemias. These therapies can lead to the emergence of CD 19- leukemic cells as a mechanism of disease escape, necessitating a reliable method for their detection. To address this clinical challenge, an enhanced B-ALL flow cytometry assay was developed, which has been thoughtfully designed to incorporate the markers CD22 and CD24. This strategic inclusion is intended to enable the robust detection of CD 19- leukemic cells, thereby closing a critical diagnostic gap and offering a more targeted approach to patient assessment and treatment planning. Looking to the future, the diagnostic capabilities are planned to be further refined by extending the application of the general model, Flow ARC, to encompass a CD 19- assay. This test may be specifically devised to identify B-cell malignancies that have evaded CD 19- targeted treatments. By adapting and applying the versatile Flow ARC algorithm to this novel assay, its sensitivity and specificity may be enhanced. This forward-thinking approach aims to provide a more comprehensive diagnostic tool, ensuring that even as treatment strategies evolve, the ability to detect and monitor B-cell malignancies remains at the cutting edge of hematologic diagnostics.4. Methods4.1 Data collection

[0071] A total of 1681 bone marrow biopsy specimen tested exclusively for B- lymphoblastic leukemia using flow cytometry were included in this retrospective study. The B-ALL flow assay includes the following surface markers: CD20, CD34, CD10, CD33, CD58, CD45, CD19, and CD38. The forward scattering channels (FSC-H and FSC-A), side scattering channels (SSC-H and SSC-A) are also measured by the flow cytometer. For each specimen, a flow cytometer measured these 12 channels for all cells and saved the results and metadata as an FCS file.-19-4891 -3791 -9722.1Atty. Dkt. No.: 115872-30094.2 Initial data preprocessing

[0072] The preprocessing pipeline is created using R programming language. The FCS files are read and subject to compensation and transformation using flowCore package. Here, compensation eliminates fluorescence spillover from other dyes for each marker, while transformation condenses the data's large dynamic range into a standardized range, reducing the impact of intrinsic variance in fluorescence signals and aiding in the visualization of discrete positive and negative populations. Compensation is applied using the spillover matrix obtained from the FCS file metadata. Subsequently, the SSC- H channel is log-transformed, the other three scattering channels undergo linear transformation, and the eight fluorescence channels are transformed with the logicle function. This initial preprocessing of the FCS files is followed by B-cell population gating.4.3 B-cell population gating

[0073] The CD 19+ B-cell gating is either be done manually by a trained pathologist or using automated density-based gating tool, flowDensity.

[0074] The manual gating procedure in this flow assay follows a series of steps. Firstly, doublets and margin events are removed using the FSC-A vs FSC-H plot. Next, a viable mononuclear cell population is selected using the SSC-H vs FSC-A plot. Finally, the CD19+ B-cell population is identified using the SSC-H vs CD19 plot.

[0075] In flowDensity, this manual gating process was replicated for automated gating. It requires the implementation of a gating strategy with various thresholds based on the density distribution of markers, a four-step approach was adopted, as illustrated in FIG. 7. Step 1 involves the exclusion of margin events and doublets. Steps 2 and 3 focus on selecting the viable mononuclear population. In Step 4, the CD 19+ B-cell population is specifically chosen. Exclusion and selection of populations can be improved by changing the flowDensity parameter thresholds. Using 15 randomly chosen cases with available manual gate, the parameter thresholds based on visual comparison with the manual gate was iteratively optimized. Thresholds that generate comparable results between automated and manual gates for 10 out of 15 cases were able to be optimized. These parameters are frozen-20-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 and incorporated into a single R script. This script can then be applied to the flow cases, ensuring consistent and efficient analysis.4.4 Datasets[0076 J Developing a Flow ARC model requires a group of patient cases and set of labeled cells. This study trains two specific Flow ARC models for manually gated and autogated cases. Flow ARC and Flow ARC -Manual are trained and tested on the clinical archive and the annotated cohort respectively. All 1681 cases from 333 patients make up the clinical archive dataset. A sub-cohort of 218 cases from the clinical archive make up the annotated cohort.

[0077] The annotated cohort includes 81 normal (B-ALL negative) cases and 137 tumor (B-ALL positive) cases. This cohort is manually gated by an experienced hematopathologist to obtain their B-cells. Then, abnormal blasts are extracted manually from the B-ALL positive cases are annotated as tumor cells while B-cells of normal cases are annotated as normal cells. A random 75% / 25% train / test split was performed, separately for the tumor and normal patient cases. The annotated cells within these train / test cases are the train / test labeled cells. The visualized breakdown of this annotated cohort is shown in FIG. 6. The cell distribution of the labeled cell set of this cohort is shown in Supplementary figure, panel D.

[0078] The clinical archive contains 1150 normal cases and 531 tumor cases. This cohort is auto-gated using the flowDensity script to obtain their B-cells. Then, the B-cells of the normal cases are annotated as normal cells. A random 65% / 35% train / test split of the normal patient cases was performed. As the extraction of abnormal blasts must be done manually, the previous manually gated tumor train / test cases was used for the annotated tumor cells. Now, the auto-gated tumor cases from this dataset, excluding the tumor cases shared with annotated cohort, are added to the patient test cases. The visualized breakdown of this clinical archive is also shown in FIG. 6. The cell distribution of the labeled cell set of this archive is also shown in Supplementary figure, panel D.4.5 Synthetic case generation-21-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0079] Flow ARC utilizes synthetic datasets with huge number of synthetic cases to train, test and validate the sample-level module. These cases have a broad-range of MRD levels and mimic the parent cell-population (CD 19+ B-cells). The training, testing and validation synthetic datasets are generated using the labeled train, test, and train cases / cells respectively.

[0080] In a synthetic dataset, half of the cases are tumor synthetic cases and half of the tumor cases are MRD cases. Each MRD and advanced tumor cases are further divided into low and high categories. Low MRD and high MRD cases contain 25 up to 500 and 500 up to 25K tumor cells. Similarly, low tumor and high tumor cases contain 25K up to 50K and 50K up to 500K tumor cells. This categorical separation of the synthetic dataset is shown in FIG. 8A, panel A.

[0081] For the normal synthetic cases, only normal B-cell population is needed. A subset of normal cases from the labeled set with combined cell count exceeding 10K was randomly selected. These selected cases are then coalesced into a single table. From this table, A random cell subset with cell count (N) randomly chosen between 10K and 50K was randomly chosen. This subset is then saved as a normal synthetic case. This process, which aims to replicate the variety and variable cell count observed in real B- cell populations, is iterated until the desired number of normal synthetic cases are created. The process of normal synthetic case generation is illustrated in FIG. 8B.

[0082] For the tumor synthetic cases, both normal and tumor B-cell populations are required, the normal population was generated, with cell count (NN) randomly chosen between 10K and 50K, in the exact way as the normal synthetic cases. For the abnormal cell population, there are four different categories: Low MRD, high MRD, low tumor and high tumor. For each category, A cell count (NA) from their category- specific range was randomly chosen. Now, A tumor case with number of cells greater than is randomly selected, but not exceeding, three times NA. From this chosen tumor case, NA number of tumor cells are randomly chosen to get the tumor population. Both normal and abnormal populations are then combined and saved as a tumor synthetic case (N=NN+NA). This process is iterated for each category until the desired number of tumor synthetic cases are created. The process of tumor synthetic case generation is illustrated in FIG. 7.-22-4891 -3791 -9722.1Atty. Dkt. No.: 115872-30094.6 Cell-level module architecture and training

[0083] The architecture of the cell-level module is based on ResNet-18 model, adapted to ID where every 2D layer is replaced with its ID pendant. This ID ResNet-18 implementation is not pretrained and can handle any input size without modification. So, no further modifications are made for its utilization in this work. The cell-level module is trained on the annotated normal and tumor training cells, with an input size of (1 x 12) representing the 12 markers of each cell, and an output of (2 x 1) providing probabilities for the normal and tumor classes.

[0084] During training, a combination of the cross-entropy (CE) loss and Orthogonal Projection Loss (OPL) is employed as the loss function.

[0085] The cell-level module is trained on the labeled training cells for 100 epochs with a batch size of 800 cells. As a validation set, a subset of 2,000,000 training cells (1 : 1 normal and tumor cells) is randomly selected. A stochastic gradient descent optimizer with a momentum of 0.9 and a weight decay of le-3 is used. A decaying learning rate, starting at 0.1 and decreasing by 10% at epoch 5, 10, 30, 50, and 70 is utilized. The hyperparameter selection is based on the best validation area under the ROC curve. The model at 100 epoch is saved as the trained model.

[0086] Different cell-level modules for Flow ARC and Flow ARC -Manual were trained. The labeled cells for clinical archive contains significantly more normal cells compared to the labeled cells for annotated cohort, see Supplementary Figure, panel D. To mitigate this bias towards the normal cells, An additional imbalanced dataset sampler for Flow ARC was employed that rebalances the class distributions at every batch during training.4.7 Sample-level module architecture and training

[0087] The sample-level module architecture is based on the ResNet-101 model with random weight initialization. This ResNet-101 model can handle any three channel 2D input. No further modifications are made for its utilization in this work.-23-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0088] The reordered and truncated synthetic cases has input size of (1 x 7500 x 12). This input matrix is Fast Fourier transformed. The input matrix, alongside the real and imaginary components of its Fourier transform, are compiled into a three-channel input of size (3 x 7500 x 12) compatible with ResNet-101. The model output is of size (2 x 1) providing probabilities for the normal and tumor classes.

[0089] Different sample-level modules for Flow ARC and Flow ARC- Manual were trained. The sample-level modules are trained on the synthetic training cases for 50 epochs with a batch size of 10 cases using the combined CE + OPL loss function. A stochastic gradient descent optimizer with a momentum of 0.9 was used. A decaying learning rate, starting at 0.05 and decreasing by 10% at epoch 5, 10, 15, 20, 25, 30 and 35 is utilized. Only for Flow ARC -Manual, A weight decay of le-4 in the optimizer was added. The hyperparameter selection is based on the best validation area under the ROC curve. The model at 50 epoch is saved as the trained model.4.8 Tumor estimator module architecture and training

[0090] The architecture of the tumor estimator module is the exact ResNet-101 model used in the sample-level module, with the output size changed from (2 x 1) to (1 x 1). This single value between 0 and 1 is further multiplied by 7500 as the final output.

[0091] Tumor estimator modules are trained for the Flow ARC -Manual. The tumor estimator module is trained on the synthetic B-ALL positive training cases for 50 epochs with a batch size of 10 cases using the huber loss function. A stochastic gradient descent optimizer with a momentum of 0.9 is used. A decaying learning rate, starting at 0.05 and decreasing by 10% at epoch 5, 10, 15, 20, 25, 30 and 35 is utilized. The hyperparameter selection is based on the best validation area under the ROC curve. The model at 50 epoch is saved as the trained model.4.9 CellCNN

[0092] The CellCNN source code for Python 3 is obtained from the CellCNN github. To make it comparable to FlowARC, CellCNN is trained and evaluated on the B- cell population from patient cases. CellCNN is operated in the 'outlier' subset selection-24-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 mode, which is recommended for MRD detection. To optimize the hyperparameters of CellCNN, the integrated CellCNN hyperparameter exploration function was utilized. This function selects the model that achieves the best predictive performance on the autogenerated validation cases. Similar to Flow ARC, each CellCNN model is generated five times with different random initializations, and these iterations are then subjected to bootstrapping to obtain a Confidence Interval (CI).References[0093| 1 Herzenberg, L. A., Tung, J., Moore, W. A., Herzenberg, L. A. &Parks, D. R. Interpreting flow cytometry data: a guide for the perplexed. Nat Immunol 7, 681- 685 (2006). https: / / doi.org: 10.1038 / ni0706-681

[0094] 2 Thorn, I. et al. Minimal residual disease assessment in childhood acute lymphoblastic leukaemia: a Swedish multi-centre study comparing real-time polymerase chain reaction and multicolour flow cytometry. Br J Haematol 152, 743-753 (2011). https: / / doi.org: 10.1111 / j.1365-2141.2010.08456.x

[0095] 3 Denys, B. et al. Improved flow cytometric detection of minimal residual disease in childhood acute lymphoblastic leukemia. Leukemia 27, 635-641 (2013). https: / / doi.org: 10.1038 / leu.2012.231

[0096] 4 Gaipa, G. et al. Time point-dependent concordance of flow cytometry and real- time quantitative polymerase chain reaction for minimal residual disease detection in childhood acute lymphoblastic leukemia. Haematologica 97, 1582- 1593 (2012). https: / / doi.org: 10.3324 / haematol.2011.060426[0097[ 5 Van Gassen, S. et al. FlowSOM: Using self-organizing maps for visualization and interpretation of cytometry data. Cytometry A 87, 636-645 (2015). https: / / doi.org: 10.1002 / cyto.a.22625

[0098] 6 Aghaeepour, N., Nikolic, R., Hoos, H. H. & Brinkman, R. R.Rapid cell population identification in flow cytometry data. Cytometry A 79, 6-13 (2011). https: / / doi.org: 10.1002 / cyto.a.21007-25-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0099] 7 Qian, Y. et al. Elucidation of seventeen human peripheral blood B- cell subsets and quantification of the tetanus response using a density-based method for the automated identification of cell populations in multidimensional flow cytometry data.Cytometry B Clin Cytom 78 Suppl 1, S69-82 (2010). https: / / doi.org: 10.1002 / cyto.b.20554

[0100] 8 Naim, I. el al. SWIFT-scalable clustering for automated identification of rare cell populations in large, high-dimensional flow cytometry datasets, part 1 : algorithm design. Cytometry A 85, 408-421 (2014). https: / / doi.org: 10.1002 / cyto.a.22446

[0101] 9 Arvaniti, E. & Claassen, M. Sensitive detection of rare disease- associated cell subsets via representation learning. Nat Commun 8, 14825 (2017). https: / / doi.org: 10.1038 / ncommsl4825

[0102] 10 Hu, Z., Tang, A., Singh, J., Bhattacharya, S. & Butte, A. J. A robust and interpretable end-to-end deep learning model for cytometry data. Proc Natl Acad Sci U SA 117, 21373-21380 (2020). https: / / doi.org: 10.1073 / pnas.2003026117

[0103] 11 Robles, E. E. et al. A cell-level discriminative neural network model for diagnosis of blood cancers. Bioinformatics 39 (2023). https: / / doi.org: 10.1093 / bioinformatics / btad585

[0104] 12 Malek, M. et al. flowDensity: reproducing manual gating of flow cytometry data by automated density-based cell population identification. Bioinformatics 31, 606- 607 (2015). https: / / doi.org: 10.1093 / bioinformatics / btu677

[0105] 13 Lang, N. et al. Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles. Remote Sensing of Environment 268, 112760 (2022).

[0106] 14 He, K., Zhang, X., Ren, S. & Sun, J. in Proceedings of WIQ IEEE conference on computer vision and pattern recognition. 770-778.

[0107] 15 Hinton, L. v. d. M. a. G. Visualizing Data using t-SNE. Journal of Machine Learning Research 9, 2579-2605 (2008).-26-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0108] 16 Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. Advances in neural information processing systems 30 (2017).B. Deep Learning-Based Automatic Detection of B-Cell Neoplasms In Clinical Flow Cytometry1. Introduction

[0109] Flow cytometry plays a pivotal role in diagnosing B-cell lymphoma in clinical settings. However, its data analysis is labor-intensive and requires highly specialized personnel. Rapid flow cytometry results are in many clinical scenarios, assisting in immunohistochemical staining, molecular / cytogenetic testing, and patient management. While automated flow cytometry data analysis has become more common in research, clinical -grade applications remain limited due to the inability of existing software to detect clinically significant abnormal populations reliably. This disclosure addresses this gap by developing a deep learning model designed to automate the detection of abnormal B- cell populations in clinical flow cytometry data.2. Design

[0110] A retrospective cohort of 1109 tumor-positive and 823 tumor-negative cases from 1742 unique patients tested for B-cell flow cytometry were collected. Manual celllevel annotations were gathered for each case. Data preprocessing included compensation, transformation, and automated gating for B cells. The dataset was divided into training (70%) and testing (30%) groups for both tumor and normal cases. A cell-level classifier was trained to identify tumor cells using the manually annotated training dataset, followed by a case-level classifier trained with synthetic sample sets generated by randomly mixing normal and tumor cells from the training dataset in varying proportions, mimicking authentic clinical samples. The final model was evaluated using the testing dataset. (FIG.11 A)3. Results-27-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0111] There are 53 million abnormal and 11 million normal B cells for training, 22 million abnormal and 5 million normal cells for testing (FIG. 10A). At the individual cell level, the model achieved an area under the ROC curve (AUC) of 0.953 on the testing set of 27 million cells. The abnormal cells predicted by the model form clusters in the abnormal sample while visualizing by a dimension reduction plot. (FIG. 10B) When applied to a test dataset of 525 clinical cases, the model demonstrated an AUC of 0.985. (FIG. 1 IB). The confusion matrix for case-level classification showed correct prediction on 286 normal and 208 tumor cases, 12 false negative and 19 false positive predictions. (FIG. 11C)4. Conclusion

[0112] An automated flow cytometry analysis tool may be used for B-cell neoplasms. This model has can be incorporated into clinical workflow, vastly reduce the turnaround time, and increase the accuracy of flow cytometry testing.C. Systems and Methods for Classifying Samples Based on Flow Cytometry Data

[0113] Referring now to FIG. 12, depicted is a block diagram of a system 100 for classifying samples based on flow cytometry data. In overview, the system 100 may include at least one data processing system 105, at least one flow cytometer 110, at least one administrative device 115, and at least one database 120, among others, communicatively coupled with one another via at least one network 125. The data processing system 105 may include at least one dataset indexer 130, at least one model trainer 135, at least one event evaluator 140, at least one sample evaluator 145, at least one output generator 150, at least one cell-level machine learning (ML) model 160, and at least one sample-level ML model 165, among others. Each of the components in the system 100 (such as the data processing system 105 and its subcomponents and the administrative device 115) as detailed herein may be implemented using hardware (e.g., one or more processors coupled with memory), or a combination of hardware and software as detailed herein in Section C. The system 100 may be used to implement the functionalities described herein in Section A and B, such as the Flow ARC.-28-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0114] In further detail, the data processing system 105 may be any computing device including one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 105 may be associated with an entity to process tomograms and data associated with subjects with glioma to provide output. The data processing system 105 can be in communication with the flow cytometer 110, the administrative device 115, the database 120, and other devices, via the network 125. The data processing system 105 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a branch office, or a site at which one or more servers corresponding to the data processing system 105 is situated

[0115] The data processing system 105 may include one or more modules, components, or subsystems to perform the various processes and tasks described herein. The dataset indexer 130 may receive datasets acquired for a sample from the flow cytometer 110. The model trainer 135 may initialize, train, and establish the cell-level ML model 160 and the sample level ML model 165. The event evaluator 140 may use the cell-level ML model 160 to determine probabilities of tumor associated with cancer for a given cell type. The sample evaluator 145 may use the sample-level ML model 165 to determine a probability of tumor associated with cancer in a given sample. The output generator 150 may produce output information based on the determinations using the cell-level ML model 160 and the sample-level ML model 165.

[0116] The cell-level ML model 160 may include any type of machine learning (ML) or artificial intelligence (Al) architecture to process flow cytometry datasets. The architecture for the cell-level ML model 160 may include, for example, a deep learning artificial neural network (ANN) (e.g., an encoder of a convolutional neural network (CNN)), a Markov chain, a support vector machine (SVM), a clustering algorithm, a Bayesian classifier, or a decision tree, among others. In general, the cell-level ML model 160 may include at least one input, at least one output, and a set of weights relating the input with the output. The input may include flow cytometry datasets for a given cell type. The output may include a value indicating a probability of tumor associated with cancer for-29-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 the given cell type. The set of weights may be arranged in accordance with the ML or Al architecture.

[0117] The sample-level ML model 165 may include any type of machine learning (ML) or artificial intelligence (Al) architecture. The sample-level ML model 165 may include any type of machine learning (ML) or artificial intelligence (Al) architecture to process flow cytometry datasets. The architecture for the cell-level ML model 160 may include, for example, a deep learning artificial neural network (ANN) (e.g., an encoder of a convolutional neural network (CNN)), a Markov chain, a support vector machine (SVM), a clustering algorithm, a Bayesian classifier, or a decision tree, among others. In general, the sample-level ML model 165 may include at least one input, at least one output, and a set of weights relating the input with the output. The input may include flow cytometry datasets for a given sample with multiple cell types. The output may include a value indicating a probability of tumor associated with cancer for the sample. The set of weights may be arranged in accordance with the ML or Al architecture.

[0118] The flow cytometer 110 may be an instrument or device to analyze properties of cells within a given sample. The flow cytometer 110 may include a fluidics component, an optics component, and a processing component. The fluidics component may transport cells from the sample into a single stream (or file) for analysis within the flow cytometer 110. The fluidic component may include a sheath fluid to surround the cells of the sample to focus the cells into a narrow stream. The fluidics component may include a flow chamber in which the cells are funneled into the stream of cells at a controlled rate and pressure. In some embodiments, the fluidics component may include a cell sorter to select B-cells for analysis from the sample. The cell sorter may be, for example, in accordance with fluorescence-activated cell sorting, to physically arrange and separate cells based on characteristics, such as size, granularity, or fluorescence, among others.

[0119] The optics component may include a light source to illuminate the cells in the stream and a set of photodetectors to collect signals as the light scatted about the cells. The light source may emit a laser at any frequencies (e.g., 488 nm for blue, 633 nm for red, and 405 nm for violet, within 10%). The set of detectors may include: a forward scatter (FSC) detector to measure amount of light scattered in a forward direction relative to the -30-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 flow of cells; a side scatter detector (SSC) to measure amount of light scattered in an oblique direction (e.g., 90 degrees, within 10%) relative to the flow of cells; and a fluorescence detector (FL) to measure light of a particular wavelength emitted by fluorescent agent (e.g., a fluorescent dye) in the cells.

[0120] In the flow cytometer 110, the processing component may convert the light detected or acquired via the detectors of the optics component to datasets to be processed by the data processing system 105. The processing component may receive the electric signals converted by the set of photodetectors from the light acquired by the optics component.The processing component may include an amplifier to amplify the electric signals from the photodetector. The processing component may perform analog-to-digital conversion (ADC), and generate the dataset using the digitized and quantized flow cytometry data. The dataset may be stored and maintained in accordance with flow cytometer standard (FCS) or comma-separated values (CSV) file format. The flow cytometer 110 may be in communication with the data processing system 105, the administrative device 115, the database 120, and other devices, via the network 125. The flow cytometer 110 may be associated with an entity (e.g., a clinician) examining the subject for cancer or a vendor providing flow cytometry data for such an entity.

[0121] The administrative device 115 (sometimes herein referred to as an end user computing device) may be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The administrative device 115 may be in communication with the data processing system 105, the flow cytometer 110, and the database 120 via the network 125. The administrative device 115 may have at least one display. The administrative device 115 may be associated with an entity (e.g., a clinician) examining the subject for cancer. The display may present information about the subject provided by the data processing system 105.

[0122] The database 120 may store and maintain various resources and data associated with the data processing system 105, the flow cytometer 110, and the administrative device 115, among others. The database 120 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The -31-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 database 120 may be in communication with the data processing system 105, the flow cytometer 110, and the administrative device 115, via the network 125. While running various operations, the data processing system 105, the flow cytometer 110, and the administrative device 115 may access the database 120 to retrieve identified data therefrom. The data processing system 105, the flow cytometer 110, and the administrative device 115 may also write data onto the database 120 from running such operations.[01231 Referring now to FIG. 13, depicted is a block diagram of a process 200 to receive event datasets from flow cytometers in the system for classifying samples. The process 200 may include or correspond to operations performed in the system 100 to acquire flow cytometry data. Under the process 200, the flow cytometer 110 may carry out, execute, or otherwise perform flow cytometry on at least one sample 210 obtained from at least subject 205. The subject 205 may be a human or animal subject, among others. The subject 205 may be at risk of or diagnosed with cancer. The subject 205 may be under evaluation for cancer. For instance, the subject 205 may be under evaluation for minimum residual disease (MRD) of cancer, subsequent to administration of anti-cancer therapy. The subject 205 may be also evaluation for presence of neoplasms (e.g., abnormal growths of tissue in the body), and may be under evaluation for treatment for malignant or precancerous neoplasms.

[0124] In some embodiments, the cancer may include leukemia. The leukemia may include, for example, acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL), among others. The sample 210 may be obtained from any organ in the subject 205 to evaluate the subject 205 for leukemia. The organ for the sample 210 may include, for example, bone marrow, blood, lymph node, spleen, or spinal fluid, among others. The sample 210 may include any number of white cell blood types. The white cell blood types may include, for example, CD10+, CD19+, CD20+, CD34+, CD38+, or CD45+, among others. For instance, the sample 210 may include B lymphocytes (B -cells) with protein markers corresponding to any one or more of the white blood cell types, among others.-32-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0125] In some embodiments, the cancer or tumor is a carcinoma, sarcoma, a melanoma, or a hematopoietic cancer. In some embodiments, the cancer or tumor is selected from among adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin's disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, non-Hodgkin's lymphoma, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.

[0126] To perform flow cytometry, the flow cytometer 110 may pass cells from the sample 210 through a flow chamber. In some embodiments, the flow cytometer 110 may carry out cell sorting (e.g., using fluorescence-activated cell sorting (FACS)) on the cells of the sample 210. For instance, the flow cytometer 110 may perform cell sorting on the cells to select the B-cells to evaluate for cancer (e.g., leukemia or neoplasm), while excluding non-B-cells (e.g., cells unrelated to leukemia or neoplasm). The flow cytometer 110 may illuminate the cells in the flow cytometer with a light source. The flow cytometer 110 may use the set of photodetectors to acquire light interacting with the cells of the sample 210 into a set of flow channels 220A-N (hereinafter referred generally as flow channels 220). The set of flow channels 220 may correspond to the set of photodetectors in the flow cytometer 110. Each flow channel 220 may correspond to a different type of light signal, such as forward scatter channel height (FSC-H), forward scatter channel area (FSC-A), side scatter channel height (SSC-H), side scatter channel area (SSC-A), and any number of fluorescence channel (e.g., green, yellow, orange, or red fluorescence), among others.

[0127] In performing the flow cytometer, the flow cytometer 110 may create, produce, or otherwise generate at least one dataset 225 based on the light detected from the cells of the sample 210. The dataset 225 may identify or include a set of event sequences 230A-N (hereinafter generally referred to as set of event sequences 230) for the set of flow-33-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 channels 220. Each event sequence 230 may correspond to a respective flow channel 220. Each event sequence 230 may identify or include a set of events 235A-N (hereinafter generally referred to events 235). The events 235 may correspond to the cells detected in the flow channel 220 associated with the event sequence 230. Each event sequence 230 may include or identify cells in the flow channel 220. The cells may correspond to (e.g., be marked with) at least one of the white blood cell types in the sample 210. In each event sequence 230, the set of event 235 may be arranged in an initial order. The initial order may, for example, correspond to an order of measurement or acquisition by the flow cytometer 110.

[0128] Within the event sequence 230, each event 235 may correspond to an occurrence or a detection of at least one cell within a sample interval. The event sequence 230 for a given flow channel 220 may be generated or acquired at a sampling rate (e.g., at a rate of 10 to 100 MHz). The event sequence 230 may have any number of events over a given period of time, ranging between 500 events (or cells) to 50,000 events (or cells) per second. The dataset 225 may be generated as one or more files, such as using flow cytometer standard (FCS) or comma-separated values (CSV) file format. In some embodiments, the dataset 225 may include or identify metadata, such as an identifier for the subject 205, an identifier of the sample 210, and date and time of acquisition of the flow cytometry data, among others. The flow cytometer 110 may send, transmit, or otherwise provide the dataset 225 to the data processing system 105 or the database 120 for storage.

[0129] The dataset indexer 130 on the data processing system 105 may retrieve, identify, or otherwise receive the dataset 225 from the flow cytometer 110. With receipt from the flow cytometer 110, the dataset indexer 130 may process or parse the dataset 225 to extract or identify the set of event sequences 230. The dataset indexer 130 may identify the set of event sequences 230 by the corresponding set of flow channels 220, such as the forward scatter channel, side scatter channel, and the fluorescence channel, among others. In some embodiments, the dataset indexer 130 may determine or identify the initial order of the events in the set of event sequences 230 from the dataset 225. The dataset indexer 130 may store and maintain the dataset 225 the database 120.-34-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0130] Referring now to FIG. 14, depicted is a block diagram of a process 300 to train machine learning models in the system for classifying samples. The process 300 may correspond to or include operations performed in the system 100 to initial, train, and establish machine learning models used to evaluate the samples for cancer. The process 300 may be performed or carried out independently of the acquisition of the new flow cytometry data. Under the process 300, the model trainer 135 may initialize, train, and establish the cell-level ML model 160 and the sample-level ML model 165. To initialize, the model trainer 135 may assign or set the set of weights in each of the cell-level ML model 160 and the sample-level ML model 165 to a defined value (e.g., a random value). In addition, the model trainer 135 on the data processing system 105 may retrieve, obtain, or otherwise identify a set of training samples 305A-N (hereinafter generally as training samples 305). The set of training samples 305 may be training data used to train the cell-level ML model 160 and the sample-level ML model 165.

[0131] Each training sample 305 may identify or include a sample dataset 310. The sample dataset 310 may identify or include a set of event sequences 315A-N (hereinafter generally referred to as set of event sequences 315) for a set of flow channels (e.g., same as or different from the flow channels 220). Each event sequence 315 may correspond to a respective flow channel. Each event sequence 315 may include or identify cells in the flow channel. The cells may correspond to (e.g., be marked with) at least one of the white blood cell types in a respective biological sample.

[0132] Each event sequence 315 may identify or include a set of events 320A-N (hereinafter generally referred to events 320). The events 320 may correspond to the cells detected in the flow channel associated with the event sequence. In each event sequence 315, the set of event 320 may be arranged in an initial order. The initial order may, for example, correspond to an order of measurement or acquisition by the flow cytometer. Within the event sequence 315, each event 320 may correspond to an occurrence or a detection of at least one cell within a sample interval. The event sequence 315 for a given flow channel may be generated or acquired at a sampling rate (e.g., at a rate of 10 to 100 MHz). The event sequence 315 may have any number of events over a given period of time, ranging between 500 events (or cells) to 50,000 events (or cells) per second. The-35-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 sample dataset 310 may be generated as one or more files, such as using flow cytometer standard (FCS) or comma-separated values (CSV) file format.

[0133] In some embodiments, the sample dataset 310 may be derived or generated directly from flow cytometry performed on the biological sample from the subject (e.g., in a similar manner as the dataset 225). For example, the flow cytometer 110 may perform flow cytometry acquisition on the biological sample from a given subject to generate the sample dataset 310. In some embodiments, the sample dataset 310 may be synthetically derived or generated. For instance, the sample dataset 310 may be a combination of portions from different flow cytometry datasets. At least one of the flow cytometry datasets may be from a tumorous sample. At least one of the flow cytometry datasets may be from a normal sample. The sample dataset 310 may be generated using a combination of other sample datasets 310.

[0134] In each training sample 305, the sample dataset 310 may be labeled or annotated. The training sample 305 may identify or include a set of cell-level labels 325A- N (hereinafter generally referred to as cell-level labels 325). The set of cell-level labels 325 may correspond to the set of events 320 across the set of event sequences 315 of the sample dataset 310. Each cell-level label 325 may define or identify the corresponding event 320 as one of presence or absence of tumor associated with cancer (e.g., leukemia or malignant neoplasm). In some embodiments, each cell-level label 325 may identify the individual cell in the corresponding event 320 as a normal or tumorous cell. The set of cell-level labels 325 may be inputted or generated by a clinician examining the flow cytometry data of the sample dataset 310 from the subject. In addition, the training sample 305 may include at least one sample-level label 330 for the sample associated with the sample dataset 310. The sample-level label 330 may define or identify the sample as one of presence or absence of tumor associated with cancer (e.g., leukemia or malignant neoplasm). The sample-level label 330 may define or identify the sample as a normal or tumor sample.

[0135] With the identification, the model trainer 135 may input, feed, or otherwise apply each event 320 across the set of event sequences 315 to the cell-level ML model 160. The model trainer 135 may process each event 320 in accordance with the set of weights of the cell-level ML model 160. From applying the set of events 320 across the event-36-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 sequences 315, the model trainer 135 may calculate, generate, or determine a corresponding set of cell-level value 335A-N (hereinafter generally referred to as cell-level values 335). Each cell-level value 335 may identify or indicate a probability of tumor associated with cancer for the cell corresponding to the event 320. In some embodiments, the model trainer 135 may determine each cell-level value 335 to identify or indicate a probability of normal (or benign) corresponding to the event 320. The model trainer 135 may iterate through all the events 320 across the set of event sequences 315 to generate the set of cell-level values 335.

[0136] For each event 320, the model trainer 135 may compare each cell-level value 335 with a corresponding cell-level label 325 for the event 320. Based on the comparison, the model trainer 135 may calculate, generate, or otherwise determine at least one error metric. The error metric may indicate a degree of deviation of the output cell-level value 335 from the expected result as indicated in the cell-level label 325 for the event 320. The error metric in accordance with any number of loss functions, such as any number of loss functions, such as a norm loss (e.g., LI or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, and a Huber loss, among others.

[0137] Using the error metric, the model trainer 135 may modify, change, or otherwise update at least one of the set of weights in the cell-level ML model 160. The updating may be accordance with a backpropagation algorithm and an objective function. The objective function may define one or more rates at which the weights are to be updated. The objective function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The updating of the weights of the cell-level ML model 160 may be repeated until convergence.

[0138] In addition, the model trainer 135 may rearrange or sort the set of events 320 in the initial order in each event sequence 315 of the sample dataset 310 using the set of cell-level values 335. The sorting may arrange the events 320 with higher cell-level values 335 prior to events 320 with lower cell-level values 335. From sorting, the model trainer 135 may generate at least one sample dataset 310’ to include a set of event sequences315’ A-N (hereinafter generally referred to as event sequences 315’). Each event sequence-37-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009315’ may correspond to a respective flow channel. Each event sequence 315’ may include the set of events 320 in a rearranged order based on the cell-level values 335. In some embodiments, the model trainer 135 may exclude, remove, or otherwise truncate at least a portion of the events 320 based on the cell-level values 335. By truncating, the model trainer 135 may select a subset of events 320 in the set of events 320 in each of the set of event sequences 230 to include in the dataset 310’. For instance, the model trainer 135 may limit the number of events 320 after sorting to a predefined number of events (e.g., 5,000- 10,000 events). The number of remaining events 320 may correspond to an input size for the sample-level ML model 165.

[0139] The model trainer 135 may input, feed, or otherwise apply the set of event sequences 315’ of the sample dataset 310’ (e.g., in its entirety) into the sample-level ML model 165. In some embodiments, the model trainer 135 may carry out or perform a transformation on the sample data 310’ from a time-series domain to a target domain (e.g., frequency domain). The transformation may include, for example, a Fourier transform, a Wavelet transform, discrete cosine transform (DCT), or a Hilbert transform, among others. The model trainer 135 may apply the transformed sample dataset 310’ to the sample-level ML model 165. The model trainer 135 may process the set of event sequences 315’ of the sample dataset 310’ in accordance with the set of weights of the sample-level ML model 165. From applying the set of events 320 across the event sequences 315, the model trainer 135 may calculate, generate, or determine at least one sample value 340. The sample value 340 may identify or indicate a probability of tumor associated with cancer (e.g., leukemia or neoplasm) for the sample associated with the sample dataset 310’. In some embodiments, the model trainer 135 may determine the sample value 340 to identify or indicate a probability of normal (or benign) for the sample.

[0140] The model trainer 135 may compare the sample value 340 with the corresponding sample-level label 330 for the sample dataset 310. Based on the comparison, the model trainer 135 may calculate, generate, or otherwise determine at least one error metric. The error metric may indicate a degree of deviation of the output sample value 340 from the expected result as indicated in the sample-level label 330 for the sample dataset 310. The error metric in accordance with any number of loss functions, such as any number-38-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 of loss functions, such as a norm loss (e.g., LI or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, and a Huber loss, among others.[0141 J Using the error metric, the model trainer 135 may modify, change, or otherwise update at least one of the set of weights in the sample-level ML model 165. The updating may be accordance with a backpropagation algorithm and an objective function. The objective function may define one or more rates at which the weights are to be updated. The objective function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The updating of the weights of the sample-level ML model 165 may be repeated until convergence. In some embodiments, the model trainer 135 may store and maintain the set of weights for the cell-level ML model 160 and the sample-level ML model 165 upon completion of training.[01421 Referring now to FIG. 15, depicted is a block diagram of a process 400 to determine probabilities of tumor in samples using the machine learning models in the system for classifying samples. The process 400 may correspond to or include operations performed in the system 100 to process the flow cytometry data using machine learning models to determine probabilities of tumors in samples. Under the process 400, the event evaluator 140 executing on the data processing system 105 may input, feed, or otherwise apply each event 235 across the set of event sequences 230 of the dataset 225 to the celllevel ML model 160. The event evaluator 140 may process each event 235 in accordance with the set of weights of the cell-level ML model 160. From applying the set of events 235 across the event sequences 230, the event evaluator 140 may calculate, generate, or determine a corresponding set of cell-level value 405A-N (hereinafter generally referred to as cell-level values 405). Each cell-level value 405 may identify or indicate a probability of tumor associated with cancer for the cell corresponding to the event 235. In some embodiments, the event evaluator 140 may determine each cell-level value 405 to identify or indicate a probability of normal (or benign) corresponding to the event 235. The event evaluator 140 may iterate through all the events 235 across the set of event sequences 230 to generate the set of cell-level values 405.-39-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0143] The sample evaluator 145 executing on the data processing system 105 may rearrange or sort the set of events 235 in the initial order in each event sequence 230 of the dataset 225 using the set of cell-level values 405. The sorting may arrange the events 235 with higher cell-level values 405 prior to events 235 with lower cell-level values 405. From sorting, the model trainer 135 may generate at least one dataset 225’ to include a set of event sequences 230’ A-N (hereinafter generally referred to as event sequences 230’). Each event sequence 230’ may correspond to the respective flow channel 220. Each event sequence 230’ may identify or include the set of events 235 in a rearranged order based on the cell-level values 405. In some embodiments, the model trainer 135 may exclude, remove, or otherwise truncate at least a portion of the events 235 based on the cell-level values 405. By truncating, the model trainer 135 may select a subset of events 235 in the set of events 235 in each of the set of event sequences 230 to include in the dataset 225’. For instance, the model trainer 135 may limit the number of events 235 after sorting to a predefined number of events (e.g., 5,000-10,000 events). The number of remaining events 235 may correspond to an input size for the sample-level ML model 165.

[0144] The sample evaluator 145 may input, feed, or otherwise apply the set of event sequences 230’ of the dataset 225’ (e.g., in its entirety) into the sample-level ML model 165. In some embodiments, the sample evaluator 145 may carry out or perform a transformation on the sample data 310’ from a time-series domain to a target domain (e.g., frequency domain). The transformation may include, for example, a Fourier transform, a Wavelet transform, discrete cosine transform (DCT), or a Hilbert transform, among others. The sample evaluator 145 may apply the transformed dataset 225’ to the sample-level ML model 165. The sample evaluator 145 may process the set of event sequences 230’ of the dataset 225’ in accordance with the set of weights of the sample-level ML model 165. Based on applying the set of events 235 across the event sequences 230 to the sample-level ML model 165, the sample evaluator 145 may calculate, generate, or determine at least one sample value 410. The sample value 410 may identify or indicate a probability of tumor associated with cancer for the sample 210 associated with the dataset 225’. In some embodiments, the sample evaluator 145 may determine the sample value 410 to identify or indicate a probability of normal (or benign) for the sample 210.-40-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0145] Referring now to FIG. 16, depicted is a block diagram of a process 500 to generate outputs in the system 100 for classifying samples. The process 500 may correspond to or include operations performed in the system 100 to use the values from the machine learning models to derive outputs. Under the process 500, the output generator 155 on the data processing system 105 may produce, determine, or otherwise generate at least one classification 505 in accordance with the sample value 410. The classification 505 may identify or indicate a presence or absence of tumor associated with cancer in the sample 210 obtained from the subject 205. To generate the classification 505, the output generator 155 may compare the sample value 410 with a threshold. The threshold may delineate, identify, or otherwise define value for the sample value 410 at which to identify the presence or absence of tumor associated with cancer in the tumor. In some embodiments, the threshold may be indicative of a presence (or absence) of minimum residual disease (MRD) associated with cancer in the subject 205.

[0146] If the sample value 410 satisfies (e.g., greater than or equal to) the threshold, the output generator 155 may generate the classification 505 to indicate the presence of cancer in the sample 210. In some embodiments, when the sample value 410 satisfies the threshold indicative of MRD, the output generator 155 may generate the classification 505 to indicate presence of MRD in the subject 205. In some embodiments, the output generator 155 may generate the classification 505 to identify that the subject 205 is to be administered with an anti-cancer therapy for cancer. When the cancer is leukemia, the anti-cancer therapy may include anti-leukemia therapy.

[0147] The anti-leukemia therapy may include, for example, a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor, among others. Examples of FLT inhibitors include, but are not limited to, sunitinib, sorafenib, midostaurin, lestaurtinib, quizartinib, gilteritinib, crenolanib, and gilteritini. Examples of IDH inhibitors include, but are not limited to, ivosidenib, olutasidenib, vorasidenib, and enasidenib. Examples of BTK inhibitors include, but are not limited to, ibrutinib, acalabrutinib, Zanubrutinib, and pirtobrutinib. Examples of BCL2 inhibitors include, but are not limited to, venetoclax, navitoclax, obatoclax, oblimersen sodium (G3139), Palcitoclax, AT-101, and LP-118.-41-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009Examples of hedgehog pathway inhibitors include, but are not limited to, glasdegib, vismodegib, sonidegib, LEQ-506 (NVP-LEQ506), itraconazole, saridegib, BMS-833923 / XL139, Arsenic Trioxide (ATO) and Taladegib.

[0148] Examples of anti -cancer therapies include but are not limited to alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogen drugs, aromatase inhibitors, ovarian suppression agents, VEGF / VEGFR inhibitors, EGFZEGFR inhibitors, PARP inhibitors, cytostatic alkaloids, cytotoxic antibiotics, antimetabolites, endocrine / hormonal agents, bisphosphonate therapy agents, immuno-modulating / stimulating antibodies and targeted biological therapy agents (e.g., therapeutic peptides described in US 6306832, WO 2012007137, WO 2005000889, WO 2010096603 etc.). In some embodiments, the at least one additional therapeutic agent is a chemotherapeutic agent. Specific chemotherapeutic agents include, but are not limited to, cyclophosphamide, fluorouracil (or 5 -fluorouracil or 5-FU), methotrexate, edatrexate (10-ethyl-10-deaza-aminopterin), thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabepilone, temozolmide, topotecan, vincristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane, letrozole, leuprolide, abarelix, buserlin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, denosumab, zoledronate, trastuzumab, tykerb, anthracyclines (e.g., daunorubicin and doxorubicin), bevacizumab, oxaliplatin, melphalan, etoposide, mechlorethamine, bleomycin, microtubule poisons, annonaceous acetogenins, or combinations thereof. Examples of immuno-modulating / stimulating antibody include but are not limited to anti-PD-1 antibody, anti-PD-Ll antibody, anti-PD- L2 antibody, anti-CTLA-4 antibody, anti-TIM3 antibody, anti -4- IBB antibody, anti-CD73 antibody, anti-GITR antibody, and anti-LAG-3 antibody.

[0149] Otherwise, if the sample value 410 does not satisfy (e.g., less than) the threshold, the output generator 155 may generate the classification 505 to indicate the absence of cancer in the sample 210. In some embodiments, when the sample value 410 does not satisfy the threshold indicative of MRD, the output generator 155 may generate the classification 505 to indicate absence of MRD in the subject 205. With the generation, the output generator 155 may store and maintain an association between the subject 205 and the-42-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 classification 505 using one or more data structures. The data structures may include, for example, an array, a matrix, a linked list, a stack, a tree, a hash table, an object, among others.

[0150] Using the classification 505, the output generator 155 may produce, create, or otherwise generate at least one output 510. The output 510 may include information associated with the classification 505. When the classification 505 indicates the presence of the cancer, the output 510 may identify or include the indication of the presence of cancer. When the classification 505 indicates the presence of MRD, the output 510 may identify or include the indication of the presence of MRD associated with cancer. In some embodiments, the output 510 may identify the subject 205 as to be administered with an anti-cancer for cancer or MRD associated with cancer in the subject 205. When the classification 505 indicates the absence of the cancer, the output 510 may identify or include the indication of the absence of cancer. When the classification 505 indicates the presence of MRD, the output 510 may identify or include the indication of the absence of MRD associated with cancer.

[0151] The output generator 155 may transmit, send, or otherwise provide the output 510 to the administrative device 115. With receipt, the administrative device 115 may display, render, or otherwise present the output 510 including the information on the classification for the subject 205. The information of the output 510 may be presented via a graphical user interface on a display of the administrative device 115. The information of the output 510 may be used to make clinical decisions. For example, a clinician operating the administrative device 115 may use the information to identify whether the sample 210 of the subject 205 has or lacks tumor associated with cancer. When the output 510 indicates the presence of MRD associated leukemia in the subject 205, the clinician may also use the information of the output 510 to select administration of an anti -leukemia therapy for the subject 205.

[0152] In this manner, by leveraging the cell-level ML model 160 and the samplelevel ML model 165, the data processing system 105 may facilitate derivation of clinically relevant data and information from flow cytometry data. The cell-level ML model 160 may individually process events 235 to provide corresponding cell-level values 405. With the -43-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 dataset 225’ including the events 235 rearranged based on the cell-level values 405, the sample-level ML model 165 may determine the sample-level value 410 for the sample 210 from which the flow cytometry data is derived. The bifurcation of the ML architecture may provide a more refined and accurate assessment of the individual cells corresponding to the events 235, which in turn can be used to evaluate the overall sample 210.

[0153] The architecture of the cell-level ML model 160 and the sample-level ML model 165 may provide performance advantages relative to other approaches, in terms of accuracy and precision. The higher accuracy and precision can facilitate the generation of clinically relevant information, such as the classification 505 regarding presence or absence of leukemia or neoplasm and MRD as well as identification of the subject 205 for additional administration of anti-leukemia therapy. From a computer resource perspective, by using the cell-level ML model 160 and the sample-level ML model 165, the data processing system 105 may allow for more efficient use of computing resources (e.g., processing, memory, and network bandwidth) that otherwise would have been wasted in providing inaccurate or useless output from flow cytometry data.

[0154] Referring now to FIG. 17, depicted is a flow diagram of a method 600 of classifying samples based on flow cytometry data. The method 600 may be implemented or performed using any of the components described herein, such as the data processing system 105, or any combination thereof, or the system 700 described in Section B. Under the method 600, a computing system may receive a dataset including a set of event sequences from a flow cytometer (605). The computing system may apply each event sequence in the dataset to a cell-level machine learning (ML) model (610). The computing system may determine a set of cell-level values for the set of event sequences based on applying to the cell-level ML model (615). The computing system may sort the set of event sequences in the dataset based on the cell-level values (620). The computing system may apply the sorted set of event sequences to the sample-level ML model (625). The computing system may determine a sample-level value based on applying to the samplelevel ML model (630). The computing system may determine whether the sample-level value satisfies a threshold (635). If the sample-level value satisfies (e.g., greater than or equal to) the threshold, the computing system may classify to indicate a presence of tumor-44-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009(640). Otherwise, if the sample-level value does not satisfy (e.g., less than) the threshold, the computing system may classify to indicate an absence of tumor (645). The computing system may provide an output based on the classification (650).D. Computing and Network Environment

[0155] Various operations described herein can be implemented on computer systems. FIG. 18 shows a simplified block diagram of a representative server system 700, client computing system 714, and network 726 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 700 or similar systems can implement services or servers described herein or portions thereof. Client computing system 714 or similar systems can implement clients described herein. The system 700 described herein can be similar to the server system 700. Server system 700 can have a modular design that incorporates a number of modules 702 (e.g., blades in a blade server embodiment); while two modules 702 are shown, any number can be provided. Each module 702 can include processing unit(s) 704 and local storage 706.

[0156] Processing unit(s) 704 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 704 can include a general-purpose primary processor as well as one or more special-purpose coprocessors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 704 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 704 can execute instructions stored in local storage 706. Any type of processors in any combination can be included in processing unit(s) 704.

[0157] Local storage 706 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 706 can be fixed, removable or upgradeable as desired. Local storage 706 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), and a-45-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 704 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 704. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 702 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections.

[0158] In some embodiments, local storage 706 can store one or more software programs to be executed by processing unit(s) 704, such as an operating system and / or programs implementing various server functions such as functions of the system 100 of FIG. 12 or any other system described herein, or any other server(s) associated with system 100 or any other system described herein.

[0159] “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 704 cause server system 700 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 704. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 706 (or non-local storage described below), processing unit(s) 704 can retrieve program instructions to execute and data to process in order to execute various operations described above.

[0160] In some server systems 700, multiple modules 702 can be interconnected via a bus or other interconnect 708, forming a local area network that supports communication between modules 702 and other components of server system 700. Interconnect 708 can be implemented using various technologies including server racks, hubs, routers, etc.-46-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0161] A wide area network (WAN) interface 710 can provide data communication capability between the local area network (interconnect 708) and the network 726, such as the Internet. Technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.24 standards).

[0162] In some embodiments, local storage 706 is intended to provide working memory for processing unit(s) 704, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 708. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 712 that can be connected to interconnect 708. Mass storage subsystem 712 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 712. In some embodiments, additional data storage resources may be accessible via WAN interface 710 (potentially with increased latency).

[0163] Server system 700 can operate in response to requests received via WAN interface 710. For example, one of modules 702 can implement a supervisory function and assign discrete tasks to other modules 702 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 710. Such operation can generally be automated. Further, in some embodiments, WAN interface 710 can connect multiple server systems 700 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation.

[0164] Server system 700 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 7 as client computing system 714. Client computing system 714 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on.-47-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0165] For example, client computing system 714 can communicate via WAN interface 710. Client computing system 714 can include computer components such as processing unit(s) 716, storage device 718, network interface 720, user input device 722, and user output device 724. Client computing system 714 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like.

[0166] Processing unit(s) 716 and storage device 718 can be similar to processing unit(s) 704 and local storage 706 described above. Suitable devices can be selected based on the demands to be placed on client computing system 714; for example, client computing system 714 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 714 can be provisioned with program code executable by processing unit(s) 716 to enable various interactions with server system 700.

[0167] Network interface 720 can provide a connection to the network 726, such as a wide area network (e.g., the Internet) to which WAN interface 710 of server system 700 is also connected. In various embodiments, network interface 720 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc ).

[0168] User input device 722 can include any device (or devices) via which a user can provide signals to client computing system 714; client computing system 714 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 722 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on.

[0169] User output device 724 can include any device via which client computing system 714 can provide information to a user. For example, user output device 724 can include a display to present images generated by or delivered to client computing system-48-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009714. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital -to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 724 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.

[0170] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer-readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 704 and 716 can provide various functionality for server system 700 and client computing system 714, including any of the functionality described herein as being performed by a server or client, or other functionality.

[0171] It will be appreciated that server system 700 and client computing system 714 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 700 and client computing system 714 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be-49-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software.[0172J While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to the specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa.[01731 Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer-readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer-readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium).-50-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009

[0174] Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims.4891 -3791 -9722.1

Claims

Atty. Dkt. No.: 115872-3009WHAT IS CLAIMED IS:

1. A method of classifying samples based on flow cytometry data, comprising: receiving, by one or more processors, from a flow cytometer, a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels, each of the plurality of event sequences identifying a respective set of events in a corresponding flow channel of the plurality of flow channels arranged in a first order, the cells corresponding to at least one of a plurality of white blood cell types in a sample obtained from a subject at risk of or diagnosed with cancer; applying, by the one or more processors, the first dataset to a first machine learning (ML) model to determine a plurality of first values corresponding to the set of events in each of the plurality of event sequences, each first value of the plurality of first values indicating a respective probability of tumor associated with cancer; sorting, by the one or more processors, the set of events in the plurality of event sequences arranged in the first order in the first dataset by the plurality of first values to generate a second dataset comprising the plurality of event sequences, each of the plurality of event sequences in the second dataset including the respective set of events in a second order; applying, by the one or more processors, the second dataset to a second ML model, wherein the second ML model is established using a plurality of training samples, each training sample of the plurality of training samples comprising a respective dataset including a corresponding plurality of event sequences labeled with presence or absence of tumor associated with cancer; determining, by the one or more processors, based on applying the second dataset to the second ML model, a second value indicating a probability of tumor associated with cancer in the sample of the subject; generating, by the one or more processors, a classification indicating a presence or absence of tumor associated with cancer in the sample in accordance with the second value; and storing, by the one or more processors, using one or more data structures, an association between the subject and the classification.-52-4891 -3791 -9722.1Atty. Dkt. No.: 115872-30092. The method of claim 1, wherein generating the classification further comprises generating the classification to indicate the presence of the tumor associated with cancer in the sample, responsive to the second value satisfying a threshold indicative of a presence of a minimum residual disease (MRD) in the subject; and providing, by the one or more processors, an output identifying at least one of (i) the classification to indicate the presence of the tumor, (ii) the presence of the MRD in the subject, or (iii) the subject as to be administered with an anti-cancer therapy for cancer.

3. The method of claim 2, wherein the cancer includes leukemia, and the cancer therapy includes an anti-leukemia therapy comprising at least one of a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor.

4. The method of claim 1, wherein generating the classification further comprises generating the classification to indicate the absence of the tumor associated with cancer in the sample, responsive to the second value not satisfying a threshold indicative of a MRD in the subject; and providing, by the one or more processors, an output identifying at least one of (i) the classification to indicate the absence of the tumor or (ii) the absence of the MRD in the subject5. The method of claim 1, further comprising transforming, by the one or more processors, the plurality of event sequences in the second dataset from a time-series domain to a frequency domain, wherein applying the second dataset to the second ML model further comprises applying the second dataset in the frequency domain to second ML model to determine the second value for the sample.

6. The method of claim 1, wherein applying the first dataset to the first ML model further comprises applying each event in each of the set of events in the plurality of event-53-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 sequences to the first ML model to generate a corresponding first value of the first plurality of first values, and wherein applying the second dataset to the second ML model further comprises applying the plurality of event sequences of the second dataset in entirety to the second ML model to determine the second value for the sample.

7. The method of claim 1, wherein sorting the first dataset further comprises truncating the plurality of event sequences based on the plurality of first values to select a subset of the plurality of event sequences to include in the second dataset.

8. The method of claim 1, wherein the first ML model is established using a plurality of second training samples, each second training sample of the plurality of second training samples comprising a respective dataset including a corresponding plurality of event sequences for a respective plurality of flow channels, the corresponding plurality of event sequences including a set of events each labeled as one of a normal cell or a tumorous cell.

9. The method of claim 1, wherein at least one of the plurality of training samples used to train the second ML model further comprises the respective dataset generated from a combination of a plurality of portions from a plurality of datasets, the plurality of datasets corresponding at least one of a normal sample or tumor sample.

10. The method of claim 1, wherein receiving the first dataset further comprise receiving the first dataset comprising the plurality of event sequences for the corresponding plurality of flow channels including at least one of a forward scatter channel, a side scatter channel, or a fluorescence channel.

11. The method of claim 1, wherein the cancer includes leukemia comprising at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL).-54-4891 -3791 -9722.1Atty. Dkt. No.: 115872-300912. The method of claim 1, wherein the plurality of white cell blood types comprises at least one of CD 10+ CD 19+ CD20+, CD34+, CD38+, or CD45+.

13. A system for classifying samples based on flow cytometry data, comprising: one or more processors coupled with memory, configured to: receive, from a flow cytometer, a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels, each of the plurality of event sequences identifying a respective set of events in a corresponding flow channel of the plurality of flow channels arranged in a first order, the cells corresponding to at least one of a plurality of white blood cell types in a sample obtained from a subject at risk of or diagnosed with cancer; apply the first dataset to a first machine learning (ML) model to determine a plurality of first values corresponding to the set of events in each of the plurality of event sequences, each first value of the plurality of first values indicating a respective probability of tumor associated with cancer; sort the set of events in the plurality of event sequences arranged in the first order in the first dataset by the plurality of first values to generate a second dataset comprising the plurality of event sequences, each of the plurality of event sequences in the second dataset including the respective set of events in a second order; apply the second dataset to a second ML model, wherein the second ML model is established using a plurality of training samples, each training sample of the plurality of training samples comprising a respective dataset including a corresponding plurality of event sequences labeled with presence or absence of tumor associated with cancer; determine, based on applying the second dataset to the second ML model, a second value indicating a probability of tumor associated with cancer in the sample of the subject; generate a classification indicating a presence or absence of tumor associated with cancer in the sample in accordance with the second value; and store, using one or more data structures, an association between the subject and the classification.-55-4891 -3791 -9722.1Atty. Dkt. No.: 115872-300914. The system of claim 13, wherein the one or more processors are further configured to: generate the classification to indicate the presence of the tumor associated with cancer in the sample, responsive to the second value satisfying a threshold indicative of a presence of a minimum residual disease (MRD) in the subject; and provide an output identifying at least one of (i) the classification to indicate the presence of the tumor, (ii) the presence of the MRD in the subject, or (iii) the subject as to be administered with an anti -cancer therapy for cancer.

15. The system of claim 14, wherein the cancer includes leukemia, and the cancer therapy includes an anti-leukemia therapy comprising at least one of a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor.

16. The system of claim 13, wherein the one or more processors are further configured to: generate the classification to indicate the absence of the tumor associated with cancer in the sample, responsive to the second value not satisfying a threshold indicative of a MRD in the subject; and provide an output identifying at least one of (i) the classification to indicate the absence of the tumor or (ii) the absence of the MRD in the subject.

17. The system of claim 13, wherein the one or more processors are further configured to: transform the plurality of event sequences in the second dataset from a time-series domain to a frequency domain, apply the second dataset in the frequency domain to second ML model to determine the second value for the sample.

18. The system of claim 13, wherein the one or more processors are further configured to: apply each event in each of the set of events in the plurality of event sequences to the first ML model to generate a corresponding first value of the first plurality of first values, and-56-4891 -3791 -9722.1Atty. Dkt. No.: 115872-3009 apply the plurality of event sequences of the second dataset in entirety to the second ML model to determine the second value for the sample.

19. The system of claim 13, wherein the one or more processors are further configured to truncate the plurality of event sequences based on the plurality of first values to select a subset of the plurality of event sequences to include in the second dataset.

20. The system of claim 13, wherein the first ML model is established using a plurality of second training samples, each second training sample of the plurality of second training samples comprising a respective dataset including a corresponding plurality of event sequences for a respective plurality of flow channels, the corresponding plurality of event sequences including a set of events each labeled as one of a normal cell or a tumorous cell.

21. The system of claim 13, wherein at least one of the plurality of training samples used to train the second ML model further comprises the respective dataset generated from a combination of a plurality of portions from a plurality of datasets, the plurality of datasets corresponding at least one of a normal sample or tumor sample.

22. The system of claim 13, wherein the one or more processors are further configured to receive the first dataset comprising the plurality of event sequences for the corresponding plurality of flow channels including at least one of a forward scatter channel, a side scatter channel, or a fluorescence channel.

23. The system of claim 13, wherein the cancer includes leukemia comprising at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL).

24. The system of claim 13, wherein the plurality of white cell blood types comprises at least one of CD10+, CD19+, CD20+, CD34+, CD38+, or CD45 +.-57-4891 -3791 -9722.1

Citation Information

Patent Citations

  • Artificial intelligence for early cancer detection

    US20220252602A1

  • Machine learning techniques for cytometry

    US20230245479A1

  • Methods of diagnosing cancer using multiple artificial neural networks to analyze flow cytometry data

    WO2020081582A1