Detection of cancer using light-chain expression features extracted from flow cytometry data

WO2026206849A1PCT designated stage Publication Date: 2026-10-01MEMORIAL SLOAN KETTERING CANCER CENT +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020382
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure US2026020382_01102026_PF_FP_ABST
    Figure US2026020382_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are systems and methods of classifying samples based on light chain (LC) expression features in flow cytometry data. A computing system can receive from a flow cytometer or a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels. The computing system can apply the first dataset to a first machine learning (ML) model to generate a plurality of LC expression classifications. The computing system can generate a plurality of spatial metrics using the plurality of LC expression classifications. The computing system can apply the first dataset and the plurality of spatial metrics to a second ML model to generate a second dataset comprising of a first plurality of values. The computing system can apply the second dataset to a third ML model to determine a second value indicating a probability of tumor associated with cancer in the sample of the subject. Presented herein are systems and methods of classifying samples based on light chain (LC) expression features in flow cytometry data. A computing system can receive from a flow cytometer or a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels. The computing system can apply the first dataset to a first machine learning (ML) model to generate a plurality of LC expression classifications. The computing system can generate a plurality of spatial metrics using the plurality of LC expression classifications. The computing system can apply the first dataset and the plurality of spatial metrics to a second ML model to generate a second dataset comprising of a first plurality of values. The computing system can apply the second dataset to a third ML model to determine a second value indicating a probability of tumor associated with cancer in the sample of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt. No.: 115872-3451DETECTION OF CANCER USING LIGHT-CHAIN EXPRESSION FEATURES EXTRACTED FROM FLOW CYTOMETRY DATA CROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 776,753, filed March 24, 2025, which is incorporated by reference in its entirety.BACKGROUND

[0002] A computing system may receive input data. Using a machine learning model, the computing system may process the input data to generate output data.SUMMARY

[0003] Aspects of the present disclosure are directed to systems, methods, devices, and non-transitory computer readable media for classifying samples based on light chain (LC) expression features in flow cytometry data. One or more processors coupled with memory can receive, from a flow cytometer, a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels. The plurality of event sequences each identify a respective set of events in a corresponding flow channel of the plurality of flow channels. The set of events can correspond to at least one of a plurality of white blood cells in a sample obtained from a subject at risk of or diagnosed with cancer. The one or more processors can apply the first dataset to a first machine learning (ML) model to generate a plurality of LC expression classifications. Each of the plurality of LC expression classifications can identify a respective class for a respective event in the respective set of events in each of the plurality of event sequences. The one or more processors can generate a plurality of spatial metrics using the plurality of LC expression classifications. Each of the plurality of spatial metrics can be for the respective event in the respective set of events in each of the plurality of event sequences. The one or more processors can apply the first dataset and the plurality of spatial metrics to a second ML model to generate a second dataset comprising a first plurality of values. Each first value of the plurality of first values can indicate a respective probability of a tumor associated with cancer in the respective -1- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451event in the respective set of events in each of the plurality of event sequences. The one or more processors can apply the second dataset to a third ML model to determine a second value indicating a probability of a tumor associated with cancer in the sample of the subject. The one or more processors can generate a tumor classification indicating a presence or absence of a tumor associated with cancer in the sample in accordance with the second value. The one or more processors can store, using one or more data structures, an association between the subject and the classification.

[0004] In some implementations, the one or more processors can generate the tumor classification to indicate the presence of the tumor associated with cancer in the sample, responsive to the second value satisfying a threshold. In some implementations, the one or more processors can provide output identifying at least one of (i) the tumor classification to indicate the presence of the tumor, (ii) the presence of the cancer in the subject, or (iii) the subject to be administered with an anti-cancer therapy for cancer. In some implementations, the cancer can include leukemia. In some implementations, the cancer therapy can include an anti-leukemia therapy comprising at least one of a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, a BCL-2 inhibitor, or a hedgehog pathway inhibitor. In some implementations, the subject can be administered with the anti-cancer therapy based on the output identifying the subject to be administered. In some implementations, the one or more processors can generate the tumor classification to indicate the absence of the tumor associated with cancer in the sample, responsive to the second value not satisfying a threshold. In some implementations, the one or more processors can provide an output identifying at least one of (i) the tumor classification to indicate the absence of the tumor or (ii) the absence of the cancer in the subject.

[0005] In some implementations, the one or more processors can identify a plurality of embedding sets based on applying the first dataset to the first ML model. Each embedding set of the plurality of embedding sets can correspond to the respective event in the respective set of events in each of the plurality of event sequences. In some implementations, the one or more processors can select from the plurality of embedding sets a subset of embedding sets corresponding to at least one of a predefined plurality of feature types. In some implementations, the one or more processors can identify, for the subset of embedding sets, a corresponding subset of LC expression classifications from the plurality -2- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451of LC expression classifications. In some implementations, the one or more processors can generate the plurality of spatial metrics using the subset of embedding sets and subset LC expression classifications. In some implementations, the predefined plurality of feature types can comprise at least one of FSC-A, FSC-H, SSC-H, CD10, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD20, CD22, or CD19.

[0006] In some implementations, the one or more processors can select, for each respective event of the respective set of events in each of the plurality of event sequences, a subset of events from the respective set of events in each of the plurality of event sequences based on a corresponding embedding set of the plurality of embedding sets and a subset of embedding sets of the plurality of embedding sets within a feature space. In some implementations, the one or more processors can identify from the plurality of LC expression classifications a subset of LC expression classifications corresponding to the subset of events. In some implementations, the one or more processors can determine a respective similarity metric based on a respective LC expression classification and the subset of LC expression classifications. In some implementations, the one or more processors can create, in a feature space, a boundary defining at least a subset of the plurality of embedding sets corresponding to at least one of the plurality of LC expression classifications. In some implementations, the one or more processors can determine, for each respective event of the respective set of events in each of the plurality of event sequences, a respective edge metric between a corresponding embedding set of the plurality of embedding sets and the boundary in the feature space.

[0007] In some implementations, the one or more processors can identify a second plurality of embedding sets generated from applying the first dataset to the first ML model. In some implementations, the one or more processors can generate the plurality of embedding sets using the second plurality of embedding sets in accordance with a dimension reduction. In some implementations, the first ML model can be trained using a plurality of training examples. Each of the plurality of training examples can comprise a respective dataset including a respective plurality of event sequences. Each of the respective plurality of event sequences can identify a corresponding set of events in the corresponding flow channel of the plurality of flow channels. Each event of the corresponding set of events can be labeled as one of a set of LC classifications.-3- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0008] In some implementations, the one or more processors can receive the first dataset where each of the plurality of event sequences identifies the respective set of events in the corresponding flow channel of the plurality of flow channels arranged in a first order. In some implementations, the one or more processors can sort the set of events in the plurality of event sequences arranged in the first order in the first dataset by the plurality of first values to define a second order. In some implementations, the one or more processors can apply the second data set arranged in the second order to the third ML model. In some implementations, the one or more processors can apply the second dataset to the third ML model to generate a second plurality of embedding sets. In some implementations, the one or more processors can generate a cancer classification indicating at least one of a plurality of cancer types for the subject based on the second plurality of embedding sets. In some implementations, the cancer can include leukemia. The plurality of cancer types can comprise at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL).

[0009] In some implementations, the one or more processors can identify the first dataset comprising the plurality of event sequences for the corresponding plurality of flow channels. Each of the plurality of event sequences can correspond to at least one biomarker of a plurality of biomarkers in the plurality of white blood cells in the sample. In some implementations, the plurality of biomarkers can comprise at least one of CD8, TCRyb, TCR pi, CD10, CD7, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD279 (PD-1), CD20, CD2, CD56, CD25, CD22, or CD1. In some implementations, each event in the respective set of events of the plurality of event sequences can comprise at least one of a forward scatter channel, a side scatter channel, or a fluorescence channel. In some implementations, each LC expression classification of the plurality of LC expression classifications can identify one of a kappa class, a lambda class, or a neither class. In some implementations, the sample can comprise at least one of a blood sample, a bone marrow sample, or a tissue sample. In some implementations, the plurality of white blood cells can comprise B-cells in the sample.

[0010] In some implementations, the one or more processors may determine, for each event of the set of events in each of the plurality of event sequences, a similarity metric -4- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451based on a first expression classification of the plurality of expression classifications corresponding to the event and a second expression classification of the plurality of expression classifications of another event. The one or more processors may apply the similarity metric for each event of the set of events in each of the plurality of event sequences to the second model. In some implementations, the one or more processors may determine, for each event of the set of events in each of the plurality of event sequences, an edge metric based on a first expression classification of the plurality of expression classifications of the event and a boundary defining a subset of the plurality of expression classifications. The one or more processors may apply the edge metric for each event of the set of events in each of the plurality of event sequences.

[0011] In some embodiments, the cancer or tumor is a carcinoma, a sarcoma, a melanoma, or a hematopoietic cancer. In some embodiments, the cancer or tumor is selected from among adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinomas, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin’s disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, lung cancers, melanomas, mesotheliomas, myelomas, nasopharynx cancers, neuroblastomas, non-Hodgkin’s lymphomas, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcomas, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:

[0013] FIGs. 1 A-1D: Block diagram of FlowARC, a multi-module deep-leaming-based model, follows a four-step approach for detecting tumor presence in flow cytometry data. The initial step (A) is a pipeline to preprocess the data. Second, in the Audit stage (B), a subset of data is meticulously reviewed by pathologists, assigning each cell as either -5- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451normal or tumor. These labeled cells are then used to train the Cell-level module, a deeplearning model that takes a single cell measurement as input and assigns a probability of it being a tumor cell. Next, in the Reorder stage (C), all cells within a flow sample are rearranged based on their tumor probabilities as determined by the Cell-level module. This reordering places cells with a higher likelihood of being tumors towards the top of the sample, allowing for the truncation of hundreds of thousands of cells into a few thousand relevant cells for tumor detection in each sample. Finally, in the Classify stage (D), the reordered and truncated flow samples, labeled as either normal or tumor, are used to train the Sample-level module. This module determines if a given flow sample is tumorous.

[0014] FIGs. 2A-2E: (A) The results of Flow ARC -Manual on the Synthetic test dataset of the annotated cohort as a confusion matrix. (B) The results of Flow ARC -Manual on the patient test cases of the annotated cohort as a confusion matrix. (C) The results of FlowARC-Auto on the Synthetic test dataset of the clinical archive as a confusion matrix. (D) The results of Flow ARC -Auto on the patient test cases of the clinical archive as a confusion matrix. (E) The comparison between Flow ARC and CellCNN on the patient test cases of the annotated cohort (left) and the clinical archive (right) is shown using the area under the ROC curve metric with 95% confidence interval.

[0015] FIGs. 3A-3C: The visualization chart using the FlowARC-Auto cell-level module on six patient test cases of the clinical archive is shown. In this chart, each cells are color-coded according to tumor likelihood obtained from the cell-level module. (A) shows two cases that were correctly predicted by FlowARC-Auto. (B) shows two cases that were incorrectly predicted by FlowARC-Auto. (C) shows two cases that could not be injected onto the FlowARC-Auto sample-level module due to their B-cell population being less than 7500.

[0016] FIGs. 4A and 4B: The performance of Flow ARC -Manual on predicting the tumor cell population on the B- ALL positive (tumor) synthetic test samples (A) and patient test cases (B) is shown. On each panel, the left graph shows the results from the FlowARC-Manual quantification module. For cases with a quantification module prediction of 7200 and above, the right graph shows the results from the averaged Flow ARC -Manual cell-level module predictions. The perfect prediction for each graph is denoted by a dashed line.-6- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0017] FIGs. 5A-5C: (A) shows the FlowARC -Manual cell-level module results as a set of bi-parametric plots on a random representative subset of test cells. Each cell is represented by a dot whose color and size reflect the classification outcomes, as shown in the legend at the bottom-right. (B) shows the SHAP summary plot for cell predictions, tumor (top) and normal (bottom). The cell markers are arranged in rows according to their significance for each category. The overlapping SHAP values are jittered along the y-axis. The color bar encodes the input value of the feature. (C) shows the t-SNE projection of the penultimate layer (features) of the Flow ARC -Manual sample-level module predictions on the synthetic test dataset. Each synthetic test sample is represented by a dot whose color and size reflect the classification outcomes, as shown in the legend at the bottom.

[0018] FIG. 6: The distribution of patient cases and the process of obtaining labeled cells for datasets used in this study, the annotated cohort, and the clinical archive, are shown.

[0019] FIG. 7: The four-step approach to auto-gating the CD 19+ B-cells from patient cases using flowDensity is shown. Step 1 removes doublets and margin events. Steps 2 and 3 select the viable mononuclear population. Step 4 selects the B-cell population.

[0020] FIGs. 8A-8C: (A) shows the distribution of cases within a synthetic dataset and their respective tumor population ranges as a pie-chart. (B) illustrates the process of generating a normal synthetic case containing N normal B-cells. (C) illustrates the process of generating a tumor synthetic case containing a total of N cells with NN normal B-cells and NA tumor cells.

[0021] FIGs. 9A-9D: (A) shows the heatmap of the cell markers and cell diagnosis (Dx) in a random synthetic flow sample with 65,000 tumor cells before and after reordering using the FlowARC- Manual cell-level module. (B) is a histogram showing the percentage of actual tumor population within the top 7,500 cells of the synthetic tumor test set. (C) is a histogram showing the false negative rate for different minimum tumor populations within a synthetic tumor test set. (D) shows the cell distributions of the labeled cells of the annotated cohort (left) and the clinical archive (right) as pie-charts.-7- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0022] FIGs. 10A and 10B: (A) B cell-count distribution among the annotated cases after preprocessing and the train / test split. (B) The cell-level classifier prediction is shown on selected bi-parametric flow cytometry plots and a tSNE plot. Each cell is color-coded according to tumor likelihood obtained from the cell-level module. The upper panel represents a normal case, and the lower panel represents an abnormal case.

[0023] FIGs. 11 A-l 1C: (A) A process is depicted. The manual annotations for all cases were obtained. The cases were then subjected to automated preprocessing for compensation, transformation, and B-cell gating using flowDensity. The cases were split into training and testing sets in a 70 / 30 split for normal and tumor groups (Subpanel 1). The cells from the training set were used to train a cell-level classifier (Subpanel 2). The process of generating a Normal Synthetic sample (Subpanel 3) and a Tumor Synthetic sample (Subpanel 4) is shown. 7500 synthetic training sets were generated using the training cells. The ratio of normal to tumor samples is 1:1. Each cell on those synthetic samples was passed through the Cell-level classifier and sorted in the descending order of tumor probability (Subpanel 5). The cells on the testing set were also passed through the Cell-level module and sorted in descending order of tumor probability (Subpanel 6). The synthetic training set was then used to train a sample-level classifier, which was then tested on the testing set. Panel B) The result from the Sample-level module, applied to the testing set, is shown as a Confusion matrix. Panel C) The ROC curve of the testing set predictions shows an AUC of 0.985.

[0024] FIGs. 12A-12F: shows the overview of the automated flow cytometry analysis system. Clinical specimens undergo flow cytometry acquisition and preprocessing, including compensation, transformation, and density-based B-cell isolation, with expert annotations providing ground truth for case-level diagnosis and cell-level labels (FIG. 12A). The Light-chain module uses a ID ResNet-18 to classify each cell's surface immunoglobulin as kappa-positive, lambda-positive, double-negative, or contaminating cell (FIG. 12B). Three engineered features are then computed from case-specific Large Vis embeddings: light-chain neighborhood (LC-N; count of 30 nearest neighbors with matching light-chain prediction), edge distance (normalized distance to population boundary), and abnormal neighborhood (mean abnormality score of 30 nearest neighbors) (FIG. 12C). The Cell-level module combines original flow parameters with these engineered features to -8- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451predict cell-level abnormality using a ID ResNet-18 (FIG. 12D), and the Sample-level module uses a 2D ResNet-101 to render a case-level diagnosis from the top 2,000 cells ranked by abnormality score (FIG. 12E). Visualization outputs display biaxial marker plots and Large Vis embeddings colored by abnormality score for pathologist review (FIG. 12F).

[0025] FIGs. 13 A and 13B: show classification performance of the automated system. (A) Receiver operating characteristic (ROC) curves demonstrate progressively improved discrimination across the three modules, with the Light-chain module achieving an AUROC of 0.896 (95% CL 0.890-0.903), the Cell-level module 0.965 (95% CL 0.961-0.969), and the Sample-level module 0.988 (95% CL 0.982-0.993). Shaded regions indicate 95% confidence intervals across five independent runs. (B) Confusion matrices display case-level classification performance on the held-out test set overall and stratified by specimen type, with values representing median counts and ranges across five runs in brackets. The system achieved 97.5% overall accuracy (95.6% sensitivity, 98.9% specificity), with consistent performance across peripheral blood (97.1% accuracy), bone marrow (97.4% accuracy), and tissue (98.0% accuracy) specimens.

[0026] FIGs. 14A and 14B: show feature contribution analysis and cell-level validation. (A) Ablation experiments assess the contribution of engineered features to model performance. The base model (B) uses only original flow cytometry parameters; B+KL adds light-chain classification and LC-N features; B+Edge adds edge distance without light-chain information; B+Edge+AN further adds abnormal neighborhood; and Full combines all features. Adding edge distance and abnormal neighborhood without lightchain grounding substantially reduced specificity (0.693 and 0.445, respectively), whereas the full model achieved optimal performance across all metrics. (B) Predicted abnormal cell counts from the Cell-level module correlate strongly with expert manual counts across the full range of tumor burden, with R2= 0.964 (95% CL 0.939-0.988) and slope = 1.021 (95% CL 1.005-1.038) on log-transformed counts. Accurate quantification is demonstrated from fewer than 100 to more than 100,000 abnormal cells across peripheral blood (n=155), bone marrow (n=19), and tissue (n=177) specimens.

[0027] FIGs. 15A-15C: show interpretable visualization outputs that enable clinical verification. (A) Follicular lymphoma case (top row). The spider plot shows the predicted-9- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451abnormal population (red) with lambda light-chain restriction and elevated CD 10 expression compared to normal B-cells (blue); the outer ring indicates SHAP-derived feature importance. Biaxial plots reveal a discrete lambda-restricted, CDlO-bright population, and the Large Vis embedding shows flagged cells forming a tight, spatially distinct cluster. (B) Large B-cell lymphoma case (middle row). The spider plot shows more subtle separation between abnormal and normal populations, with SHAP importance distributed across multiple markers rather than dominated by a single feature. Despite less pronounced marker differences in biaxial plots, the Large Vis embedding demonstrates that flagged cells form a coherent cluster. (C) Chronic lymphocytic leukemia case (bottom row). The spider plot reveals the characteristic CD5-positive, CD20-dim phenotype of the abnormal population. SHAP analysis identifies CD5 and CD3 as important features, reflecting the model's use of CD3 negativity to distinguish CLL cells from CD5-positive T-cells. The Large Vis embedding shows a tightly clustered abnormal population clearly separated from normal B-cells.

[0028] FIGs. 16A and 16B : show abnormal cell features that distinguish disease subtypes. (A) Dimensionality reduction of averaged features from predicted abnormal populations in tissue specimens reveals interpretable clustering patterns across five B-cell neoplasm categories. CLL cases (green) form a tight, distinct cluster, FL cases (orange) group separately, and MCL cases (purple) occupy their own region. DLBCL (red) and other low-grade B-cell lymphomas (blue) show more overlap with each other and with other subtypes, consistent with the known phenotypic heterogeneity of these entities. (B) Confusion matrix for subtype classification using a classifier trained on features extracted from predicted abnormal populations (AUROC = 0.925). CLL achieved perfect classification (32 / 32), followed by FL (43 / 49) and MCL (8 / 12). Most misclassifications occur between DLBCL and other low-grade B-cell lymphoma categories, supporting that the system captures disease-relevant immunophenotypic patterns that could guide subsequent diagnostic workup.

[0029] FIG. 17 shows high-dimensional clustering methods on unannotated flow cytometry data.-10- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0030] FIG. 18 shows cell classifiers trained on annotated cells and aggregated for the sample-level prediction.

[0031] FIG. 19 shows a block diagram of an overview of a process for processing flow cytometry data using cell-level and sample-level modules.

[0032] FIG. 20 shows a block diagram of an example architecture for classifiers and modules.

[0033] FIGs. 21 A-21C: show graphs related to transformation and compensation preprocessing flow cytometry data.

[0034] FIG. 22 shows graphs of automated B-cell gating to preprocess clinical flow cytometry data before deep-learning inference.

[0035] FIGs. 23 A-23F : depict block diagrams of a process for processing flow cytometry data using a kappa-lambda (KL) classifier, cell-level classifier, and sample-level classifier modules.

[0036] FIG. 24A depicts graphs of annotation labels and model predictions by the KL classifier in a two-dimensional feature space. The label and the prediction may classify individual cells as kappa-expressing, lambda-expressing, or neither (double-negative).

[0037] FIG. 24B depicts a graph of a receiver operating characteristic (ROC) curve demonstrating the KL classifier model’s ability to discriminate.

[0038] FIG. 25A depicts graphs of annotation labels and model predictions by the KL classifier in a two-dimensional feature space. The label and the prediction may classify individual cells as normal or abnormal.

[0039] FIG. 25B depicts a graph of a receiver operating characteristic (ROC) curve demonstrating cell-level model’s ability to discriminate.

[0040] FIG. 26A depicts a confusion matrix comparing predicted versus true case labels (tumor vs. normal) generated by the sample-level classifier operating on cell-level events.-11- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0041] FIG. 26B depicts a receiver operating characteristic (ROC) curve demonstrating high discriminative performance (AUC).

[0042] FIGs. 27A-27C: depict confusion matrices of specimen-type level performance. The specimens include peripheral blood, bone marrow, and tissue specimens.

[0043] FIG. 28 depicts a graph of an abnormal population estimator based on celllevel abnormality predictions. The graph plots, on a log scale, the number of abnormal (tumor) B cells estimated by the cell-level module versus expert-annotated tumor cell counts across peripheral blood, bone marrow, and tissue specimens, with the identity line indicating perfect agreement and the reported R2demonstrating strong concordance for tumor burden quantification.

[0044] FIGs. 29A and 29B: depict graphs of biomarker outputs for clinical verification of model predictions. The graphs in FIG. 29A show representative marker-versus-marker dot plots and a two-dimensional embedding for a normal specimen. The graphs in FIG. 29B show an abnormal specimen. Each events colored by the predicted tumor likelihood score.

[0045] FIG. 30 depicts a graph of a cell-level module performance for abnormal B-cell identification. The graph shows the receiver operating characteristic (ROC) curve on held-out test cells for the cell-level classifier that distinguishes abnormal (tumor) B cells from normal B cells.

[0046] FIG. 31A depicts a confusion matrix comparing the predicted case labels (tumor vs. normal) with ground-truth labels for the sample-level classifier module.

[0047] FIG. 3 IB depicts a graph of a receiver operating characteristic (ROC) curve with area under the curve (AUC) for the sample-level classifier module.

[0048] FIG. 32 depicts a graph of a receiver operating characteristic (ROC) curve on held-out test cells for the cell-level classifier distinguishing abnormal (tumor) B cells from normal B cells with the incorporation of the KL light-chain classifier output and features, with the reported AUC indicating increased discriminative performance.-12- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0049] FIG. 33A depicts a confusion matrix comparing predicted versus true case labels (tumor vs. normal) for the sample-level classifier module.

[0050] FIG. 33B depicts a graph of a receiver operating characteristic (ROC) curve with AUC for the sample-level module with the incorporation of light-chain-informed celllevel scoring and reordering of events.

[0051] FIG. 34 depicts a block diagram of a system for classifying samples based on flow cytometry data in accordance with an illustrative embodiment.

[0052] FIG. 35 depicts a block diagram of a process to receive event datasets from flow cytometers in the system for classifying samples in accordance with an illustrative embodiment.

[0053] FIGs. 36A and 36B depict block diagrams of a process to train machine learning models in the system for classifying samples in accordance with an illustrative embodiment.

[0054] FIGs. 37A and 37B depict block diagrams of a process to determine probabilities of tumors in samples using the machine learning models in the system for classifying samples in accordance with an illustrative embodiment.

[0055] FIG. 38 depicts a block diagram of a process to generate outputs in the system for classifying samples in accordance with an illustrative embodiment.

[0056] FIG. 39 depicts a flow diagram of a method of classifying samples based on flow cytometry data in accordance with an illustrative embodiment.

[0057] FIG. 40 depicts a block diagram of a server system and a client computer system, in accordance with one or more implementations.DETAILED DESCRIPTION

[0058] Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for classifying samples based on light chain expression features in flow cytometry data. It should be appreciated that various concepts -13- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451introduced above and discussed in greater detail below may be implemented in any of numerous ways, as the disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.

[0059] Section A describes a deep learning based automatic detection of minimal residual disease in acute lymphoblastic leukemia / lymphoma.

[0060] Section B describes a deep learning-based automatic detection of B-cell neoplasms in clinical flow cytometry.

[0061] Section C describes an automated detection of mature B-cell neoplasms by flow cytometry using deep learning.

[0062] Section D describes a deep learning-based automatic detection of B-cell neoplasms in clinical flow cytometry for diagnosis and monitoring of hematological diseases.

[0063] Section E describes systems and methods of classifying samples based on light chain expression features in flow cytometry data.

[0064] Section F describes a network environment and computing environment, which may be useful for practicing various computing related embodiments described herein.A. Deep Learning Based Automatic Detection of Minimal Residual Disease in Acute Lymphoblastic Leukemia / Lymphoma

[0065] The evaluation of minimal residual disease (MRD) in B-lymphoblastic leukemia (B-ALL) using flow cytometry methods has been widely used in clinical practice. However, this analysis is resource-intensive, time-consuming, and relies on extensively trained personnel. To address these challenges, a stepwise deep learning model may be used (referred herein as Flow ARC) to automatically detect MRD in B-ALL. First, the celllevel module, a cell classifier, is used to extract suspicious cells from the CD 19+ B-cell population of a case. From these cells, the sample-level module, a sample classifier, then-14- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451accurately detects the presence of MRD. Retrospectively from a clinical archive of 1681 cases from 333 patients, cell annotations were obtained on 137 B-ALL positive and 81 B-ALL negative cases by an experienced practicing pathologist.

[0066] Two Flow ARC models, Flow ARC -Manual and FlowARC-Auto, were trained based on manual or automated extraction of CD 19+ B-cells pre-analysis.Flow ARC -Manual achieved 100% accuracy in identifying all 24 B-ALL positive and 14 B-ALL negative cases out of 38 clinical test cases. FlowARC-Auto correctly classified 239 out of 248 B-ALL positive and 229 out of 232 B-ALL negative cases out of 480 clinical test cases, achieving an accuracy of 97.5%, sensitivity of 96.4%, and a specificity of 98.7%. In addition, Flow ARC includes an automated preprocessing pipeline, a visualization chart for tumor cell assessment and verification, and a quantification module for estimating the size of the tumor population. This groundbreaking MRD detection model for B-ALL surpasses existing models and has the potential to revolutionize MRD assessments by serving as an assistant tool for hematopathologists, resulting in significant cost reduction and faster turnaround time.1. Introduction

[0067] Minimal / measurable residual disease (MRD) serves as an independent prognostic metric, indicating the residual cancer cells that may persist post-treatment, even when the patient appears to be in remission. MRD is defined as any detectable disease below a 5% abnormal cell threshold that could not be identified by only morphological assessment. MRD detection is crucial for assessing the efficacy of treatment and guiding subsequent therapeutic strategies, especially for patients with hematopoietic malignancies such as acute lymphoblastic leukemia, acute myeloid leukemia, and multiple myeloma. Consequently, an accurate MRD detection workflow is vital for any clinical hematopathology laboratory.

[0068] MRD detection can be carried out using either multi-parameter flowcytometry (MFC) or PCR-based molecular techniques that provide assay detection sensitivities as low as 10'6. MFC allows for the simultaneous evaluation of multiple markers on individual cells, generating data on thousands to millions of cells from a single case1. Compared to molecular MRD diagnostics, MFC is faster, more cost-effective, and -15- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451does not require access to the patient’s diagnostic sample. However, its specificity and sensitivity were not as good as what can be achieved through molecular techniques historically2'4. In recent years, the performance of MFC analysis has been enhanced by incorporating additional concurrent markers, particularly in B-lymphoblastic leukemia (B-ALL), making it a preferred technique.

[0069] Regrettably, these advancements in MFC analysis have made it more labor-intensive and reliant on the experience of the analysts. Therefore, development of automated solutions for accurate MRD analysis using flow cytometry data is a critical clinical need to broaden the accessibility and availability of this technology.

[0070] Machine learning methods can automatically extract relevant patterns directly from the data and provide objective interpretation. Therefore, automation tools are widely utilized. In MFC analysis, ML-based tools can be used to automatically extract complex patterns from the vast datasets. This, in turn, enables the detection of subtle shifts in cell populations and the identification of unknown or rare cell populations. Unsupervised ML methods like flowSOM5, FlowMeans6, FLOCK7, and SWIFT8have demonstrated their ability to enhance cellular heterogeneity resolution and group similar cells by projecting data in high-dimensional space. However, these methods primarily focus on grouping cells with similar characteristics and have not yet demonstrated the capability to adequately account for rare cell populations, which are crucial for MRD detection.

[0071] Supervised ML methods that utilize the available diagnostic sample labels have been shown to improve the detection of the rare populations using MFC data. For example, CellCNN, a deep learning model adapted for cellular data, can identify disease-associated cell subsets at very low frequencies, sometimes as rare as 0.01% within small flow cytometry datasets9. Deep CNN10and Cell Scoring Neural Network (CSNN)11, like CellCNN, use the diagnostic sample labels to extract scores at the cell-level that are then aggregated for a sample-level prediction. These supervised models represent a significant advancement, compared to fully unsupervised models, in automated MRD detection for leukemias. However, their reported accuracies still fall short of the clinically acceptable level required for MRD detection.-16- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0072] Flow ARC may be used for accurate and automated MRD detection using flow cytometry data. In contrast to previous methods, Flow ARC uses both cell annotations and sample labels. It utilizes a multi-level classification strategy that combines cell classification with subsequent sample classification to improve MRD detection accuracy. Flow ARC was implemented and evaluated on the in-house B-ALL clinical archive containing 1681 flow cytometry cases.2. Results2.1 Study design and overview

[0073] A retrospective cohort of patient cases tested for B-ALL by an 8-color flow cytometry assay at the institution between 2015 and 2019 was gathered. Cases with multiple diagnoses were then filtered out, resulting in a total of 1681 cases from 333 patients (Male: 184, Female: 149). The average patient age was 32.8 with a standard deviation of 22.2. Flow ARC was applied to this clinical archive with 1150 normal cases and 531 B-ALL positive cases.

[0074] As Flow ARC requires at least a subset of the initial dataset to have explicit cell-level annotations, a sub-cohort of 218 cases from the clinical archive were manually annotated by an experienced pathologist to obtain their CD 19+ B-cell population and abnormal / tumor B-cell population. This manually annotated cohort includes 81 normal cases and 137 positive cases for BALL.

[0075] The workflow of Flow ARC includes four stages. In the first stage, an initial data preprocessing pipeline performs standardization of the flow data, followed by gating to obtain the CD19+ B cell population (FIG. 1 A). In the second stage, a cell-level module is trained to classify cells as either tumor or normal (FIG. IB). In the third stage, the cells within the preprocessed flow cases are ranked based on their probability of being tumorous, and the input data was truncated to a set number of cells which are most likely tumorous (FIG. 1C). Finally, in the fourth stage, a sample-level module is trained to classify these reordered and truncated cases for MRD diagnosis (FIG. ID).

[0076] In the preprocessing stage, two approaches are employed for CD 19+ B-cell gating: manual gating by pathologists and automated gating using flowDensity12, a -17- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451published algorithm for density-based cell population identification. The manual B-cell gates are only available for the annotated cohort cases. Hence, for all clinical archive cases, automated B-cell gating is performed. Recognizing that different gating techniques may introduce systematic differences, two models were built: Flow ARC -Manual for manually gated patient cases from the annotated cohort, and FlowARC-Auto for auto-gated patient cases from the clinical archive.

[0077] The normal and B-ALL positive (tumor) cases in the annotated cohort are split (75% / 25%) into training / testing for Flow ARC -Manual. All cells within each of these sets have manual annotations and become the training / testing set for the cell-level module of Flow ARC -Manual. For FlowARC-Auto, the normal cases in the clinical archive are split (65% / 35%) into normal training / testing cases, whereas all B-ALL positive cases lacking cell annotations (75%) are set as the B-ALL positive testing cases. For the cell-level module of FlowARC-Auto, all cells within the auto-gated normal training / testing cases are annotated as normal training / testing, and the B-ALL positive cases with manual cell annotations are split (75% / 25%) into tumor training / testing cells.

[0078] The annotated cells for both Flow ARC models train and validate their respective cell-level module. Further, these annotated cells also generate synthetic datasets to train and validate their sample-level module. The synthetic datasets encompass thousands of cases (training / validation / testing: 15,000 / 6,000 / 9,000) with a vast range of MRD levels that are generated by randomly mixing annotated (training / training / testing) cells in varying proportions to mimic the composition of real-world B-cell population. Half of the cases in each synthetic dataset are B-ALL positive, which are further divided into four equal subsets with their total tumor cells in the ranges: 25 to 500 (low MRD), 500 to 25K (high MRD), 25K to 50K (low tumor), and 50K to 500K (high tumor).

[0079] Additionally, it is noted that only a set number of cells per case are input to the sample-level module rather than the whole parent cell population. This number is considered a hyperparameter and has been set at 7500, as it yielded the optimal validation area under the ROC curve. Consequently, a minimum of 7500 B-cells per flow case is required for obtaining the sample-level module prediction.-18- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0080] The cell-level module is based on the ID ResNet-18 architecture, as implemented by Lang et al.13, without modification. The sample-level module is based on the ResNet-10114 architecture without modification. Both modules output two probabilities: normal and B-ALL positive / tumor, and are optimized using the area under the ROC curve of the validation set. To evaluate the performance of each module, metrics including the area under the ROC curve and the R2are used. Each module is generated five times using different random initializations. These iterations are then bootstrapped to obtain a Confidence Interval (CI). The results obtained from both Flow ARC models are presented for comprehensive analysis.2.2 Flow ARC makes accurate automated prediction of MRD.

[0081] On the synthetic testing set of 9000 cases, Flow ARC -Manual achieves an area under the ROC curve, AUCROC of 0.988 (95% CI, 0.980-0.996). Using the model with the best validation AUCROC, an exemplary 99.0% accuracy, 99.0% sensitivity, and 99.0% specificity were achieved. The associated confusion matrix with 4457 true normal, 4457 true tumor, 43 false normal, and 43 false tumor cases is displayed in FIG. 2A. This FlowARC-manual model is then tested on the 38 manually gated patient cases of the annotated cohort test set with at least 7500 B-cells. All cases were correctly identified, as shown by the confusion matrix in FIG. 2B.

[0082] On the synthetic testing set of 9000 cases, FlowARC-Auto achieves an AUCROC of 0.994 (95% CI, 0.992-0.996). Using the model with the best validation AUCROC, a 98.8% accuracy, 98.5% sensitivity, and 99.1% specificity are achieved with the associated confusion matrix with 4459 true normal, 4431 true tumor, 41 false normal, and 69 false tumor cases displayed in FIG. 2C. This FlowARC-Auto model is then subsequently tested on the auto-gated original patient cases from the test set of the clinical archive. The test set encompasses a total of 480 cases — 232 normal and 248 as tumorcontaining a minimum of 7.5K cells within their B-cell gate. The model achieves an AUCROC of 0.990, with an accuracy of 97.5%, a sensitivity of 96.4%, and a specificity of 98.7%. The corresponding confusion matrix with 229 true normal, 239 true tumor, 3 false normal, and 9 false tumor cases is presented in FIG. 2D.-19- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0083] Both Flow ARC -Manual and FlowARC-Auto exhibit excellent performance in predicting MRD. To further evaluate the performance of Flow ARC, its prediction results with those of CellCNN were compared, the most sensitive previously reported model for detection rare abnormal cell detection in the literatures. For a fair comparison, use the exact B-cell population of the training and testing patient cases that were employed forFlow ARC -Manual and FlowARC-Auto to train two CellCNN models: one for manually gated cases and another for auto-gated cases.

[0084] For the manually gated test cases, CellCNN achieves an AUCROC of 0.942 (95% CI, 0.922-0.963). For the auto-gated test cases, CellCNN achieves an AUCROC of 0.869 (95% CI, 0.860-0.879). A comparison of the AUCROC values between CellCNN and Flow ARC is presented in FIG. 2E. Clearly, Flow ARC shows a significant improvement compared to CellCNN.2.3 Flow ARC shows capability to assist pathologists to visualize and verify suspicious cell population.

[0085] Flow cytometry data is visualized on many two-dimensional plots by demonstrating two markers on each plot. Pathologists diagnose MRD in B-ALL by visualizing a tumor cell cluster separated from the normal B-cells population on these plots. Here, similar plots were used to visualize and verify the most suspicious B-cell population within flow cases as predicted by the cell-level module. Dimension reduction is performed using the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm15to visualize the total high-dimensional information of all cells on a two dimensional plot. These CD 10 vs. CD20, CD45 vs. CD20, CD34 vs. CD38, CD 10 vs. CD38, CD34 vs. CD45, and t-SNE comp-1 vs. comp-2 plots are grouped together to make a visualization chart.

[0086] The prediction performance of the cell-level module is important for this chart. The Flow ARC -Manual cell-level module achieves a test area under the ROC curve of 0.978 (95% CI, 0.978-0.978) on the test cells. Similarly, the FlowARC-Auto cell-level module achieves a test area under the ROC curve of 0.948 (95% CI, 0.948-0.948). These results show that both of the cell-level modules can separate normal and tumor cells well. As FlowARC-Auto can be applied to any patient case, the capability of the visualization chart is demonstrated using its cell-level module.-20- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0087] The FlowARC-Auto cell-level module with the best area under the ROC curve on validation cells (random subset of training cells) is used on patient test cases of the clinical archive to generate their visualization chart. For true tumor cases, there is a distinct clustering of cells with very high tumor likelihood (darker cells shown in FIG. 3 A) in most of the plots. In true normal cases, even though there can be cells with higher tumor likelihood, their arrangement is relatively diffuse in most plots (FIG. 3B). With this expected pattern, the visualization chart can be used as a prediction verification tool.Furthermore, the cases mentioned that with less than 7500 cells in the B-cell gate, it cannot be injected into the sample-level module. However, these cases can be evaluated by visualizing the clustering pattern for the true tumor cases and the scattered pattern for the true normal cases. (FIG. 3C). This method allows us to be able to assign a prediction to the cases with fewer B-cells, which is not uncommon post therapy.2.4 Flow ARC shows capability to predict the tumor content.

[0088] The enumeration of tumor cells within a positive flow cytometry case is an essential clinical parameter for disease management. To enhance the capability of Flow ARC in providing precise estimations of tumor cell prevalence, a quantification module may be designed focused on population estimation. The architecture of the quantification module is the exact ResNet-10114model used in the sample-level module, with the output changed to a number between 0 and 7500. This module is trained and validated on the B-ALL positive cases of the synthetic dataset, and is optimized using the validation coefficient of determination, R2.

[0089] To achieve the best performance, an integrated approach may be designed for cases with different tumor levels. Every B-ALL positive case is first analyzed by the quantification module. If the subsequent prediction is below 7200 cells, the total tumor burden is accepted. For the cases predicted to have more than 7200 cells, the number of cells predicted to be positive were aggregated by cell level model to provide an estimate of the total tumor burden.

[0090] For the Flow ARC -Manual synthetic testing tumor set, R2is 0.998 (CL 0.998 - 0.999), indicating an exceptional alignment between the true and predicted tumor population. Additionally, the mean absolute error (MAE) is found to be 0.012 (CL 0.010 - -21- 4930-6229-2378.1Atty. Dkt. No.: 115872-34510.014), further attesting to the accuracy and precision of the quantification module. Using the model with the best validation R2, the comprehensive evaluation of the true and a predicted tumor population is plotted in FIG. 4A. For cases with predicted tumor population above 7200 cells, the estimation from the cell-level module achieves R2of 0.998 as shown in FIG. 4A.

[0091] For the manually gated real -world test tumor cases, the quantification module with the best validation R2achieves, coefficient of determination, R2, of 0.922 and a mean absolute error (MAE) of 0.048, signifying a good estimation of tumor burden. The correlation between the true and predicted tumor population is graphically depicted in FIG.6 (left). For cases with a predicted tumor population above 7200 cells, the estimation from the cell-level module achieves R2of 0.997 as shown in FIG. 4B (right).2.5 FlowARC explainability analysis

[0092] Here, the inner workings of the cell-level and sample-level modules were explored, and analyze their prediction pattern in more detail.

[0093] First, the classifications of the cell-level module (Flow ARC -Manual) are visualized in the set of marker vs. marker plots using a random representative subset of the test cells in FIG. 5A. The visual representation employs color and dot size to reflect the classification outcomes from the Cell-level module, as shown in the legend at the bottomright of the figure. It becomes apparent upon examination that most misclassifications are localized around the boundaries separating the two cell populations.

[0094] For a more nuanced understanding of the operational mechanisms of the Cell-level module, the Shapley Additive exPlanations (SHAP) algorithm16a widely used method for elucidating the decision-making processes of deep-learning models were applied. The SHAP algorithm dissects the model’s predictions by pinpointing the contributing input features and appraising their relative significance in shaping the confidence of the classification. By utilizing SHAP values that are rooted in game theory, the algorithm imparts a quantifiable importance to each feature within the model. Features with positive SHAP values positively influence tumor predictions, while those with negative values have a negative influence.-22- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0095] The SHAP summary plot offers a synthesis of feature importance and their respective effects. On this plot, each data point signifies a SHAP value for a feature in relation to a specific instance (cell). The summary plot, presented in FIG. 7, showcases these SHAP values for two distinct scenarios: Tumor prediction and Normal prediction. The features, or cell markers, are arrayed in rows according to their significance, with SHAP values that overlap being jittered along the y-axis. The color bar encodes the input value of the feature. From the plot, a lower forward scatter area (FSC-A) coupled with a higher forward scatter height (FSC-H) is more indicative of a normal cell prediction, while lower expression of CD38 is more suggestive of a tumor cell prediction that may be discerned.

[0096] The sample-level module is a robust neural network capable of extracting rich and valuable features from its input. These extracted features are then used to classify each case in the model’s final layer. Understanding how these features differ for each class may provide us insight into the model itself. To this end, t-SNE is applied to project the penultimate 2048-dimensional feature layer into a lower-dimensional space and plotted in FIG. 7. The normal and tumor cases are well separated, signifying the Sample-level module’s capacity to accurately differentiate and predict the two distinct classes.3. Discussion

[0097] Flow ARC represents a significant advancement in the automatic and accurate detection of minimal residual disease (MRD) using flow cytometry data. The distinguishing feature of this model lies in its multi-step approach, wherein the challenge posed by a large number of randomly ordered cells within each flow case is first addressed by a ID cell-level module. This module truncates each case into an ordered table, consisting of the most tumor-like cell subset, which is then thoroughly analyzed by a 2D sample-level module to provide an accurate prediction. This approach overcomes the limitations of previous works, which either rely on unsupervised clustering methods that may miss crucial rare cell populations or use neural networks to extract single-cell information without effectively utilizing the 2D information within the cases. In contrast, Flow ARC leverages the 2D information among the relevant subset of cells to accurately predict the presence of minimal residual disease.-23- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0098] Flow ARC comprises a ID cell-level module trained on manually annotated cells and a 2D sample-level module trained on annotated cases. Both modules make use of the powerful ResNet architecture, a state-of-the-art deep learning model. Additionally, Flow ARC incorporates a clinically relevant estimator for the tumor population. As a supervised deep learning model, Flow ARC requires a substantial amount of annotated data for training, including millions of annotated cells and thousands of annotated cases.However, the process of obtaining these annotations can be time-consuming and may present a bottleneck. To address this issue, FlowARC can be trained using synthetic cases generated from a smaller set of annotated cases. These synthetic cases closely mimic real-world cases and cover a wide range of tumor population sizes, including low levels of MRD and advanced cases of the disease. This approach significantly reduces the training requirements and enhances the accessibility of FlowARC.

[0099] Using a clinical archival dataset of B-ALL cases comprising 1681 cases, a FlowARC model was developed for the automatic detection of MRD. The model achieved highly accurate results on real-world patient cases, demonstrating an accuracy of 97% on a test set of 480 cases. The model was trained on manual annotations of 103 B-ALL cases, while the remaining training set was auto-gated. These results highlight that even a small set of annotated flow cases for a specific disease panel is sufficient to train a FlowARC model with clinical-grade performance. This has the potential to significantly reduce the turnaround time for flow MRD analysis, ultimately leading to improved patient outcomes.

[0100] It is important to note that MRD detection using flow cytometry panels requires a minimum prevalence of tumor cells, known as the sensitivity of the flow test. In the synthetic case generation process, a minimum tumor population of 25 cells was set, which corresponds to a test sensitivity of up to 10'5depending on the number of cells tested. It should be acknowledged that in FlowARC, the false negative rate tends to increase as the tumor population decreases. Therefore, this minimum tumor population can be adjusted to favor either higher test sensitivity or a lower false negative rate in FlowARC. Additionally, the model can be fine-tuned during training to achieve a lower false negative rate by assigning greater weight to the tumor cases. This flexibility to adapt to different use cases is a significant strength of FlowARC.-24- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0101] Furthermore, Flow ARC is a generic model that can be trained to detect and estimate MRD for any leukemia flow cytometry test. It can also be applied to modem flow cytometer panels with a larger number of features, including mass flow cytometers.Consequently, Flow ARC has the potential to automate MRD detection workflows and find utility in various clinical settings.

[0102] The proficiency of the Flow ARC system was successfully highlighted in pinpointing aberrant precursor B cells within the confines of the CD 19+ B cell population. A pivotal need for the identification of the infrequent CD19-B-cell acute lymphoblastic leukemia (B-ALL) phenotype was recognized, which has become increasingly relevant due to the escalated adoption of anti-CD19 immunotherapies in the treatment of patients with relapsed and refractory B-cell acute leukemias. These therapies can lead to the emergence of CD19-leukemic cells as a mechanism of disease escape, necessitating a reliable method for their detection. To address this clinical challenge, an enhanced B-ALL flow cytometry assay was developed, which has been thoughtfully designed to incorporate the markers CD22 and CD24. This strategic inclusion is intended to enable the robust detection of CD 19- leukemic cells, thereby closing a critical diagnostic gap and offering a more targeted approach to patient assessment and treatment planning. Looking to the future, the diagnostic capabilities are planned to be further refined by extending the application of the general model, Flow ARC, to encompass a CD19-assay. This test may be specifically devised to identify B-cell malignancies that have evaded CD19-targeted treatments. By adapting and applying the versatile Flow ARC algorithm to this novel assay, its sensitivity and specificity may be enhanced. This forward-thinking approach aims to provide a more comprehensive diagnostic tool, ensuring that even as treatment strategies evolve, the ability to detect and monitor B-cell malignancies remains at the cutting edge of hematologic diagnostics.4. Methods4.1 Data collection

[0103] A total of 1681 bone marrow biopsy specimen tested exclusively for B-lymphoblastic leukemia using flow cytometry were included in this retrospective study. The B-ALL flow assay includes the following surface markers: CD20, CD34, CD 10, -25- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451CD33, CD58, CD45, CD19, and CD38. The forward scattering channels (FSC-H and FSC-A), side scattering channels (SSC-H and SSC-A) are also measured by the flow cytometer. For each specimen, a flow cytometer measured these 12 channels for all cells and saved the results and metadata as an FCS file.4.2 Initial data preprocessing

[0104] The preprocessing pipeline is created using R programming language. The FCS files are read and subject to compensation and transformation using flowCore package. Here, compensation eliminates fluorescence spillover from the other dyes for each marker, while transformation condenses the data’s large dynamic range into a standardized range, reducing the impact of intrinsic variance in fluorescence signals and aiding in the visualization of discrete positive and negative populations. Compensation is applied using the spillover matrix obtained from the FCS file metadata. Subsequently, the SSC-H channel is log-transformed, the other three scattering channels undergo linear transformation, and the eight fluorescence channels are transformed with the logical function. This initial preprocessing of the FCS files is followed by B-cell population gating.4.3 B-cell population gating

[0105] The CD 19+ B-cell gating is either done manually by a trained pathologist or by using automated density-based gating tool, flowDensity.

[0106] The manual gating procedure in this flow assay follows a series of steps. Firstly, doublets and margin events are removed using the FSC-A vs. FSC-H plot. Next, a viable mononuclear cell population is selected using the SSC-H vs. FSC-A plot. Finally, the CD19+ B-cell population is identified using the SSC-H vs. CD19 plot.

[0107] In flowDensity, this manual gating process was replicated for automated gating. It requires the implementation of a gating strategy with various thresholds based on the density distribution of markers, a four-step approach was adopted, as illustrated in FIG.7. Step 1 involves the exclusion of margin events and doublets. Steps 2 and 3 focus on selecting the viable mononuclear population. In Step 4, the CD 19+ B-cell population is specifically chosen. Exclusion and selection of populations can be improved by changing the flowDensity parameter thresholds. Using 15 randomly chosen cases with available -26- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451manual gate, the parameter thresholds based on visual comparison with the manual gate was iteratively optimized. Thresholds that generate comparable results between automated and manual gates for 10 out of 15 cases were able to be optimized. These parameters are frozen and incorporated into a single R script. This script can then be applied to the flow cases, ensuring consistent and efficient analysis.4.4 Datasets

[0108] Developing a Flow ARC model requires a group of patient cases and set of labeled cells. This study trains two specific Flow ARC models for manually gated and autogated cases. Flow ARC and Flow ARC -Manual are trained and tested on the clinical archive and the annotated cohort respectively. All 1681 cases from 333 patients make up the clinical archive dataset. A sub-cohort of 218 cases from the clinical archive make up the annotated cohort.

[0109] The annotated cohort includes 81 normal (B-ALL negative) cases and 137 tumor (B-ALL positive) cases. This cohort is manually gated by an experienced hematopathologist to obtain their B-cells. Then, abnormal blasts are extracted manually from the B-ALL positive cases are annotated as tumor cells while B-cells of normal cases are annotated as normal cells. A random 75% / 25% train / test split was performed, separately for the tumor and normal patient cases. The annotated cells within these train / test cases are the train / test labeled cells. The visualized breakdown of this annotated cohort is shown in FIG. 6.

[0110] The clinical archive contains 1150 normal cases and 531 tumor cases. This cohort is auto-gated using the flowDensity script to obtain their B-cells. Then, the B-cells of the normal cases are annotated as normal cells. A random 65% / 35% train / test split of the normal patient cases was performed. As the extraction of abnormal blasts must be done manually, the previous manually gated tumor train / test cases was used for the annotated tumor cells. Now, the auto-gated tumor cases from this dataset, excluding the tumor cases shared with annotated cohort, are added to the patient test cases. The visualized breakdown of this clinical archive is also shown in FIG. 6.4.5 Synthetic case generation-27- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451[01U| Flow ARC utilizes synthetic datasets with huge number of synthetic cases to train, test and validate the sample-level module. These cases have a broad-range of MRD levels and mimic the parent cell-population (CD 19+ B-cells). The training, testing and validation synthetic datasets are generated using the labeled train, test, and train cases / cells respectively.

[0112] In a synthetic dataset, half of the cases are tumor synthetic cases and half of the tumor cases are MRD cases. Each MRD and advanced tumor cases are further divided into low and high categories. Low MRD and high MRD cases contain 25 up to 500 and 500 up to 25K tumor cells. Similarly, low tumor and high tumor cases contain 25K up to 50K and 50K up to 500K tumor cells. This categorical separation of the synthetic dataset is shown in FIG. 8A, panel A.

[0113] For the normal synthetic cases, only normal B-cell population is needed. A subset of normal cases from the labeled set with combined cell count exceeding 10K was randomly selected. These selected cases are then coalesced into a single table. From this table, A random cell subset with cell count (N) randomly chosen between 10K and 50K was randomly chosen. This subset is then saved as a normal synthetic case. This process, which aims to replicate the variety and variable cell count observed in real B-cell populations, is iterated until the desired number of normal synthetic cases are created. The process of normal synthetic case generation is illustrated in FIG. 8B.[0H4] For the tumor synthetic cases, both normal and tumor B-cell populations are required, the normal population was generated, with cell count (NN) randomly chosen between 10K and 50K, in the exact way as the normal synthetic cases. For the abnormal cell population, there are four different categories: Low MRD, high MRD, low tumor and high tumor. For each category, A cell count (NA) from their category- specific range was randomly chosen. Now, A tumor case with number of cells greater than is randomly selected, but not exceeding, three times NA. From this chosen tumor case, NA number of tumor cells are randomly chosen to get the tumor population. Both normal and abnormal populations are then combined and saved as a tumor synthetic case (N=NN+NA). This process is iterated for each category until the desired number of tumor synthetic cases are created. The process of tumor synthetic case generation is illustrated in FIG. 7.-28- 4930-6229-2378.1Atty. Dkt. No.: 115872-34514.6 Cell-level module architecture and training

[0115] The architecture of the cell-level module is based on ResNet-18 model, adapted to ID where every 2D layer is replaced with its ID pendant. This ID ResNet-18 implementation is not pretrained and can handle any input size without modification. So, no further modifications are made for its utilization in this work. The cell-level module is trained on the annotated normal and tumor training cells, with an input size of (1 x 12) representing the 12 markers of each cell, and an output of (2 x 1) providing probabilities for the normal and tumor classes.

[0116] During training, a combination of the cross-entropy (CE) loss and Orthogonal Projection Loss (OPL) is employed as the loss function.

[0117] The cell-level module is trained on the labeled training cells for 100 epochs with a batch size of 800 cells. As a validation set, a subset of 2,000,000 training cells (1 : 1 normal and tumor cells) is randomly selected. A stochastic gradient descent optimizer with a momentum of 0.9 and a weight decay of le-3 is used. A decaying learning rate, starting at 0.1 and decreasing by 10% at epoch 5, 10, 30, 50, and 70 is utilized. The hyperparameter selection is based on the best validation area under the ROC curve. The model at 100 epoch is saved as the trained model.

[0118] Different cell-level modules for Flow ARC and Flow ARC -Manual were trained. The labeled cells for clinical archive contains significantly more normal cells compared to the labeled cells for annotated cohort. To mitigate this bias towards the normal cells, An additional imbalanced dataset sampler for Flow ARC was employed that rebalances the class distributions at every batch during training.4.7 Sample-level module architecture and training

[0119] The sample-level module architecture is based on the ResNet-101 model with random weight initialization. This ResNet-101 model can handle any three channel 2D input. No further modifications are made for its utilization in this work.

[0120] The reordered and truncated synthetic cases has input size of (1 x 7500 x 12). This input matrix is Fast Fourier transformed. The input matrix, alongside the real and -29- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451imaginary components of its Fourier transform, are compiled into a three-channel input of size (3 x 7500 x 12) compatible with ResNet-101. The model output is of size (2 x 1) providing probabilities for the normal and tumor classes.

[0121] Different sample-level modules for Flow ARC and Flow ARC -Manual were trained. The sample-level modules are trained on the synthetic training cases for 50 epochs with a batch size of 10 cases using the combined CE + OPL loss function. A stochastic gradient descent optimizer with a momentum of 0.9 was used. A decaying learning rate, starting at 0.05 and decreasing by 10% at epoch 5, 10, 15, 20, 25, 30 and 35 is utilized. Only for Flow ARC -Manual, a weight decay of le-4 in the optimizer was added. The hyperparameter selection is based on the best validation area under the ROC curve. The model at 50 epoch is saved as the trained model.4.8 Tumor estimator module architecture and training

[0122] The architecture of the tumor estimator module is the exact ResNet-101 model used in the sample-level module, with the output size changed from (2 x 1) to (1 x 1). This single value between 0 and 1 is further multiplied by 7500 as the final output.

[0123] Tumor estimator modules are trained for the Flow ARC -Manual. The tumor estimator module is trained on the synthetic B-ALL positive training cases for 50 epochs with a batch size of 10 cases using the Huber loss function. A stochastic gradient descent optimizer with a momentum of 0.9 is used. A decaying learning rate, starting at 0.05 and decreasing by 10% at epoch 5, 10, 15, 20, 25, 30 and 35 is utilized. The hyperparameter selection is based on the best validation area under the ROC curve. The model at 50 epoch is saved as the trained model.4.9 CellCNN

[0124] To make it comparable to Flow ARC, CellCNN is trained and evaluated on the B-cell population from patient cases. CellCNN is operated in the ‘outlier’ subset selection mode, which is recommended for MRD detection. To optimize the hyperparameters of CellCNN, the integrated CellCNN hyperparameter exploration function was utilized. This function selects the model that achieves the best predictive performance on the auto-generated validation cases. Similar to Flow ARC, each CellCNN model is -30- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451generated five times with different random initializations, and these iterations are then subjected to bootstrapping to obtain a Confidence Interval (CI).References

[0125] 1 Herzenberg, L. A., Tung, J., Moore, W. A., Herzenberg, L. A. & Parks, D. R. Interpreting flow cytometry data: a guide for the perplexed. Nat Immunol 7 , 681- 685 (2006). doi.org: 10.1038 / ni0706-681

[0126] 2 Thorn, I. et al. Minimal residual disease assessment in childhood acute lymphoblastic leukaemia: a Swedish multi-centre study comparing real-time polymerase chain reaction and multicolour flow cytometry. Br J Haematol 152, 743-753 (2011). doi.org: 10.1111 / j.1365-2141.2010.08456.x

[0127] 3 Denys, B. et al. Improved flow cytometric detection of minimal residual disease in childhood acute lymphoblastic leukemia. Leukemia 27, 635-641 (2013). doi.org: 10.1038 / leu.2012.231

[0128] 4 Gaipa, G. et al. Time point-dependent concordance of flow cytometry and real- time quantitative polymerase chain reaction for minimal residual disease detection in childhood acute lymphoblastic leukemia. Haematologica 97, 1582-1593 (2012). doi.org: 10.3324 / haematol.2011.060426

[0129] 5 Van Gassen, S. etal. FlowSOM: Using self-organizing maps for visualization and interpretation of cytometry data. Cytometry A 87, 636-645 (2015). doi.org: 10.1002 / cyto.a.22625

[0130] 6 Aghaeepour, N., Nikolic, R., Hoos, H. H. & Brinkman, R. R. Rapid cell population identification in flow cytometry data. Cytometry A 79, 6-13 (2011). doi . org : 10.1002 / cyto. a.21007

[0131] 7 Qian, Y. et al. Elucidation of seventeen human peripheral blood B-cell subsets and quantification of the tetanus response using a density-based method for the automated identification of cell populations in multidimensional flow cytometry data. Cytometry B Clin Cytom 78 Suppl 1, S69-82 (2010). doi.org: 10.1002 / cyto.b.20554-31- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0132] 8 Naim, I. et al. SWIFT-scalable clustering for automated identification of rare cell populations in large, high-dimensional flow cytometry datasets, part i: algorithm design. Cytometry A 85, 408-421 (2014). doi.org: 10.1002 / cyto.a.22446

[0133] 9 Arvaniti, E. & Claassen, M. Sensitive detection of rare disease-associated cell subsets via representation learning. Nat Commun 8, 14825 (2017). doi.org: 10.1038 / ncommsl4825

[0134] 10 Hu, Z., Tang, A., Singh, J., Bhattacharya, S. & Butte, A. J. A robust and interpretable end-to-end deep learning model for cytometry data. Proc Natl Acad Set U SA 117, 21373-21380 (2020). doi.org: 10.1073 / pnas.2003026117

[0135] 11 Robles, E. E. et al. A cell-level discriminative neural network model for diagnosis of blood cancers. Bioinformatics 39 (2023).doi.org: 10.1093 / bioinformatics / btad585

[0136] 12 Malek, M. etal. flowDensity: reproducing manual gating of flow cytometry data by automated density -based cell population identification. Bioinformatics 31, 606- 607 (2015). doi.org: 10.1093 / bioinformatics / btu677

[0137] 13 Lang, N. et al. Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles. Remote Sensing of Environment 268, 112760 (2022).

[0138] 14 He, K., Zhang, X., Ren, S. & Sun, J. in Proceedings of the IEEE conference on computer vision and pattern recognition. 770-778.

[0139] 15 Hinton, L. v. d. M. a. G. Visualizing Data using t-SNE. Journal of Machine Learning Research 9, 2579-2605 (2008).

[0140] 16 Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. Advances in neural information processing systems 30 (2017).B. Deep Learning-Based Automatic Detection of B-Cell Neoplasms In Clinical Flow Cytometry-32- 4930-6229-2378.1Atty. Dkt. No.: 115872-34511. Introduction

[0141] Flow cytometry plays a pivotal role in diagnosing B-cell lymphoma in clinical settings. However, its data analysis is labor-intensive and requires highly specialized personnel. Rapid flow cytometry results are in many clinical scenarios, assisting in immunohistochemical staining, molecular / cytogenetic testing, and patient management. While automated flow cytometry data analysis has become more common in research, clinical-grade applications remain limited due to the inability of existing software to detect clinically significant abnormal populations reliably. This disclosure addresses this gap by developing a deep learning model designed to automate the detection of abnormal B-cell populations in clinical flow cytometry data.2. Design

[0142] A retrospective cohort of 1109 tumor-positive and 823 tumor-negative cases from 1742 unique patients tested for B-cell flow cytometry were collected. Manual celllevel annotations were gathered for each case. Data preprocessing included compensation, transformation, and automated gating for B cells. The dataset was divided into training (70%) and testing (30%) groups for both tumor and normal cases. A cell-level classifier was trained to identify tumor cells using the manually annotated training dataset, followed by a case-level classifier trained with synthetic sample sets generated by randomly mixing normal and tumor cells from the training dataset in varying proportions, mimicking authentic clinical samples. The final model was evaluated using the testing dataset. (FIG.11 A)3. Results

[0143] There are 53 million abnormal and 11 million normal B cells for training, 22 million abnormal and 5 million normal cells for testing (FIG. 10A). At the individual cell level, the model achieved an area under the ROC curve (AUC) of 0.953 on the testing set of 27 million cells. The abnormal cells predicted by the model form clusters in the abnormal sample while visualizing by a dimension reduction plot. (FIG. 10B) When applied to a test dataset of 525 clinical cases, the model demonstrated an AUC of 0.985. (FIG. 1 IB). The confusion matrix for case-level classification showed correct prediction -33- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451on 286 normal and 208 tumor cases, 12 false negative and 19 false positive predictions. (FIG. 11C)4. Conclusion

[0144] An automated flow cytometry analysis tool may be used for B-cell neoplasms. This model has can be incorporated into clinical workflow, vastly reduce the turnaround time, and increase the accuracy of flow cytometry testing.C. Automated Detection of Mature B-Cell Neoplasms by Flow Cytometry Using Deep Learning

[0145] Multiparameter flow cytometry is essential for diagnosing mature B-cell lymphomas, yet analysis remains largely manual, time-consuming, and subject to interoperator variability. An automated system was developed for B-cell neoplasm detection using a three-stage deep learning architecture that produces interpretable intermediate outputs. The Light-chain module classifies each cell’s surface immunoglobulin expression. The Cell-level module combines original flow parameters with engineered features capturing light-chain neighborhood and spatial position to identify abnormal cells. The Sample-level module renders case-level diagnoses from cells ranked by abnormality score. The system was evaluated on 3,070 clinical specimens from Memorial Sloan Kettering Cancer Center, comprising peripheral blood (n=l,369), bone marrow (n=316), and tissue (n=l,385) samples. The system achieved 97.5% case-level accuracy with 95.6% sensitivity and 98.9% specificity. Predicted abnormal cell counts correlated with expert manual counts (R2= 0.964 [95% CI: 0.939-0.988], slope=1.021 [1.005-1.038], intercepts.121 [-0.213 --0.030]). Ablation experiments demonstrated that engineered features provide complementary information: light-chain neighborhood grounds spatial features that otherwise introduce noise, enabling the full model to outperform any feature subset. The system generates visualization outputs including marker expression comparisons and dimensionality-reduced embeddings colored by abnormality score, allowing pathologists to verify that flagged populations form coherent clusters with phenotypes consistent with disease. Features extracted from abnormal populations also supported lymphoma subtypes prediction (AUROC=0.925). This work demonstrates that automated flow cytometry-34- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451interpretation can achieve high diagnostic accuracy while providing transparent, verifiable outputs suitable for clinical integration.

[0146] Flow cytometry analysis for B-cell lymphoma diagnosis currently requires expert manual review of complex multidimensional data, limiting throughput and introducing variability between operators and laboratories. An automated system was presented that achieves diagnostic accuracy comparable to expert analysis while producing intermediate outputs that pathologists can verify. Unlike previous approaches that function as opaque classifiers, the system shows which cells it flagged as abnormal and why, enabling pathologists to confirm that predictions correspond to biologically coherent populations. This transparency addresses a key barrier to clinical adoption of automated diagnostic tools and provides a framework for extending such systems to other areas of laboratory medicine.1. Introduction

[0147] Mature B-cell neoplasms as a group are the most common subtype of nonHodgkin lymphomas and are a significant source of cancer-related morbidity and mortality (1). Treatment decisions require accurate diagnosis: targeted therapies, risk stratification, and monitoring all depend on precise characterization of the underlying disease (1).Multiparameter flow cytometry is central to this process. By measuring surface antigen expression on individual cells, it enables rapid detection of clonal B-cell populations, distinction between reactive and neoplastic infiltrates, provides immunophenotypic basis for specific entity classification, and quantification of tumor burden (2). A particularly informative readout in routine B-cell immunophenotyping is surface immunoglobulin light chain expression. Through allelic exclusion, each normal mature B cell expresses either kappa or lambda light chains, so non-neoplastic B-cell populations are normally polytypic and typically show a consistent kappa: lambda distribution (approximately 1.5:1 to 2:1) (3). In contrast, mature B-cell neoplasms commonly show light chain restriction to either kappa or lambda, or in some cases absent or very dim surface light chains, and this uniformity is widely used as a practical proxy for clonality in clinical flow cytometry (2, 4).Interpretation requires context because germinal center and activated B-cell subsets can have reduced surface immunoglobulin, and some reactive conditions can yield skewed-35- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451ratios that mimic restriction (5). In settings ranging from initial evaluation of lymphadenopathy to assessment of treatment response, flow cytometry provides information that directly shapes clinical decisions (4).

[0148] Despite its central role in evaluating mature B-cell lymphomas, routine flow cytometry analysis remains largely a manual process. Analysts and pathologists examine dozens of two-dimensional dot plots, sequentially gating cell populations and integrating information across multiple markers to identify abnormal cells. This is time-consuming, labor-intensive, and difficult to scale in high-volume laboratories (6). The process is also subjective. Gating strategies vary across operators, and studies have documented substantial inter-laboratory variability that affects diagnostic consistency, staging, and assessment of treatment response (7). These problems are exacerbated with expanding panel sizes that include more markers and specimens containing more cells. A single case may now involve up to 50 parameters measured on 500,000 or more cells, which is far more information than can be efficiently evaluated by visual inspection of pairwise plots.

[0149] Efforts to automate flow cytometry analysis have progressed through several phases. Early work applied traditional machine learning methods, such as support vector machines and random forests, to summary statistics derived from manually gated populations (8). These approaches improved consistency but remained dependent on the same subjective gating they aimed to replace. Unsupervised clustering and dimensionality reduction tools, including FlowSOM, PhenoGraph, and t-SNE, enabled visualization of high-dimensional data and discovery of novel cell subsets (9-12). However, these methods required manual interpretation of results and lacked the reproducibility needed for standardized clinical reporting. More recently, deep learning models have achieved strong performance on disease classification tasks, including detection and subtyping of B-cell neoplasms (13, 14). Convolutional neural networks applied to transformed flow data have demonstrated accuracy comparable to expert analysis in controlled settings (15).

[0150] Despite these advances, no automated system has achieved accuracy and precision required for routine clinical deployment for B-cell neoplasm diagnosis due to multiple gaps: (a) current methods were developed on single-center cohorts with limited sample sizes, often restricted to cases with substantial tumor burden that are relatively -36- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451straightforward to classify. Few studies have validated performances across the full range of specimen types encountered in practice: peripheral blood, bone marrow, and tissue each present distinct challenges in background composition and cellular debris. Existing approaches typically produce case-level predictions without interpretable intermediate outputs, making it difficult for pathologists to verify why a specimen was flagged as abnormal. Lack of interpretable methods is a significant barrier to clinical adoption, where decisions must be defensible and errors traceable.

[0151] To address these gaps, an automated system was developed with three sequential stages. First, a Light-chain module classifies each B-cell’s surface immunoglobulin expression as kappa-positive, lambda-positive, or double-negative.Second, a Cell-level module identifies potentially abnormal cells by combining the original flow cytometry parameters with engineered features capturing light-chain neighborhood and position relative to population boundaries. Third, a Sample-level module evaluates cells ranked by abnormality score to render a case-level diagnosis. This design produces verifiable outputs at each stage. Predicted abnormal cell counts can be compared directly to manual assessment. Cells can be displayed on standard marker-versus-marker plots colored by abnormality score, allowing a review of whether flagged populations cluster coherently. Cases where predictions scatter across the plot rather than forming tight clusters warrant closer examination.

[0152] Ultimately, the end-to-end approach consists of a staged and interpretable system for automated B-cell neoplasm detection that produces verifiable intermediate outputs at the light-chain, cell, and specimen levels. The approach was evaluated on a large retrospective cohort spanning peripheral blood, bone marrow, and tissue specimens and show that it achieves clinical-grade diagnostic performance while providing visual and quantitative outputs that enable pathologist review. In addition, the system’s cell-derived features capture phenotypic structure consistent with known B-cell biology and support exploratory separation of major disease subtypes.2. ResultsStudy Cohort and System Architecture-37- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0153] A three-module system was developed that automates the full arc of flow cytometry analysis, from raw cellular measurements to case-level diagnosis (FIGs. 12A-F). Unlike other workflows that depend on operator-defined gates, the system learns directly from data, making it robust to the specimen quality variation common in clinical practice.

[0154] The system was evaluated on 3,070 clinical flow cytometry specimens collected: peripheral blood (n=l,369; 515 tumor, 854 normal), bone marrow (n=316; 66 tumor, 250 normal), and tissue (n=l,385; 600 tumor, 785 normal). Tumor cases required a minimum of 50 abnormal cells; cases with multiple concurrent diagnoses were excluded. Expert clinical annotations served as ground truth for both case-level diagnosis and celllevel phenotypic labels.

[0155] Raw flow cytometry files first undergo preprocessing to isolate B-cells via density-based gating (16) on CD 19 and CD3, followed by removal of doublets and margin events (FIG. 12A). The Light-chain module, a ID ResNet-18 classifier (17), assigns each cell’s surface immunoglobulin expression to one of four categories: kappa-positive, lambda-positive, double-negative, or contaminating cell (FIG. 12B). These predictions enable calculation of two engineered features within a two-dimensional LargeVis(18) embedding computed independently for each case (FIG. 12C). Light-chain neighborhood (LC-N) counts how many of a cell’s 30 nearest neighbors share the same light-chain prediction. This feature encodes a key diagnostic principle: abnormal B-cells typically cluster into light-chain restricted populations, while normal B-cells do not. By quantifying this pattern, LC-N translates expert reasoning into a form that the model can use. Edge distance measures each cell’s Euclidean distance to the boundary of the embedded population, normalized by total bounded area.

[0156] The Cell-level module, also a ID ResNet-18, combines the original flow parameters with LC-N and edge distance to predict whether each cell is abnormal (FIG. 12D). From these predictions, a third engineered feature is computed: abnormal neighborhood defined as the mean abnormality score among each cell’s 30 nearest neighbors.

[0157] The Sample-level module uses a 2D ResNet-101 for case-level diagnosis (FIG. 12E). Cells are ranked by their abnormality scores, and the top 2,000 are arranged -38- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451into a structured matrix. This intermediate representation distills each specimen into a fixed-size input, allowing the final classifier to focus on diagnostic patterns rather than raw cellular complexity. The system also generates visualization outputs including marker-versus-marker dot plots and Large Vis embeddings colored by abnormality score, allowing pathologists to review flagged populations (FIG. 12F).Deep Learning Model Performance

[0158] Across five independent runs, all three modules showed strong discriminatory performance (FIG. 13 A). The Light-chain module achieved an AUROC of 0.896 (95% CL 0.890-0.903), the Cell-level module achieved an AUROC of 0.965 (95% CL 0.961-0.969), and the Sample-level module achieved an AUROC of 0.988 (95% CL 0.982-0.993), indicating progressively improved discrimination from Light-chain to Celllevel to Sample-level inference.

[0159] Case-level performance on the held-out test set is shown in FIG. 13B. On the best-performing seed, the system achieved 97.5% overall accuracy (306 true positives, 430 true negatives, 14 false negatives, 5 false positives), with 95.6% sensitivity and 98.9% specificity. Across all five seeds, accuracy ranged from 96.8% to 97.9%. Performance was consistent across specimen types. Peripheral blood achieved 97.1% accuracy (94.2% sensitivity, 99.1% specificity). Tissue achieved 98.0% accuracy (98.0% sensitivity, 98.1% specificity). Bone marrow achieved 97.4% accuracy (88.2% sensitivity, 100% specificity).Feature Contribution and Cell-level Validation

[0160] Ablation experiments assessed how each engineered feature contributes to case and cell-level performance (FIG. 14A). The base model (B) using only original flow parameters achieved 95.9% case accuracy with 92.5% sensitivity and 98.4% specificity. Adding light-chain classification and LC-N features (B+KL) maintained accuracy at 95.9% while modestly increasing sensitivity to 93.7%.

[0161] Adding edge distance without light-chain information (B+Edge) reduced accuracy to 78.4%, with specificity dropping from 98.4% to 69.3%. Adding abnormal neighborhood on top of edge distance, still without light-chain grounding (B+Edge+AN),-39- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451further reduced accuracy to 65.7% and specificity to 44.5%. These spatial features capture real patterns in the data but introduce noise when not anchored by light-chain information.

[0162] The full model (Full) combining all features achieved the best performance: 97.5% accuracy, 95.6% sensitivity, 98.8% specificity. LC-N provides biological grounding that allows the spatial features to contribute positively. Edge distance identifies cells at population boundaries; LC-N determines whether those cells lie in a light-chain restricted region; abnormal neighborhood captures whether a cell resides within a cluster of other high-abnormality cells. Together these features provide complementary information that no single feature captures alone.

[0163] At the cell level, predicted abnormal cell counts correlated with expert manual counts across the full range of tumor burden (FIG. 14B). On log-transformed counts, the model achieved R2= 0.964 (95% CI: 0.939-0.988) with a slope of 1.021 (95% CI: 1.005-1.038) and an intercept of -0.121 (95% CI: -0.213 to -0.030), indicating accurate quantification from fewer than 100 to more than 100,000 abnormal cells per specimen. This correlation held across peripheral blood (n=155), bone marrow (n=19), and tissue (n=177) specimens.Interpretable Outputs Enable Clinical Verification

[0164] The system produces visualization outputs designed for pathologist review (FIGs. 15A-C). For each case, a spider plot compares average marker expression between predicted normal and abnormal populations, with line thickness scaled by population size. The outer ring indicates feature importance derived from SHAP (19) values computed on a random subsample of up to 5,000 cells per case and aggregated to a case-level summary. Standard biaxial dot plots and Large Vis (18) embeddings display all cells colored by predicted status or abnormality score, allowing pathologists to assess whether flagged cells form coherent populations consistent with known disease phenotypes.

[0165] Figure 4 shows representative examples from three tumor cases with distinct immunophenotypes. In a follicular lymphoma case (FIG. 15 A), the predicted abnormal population shows lambda light chain restriction and bright CD 10 expression. The spider plot highlights separation in CD 10, and SHAP analysis identifies CD 10 as the primary -40- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451driver. The biaxial plots confirm a discrete lambda-restricted population with uniformly high CD 10, and the Large Vis embedding shows these cells form a tight cluster separated from normal B-cells.

[0166] In a large B-cell lymphoma case (FIG. 15B), the predicted abnormal population shows more overlap with normal cells across individual markers. The spider plot reveals modest separation in CD38 and CD20, with SHAP importance distributed across multiple features rather than dominated by a single marker. The Large Vis embedding shows the abnormal population forms a coherent cluster despite the less pronounced marker differences visible in biaxial plots.

[0167] In a chronic lymphocytic leukemia case (FIG. 15C), the predicted abnormal population displays the characteristic CD5-positive, CD20-dim phenotype. The spider plot shows CD5 expression in the abnormal population exceeds that of normal B-cells, while CD20 is diminished. SHAP analysis identifies CD3 and CD5 as important features, reflecting the model’s use of CD3 negativity to distinguish CLL cells from T-cells that share CD5 expression. The Large Vis embedding reveals a distinct, tightly clustered abnormal population.

[0168] These visualizations allow pathologists to verify that flagged populations exhibit immunophenotypes consistent with the suspected diagnosis, or to identify cases where scattered predictions suggest potential misclassification.Abnormal Cell Features Distinguish Disease Subtypes

[0169] Flow cytometry alone does not provide definitive lymphoma subtype classification, which requires integration with morphology, immunohistochemistry, and molecular studies. However, the features was tested whether extracted from predicted abnormal cells contain subtype-relevant information.

[0170] Using averaged features from predicted abnormal populations in tissue specimens, a classifier was trained to distinguish five categories: chronic lymphocytic leukemia (CLL), diffuse large B-cell lymphoma (DLBCL), follicular lymphoma (FL), mantle cell lymphoma (MCL), and other low-grade B-cell lymphomas (including marginal zone lymphoma). This classifier achieved an AUROC of 0.925 (FIG. 16A and 16B).-41- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451[01711 Dimensionality reduction of the feature space revealed interpretable clustering patterns (FIG. 16A). CLL cases formed a tight, distinct cluster with no misclassifications (32 / 32 correct). FL cases grouped separately (43 / 49 correct). MCL occupied its own region (8 / 12 correct). DLBCL and other low-grade B-cell lymphomas showed more overlap with each other and with other subtypes, consistent with the known phenotypic heterogeneity of these entities. The confusion matrix (FIG. 16B) confirms that most misclassifications occur between DLBCL and the other low-grade B-cell lymphoma categories.

[0172] While not a substitute for the molecular and immunohistochemical studies required for definitive diagnosis, these results suggest the system captures phenotypic patterns that could help guide subsequent workup.3. Discussion

[0173] Automated flow cytometry interpretation has remained difficult despite two decades of effort. Previous approaches treated flow data as a generic classification task, achieving reasonable performance on curated datasets but struggling with the variability of routine clinical specimens. The system achieves 97.5% diagnostic accuracy across 3,070 clinical cases by combining domain-informed feature engineering with a staged architecture that produces verifiable intermediate outputs.

[0174] This study is innovated by bringing domain-informed feature engineering. The ablation results clarify why engineered features matter. Spatial features (edge distance and abnormal neighborhood) alone actually degraded performance when added without light-chain information, dropping accuracy from 95.9% to 65.7%. These features capture meaningful patterns: abnormal cells often cluster at population boundaries, and malignant populations tend to be spatially cohesive. But without the LC-N feature to distinguish clonal expansion from benign polyclonal variation, these spatial signals introduce more noise than signal. The LC-N feature provides the necessary grounding. It transforms edge position from an ambiguous observation into a meaningful indicator: a cell at a cluster boundary with high LC-N likely represents clonal disease, while a boundary cell with low LC-N is more likely normal variation. This interaction explains why the full model outperforms any subset of features.-42- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0175] The three-stage architecture serves a practical purpose beyond performance. Each module produces outputs that can be independently verified. Light-chain predictions can be checked against conventional kappa / lambda plots. Abnormal cell counts can be compared to manual assessment, and the R2of 0.964 indicates this comparison will generally agree. Case-level predictions can be reviewed alongside marker-versus-marker plots showing which cells were flagged and whether they cluster coherently. Pathologists are unlikely to trust a system that provides only a binary tumor / normal output without supporting evidence. The visualizations also provide a safeguard against false positives: pathologists can directly inspect whether flagged cells form coherent, phenotype-consistent populations versus scattered high scores that do not correspond to a plausible abnormal cluster.

[0176] Several design choices overcome the constraints of clinical flow cytometry. Large Vis was selected for dimensionality reduction because it scales to 30,000-50,000 cells without prohibitive memory requirements, unlike t-SNE or UMAP. The 30-neighbor parameter for light chain neighborhood features reflects the convention that 20 or more clustered cells typically define a distinct population; empirical testing confirmed this choice. The 2,000-cell case classifier input balances performance against the need to include cases with lower B-cell counts.

[0177] An implementation was envisioned as decision support rather than autonomous diagnosis. The system would generate preliminary reports containing the prediction, visualizations with cells colored by abnormality score, and immunophenotypic summaries of flagged populations. Pathologists would review this output and determine whether additional examination is needed. Analysis time would decrease from the current 10-20 minutes for complete manual gating to under 5 minutes, enabling same-day reporting while preserving clinical judgment. Automation would also standardize analysis across operators, reducing the inter-analyst variability that currently limits reproducibility.

[0178] The subtype classification results, while exploratory, suggest the system captures disease-relevant phenotypic information. CLL, follicular lymphoma, and mantle cell lymphoma formed distinct clusters in feature spaces, consistent with their known immunophenotypic signatures. Flow cytometry is usually not sufficient for a definitive -43- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451subtype diagnosis, which requires integration with morphology, immunohistochemistry, and molecular studies; except a few specific disease types, such as chronic lymphocytic leukemia. However, a preliminary phenotypic characterization from automated analysis could guide subsequent workup and help prioritize cases for expert review.

[0179] In summary, this work demonstrates that automated flow cytometry interpretation can achieve high accuracy when algorithmic design incorporates domain knowledge. The key is not simply applying deep learning to flow data, but engineering features that capture diagnostically relevant patterns and structuring the system to produce outputs that clinicians can verify. The goal is not to replace expertise but to extend it: accurate enough to reduce workload, transparent enough to permit verification, and practical enough to integrate into existing workflows.4. Materials and MethodsStudy Design and Dataset

[0180] This retrospective study analyzed 3,070 flow cytometry specimens from patients evaluated for mature B-cell neoplasms. The cohort comprised 1,369 peripheral blood samples (515 tumor, 854 normal), 316 bone marrow samples (66 tumor, 250 normal), and 1,385 tissue samples (600 tumor, 785 normal).

[0181] Tumor cases were restricted to specimens containing a single abnormal mature B-cell population identified during routine clinical analysis. Cases was excluded with immature B-cell neoplasms, those harboring more than one distinct abnormal mature B-cell population, and those with abnormalities in other lineages including T cells or plasma cells. Normal cases contained no detectable abnormal B-cell populations. These criteria yielded a dataset specific to mature B-cell lymphoma detection.

[0182] Expert annotations from prior clinical analyses served as ground truth for model training. Case-level labels reflected binary tumor / normal classifications from original clinical diagnoses. Cell-level annotations provided two types of classification: phenotypic assignment (abnormal B-cell, normal B-cell, plasma cell, or contaminating cells) and light chain expression pattern (kappa-positive, lambda-positive, double-negative, or junk). While automated gating using FlowDensity(16) aimed to isolate B-cells, plasma -44- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451cells and other contaminating events inevitably passed through, necessitating these additional phenotypic and junk categories. These multi-level annotations enabled training of the sequential classifiers.Data Preprocessing

[0183] Specimens were analyzed using the MSKCC 19-color BT tube panel (20): FSC-A, FSC-H, SSC-A, SSC-H, KAPPA / CD8, TCRyb / pi BB700, CD 10, CD7 APC-A700, CD45 APC-H7, CD26, CD4, CD3, CD14, CD5, CD38 BV480, CD279, CD20, TCRyb, CD2, LAMBDA / CD56, CD25 PE-Dazzle 594, CD22 PC5-5, CD 19 PC7. The raw cellular measurements were output as FCS files from the flow cytometer. These FCS files underwent compensation for spectral overlap using the available compensation matrices for the associated batch.

[0184] Transformation strategies matched each parameter’s properties. Forward scatter parameters (FSC-A, FSC-H) and SSC-A remained on linear scales to preserve cell size relationships. SSC-H used logarithmic transformation to better separate lymphocyte populations from debris. All fluorescence channels underwent biexponential transformation using the logicle (21) function. This approach maintained resolution near zero fluorescence while accommodating the full 5-decade dynamic range of modern cytometers.

[0185] Quality control removed several categories of problematic events. Doublets were excluded when FSC-A / FSC-H ratios exceeded population-specific thresholds calculated as median plus four median absolute deviations. Margin events at acquisition boundaries were filtered as potentially incomplete measurements.

[0186] A FlowDensity(16) (version 1.0) was used and implemented in R (version 4.3) to isolate B-cell populations from all blood cells through sequential gating steps. First, mononuclear cells were selected using SSC-H versus FSC-A parameters based on characteristic light scatter properties. B lymphocytes were then isolated using CD3 versus CD19 expression, selecting for CD 19-positive, CD3-negative populations. This algorithm applies threshold-based gates using data density distributions, providing more reproducible results than unsupervised clustering while following the sequential gating logic used by clinical laboratories. This preprocessing step removed most irrelevant cells while retaining -45- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451some contaminating cells for subsequent classifier management. Cases with fewer than 2000 B-cells were kept for the Light-chain and Cell-level module dataset but excluded from the Sample-level module dataset.Feature Engineering

[0187] LargeVis(18) dimension reduction generated case-specific embeddings capturing population structure within each specimen’s multidimensional flow cytometry space. The following was selected B-cell relevant features for embedding: FSC-A, FSC-H, SSC-H, CD10, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD20, CD22, and CD19. Parameters included: max iterations=1000, n_negatives=50, learning rate=1.0, optimizer Adam. The algorithm preserved local cellular neighborhoods while revealing global population architecture in two dimensions.

[0188] Edge features quantified each cell’s position relative to population boundaries. After Large Vis embedding, the alpha-shape was computed encompassing all cells using tolerance parameter=2.5. For each cell, a minimum Euclidean distance was calculated to the nearest hull boundary, then normalized by total hull area. Higher values indicated peripheral positions where abnormal cells often reside.

[0189] Light-chain neighborhood (LC-N) features measured local clonality. Using ball tree search in Large Vis space, each cell’s 30 was identified in nearest neighbors. The LC-N score equaled the count of neighbors sharing the same light chain prediction from the Light-chain module. Scores approaching 30 indicated light chain restriction suggestive of clonality, while scores near 15 reflected normal polyclonal distributions.

[0190] Abnormal nearest neighbor features provided local context after the initial classification. For each cell, the mean abnormality score across was computed to its 30 nearest neighbors, capturing the principle that malignant cells cluster together while scattered abnormal predictions likely represent classification noise.Deep Learning Architecture

[0191] The system comprised three sequential modules. The Light-chain module used a ID ResNet-18 (17) architecture to classify each B-cell’ s light-chain expression into -46- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451four categories: kappa-positive, lambda-positive, double-negative, or non-B-cell contaminant. Input consisted of the 23 flow cytometry parameters per cell.

[0192] The Cell-level module used a ID ResNet-18 architecture to predict whether each cell was abnormal. Input included the original flow parameters plus LC-N and edge distance features. From these predictions, a third engineered feature was computed: abnormal neighborhood, defined as the mean abnormality score among each cell’s 30 nearest neighbors. This feature was used by the Sample-level module but not by the Celllevel module itself.

[0193] The Sample-level module used a 2D ResNet-101 architecture. Cells were ranked by abnormality score from the Cell-level module, and the top 2,000 cells were selected to form a structured matrix of dimensions 2000 x [number of features]. This matrix was augmented with real and imaginary components of the 2D Fast Fourier Transform as additional input channels.Training Procedures

[0194] Models were trained sequentially, with outputs from earlier stages serving as inputs for subsequent stages. Data were split using stratified sampling to maintain consistent tumor / normal ratios and specimen type distributions: 50% training, 20% validation, 30% testing. Five independent runs used different random seeds (11, 12, 13, 14, 15) to generate distinct splits for performance estimation.

[0195] The Light-chain module was trained for 10 epochs with batch size 800. The SGD optimizer used learning rate 0.005 with momentum 0.9 and cosine annealing schedule (T max = 5). The Cell-level module was trained for 25 epochs with batch size 800 and the same optimizer settings (T max = 10). The Sample-level module was trained for 150 epochs with batch size 25, learning rate 0.005, and T max = 10. At the end of the training, the model was chosen at the epoch that achieved the best validation metric. For ID modules (Light-chain and Cell-level), (Sensitivity + Specificity -1) the validation metric was used, whereas for the 2D Sample-level module, was used (Accuracy + Specificity - 1) as the validation metric.-47- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0196] All modules used a combined loss function of cross-entropy plus orthogonal projection loss (22) (gamma = 0.5, alpha = 1.0). The orthogonal projection loss encourages learned feature representations of same-class cells to cluster together while pushing different-class representations apart, complementing the cross-entropy classification objective. For the Sample-level module, top 2,000 cells post-reordering were sampled per case. Performance was evaluated on the model achieving the best validation metric, with confidence intervals derived from all five runs.Statistical Analysis

[0197] Sensitivity, specificity, accuracy, and area under the receiver operating characteristic curve were calculated with 95% confidence intervals from bootstrap resampling. Cell-level predictions were evaluated using R2and mean absolute error comparing predicted to manual abnormal cell counts. Linear regression assessed the relationship between predicted and true counts. Feature importance was quantified using SHAP values for individual case interpretation.References

[0198] 1. R. Alaggio et al., The 5th edition of the World Health Organization Classification of Haematolymphoid Tumours: Lymphoid Neoplasms. Leukemia 6, 1720-1748 (2022).

[0199] 2. A. C. Seegmiller, E. D. Hsi, F. E. Craig, The current role of clinical flow cytometry in the evaluation of mature B-cell neoplasms. Cytometry B Clin Cytom 96, 20-29 (2019).

[0200] 3. D. C. Gajzer, J. R. Fromm, Flow Cytometry for B-Cell Non-Hodgkin and Hodgkin Lymphomas. Cancers (Basel) 17 (2025).

[0201] 4. F. E. Craig, K. A. Foon, Flow cytometric immunophenotyping for hematologic neoplasms. Blood 111, 3941-3967 (2008).-48- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0202] 5. S. L. Huang, T. Fennell, X. Chen, J. Z. Huang, The patterns and diagnostic significance of the lack of surface immunoglobulin light chain on mature B cells in clinical samples for lymphoma workup. Cytometry B Clin Cytom 104, 263-270 (2023).

[0203] 6. H. T. Maecker, J. P. McCoy, R. Nussenblatt, Standardizing immunophenotyping for the Human Immunology Project. Nat Rev Immunol 12, 191-200 (2012).

[0204] 7. G. Finak el al., Standardizing Flow Cytometry Immunophenotyping Analysis from the Human ImmunoPhenotyping Consortium. Sci Rep 6, 20686 (2016).

[0205] 8. N. Aghaeepour et al., Critical assessment of automated flow cytometry data analysis techniques. Nat Methods 10, 228-238 (2013).

[0206] 9. J. H. Levine et al., Data-Driven Phenotypic Dissection of AML Reveals Progenitor-like Cells that Correlate with Prognosis. Cell 162, 184-197 (2015).

[0207] 10. L. van der Maaten, G. Hinton, Visualizing Data using t-SNE.Journal of Machine Learning Research 9, 2579-2605 (2008).

[0208] 11. S. Van Gassen et al., FlowSOM: Using self-organizing maps for visualization and interpretation of cytometry data. Cytometry A 87, 636-645 (2015).

[0209] 12. L. M. Weber, M. D. Robinson, Comparison of clustering methods for high-dimensional single-cell flow and mass cytometry data. Cytometry A 89, 1084-1096 (2016).

[0210] 13. N. Mallesh et al, Knowledge transfer to enhance the performance of deep learning models for automated classification of B cell neoplasms. Patterns (N Y) 2, 100351 (2021).

[0211] 14. M. Zhao et al., Hematologist-Level Classification of Mature B-Cell Neoplasm Using Deep Learning on Multiparameter Flow Cytometry Data. Cytometry A 97 , 1073-1080 (2020).-49- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0212] 15. Z. Hu, A. Tang, J. Singh, S. Bhattacharya, A. J. Butte, A robust and interpretable end-to-end deep learning model for cytometry data. Proc Natl Acad Set USA 117, 21373-21380 (2020).

[0213] 16. M. Malek et al., flowDensity: reproducing manual gating of flow cytometry data by automated density -based cell population identification. Bioinformatics 31, 606-607 (2015).

[0214] 17. K. He, X. Zhang, S. Ren, J. Sun (2016) Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770-778.

[0215] 18. J. Tang, J. Z. Liu, M. Zhang, Q. Z. Mei, Visualizing Large-scale and High-dimensional Data. Proceedings of the 25th International Conference on World Wide Web (Www ’16) 10.1145 / 2872427.2883041, 287-297 (2016).

[0216] 19. S. M. Lundberg, S.-I. Lee, A unified approach to interpreting model predictions. Advances in neural information processing systems 30 (2017).

[0217] 20. A. Chan, Q. Gao, M. Roshal, 19-color, 21 -Antigen Single Tube for Efficient Evaluation of B- and T-cell Neoplasms. Curr Protoc 3, e884 (2023).

[0218] 21. D. R. Parks, M. Roederer, W. A. Moore, A new “Logicle” display method avoids deceptive effects of logarithmic scaling for low signals and compensated data. Cytom Part A 69a, 541-551 (2006).

[0219] 22. K. Ranasinghe, M. Naseer, M. Hayat, S. Khan, F. S. Khan (2021) Orthogonal projection loss. Proceedings of the IEEE / CVF international conference on computer vision, pp 12333-12343.D. Deep Learning-Based Automatic Detection of B-Cell Neoplasms in Clinical Flow Cytometry for Diagnosis And Monitoring of Hematological DiseasesIntroduction-50- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0220] Flow cytometry rapidly generates multiparametric, single-cell measurements for millions of cells per specimen. This high-dimensional data is interpreted by reviewing a series of two-dimensional plots and applying hierarchical gating strategies to identify relevant cell populations. This manual workflow requires the expertise of an experienced technician and a pathologist and demands substantial time investment. Computational approaches are therefore pursued to reduce this resource burden by automating and standardizing key parts of the analysis.

[0221] There may be two computation approaches to process flow cytometry data. Referring now to FIG. 17, depicted are high-dimensional clustering methods on unannotated flow cytometry data. Unsupervised learning approaches attempt to identify cell populations directly from unannotated, high-dimensional flow cytometry data using clustering methods. These may include:Methods Identifies cell populations using:FLOCK Density-based clusteringflowSOM self-organizing maps and hierarchicalflowMeans K-means clustering and a change point detection algorithm PhenoGraph Community-detection-algorithm-based clustering of nearest neighbor graphof the single cells based on their phenotypic similarity

[0222] Referring now to FIG. 18, depicted are cell classifiers trained on annotated cells and aggregated for the sample-level prediction. Supervised learning approaches train cell-level classifiers using annotated cells and then aggregate cell-level outputs to generate sample-level predictions. Cell classifiers trained on annotated cells then aggregated for the sample-level prediction may include logistic regression, naive Bayes, decision trees, and random forests; while deep-learning-based approaches include architectures such as CellCNN, DeepCellCNN, and CSNN, among others.Methods and Materials for Deep Learning Approach for Diagnosing B-Cell Neoplasm

[0223] Flow cytometry may be used for diagnosing mature B-cell neoplasms by measuring antigen expression patterns across large numbers of individual cells, enabling detection and quantification of abnormal B-cell populations. A key cue in routine B-cell immunophenotyping is immunoglobulin light-chain restriction (kappa vs. lambda), which -51- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451serves as a practical proxy for clonality. Capturing these features computationally, however, may be challenging because background polytypic B cells, variable signal intensity, and specimen-to-specimen variability can obscure restricted populations. Other automated methods have therefore often fallen short of the sensitivity required for clinical deployment.

[0224] To address these and other technique challenges, presented herein a deeplearning model that approaches clinical sensitivity for automated detection of B cell neoplasms.Data and Cohort information

[0225] Retrospective cohort of 3,070 patient specimen tested for B-cell neoplasms:Specimen type Case Patients Tumor case (test split) Normal cases (test split) Peripheral Blood 1,369 1,357 515 (30%) 854 (30%) Bone Marrow 316 312 66 (30%) 250 (30%)Tissue 1,385 1,225 600 (30%) 785 (30%)

[0226] The cellular features may include one or more of: FSC-A, FSC-H, SSC-A, SSC-H, KAPPA / CD8, TCRCbetal BB700, CD10, CD7 APC-A700, CD45 APC-H7, CD26, CD4, CD3, CD14, CD5, CD38 BV480, CD279, CD20, TCRgd, CD2, LAMBDA / CD56, CD25 PE-Dazzle 594, CD22 PC5-5, or CD 19 PC7, among others. The labels may include Abnormal or Normal B-cell label; Kappa, Lambda, or Neither Light-chain expression label; and Tumor or Normal case label.Process for Processing Flow Cytometry Data

[0227] Referring now to FIG. 19, depicted is a block diagram of an overview of a process for processing flow cytometry data using cell-level and sample-level modules. First, raw flow cytometry (.fcs) data may be preprocessed (e.g., compensation, transformations and cleaning, gating steps) to produce standardized per-cell feature inputs. Next, expert-annotated cells may be used to train a cell-level deep learning module that assigns each cell an abnormality-related score or label (normal vs. abnormal), which is the key intermediate output. The trained cell-level module is then applied back to each case to reorder or sort cells by predicted abnormality, creating a consistent, fixed- structure-52- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451representation that emphasizes the most suspicious events and reduces the burden of variable cell counts across specimens. Finally, expert-annotated cases (tumor vs. normal at the specimen level) may be used to train a sample-level module on these reordered cases, enabling the system to render a case-level diagnosis while remaining grounded in the celllevel evidence that can later be visualized and reviewed by pathologists.

[0228] Referring now to FIG. 20, depicted is a block diagram of an example architecture for classifiers and modules. Residual Network (ResNet) architecture may be built from convolutional blocks connected by skip connections may be used to implement the modules and classifiers. The skip connection adds the input of a block directly to its output after one or more learned transformations, allowing the network to learn residual functions and improving training stability as depth increases. The architecture may include stacked convolution, normalization, and activation layers form residual blocks (including shortcut connections and, in some variants, squeeze-and-excitation (SE) blocks), followed by global average pooling and fully connected layers to produce the final classification output. The ResNet model may be used across modules (e.g., light-chain classification, cell-level abnormality detection, and sample-level case classification) to provide a consistent, high-capacity modeling framework for flow cytometry data.Data processingTransformation and Compensation

[0229] Referring now to FIGs. 21 A-C, depicted are graphs related to transformation and compensation preprocessing flow cytometry data. Transformation: Flow Cytometry data has a large dynamic range. Linear and Log scale might not be helpful for many channels (see FIG. 2 IB). Transformations may be used to aid visualization of the data in each channel (see FIG. 21 A).

[0230] Compensation: The emission spectra of fluorescent dyes are broad, and there can be spectral overlap (see FIG. 21C). This can cause false channel expression and need to be compensated for. Together, these steps standardize the input data so the subsequent deep-learning modules can learn biologically meaningful patterns rather than artifacts of scale or spectral spillover.-53- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451Automated B Cell Gating

[0231] Referring now to FIG. 22, depicted are graphs of automated B-cell gating to preprocess clinical flow cytometry data before deep-learning inference.Process

[0232] Referring now to FIGs. 23 A-F, depicted are block diagrams of a process for processing flow cytometry data using a kappa-lambda (KL) classifier, cell-level classifier, and sample-level classifier modules. Starting with FIG. 23 A, a KL Classifier may be trained using the light-chain labels (Kappa / Lambda / Neither). In FIG. 23B, ‘Local KL similarity’ feature may be generated for all cells. For each cell, get 30 Nearest Neighbors and find number of neighbors with same KL prediction. Moving onto FIG. 23 C, ‘Edge distance’ feature may be generated for all cells. For each cell, the distance to the estimated boundary may be calculated. Moving to FIG. 23D, the Cell-level module may be trained using the abnormality labels (Normal / Abnormal). As seen on FIG. 23E, cell (events) may be reordered using the Cell-level module. Moving onto FIG. 23F, annotated cases may be used to train the Sample-level module.Results

[0233] Referring now to FIG. 24A, depicted are graphs of annotation labels and model predictions by the KL classifier in a two-dimensional feature space. The label and the prediction may classify individual cells as kappa-expressing, lambda-expressing, or neither (double-negative). Referring now to FIG. 24B, depicted is a graph of a receiver operating characteristic (ROC) curve demonstrating KL classifier model’s ability to discriminate. These demonstrate that the KL classifier can reliably capture light-chain restriction patterns that are central to identifying clonal mature B-cell populations and that are subsequently used by downstream cell-level and sample-level classifiers as biologically grounded features for abnormal cell detection and case-level diagnosis.

[0234] Referring now to FIG. 25A, depicted are graphs of annotation labels and model predictions by the KL classifier in a two-dimensional feature space. The label and the prediction may classify individual cells as normal or abnormal. Referring now to FIG.-54- 4930-6229-2378.1Atty. Dkt. No.: 115872-345125B, depicted is a graph of a receiver operating characteristic (ROC) curve demonstrating cell-level model’s ability to discriminate between abnormal and normal cells.

[0235] Referring now to FIG. 26A, depicted is a confusion matrix comparing predicted versus true case labels (tumor versus normal) generated by the sample-level classifier operating on cell-level events. Referring now to FIG. 26B, depicted is a receiver operating characteristic (ROC) curve demonstrating high discriminative performance (AUC). These results demonstrate that the sample-level classifier can convert complex single-cell measurements into clinically relevant specimen-level diagnoses with high discriminative performance. As seen in the graphs, the sample-level classifier demonstrates high diagnostic performance despite differences in background composition among the different specimen types, such as blood, marrow, and tissue.

[0236] Referring now to FIG. 28, depicted is a graph abnormal population estimator based on cell-level abnormality predictions. The graph plots, on a log scale, the number of abnormal (tumor) B cells estimated by the cell-level module versus expert-annotated tumor cell counts across peripheral blood, bone marrow, and tissue specimens, with the identity line indicating perfect agreement and the reported R2demonstrating strong concordance for tumor burden quantification. Referring now to FIGs. 29A and 29B, depicted are graphs of biomarker outputs for clinical verification of model predictions. The graphs in FIG. 29A show representative marker-versus-marker dot plots and a two-dimensional embedding for a normal specimen. The graphs in FIG. 29B show an abnormal specimen. Each events colored by the predicted tumor likelihood score.Comparison of Model PerformanceResults for Cell-level Module Without KL Classifier

[0237] Referring now to FIG. 30, depicted is a graph of a cell-level module performance for abnormal B-cell identification. The graph shows the receiver operating characteristic (ROC) curve on held-out test cells for the cell-level classifier that distinguishes abnormal (tumor) B cells from normal B cells. The ROC curve shows the cell-level classifier achieved an AUROC of 0.952 using only original flow cytometry parameters, without the KL model or its derived engineered features.-55- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451Results for Sample-level Modules Without KL Classifier

[0238] Referring now to FIG. 31 A, depicted is a confusion matrix comparing the predicted case labels (tumor vs. normal) with ground-truth labels for the sample-level classifier module. The confusion matrix shows 413 true negatives, 30 false positives, 29 false negatives, and 289 true positives. Referring now to FIG. 3 IB, depicted is a graph of a receiver operating characteristic (ROC) curve with an area under the curve (AUC) for the sample-level classifier module. The ROC curve shows an AUROC of 0.964. Overall, the system achieved 95.9% accuracy, 92.5% sensitivity, and 98.4% specificity.Model with KL Classifier

[0239] The KL classifier provides a classification of each cell as kappa, lambda, or neither, a local similarity, and a distance to edge feature.Results for Cell-level Module with KL Classifier

[0240] Referring now to FIG. 32, depicted is a graph of receiver operating characteristic (ROC) curve on held-out test cells for the cell-level classifier distinguishing abnormal (tumor) B cells from normal B cells with the incorporation of the KL light-chain classifier output and features, with the reported AUC indicating increased discriminative performance. The ROC curve shows the cell-level classifier achieved an improved AUROC of 0.981, up from 0.952 before the KL model. The KL model provides biological grounding through light-chain classification, enabling engineered features (LC-N, edge distance, abnormal neighborhood) to substantially enhance cell-level discrimination.Results for Sample-Level Module with KL Classifier

[0241] Referring now to FIG. 33 A, depicted is a confusion matrix comparing predicted vs. true case labels (tumor vs. normal) for the sample-level classifier module. The confusion matrix shows 435 true negatives, 8 false positives, 9 false negatives, and 309 true positives. Referring now to FIG. 33B, the ROC curve shows an AUROC of 0.992. The overall model achieved 97.5% accuracy, 95.6% sensitivity, and 98.8% specificity. Compared to the base model (e.g., without KL classifier), the model with the KL classifier-56- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451reduced false positives from 30 to 8 and false negatives from 29 to 9, demonstrating that light-chain classification is essential for achieving clinical-grade diagnostic performance.

[0242] The automated method for identifying abnormal B cells in flow cytometry was incorporated with three new features that significantly boost accuracy. First, a kappa / lambda (K / X) classifier was trained to assign each cell as K- or X-light-chain-expressing, and this label was used as an additional input to the cell-level module. Second, a “local K / X similarity” feature, which measures how many neighboring cells in the multidimensional space share the same light chain expression, was incorporated. Third, an “edge distance” feature was added to quantify how close a given cell is to the boundary of its population cluster.

[0243] With these enhancements, the model’s cell -level AUROC rose from 0.95 to 0.98. At the sample level, the AUROC improved from 0.96 to 0.99, and accuracy increased from 0.959 to 0.975. The improvement in accuracy from 95.9% to 97.5% after incorporating the KL model represents a clinically meaningful gain. At a high-volume institution that processes approximately 10,000 flow cytometry diagnoses per year, this approximately 1.7% increase in accuracy translates to approximately 170 additional patients who receive correct diagnoses. In clinical practice, each of these cases represents a patient who could benefit from earlier and more accurate treatment decisions, underscoring the real-world impact of integrating light-chain classification into automated flow cytometry analysis.E. Systems and Methods for Classifying Samples Based on Light Chain Expression Features in Flow Cytometry Data

[0244] Multiparameter flow cytometry may be used to evaluate and analyze individual cells and particles in samples. The interpretation of cytometry data may be largely manual, with experts meticulously reviewing the readouts from the cytometry data. This process may be time-consuming, impractical, and difficult to scale to high-parameter panels with large counts. Manual review may also be subject to inter-operator and interlaboratory variability. Other attempts to automate the analysis of flow cytometry data may involve the use of unsupervised clustering models or supervised deep-learning models by feeding all the cytometry data at once to generate an output. Without any particular -57- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451consideration of flow cytometry data or the nature of cells, these models may not meet the sensitivity, specificity, or accuracy specifications for clinical deployment. Furthermore, such models may not generalize across specimen types, such as blood, bone marrow, and tissue samples.

[0245] To address these and other technical challenges, a machine learning (ML) architecture including a kappa-lambda (KL) classifier, a cell-level model, and a samplelevel model may be used to process flow cytometry data. The ML architecture may have been trained using labeled flow cytometry dataset in the sample in accordance with supervised learning. The labels may include an annotation of a light chain expression for each cell in a sample, an annotation of whether a given cell is abnormal or normal, and an annotation of whether a tumor is present or absent in the sample. With the training and establishment of the ML architecture, a computing system may receive a flow cytometry dataset. The flow cytometry dataset may include a set of events, each corresponding to a detection of a cell within a sample from a subject. The computing system may perform preprocessing on the flow cytometry dataset, such as select particular types of feature types (e.g., to enrich B-cell data within the remaining dataset).

[0246] The computing system may apply the set of events of the flow cytometry dataset to the KL classifier of the ML architecture. For each event, the computing system applying the KL classifier may generate a light chain expression for the corresponding cell as kappa, lambda, or neither. Using the embedding sets used to generate light-chain expression classifications, the computing system may generate a cluster in a feature space. Based on the clustering in the feature space, the computing system may determine a set of spatial metrics for the light-chain expression classifications. The computing system may execute the cell-level model using the spatial metrics and the set of events. For each event, the computing system executing the cell-level model may determine a cell-level value indicating whether the corresponding cell is normal or abnormal.

[0247] In accordance with the cell-level values, the computing system may sort the set of events in the flow cytometry dataset and may select a subset of events. The computing system may apply the selected subset of events to the sample-level model. From applying the events, the computing system may generate a sample-level value indicating a-58- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451likelihood of a presence of tumor in the sample. Based on the sample-level value, the computing system may classify whether tumor is present or absent in the sample. The output from the ML architecture may be used to facilitate clinical decisions, including identifying the cancer subtype and whether to select the subject as a candidate for anticancer therapy.

[0248] In this manner, by staged processing the flow cytometry data at the cell level and the sample level, the ML architecture may achieve high performance in terms of sensitivity, specificity, and accuracy. Instead of feeding the events of the flow cytometry data all at once, the ML architecture may use the KL classifier and cell-level model to individually process each event and then process a subset of the events through the samplelevel model. This setup may allow for reduced input size at each of the KL classifier, the cell-level model and the sample-level model, thereby allowing for computational tractability and improving resource allocation efficiency. The addition of the spatial metrics using the outputs of the KL classifier may provide for extraction of features that would otherwise be missed in other approaches. From a clinical perspective, the ML architecture may reduce or decrease the amount of manual effort and time that would have been spent by a human expert in examining flow cytometry data. The higher sensitivity, specificity, and accuracy exhibited by the ML architecture may allow the use of the ML architecture to facilitate in clinical decisions.

[0249] Referring now to FIG. 34, depicted is a block diagram of a system 100 for classifying samples based on flow cytometry data. In overview, the system 100 may include at least one data processing system 105, at least one flow cytometer 110, at least one computing device 115, and at least one database 120, among others, communicatively coupled with one another via at least one network 125. The data processing system 105 may include at least one dataset indexer 130, at least one model trainer 135, at least one expression evaluator 140, at least one event evaluator 145, at least one sample evaluator 150, at least one output generator 155, at least one kappa-lambda (KL) classifier model 160, at least one cell-level machine learning (ML) model 165, and at least one sample-level ML model 170, among others. Each of the components in the system 100 (such as the data processing system 105 and its subcomponents and the computing device 115) as detailed herein may be implemented using hardware (e.g., one or more processors coupled with -59- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451memory), or a combination of hardware and software as detailed herein in Section F. The system 100 may be used to implement the functionalities described herein in Sections A-D.

[0250] In further detail, the data processing system 105 may be any computing device including one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 105 may be associated with an entity to process flow cytometry data. The data processing system 105 can be in communication with the flow cytometer 110, the computing device 115, the database 120, and other devices, via the network 125. The data processing system 105 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a branch office, or a site at which one or more servers corresponding to the data processing system 105 is situated.

[0251] The data processing system 105 may include one or more modules, components, or subsystems to perform the various processes and tasks described herein. The dataset indexer 130 may receive datasets acquired for a sample from the flow cytometer 110. The model trainer 135 may initialize, train, and establish the KL classifier model 160, the cell-level ML model 165, and the sample-level ML model 170. The expression evaluator 140 may use the KL classifier model 160 to determine a light-chain (LC) expression classification of each cell as one of kappa (kappa-positive), lambda (lambdapositive), or neither. The event evaluator 145 may use the cell-level ML model 165 to determine probabilities of tumor associated with cancer for each cell. The sample evaluator 150 may use the sample-level ML model 170 to determine a probability of a presence of tumor associated with cancer in a given sample. The output generator 155 may produce output information based on the determinations using the KL classifier model 160, cell-level ML model 165 and the sample-level ML model 170.

[0252] The KL classifier model 160 may include any type of machine learning (ML) or artificial intelligence (Al) architecture to process flow cytometry datasets. The architecture for the KL classifier model 160 may include, for example, a deep learning artificial neural network (ANN) (e.g., a convolutional neural network (CNN), residual neural network (RNN), feedforward neural network (FNN), transformer network, or an autoencoder), a Markov chain, a support vector machine (SVM), a clustering algorithm, a-60- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451Bayesian classifier, or a decision tree, among others. In general, the KL classifier model 160 may include at least one input, at least one output, and a set of weights relating the input with the output. The input may include flow cytometry datasets. The output may include a light-chain (LC) expression classification for each event. The set of weights may be arranged in accordance with the ML or Al architecture.

[0253] The cell-level ML model 165 may include any type of machine learning (ML) or artificial intelligence (Al) architecture to process flow cytometry datasets. The architecture for the cell-level ML model 165 may include, for example, a deep learning artificial neural network (ANN) (e.g., convolutional neural network (CNN), residual neural network (RNN), feedforward neural network (FNN), transformer network, or an autoencoder), a Markov chain, a support vector machine (SVM), a clustering algorithm, a Bayesian classifier, or a decision tree, among others. In general, the cell-level ML model 165 may include at least one input, at least one output, and a set of weights relating the input with the output. The input may include flow cytometry datasets. The output may include a value indicating a probability of tumor associated with cancer for each event. The set of weights may be arranged in accordance with the ML or Al architecture.

[0254] The sample-level ML model 170 may include any type of machine learning (ML) or artificial intelligence (Al) architecture. The sample-level ML model 170 may include any type of machine learning (ML) or artificial intelligence (Al) architecture to process flow cytometry datasets. The architecture for the cell-level ML model 165 may include, for example, a deep learning artificial neural network (ANN) (e.g., a convolutional neural network (CNN), residual neural network (RNN), feedforward neural network (FNN), transformer network, or an autoencoder), a Markov chain, a support vector machine (SVM), a clustering algorithm, a Bayesian classifier, or a decision tree, among others. In general, the sample-level ML model 170 may include at least one input, at least one output, and a set of weights relating the input with the output. The input may include flow cytometry datasets for a given sample with multiple cell types. The output may include a value indicating a probability of tumor associated with cancer for the sample. The set of weights may be arranged in accordance with the ML or Al architecture.-61- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0255] The flow cytometer 110 may be an instrument or device to analyze properties of cells within a given sample. The flow cytometer 110 may include a fluidics component, an optics component, and a processing component. The fluidics component may transport cells from the sample into a single stream (or file) for analysis within the flow cytometer 110. The fluidic component may include a sheath fluid to surround the cells of the sample to focus the cells into a narrow stream. The fluidics component may include a flow chamber in which the cells are funneled into the stream of cells at a controlled rate and pressure. In some embodiments, the fluidics component may include a cell sorter to select B-cells for analysis from the sample. The cell sorter may be, for example, in accordance with fluorescence-activated cell sorting, to physically arrange and separate cells based on characteristics, such as size, granularity, or fluorescence, among others.

[0256] The optics component may include a light source to illuminate the cells in the stream and a set of photodetectors to collect signals as the light scatted about the cells. The light source may emit a laser at any frequencies (e.g., 488 nm for blue, 633 nm for red, and 405 nm for violet, within 10%). The set of detectors may include: a forward scatter (FSC) detector to measure amount of light scattered in a forward direction relative to the flow of cells; a side scatter detector (SSC) to measure amount of light scattered in an oblique direction (e.g., 90 degrees, within 10%) relative to the flow of cells; and a fluorescence detector (FL) to measure light of a particular wavelength emitted by fluorescent agent (e.g., a fluorescent dye) in the cells.

[0257] In the flow cytometer 110, the processing component may convert the light detected or acquired via the detectors of the optics component to datasets to be processed by the data processing system 105. The processing component may receive the electric signals converted by the set of photodetectors from the light acquired by the optics component. The processing component may include an amplifier to amplify the electric signals from the photodetector. The processing component may perform analog-to-digital conversion (ADC), and generate the dataset using the digitized and quantized flow cytometry data. The dataset may be stored and maintained in accordance with flow cytometer standard (FCS) or comma-separated values (CSV) file format. The flow cytometer 110 may be in communication with the data processing system 105, the computing device 115, the database 120, and other devices, via the network 125. The flow cytometer 110 may be -62- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451associated with an entity (e.g., a clinician) examining the subject for cancer or a vendor providing flow cytometry data for such an entity.

[0258] The computing device 115 (sometimes herein referred to as an end user computing device) may be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The computing device 115 may be in communication with the data processing system 105, the flow cytometer 110, and the database 120 via the network 125. The computing device 115 may have at least one display. The computing device 115 may be associated with an entity (e.g., a clinician) examining the subject for cancer. The display may present information about the subject provided by the data processing system 105.

[0259] The database 120 may store and maintain various resources and data associated with the data processing system 105, the flow cytometer 110, and the computing device 115, among others. The database 120 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The database 120 may be in communication with the data processing system 105, the flow cytometer 110, and the computing device 115, via the network 125. While running various operations, the data processing system 105, the flow cytometer 110, and the computing device 115 may access the database 120 to retrieve identified data therefrom. The data processing system 105, the flow cytometer 110, and the computing device 115 may also write data onto the database 120 from running such operations.

[0260] Referring now to FIG. 35, depicted is a block diagram of a process 200 to receive event datasets from flow cytometers in the system for classifying samples. The process 200 may include or correspond to operations performed in the system 100 to acquire flow cytometry data. Under the process 200, the flow cytometer 110 may carry out, execute, or otherwise perform flow cytometry on at least one sample 204 obtained from at least subject 202. The subject 202 may be a human or animal subject, among others. The subject 202 may be at risk of, or diagnosed with, cancer. The subject 202 may be under evaluation for cancer. For instance, the subject 202 may be under evaluation for minimum residual disease (MRD) of cancer, subsequent to administration of anti-cancer therapy. The subject 202 may be also evaluation for presence of neoplasms (e.g., abnormal growths of-63- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451tissue in the body), and may be under evaluation for treatment for malignant or precancerous neoplasms.

[0261] In some embodiments, the cancer may include leukemia. The leukemia may include, for example, acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL), among others. The sample 204 may be obtained from any organ in the subject 202 to evaluate the subject 202 for leukemia. The organ for the sample 204 may include, for example, bone marrow, blood, lymph node, spleen, or spinal fluid, among others. The sample 204 may include any number of white blood cells, such as T-cells or B-cells. For instance, the sample 204 may include B lymphocytes (B-cells) with protein markers corresponding to any one or more of the white blood cells, among others. In some embodiments, the white blood may include any number of biomarkers, such as CD8, TCRyb, TCR pi, CD10, CD7, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD279 (PD-1), CD20, CD2, CD56, CD25, CD22, or CD1, among others.

[0262] In some embodiments, the cancer or tumor is a carcinoma, sarcoma, a melanoma, or a hematopoietic cancer. In some embodiments, the cancer or tumor is selected from among adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin’s disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, non-Hodgkin’s lymphoma, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.

[0263] To perform flow cytometry, the flow cytometer 110 may pass cells from the sample 204 through a flow chamber. In some embodiments, the flow cytometer 110 may carry out cell sorting (e.g., using fluorescence-activated cell sorting (FACS)) on the cells of-64- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451the sample 204. For instance, the flow cytometer 110 may perform cell sorting on the cells to select the B-cells to evaluate for cancer (e.g., leukemia or neoplasm), while excluding non-B-cells (e.g., cells unrelated to leukemia or neoplasm). The flow cytometer 110 may illuminate the cells in the flow cytometer with a light source. The flow cytometer 110 may use the set of photodetectors to acquire light interacting with the cells of the sample 204 into a set of flow channels 206A-N (hereinafter referred generally as flow channels 206). The set of flow channels 206 may correspond to the set of photodetectors in the flow cytometer 110. Each flow channel 206 may correspond to a different type of light signal, such as forward scatter channel height (FSC-H), forward scatter channel area (FSC-A), side scatter channel height (SSC-H), side scatter channel area (SSC-A), and any number of fluorescence channel (e.g., green, yellow, orange, or red fluorescence), among others.

[0264] In performing the flow cytometer, the flow cytometer 110 may create, produce, or otherwise generate at least one dataset 210 based on the light detected from the cells of the sample 204. The dataset 210 may identify or include a set of event sequences 212A-N (hereinafter generally referred to as set of event sequences 212) for the set of flow channels 206. Each event sequence 212 may correspond to a respective flow channel 206. Each event sequence 212 may identify or include a set of events 214A-N (hereinafter generally referred to events 214). The events 214 may correspond to the cells detected in the flow channel 206 associated with the event sequence 212. Each event sequence 212 may include or identify cells in the flow channel 206. The cells may correspond to (e.g., be marked with) at least one of the white blood cell types in the sample 204. In each event sequence 212, the set of event 214 may be arranged in an initial order. The initial order may, for example, correspond to an order of measurement or acquisition by the flow cytometer 110.

[0265] Within the event sequence 212, each event 214 may correspond to an occurrence or a detection of at least one cell within a sample interval. The event 214 may correspond to a detection of at least one biomarker (e.g., protein biomarker) in the cell of the sample 204. The biomarker may include, for example, at least one of CD8, TCRyb, TCR pi, CD10, CD7, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD279 (PD-1), CD20, CD2, CD56, CD25, CD22, or CD1, among others. The event sequence 212 for a given flow channel 206 may be generated or acquired at a sampling rate (e.g., at a rate of 10 to 100 -65- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451MHz). The event sequence 212 may have any number of events over a given period of time, ranging between 500 events (or cells) to 50,000 events (or cells) per second. The dataset 210 may be generated as one or more files, such as using flow cytometer standard (FCS) or comma-separated values (CSV) file format. In some embodiments, the dataset 210 may include or identify metadata, such as an identifier for the subject 202, an identifier of the sample 204, and date and time of acquisition of the flow cytometry data, among others. The flow cytometer 110 may send, transmit, or otherwise provide the dataset 210 to the data processing system 105 or the database 120 for storage.

[0266] The dataset indexer 130 on the data processing system 105 may retrieve, identify, or otherwise receive the dataset 210 from the flow cytometer 110. With receipt from the flow cytometer 110, the dataset indexer 130 may process or parse the dataset 210 to extract or identify the set of event sequences 212. The dataset indexer 130 may identify the set of event sequences 212 by the corresponding set of flow channels 206, such as the forward scatter channel, side scatter channel, and the fluorescence channel, among others. In some embodiments, the dataset indexer 130 may perform gating on the dataset 210 from the flow cytometer 110. In some embodiments, the dataset indexer 130 may determine or identify the initial order of the events in the set of event sequences 212 from the dataset 210. The dataset indexer 130 may store and maintain the dataset 210 the database 120.

[0267] Referring now to FIGs. 36A and 36B, depicted is a block diagram of a process 300 to train machine learning models in the system for classifying samples. The process 300 may correspond to or include, operations performed in the system 100 to initial, train, and establish machine learning models used to evaluate the samples for cancer. The process 300 may be performed or carried out independently of the acquisition of the new flow cytometry data. Starting with FIG. 36 A, under the process 300, the model trainer 135 may initialize, train, and establish the cell-level ML model 165 and the sample-level ML model 170. To initialize, the model trainer 135 may assign or set the set of weights in each of the cell-level ML model 165 and the sample-level ML model 170 to a defined value (e.g., a random value). In addition, the model trainer 135 on the data processing system 105 may retrieve, obtain, or otherwise identify a set of training samples 302A-N (hereinafter generally as training samples 302). The set of training samples 302 may be training data used to train the cell-level ML model 165 and the sample-level ML model 170.-66- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0268] Each training sample 302 may identify or include at least one sample dataset 310. The sample dataset 310 may identify or include a set of event sequences 312A-N (hereinafter generally referred to as set of event sequences 312) for a set of flow channels (e.g., same as or different from the flow channels 206). Each event sequence 312 may correspond to a respective flow channel. Each event sequence 312 may include or identify cells in the flow channel. The cells may correspond to (e.g., be marked with) at least one of the white blood cell types in a respective biological sample. The biological sample may include at least one of a blood sample, a bone marrow sample, or a tissue sample from a subject.

[0269] The subject may be under evaluation for cancer. The subject may be also under evaluation for presence of neoplasms (e.g., abnormal growths of tissue in the body), and may be under evaluation for treatment for malignant or precancerous neoplasms. In some embodiments, the cancer may include leukemia. The leukemia may include, for example, acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL), among others. The sample may be obtained from any organ in the subject to evaluate the subject for leukemia. The sample may include any number of white cell blood types. The white cell may have any number of biomarkers, such as CD8, TCRyb, TCR (31, CD10, CD7, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD279 (PD-1), CD20, CD2, CD56, CD25, CD22, or CD1, among others. In some embodiments, the biomarkers may include CD10+, CD19+, CD20+, CD34+, CD38+, or CD45+, among others.

[0270] In some embodiments, the cancer or tumor may be a carcinoma, sarcoma, a melanoma, or a hematopoietic cancer. In some embodiments, the cancer or tumor is selected from among adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin’s disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, non-Hodgkin’s lymphoma, oral cancers, ovarian cancers, pancreatic -67- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.

[0271] Each event sequence 312 may identify or include a set of events 314A-N (hereinafter generally referred to events 314). The events 314 may correspond to the cells detected in the flow channel associated with the event sequence. In each event sequence 312, the set of event 314 may be arranged in an initial order. The initial order may, for example, correspond to an order of measurement or acquisition by the flow cytometer. Within the event sequence 312, each event 314 may correspond to an occurrence or a detection of at least one cell within a sample interval. The event sequence 312 for a given flow channel may be generated or acquired at a sampling rate (e.g., at a rate of 10 to 100 MHz). The event sequence 312 may have any number of events over a given period of time, ranging between 500 events (or cells) to 50,000 events (or cells) per second. The sample dataset 310 may be generated as one or more files, such as using flow cytometer standard (FCS) or comma-separated values (CSV) file format.

[0272] In some embodiments, the sample dataset 310 may be derived or generated directly from flow cytometry performed on the biological sample from the subject (e.g., in a similar manner as the dataset 210). For example, the flow cytometer 110 may perform flow cytometry acquisition on the biological sample from a given subject to generate the sample dataset 310. In some embodiments, the sample dataset 310 may be synthetically derived or generated. For instance, the sample dataset 310 may be a combination of portions from different flow cytometry datasets. At least one of the flow cytometry datasets may be from a tumorous sample. At least one of the flow cytometry datasets may be from a normal sample. The sample dataset 310 may be generated using a combination of other sample datasets 310.

[0273] In each training sample 302, the sample dataset 310 may be labeled or annotated. The training sample 302 may identify or include a set of light-chain (LC) class labels 304A-N (hereinafter generally referred to as LC class labels 304). The set of celllevel labels 306 may correspond to the set of events 314 across the set of event sequences 312 of the sample dataset 310. Each LC class label 304 may indicate or identify the-68- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451corresponding event 314 as one of a kappa class (kappa-positive), a lambda class (lambda positive), or a neither class (kappa negative and lambda negative), among others. The LC class labels 304 may be inputted or generated by a clinician examining the flow cytometry data of the sample dataset 310 from the subject.

[0274] The training sample 302 may identify or include a set of cell-level labels 306A-N (hereinafter generally referred to as cell-level labels 306). The set of cell-level labels 306 may correspond to the set of events 314 across the set of event sequences 312 of the sample dataset 310. Each cell-level label 306 may define or identify the corresponding event 314 as one of presence or absence of tumor associated with cancer (e.g., leukemia or malignant neoplasm). In some embodiments, each cell-level label 306 may identify the individual cell in the corresponding event 314 as a normal or tumorous cell. Each cell-level label 306 may be associated with the same event 314 (or cell) as a corresponding LC class label 304. The set of cell-level labels 306 may be inputted or generated by a clinician examining the flow cytometry data of the sample dataset 310 from the subject.

[0275] In addition, the training sample 302 may include at least one sample-level label 308 for the sample associated with the sample dataset 310. The sample-level label 308 may define or identify the sample as one of presence or absence of tumor associated with cancer (e.g., leukemia or malignant neoplasm). The sample-level label 308 may define or identify the sample as a normal or tumor sample. In some embodiments, the sample-level label 308 may indicate or identify a cancer subtype for the sample. The cancer subtype may include, for example, at least one of the plurality of cancer types may include at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL), among others. The sample-level label 308 may be inputted or generated by a clinician examining the flow cytometry data of the sample dataset 310 from the subject.

[0276] With the identification, the expression evaluator 140 may input, feed, or otherwise apply each event 314 across the set of event sequences 312 to the KL classifier model 160. The expression evaluator 140 may process each event 314 in accordance with the set of weights of the KL classifier model 160. From applying the set of events 314 -69- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451across the event sequences 312, the expression evaluator 140 may extract or generate a set of embedding sets 320A-N (hereinafter generally referred to as embedding sets 320). Each embedding set 320 may correspond to a reduced dimensional representation of the corresponding event 314 from the event sequences 312.

[0277] The expression evaluator 140 may process the embedding sets 320 using the KL classifier model 160. From applying the embedding sets 320 using the KL classifier model 160, the expression evaluator 140 may calculate, generate, or determine a corresponding set of LC expression classifications 322A-N (hereinafter generally referred to as LC expression classifications 322). Each LC expression classification 322 may indicate or identify the event 314 as one of the kappa class (e.g., kappa positive), the lambda class (e.g., lambda positive), or the neither class, among others. In some embodiments, each LC expression classification 322 may identify or indicate one of a probability of the kappa class, a probability of the lambda class (e.g., lambda positive), or a probability of the neither class, among others. The expression evaluator 140 may iterate through all the events 314 across the set of event sequences 312 to generate the set of LC expression classifications 322.

[0278] For each event 314, the model trainer 135 may compare each LC expression classification 322 with a corresponding LC class label 304 for the event 314. Based on the comparison, the model trainer 135 may calculate, generate, or otherwise determine at least one error metric. The error metric may indicate a degree of deviation of the output LC expression classification 322 from the expected result as indicated in the LC class label 304 for the event 314. The error metric in accordance with any number of loss functions, such as any number of loss functions, such as a norm loss (e.g., LI or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, or a Huber loss, among others.

[0279] Using the error metric, the model trainer 135 may modify, change, or otherwise update at least one of the sets of weights in the KL classifier model 160. The updating may be accordance with a backpropagation algorithm and an objective function. The objective function may define one or more rates at which the weights are to be updated. The objective function may be in accordance with stochastic gradient descent, and may-70- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The updating of the weights of the KL classifier model 160 may be repeated until convergence.

[0280] Using the set of embedding sets 320, the expression evaluator 140 may define, create, or otherwise generate at least one feature space 330. The feature space 330 may have a set of dimensions corresponding to the dimensions of the embedding sets 320. To generate, the expression evaluator 140 may extract or identify the set of embedding sets 320 from applying the sample dataset 310 to the KL classifier model 160. Each embedding set 320 may correspond to a respective event 314 in the event sequences 312. The dimension reduction may decrease the number of dimensions of the embedding sets 320. From the set of embedding sets 320, the expression evaluator 140 may identify or select a subset of embedding sets 320’ A-N (hereinafter generally referred to as embedding sets 320’) corresponding to a defined set of feature types. The defined feature types may include, for example, at least one of FSC-A, FSC-H, SSC-H, CD10, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD20, CD22, or CD19, among others. In some embodiments, the expression evaluator 140 may perform a dimension reduction (e.g., Large Vis, principal component analysis (PCA), t-SNE, or uniform manifold approximation and projection (UMAP)) on the set of embedding sets 320 to generate the set of embedding sets 320’. For each embedding set 320’, the expression evaluator 140 may determine or identify the corresponding LC expression classification 322.

[0281] Using the LC expression classifications 322, the expression evaluator 140 may calculate, determine, or otherwise generate a set of spatial metrics 332A-N (hereinafter generally referred to as spatial metrics 332). In some embodiments, the expression evaluator 140 may generate the set of spatial metrics 332 using the selected embedding sets 320’ and the corresponding LC expression classifications 322. Each spatial metric 332 may correspond to or be associated with a respective event 314 in the event sequences 312 of the sample dataset 310. In generating the spatial metrics, the expression evaluator 140 may assign or map the set of embedding sets 320 (or the selected embedding sets 320’) onto the feature space 330. The feature space 330 may define a coordinate system with a number of dimensions corresponding to the dimensions of the embedding sets 320 (or 320’). Based on the mapping, the expression evaluator 140 may generate the spatial metrics 332.-71- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0282] In some embodiments, for each event 314, the expression evaluator 140 may calculate, generate, or otherwise determine a respective similarity metric (also referred herein as the local similarity metric) as the spatial metric 332. To determine, the expression evaluator 140 may identify a respective embedding set 320 (or 320’) for the event 314. From the set of events 314, the expression evaluator 140 may identify or select a neighboring subset of events 314 based on the embedding set 320 (or 320’) corresponding to the event 314 and a neighboring subset of embedding sets 320 (or 320’) corresponding to the subset of events 314. For example, the expression evaluator 140 may find neighboring embedding sets 320 (or 320’) that are within a predefined distance relative to the embedding set 320 (e.g., for which the similarity metric is being calculated) in the feature space 330. In some embodiments, the expression evaluator 140 may identify the LC expression classifications 322 corresponding to the neighboring subset of embedding sets 320 and the LC expression classification 322 of the respective embedding set 320. Based on the LC expression classifications 322, the expression evaluator 140 may calculate or determine the respective similarity metric to indicate a degree of similarity (or difference) between the LC expression classifications 322 of the event 314 and the neighboring subset of events 314. The expression evaluator 140 may add or include the similarity metric as the spatial metric 332 for the event 314.

[0283] In some embodiments, for each event 314, the expression evaluator 140 may calculate, generate, or otherwise determine a respective edge metric (also referred herein as the edge distance) as the spatial metric 332. To determine, the expression evaluator 140 may determine or create at least one boundary 334 in the feature space 330. The boundary 334 may identify or define an outline enveloping or surrounding the embedding sets 320 (or 320’) in the feature space 330. With the determination of the boundary 334, for each event 314, the expression evaluator 140 may identify a corresponding embedding set 320 (or 320’) in the feature space 330. The expression evaluator 140 may determine the edge metric between the embedding set 320 for the event 314 and the boundary 334. The edge metric may identify a distance within the feature space 330 between the coordinates of the embedding set 320 and the coordinates of a point (e.g., closest to the embedding set 320) of the boundary 334. The expression evaluator 140 may add or include the edge metric (along with the similarity metric) as the spatial metric 332 for the event 314.-72- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0284] Moving onto FIG. 36B, the event evaluator 145 may input, feed, or otherwise apply each event 314 across the set of event sequences 312 to the cell-level ML model 165. In some embodiments, for each event 314, the event evaluator 145 may apply a corresponding spatial metric 332 and corresponding event 314 as input to the cell-level model 160. The event evaluator 145 may process each event 314 and spatial metric 332 in accordance with the set of weights of the cell-level ML model 165. From applying the set of events 314 across the event sequences 312 and the corresponding set of spatial metrics 332, the event evaluator 145 may calculate, generate, or determine a corresponding set of cell-level values 340A-N (hereinafter generally referred to as cell-level values 340340). Each cell-level value 340 may identify or indicate a probability of tumor associated with cancer for the cell corresponding to the event 314. In some embodiments, the event evaluator 145 may determine each cell-level value 340 to identify or indicate a probability of normal (or benign) corresponding to the event 314. The event evaluator 145 may iterate through all the events 314 across the set of event sequences 312 to generate the set of celllevel values 340.

[0285] For each event 314, the model trainer 135 may compare each cell-level value 340 with a corresponding cell-level label 306 for the event 314. Based on the comparison, the model trainer 135 may calculate, generate, or otherwise determine at least one error metric. The error metric may indicate a degree of deviation of the output cell-level value 340 from the expected result as indicated in the cell-level label 306 for the event 314. The error metric in accordance with any number of loss functions, such as any number of loss functions, such as a norm loss (e.g., LI or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, or a Huber loss, among others.

[0286] Using the error metric, the model trainer 135 may modify, change, or otherwise update at least one of the set of weights in the cell-level ML model 165. The updating may be accordance with a backpropagation algorithm and an objective function. The objective function may define one or more rates at which the weights are to be updated. The objective function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The updating of the weights of the cell-level ML model 165 may be repeated until convergence.-73- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0287] In addition, the event evaluator 145 may rearrange or sort the set of events 314 in the initial order in each event sequence 312 of the sample dataset 310 using the set of cell-level values 340. The sorting may arrange the events 314 with higher cell-level values 340 prior to events 314 with lower cell-level values 340. From sorting, the event evaluator 145 may generate at least one sample dataset 310’ to include a set of event sequences 312’ A-N (hereinafter generally referred to as event sequences 312’). Each event sequence 312’ may correspond to a respective flow channel. Each event sequence 312’ may include the set of events 314 in a rearranged order based on the cell-level values 340. In some embodiments, the event evaluator 145 may exclude, remove, or otherwise truncate at least a portion of the events 314 based on the cell-level values 340. By truncating, the event evaluator 145 may select a subset of events 314’ in the set of events 314 in each of the set of event sequences 312 to include in the dataset 310’. For instance, the event evaluator 145 may limit the number of events 314 after sorting to a predefined number of events (e.g., 5,000-10,000 events). The number of remaining events 314 may correspond to an input size for the sample-level ML model 170.

[0288] The sample evaluator 150 may input, feed, or otherwise apply the set of event sequences 312’ of the sample dataset 310’ (e.g., in its entirety) into the sample-level ML model 170. In some embodiments, the sample evaluator 150 may carry out or perform a transformation on the sample data 310’ from a time-series domain to a target domain (e.g., frequency domain). The transformation may include, for example, a Fourier transform, a Wavelet transform, discrete cosine transform (DCT), or a Hilbert transform, among others. The sample evaluator 150 may apply the transformed sample dataset 310’ to the samplelevel ML model 170. The sample evaluator 150 may process the set of event sequences 312’ of the sample dataset 310’ in accordance with the set of weights of the sample-level ML model 170.

[0289] From applying the set of events 314 across the event sequences 312, the sample evaluator 150 may calculate, generate, or determine at least one sample value 350. The sample value 350 may identify or indicate a probability of tumor associated with cancer (e.g., leukemia or neoplasm) for the sample associated with the sample dataset 310’. In some embodiments, the sample evaluator 150 may determine the sample value 350 to identify or indicate a probability of normal (or benign) for the sample. In some-74- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451embodiments, the sample evaluator 150 may generate a set of sample values 350 for the candidate set of cancer subtypes, such as ALL, CLL, AML, CML, HCL, or APL. Each sample value 350 may indicate a probability of a presence (or absence) of a respective cancer subtype.

[0290] The model trainer 135 may compare the sample value 350 with the corresponding sample-level label 308 for the sample dataset 310. Based on the comparison, the model trainer 135 may calculate, generate, or otherwise determine at least one error metric. The error metric may indicate a degree of deviation of the output sample value 350 from the expected result as indicated in the sample-level label 308 for the sample dataset 310. The error metric in accordance with any number of loss functions, such as any number of loss functions, such as a norm loss (e.g., LI or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, and a Huber loss, among others.

[0291] Using the error metric, the model trainer 135 may modify, change, or otherwise update at least one of the set of weights in the sample-level ML model 170. The updating may be accordance with a backpropagation algorithm and an objective function. The objective function may define one or more rates at which the weights are to be updated. The objective function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The updating of the weights of the sample-level ML model 170 may be repeated until convergence. In some embodiments, the model trainer 135 may store and maintain the set of weights for the KL classifier model 160, the cell-level ML model 165 and the sample-level ML model 170 upon completion of training.

[0292] Referring now to FIGs. 37A and 37B, depicted is a block diagram of process 400 to determine probabilities of tumor in samples using the machine learning models in the system for classifying samples. The process 400 may correspond to or include operations performed in the system 100 to process the flow cytometry data using machine learning models to determine probabilities of tumors in samples. Starting from FIG. 37A, under the process 400, the expression evaluator 140 may input, feed, or otherwise apply each event 214 across the set of event sequences 212 to the KL classifier model 160. The expression-75- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451evaluator 140 may process each event 214 in accordance with the set of weights of the KL classifier model 160. From applying the set of events 214 across the event sequences 212, the expression evaluator 140 may extract or generate a set of embedding sets 420 A-N (hereinafter generally referred to as embedding sets 420). Each embedding set 420 may correspond to a reduced dimensional representation of the corresponding event 214 from the event sequences 212.

[0293] The expression evaluator 140 may process the embedding sets 420 using the KL classifier model 160. From applying the embedding sets 420 using the KL classifier model 160, the expression evaluator 140 may calculate, generate, or determine a corresponding set of LC expression classifications 422 A-N (hereinafter generally referred to as LC expression classifications 422). Each LC expression classification 422 may indicate or identify the event 214 as one of the kappa class (e.g., kappa positive), the lambda class (e.g., lambda positive), or the neither class, among others. In some embodiments, each LC expression classification 422 may identify or indicate one of a probability of the kappa class, a probability of the lambda class , or a probability of the neither class, among others. The expression evaluator 140 may iterate through all the events 214 across the set of event sequences 212 to generate the set of LC expression classifications 422.

[0294] Using the set of embedding sets 420, the expression evaluator 140 may define, create, or otherwise generate at least one feature space 430. The feature space 430 may have a set of dimensions corresponding to the dimensions of the embedding sets 420. To generate, the expression evaluator 140 may extract or identify the set of embedding sets 420 from applying the sample dataset 210 to the KL classifier model 160. Each embedding set 420 may correspond to a respective event 214 in the event sequences 212. The dimension reduction may decrease the number of dimensions of the embedding sets 420. From the set of embedding sets 420, the expression evaluator 140 may identify or select a subset of embedding sets 420’ A-N (hereinafter generally referred to as embedding sets 420’) corresponding to a defined set of feature types. The defined feature types may include, for example, at least one of FSC-A, FSC-H, SSC-H, CD10, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD20, CD22, or CD19, among others. In some embodiments, the expression evaluator 140 may perform a dimension reduction (e.g., Large Vis, principal component analysis (PCA), t-SNE, or uniform manifold approximation and projection -76- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451(UMAP)) on the set of embedding sets 420 to generate the set of embedding sets 420’. For each embedding sets 420’, the expression evaluator 140 may determine or identify the corresponding LC expression classification 422.

[0295] Using the LC expression classifications 422, the expression evaluator 140 may calculate, determine, or otherwise generate a set of spatial metrics 432A-N (hereinafter generally referred to as spatial metrics 432). In some embodiments, the expression evaluator 140 may generate the set of spatial metrics 432 using the selected embedding sets 420’ and the corresponding LC expression classifications 422. Each spatial metric 432 may correspond to or be associated with a respective event 214 in the event sequences 212 of the sample dataset 210. In generating the spatial metrics, the expression evaluator 140 may assign or map the set of embedding sets 420 (or the selected embedding sets 420’) onto the feature space 430. The feature space 430 may define a coordinate system with a number of dimensions corresponding to the dimensions of the embedding sets 420 (or 420’). Based on the mapping, the expression evaluator 140 may generate the spatial metrics 432.

[0296] In some embodiments, for each event 214, the expression evaluator 140 may calculate, generate, or otherwise determine a respective similarity metric as the spatial metric 432. To determine, the expression evaluator 140 may identify a respective embedding set 420 (or 420’) for the event 214. From the set of events 214, the expression evaluator 140 may identify or select a neighboring subset of events 214 based on the embedding set 420 (or 420’) corresponding to the event 214 and a neighboring subset of embedding sets 420 (or 420’) corresponding to the subset of events 214. For example, the expression evaluator 140 may find neighboring embedding sets 420 (or 420’) that are within a predefined distance relative to the embedding set 420 (e.g., for which the similarity metric is being calculated) in the feature space 430. In some embodiments, the expression evaluator 140 may identify the LC expression classifications 422 corresponding to the neighboring subset of embedding sets 420 and the LC expression classification 422 of the respective embedding set 420. Based on the LC expression classifications 422, the expression evaluator 140 may calculate or determine the respective similarity metric to indicate a degree of similarity (or difference) between the LC expression classifications 422 of the event 214 and the neighboring subset of events 214. The expression evaluator 140 may add or include the similarity metric as the spatial metric 432 for the event 214.-77- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0297] In some embodiments, for each event 214, the expression evaluator 140 may calculate, generate, or otherwise determine a respective edge metric as the spatial metric 432. To determine, the expression evaluator 140 may determine or create at least one boundary 434 in the feature space 430. The boundary 434 may identify or define an outline enveloping or surrounding the embedding sets 420 (or 420’) in the feature space 430. With the determination of the boundary 434, for each event 214, the expression evaluator 140 may identify a corresponding embedding set 420 (or 420’) in the feature space 430. The expression evaluator 140 may determine the edge metric between the embedding sets 420 for the event 214 and the boundary 434. The edge metric may identify a distance within the feature space 430 between the coordinates of the embedding set 420 and the coordinates of a point (e.g., closest to the embedding set 420) of the boundary 434. The expression evaluator 140 may add or include the edge metric (along with the similarity metric) as the spatial metric 432 for the event 214.

[0298] The event evaluator 145 may input, feed, or otherwise apply each event 214 across the set of event sequences 212 of the dataset 210 to the cell-level ML model 165. In some embodiments, for each event 214, the event evaluator 145 may apply a corresponding spatial metric 432 and corresponding event 214 as input to the cell-level model 160. The event evaluator 145 may process each event 214 in accordance with the set of weights of the cell-level ML model 165. From applying the set of events 214 across the event sequences 212, the event evaluator 145 may calculate, generate, or determine a corresponding set of cell-level value 405A-N (hereinafter generally referred to as cell-level values 405). Each cell-level value 405 may identify or indicate a probability of tumor associated with cancer for the cell corresponding to the event 214. In some embodiments, the event evaluator 145 may determine each cell-level value 405 to identify or indicate a probability of normal (or benign) corresponding to the event 214. The event evaluator 145 may iterate through all the events 214 across the set of event sequences 212 to generate the set of cell-level values 405.

[0299] The sample evaluator 150 executing on the data processing system 105 may rearrange or sort the set of events 214 in the initial order in each event sequence 212 of the dataset 210 using the set of cell-level values 440. The sorting may arrange the events 214 with higher cell-level values 440 prior to events 214 with lower cell-level values 440. From sorting, the model trainer 135 may generate at least one dataset 210’ to include a set of -78- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451event sequences 212’ A-N (hereinafter generally referred to as event sequences 212’). Each event sequence 212’ may correspond to the respective flow channel 206. Each event sequence 212’ may identify or include the set of events 214 in a rearranged order based on the cell-level values 440. In some embodiments, the model trainer 135 may exclude, remove, or otherwise truncate at least a portion of the events 214 based on the cell-level values 440. By truncating, the model trainer 135 may select a subset of events 214 in the set of events 214 in each of the set of event sequences 212 to include in the dataset 210’. For instance, the event evaluator 145 may limit the number of events 214 after sorting to a predefined number of events (e.g., 5,000-10,000 events). The number of remaining events 214 may correspond to an input size for the sample-level ML model 170.

[0300] The sample evaluator 150 may input, feed, or otherwise apply the set of event sequences 212’ of the dataset 210’ (e.g., in its entirety) into the sample-level ML model 170. In some embodiments, the sample evaluator 150 may carry out or perform a transformation on the dataset 210’ from a time-series domain to a target domain (e.g., frequency domain). The transformation may include, for example, a Fourier transform, a Wavelet transform, discrete cosine transform (DCT), or a Hilbert transform, among others. The sample evaluator 150 may apply the transformed dataset 210’ to the sample-level ML model 170. The sample evaluator 150 may process the set of event sequences 212’ of the dataset 210’ in accordance with the set of weights of the sample-level ML model 170.

[0301] Based on applying the set of events 214 across the event sequences 212 to the sample-level ML model 170, the sample evaluator 150 may calculate, generate, or determine at least one sample value 450. The sample value 450 may identify or indicate a probability of tumor associated with cancer for the sample 204 associated with the dataset 210’. In some embodiments, the sample evaluator 150 may determine the sample value 450 to identify or indicate a probability of normal (or benign) for the sample 204. In some embodiments, the sample evaluator 150 may generate a set of sample values 450 for the candidate set of cancer subtypes, such as ALL, CLL, AML, CML, HCL, or APL. Each sample value 450 may indicate a probability of a presence (or absence) of a respective cancer subtype in the sample 204.-79- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0302] Referring now to FIG. 38, depicted is a block diagram of a process 500 to generate outputs in the system 100 for classifying samples. The process 500 may correspond to or include operations performed in the system 100 to use the values from the machine learning models to derive outputs. Under the process 500, the output generator 155 on the data processing system 105 may produce, determine, or otherwise generate at least one classification 505 (sometimes herein referred to as a tumor classification or sample classification) in accordance with the sample value 450. The classification 505 may identify or indicate a presence or absence of tumor associated with cancer in the sample 204 obtained from the subject 202. To generate the classification 505, the output generator 155 may compare the sample value 450 with a threshold. The threshold may delineate, identify, or otherwise define value for the sample value 450 at which to identify the presence or absence of tumor associated with cancer in the tumor. In some embodiments, the threshold may be indicative of a presence (or absence) of minimum residual disease (MRD) associated with cancer in the subject 202. In some embodiments, the threshold may be indicative of a presence (or absence) of a respective cancer subtype.

[0303] If the sample value 450 satisfies (e.g., greater than or equal to) the threshold, the output generator 155 may generate the classification 505 to indicate the presence of cancer in the sample 204. In some embodiments, when the sample value 450 satisfies (e.g., greater than or equal to) the threshold indicative of MRD, the output generator 155 may generate the classification 505 to indicate presence of MRD in the subject 202. In some embodiments, when the sample value 450 satisfies the threshold indicative of the respective cancer subtype, the output generator 155 may generate the classification 505 to indicate a presence of the cancer subtype in the sample 204. In some embodiments, the output generator 155 may generate the classification 505 to identify that the subject 202 is to be administered with an anti-cancer therapy for cancer. When the cancer is leukemia, the anticancer therapy may include anti-leukemia therapy.

[0304] The anti-leukemia therapy may include, for example, a cytosine arabinoside, a FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor, among others. Examples of FLT inhibitors include, but are not limited to; sunitinib, sorafenib, midostaurin, lestaurtinib, quizartinib, gilteritinib, crenolanib, and gilteritini. Examples of IDH inhibitors include, but are not -80- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451limited to; ivosidenib, olutasidenib, vorasidenib, and enasidenib. Examples of BTK inhibitors include, but are not limited to; ibrutinib, acalabrutinib, Zanubrutinib, and pirtobrutinib. Examples of BCL2 inhibitors include, but are not limited to; venetoclax, navitoclax, obatoclax, oblimersen sodium (G3139), Palcitoclax, AT-101, and LP-118. Examples of hedgehog pathway inhibitors include, but are not limited to; glasdegib, vismodegib, sonidegib, LEQ-506 (NVP-LEQ506), itraconazole, saridegib, BMS-833923 / XL 139, Arsenic Tri oxi de (ATO) and Taladegib.

[0305] Examples of anti-cancer therapies include, but are not limited to; alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogen drugs, aromatase inhibitors, ovarian suppression agents, VEGF / VEGFR inhibitors, EGFZEGFR inhibitors, PARP inhibitors, cytostatic alkaloids, cytotoxic antibiotics, antimetabolites, endocrine / hormonal agents, bisphosphonate therapy agents, immuno-modulating / stimulating antibodies and targeted biological therapy agents (e.g., therapeutic peptides described in US 6306832, WO 2012007137, WO 2005000889, WO 2010096603, etc.). In some embodiments, the at least one additional therapeutic agent is a chemotherapeutic agent. Specific chemotherapeutic agents include, but are not limited to; cyclophosphamide, fluorouracil (or 5 -fluorouracil or 5-FU), methotrexate, edatrexate (10-ethyl-10-deaza-aminopterin), thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabepilone, temozolmide, topotecan, vincristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane, letrozole, leuprolide, abarelix, buserlin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, denosumab, zoledronate, trastuzumab, tykerb, anthracy clines (e.g., daunorubicin and doxorubicin), bevacizumab, oxaliplatin, melphalan, etoposide, mechlorethamine, bleomycin, microtubule poisons, annonaceous acetogenins, or combinations thereof. Examples of immuno-modulating / stimulating antibody include, but are not limited to; anti-PD-1 antibody, anti-PD-Ll antibody, anti-PD-L2 antibody, anti-CTLA-4 antibody, anti-TIM3 antibody, anti-4-lBB antibody, anti-CD73 antibody, anti-GITR antibody, and anti-LAG-3 antibody.

[0306] Otherwise, if the sample value 450 does not satisfy (e.g., less than) the threshold, the output generator 155 may generate the classification 505 to indicate the absence of cancer in the sample 204. In some embodiments, when the sample value 450 -81- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451does not satisfy (e.g., less than) the threshold indicative of MRD, the output generator 155 may generate the classification 505 to indicate absence of MRD in the subject 202. In some embodiments, when the sample value 450 does not satisfy (e.g., less than) the threshold indicative of the respective cancer subtype, the output generator 155 may generate the classification 505 to indicate an absence of the cancer subtype in the sample 204. With the generation, the output generator 155 may store and maintain an association between the subject 202 and the classification 505 using one or more data structures. The data structures may include, for example, an array, a matrix, a linked list, a stack, a tree, a hash table, an object, among others.

[0307] Using the classification 505, the output generator 155 may produce, create, or otherwise generate at least one output 510. The output 510 may include information associated with the classification 505. In some embodiments, the output 510 may include information based on the set of LC expression classifications 422, the cell-level values 440, and the sample value 450, among others. When the classification 505 indicates the presence of the cancer, the output 510 may identify or include the indication of the presence of cancer. When the classification 505 indicates the presence of MRD, the output 510 may identify or include the indication of the presence of MRD associated with cancer. In some embodiments, the output 510 may identify the subject 202 as to be administered with an anti-cancer for cancer or MRD associated with cancer in the subject 202. When the classification 505 indicates the absence of the cancer, the output 510 may identify or include the indication of the absence of cancer. When the classification 505 indicates the absence of MRD, the output 510 may identify or include the indication of the absence of MRD associated with cancer.

[0308] The output generator 155 may transmit, send, or otherwise provide the output 510 to the computing device 115. With receipt, the computing device 115 may display, render, or otherwise present the output 510 including the information on the classification for the subject 202. In some embodiments, the information may include or may be based on the set of LC expression classifications 422, the cell-level values 440, and the sample value 450, among others. The information of the output 510 may be presented via a graphical user interface on a display of the computing device 115. The information of the output 510 may be used to make clinical decisions. For example, a clinician operating the computing -82- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451device 115 may use the information to identify whether the sample 204 of the subject 202 has or lacks tumor associated with cancer. When the output 510 identifies the subject 202 as to be administered with an anti-cancer therapy, the clinician may also use the information of the output 510 to select administration of an anti-cancer therapy (e.g., anti-leukemia therapy) for the subject 202. With the selection, the subject 202 may be administered with a therapeutically effective amount of the anti-cancer (e.g., anti-leukemia) therapy based on the output 510 identifying the subject 202 as to be administered. Otherwise, when the output 510 does not identify the subject 202 as to be administered with an anti-cancer therapy, the subject 202 may be withheld from the administration of the anti-cancer therapy.

[0309] In this manner, by leveraging the KL classifier model 160, the cell-level ML model 165 and the sample-level ML model 170, the data processing system 105 may extract and derivate clinically relevant data and information from flow cytometry data. The celllevel ML model 165 may individually process events 214 to provide corresponding LC expression classifications 422. The LC expression classifications 422 may be used to derive or generate spatial metrics 432. The events 214 along with the corresponding spatial metrics 432 may be used by the cell-level ML model 165 to generate cell-level values 440. With the dataset 210’ including the events 214 rearranged based on the cell-level values 440, the sample-level ML model 170 may determine the sample-level value 410 for the sample 204 from which the flow cytometry data is derived. The bifurcation of the ML architecture may provide a more refined and accurate assessment of the individual cells corresponding to the events 214, which in turn can be used to evaluate the overall sample 204.

[0310] The architecture of the KL classifier model 160, the cell-level ML model 165 and the sample-level ML model 170 may provide performance advantages relative to other approaches, in terms of accuracy and precision. The higher accuracy and precision can facilitate the generation of clinically relevant information, such as the classification 505 regarding presence or absence of leukemia or neoplasm and MRD as well as identification of the subject 202 for additional administration of anti-leukemia therapy. From a computer resource perspective, by using the KL classifier model 160, cell-level ML model 165 and the sample-level ML model 170, the data processing system 105 may allow for more efficient use of computing resources (e.g., processing, memory, and network bandwidth) -83- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451that otherwise would have been wasted in providing inaccurate or useless output from flow cytometry data.

[0311] As explained in Section D, the machine learning architecture in flow cytometry may be incorporated with three new features that significantly boost accuracy. First, a kappa / lambda (K / X) classifier was trained to assign each cell as K- or X-light-chain-expressing, and this label was used as an additional input to the cell-level module. Second, a “local K / X similarity” feature, which measures how many neighboring cells in the multidimensional space share the same light chain expression, was incorporated. Third, an “edge distance” feature was added to quantify how close a given cell is to the boundary of its population cluster. With these enhancements, the model’s cell-level AUROC rose from 0.95 to 0.98. At the sample level, the AUROC improved from 0.96 to 0.99, and accuracy increased from 0.92 to 0.978, making the model suitable for clinical settings.

[0312] Referring now to FIG. 39, depicted is a flow diagram of a method 600 of classifying samples based on light chain expression features in flow cytometry data The method 600 may be implemented or performed using any of the components described herein, such as the data processing system 105, or any combination thereof, or the system 700 described in Section B. Under the method 600, a computing system may receive a dataset including a set of event sequences from a flow cytometer (605). The computing system may apply each event sequence in the dataset to a kappa-lambda (KL) classifier model (610). The computing system may determine a KL class for each event (615). The computing system may generate light chain (LC) expression features using the KL classes (620).

[0313] The computing system may apply the LC expression features and the events to a cell-level machine learning (ML) model (625). The computing system may determine a set of cell-level values for the set of event sequences based on applying to the cell-level ML model (630). The computing system may sort the set of event sequences in the dataset based on the cell-level values (635). The computing system may apply the sorted set of event sequences to the sample-level ML model (640). The computing system may determine a sample-level value based on applying to the sample-level ML model (645). The computing system may determine whether the sample-level value satisfies a threshold-84- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451(650). If the sample-level value satisfies (e.g., greater than or equal to) the threshold, the computing system may classify to indicate a presence of tumor (655). Otherwise, if the sample-level value does not satisfy (e.g., less than) the threshold, the computing system may classify to indicate an absence of tumor (660). The computing system may provide an output based on the classification (665).F. Computing and Network Environment

[0314] Various operations described herein can be implemented on computer systems. FIG. 40 shows a simplified block diagram of a representative server system 700, client computing system 714, and network 726 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 700 or similar systems can implement services or servers described herein or portions thereof. Client computing system 714 or similar systems can implement clients described herein. The system 700 described herein can be similar to the server system 700. Server system 700 can have a modular design that incorporates a number of modules 702 (e.g., blades in a blade server embodiment); while two modules 702 are shown, any number can be provided. Each module 702 can include processing unit(s) 704 and local storage 706.

[0315] Processing unit(s) 704 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 704 can include a general-purpose primary processor as well as one or more special-purpose coprocessors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 704 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 704 can execute instructions stored in local storage 706. Any type of processors in any combination can be included in processing unit(s) 704.

[0316] Local storage 706 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 706 can be fixed, removable or upgradable as desired. Local storage 706 can be physically or logically -85- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451divided into various subunits such as a system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 704 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 704. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 702 is powered down. The term “storage medium” as used herein, includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections.

[0317] In some embodiments, local storage 706 can store one or more software programs to be executed by processing unit(s) 704, such as an operating system and / or programs implementing various server functions such as functions of the system 100 or any other system described herein, or any other server(s) associated with system 100 or any other system described herein.

[0318] Software” refers generally to sequences of instructions that, when executed by processing unit(s) 704 cause server system 700 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 704. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 706 (or non-local storage described below), processing unit(s) 704 can retrieve program instructions and data to process in order to execute various operations described above.

[0319] In some server systems 700, multiple modules 702 can be interconnected via a bus or other interconnect 708, forming a local area network that supports communication between modules 702 and other components of server system 700. Interconnect 708 can be implemented using various technologies including server racks, hubs, routers, etc.-86- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0320] A wide area network (WAN) interface 710 can provide data communication capability between the local area network (interconnect 708) and the network 726, such as the Internet. Technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.24 standards).

[0321] In some embodiments, local storage 706 is intended to provide working memory for processing unit(s) 704, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 708. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 712 that can be connected to interconnect 708. Mass storage subsystem 712 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 712. In some embodiments, additional data storage resources may be accessible via WAN interface 710 (potentially with increased latency).

[0322] Server system 700 can operate in response to requests received via WAN interface 710. For example, one of modules 702 can implement a supervisory function and assign discrete tasks to other modules 702 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 710. Such operation can generally be automated. Further, in some embodiments, WAN interface 710 can connect multiple server systems 700 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation.

[0323] Server system 700 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 7 as client computing system 714. Client computing system 714 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on.-87- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451

[0324] For example, client computing system 714 can communicate via WAN interface 710. Client computing system 714 can include computer components such as processing unit(s) 716, storage device 718, network interface 720, user input device 722, and user output device 724. Client computing system 714 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like.

[0325] Processing unit(s) 716 and storage device 718 can be similar to processing unit(s) 704 and local storage 706 described above. Suitable devices can be selected based on the demands to be placed on client computing system 714; for example, client computing system 714 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 714 can be provisioned with program code executable by processing unit(s) 716 to enable various interactions with server system 700.

[0326] Network interface 720 can provide a connection to the network 726, such as a wide area network (e.g., the Internet) to which WAN interface 710 of server system 700 is also connected. In various embodiments, network interface 720 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc ).

[0327] User input device 722 can include any device (or devices) via which a user can provide signals to client computing system 714; client computing system 714 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 722 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on.

[0328] User output device 724 can include any device via which client computing system 714 can provide information to a user. For example, user output device 724 can include a display to present images generated by or delivered to client computing system 714. The display can incorporate various image generation technologies, e.g., a liquid -88- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451crystal display (LCD), light-emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 724 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.

[0329] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer-readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 704 and 716 can provide various functionality for server system 700 and client computing system 714, including any of the functionality described herein as being performed by a server or client, or other functionality.

[0330] It will be appreciated that server system 700 and client computing system 714 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 700 and client computing system 714 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable -89- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software.

[0331] While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to the specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa.

[0332] Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer-readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer-readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium).

[0333] Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims.-90- 4930-6229-2378.1

Claims

Atty. Dkt. No.: 115872-3451WHAT IS CLAIMED IS:

1. A method of classifying samples based on light chain (LC) expression features in flow cytometry data, comprising:receiving, by one or more processors, from a flow cytometer, a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels, each of the plurality of event sequences identifying a respective set of events in a corresponding flow channel of the plurality of flow channels, the set of events corresponding to at least one of a plurality of white blood cells in a sample obtained from a subject at risk of or diagnosed with cancer;applying, by the one or more processors, the first dataset to a first machine learning (ML) model to generate a plurality of LC expression classifications, each of the plurality of LC expression classifications identifying a respective class for a respective event in the respective set of events in each of the plurality of event sequences;generating, by the one or more processors, a plurality of spatial metrics using the plurality of LC expression classifications, each of the plurality of spatial metric for the respective event the respective set of events in each of the plurality of event sequences; applying, by the one or more processors, the first dataset and the plurality of spatial metrics to a second ML model to generate a second dataset comprising a first plurality of values, each first value of the plurality of first values indicating a respective probability of tumor associated with cancer in the respective event in the respective set of events in each of the plurality of event sequences;applying, by the one or more processors, the second dataset to a third ML model to determine a second value indicating a probability of tumor associated with cancer in the sample of the subject;generating, by the one or more processors, a tumor classification indicating a presence or absence of tumor associated with cancer in the sample in accordance with the second value; andstoring, by the one or more processors, using one or more data structures, an association between the subject and the classification.-91- 4930-6229-2378.1Atty. Dkt. No.: 115872-34512. The method of claim 1, wherein generating the tumor classification further comprises generating the tumor classification to indicate the presence of the tumor associated with cancer in the sample, responsive to the second value satisfying a threshold; and providing, by the one or more processors, an output identifying at least one of (i) the tumor classification to indicate the presence of the tumor, (ii) the presence of the cancer in the subject, or (iii) the subject as to be administered with an anti-cancer therapy for cancer.

3. The method of claim 2, wherein the cancer includes leukemia, and the cancer therapy includes an anti-leukemia therapy comprising at least one of a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor.

4. The method of any one of claims 2 or 3, wherein the subject is administered with the anticancer therapy based on the output identifying the subject as to be administered.

5. The method of claim 1, wherein generating the tumor classification further comprises generating the tumor classification to indicate the absence of the tumor associated with cancer in the sample, responsive to the second value not satisfying a threshold; and providing, by the one or more processors, an output identifying at least one of (i) the tumor classification to indicate the absence of the tumor or (ii) the absence of the cancer in the subject.

6. The method of any one of claims 1-5, further comprising:identifying, by the one or more processors, a plurality of embedding sets based on applying the first dataset to the first ML model, each embedding set of the plurality of embedding sets corresponding to the respective event in the respective set of events in each of the plurality of event sequences;selecting, by the one or more processors, from the plurality of embedding sets, a subset of embedding sets corresponding to at least one of a predefined plurality of feature types;-92- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451identifying, by the one or more processors, for each embedding set of the subset of embedding sets, a corresponding LC expression classification from the plurality of LC expression classifications; andwherein generating the plurality of spatial metrics further comprises generating the plurality of spatial metrics using the subset of embedding sets and subset LC expression classifications.

7. The method of claim 6, wherein the predefined plurality of feature types comprises at least one of FSC-A, FSC-H, SSC-H, CD10, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD20, CD22, or CD 19.

8. The method of any one of claims 6 or 7, wherein the generating the plurality of spatial metrics further comprises, for each respective event of the respective set of events in each of the plurality of event sequences:selecting, from the respective set of events in each of the plurality of event sequences, a subset of events based on a corresponding embedding set of the plurality of embedding sets and a subset of embedding sets of the plurality of embedding sets within a feature space;identifying, from the plurality of LC expression classifications, a subset of LC expression classifications corresponding to the subset of events; anddetermining a respective similarity metric based on a respective LC expression classification and the subset of LC expression classifications.

9. The method of any one of claims 6-8, further comprising creating, by the one or more processors, in a feature space, a boundary defining at least a subset of the plurality of embedding sets corresponding to at least one of the plurality of LC expression classifications; andwherein the generating the plurality of spatial metrics further comprises, for each respective event of the respective set of events in each of the plurality of event sequences, determining a respective edge metric between a corresponding embedding set of the plurality of embedding sets and the boundary in the feature space.-93- 4930-6229-2378.1Atty. Dkt. No.: 115872-345110. The method of any one of claims 6-9, wherein identifying the plurality of embedding sets further comprises:identifying a second plurality of embedding sets generated from applying the first dataset to the first ML model; andgenerating the plurality of embedding sets using the second plurality of embedding set in accordance with a dimension reduction.

11. The method of any one of claims 1-10, wherein the first ML model is trained using a plurality of training examples, each of the plurality of training examples comprising a respective dataset including a respective plurality of event sequences, each of the respective plurality of event sequences identifying a corresponding set of events in the corresponding flow channel of the plurality of flow channels, each event of the corresponding set of events labeled as one of a set of LC classifications.

12. The method of any one of claims 1-11, wherein receiving the first dataset further comprises receiving the first dataset, each of the plurality of event sequences identifying the respective set of events in the corresponding flow channel of the plurality of flow channels arranged in a first order, and further comprises:sorting, by the one or more processors, the set of events in the plurality of event sequences arranged in the first order in the first dataset by the plurality of first values to define a second order; andwherein applying the second dataset further comprises applying the second dataset arranged in the second order to the third ML model.

13. The method of any one of claims 1-12, wherein applying the second dataset further comprises applying the second dataset to the third ML model to generating a second plurality of embedding sets, and further comprising:generating, by the one or more processors, a cancer classification indicating at least one of a plurality of cancer types for the subject based on the second plurality of embedding sets.-94- 4930-6229-2378.1Atty. Dkt. No.: 115872-345114. The method of claim 13, wherein the cancer includes leukemia, and wherein the plurality of cancer types comprises at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL).

15. The method of any one of claim 1-14, wherein receiving the first dataset further comprises identifying the first dataset comprising the plurality of event sequences for the corresponding plurality of flow channels, each of the plurality of event sequences corresponding to at least one biomarker of a plurality of biomarkers in the plurality of white blood cells in the sample.

16. The method of claim 15, wherein plurality of biomarkers comprises at least one of CD8, TCRyb, TCR pi, CD10, CD7, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD279 (PD-1), CD20, CD2, CD56, CD25, CD22, or CD1.

17. The method of any one of claims 1-16, wherein each event in the respective set of events of the plurality of event sequences comprises at least one of a forward scatter channel, a side scatter channel, or a fluorescence channel.

18. The method of any one of claims 1-17, wherein each LC expression classification of the plurality of LC expression classifications identifies one of a kappa class, a lambda class, or a neither class.

19. The method of any one of claims 1-18, wherein the sample comprises at least one of a blood sample, a bone marrow sample, or tissue sample.

20. The method of any one of claims 1-19, wherein the plurality of white blood cells comprises B-cells in the sample.

21. A system for classifying samples based on light chain (LC) expression features in flow cytometry data, comprising:one or more processors coupled with memory, configured to:-95- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451receive, from a flow cytometer, a first dataset comprising a plurality of event sequences for a corresponding plurality of flow channels, each of the plurality of event sequences identifying a respective set of events in a corresponding flow channel of the plurality of flow channels, the set of events corresponding to at least one of a plurality of white blood cells in a sample obtained from a subject at risk of or diagnosed with cancer;apply the first dataset to a first machine learning (ML) model to generate a plurality of LC expression classifications, each of the plurality of LC expression classifications identifying a respective class for a respective event in the respective set of events in each of the plurality of event sequences;generate a plurality of spatial metrics using the plurality of LC expression classifications, each of the plurality of spatial metric for the respective event the respective set of events in each of the plurality of event sequences;apply the first dataset and the plurality of spatial metrics to a second ML model to generate a second dataset comprising a first plurality of values, each first value of the plurality of first values indicating a respective probability of tumor associated with cancer in the respective event in the respective set of events in each of the plurality of event sequences;apply the second dataset to a third ML model to determine a second value indicating a probability of tumor associated with cancer in the sample of the subject;generate a tumor classification indicating a presence or absence of tumor associated with cancer in the sample in accordance with the second value; andstore, using one or more data structures, an association between the subject and the classification.

22. The system of claim 21, wherein the one or more processors are further configured to:generate the tumor classification to indicate the presence of the tumor associated with cancer in the sample, responsive to the second value satisfying a threshold; and provide output identifying at least one of (i) the tumor classification to indicate the presence of the tumor, (ii) the presence of the cancer in the subject, or (iii) the subject as to be administered with an anti-cancer therapy for cancer.-96- 4930-6229-2378.1Atty. Dkt. No.: 115872-345123. The system of claim 22, wherein the cancer includes leukemia, and the cancer therapy includes an anti-leukemia therapy comprising at least one of a cytosine arabinoside, an FLT inhibitor, an IDH inhibitor, a BTK inhibitor, gemtuzumab ozogamicin, BCL-2 inhibitor, or hedgehog pathway inhibitor.

24. The system of any one of claims 22 or 23, wherein the subject is administered with the anti-cancer therapy based on the output identifying the subject as to be administered.

25. The system of claim 22, wherein the one or more processors are further configured to:generate the tumor classification to indicate the absence of the tumor associated with cancer in the sample, responsive to the second value not satisfying a threshold; and provide an output identifying at least one of (i) the tumor classification to indicate the absence of the tumor or (ii) the absence of the cancer in the subject.

26. The system of any one of claims 21-24, wherein the one or more processors are further configured to:identify a plurality of embedding sets based on applying the first dataset to the first ML model, each embedding set of the plurality of embedding sets corresponding to the respective event in the respective set of events in each of the plurality of event sequences;select, from the plurality of embedding sets, a subset of embedding sets corresponding to at least one of a predefined plurality of feature types;identify, for the subset of embedding sets, a corresponding subset of LC expression classifications from the plurality of LC expression classifications; andgenerate the plurality of spatial metrics using the subset of embedding sets and subset LC expression classifications.

27. The system of claim 26 wherein the predefined plurality of feature types comprises at least one of FSC-A, FSC-H, SSC-H, CD10, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD20, CD22, or CD 19-97- 4930-6229-2378.1Atty. Dkt. No.: 115872-345128. The system of any one of claims 26 or 27, wherein the one or more processors are further configured to, for each respective event of the respective set of events in each of the plurality of event sequences:select, from the respective set of events in each of the plurality of event sequences, a subset of events based on a corresponding embedding set of the plurality of embedding sets and a subset of embedding sets of the plurality of embedding sets within a feature space; identify, from the plurality of LC expression classifications, a subset of LC expression classifications corresponding to the subset of events; anddetermine a respective similarity metric based on a respective LC expression classification and the subset of LC expression classifications.

29. The system of any one of claims 26-28, wherein the one or more processors are further configured to:create, in a feature space, a boundary defining at least a subset of the plurality of embedding sets corresponding to at least one of the plurality of LC expression classifications; anddetermine, for each respective event of the respective set of events in each of the plurality of event sequences, a respective edge metric between a corresponding embedding set of the plurality of embedding sets and the boundary in the feature space.

30. The system of any one of claims 26-29, wherein the one or more processors are further configured to:identify a second plurality of embedding sets generated from applying the first dataset to the first ML model; andgenerate the plurality of embedding sets using the second plurality of embedding set in accordance with a dimension reduction.

31. The system of any one of claims 21-30, wherein the first ML model is trained using a plurality of training examples, each of the plurality of training examples comprising a respective dataset including a respective plurality of event sequences, each of the respective plurality of event sequences identifying a corresponding set of events in the corresponding-98- 4930-6229-2378.1Atty. Dkt. No.: 115872-3451flow channel of the plurality of flow channels, each event of the corresponding set of events labeled as one of a set of LC classifications.

32. The system of any one of claims 21-31, wherein the one or more processors are further configured to:receive the first dataset, each of the plurality of event sequences identifying the respective set of events in the corresponding flow channel of the plurality of flow channels arranged in a first order, and further comprises:sort the set of events in the plurality of event sequences arranged in the first order in the first dataset by the plurality of first values to define a second order; andapply the second dataset arranged in the second order to the third ML model.

33. The system of any one of claims 21-32, wherein the one or more processors are further configured to:apply the second dataset to the third ML model to generating a second plurality of embedding sets, and further comprising; andgenerate a cancer classification indicating at least one of a plurality of cancer types for the subject based on the second plurality of embedding sets.

34. The system of claim 33, wherein the cancer includes leukemia, and wherein the plurality of cancer types comprises at least one of acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), or acute promyelocytic leukemia (APL).

35. The system of any one of claim 21-34, wherein the one or more processors are further configured to identify the first dataset comprising the plurality of event sequences for the corresponding plurality of flow channels, each of the plurality of event sequences corresponding to at least one biomarker of a plurality of biomarkers in the plurality of white blood cells in the sample.-99- 4930-6229-2378.1Atty. Dkt. No.: 115872-345136. The system of claim 35, wherein plurality of biomarkers comprises at least one of CD8, TCRyb, TCR pi, CD10, CD7, CD45, CD26, CD4, CD3, CD14, CD5, CD38, CD279 (PD-1), CD20, CD2, CD56, CD25, CD22, or CD1.

37. The system of any one of claims 21-36, wherein each event in the respective set of events of the plurality of event sequences comprises at least one of a forward scatter channel, a side scatter channel, or a fluorescence channel.

38. The system of any one of claims 21-37, wherein each LC expression classification of the plurality of LC expression classifications identifies one of a kappa class, a lambda class, or a neither class.

39. The system of any one of claims 21-38, wherein the sample comprises at least one of a blood sample, a bone marrow sample, or tissue sample.

40. The system of any one of claims 21-39, wherein the plurality of white blood cells comprises B-cells in the sample.-100- 4930-6229-2378.1