Systems and methods for comprehensive and standardized immune system phenotyping and automated cell classification

By combining full spectrum flow cytometry and integrated machine learning methods, the problem of labor-consuming and incomparable immunological spectrum analysis in the prior art is solved, and high-throughput, standardized immunological spectrum generation is achieved, supporting clinical diagnosis and treatment decisions.

CN120359403APending Publication Date: 2025-07-22MELIO HEALTHCARE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083148.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-07
Filing Date
2023-07-14
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing immuno spectrum analysis methods are laborious and time-consuming, making it difficult to achieve standardization and comparableity in high-throughput sample processing, resulting in the ingenerous experimental results being ungeneralized and the inability to generate translatable immune spectrum within clinically relevant time.

Method used

Combining full-spectrum flow cytometry and integrated machine learning methods, fluorescently labeled cells contact the immunophenotypic typing set, individual cells are classified and counted using an integrated machine learning model of cascaded hierarchical tree structure to achieve automated data processing and standardized immune spectrum generation.

Benefits of technology

It realizes the generation of standardized and comprehensive immune spectrums within clinically relevant time, supports high-throughput sample processing, and improves the comparability of experimental results and diagnostic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359403A_ABST
    Figure CN120359403A_ABST
Patent Text Reader

Abstract

Methods and systems for generating an immune profile of a subject are described. In some examples, the method includes contacting at least a first aliquot of a sample from the subject with at least a first immunophenotyping kit to fluorescently label cells contained in the sample; processing the fluorescently labeled cells using a full spectrum flow cytometer to generate fluorescence intensity data from or derived from the fluorescently labeled cells of the sample; providing, as input to an integrated machine learning model, at least a subset of the fluorescence intensity data of the fluorescently labeled cells or data derived therefrom, the integrated machine learning model configured to process the data and classify individual cells as belonging to a different immune cell subset of a plurality of different immune cell subpopulations; and outputting the total cell count or cell frequency of each of the plurality of different immune cell subpopulations in the sample as part of an immune profile of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 386,476, filed on December 7, 2022, the content of which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates to methods and systems for generating a standardized and comprehensive immune profile of a mammalian (e.g., human) subject. Background Art

[0004] Immune profiling, i.e., analyzing a subject's immune health at a given time point at the serological or cellular level, can help diagnose immune - related diseases and conditions (e.g., allergies, hyper - reactive immune responses (such as in asthma and Crohn's disease (inflammatory bowel disease)) or autoimmune diseases (such as certain aspects of autoimmune polyendocrine syndrome and diabetes; see, e.g., “Diseases of the Immune System”, National Center for Biotechnology Information (US), in “Genes and Disease” [Internet], 1998 -, Bethesda (MD)). Immune profiling can also be used, for example, to identify an individual's specific response to an infectious disease (e.g., viral, bacterial, fungal, or parasitic infection), to monitor a patient's (e.g., a cancer patient's) response to treatment (see, e.g., Lyons, et al. (2017), “Immune Cell Profiling in Cancer: Molecular Approaches to Cell - Specific Identification”, Precision Oncology 1, 26), and potentially to predict healthcare outcomes.

[0005] Conventional methods for generating an immune profile include, for example, enzyme - linked immunosorbent assay (ELISA), immunoblotting techniques, and flow - cytometry - based techniques, including using a panel of fluorescently labeled antibodies against multiple cell - surface receptors and manual gating of flow - cytometry data. However, these techniques are generally labor - intensive and time - consuming and are not easily scalable to a level capable of handling hundreds or thousands of samples.

[0006] Recently, automation technology has been applied to flow cytometry-based methods. There are several programming methods to find high-dimensional clusters of labeled immune cells and identify them accordingly using manual expert guidance. Off-the-shelf clustering algorithms are commonly used for this purpose, including tSNE, FlowSOM, UMAP, etc. There are also other specially constructed algorithms for defining labeled cell clusters in high-dimensional immune space, such as flowType and Phenograph. In addition, companies such as Cytapex and Dotmatics OMIQ use similar methods to generate automated cell gating. However, despite the established immunoprofiling platforms and methods, the immunophenotyping panels used in various experiments often vary, which requires users to configure the existing automated gating methods on a per-experiment basis. As a result, the direct comparability between experiments is reduced or eliminated. Therefore, there is a need for improved methods and systems that can provide a standardized and comprehensive immunoprofile for a subject (e.g., a patient), which will facilitate biomedical research and enable improved healthcare outcomes. Summary of the Invention

[0007] Disclosed herein are methods and systems for processing a sample (such as a blood sample) and generating a standardized and comprehensive immunoprofile of a subject. The disclosed methods and systems combine in-sample cell analysis based on full-spectrum flow cytometry (FSFC) with an ensemble machine learning-based method to filter and classify individual immune cells into multiple distinct immune cell subsets. The key advantages of the disclosed methods and systems are achieved through a standardized immunoprofiling platform (including an immunophenotyping panel for fluorescently labeling cells) and by implementing automated data processing. This in turn enables the processing of highly complex FSFC data to generate a translatable immunoprofile within a clinically relevant time frame (i.e., hours).

[0008] Disclosed herein are methods for generating an immunoprofile of a subject, the methods comprising: contacting at least a first aliquot of a sample from the subject with at least a first immunophenotyping panel to fluorescently label cells contained in the sample; processing the fluorescently labeled cells using a full-spectrum flow cytometer to generate fluorescence intensity data of a plurality of fluorescently labeled cells from the sample or data derived therefrom; providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an ensemble machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of distinct immune cell subsets; and outputting a total cell count or cell frequency of each of the plurality of distinct immune cell subsets in the sample as part of the immunoprofile of the subject.

[0009] In some embodiments, the integrated machine learning model is organized in a cascaded hierarchical tree structure including a plurality of nodes, and each of the nodes includes a separate machine learning model. In some embodiments, each separate machine learning model includes an input data set and one to eight output data sets corresponding to branches of the cascaded hierarchical tree structure. In some embodiments, each separate machine learning model includes a neural network model. In some embodiments, each separate machine learning model includes a gradient boosting tree model. In some embodiments, the plurality of nodes includes at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 nodes. In some embodiments, the number of separate machine learning models in the integrated machine learning model is equal to the number of different immune cell subsets among the plurality of different immune cell subsets. In some embodiments, the design of the cascaded hierarchical tree structure is at least partially based on an expert analysis of manually gated fluorescence intensity data of one or more control samples or data derived therefrom.

[0010] In some embodiments, individual cells are classified independently of all other cells among the plurality of fluorescently labeled cells. In some embodiments, individual cells are recursively classified together with all other cells among the plurality of fluorescently labeled cells.

[0011] In some embodiments, one or more labeled training data sets are used to train the integrated machine learning model, the one or more labeled training data sets being generated by an expert by manually gating fluorescence intensity data of one or more control samples or data derived therefrom. In some embodiments, the one or more labeled training data sets are used to individually train the separate machine learning models in the integrated machine learning model. In some embodiments, during training, the predictions of the individual models are used to validate the individual models, but are not propagated forward through the integrated machine learning model, thereby eliminating error propagation during training. In some embodiments, a recursive training method is used to jointly train the separate machine learning models in the integrated machine learning model. In some embodiments, the training of the integrated machine learning model is controlled by one or more hyperparameter values that are the same for each node in the cascaded hierarchical tree structure. In some embodiments, the training of the integrated machine learning model is controlled by one or more hyperparameter values that are different for a subset of nodes in the cascaded hierarchical tree structure. In some embodiments, the training of the integrated machine learning model is controlled by one or more hyperparameter values that are determined by performing a random grid search over a range of values of the one or more hyperparameters.

[0012] In some embodiments, the method further includes performing a mathematical transformation on the fluorescence intensity data or data derived therefrom, and then using the transformed fluorescence intensity data as an input for the integrated machine learning model.

[0013] In some embodiments, the fluorescence intensity data or data derived therefrom includes fluorescence intensity data for at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, or 40 fluorescence detection channels. In some embodiments, the fluorescence intensity data or data derived therefrom further includes forward scatter height data, forward scatter area data, side scatter height data, side scatter area data, autofluorescence data, or any combination thereof.

[0014] In some embodiments, the sample includes a blood sample, a buffy coat sample, or a cell suspension.

[0015] In some embodiments, at least one immunophenotyping panel includes a panel of fluorescently labeled antibodies against cell surface proteins associated with antigen-presenting cells (APCs). In some embodiments, the panel of fluorescently labeled antibodies includes fluorescently labeled antibodies against IgM, CD5, CD62L, CD294, CD69, CD38, PD1, CD11C, CD3, CD8, HLA-DR, CD24, CD337, CD123, CD141, CD1C, CD4, TACI, CD319, CD335, PDL1, CD10, CD45, CD16, IgD, CD40, CD19_TCRγδ, CD43, CD14, CD138, CD15, CD56, CD86, CD303, CD27, or any combination thereof. In some embodiments, the panel of fluorescently labeled antibodies further includes a fluorescently labeled antibody against a cell surface marker indicative of live cells, dead cells, or both.

[0016] In some embodiments, the at least one immunophenotyping panel comprises a panel of fluorescently labeled antibodies against cell surface proteins associated with T cells. In some embodiments, the panel of fluorescently labeled antibodies comprises fluorescently labeled antibodies against TIGIT, CD5, CD28, CXCR5, CD39, TIM3, CD38, PD1, TCRVA7_2_TCRVD1, CD95, CD3, CD8, HLADR, CD31, CCR4, CCR6, CCR7, CD57, ICOS, CD4, KLRG1, TCRVA24_JA18, CD122, CD103, CXCR3, TCRVD2, CD45, CCR10, CD16, CD25, CD161, CD19_TCRGD, LAG3, CD14, CD45RO, CD56, CD127, CD45RA, CD27, or any combination thereof.

[0017] In some embodiments, the plurality of different immune cell subsets includes at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 different immune cell subsets. In some embodiments, the plurality of different immune cell subsets includes white blood cells (WBC), eosinophils, eosinophils / CD5+, neutrophils, neutrophils / large, neutrophils / CD5+, neutrophils / small, B cells, B cells / CD5 - CD27 -, monocytes / CD56+, monocytes / CD56 -, NK cells, dendritic cells (DC), T cells, iNKT cells, γδ T cells (total GD), Vd1 cells, Vd2 cells, Vdx cells, mucosal-associated invariant T (MAIT) cells, TEMRA cells, CD4 naive cells, T helper cells, CD4 effector memory cells, Treg cells, or any combination thereof.

[0018] In some embodiments, the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample is output as part of the immune profile in less than 24 hours, 12 hours, 10 hours, 8 hours, 6 hours, or 4 hours.

[0019] In some embodiments, the immune profile is used to diagnose an immune-related disease or disorder in the subject, monitor the progression of an immune-related disease or disorder, or monitor the response to treatment of an immune-related disease or disorder.

[0020] The present disclosure provides computer-implemented methods for generating an immune profile of a subject, the method comprising: receiving fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from the subject using a full-spectrum flow cytometer or data derived therefrom; providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells into one of a plurality of different immune cell subsets; and outputting the total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0021] In some embodiments, the integrated machine learning model is organized in a cascading hierarchical tree structure comprising a plurality of nodes, and wherein each node comprises an individual machine learning model. In some embodiments, each individual machine learning model comprises an input data set and one to eight output data sets corresponding to branches of the cascading hierarchical tree structure. In some embodiments, each individual machine learning model comprises a neural network model. In some embodiments, each individual machine learning model comprises a gradient boosting tree model. In some embodiments, the plurality of nodes comprises at least 1000, 1200, 1400, 1600, 1800, 2000, 2200 or 2400 nodes. In some embodiments, the number of individual machine learning models in the integrated machine learning model is equal to the number of different immune cell subsets among the plurality of different immune cell subsets. In some embodiments, the design of the cascading hierarchical tree structure is at least partially based on expert analysis of manually gated fluorescence intensity data of one or more control samples or data derived therefrom.

[0022] In some embodiments, individual cells are classified independently of all other cells among the plurality of fluorescently labeled cells. In some embodiments, individual cells are recursively classified together with all other cells among the plurality of fluorescently labeled cells.

[0023] In some embodiments, one or more labeled training datasets are used to train an ensemble machine learning model, the one or more labeled training datasets being generated by an expert by manually gating the fluorescence intensity data of one or more control samples or data derived therefrom. In some embodiments, the one or more labeled training datasets are used to train the individual machine learning models in the ensemble machine learning model individually. In some embodiments, during training, the predictions of the individual models are used to validate the individual models, but are not propagated forward through the ensemble machine learning model, thereby eliminating error propagation during training. In some embodiments, a recursive training method is used to co-train the individual machine learning models in the ensemble machine learning model. In some embodiments, the training of the ensemble machine learning model is controlled by one or more hyperparameter values, the one or more hyperparameter values being the same for each node in the cascaded hierarchical tree structure. In some embodiments, the training of the ensemble machine learning model is controlled by one or more hyperparameter values, the one or more hyperparameter values being different for a subset of the nodes in the cascaded hierarchical tree structure. In some embodiments, the training of the ensemble machine learning model is controlled by one or more hyperparameter values, the one or more hyperparameter values being determined by performing a random grid search over a range of values of the one or more hyperparameters.

[0024] In some embodiments, the method further comprises performing a mathematical transformation on the fluorescence intensity data or data derived therefrom, and then using the transformed fluorescence intensity data as an input for the ensemble machine learning model.

[0025] In some embodiments, the fluorescence intensity data or data derived therefrom includes fluorescence intensity data of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, or 40 fluorescence detection channels. In some embodiments, the fluorescence intensity data or data derived therefrom further includes forward scatter height data, forward scatter area data, side scatter height data, side scatter area data, autofluorescence data, or any combination thereof.

[0026] In some embodiments, the plurality of different immune cell subsets includes at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 different immune cell subsets. In some embodiments, the plurality of different immune cell subsets includes white blood cells (WBC), eosinophils, eosinophils / CD5+, neutrophils, neutrophils / large, neutrophils / CD5+, neutrophils / small, B cells, B cells / CD5-CD27-, monocytes / CD56+, monocytes / CD56-, NK cells, dendritic cells (DC), T cells, iNKT cells, γδ T cells (total GD), Vδ1 cells, Vδ2 cells, Vδx cells, mucosal-associated invariant T (MAIT) cells, TEMRA cells, CD4 naive cells, T helper cells, CD4 effector memory cells, Treg cells, or any combination thereof.

[0027] In some embodiments, the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample is output as part of an immune profile within less than 24 hours, 12 hours, 10 hours, 8 hours, 6 hours, or 4 hours.

[0028] In some embodiments, the immune profile is used to diagnose an immune-related disease or disorder in the subject, monitor the progression of an immune-related disease or disorder, or monitor the response to treatment of an immune-related disease or disorder.

[0029] The present disclosure provides a system that includes: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom; provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of different immune cell subsets; and output the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject. In some embodiments, the system further includes a full-spectrum flow cytometer.

[0030] The present disclosure relates to a non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to: receive fluorescence intensity data generated by processing a fluorescence-labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom; provide at least a subset of the fluorescence intensity data of the plurality of fluorescence-labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescence-labeled cells as belonging to a different immune cell subset among a plurality of different immune cell subsets; and output a total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0031] Incorporated by reference

[0032] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety to the extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference in its entirety. If a term in this document conflicts with a term in the incorporated reference, the term in this document controls. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Aspects of the disclosed methods, apparatuses, and systems are set forth in the appended claims. A better understanding of the features and advantages of the disclosed methods, apparatuses, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, in which:

[0034] Figure 1 A non-limiting example of a process flow diagram of a method for processing a blood sample and generating an immune profile according to one embodiment described herein is provided.

[0035] Figure 2 A non-limiting example of a process flow diagram of a method for training an integrated machine learning (ML) model according to one embodiment described herein is provided.

[0036] Figure 3 An exemplary illustration of machine learning model-based prediction and immune profile generation according to one embodiment of the methods and systems described herein is provided.

[0037] Figure 4 An exemplary computing system according to some embodiments of the methods and systems described herein is shown.

[0038] Figure 5 A non-limiting schematic diagram of a manual gating process for processing full-spectrum flow cytometry data is provided.

[0039] Figure 6 Provides a non - limiting example of a simplified gating hierarchy for constructing a neural network for immune cell classification.

[0040] Figure 7 Provides a simplified schematic diagram of a neural network.

[0041] Figure 8 Provides a non - limiting example of test data generated by a trained neural network classifier.

[0042] Figure 9 Provides another non - limiting example of test data generated by a trained neural network classifier.

[0043] Figure 10 Provides a non - limiting example of validation data generated by a trained neural network classifier.

[0044] Figure 11 Provides another non - limiting example of validation data generated by a trained neural network classifier.

[0045] Figure 12 Provides a non - limiting example of the data flow of cascaded predictions from node to node (i.e., from the output of a parent ML model to the input of a child / chilrden ML model) in an ensemble machine learning model.

[0046] Figure 13A Provides a non - limiting example of data on the frequency of T - helper 17 (T17) cells and vitamin D levels in different age groups considering gender as a covariate.

[0047] Figure 13B Provides a non - limiting example of data on the frequency of T - helper 2 (T2) cells and vitamin D levels in different age groups considering gender as a covariate. Detailed Description

[0048] Disclosed herein are methods and systems for processing a sample (such as a blood sample) and generating a standardized and comprehensive immune profile of a subject. The disclosed methods and systems combine in - sample cell analysis based on full - spectrum flow cytometry (FSFC) with an ensemble machine - learning - based method to filter and classify individual immune cells into multiple distinct immune cell subsets. The key advantages of the disclosed methods and systems are achieved through a standardized immune profiling platform (including an immunophenotyping panel for fluorescently labeling cells) and automated data processing. This in turn enables the processing of highly complex FSFC data to generate a translatable immune profile within a clinically relevant time frame (i.e., hours).

[0049] In some examples, for instance, methods for generating an immune profile of a subject are described, the methods comprising: contacting at least a first aliquot of a sample from the subject with at least a first immunophenotyping panel to fluorescently label cells contained in the sample; processing the fluorescently labeled cells using a spectral flow cytometer to generate fluorescence intensity data of a plurality of fluorescently labeled cells from the sample or data derived therefrom; providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to a different immune cell subset among a plurality of different immune cell subsets; and outputting a total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0050] Systems are also described, the systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive fluorescence intensity data of a fluorescently labeled cell sample collected from a subject or data derived therefrom obtained by processing the fluorescently labeled cell sample using a spectral flow cytometer; provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to a different immune cell subset among a plurality of different immune cell subsets; and output a total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0051] A non-transitory computer-readable storage medium storing one or more programs is also described, the one or more programs comprising instructions that, when executed by one or more processors of a system, cause the system to: receive fluorescence intensity data of a fluorescently labeled cell sample collected from a subject or data derived therefrom obtained by processing the fluorescently labeled cell sample using a spectral flow cytometer; provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to a different immune cell subset among a plurality of different immune cell subsets; and output a total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0052] Definition

[0053] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0054] As used in this specification and the appended claims, the singular forms "a / an" and "the" include plural referents unless the context clearly indicates otherwise. Any reference to "or" herein is intended to cover "and / or" and to cover any and all possible combinations of one or more of the associated listed items.

[0055] As used herein, the terms "includes", "including", "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, components and / or units, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, units and / or groups thereof.

[0056] Throughout this application, various parameter values may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of this disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all possible sub-ranges as well as the individual numerical values within that range, whether or not a particular numerical value or sub-range is explicitly stated. For example, a description of a range such as 1 to 6 should be considered to have specifically disclosed sub-ranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as the individual numbers within that range, such as 1, 1.4, 2, 3, 3.6, 4, 5, 5.8 and 6. This applies regardless of the width of the range.

[0057] Numbers may be expressed herein as "about" a particular value. Similarly, ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. The terms "about" and "approximately" should generally mean an acceptable degree of error or variation of a given value or range of values, such as an error or variation within 20 percent (%), 15%, 10% or 5% of the given value or range of values.

[0058] It should be recognized that the use of ordinal terms such as "first" and "second" in the description of the methods and systems disclosed herein does not in and of itself imply any priority, order of importance of one system component relative to another, or chronological order of performing the acts of a method, but is merely used as a label to distinguish, for example, one system component having a particular name from another system component having the same name, with the ordinal terms being used solely to distinguish the two system components.

[0059] In addition, the various embodiments of the methods and systems set forth herein may be described with reference to exemplary block diagrams, process flowcharts, and other illustrations. After reading this document, it will be apparent to those of ordinary skill in the art that the various embodiments set forth herein may be implemented without being limited to the examples shown. For example, the block diagrams and their accompanying descriptions should not be construed as requiring a particular architecture or configuration. Similarly, in an exemplary process flowchart, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in conjunction with the exemplary process. Thus, the methods and systems described and illustrated in more detail below are exemplary in nature and should not be considered limiting.

[0060] As used herein, the terms "full-spectrum flow cytometry" and "full-spectrum flow cytometer" refer, respectively, to a technique and an instrument for performing flow cytometry, wherein the instrument is configured to capture the full emission spectrum of fluorescent molecules using a high-sensitivity photodetector array, thereby enabling the capture of highly multiplexed fluorescence intensity datasets.

[0061] As used herein, the term "immunophenotyping panel" refers to a panel of antibodies (e.g., fluorescently labeled antibodies) that are used to identify cells based on the type of antigen or marker (e.g., cell surface receptor protein) present on the cell surface.

[0062] The subsection headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. This description is presented to enable one of ordinary skill in the art to make and use the invention, and this description is provided in the context of a patent application and its claims.

[0063] Methods for Immune System Phenotyping and Automated Cell Sorting

[0064] As described above, there are certain programming methods for finding highly dimensional clusters of labeled immune cells and specially constructed algorithms for defining highly dimensional clusters in immune space. However, these methods and algorithms have many drawbacks, as described, for example, in several FlowCAP challenge publications (see, e.g., Aghaeepour, et al. (2013), “Critical Assessment of Automated Flow Cytometry Data Analysis Techniques”, Nature Methods 10(3):228-239).

[0065] As further described above, certain automated cell gating methods for flow cytometry applications are known. However, these existing methods rely on the recognition of patterns in the distribution of aggregated flow cytometry events and on clustering-first methods and do not utilize the integration of machine learning models for classifying the identity of individual cells.

[0066] Overall, current immunophenotyping platforms (including high-throughput methods) have several limitations, including the following two problems: (1) the phenotyping platforms are built for specific purposes and are sufficiently different in design to confound data comparisons between different experiments and across different immunophenotyping panels, and (2) the manual labor required to establish the counts and / or frequencies of immune cells in the different immune cell subsets identified by an immunophenotyping panel is prohibitive for scaling the analysis to more than a few thousand samples. For existing platforms and methods, the immunophenotyping panels used in different experiments are not comparable, and the existing automated gating methods must be configured on a per-experiment basis that is only applicable to that experiment. Thus, the results of immunoprofiling experiments are not generalizable. Accordingly, improved immunophenotyping platforms and methods are needed.

[0067] Disclosed herein are immunophenotyping platforms (also referred to as immunoprofiling platforms) and methods that can address one or more of the above drawbacks and needs. The immunophenotyping methods disclosed herein include performing high-throughput, full-spectrum flow cytometry in the context of an immunophenotyping processing pipeline that includes several stages: sample processing, data generation, raw data analysis, and result analysis. The disclosed immunophenotyping platforms and methods, in combination with the novel application of machine learning algorithms as disclosed herein, can provide scalable, high-throughput, and automated methods for addressing the above drawbacks and needs.

[0068] The immunophenotyping platform disclosed herein provides the ability to process a blood sample from an individual (or any other single-cell suspension of immune cells extracted from a tissue sample) and generate an immune profile of the individual using full-spectrum flow cytometry. Although described primarily in the context of immunoprofiling, the disclosed immunophenotyping platform and methods can also be used more generally to generate a cell type profile based on FSFC analysis of any blood sample (or other single-cell suspension), for which a suitable fluorescently labeled antibody panel against an appropriate set of discriminative cell surface antigens can be assembled.

[0069] In some instances, whole blood samples can be processed using FSFC to generate a fluorescence spectrum for each cell (e.g., each immune cell) within the sample. In some instances, a blood sample can be processed to extract immune cells, and then the immune cells can be processed using FSFC to generate a fluorescence spectrum for each immune cell extracted from the sample. These spectra can then be used to generate an immune profile of the individual. This can be a highly standardized process that produces directly comparable results for analyses performed in different laboratories or at different times. In some instances, these results are used to train machine learning models that are used to predict to which category of immune cell type (i.e., which distinct immune cell subset) each detected cell belongs. In summary, the disclosed immunophenotyping platform and machine learning training and prediction framework enable the generation of clinically relevant reference ranges for immune cell subtypes and support diagnostic decision-making in a scalable and less-biased manner by leveraging the standardization of sample processing and the automation of machine learning-based flow cytometry data processing.

[0070] The immunophenotyping platform and methods disclosed herein differ from existing platforms and methods in that other platforms do not provide comprehensive immune system-level cell phenotyping at high sample processing throughput. Existing platforms focus on specific solutions to specific scientific questions. The methods disclosed herein allow for an integrated, automated, and standardized approach to generate high-resolution immune profiles from blood samples at an unprecedented speed and scale. The standardized and high sample throughput nature of the platform (e.g., processing up to 100, 150, 200, or 250 samples per FSFC instrument per 8-hour workday) enables the establishment of biological reference ranges for cell types or subtypes (e.g., immune cell subtypes) in human populations, which will lead to better patient stratification and improved clinical diagnostic applications.

[0071] The platforms and methods disclosed herein utilize full-spectrum flow cytometry (FSFC) combined with machine learning algorithms to analyze cells labeled with a standardized panel of antibodies for immune system states. This represents the first comprehensive and standardized immunophenotyping platform with automated cell classification. First, all biological samples processed by the FSFC instrument are handled using a highly standardized and tightly controlled protocol that is subject to quality control (QC) criteria. Second, training and application of a customized machine learning model to analyze the resulting flow cytometry data enables generation of an immune profile of a blood sample on a timeline measured in hours.

[0072] Previous attempts to classify immune cell identities were based on clustering methods similar to those used in manual gating. The ML-based data analysis pipeline disclosed herein differs from previous methods in that it examines the fluorescence characteristics of each cell and classifies it based on an ensemble of gradient boosting machine learning models. Thus, identification of immune cell clusters and their hierarchical localization are by-products of individual cell identification and classification through application of a trained ensemble model that includes thousands of individual machine learning models and are not based on first identifying clusters and then measuring the correspondence of individual cells to the identified clusters.

[0073] Figure 1A non-limiting example of a process flow diagram of the immunospectrum analysis method 100 as disclosed herein is provided. As shown, a blood sample (or in some instances, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a cell suspension, etc.) is received at a laboratory facility, and the sample is prepared 102 (e.g., including performing one or more of a dilution step, a centrifugation step, a staining step (using one or more fluorescently labeled antibody panels), and / or a washing step) and analyzed on a full-spectrum flow cytometer (FSFC) 104 to create a flow cytometry standard FCS file (including flow cytometry data) for the sample stored in a database 106. The FCS file may include, for example, fluorescence intensity data for one or more fluorescence detection channels (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 50, or more than 50 fluorescence detection channels), and data derived therefrom (e.g., forward scatter height data, forward scatter area data, side scatter height data, side scatter area data, autofluorescence data, or any combination thereof). In some instances, the number of available fluorescence detection channels may be determined by, for example, a combination of detection hardware available as part of the flow cytometry instrument (e.g., including 5, 10, 20, 25, 50, 75, 100, 125, 150, 175, 200, or more than 200 detectors) and the number of spectrally distinct fluorophores (e.g., 5, 10, 20, 25, 30, 35, 40, 45, 50, 60, or more than 60 spectrally distinct fluorophores).

[0074] During the training phase, data for a number of FCS files may be manually gated (e.g., by an immunologist or other expert) to generate a labeled training data set for one or more samples (e.g., control samples), and the one or more labeled training data sets are used to train 108 a machine learning model (e.g., an individual model in an ensemble machine learning model). These training sets are complete instances of the gating hierarchy implemented on the FSFC output of the samples (e.g., blood samples). The output of the gating process is a set of industry-standard FlowJo workspace files that describe the immunophenotyping classification hierarchy associated with their corresponding samples. The gating data encoded in these files is then extracted and fed through a cascaded ensemble machine learning hierarchy that follows the same gating procedure. Within this hierarchy, each node represents an execution pipeline that trains and maintains a single ML model that predicts events for the corresponding biological cell type specific to that location in the hierarchy. These models are trained using specified fluorescence channel data and gating position parameters provided by domain experts.

[0075] In some instances, the disclosed immunophenotyping and automated data analysis platform can use a single machine learning model to classify all cells within a cell population. This can work well for some cell populations, but for populations where the true number of positive classes (cell subtypes) is small (e.g., less than about 100), class imbalance may overwhelm the ability of a single ML model to reliably classify cell detection events. In cases of misclassification (which is particularly evident as class imbalance grows), weighting the fewer class detection events can result in extreme cases that lead to a decline in model performance. In some instances, gradient boosting machine learning methods (e.g., gradient boosting ensemble models) can provide superior performance.

[0076] The trained machine learning model 110 is then used to generate predictions of the cell type or subtype (e.g., immune cell subset) for individual cell detection events and to determine the cell count (or frequency) of each of the various different cell types or subtypes, which are then combined into an immune profile 112.

[0077] As described above, the processing of a sample (e.g., a blood sample) 102 can include one or more steps, including a staining step that includes contacting the cells within at least a first aliquot of the sample with at least a first fluorescently labeled antibody panel (i.e., an immunophenotyping panel or an FSFC panel) of antibodies directed against a set of specific cell surface antigens (e.g., cell surface proteins) that together enable the discrimination of cell types or cell subtypes of interest. Sample processing can also include immunophenotyping panel design. The sample processing platform can include contacting each of one or more sample aliquots (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more sample aliquots) with one or more FSFC panels (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more FSFC panels).

[0078] For example, in some instances, each of two sample aliquots can be stained with a different FSFC panel, one panel focused on the antigen-presenting cell (APC) arm of the immune system (Panel A), which contains antibodies against 36 different cell surface proteins, and the other panel focused on the adaptive arm of the immune system (Panel T), which contains antibodies against 41 different cell surface proteins. In some instances, the panel can also include cell viability staining to distinguish live cells from dead cells. In some instances, the panel can also include autofluorescence measurements as a "marker". Specific examples of cell surface proteins and additional markers that can be included in these panels are listed in Table 1.

[0079] Table 1. Non-limiting examples of cell surface receptor proteins and other markers for discriminating immune cell subsets.

[0080]

[0081] These panels are customized for the immunophenotyping platform disclosed herein and are designed for the most comprehensive observation of the state of the sample donor's immune system. These panels include markers for determining immune cell types, immune system activation, lineage (e.g., primary markers typically used to define a cell population before further subsetting cell types; examples include but are not limited to CD3 to define total T cells and CD56 and CD16 to define natural killer cells), and exhaustion (cells expressing markers associated with "cell exhaustion" (e.g., PD-1, TIGIT) no longer proliferate and lose their function due to chronic stimulation / extended activation of the immune response). The immunophenotyping platform also defines a common hierarchy (also known as a gating tree) that is used to process the fluorescence spectral data of each detected cell and determine which cells and how many cells belong to each measured population (e.g., immune cell subsets) in the hierarchy. For the current configuration of the platform, this includes 200+ gates for the APC panel and 2000+ gates for the T cell panel.

[0082] As part of the panel design, it may be advantageous to define the maximum number of different fluorescent dyes that can be used within a single panel (which is related to the number of different fluorescence detection channels available in the FSFC instrument) to maximize coverage of the available spectrum while also providing clearly distinguishable signals to differentiate cell surface markers from one another. Multiple iterations can be utilized to determine panel design parameters, and biological constraints can be exploited to reuse dye usage (e.g., the gamma-delta (GD) T cell receptor (TCR GD) is likely not expressed on B cells, so the same fluorophore conjugated to anti-CD19 and anti-TCR GD antibodies can be used to identify B cells and GD T cells, respectively).

[0083] In addition, for panel design, it may be advantageous to minimize the volume of sample (e.g., blood) required to process the sample. In some instances, the methods described herein directly use whole blood rather than isolated peripheral blood mononuclear cells. This can allow for determination of granulocyte counts and frequencies.

[0084] Return reference Figure 1, data generation can refer to a largely automated data generation pipeline that includes performing full-spectrum flow cytometry 104 on a prepared blood sample to generate raw immunofluorescence intensities for each detected cell (e.g., immune cell) and recording the output in a database 106. In some instances, e.g., when two immunophenotyping panels are used to stain aliquots of a blood sample (such as the A panel and the T panel described above), the process generates two data sets, e.g., data sets corresponding to each of the two immunophenotyping panels. In some instances, automated data transformation (e.g., mathematical transformation of fluorescence intensity data or data derived therefrom to minimize the effect of aberrant data points) can be triggered by deposition of an FCS file onto the FSFC instrument hard drive. The transformed data can then be used as input to a machine learning model configured to classify individual cells according to cell type or subtype.

[0085] Any of a variety of supervised machine learning models can be used to implement the methods and systems described herein. Examples include, but are not limited to, neural networks (e.g., deep neural networks), decision trees, and random forests.

[0086] In some instances, an ensemble machine learning model including multiple individual machine learning models can be used to implement the disclosed methods and systems. For example, in some instances, an ensemble machine learning model can be used to implement the disclosed methods and systems, the ensemble machine learning model being configured to process fluorescence intensity data or data derived therefrom and classify individual cells among multiple fluorescently labeled cells as belonging to one of multiple different immune cell subsets.

[0087] In some instances, the ensemble machine learning model can be organized in a cascading hierarchical tree structure including, e.g., multiple nodes, where each node includes an individual machine learning model. In some instances, each individual machine learning model includes an input data set and up to eight output data sets (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 output data sets) corresponding to branches of the cascading hierarchical tree structure. In some instances, each individual machine learning model can include a neural network model. In some instances, each individual machine learning model can include a gradient-boosted tree model. In some instances, the multiple nodes (i.e., the number of individual machine learning models in the ensemble) include at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 nodes. In some instances, the number of individual machine learning models in the ensemble machine learning model is equal to the number of different immune cell subsets among the multiple different immune cell subsets.

[0088] In some instances, the design of the cascading hierarchical tree structure can be at least partially based on expert analysis of manual gating fluorescence intensity data of one or more samples (e.g., control samples as described elsewhere herein) or data derived therefrom. In some instances, individual cells can be classified by a machine learning model (e.g., an ensemble machine learning model) independent of all other cells among the plurality of fluorescently labeled cells contained within the sample. In some instances, individual cells can be recursively classified together with all other cells among the plurality of fluorescently labeled cells within the sample.

[0089] As described above, during the training phase 108, a labeled training dataset is used to train a machine learning model (e.g., an individual model within an ensemble machine learning model), the labeled training dataset being generated using one or more FCS files of one or more samples (e.g., control samples) manually gated, for example, by an immunologist or other expert. In some instances, the one or more samples can include whole blood samples. In some instances, the one or more control samples can include, for example, a cell suspension that contains purified or partially purified cells of a single or only a few cell types or subtypes. Training and prediction can be updated periodically as the model architecture advances, more appropriate hyperparameters are identified, or new data becomes available.

[0090] In some instances, one or more labeled training datasets can be used to train individual machine learning models within an ensemble machine learning model separately. For example, during training, the predictions of an individual model can be used to validate the individual model, but may not be propagated forward through the ensemble machine learning model in order to minimize or eliminate error propagation during the training process.

[0091] In some instances, individual machine learning models within an ensemble machine learning model can be co-trained, for example, using a recursive training method.

[0092] In some instances, the training of the ensemble machine learning model (or individual models contained therein) can be controlled by one or more hyperparameter values that are the same for each node (individual model) in the cascading hierarchical tree structure. In some instances, the training of the ensemble machine learning model can be controlled by one or more hyperparameter values that are different for each node (individual model) in the cascading hierarchical tree structure, or different for a subset of the nodes. In some instances, the training of the ensemble machine learning model can be controlled by one or more hyperparameter values that are determined by performing a random grid search over a range of values of one or more hyperparameters.

[0093] Return to reference Figure 1, Data analysis can refer to the automated application of a trained machine learning model (e.g., an ensemble machine learning model including 2200+ individual machine learning models) 110 to the raw data files generated by an FSFC instrument and stored in a database 106. The cascading hierarchical structure of the ensemble machine learning model processes all or part of the fluorescence intensity readings (or data derived therefrom) from each cell and predicts the identity of the cell (i.e., the cell type or subtype to which it belongs (also referred to herein as a cell subset)), as determined by the gating hierarchy defined by an immunophenotyping panel design.

[0094] The input to each ML model can use the same set of cell surface markers that an expert would use to analyze the data at that level of the hierarchy. For example, to determine an event identified as a neutrophil, an expert would view a two-dimensional plot where the side scatter area is on one axis and the fluorescence intensity of the CD16 marker is on the second axis. The machine learning model at this node in the hierarchy can use the same two inputs. These inputs represent the encoded values (measured as intensity) representative of the size and quantity of the marker. These input channels are optionally chosen to mimic as closely as possible those used by the expert, without including noise from other channels that may drown out the signal provided. Biological expertise guides the selection of each channel for each subset of the immunological population used in the ML ensemble. The input to each ML model can also be a complete set of fluorescence channels available in the panel.

[0095] The ensemble ML models work together to produce an accurate prediction of cell types for the entire gating hierarchy. During prediction, the ML models are programmatically arranged such that the top-level model forwards its prediction to the next-level model. For example, the white blood cell (WBC) model uses two approximate size parameters (side scatter area vs. side scatter height) to determine whether a cell detection event is likely due to the cell being a WBC. Then, one or two ML models can be used to further predict all events classified as WBC events and classify them into side scatter high (SSChi) and side scatter low (SSClo) event types using the forward scatter (FSC) and side scatter (SSC) channels. Then, the cell detection events predicted to be SSChi are further characterized to determine whether they are positive for the expression of the CD15 cell surface marker. Each step of this process is performed by a specific ML model trained for this explicit purpose. This process is followed for each node in the gating hierarchy.

[0096] The machine learning platform 110 uses one or more machine learning models to independently generate cell type or subtype predictions at each leaf (i.e., the final node without child nodes) in a defined gating hierarchy. The one or more machine learning models can have the same or different architectures and work together in an ensemble. Predictions from these models can be cascaded recursively through the hierarchy from more general to more specific (e.g., from parent to child), or each cell can be classified independently of all other cells and all other gating hierarchy decisions.

[0097] The trained machine learning models can be configured to classify individual cells as belonging to one of multiple different cell subsets (e.g., immune cell subsets). In some instances, the multiple different cell subsets (e.g., different immune cell subsets) can include at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 different immune cell subsets.

[0098] Non-limiting examples of immune cell subsets (or subtypes) that can be identified using, for example, the Panel A and Panel T sets of fluorescently labeled antibodies described above are shown in Table 2.

[0099] Table 2. Non-limiting examples of immune cell subsets identified using Panel A and Panel T antibodies.

[0100]

[0101] Figure 2 A non-limiting example of a flowchart of the machine learning training process 200 is provided. Domain experts can generate and store manually labeled training data 206 from the selected FCS files. Machine learning professionals can define the machine learning architecture, training algorithms, and hyperparameters 202 to generate an ML script. The ML script can then be executed to train a machine learning model 204 on the selected manually labeled data 206 and subsequently store it in a database 208. A held-out dataset 212, also manually labeled by experts, can be used to validate the output model. The results of the validation can be manually reviewed by data science experts and / or biologists 210. Models that do not pass the validation can be adjusted and retrained.

[0102] Training data for training machine learning models can be manually generated using a subset of the exemplary immunophenotyping output of professional gating. These manually gated cells can be used to train machine learning models to be able to identify similar cells based on the fluorescence spectra characterized by a spectral flow cytometer for a given cell type or subtype.

[0103] In some instances, using training data generated by multiple experts can remove bias from the self-gating process, so replacing the automated gating process with manual gating may produce different (and potentially incorrect) results.

[0104] As described above, the holdout set of the manual gating file generated by domain experts can be compared, and prediction metrics relative to the manual gating holdout set can be used to validate machine learning model predictions. Examples of prediction metrics include, but are not limited to, comparisons where the mean of the cell population distribution from the prediction set is within 5%, 10%, 20%, 30% of the mean of the holdout set; comparisons where the standard deviation of the cell population distribution of the prediction set is within 5%, 10%, 20% of the standard deviation of the holdout set; and / or the correlation between the cell population ratio of the prediction set and the cell population ratio of the holdout set is greater than 85%, 90%, 95% or 98%. A 100% correlation means that the cell population ratio predicted by the ML model is exactly the same as the cell population ratio determined by manual gating.

[0105] In some instances, the threshold for determining the validity of ML model predictions can be adjusted by thresholding the standard deviation difference using, for example, a standard deviation difference threshold of less than 5%, 10%, 15% or 20%.

[0106] In some instances, the threshold for determining the validity of ML model predictions can be adjusted by thresholding the correlation using a correlation threshold greater than 80%, 85%, 90%, 95% or 98%.

[0107] The validated model can be accepted as a good predictor and stored in database 208. The remaining models can be cycled by data science experts and / or biologists 210 through one or more additional rounds of training, with more training data, algorithm changes or hyperparameter tuning in each round until they are validated.

[0108] In some instances, the machine learning model architecture can include, for example:

[0109] · A gradient descent boosted tree machine learning algorithm with up to one thousand trees;

[0110] · A deep neural network, typically having between 2 and 4 hidden layers with up to 250 nodes per layer;

[0111] · A convolutional neural network, which has; and / or

[0112] · An autoencoder, having, for example, up to 7 layers (e.g., 5 convolutional, 2 pooling).

[0113] Model hyperparameters can be adjusted according to the specific application for which the model is trained.

[0114] The trained and validated models can be used to predict immune cell subset counts and frequencies on all suitable samples. The prediction results for each cell subset in a given sample can be aggregated to produce a systematic view of the immune status of the individual from whom the sample was collected, which can form all or part of an immune profile.

[0115] Figure 3 An exemplary illustration of a process 300 for prediction based on a machine learning model and immune profile generation is provided. The validated machine learning model (stored in database 302) is applied to the raw FCS file 306 to produce immune cell population count predictions 304. These predictions are collected into an immune profile (stored in database 308) and made available for further analysis 310, including, for example, querying an aggregated immune profile stored in database 312 (e.g., aggregated according to the demographics or clinical data of an individual or patient, including but not limited to age, gender, race, family history, disease diagnosis, etc.).

[0116] Identifying the most suitable machine learning algorithms and metadata parameters can help in generating high-quality predictions about immune data. The system can be configured to use the most suitable model architecture, and the model architecture can continue to iterate. When new training data becomes available, the training of the model can also be updated periodically or continuously.

[0117] In some instances, an ensemble machine learning model including more than 2000 individual machine learning models can be used to perform automated analysis on samples. Training and managing such a large number of machine learning models can be accomplished through a programming infrastructure. The platforms and methods disclosed herein can use a software platform to manage the application of trained machine learning models to raw flow cytometry data to produce cell count predictions.

[0118] The advantages of the platforms and methods disclosed herein can be provided by standardizing the immunophenotyping platform and automating machine learning-based data processing. This enables the disclosed platforms and methods to transform highly complex FSFC data into a translatable immune profile for use within a clinically relevant time frame (e.g., hours rather than days). In some instances, an immune profile of a sample can be generated in less than 24 hours, 12 hours, 10 hours, 8 hours, 6 hours, or 4 hours.

[0119] The output of the prediction process can include an immune profile. The profile can include immune cell counts and / or subset frequencies for all measured cell subsets. These immune profiles represent a snapshot of the donor's immune system at the time of sample collection, and several longitudinal samples can be compiled to show immune trajectories. Additionally, several different donor immune signatures can be compiled to discover statistically significant population-level immune signatures. The display of these different use cases may depend on specific project requirements.

[0120] Compared to using different panels, different algorithms, and / or different automated gating methods, the methods described herein can have advantages. The advantages can include generating more accurate and efficient cell clustering, cell type or subtype prediction, and immune profiles. The accuracy of the disclosed machine learning models can be evaluated, for example, based on precision (e.g., the percentage of cells classified as belonging to a given cell type that actually belong to that type), recall (e.g., the percentage of cells in the dataset that are correctly classified as belonging to a given cell type), and the F1 score (i.e., an accuracy metric that combines the precision and recall of the model to assess the number of times the model makes correct predictions across the entire dataset).

[0121] Applications

[0122] The disclosed methods and systems for generating immune profiles can be used in a variety of biomedical research and clinical diagnostic applications. Examples include, but are not limited to, the diagnosis of immune-related diseases and disorders, the diagnosis of autoimmune diseases and cancer, the diagnosis and / or identification of individual-specific responses to infectious diseases, the prediction of responses to treatments before or during treatment, and the characterization of donors and products in cell therapy manufacturing.

[0123] Systems for immune system phenotyping and automated cell classification

[0124] The present disclosure also discloses a system designed to implement any one of the disclosed methods for generating an immune profile of a sample from a subject. The system can include, for example, one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive fluorescence intensity data obtained by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom; provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to a different immune cell subset among a plurality of different immune cell subsets; and output the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject. In some instances, the system can further include a full-spectrum flow cytometer (FSFC) instrument.

[0125] Similarly, a non-transitory computer-readable storage medium is disclosed that can include instructions for an operating system, the system being configured to perform any one of the disclosed methods for generating an immune profile of a sample from a subject. For example, a non-transitory computer-readable storage medium storing one or more programs is described, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to: receive fluorescence intensity data obtained by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom; provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to a different immune cell subset among a plurality of different immune cell subsets; and output the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0126] Computer processors and computing systems

[0127] Figure 4 An exemplary computing system according to some embodiments is shown. Computing system 400 can be a component of a system for generating an immune profile of a sample from a subject.

[0128] The computing system 400 may include a host computer connected to a network. The computing system 400 may be a client computer or a server. As Figure 4 shown, the computing system 400 may include any suitable type of microprocessor-based device, such as a personal computer; a workstation; a server; or a handheld computing device, such as a phone or a tablet. The computer may include, for example, one or more of a processor 410, an input device 420, an output device 430, a memory storage device 440, and a communication device 460.

[0129] The input device 420 may be any suitable device that provides input, such as a touchscreen or monitor, a keyboard, a mouse, or a voice recognition device. The output device 430 may be any suitable device that provides output, such as a touchscreen, a monitor, a printer, a disk drive, or a speaker.

[0130] The memory storage device 440 may be any suitable device that provides storage, such as an electrical, magnetic, or optical memory, including RAM, cache, a hard disk drive, a CD-ROM drive, a tape drive, or a removable storage disk. The communication device 460 may include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or card. The components of the computer may be connected in any suitable manner, such as via a physical bus or wirelessly. The memory storage device 440 may be a non-transitory computer-readable storage medium including one or more programs that, when executed by one or more processors (such as processor 410), cause the one or more processors to perform any of the methods described herein.

[0131] The software 450 that may be stored in the memory storage device 440 and executed by the processor 410 may include, for example, programming embodying the functions of the present disclosure (e.g., as embodied in the methods, systems, computers, servers, and / or devices described above). In some embodiments, the software 450 may be implemented and executed on a combination of servers such as an application server and a database server.

[0132] The software 450 may also be stored and / or transmitted within any computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (such as those described above) that can obtain instructions associated with the software and execute the instructions. In the context of the present disclosure, a computer-readable storage medium may be any medium, such as the storage device 440, that can contain or store a program for use by or in connection with an instruction execution system, device, or apparatus.

[0133] The software 450 can also be propagated in any transmission medium for use by or in conjunction with an instruction execution system, apparatus, or device (such as those described above), which can obtain instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, a transmission medium can be any medium capable of conveying, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. A transmission readable medium can include, but is not limited to, electrical, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0134] The computing system 400 can be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links that can implement the transmission and reception of network signals, such as a wireless network connection, Tl or T3 lines, a cable network, DSL, or a telephone line.

[0135] The computing system 400 can implement any operating system suitable for operating on a network. The software 450 can be written in any suitable programming language such as C, C++, Java, or Python. In various embodiments, the application software embodying the functions of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or as a web-based application or web service through a web browser, for example.

[0136] Examples

[0137] The following examples are included for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0138] Example 1 – Manual gating of full-spectrum flow cytometry data.

[0139] One fundamental principle of flow cytometry data analysis is "gating", which is based on the sequential identification and refinement of cell populations of interest using a panel of cell type-specific molecules (also known as markers) labeled with fluorescently labeled antibodies detected, for example, by fluorescence (Verschoor, et al. (2015), "An Introduction to Automated Flow Cytometry Gating Tools and Their Implementation", Frontiers in Immunology, Volume 6, Article 380). Figure 5A non-limiting schematic illustration of a manual gating process for handling full-spectrum flow cytometry data is provided. Fluorescence intensity data or data derived therefrom (e.g., forward scatter data, side scatter data, live / dead cell staining, autofluorescence, etc.) are generated by an FSFC instrument and reviewed by expert practitioners. Cell types or subtypes are identified by selecting portions or sub-portions of the plotted fluorescence data (e.g., fluorescence intensity data plotted as a heat map in the respective plots of Figure 5 and further refining using additional criteria, where each additional criterion applied to the analysis constitutes a "gate". The number of cell detection events defined within each gate is equal to the number of cells identified for that cell population or subset. As Figure 5 shown, exemplary gating criteria include SSC-A (side scatter – area), FSC-A (forward scatter – area), SSC-H (side scatter – height), FSC-H (forward scatter – height), CD45 (fluorescent signal generated by a fluorescently labeled anti-CD45 monoclonal antibody), etc. The numbers associated with the designated cell types in some plots represent the percentage of cells from the previous gate that meet the current gating criteria.

[0140] Example 2 – Training of a neural network model for immune cell type classification.

[0141] This example provides an illustration of training a machine learning model (e.g., a neural network) that accurately predicts immune cell types using predefined gates from experts. The initial assumptions on which the model is developed include: (i) the FCS file includes data on hundreds of thousands of cell detection events, (ii) each cell detection event represents a training opportunity, and (iii) the type, subtype, and status of a cell can be clearly identified by its inclusion in a set of labeled gates.

[0142] Figure 6 A non-limiting example of a simplified gating hierarchy for constructing a neural network for immune cell classification is provided. Each node of the gating hierarchy represents a different immune cell type, subtype, or cell status. Starting from the root, high side scatter (SSChi) and low side scatter (SSClow) signals can be used, for example, to distinguish white blood cells (eosinophils and neutrophils) from other immune cells. Then, fluorescent signals associated with the labeling of appropriate cell surface markers can be used to distinguish eosinophils from neutrophils. As Figure 6 indicated in the gating hierarchy shown, the detection of fluorescent signals associated with the labeling of other cell surface markers (e.g., CD19, CD3T, GDT, CD4, CD8, etc.) can be used to classify immune cells into many different immune cell subtypes.

[0143] Table 3 provides non-limiting examples of a panel of fluorescently labeled monoclonal antibodies that can be used to detect cell surface markers (adapted from Mahnke, et al. (2012), "OMIP-013: Differentiation of Human T-Cells", Cytometry Part A 81A:935-936).

[0144] Table 3. Non-limiting examples of fluorescently labeled monoclonal antibodies.

[0145] Cell surface marker Monoclonal antibody clone Fluorescent dye Purpose CD3 SK7 APC-H7 Lineage CD4 M-T477 QD605 CD8 RPA-T8 QD585 CCR7 150503 Ax680 Memory / differentiation CD27 O323 FITC CD28 CD28.2 PE-Cy5 CD31 WM59 PE-Cy7 CD45RA HI100 APC CD57 NK-1 QD705 CD95 DX2 PE CD127 A019D5 BV421 CD244 C1.7 PE-Cy5.5 Dead cell __ AqBlu Dump

[0146] APC, allophycocyanin; H7, Highlight 750; QD, quantum dot; Ax, Alexa; FITC, fluorescein isothiocyanate; PE, R-phycoerythrin; Cy, cyanine; BV, bright violet; AqBlu, LIVE / DEAD Fixable Aqua dead cell stain.

[0147] Figure 7 A simplified schematic of a neural network is provided. The neural network includes an input layer, at least one hidden layer (including four or more nodes), and an output layer (including a single node in this non-limiting example), and the input layer includes two or more nodes (or "perceptrons"). Generally, a neural network can include any total number of layers and any number of hidden layers, where the hidden layers serve as trainable feature extractors that allow mapping a set of input data to a preferred output value or a set of output values. Each layer of the neural network includes a plurality of nodes (or perceptrons). The nodes receive inputs directly from the input data (e.g., flow cytometry data or data derived therefrom) or from the outputs of nodes in a previous layer and perform a specific operation, such as a summation operation. In some cases, the connections from the input to the nodes are associated with weights (or weighting factors). In some cases, a node can, for example, sum the products of all inputs from the previous layer and their associated weights. In some cases, the weighted sum is offset by a bias b. In some cases, a threshold or activation function f can be used to gate the output of the node, and the threshold or activation function can be a linear or non-linear function. The activation function can be, for example, a rectified linear unit (ReLU) activation function or other functions, such as saturated hyperbolic tangent, identity, binary step, logistic, arctangent, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, standard sine, sine, Gaussian, or sigmoid function or any combination thereof.

[0148] One or more sets of training data can be used to "teach" or "learn" the weighting factors, bias values, and thresholds or other computational parameters of a neural network during the training phase. For example, input data from a training data set and gradient descent or backpropagation methods can be used to train the parameters such that the output values predicted by the neural network (e.g., immune cell classification) are consistent with the examples included in the training data set. A backpropagation neural network training process, for example, can be used to obtain the tunable parameters of the model, and this backpropagation neural network training process can be performed using the same hardware as or different hardware from that used to perform immune cell classification.

[0149] Use labeled training data divided into training (usually 80% of the total) and test (usually 20% of the total) data sets to train a neural network such as Figure 7 the neural network schematically shown in. For each cell detection event in the training data set, the data is matched to the appropriate input nodes, the neural network model performs an initial classification, the output is compared to the actual (known) output of the detection event, the internal model weights are adjusted, and the cycle is repeated. The trained model is then tested by processing the held-out test data and comparing the model predictions for cell type to the actual (known) cell types. Metrics such as precision and recall are used to evaluate the performance of the model.

[0150] Figure 8 Non-limiting examples of test data generated by a trained neural network classifier are provided. The support numbers indicated in the rightmost column are the number of cells manually gated as belonging to the designated cell type. As described elsewhere herein, the performance of the model is evaluated by calculating precision, recall, and F1 scores for 16 immune cell types. Micro-average: The metric is calculated globally by counting the total number of true positives, false negatives, and false positives. Macro-average: The metric is calculated for each label (cell type) to determine its unweighted average (which does not account for label imbalance). Weighted average: The metric is calculated for each label to determine their average weighted by the support number (the number of true instances for each label). This weighted average modifies the "macro" calculation to account for label imbalance and can result in an F-score that is not between precision and recall. Sample average: The metric is calculated for each instance to determine their average (this calculation is only meaningful for multi-label classification, where this is different from accuracy_score). It can be seen that in this study, the weighted averages of precision and recall are 0.94 and 0.92, respectively.

[0151] Figure 9 Non-limiting examples of test data generated by a trained neural network classifier after training the model on a much larger training data set are provided. It can be seen that in this study, the weighted averages of precision and recall are 0.98 and 0.98, respectively.

[0152] Figure 10 Non-limiting examples of validation data (i.e., model prediction data for input data not previously provided to the model) for a trained neural network classifier are provided. In this validation run, the weighted average precision and recall were 0.99 and 0.98, respectively.

[0153] Figure 11 Another non-limiting example of validation data generated by a trained neural network classifier is provided. In this validation run, the weighted average precision and recall were 0.98 and 0.91, respectively.

[0154] Example 3 – Application of an ensemble machine learning model to automated cell counting.

[0155] An FCS file consists of metadata about the fluorophores and associated markers used when performing a flow cytometry experiment. It also includes data for each fluorescence channel corresponding to each event detected by the flow cytometer. Conventional methods for extracting the count and frequency of an immune cell population from a sample from an FCS file (such as determining the percentage of neutrophils identified among all white blood cells detected) involve using specialized software (e.g., FlowJo, BD Biosciences, Ashland OR). This manual method requires opening the file in the software, visually inspecting the events in multiple bivariate plots, and identifying cell clusters that conform to well-established biological phenotypes. It also requires leveraging the association between fluorescent dyes and immune cell surface proteins, as defined by an immunophenotyping panel design. Subsequently, in accordance with the background immunology expertise of observing the distribution of cell surface markers, closed polygons are manually created to demarcate immune cell clusters in two-dimensional space. When all immune cell subsets of interest (more than 2000 in the case of the IMU (IMU Biosciences, London, UK) immunophenotyping panel) have been identified using this method, the characteristics of the cell population ratios can be extracted.

[0156] To improve the speed and accuracy of this cell counting and classification process, automated cell classification can be applied. Given a standard gating hierarchy, a set of FCS files, and a set of reference configurations for each gate stored in a manual gating workspace file, an ensemble machine learning model can be trained to determine cell population characteristics. The ensemble machine learning model comprises a set of individual ML models, where each ML model in the set is trained to predict whether an event falls within or outside of a configured gate. Once the set of models has been trained, the characteristics can be constructed without human intervention as follows.

[0157] For each event in the FCS file, a machine learning ensemble is called. Each node in the ensemble can consist of, for example, a machine learning model that is configured to receive inputs from each fluorescence channel, where the machine learning model has an internal state learned during a training phase and outputs a set of events that are predicted to be enclosed by a defined gate represented in the ML model. These ML-filtered events are then cascaded into one or more predicted output classes corresponding to the number of sub-gates at the next level of the gating hierarchy. A set of machine learning models in the ensemble is organized in a directed acyclic graph structure where the root node is the first filter in the gating hierarchy. To begin the process of predicting immune signatures, a complete set of events in the FCS file is passed as input into the root machine learning model. Events that conform to the gating definitions contained in the ML model are retained and provided as input to the next model (or models) in the hierarchy. Each node in the hierarchy represents an immune cell type or state, and when the process is complete, all counts and frequencies of all nodes can be extracted to comprise the complete immune signature of the data file.

[0158] Figure 12 A non-limiting example of the data flow of the cascaded predictions from node to node (i.e., from the output of the parent ML model to the input of the child ML model) is provided. In this example, fluorescence intensities are input into the root node (white blood cell (WBC)) model and individual detection events are classified into two output classes (WBC subtypes). The fluorescence intensity data of the latter are then input into the next set of models in the gating hierarchy, such as the eosinophil ML model and the neutrophil ML model, or the B cell ML model and the T cell ML model. As indicated in the illustration, the fluorescence intensity data of a subset of the T cell-gated events can then be input into, for example, the iNKT cell ML model or the γδ T cell ML model.

[0159] Example 4 – Applying immune signatures to highlight biologically meaningful insights.

[0160] To demonstrate the effectiveness of the disclosed immunophenotyping method in identifying and correcting changes in the general healthy lifestyle and health factors of donor samples, an analysis was performed to determine the extent to which vitamin D fluctuations affect immune system parameters. A dataset of 609 donors (58% female, age 38.7 ± 12.5 years) was integrated, and fresh blood was collected and analyzed using the FCSC method defined herein. Additionally, comprehensive general health blood tests (including determination of lipid profiles, vitamin levels, and approximately 50 other biomarkers) were performed on the same group of donors at the same experimental time points. An automated cell sorting method generated immune profiles for all donors. These profiles were then evaluated using a linear regression model to detect significant changes in immune cell populations based on vitamin D levels, with gender and age groups (<30 years: young, >=30 years <60 years: middle-aged, >=60 years, elderly) as covariates. After applying multiple testing correction, a total of 60 immune populations in both genders and across age groups were identified as being affected by changes in vitamin D levels.

[0161] Restoring the significant interactions and trajectories of vitamin D with Th2 cells involved in allergic responses (Georas, et al. (2005), “T-helper cell type-2 regulation in allergic disease”, Eur Respir J. 26(6):1119-37) and Th17 cells involved in autoimmune diseases and infection response pathways (Zambrano-Zaragoza, et al. (2014), “Th17 cells in autoimmune and infectious diseases”, Int J Inflam. 2014:651503) provides a reference range for the baseline phenotypes of these two key populations in healthy cohorts, which can be used as a monitoring and diagnostic metric for patients undergoing treatment for any of these conditions.

[0162] Figures 13A to 13B Non-limiting example graphs are provided that show, in 609 healthy individuals across age groups, the trajectories of T helper 2 ( Figure 13B ) and T helper 17 ( Figure 13A ) cell populations decreasing as vitamin D levels increase, considering gender as a covariate.

[0163] Exemplary embodiments

[0164] Exemplary methods, systems, and computer-readable storage media are listed in the following:

[0165] 1. A method for generating an immune profile of a subject, the method comprising:

[0166] Contact at least a first aliquot of a sample from the subject with at least a first immunophenotyping panel to fluorescently label cells contained in the sample;

[0167] Process the fluorescently labeled cells using a full-spectrum flow cytometer to generate fluorescence intensity data of a plurality of fluorescently labeled cells from the sample or data derived therefrom;

[0168] Provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model, the integrated machine learning model being configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of different immune cell subsets; and

[0169] Output the total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0170] 2. The method according to embodiment 1, wherein the integrated machine learning model is organized in a cascaded hierarchical tree structure including a plurality of nodes, and wherein each node includes an individual machine learning model.

[0171] 3. The method according to embodiment 2, wherein each individual machine learning model includes an input data set and one to eight output data sets corresponding to branches of the cascaded hierarchical tree structure.

[0172] 4. The method according to embodiment 2 or embodiment 3, wherein each individual machine learning model includes a neural network model.

[0173] 5. The method according to embodiment 2 or embodiment 3, wherein each individual machine learning model includes a gradient boosting tree model.

[0174] 6. The method according to any one of embodiments 2 to 5, wherein the plurality of nodes includes at least 1000, 1200, 1400, 1600, 1800, 2000, 2200 or 2400 nodes.

[0175] 7. The method according to any one of embodiments 2 to 6, wherein the number of individual machine learning models in the integrated machine learning model is equal to the number of different immune cell subsets among the plurality of different immune cell subsets.

[0176] 8. The method according to any one of embodiments 2 to 7, wherein the design of the cascading hierarchical tree structure is at least partially based on an expert analysis of the manually gated fluorescence intensity data of one or more control samples or data derived therefrom.

[0177] 9. The method according to any one of embodiments 1 to 8, wherein individual cells are sorted independently of all other cells in the plurality of fluorescently labeled cells.

[0178] 10. The method according to any one of embodiments 1 to 8, wherein individual cells are recursively sorted together with all other cells in the plurality of fluorescently labeled cells.

[0179] 11. The method according to any one of embodiments 1 to 10, wherein one or more labeled training datasets are used to train the ensemble machine learning model, the one or more labeled training datasets being generated by an expert by manually gating the fluorescence intensity data of one or more control samples or data derived therefrom.

[0180] 12. The method according to embodiment 11, wherein the one or more labeled training datasets are used to individually train the individual machine learning models in the ensemble machine learning model.

[0181] 13. The method according to embodiment 11 or embodiment 12, wherein during training, the predictions of the individual models are used to validate the individual models but are not propagated forward through the ensemble machine learning model, thereby eliminating error propagation during training.

[0182] 14. The method according to embodiment 11, wherein a recursive training method is used to co-train the individual machine learning models in the ensemble machine learning model.

[0183] 15. The method according to any one of embodiments 11 to 14, wherein the training of the ensemble machine learning model is controlled by one or more hyperparameter values, the one or more hyperparameter values being the same for each node in the cascading hierarchical tree structure.

[0184] 16. The method according to any one of embodiments 11 to 14, wherein the training of the ensemble machine learning model is controlled by one or more hyperparameter values, the one or more hyperparameter values being different for a subset of nodes in the cascading hierarchical tree structure.

[0185] 17. The method according to any one of embodiments 11 to 16, wherein the training of the integrated machine learning model is controlled by one or more hyperparameter values, and the one or more hyperparameter values are determined by performing a random grid search on the value ranges of the one or more hyperparameters.

[0186] 18. The method according to any one of embodiments 1 to 17, the method further comprising performing a mathematical transformation on the fluorescence intensity data or data derived therefrom, and then using the transformed fluorescence intensity data as an input for the integrated machine learning model.

[0187] 19. The method according to any one of embodiments 1 to 18, wherein the fluorescence intensity data or data derived therefrom comprises fluorescence intensity data of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35 or 40 fluorescence detection channels.

[0188] 20. The method according to embodiment 19, wherein the fluorescence intensity data or data derived therefrom further comprises forward scatter height data, forward scatter area data, side scatter height data, side scatter area data, autofluorescence data or any combination thereof.

[0189] 21. The method according to any one of embodiments 1 to 20, wherein the sample comprises a blood sample, a buffy coat sample or a cell suspension.

[0190] 22. The method according to any one of embodiments 1 to 21, wherein at least one immunophenotyping panel comprises a set of fluorescently labeled antibodies against cell surface proteins associated with antigen-presenting cells (APCs).

[0191] 23. The method according to embodiment 22, wherein the set of fluorescently labeled antibodies comprises fluorescently labeled antibodies against IGM, CD5, CD62L, CD294, CD69, CD38, PD1, CD11C, CD3, CD8, HLADR, CD24, CD337, CD123, CD141, CD1C, CD4, TACI, CD319, CD335, PDL1, CD10, CD45, CD16, IGD, CD40, CD19_TCRGD, CD43, CD14, CD138, CD15, CD56, CD86, CD303, CD27 or any combination thereof.

[0192] 24. The method according to embodiment 23, wherein the set of fluorescently labeled antibodies further comprises a fluorescently labeled antibody against a cell surface marker indicative of live cells, dead cells or both.

[0193] 25. The method according to any one of embodiments 1 to 24, wherein the at least one immunophenotyping panel comprises a panel of fluorescently labeled antibodies against cell surface proteins associated with T cells.

[0194] 26. The method according to embodiment 25, wherein the panel of fluorescently labeled antibodies comprises fluorescently labeled antibodies against TIGIT, CD5, CD28, CXCR5, CD39, TIM3, CD38, PD1, TCRVA7_2_TCRVD1, CD95, CD3, CD8, HLADR, CD31, CCR4, CCR6, CCR7, CD57, ICOS, CD4, KLRG1, TCRVA24_JA18, CD122, CD103, CXCR3, TCRVD2, CD45, CCR10, CD16, CD25, CD161, CD19_TCRGD, LAG3, CD14, CD45RO, CD56, CD127, CD45RA, CD27, or any combination thereof.

[0195] 27. The method according to any one of embodiments 1 to 26, wherein the plurality of different immune cell subsets comprises at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 different immune cell subsets.

[0196] 28. The method according to any one of embodiments 1 to 27, wherein the plurality of different immune cell subsets comprises white blood cells (WBC), eosinophils, eosinophils / CD5+, neutrophils, neutrophils / large, neutrophils / CD5+, neutrophils / small, B cells, B cells / CD5 - CD27 -, monocytes / CD56+, monocytes / CD56 -, NK cells, dendritic cells (DC), T cells, iNKT cells, γδ T cells (total GD), Vd1 cells, Vd2 cells, Vdx cells, mucosa-associated invariant T (MAIT) cells, TEMRA cells, CD4 naive cells, T helper cells, CD4 effector memory cells, Treg cells, or any combination thereof.

[0197] 29. The method according to any one of embodiments 1 to 28, wherein the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample is output as part of the immune profile within less than 24 hours, 12 hours, 10 hours, 8 hours, 6 hours, or 4 hours.

[0198] 30. The method according to any one of embodiments 1 to 29, wherein the immune profile is used to diagnose an immune-related disease or disorder in the subject, monitor the progression of an immune-related disease or disorder, or monitor the response to treatment of an immune-related disease or disorder.

[0199] 31. A computer-implemented method for generating an immune profile of a subject, the method comprising:

[0200] Receiving fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from the subject using a full-spectrum flow cytometer or data derived therefrom;

[0201] Providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells into one of a plurality of different immune cell subsets; and

[0202] Outputting the total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0203] 32. A system, the system comprising:

[0204] One or more processors; and

[0205] A memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:

[0206] Receive fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom;

[0207] Provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells into one of a plurality of different immune cell subsets; and

[0208] Output the total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0209] 33. The system as described in embodiment 32, wherein the system further comprises a full-spectrum flow cytometer.

[0210] 34. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a system, cause the system to:

[0211] Receive fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom;

[0212] Provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells into one of a plurality of different immune cell subsets; and

[0213] Output the total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

[0214] It should be understood from the foregoing that while specific embodiments of the disclosed methods and systems have been illustrated and described, various modifications thereto are possible and are contemplated herein. The present invention is also not intended to be limited by the specific examples provided in the specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the preferred embodiments herein are not meant to be construed in a limiting sense. Further, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein, which may depend on various conditions and variables. Various modifications in form and detail of the embodiments of the present invention will be apparent to those skilled in the art. Accordingly, it is contemplated that the present invention should also cover any such modifications, variations, and equivalents.

Claims

1. A method for generating an immune profile of a subject, the method comprising: contacting at least a first aliquot of a sample from the subject with at least a first immunophenotyping panel to fluorescently label cells contained in the sample; processing the fluorescently labeled cells using a full-spectrum flow cytometer to generate fluorescence intensity data of a plurality of fluorescently labeled cells from the sample or data derived therefrom; providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of different immune cell subsets; and outputting a total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

2. The method of claim 1, wherein the integrated machine learning model is organized in a cascading hierarchical tree structure comprising a plurality of nodes, and wherein each node comprises a separate machine learning model.

3. The method of claim 2, wherein each separate machine learning model comprises an input data set and one to eight output data sets corresponding to branches of the cascading hierarchical tree structure.

4. The method of claim 2 or claim 3, wherein each separate machine learning model comprises a neural network model.

5. The method of claim 2 or claim 3, wherein each separate machine learning model comprises a gradient boosting tree model.

6. The method of any one of claims 2 to 5, wherein the plurality of nodes comprises at least 1000, 1200, 1400, 1600, 1800, 2000, 2200 or 2400 nodes.

7. The method of any one of claims 2 to 6, wherein the number of separate machine learning models in the integrated machine learning model is equal to the number of different immune cell subsets among the plurality of different immune cell subsets.

8. The method of any one of claims 2 to 7, wherein the design of the cascading hierarchical tree structure is at least partially based on an expert analysis of manually gated fluorescence intensity data of one or more control samples or data derived therefrom.

9. The method of any one of claims 1 to 8, wherein individual cells are classified independently of all other cells among the plurality of fluorescently labeled cells.

10. The method of any one of claims 1 to 8, wherein individual cells are recursively classified together with all other cells among the plurality of fluorescently labeled cells.

11. The method of any one of claims 1 to 10, wherein the integrated machine learning model is trained using one or more labeled training data sets generated by an expert by manually gating fluorescence intensity data of one or more control samples or data derived therefrom.

12. The method according to claim 11, wherein the one or more labeled training datasets are used to individually train the individual machine learning models in the ensemble machine learning model.

13. The method according to claim 11 or claim 12, wherein during training, the predictions of the individual models are used to validate the individual models, but are not propagated forward through the ensemble machine learning model, thereby eliminating error propagation during training.

14. The method according to claim 11, wherein a recursive training method is used to jointly train the individual machine learning models in the ensemble machine learning model.

15. The method according to any one of claims 11 to 14, wherein the training of the ensemble machine learning model is controlled by one or more hyperparameter values, and the one or more hyperparameter values are the same for each node in the cascaded hierarchical tree structure.

16. The method according to any one of claims 11 to 14, wherein the training of the ensemble machine learning model is controlled by one or more hyperparameter values, and the one or more hyperparameter values are different for a subset of the nodes in the cascaded hierarchical tree structure.

17. The method according to any one of claims 11 to 16, wherein the training of the ensemble machine learning model is controlled by one or more hyperparameter values, and the one or more hyperparameter values are determined by performing a random grid search over a range of values of the one or more hyperparameters.

18. The method according to any one of claims 1 to 17, the method further comprising performing a mathematical transformation on the fluorescence intensity data or data derived therefrom, and then using the transformed fluorescence intensity data as input for the ensemble machine learning model.

19. The method according to any one of claims 1 to 18, wherein the fluorescence intensity data or data derived therefrom comprises fluorescence intensity data of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35 or 40 fluorescence detection channels.

20. The method according to claim 19, wherein the fluorescence intensity data or data derived therefrom further comprises forward scatter height data, forward scatter area data, side scatter height data, side scatter area data, autofluorescence data or any combination thereof.

21. The method according to any one of claims 1 to 20, wherein the sample comprises a blood sample, a buffy coat sample or a cell suspension.

22. The method according to any one of claims 1 to 21, wherein at least one immunophenotyping panel comprises a set of fluorescently labeled antibodies against cell surface proteins associated with antigen-presenting cells (APCs).

23. The method according to claim 22, wherein the fluorescently labeled antibody panel comprises fluorescently labeled antibodies against IgM, CD5, CD62L, CD294, CD69, CD38, PD1, CD11C, CD3, CD8, HLA-DR, CD24, CD337, CD123, CD141, CD1C, CD4, TACI, CD319, CD335, PDL1, CD10, CD45, CD16, IgD, CD40, CD19_TCRγδ, CD43, CD14, CD138, CD15, CD56, CD86, CD303, CD27, or any combination thereof.

24. The method according to claim 23, wherein the fluorescently labeled antibody panel further comprises a fluorescently labeled antibody against a cell surface marker indicative of live cells, dead cells, or both.

25. The method according to any one of claims 1 to 24, wherein the at least one immunophenotyping panel comprises a fluorescently labeled antibody panel against cell surface proteins associated with T cells.

26. The method according to claim 25, wherein the fluorescently labeled antibody panel comprises fluorescently labeled antibodies against TIGIT, CD5, CD28, CXCR5, CD39, TIM3, CD38, PD1, TCRVα7_2_TCRVδ1, CD95, CD3, CD8, HLA-DR, CD31, CCR4, CCR6, CCR7, CD57, ICOS, CD4, KLRG1, TCRVα24_Jα18, CD122, CD103, CXCR3, TCRVδ2, CD45, CCR10, CD16, CD25, CD161, CD19_TCRγδ, LAG3, CD14, CD45RO, CD56, CD127, CD45RA, CD27, or any combination thereof.

27. The method according to any one of claims 1 to 26, wherein the plurality of different immune cell subsets comprises at least 1000, 1200, 1400, 1600, 1800, 2000, 2200, or 2400 different immune cell subsets.

28. The method according to any one of claims 1 to 27, wherein the plurality of different immune cell subsets comprises white blood cells (WBC), eosinophils, eosinophils / CD5+, neutrophils, neutrophils / large, neutrophils / CD5+, neutrophils / small, B cells, B cells / CD5 - CD27 -, monocytes / CD56+, monocytes / CD56 -, NK cells, dendritic cells (DC), T cells, iNKT cells, γδT cells (total γδ), Vδ1 cells, Vδ2 cells, Vδx cells, mucosa-associated invariant T (MAIT) cells, TEMRA cells, CD4 naive cells, T helper cells, CD4 effector memory cells, Treg cells, or any combination thereof.

29. The method according to any one of claims 1 to 28, wherein the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample is output as part of an immune profile within less than 24 hours, 12 hours, 10 hours, 8 hours, 6 hours, or 4 hours.

30. The method according to any one of claims 1 to 29, wherein the immune profile is used to diagnose an immune-related disease or disorder, monitor the progression of an immune-related disease or disorder, or monitor the response to treatment of an immune-related disease or disorder in the subject.

31. A computer-implemented method for generating an immune profile of a subject, the method comprising: receiving fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from the subject using a full-spectrum flow cytometer or data derived therefrom; providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of different immune cell subsets; and outputting the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

32. A system, the system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom; provide at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of different immune cell subsets; and output the total cell count or cell frequency of each different immune cell subset among the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.

33. The system according to claim 32, wherein the system further comprises a full-spectrum flow cytometer.

34. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a system, cause the system to: receive fluorescence intensity data generated by processing a fluorescently labeled cell sample collected from a subject using a full-spectrum flow cytometer or data derived therefrom; Providing at least a subset of the fluorescence intensity data of the plurality of fluorescently labeled cells or data derived therefrom as an input to an integrated machine learning model configured to process the fluorescence intensity data or data derived therefrom and classify individual cells among the plurality of fluorescently labeled cells as belonging to one of a plurality of different immune cell subsets; and Outputting the total cell count or cell frequency of each of the plurality of different immune cell subsets in the sample as part of the immune profile of the subject.