Comprehensive immune profiling of peripheral blood
By analyzing cytometry and RNA expression data to generate leukocyte signatures and apply machine learning classifiers, the method addresses the limitations of existing immune profiling methods, enabling precise prediction of cancer prognosis and immunotherapy response.
Patent Information
- Application Number
- JP2025528233
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-14
- Filing Date
- 2023-11-17
- Publication Date
- 2025-12-23
AI Technical Summary
Existing immune profiling methods, such as RNA sequencing and cytometry, are inadequate for determining leukocyte immune profiles independently of a patient's health status and predicting response to immunotherapy, particularly in cancer patients.
A method involving cytometry data or RNA expression data analysis to determine cellular composition percentages for at least 20 cell types, generating a leukocyte signature, and identifying a leukocyte immune profile type using machine learning classifiers, such as a tabular pre-data fitted network transformer (TabPFN) classifier, to predict cancer prognosis and immunotherapy response.
Enables accurate characterization of leukocyte immune profiles, predicting cancer prognosis and likelihood of response to immunotherapy, facilitating personalized treatment decisions.
Smart Images

Figure 2025541671000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit under 35 U.S.C. § 119(e) of the filing dates of U.S. Provisional Patent Application No. 63 / 490,214, entitled "COMPREHENSIVE IMMUNOPROFILING OF PERIPHERAL BLOOD," filed March 14, 2023, and U.S. Provisional Patent Application No. 63 / 426,153, entitled "COMPREHENSIVE IMMUNOPROFILING OF PERIPHERAL BLOOD REVEALS FIVE CONSERVED IMMUNOTYPES WITH IMPLICATIONS FOR IMMUNOTHERAPY IN CANCER PATIENTS," filed November 17, 2022, the entire contents of each of which are incorporated herein by reference. [Background technology]
[0002] background Immune profiling methods include, but are not limited to, RNA sequencing (RNAseq) and cytometry. RNAseq is a method that can be used to determine the sequence and / or relative amount of RNA (e.g., RNA expressed by immune cells) in a sample. The sequence and relative expression level of RNA can be indicative of cellular characteristics. Cytometry is a laboratory technique used to analyze single cells or particles in biological samples. Cytometry is used in a variety of applications, including immunology and molecular biology. Cytometry can be used to measure characteristics of individual cells or particles. Types of cytometry include flow cytometry and mass cytometry.
[0003] Flow cytometry measures the intensity generated by fluorescent markers used to label cells in a biological sample. For example, cells labeled with one or more markers may be processed in a flow cytometry platform, where the fluorescence intensity of the markers is measured. The measured fluorescence intensity can be referred to as a "marker value" and can be used for various applications, such as cell counting, cell sorting, and / or determining various cell characteristics. Other types of cytometry (e.g., mass cytometry) can also be used for such applications. Summary of the Invention [Means for solving the problem]
[0004] overview Aspects of the present disclosure relate to methods, systems, and computer-readable storage media useful for characterizing a subject's leukocyte (e.g., white blood cell (WBC) or peripheral blood mononuclear cell (PBMC)) immune profile type. The leukocyte immune profile type can be determined independently of the patient's health status, e.g., whether the patient is healthy or has, is suspected of having, or is at risk of having cancer.
[0005] The present disclosure is based, in part, on methods for immune profiling of cancer subjects based on analysis of leukocyte populations in the subject's peripheral blood and the subject's prognosis and / or likelihood of response to immunotherapy. In some embodiments, the methods described by the present disclosure are useful for determining a leukocyte immune profile type of a subject with cancer. In some embodiments, the subject's leukocyte immune profile type is indicative of the subject's cancer prognosis (e.g., pancreatic cancer, breast cancer, non-small cell lung cancer, colorectal cancer, melanoma, prostate cancer, etc.) and / or the likelihood that the subject will respond to treatment with a particular therapeutic agent, e.g., an immunotherapeutic agent such as an immune checkpoint inhibitor (ICI). In some embodiments, the subject's leukocyte immune profile type is indicative of the subject's head and neck squamous cell carcinoma (HNSCC) prognosis and / or the likelihood that a subject with HNSCC will respond to treatment with a particular therapeutic agent, e.g., an immunotherapeutic agent such as an ICI.
[0006] Thus, in some aspects, the present disclosure provides a method of determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, the method comprising: obtaining, using at least one computer hardware processor, cytometry data or RNA expression data from a biological sample obtained from the subject; processing the cytometry data or RNA expression data to determine cellular composition percentages for at least 20 cell types listed in Table 4; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, the leukocyte signature comprising the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types.
[0007] In some aspects, the present disclosure provides systems including at least one computer hardware processor and at least one computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method of determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, the method including: obtaining, using the at least one computer hardware processor, cytometry data or RNA expression data from a biological sample obtained from the subject; processing the cytometry data or RNA expression data to determine cellular composition percentages for at least 20 cell types listed in Table 4; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, the leukocyte signature comprising the cellular composition percentages for the at least 20 cell types; and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types using the leukocyte signature.
[0008] In some aspects, the present disclosure provides at least one computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method of determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, the method including: using at least one computer hardware processor to obtain cytometry data or RNA expression data from a biological sample obtained from the subject; processing the cytometry data or RNA expression data to determine cellular composition percentages for at least 20 cell types listed in Table 4; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, the leukocyte signature comprising the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types.
[0009] Embodiments of any of the above aspects may have one or more of the following features.
[0010] In some embodiments, the cytometry data comprises flow cytometry data. In some embodiments, the flow cytometry data is obtained from a biological sample consisting of white blood cells.
[0011] In some embodiments, processing the flow cytometry data includes determining the cellular composition percentage for each cell type listed in Table 1. In some embodiments, the flow cytometry data is obtained from a biological sample consisting of peripheral blood mononuclear cells (PBMCs).
[0012] In some embodiments, processing the flow cytometry data includes determining the cellular composition percentage for each cell type listed in Table 2.
[0013] In some embodiments, processing the cytometry data to obtain cellular composition percentages for at least 20 cell types listed in Table 4 comprises applying one or more machine learning models to the cytometry data.
[0014] In some embodiments, obtaining the RNA expression data comprises obtaining sequencing data previously obtained by sequencing a biological sample obtained from the subject. In some embodiments, the sequencing data comprises at least 1 million reads, at least 5 million reads, at least 10 million reads, at least 20 million reads, at least 50 million reads, or at least 100 million reads.
[0015] In some embodiments, the method further comprises normalizing the RNA expression data to transcripts per million (TPM) units before processing the RNA expression data to determine cellular composition percentages.
[0016] In some embodiments, a plurality of leukocyte immune profile types are associated with each of the plurality of leukocyte immune profile types, and identifying a leukocyte immune profile type for the subject using the leukocyte signature and from among the plurality of leukocyte immune profile types includes associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types, and identifying the leukocyte immune profile type for the subject as a leukocyte immune profile type that corresponds to the particular one of the plurality of leukocyte immune profile types with which the subject's leukocyte signature is associated.
[0017] In some embodiments, associating the subject's leukocyte signature with a particular one of a plurality of leukocyte immune profile types comprises processing the leukocyte signature with a trained classifier to obtain an output indicative of the particular one of the plurality of leukocyte immune profile types. In some embodiments, the trained classifier comprises a trained neural network classifier, optionally a tabular pre-data fitted network transformer (TabPFN) classifier.
[0018] In some embodiments, associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types includes determining, for each particular one of the plurality of leukocyte immune profile types, a score indicating whether the subject's leukocyte signature is associated with that particular leukocyte immune profile type, and determining the score for the particular leukocyte immune profile type includes applying a linear regression model associated with the particular leukocyte immune profile type to the cellular composition percentages in the leukocyte signature.
[0019] In some embodiments, the method includes generating a plurality of leukocyte immune profile types, wherein the generating comprises obtaining a plurality of cytometry data or RNA expression datasets from biological samples obtained from a plurality of respective subjects, wherein each of the plurality of cytometry data or RNA expression datasets indicates a cellular composition percentage for at least 20 cell types listed in Table 4; generating a plurality of leukocyte signatures from the plurality of cytometry data or RNA expression datasets, wherein each of the plurality of leukocyte signatures comprises a cellular composition percentage for at least 20 cell types listed in Table 4, wherein the generating comprises, for each particular one of the plurality of leukocyte signatures, determining the leukocyte signature by using the cytometry data or RNA expression data to determine a cellular composition percentage in the particular cytometry data or RNA expression dataset from which the particular one leukocyte signature was generated; and Clustering multiple leukocyte signatures to obtain multiple leukocyte immune profile types It further includes including:
[0020] In some embodiments, the method further includes updating a plurality of leukocyte immune profile types using the subject's leukocyte signature, wherein the subject's leukocyte signature is one of a threshold number of leukocyte signatures for a threshold number of subjects, and the leukocyte immune profile type is updated when the threshold number of leukocyte signatures are generated, wherein the threshold number of leukocyte signatures are at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, or at least 5000 leukocyte signatures.
[0021] In some embodiments, updating the clusters is performed using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and an agglomerative clustering algorithm.
[0022] In some embodiments, the method further includes determining a leukocyte immune profile type of a second subject, wherein the leukocyte immune profile type of the second subject is identified using an updated leukocyte immune profile type, wherein the identifying includes determining a leukocyte signature of the second subject from cytometry data or RNA expression data from a biological sample obtained from the second subject; associating the leukocyte signature of the second subject with a particular one of the plurality of updated leukocyte immune profile types; and identifying the leukocyte immune profile type for the second subject as the leukocyte immune profile type corresponding to the particular one of the plurality of updated leukocyte immune profile types with which the leukocyte signature of the second subject is associated.
[0023] In some embodiments, the clustering is performed using a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm, hi some embodiments, the clustering is performed using a spectral clustering algorithm.
[0024] In some embodiments, the plurality of leukocyte immune profile types comprises a naive type, a primed type, a progressive type, a chronic type, and a suppressed type.
[0025] In some embodiments, the method further includes identifying the subject as a candidate for immunotherapeutic treatment based on identifying the leukocyte immune profile type for the subject, hi some embodiments, the method further includes identifying the subject as a candidate for immunotherapeutic treatment when the subject is identified as having a primed type.
[0026] In some embodiments, the method further comprises administering a therapeutic agent to the subject based on the identification of the subject's leukocyte immune profile type, hi some embodiments, the method further comprises administering an immunotherapy to the subject when the subject is identified as having a primed type.
[0027] In some aspects, the present disclosure provides a method of determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, the method comprising: using at least one computer hardware processor to obtain flow cytometry data for white blood cells (WBCs) isolated from a biological sample obtained from the subject; processing the flow cytometry data to determine cellular composition percentages for at least 20 cell types listed in Table 1; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, wherein the leukocyte signature comprises the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types.
[0028] In some embodiments, the WBCs consist of granulocytic and agranulocytic cells.
[0029] In some embodiments, processing the flow cytometry data includes applying one or more machine learning models to the flow cytometry data to obtain cellular composition percentages for at least 20 cell types listed in Table 1.
[0030] In some embodiments, processing the flow cytometry data includes determining the cellular composition percentages for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD4+ T cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
[0031] In some embodiments, processing the flow cytometry data includes identifying naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, V52+ γδ T cells, post-class-switched memory B cells, central memory CD8+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD39+ CD4+ Tregs, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, effector memory CD4+ T cells, NKT cells, CD8+ TEMRA, effector memory CD8+ T cells, CD4+ This involves determining the cellular composition percentages for TEMRA, neutrophils, granulocytes, classical monocytes, non-classical monocytes, and HLA-DR low monocytes.
[0032] In some embodiments, the leukocyte signature comprises cellular composition percentages for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD4+ T cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
[0033] In some embodiments, the leukocyte signature is selected from the group consisting of naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, V52+ γδ T cells, post-class switched memory B cells, central memory CD8+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD39+ CD4+ Tregs, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, effector memory CD4+ T cells, NKT cells, CD8+ TEMRA, effector memory CD8+ T cells, CD4+ Includes cell composition percentages for TEMRA, neutrophils, granulocytes, classical monocytes, non-classical monocytes, and HLA-DR low monocytes.
[0034] In some embodiments, a plurality of leukocyte immune profile types are associated with each of the plurality of leukocyte immune profile types, and identifying a leukocyte immune profile type for the subject using the leukocyte signature and from among the plurality of leukocyte immune profile types includes associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types, and identifying the leukocyte immune profile type for the subject as a leukocyte immune profile type that corresponds to the particular one of the plurality of leukocyte immune profile types with which the subject's leukocyte signature is associated.
[0035] In some embodiments, associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types comprises processing the leukocyte signature with a trained classifier to obtain an output indicative of the particular one of the plurality of leukocyte immune profile types. In some embodiments, the trained classifier comprises a trained neural network classifier, optionally a tabular pre-data fitted network transformer (TabPFN) classifier.
[0036] In some embodiments, associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types includes determining, for each particular one of the plurality of leukocyte immune profile types, a score indicating whether the subject's leukocyte signature is associated with that particular leukocyte immune profile type, and determining the score for the particular leukocyte immune profile type includes applying a linear regression model associated with the particular leukocyte immune profile type to the cellular composition percentages in the leukocyte signature.
[0037] In some embodiments, generating the plurality of leukocyte immune profile types includes obtaining a plurality of flow cytometry data sets from white blood cells (WBCs) isolated from biological samples obtained from a plurality of respective subjects, each of the plurality of flow cytometry data sets indicating a cellular composition percentage for at least 20 cell types listed in Table 1; generating a plurality of leukocyte signatures from the plurality of flow cytometry data sets, each of the plurality of leukocyte signatures including a cellular composition percentage for at least 20 cell types listed in Table 1, wherein generating includes, for each particular one of the plurality of leukocyte signatures, determining the leukocyte signature by determining the cellular composition percentage using flow cytometry data in the particular flow cytometry data set from which the particular one leukocyte signature was generated; and clustering the plurality of leukocyte signatures to obtain the plurality of leukocyte immune profile types.
[0038] In some embodiments, the method further includes updating a plurality of leukocyte immune profile types using the subject's leukocyte signature, wherein the subject's leukocyte signature is one of a threshold number of leukocyte signatures for a threshold number of subjects, and the leukocyte immune profile type is updated upon generation of the threshold number of leukocyte signatures, wherein the threshold number of leukocyte signatures is at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, or at least 5000 leukocyte signatures.
[0039] In some embodiments, the updating is performed using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and an agglomerative clustering algorithm.
[0040] In some embodiments, the method further includes determining a leukocyte immune profile type of a second subject, wherein the leukocyte immune profile type of the second subject is identified using an updated leukocyte immune profile type, wherein the identifying includes determining a leukocyte signature of the second subject from flow cytometry data from leukocyte cells isolated from a biological sample obtained from the second subject; associating the leukocyte signature of the second subject with a particular one of the plurality of updated leukocyte immune profile types; and identifying the leukocyte immune profile type for the second subject as the leukocyte immune profile type corresponding to the particular one of the plurality of updated leukocyte immune profile types with which the leukocyte signature of the second subject is associated.
[0041] In some embodiments, the clustering is performed using a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm, hi some embodiments, the clustering is performed using a spectral clustering algorithm.
[0042] In some embodiments, the plurality of leukocyte immune profile types comprises a naive type, a primed type, a progressive type, a chronic type, and a suppressed type.
[0043] In some embodiments, the method further comprises identifying the subject as a candidate for immunotherapeutic treatment based on identifying the leukocyte immune profile type for the subject, hi some embodiments, the method further comprises identifying the subject as a candidate for immunotherapeutic treatment when the subject is identified as having a primed type.
[0044] In some embodiments, the method further comprises administering a therapeutic agent to the subject based on the identification of the subject's leukocyte immune profile type, hi some embodiments, the method further comprises administering an immunotherapy to the subject when the subject is identified as having a primed type.
[0045] In some embodiments, the subject has head and neck squamous cell carcinoma (HNSCC).
[0046] In some aspects, the present disclosure provides methods for determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, the method comprising: using at least one computer hardware processor to obtain RNA expression data for peripheral blood mononuclear cells (PBMCs) isolated from a biological sample obtained from the subject; processing the RNA expression data to determine cellular composition percentages for at least 20 cell types listed in Table 3; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, wherein the leukocyte signature comprises the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature to identify a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types.
[0047] In some embodiments, the RNA expression data comprises bulk RNA expression data. In some embodiments, processing the RNA expression data comprises applying a cellular deconvolution technique comprising one or more machine learning models to obtain cellular composition percentages.
[0048] In some embodiments, the cellular composition percentages include cellular composition percentages for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
[0049] In some embodiments, the composition percentages include cell composition percentages for naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-switched memory IgM B cells, class-switched memory B cells, central memory, CD4+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD4+ T cells, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, CD8+ TEMRA, effector memory CD8+ T cells, neutrophils, granulocytes, classical monocytes, and non-classical monocytes.
[0050] In some embodiments, the leukocyte signature comprises the cellular composition percentages for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
[0051] In some embodiments, the leukocyte signature comprises cellular composition percentages for naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, class-switched memory B cells, central memory, CD4+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD4+ T cells, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, CD8+ TEMRA, effector memory CD8+ T cells, neutrophils, granulocytes, classical monocytes, and non-classical monocytes.
[0052] In some embodiments, a plurality of leukocyte immune profile types are associated with each of the plurality of leukocyte immune profile types, and identifying a leukocyte immune profile type for the subject using the leukocyte signature and from among the plurality of leukocyte immune profile types includes associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types, and identifying the leukocyte immune profile type for the subject as a leukocyte immune profile type that corresponds to the particular one of the plurality of leukocyte immune profile types with which the subject's leukocyte signature is associated.
[0053] In some embodiments, associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types comprises processing the leukocyte signature with a trained classifier to obtain an output indicative of the particular one of the plurality of leukocyte immune profile types. In some embodiments, the trained classifier comprises a trained neural network classifier, optionally a tabular pre-data fitted network transformer (TabPFN) classifier.
[0054] In some embodiments, associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types comprises determining, for each particular one of the plurality of leukocyte immune profile types, a score indicating whether the subject's leukocyte signature is associated with that particular leukocyte immune profile type, wherein determining the score for the particular leukocyte immune profile type comprises applying a linear regression model associated with the particular leukocyte immune profile type to the cellular composition percentages in the leukocyte signature.
[0055] In some embodiments, the method further comprises generating a plurality of leukocyte immune profile types, wherein the generating comprises obtaining a plurality of RNA expression datasets from white blood cells (WBCs) isolated from biological samples obtained from a plurality of respective subjects, each of the plurality of RNA expression datasets indicating a cellular composition percentage for at least 20 cell types listed in Table 3; generating a plurality of leukocyte signatures from the plurality of RNA expression datasets, each of the plurality of leukocyte signatures including a cellular composition percentage for at least 20 cell types listed in Table 3, wherein the generating comprises, for each particular one of the plurality of leukocyte signatures, determining the cellular composition percentage using RNA expression data in the particular RNA expression dataset from which the particular one leukocyte signature was generated; and clustering the plurality of leukocyte signatures to obtain the plurality of leukocyte immune profile types.
[0056] In some embodiments, the method further includes updating a plurality of leukocyte immune profile types using the subject's leukocyte signature, wherein the subject's leukocyte signature is one of a threshold number of leukocyte signatures for a threshold number of subjects, and the leukocyte immune profile type is updated upon generation of the threshold number of leukocyte signatures, wherein the threshold number of leukocyte signatures is at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, or at least 5000 leukocyte signatures.
[0057] In some embodiments, the updating is performed using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and an agglomerative clustering algorithm.
[0058] In some embodiments, the method further includes determining a leukocyte immune profile type of the second subject, wherein the leukocyte immune profile type of the second subject is identified using the updated leukocyte immune profile type, wherein the identifying includes: determining a leukocyte signature of the second subject from RNA expression data from leukocyte cells isolated from a biological sample obtained from the second subject; associating the leukocyte signature of the second subject with a particular one of the plurality of updated leukocyte immune profile types; and identifying the leukocyte immune profile type for the second subject as the leukocyte immune profile type corresponding to the particular one of the plurality of updated leukocyte immune profile types with which the leukocyte signature of the second subject is associated.
[0059] In some embodiments, the clustering is performed using a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm, hi some embodiments, the clustering is performed using a spectral clustering algorithm.
[0060] In some embodiments, the plurality of leukocyte immune profile types comprises a naive type (G1), a primed type (G2), a progressive type (G3), a chronic type (G4), and a suppressed type (G5).
[0061] In some embodiments, the method further includes identifying the subject as a candidate for immunotherapeutic treatment based on identifying the leukocyte immune profile type for the subject, hi some embodiments, the method further includes identifying the subject as a candidate for immunotherapeutic treatment when the subject is identified as having a primed type.
[0062] In some embodiments, the method further comprises administering a therapeutic agent to the subject based on the identification of the subject's leukocyte immune profile type, hi some embodiments, the method further comprises administering an immunotherapy to the subject when the subject is identified as having a primed type.
[0063] In some embodiments, the subject has head and neck squamous cell carcinoma (HNSCC). [Brief explanation of the drawings]
[0064] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] 1 provides a representative example of a method for identifying a subject's leukocyte immune profile type using cytometry data, according to some embodiments of the technology described herein. [Figure 2] 1 provides a representative example of a method for identifying a subject's leukocyte immune profile type using RNA expression data, according to some embodiments of the technology described herein. [Figure 3] 1 is a flowchart of an exemplary method for determining a cellular composition percentage based on a cell count determined for a plurality of cells of a biological sample, according to some embodiments of the technology described herein. [Figure 4] FIG. 1 depicts an exemplary technique for using leukocyte signatures (e.g., WBC signatures or PBMC signatures) to identify leukocyte immune profile types associated with each leukocyte signature cluster, according to some embodiments of the technology described herein. [Figure 5A] FIG. 1 shows a schematic diagram depicting exemplary cohorts used to generate leukocyte immune profile types according to some embodiments of the technology as described herein. [Figure 5B] FIG. 1 shows a schematic diagram depicting exemplary cohorts used to generate leukocyte immune profile types according to some embodiments of the technology as described herein. [Figure 6A]5A-5B are representative heat maps of flow cytometry data obtained for the cohorts depicted in FIGS. 5A-5B, according to some embodiments of the technology described herein. [Figure 6B] 5A-5B are representative heat maps of flow cytometry data obtained for the cohorts depicted in FIGS. 5A-5B, according to some embodiments of the technology described herein. [Figure 7A] 10A-10C are representative heatmaps of cellular deconvolution analysis performed on RNA-seq data from samples belonging to different leukocyte immune profile types according to some embodiments of the technology described herein. [Figure 7B] 10A-10C are representative heatmaps of cellular deconvolution analysis performed on RNA-seq data from samples belonging to different leukocyte immune profile types according to some embodiments of the technology described herein. [Figure 8] FIG. 8 is an exemplary heatmap showing differential gene expression analysis performed on RNA-seq data from samples belonging to leukocyte immune profile types, according to some embodiments of the technology described herein. [Figure 9A] 9A-9B show representative data for T cell receptor (TCR) and B cell receptor (BCR) analysis of leukocyte immune profile types, according to some embodiments of the technology described herein. The top panel of FIG. 9A represents clonality and Chao1 metric data for α T cell receptor (TCR) variants within five leukocyte immune profile types. The middle panel represents clonality and Chao1 metric data for β TCR variants within five leukocyte immune profile types. The bottom panel represents representative data for heavy, lambda, and kappa chain clonality of BCR clonotypes in samples within five leukocyte immune profile types. [Figure 9B] Representative data for TCR clonality, TCR diversity, and lambda and kappa chains of BCR clonotypes for five leukocyte immune profile types are shown. [Figure 10A] FIG. 1 shows a schematic diagram depicting an example immune profiling pipeline according to some embodiments of the technology as described herein. [Figure 10B] 1 shows representative data of a cytometry panel of cell populations shown as heatmaps of normalized signal intensities of immune cell populations and t-distributed stochastic neighbor embedding (tSNE), according to some embodiments of the technology described herein. [Figure 11A] 11A-11D show representative data for prediction of healthy and cancer samples using a machine learning (ML) classifier, according to some embodiments of the technology described herein. Figure 11A shows a representative Uniform Manifold Approximation Projection (UMAP) plot of 450 immune cell populations. [Figure 11B] Volcano plots demonstrating differentially represented populations between healthy individuals and cancer patients are shown. [Figure 11C] A comparison of UMAPs of a selected population with true and predicted labels is shown, with the gradient between cancer patients and healthy individuals indicated. [Figure 11D] Representative data of ROC-AUC for the classification quality of the model for the training (left) and validation cohorts (right) for 20 selected populations compared to a T and B natural killer cell (TBNK) panel, which includes the most common populations used in cytometry studies, are shown. [Figure 12A] 12A shows representative data for leukocyte signature clustering according to some embodiments of the technology described herein. Figure 12A shows that unsupervised spectral clustering analysis applied to normalized flow cytometry percentages revealed five distinctively distinct leukocyte immune profile types. [Figure 12B] Corresponding RNA-seq (n = 824) data are shown, which provided orthogonal validation for flow cytometry (FC) analysis of cell percentages using the cell deconvolution algorithm Kassandra. [Figure 12C]We show another example of corresponding RNA-seq data that provided orthogonal validation for flow cytometry (FC) analysis of cell percentages using the cell deconvolution algorithm Kassandra. [Figure 12D] A representative flow cytometry database pseudotime analysis graph is shown, which analyzes the connections between different peripheral blood samples using cell percentages obtained from flow cytometry data analysis. [Figure 13A] 13A-13C show exemplary data indicating differential expression of cytokine pathways across leukocyte immune profile types, according to some embodiments of the technology described herein. Figure 13A shows an exemplary heatmap indicating correlations were found between functional gene signatures for cytokine-related pathways from the MSigDB database across leukocyte immune profile types G1-G5. The health status of each patient is also shown. [Figure 13B] Representative heatmaps are shown showing correlations between functional gene signatures for cytokine-related pathways from the MSigDB database across leukocyte immune profile types G1-G5. The health status of each patient is also shown. [Figure 13C] Representative data are shown showing a comparison of differential gene expression levels of five leukocyte immune profile types, cytokine and chemokine genes FLT3LG, CCL4, CXCL16, CCR7, TGFBR3 and IL1R1. [Figure 14] 14 shows representative data for TCR and BCR analysis of leukocyte immune profile types. Analysis of TCR (for both α and β chain) landscape stratified by leukocyte immune profile type (G1-G5) according to some embodiments of the technology described herein. [Figure 15] FIG. 15 depicts an exemplary implementation of a computer system that may be used in connection with some embodiments of the technology described herein. [Figure 16A]16A-16C show an embodiment of a cluster formation workflow, model training, cytometry data composition, and an overall description of the differences between healthy and cancer patients in cytometry data, according to some embodiments of the technology described herein. Figure 16A shows an embodiment of the overall cluster formation workflow: blood draw, hematology analyzer, sample processing, flow cytometry run, manual data labeling, machine learning model training and implementation, and subsequent cohort analysis. [Figure 16B] Representative heatmaps show the signal intensities of various cell surface markers by cell population, combined with tSNE-based panel-specific immune cell population images across different panels. [Figure 16C] Representative data are shown in a polar graph-like tree-type scatter plot image, highlighting differences between healthy donors and cancer patients. [Figure 17A] 17A-17G show representative data describing a patient cohort, according to some embodiments of the technology described herein. Figure 17A shows a description of the cohort, along with disease breakdown and clinical annotations. [Figure 17B] Figure 1 shows that the cohort was analyzed for different clinical groups and divided into training and validation sets for machine learning applications. [Figure 17C] Representative uMAP analysis images of raw cytometry data are shown, with representative age, diagnosis, and treatment groups individually highlighted. [Figure 17D] Representative volcano plots showing differentially represented clusters between healthy and cancer patients, selected with the Maximum Relevance Minimum Redundancy (MRMR) algorithm. [Figure 17E] Representative box plots showing the distribution of cell populations between healthy individuals and cancer patients for naive CD4 T cells, monocytes, IgM unconverted memory B cells, and CX3CR1-CD8 TEMRA are shown. [Figure 17F] Representative uMAP analysis images of selected cell populations are shown, highlighting some features. [Figure 17G]Representative ROC AUC analyses of healthy / cancer classifiers trained on select populations and the TBNK panel, respectively. Training and validation cohort splits are shown. [Figure 18A] 18A-18E show representative data for clustering of cell populations according to some embodiments of the technology described herein. Figure 18A shows a heatmap after clustering by five selected distinct leukocyte immune type clusters (G1-naive, G2-primed, G3-progressive, G4-chronic, and G5-suppressed); healthy and cancer patients are highlighted. [Figure 18B] A bubble plot of the median cell percentages for a representative flow cytometry database and a Kassandra deconvolution database is shown. [Figure 18C] Representative heatmaps of GSEA scores for RNA-seq data samples grouped by five clusters and age distribution among clusters shown as histogram plots are shown. [Figure 18D] Representative violin plots showing the distribution of selected cytokine expression are shown. [Figure 18E] Representative pseudotime plot images are shown, highlighting samples representative of the five clusters. [Figure 19A] 19A-19C show representative data analysis of leukocyte immune profile clustering analysis according to some embodiments of the technology described herein: Figure 19A shows a three-dimensional principal component analysis (PCA) representation of a blood RNA-seq dataset processed by the Kassandra cell deconvolution algorithm. [Figure 19B] Bar graphs showing the distribution of five clusters (G1-naive, G2-primed, G3-advanced, G4-chronic, and G5-suppressed) within each disease group are shown. [Figure 19C] Representative data are shown indicating statistical differences between leukocyte immune types within each disease group compared to healthy samples. Color intensity depends on the log p-value, and circle size is proportional to the proportion within the disease group. [Figure 20A] 20A-20G show representative data for analysis of HLA alleles and TCR / BCR repertoires according to some embodiments of the technology described herein. Figure 20A shows the HLA allele distribution for HLA type I genes. Only alleles with counts of 5 or greater are shown and labeled in different colors. [Figure 20B] Representative TCR-β clonality distribution within five clusters (G1-naive, G2-primed, G3-progressive, G4-chronic, and G5-suppressed) using swarm plots is shown. [Figure 20C] Representative TCR-β chao1 (diversity metric) distribution within the five clusters using swarm plots is shown. [Figure 20D] A representative TCR-β landscape is shown, with the top row showing CDR3β coverage, the middle row highlighting the explicit TCR-β ratios, and the bottom row showing the clusters and healthy / cancer status. [Figure 20E] Representative GSEA scores for different gene groups for all five clusters are shown, as well as the distribution of differentially expressed genes (ID3, TCF7, and LEF1 are overexpressed in the G1 cluster, while TBX21, EOMES, and TOX are overexpressed in the G4 cluster). [Figure 20F] Representative gene cluster scores for PD1 signaling and cancer immunotherapy with PD1 blockade among patients receiving ICB for all five clusters are shown. [Figure 20G] Representative data of log2TPM of the PDCD1 gene cluster for all five clusters are shown. [Figure 21A] 21A-21F show exemplary data for PBMC immune profile typing of head and neck squamous cell carcinoma (HNSCC) subjects, according to some embodiments of the technology described herein. Figure 21A shows a representation of a representative HNSCC cohort. Treatments and blood collection time points are shown. [Figure 21B] A representative uMAP analysis of data from a pan-cancer patient cohort compared with a representative HNSCC patient cohort is shown. [Figure 21C] Representative heatmaps of the HNSCC cohort for all five clusters (G1-G5) are shown with the corresponding cell populations, highlighting responders and non-responders and pre- and post-treatment time points. [Figure 21D] Representative bar graphs are shown for cluster-based analysis (left) and response-based analysis (right) of pre-treatment samples for all five PBMC immune profile clusters (G1-G5). [Figure 21E] A representative Sankey plot indicating the shift of the five clusters after treatment is shown. [Figure 21F] Representative data are shown for bar graphs of cluster-based analysis (left) and response-based analysis (right) of post-treatment samples for all five clusters (G1-G5). [Figure 22A] 22A-22F show representative data for HNSCC sample analysis according to some embodiments of the technology described herein. Figure 22A shows a representative volcano plot illustrating differentially represented populations between responders and non-responders within pre-treatment samples of an HNSCC cohort. Populations with a difference of greater than 20% at p-values <0.05 are highlighted. [Figure 22B] Representative waterfall plots are shown with overrepresented cell populations with a difference of more than 20% at p-values <0.05 for pre-treatment samples. [Figure 22C] Representative boxplots showing the distribution of selected cell populations between responders and non-responders within the pre-treatment cohort are shown. [Figure 22D] Representative uMAP images of cluster (G1-G5) distribution (left) and primed (G2) cluster signature scores (right) are shown. [Figure 22E] Representative data for primed (G2) signature distribution between non-responders and responders shown as box plots for pre-treatment (left) and post-treatment (right) samples are shown. [Figure 22F] Representative ROC-AUC curves based on cohort-based median cutoffs are shown. [Figure 23A] 23A-23B show representative data for blood immune profile cluster analysis according to some embodiments of the technology described herein. Figure 23A shows a representative heatmap showing gene set enrichment analysis scores performed on RNA-seq samples aligned in five clusters using gene sets from mSigDB. [Figure 23B] Representative cell lineage diagrams for the cell populations used in the cluster analysis are shown, highlighting clinical value-added populations. [Figure 24A] Figures 24A-24M show representative data comparing five clusters (G1-G5) with TCR / BCR composition, according to some embodiments of the technology described herein. Figure 24A shows a bubble plot showing HLA-B allele distribution among the five clusters, highlighting the mean TCR β and α clonality. [Figure 24B] A bubble plot showing the HLA-B allele distribution among the five clusters is shown, highlighting the mean TCR β and α clonality. [Figure 24C] A representative distribution of the percentage of initial clonotypes for TCRβ classified by five clusters is shown. [Figure 24D] Representative TCR-α clonality distributions compared by cluster type (G1 to G5) using swarm plots are shown. [Figure 24E] Representative TCR-α chao1 (diversity metric) distributions for the five clusters using swarm plots are shown. [Figure 24F] A representative distribution of the percentage of initial clonotypes for TCRα within the five clusters is shown. [Figure 24G] A swarm plot is shown, showing a representative distribution of the percentage of common (tumor-intersecting) clonotypes in the five clusters. [Figure 24H]A swarm plot is shown, showing a representative distribution of the percentage of common (cross-blood) clonotypes in the five clusters per cluster type. [Figure 24I] Scatter plots showing representative percentages in both blood and tumor for common clonotypes (G1, G2, G3, G4, and G5) are shown. [Figure 24J] Figure 1 shows the BCR heavy chain chao1 (diversity metric) distribution within the five clusters using a swarm plot. [Figure 24K] Figure 1 shows the BCRλ chao1 (diversity metric) distribution within the five clusters using a swarm plot. [Figure 24L] The distribution of BCRκ chao1 (diversity metric) within the five clusters shown in the swarm plot is shown. [Figure 24M] Representative TCR-α landscapes are shown, with the top row showing CDR3β coverage, the middle row highlighting the explicit TCR-α ratio, and the bottom row showing clusters (G1–G5) and healthy / cancer status. [Figure 25] FIG. 25 shows that there were significant differences between healthy donors and cancer patients in absolute numbers of RBCs, platelets, neutrophils, and lymphocytes, but similar absolute numbers of monocytes, according to some embodiments of the technology described herein. [Figure 26A] 26A-26C show immune profiling analysis of a representative head and neck squamous cell carcinoma (HNSCC) cohort according to some embodiments of the technology described herein. Figure 26A shows a representative HNSCC cohort analysis workflow schema. [Figure 26B] Box plots showing representative distribution of PBMC percentage of all cell events for HNSCC frozen PBMCs compared to other blood source processing techniques are shown. [Figure 26C] Box plots showing the distribution of CD62L+ CD8 cell percentages from PBMCs in comparison with other blood source processing techniques are shown. [Figure 27A]27A-27F show exemplary data for training and validation of a machine learning model for blood immune profile cluster identification according to some embodiments of the technology described herein. Figure 27A shows an exemplary healthy / cancer classifier training workflow schema. [Figure 27B] 1 shows the correlation between model cell type predictions and manual markup. [Figure 27C] DE Volcano plots (fold change vs. false discovery rate) of genes for each of the five clusters G1 to G5 are shown. [Figure 27D] A representative Kassandra deconvolution heatmap with labels based on flow cytometry clustering is shown. [Figure 27E] Representative data for the cytometry / deconvolution comparison with values normalized to the internal cohort distribution are shown. [Figure 27F] 1 shows a representative machine learning model for flow cytometry data analysis training workflow schema. [Figure 28] FIG. 28 shows a flowchart illustrating data preparation, model training, and signature calculation of a new blood sample according to some embodiments of the technology described herein. [Figure 29A] FIG. 29A depicts an exemplary technique for determining RNA percentages based on RNA expression data according to some embodiments of the technology described herein. [Figure 29B] FIG. 29B depicts an example of determining RNA percentages based on RNA expression data using one or more machine learning models according to some embodiments of the technology described herein. [Figure 30] FIG. 30 depicts an exemplary technique for processing cytometry data for one or more cells to determine their respective types according to some embodiments of the technology described herein. DETAILED DESCRIPTION OF THE INVENTION
[0065] Detailed Description Aspects of the present disclosure relate to methods, systems, and computer-readable storage media useful for characterizing the immune profile of a subject (e.g., a healthy subject or a subject diagnosed with cancer). In some embodiments, the methods described by the present disclosure are useful for determining a subject's leukocyte immune profile. In some embodiments, the leukocyte immune profile type is determined from a biological sample comprising or consisting of (or consisting essentially of) the subject's white blood cells (WBCs). In some embodiments, the leukocyte immune profile is determined from a biological sample comprising or consisting of (or consisting essentially of) the subject's peripheral blood mononuclear cells (PBMCs). The present disclosure is based, in part, on methods for immune profiling a subject with cancer based on analysis of leukocyte populations in the subject's blood and the subject's prognosis and / or likelihood of response to immunotherapy. In some embodiments, the methods described by the present disclosure are useful for determining a leukocyte immune profile type (also referred to in some embodiments as a white blood cell (WBC) or peripheral blood mononuclear cell (PBMC) immune profile type) of a subject with cancer. In some embodiments, a subject's leukocyte immune profile type is indicative of the subject's cancer prognosis and / or the likelihood that the subject will respond to treatment with a particular therapeutic agent, e.g., an immunotherapeutic agent such as an immune checkpoint inhibitor (ICI).
[0066] The highly heterogeneous nature of cancer presents a significant therapeutic challenge. For example, different patients diagnosed with the same cancer may respond differently to the same treatment. Therefore, there is a need to identify patient and cancer characteristics that indicate the type of treatment that a patient is likely to respond to. Previous methods for identifying these characteristics have focused on classifying patients according to cancer subtypes, for example, by using cancer cell histology or RNA sequencing data and statistical analysis. This classification is then used to determine whether a given treatment is expected to be effective for a particular subject. These methods require obtaining tumor tissue samples from the subject, which is often highly invasive (e.g., requires surgery), time-consuming, and expensive.
[0067] Aspects of the present disclosure relate to methods of determining a subject's leukocyte immune profile type (e.g., WBC or PBMC immune profile type) by analyzing WBC or PBMC cytometry data (e.g., flow cytometry data or CYTOF data) or RNA-seq data obtained from a healthy or diseased subject (e.g., a subject with cancer, an infectious disease, an autoimmune disease, or an inflammatory disease) using machine learning-based techniques. The inventors have recognized that by analyzing the percentage composition of certain cell types in a subject's peripheral blood, it is possible to determine a leukocyte signature (also referred to in some embodiments as a WBC signature or PBMC signature) that characterizes the subject's immune type and whether the subject is healthy or diseased, regardless of the type of disease. The inventors also recognized that a collection of five reproducible leukocyte immune types (naive, primed, advanced, chronic, and suppressed, described further below) can be identified independent of whether a subject is healthy or diseased, and based on a subject's WBC or PBMC flow cytometry and / or RNA expression data. This is an improvement over previous immunotyping techniques because, while previous techniques focused on subclassifying patients with the same cancer type, the immune types identified by the methods described herein are conserved across healthy subjects and subjects with a variety of different cancer types. Thus, the leukocyte immune profile types described herein may have pan-cancer utility in determining effective treatments for a given patient. The inventors have also recognized that the specific immune types described herein are indicative of a positive response to certain therapies in patients diagnosed with certain cancers, such as head and neck squamous cell carcinoma (HNSCC), and therefore can be used to determine which therapies to administer (or not administer) to a given patient.
[0068] Aspects of the present disclosure relate to methods for identifying a subject as having one of five distinctively distinct leukocyte immune types (naive, primed, advanced, chronic, and suppressed) by analyzing WBC or PBMC cytometry data and / or RNA sequencing data indicative of the subject's WBC or PBMC cellular composition. The five identified leukocyte immune types are characterized by distinct distributions of immune cell types and activation states, reflecting underlying immunological processes and tissue microenvironments. Analysis of over 18,000 transcriptomes from leukocytes and PBMCs demonstrated that these immune types are highly conserved across different patient groups and diseases.
[0069] In some embodiments, the methods described herein include determining a leukocyte immune profile type (selected from among naive, primed, advanced, chronic, or suppressed immune profile types) of a subject with head and neck squamous cell carcinoma (HNSCC) and determining a treatment strategy based on the leukocyte (e.g., PBMC) immune profile type. As further described in this example, data indicate that HNSCC patients identified as having the primed immune profile type are more likely to respond to immunotherapy than patients with other immune profile types. The primed type is characterized by a higher percentage of differentiated CD4+ central and transitional memory T cells and CD39+ regulatory T cells (Tregs) compared to other leukocyte immune profile types.
[0070] Thus, in some aspects, the present disclosure provides a method for determining a leukocyte immune profile type of a subject (e.g., healthy or with cancer), comprising using at least one computer hardware processor to obtain cytometry data or RNA expression data from the subject's blood (e.g., obtaining cytometry data or RNA expression data for one or more whole blood samples comprising WBCs or PBMCs obtained from the subject); determining cellular composition for at least some cell types (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34 cell types) of a plurality of cell types listed in any one of Tables 1-3; and generating a leukocyte signature for the subject using the cytometry data or RNA expression data, the leukocyte signature including a cellular composition percentage for each cell type (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34 cell types) in at least a portion of the plurality of cell types; and using the leukocyte signature to identify a leukocyte immune profile type for the subject from among the plurality of leukocyte immune profile types. In some embodiments, the method includes obtaining WBC or PMBC flow cytometry data from the subject's whole blood. Compared to CYTOF, flow cytometry has several advantages, including, but not limited to, lower cost, widespread availability, and easier control of the measurement signal. In some embodiments, the method includes obtaining WBC or PMBC RNA-seq data from the subject's whole blood.
[0071] In some embodiments, WBC cytometry data is processed and the cellular composition percentages include the cellular composition percentages for 15 or more (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34) of the, or each, cell types listed in Table 1. In some embodiments, PMBC cytometry data is processed and the cellular composition percentages include the cellular composition percentages for 15 or more (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34) of the, or each, cell types listed in Table 2. In some embodiments, the RNA-seq data is processed and the cellular composition percentages include 15 or more (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34) of the cell types listed in Table 3, or the cellular composition percentages for each.
[0072] In some embodiments, the method comprises processing data (e.g., cytometry data or RNA expression data) to identify, for flow cytometry data, naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-switched memory IgM B cells, V52+ γδ T cells, post-class-switched memory B cells, central memory CD8+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD39+ CD4+ Tregs, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, effector memory CD4+ T cells, NKT cells, CD8+ TEMRA, effector memory CD8+ T cells, CD4+ TEMRA, neutrophils, granulocytes, classical monocytes, non-classical monocytes, HLA-DR low monocytes; and RNA-seq cell deconvolution data: naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-switched memory IgM B cells, class-switched memory B cells, central memory CD4+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD4+ T cells, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, CD8+ TEMRA, effector memory CD8+ Determining the cellular composition percentages for multiple cell types selected from T cells, neutrophils, granulocytes, classical monocytes, and non-classical monocytes.
[0073] In some embodiments, the plurality of leukocyte immune profile types are clusters identified by clustering a plurality of leukocyte signatures associated with each subject in a subject cohort. The cohort may include subjects diagnosed with cancer. The cohort may include healthy subjects. The cohort may include subjects diagnosed with cancer and with a known prognosis and / or known likelihood of responding to a particular treatment, such as immunotherapy.
[0074] The following sections provide a more detailed description of various concepts related to the cell-typing system and method developed by the inventors, and embodiments thereof. It should be understood that the various aspects described herein may be implemented in any of numerous ways. Examples of specific embodiments are provided herein for illustrative purposes only. In addition, the various aspects described in the following embodiments may be used alone or in any combination, and are not limited to the combinations expressly described herein.
[0075] 1 depicts an exemplary method 100 for determining a subject's leukocyte (e.g., PBMC) immune profile type. In some embodiments, the subject is a healthy subject. In some embodiments, the subject is a subject who has, is suspected of having, or is at risk of having cancer. In some embodiments, the subject may include any of the embodiments described herein, including those related to the "Subject" section.
[0076] In some embodiments, the exemplary method 100 may be performed in a clinical or laboratory setting. For example, the exemplary method 100 may be implemented on a computing device located within the clinical or laboratory setting. In some embodiments, the computing device may obtain the cytometry data directly from a cytometry platform located within the clinical or laboratory setting. For example, a computing device included within the cytometry platform may obtain the cytometry data directly from the cytometry platform. In some embodiments, the computing device may obtain the cytometry data indirectly from a cytometry platform located within or outside the clinical or laboratory setting. For example, a computing device located within a clinical or laboratory setting may obtain the cytometry data via a communications network, such as the Internet or any other suitable network, as aspects of the technology described herein are not limited to any particular communications network.
[0077] Additionally or alternatively, exemplary method 100 may be performed in a setting remote from a clinical or laboratory setting. For example, exemplary method 100 may be implemented on a computing device located external to the clinical or laboratory setting. In this case, the computing device may indirectly obtain cytometry data generated using a cytometry platform located within or external to the clinical or laboratory setting. For example, the cytometry data may be provided to the computing device via a communications network, such as the Internet or any other suitable network, as aspects of the technology described herein are not limited to any particular communications network. In some embodiments, the cytometry data may be obtained from a database or data store, or may be data previously obtained and stored from the cytometry platform (possibly partially processed after receipt from the cytometry platform). In some embodiments, obtaining flow cytometry data includes obtaining data from subjects with multiple different cancers (e.g., pancreatic cancer, breast cancer, non-small cell lung cancer, colorectal cancer, melanoma, prostate cancer, etc.). Thus, in some embodiments, the methods described herein have pan-cancer applicability.
[0078] As described herein, in some embodiments, method 100 begins with process 102, in which cytometry is performed on a biological sample containing PBMCs (or WBCs, in the case of determining a WBC immune profile type) obtained from a subject. In some embodiments, process 102 involves processing the biological sample using a cytometry platform, where the platform generates cytometry data. The biological sample processed in process 102 may be obtained from a subject having, suspected of having, or at risk of having cancer or any immune-related disease. The biological sample processed in process 102 may be obtained from a healthy subject. The biological sample may be obtained by performing a biopsy or by obtaining a blood sample, saliva sample, or any other suitable biological sample from a subject. The biological sample may include diseased tissue (e.g., cancerous) and / or healthy tissue. In some embodiments, the source or preparation method of the biological sample may include any of the embodiments described herein, including those related to the "Biological Sample" section.
[0079] In some embodiments, a cytometry platform includes any suitable instrument and / or system configured to perform cytometry, such that aspects of the technology described herein are not limited to any particular type of cytometry system. For example, a cytometry platform may include any suitable flow cytometry platform. Additionally or alternatively, a cytometry platform may include any suitable mass cytometry platform. In some embodiments, a biological sample may be prepared according to a manufacturer's protocol associated with the cytometry platform. In some embodiments, a biological sample may be prepared according to any suitable protocol, such that embodiments of the technology described herein are not limited to any particular preparation protocol. In some embodiments, a flow cytometry technique may include any of the embodiments described herein, including those described in the "Flow Cytometry" section. In some embodiments, a mass cytometry technique may include any of the embodiments described herein, including those described in the "Mass Cytometry" section.
[0080] Those skilled in the art will recognize that in some embodiments, operation 102 is optional and not necessary to perform method 100. For example, in some instances, cytometry has already been performed on the biological sample and cytometry data exists before method 100 begins.
[0081] Regardless of whether operation 102 is performed, method 100 either proceeds to or begins with operation 104, in which a leukocyte immune profile type for the subject is determined. Operation 104 involves operations 106, 108, 110, and 112, and method 100 proceeds through these operations sequentially, beginning with operation 106, in which cytometry data for the subject is obtained. The cytometry data typically includes information about a plurality of cells, e.g., information about a population of immune cell types (e.g., PBMCs) for the subject. In some embodiments, the cytometry data includes information about the presence, absence, and / or relative abundance of at least some (or all) of the cells in the plurality of cells, e.g., some or all of the cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data includes flow cytometry data. In some embodiments, the cytometry data includes cytometry by time-of-flight (CyTOF) data.
[0082] In some embodiments, the cytometry data includes information regarding the presence, absence, and / or relative abundance of 15 to 36 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data includes information regarding the presence, absence, and / or relative abundance of 3 to 8, 5 to 12, 10 to 20, 15 to 25, 18 to 34, or 18 to 36 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data comprises information regarding the presence, absence, and / or relative abundance of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, or 36 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data comprises information regarding the presence, absence, and / or relative abundance of additional cell types not listed in Table 1 or Table 2.
[0083] Method 100 then proceeds to operation 108, which processes the cytometry data to obtain cellular composition percentages. In some embodiments, the cytometry data is processed to obtain cellular composition percentages for at least some of the cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data is processed to obtain cellular composition percentages for 15-36 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data is processed to obtain cellular composition percentages for 2-34 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data is processed to obtain cellular composition percentages for 3-8, 5-12, 10-20, 15-25, 18-34, or 18-36 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data is processed to obtain cellular composition percentages for at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, or 36 cell types listed in Table 1 or Table 2. In some embodiments, the cytometry data is processed to obtain cellular composition percentages for additional cell types not listed in Table 1 or Table 2. Methods for obtaining cellular composition percentages by processing cytometry data are further described herein in connection with FIG.
[0084] In some embodiments, processing the cytometry data includes applying one or more machine learning models to the cytometry data to obtain cellular composition percentages for at least some (or all) of the plurality of cell types listed in Table 1 or Table 2. Examples of machine learning models that may be used to process the cytometry data to obtain cellular composition percentages are described, for example, in International Publication No. WO 2023 / 147177, filed January 31, 2023, the entire contents of which are incorporated herein by reference. In some embodiments, the machine learning model includes the Cibersort method (e.g., as described by Newman et al. Nature Methods volume 12, pages 453-457 (2015)) or the CibersortX method (e.g., as described by Newman et al. Nature Biotechnology volume 37, pages 773-782 (2019)). Aspects of machine learning models are described herein, including at least those in the section "Cytometry-Based Cell Deconvolution."
[0085] After obtaining the cellular composition percentages from the cytometry data in operation 108, method 104 proceeds to operation 110, which generates a leukocyte signature using the cytometry data. In some embodiments, the leukocyte signature includes the cellular composition percentages for at least some of the cell types listed in Table 1 or Table 2. In some embodiments, the leukocyte signature includes the cellular composition percentages for 15 to 36 cell types listed in Table 1 or Table 2. In some embodiments, the leukocyte signature includes the cellular composition percentages for 3 to 8, 5 to 12, 10 to 20, 15 to 25, 18 to 34, or 18 to 36 cell types listed in Table 1 or Table 2. In some embodiments, the leukocyte signature comprises cellular composition percentages for at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, or 36 cell types listed in Table 1 or Table 2. In some embodiments, the leukocyte signature comprises cellular composition percentages for additional cell types not listed in Table 1 or Table 2. In some embodiments, the leukocyte signature is output as a vector comprising cellular composition percentages.
[0086] In some embodiments, the cytometry data is processed using a computing device. In some embodiments, the computing device may be one or more computing devices of any suitable type. For example, the computing device may be a portable computing device (e.g., laptop, smartphone) or a stationary computing device (e.g., desktop computer, server). When the computing device includes multiple computing devices, the one or more devices may be physically co-located (e.g., in a single room) or distributed across multiple physical locations. In some embodiments, the computing device may be part of a cloud computing infrastructure. In some embodiments, one or more computers may be co-located within a facility operated by an entity (e.g., a hospital, a research institute). In some embodiments, one or more computing devices may be physically co-located with a medical device, such as a cytometry platform. For example, the cytometry platform may include a computing device.
[0087] In some embodiments, the computing device may be operated by a user, such as a physician, clinician, researcher, patient, or other individual. For example, the user may provide cytometry data as input to the computing device (e.g., by uploading a file) and / or may provide user input specifying a process or other method to be performed using the cytometry data.
[0088] In some embodiments, the computing device includes software configured to perform various functions with respect to the cytometry data. Examples of computing devices including such software are described herein, including at least those associated with FIG. 15.
[0089] Method 100 then proceeds to operation 112, where a leukocyte immune profile type for the subject is identified using the leukocyte signature generated in operation 110. This may be done in any suitable manner. For example, in some embodiments, each of the possible leukocyte immune profile types is associated with (e.g., defined or characterized by) a respective plurality of leukocyte immune profile types. In such embodiments, the leukocyte immune profile type for the subject may be identified by associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types (e.g., the identified type may be a type associated with (e.g., defined or characterized by) a leukocyte signature cluster to which the subject's leukocyte signature is closest according to a distance measure or any suitable measure of distance or similarity); and identifying the leukocyte immune profile type for the subject as the leukocyte immune profile type that corresponds to the particular one of the plurality of leukocyte immune profile types with which the subject's leukocyte signature is associated. Examples of leukocyte immune profile types are described herein. Aspects of identifying a subject's leukocyte immune profile type are described herein, including in the sections entitled "Generation of Leukocyte Signatures and Identification of Leukocyte Immune Profile Types" and "Techniques for Correlating Leukocyte Signatures with Leukocyte Immune Profile Types," and in FIG. 4.
[0090] As described above, the subject's leukocyte immune profile type is identified in operation 112. In some embodiments, the subject's leukocyte immune profile type is identified as one of the following leukocyte immune profile types: naive, primed, advanced, chronic, or suppressed. In some embodiments, method 104 ends upon completion of operation 112.
[0091] In some embodiments, method 100 proceeds to operation 114, where the leukocyte immune profile type identified in operation 112 is used to identify the subject's likelihood of responding to a therapy. In some embodiments, when operation 112 identifies the subject as having a naive leukocyte immune profile type, operation 114 identifies the subject as having an increased likelihood of responding to an immunotherapy (e.g., an immune checkpoint inhibitor such as a PD1 antibody, such as pembrolizumab) relative to subjects having other leukocyte immune profile types. In some embodiments, when operation 112 identifies the subject as having a primed leukocyte immune profile type, operation 114 identifies the subject as having an increased likelihood of responding to an immunotherapy (e.g., an immune checkpoint inhibitor such as a PD1 antibody, such as pembrolizumab) relative to subjects having other leukocyte immune profile types. In some embodiments, when the subject is identified as having a suppressed leukocyte immune profile type in process 112, the subject is identified as having a reduced likelihood of responding to immunotherapy (e.g., an immune checkpoint inhibitor such as a PD1 antibody, such as pembrolizumab) compared to subjects with other leukocyte immune profile types in process 114, and a non-immunotherapy therapy may be identified for the subject. Aspects of identifying whether a subject is likely to respond to a therapy are described herein, including in the section below entitled "Treatment Indications."
[0092] In some embodiments, method 100 is complete after completion of operation 114. In some such embodiments, the determined leukocyte signature and / or the identified leukocyte immune profile type, and / or the identified likelihood that a therapy will be effective in a subject may be stored for subsequent use, provided to one or more recipients (e.g., clinicians, researchers, etc.), and / or used to update the leukocyte immune profile type.
[0093] However, in some embodiments, one or more other operations are performed after operation 114. For example, in the illustrated embodiment of Figure 1, method 100 may include optional operation 116, which is shown with a dashed line in Figure 1. For example, in operation 116, the subject is administered one or more therapeutic agents (e.g., immunotherapy, such as an immune checkpoint inhibitor).
[0094] Examples of immunotherapy and other therapies are provided herein.
[0095] Although operations 102, 114, and 116 are indicated as optional in the example of FIG. 1, it should be understood that in other embodiments, one or more other operations may be optional (in addition to or instead of operations 102, 114, and 116).
[0096] In some aspects, the present disclosure provides a method for determining a leukocyte immune profile type of a subject, comprising: using at least one computer hardware processor to obtain RNA expression data for WBCs or PBMCs of the subject; processing the RNA expression data using a cellular deconvolution technique to obtain cellular composition percentages for at least some cell types (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28) of a plurality of cell types listed in Table 3. determining a leukocyte immune profile type for a subject using RNA expression data, the leukocyte signature comprising a cellular composition percentage for each cell type (e.g., at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28) in at least a portion of the plurality of cell types; and using the leukocyte signature to identify a leukocyte immune profile type for the subject from among the plurality of leukocyte immune profile types. In some embodiments, the method comprises determining a leukocyte immune profile type for a subject having, suspected of having, or at risk of having cancer.
[0097] 2 depicts an exemplary method 200 for determining a peripheral blood leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer. In some embodiments, the subject may include any of the embodiments described herein, including those related to the "Subject" section.
[0098] In some embodiments, exemplary method 200 may be performed in a clinical or laboratory setting. For example, exemplary method 200 may be implemented on a computing device located within the clinical or laboratory setting. In some embodiments, the computing device may obtain the RNA expression data directly from a sequencing platform (e.g., a nucleic acid sequencing platform) located within the clinical or laboratory setting. For example, a sequencing platform included within the computing device may obtain the RNA expression data directly from the sequencing platform. In some embodiments, the computing device may obtain the RNA expression data indirectly from a sequencing platform located within or outside the clinical or laboratory setting. For example, a computing device located within a clinical or laboratory setting may obtain the RNA expression data via a communications network, such as the Internet or any other suitable network, as aspects of the technology described herein are not limited to any particular communications network.
[0099] Additionally or alternatively, exemplary method 200 may be performed in a setting remote from a clinical or laboratory setting. For example, exemplary method 200 may be implemented on a computing device located outside the clinical or laboratory setting. In this case, the computing device may indirectly obtain RNA expression data generated using a sequencing platform located within or outside the clinical or laboratory setting. For example, the RNA expression data may be provided to the computing device via a communications network, such as the Internet or any other suitable network, as aspects of the technology described herein are not limited to any particular communications network.
[0100] As described herein, in some embodiments, method 200 begins with process 202, in which RNA sequencing is performed on a biological sample obtained from a subject. In some embodiments, process 202 involves processing the biological sample using a sequencing platform to generate sequencing data. In some embodiments, the RNA sequencing data is processed to obtain RNA expression data. The biological sample processed in process 202 may be obtained from a subject who has, is suspected of having, or is at risk of having cancer or any immune-related disease. The biological sample may be obtained by performing a biopsy or by obtaining a blood sample, saliva sample, or any other suitable biological sample from a subject. The biological sample may include diseased tissue (e.g., cancerous) and / or healthy tissue. In some embodiments, the source or preparation of the biological sample may include any of the embodiments described herein, including those related to the "Biological Sample" section.
[0101] In some embodiments, the sequencing platform includes any suitable instrument and / or system configured to perform nucleic acid sequencing (e.g., RNA sequencing), such that aspects of the technology described herein are not limited to any particular type of sequencing system. In some embodiments, the biological sample may be prepared according to a manufacturer's protocol associated with the sequencing platform. In some embodiments, the biological sample may be prepared according to any suitable protocol, such that embodiments of the technology described herein are not limited to any particular preparation protocol. As one illustrative example, in some embodiments, the sequencing data may include bulk sequencing data. The bulk sequencing data may include at least 1 million reads, at least 5 million reads, at least 10 million reads, at least 20 million reads, at least 50 million reads, or at least 100 million reads. In some embodiments, the sequencing data includes bulk RNA sequencing (RNA-seq) data, single-cell RNA sequencing (scRNA-seq) data, or next-generation sequencing (NGS) data. In some embodiments, the RNA sequencing techniques may include any of the embodiments described herein, including those related to the "RNA Expression Data" section.
[0102] Those skilled in the art will recognize that in some embodiments, process 202 is optional and not necessary to perform method 200. For example, in some instances, a biological sample has already been subjected to RNA sequencing and processed to generate RNA expression data, which is present prior to the start of method 200.
[0103] Regardless of whether 202 is performed, method 200 either proceeds to or begins with operation 204, in which a leukocyte immune profile type is determined for the subject. Operation 204 involves operations 206, 208, 210, and 212, through which method 200 proceeds sequentially, beginning with operation 206, in which RNA expression data for the subject is obtained. The RNA expression data, in some embodiments, includes RNA expression levels for genes expressed by a plurality of cells of the subject, e.g., a plurality of immune cell types (e.g., PBMCs). In some embodiments, the RNA expression data includes information (e.g., RNA expression levels) related to the presence, absence, and / or relative abundance of at least some (or all) of the cells of the plurality of cells, e.g., some or all of the cell types listed in Table 3.
[0104] In some embodiments, the RNA expression data comprises RNA expression levels of genes associated with (e.g., defined or characterized by) between 2 and 28 cell types listed in Table 3. In some embodiments, the RNA expression data comprises RNA expression levels of genes associated with between 3 and 8, between 5 and 12, between 10 and 20, between 15 and 25, between 18 and 26, or between 18 and 28 cell types listed in Table 3. In some embodiments, the RNA expression data comprises RNA expression levels of genes associated with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, or at least 28 cell types listed in Table 3. In some embodiments, the RNA expression data includes RNA expression levels of genes associated with (e.g., defined or characterized by) additional cell types not listed in Table 3.
[0105] Method 200 then proceeds to operation 208, where the RNA expression data is processed to obtain cellular composition percentages. In some embodiments, the RNA expression data is processed to obtain cellular composition percentages for at least some of the cell types listed in Table 3. In some embodiments, the RNA expression data is processed to obtain cellular composition percentages for 20-28 cell types listed in Table 3. In some embodiments, the RNA expression data is processed to obtain cellular composition percentages for 3-8, 5-12, 10-20, 15-25, 18-26, or 18-28 cell types listed in Table 3. In some embodiments, the RNA expression data is processed to obtain cellular composition percentages for at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, or at least 28 cell types listed in Table 3. In some embodiments, the RNA expression data is processed to obtain cellular composition percentages for additional cell types not listed in Table 3.
[0106] In some embodiments, processing 208 includes processing the RNA expression levels with a cellular deconvolution technique to determine cellular composition percentages for at least some (or all) of the plurality of cell types listed in Table 3. In some embodiments, processing the RNA expression data includes applying one or more machine learning models to the RNA expression data to obtain cellular composition percentages for at least some (or all) of the plurality of cell types listed in Table 3. Examples of machine learning models that may be used to process the RNA expression data to obtain cellular composition percentages are described, for example, in PCT / US2021 / 022155, published on September 16, 2021 as WO 2021 / 183917; and PCT / US2022 / 027088, published on November 3, 2022 as WO 2022 / 232615 (the entire contents of each of which are incorporated herein by reference). Aspects of the machine learning model are described herein, including at least those in the section "RNA-Based Cellular Deconvolution."
[0107] After obtaining the cellular composition percentages from the RNA expression data in operation 208, method 204 proceeds to operation 210, which generates a leukocyte signature using the RNA expression data. In some embodiments, the leukocyte signature includes the cellular composition percentages for at least a portion (e.g., at least 20, 21, 22, 23, 24, 25, 26, 27, or 28) of the cell types listed in Table 3. In some embodiments, the leukocyte signature includes the cellular composition percentages for 2 to 28 cell types listed in Table 3. In some embodiments, the leukocyte signature includes the cellular composition percentages for 3 to 8, 5 to 12, 10 to 20, 15 to 25, 16 to 26, or 18 to 28 cell types listed in Table 3. In some embodiments, the leukocyte signature comprises cellular composition percentages for at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, or at least 28 cell types listed in Table 3. In some embodiments, the leukocyte signature comprises cellular composition percentages for additional cell types not listed in Table 3. In some embodiments, the leukocyte signature is output as a vector comprising cellular composition percentages.
[0108] In some embodiments, the RNA expression data is processed using a computing device. In some embodiments, the computing device may be one or more computing devices of any suitable type. For example, the computing device may be a portable computing device (e.g., laptop, smartphone) or a stationary computing device (e.g., desktop computer, server). When the computing device includes multiple computing devices, the one or more devices may be physically co-located (e.g., in one room) or distributed across multiple physical locations. In some embodiments, the computing device may be part of a cloud computing infrastructure. In some embodiments, one or more computers may be co-located within a facility operated by an entity (e.g., hospital, research institute). In some embodiments, one or more computing devices may be physically co-located with a medical device, such as a sequencing platform. For example, a sequencing platform may include a computing device.
[0109] In some embodiments, the computing device may be operated by a user, such as a physician, clinician, researcher, patient, or other individual. For example, a user may provide RNA expression data as input to the computing device (e.g., by uploading a file) and / or may provide user input specifying processing or other methods to be performed using the RNA expression data.
[0110] In some embodiments, the computing device includes software configured to perform various functions on the RNA expression data. Examples of computing devices including such software are described herein, including in connection with at least FIG. 15.
[0111] Method 200 then proceeds to operation 212, where a leukocyte immune profile type for the subject is identified using the leukocyte signature generated in operation 210. This may be done in any suitable manner. For example, in some embodiments, each of the possible leukocyte immune profile types is associated with (e.g., defined or characterized by) a respective plurality of leukocyte immune profile types. In such embodiments, the leukocyte immune profile type for the subject may be identified by associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types (e.g., the identified type may be the type associated with the leukocyte cluster to which the subject's leukocyte signature is closest according to a distance measure or any suitable measure of distance or similarity); and identifying the leukocyte immune profile type for the subject as the leukocyte immune profile type corresponding to the particular one of the plurality of leukocyte immune profile types with which the subject's leukocyte signature is associated. Examples of leukocyte immune profile types are described herein. Aspects of identifying a subject's leukocyte immune profile type are described herein, including in the sections entitled "Generation of Leukocyte Signatures and Identification of Leukocyte Immune Profile Types" and "Techniques for Correlating Leukocyte Signatures with Leukocyte Immune Profile Types," and in FIG. 4.
[0112] As described above, the subject's leukocyte immune profile type is identified in operation 212. In some embodiments, the subject's leukocyte immune profile type is identified as one of the following leukocyte immune profile types: naive, primed, advanced, chronic, or suppressed. In some embodiments, method 204 ends upon completion of operation 212.
[0113] In some embodiments, method 200 proceeds to operation 214, in which the leukocyte immune profile type identified in operation 212 is used to identify the subject's likelihood of responding to a therapy. In some embodiments, when operation 212 identifies the subject as having a naive leukocyte immune profile type, the subject is identified in operation 214 as having an increased likelihood of responding to an immunotherapy (e.g., an immune checkpoint inhibitor such as a PD1 antibody, such as pembrolizumab) relative to subjects having other leukocyte immune profile types. In some embodiments, when operation 212 identifies the subject as having a primed leukocyte immune profile type, the subject is identified in operation 214 as having an increased likelihood of responding to an immunotherapy (e.g., an immune checkpoint inhibitor such as a PD1 antibody, such as pembrolizumab) relative to subjects having other leukocyte immune profile types. In some embodiments, when the subject is identified as having a suppressed leukocyte immune profile type in process 212, the subject is identified as having a reduced likelihood of responding to immunotherapy (e.g., an immune checkpoint inhibitor such as a PD1 antibody, such as pembrolizumab) compared to subjects with other leukocyte immune profile types in process 214, and a non-immunotherapy therapy may be identified for the subject. Aspects of identifying whether a subject is likely to respond to a therapy are described herein, including in the section below entitled "Treatment Indications."
[0114] In some embodiments, method 200 is complete after completion of operation 214. In some such embodiments, the determined WBC or PBMC signature and / or the identified leukocyte immune profile type, and / or the identified likelihood that a therapy will be effective in a subject may be stored for subsequent use, provided to one or more recipients (e.g., clinicians, researchers, etc.), and / or used to update the leukocyte immune profile type.
[0115] However, in some embodiments, one or more other operations are performed after operation 214. For example, in the illustrated embodiment of Figure 2, method 200 may include optional operation 216, which is shown using a dashed line in Figure 2. For example, in operation 216, the subject is administered one or more therapeutic agents (e.g., immunotherapies such as immune checkpoint inhibitors). Examples of immunotherapies and other therapies are provided herein.
[0116] Although operations 202, 214, and 216 are indicated as optional in the example of FIG. 2, it should be understood that in other embodiments, one or more other operations may be optional (in addition to or instead of operations 202, 214, and 216).
[0117] [Table 1-1]
[0118] [Table 1-2]
[0119] [Table 2-1]
[0120] [Table 2-2]
[0121] [Table 3-1]
[0122] [Table 3-2]
[0123] [Table 4]
[0124] subject Aspects of the present disclosure relate to biological samples obtained from a subject. In some embodiments, the subject is a mammal (e.g., a human, mouse, cat, dog, horse, hamster, cow, pig, or other animal). The subject may be a human. The subject may be an adult human (e.g., 18 years of age or older) or a child (e.g., under 18 years of age). The human may be diagnosed with or have been diagnosed with at least one form of cancer.
[0125] In some embodiments, the cancer afflicting the subject is a carcinoma, sarcoma, myeloma, leukemia, lymphoma, or a mixed type cancer comprising two or more of carcinoma, sarcoma, myeloma, leukemia, and lymphoma. Carcinoma refers to a malignant neoplasm of epithelial origin or a cancer of the lining or outer membranes of the body. Sarcoma refers to a cancer that arises in supportive and connective tissues, such as bone, tendon, cartilage, muscle, and fat. Myeloma is a cancer that arises in the plasma cells of the bone marrow. Leukemia ("liquid cancer" or "blood cancer") is a cancer of the bone marrow (the site of blood cell production). Lymphoma develops in the glands or nodes, vascular networks, nodes, and organs of the lymphatic system (particularly the spleen, tonsils, and thymus) that produce white blood cells, i.e., lymphocytes, that purify body fluids and fight infection. Non-limiting examples of mixed type cancers include adenosquamous carcinoma, mixed mesodermal tumor, carcinosarcoma, and teratocarcinoma. In some embodiments, the subject has a tumor. The tumor can be benign or malignant. In some embodiments, the cancer is any one of the following: skin cancer, lung cancer, breast cancer, prostate cancer, colon cancer, rectal cancer, cervical cancer, and uterine cancer. In some embodiments, the cancer is any one of the following: sarcoma, breast cancer, colorectal cancer, pancreatic cancer, non-small cell lung cancer (NSCLC), melanoma, or prostate cancer. In some embodiments, the cancer is head and neck squamous cell carcinoma (HNSCC).
[0126] In some embodiments, a subject is at risk of developing cancer because, for example, the subject has one or more genetic risk factors or has been or is being exposed to one or more carcinogens (e.g., tobacco smoke or chewing tobacco).
[0127] biological samples Any of the methods, systems, or other claimed elements may use or be used in the analysis of a biological sample from a subject. In some embodiments, the biological sample is obtained from a subject who has, is suspected of having, or is at risk of having cancer. In some embodiments, the biological sample comprises a bodily fluid (e.g., blood, urine, or cerebrospinal fluid) and / or a tumor.
[0128] A tumor sample, in some embodiments, refers to a sample containing cells from a tumor. In some embodiments, a tumor sample contains cells from a benign tumor, e.g., non-cancerous cells. In some embodiments, a tumor sample contains cells from a pre-malignant tumor, e.g., pre-cancerous cells. In some embodiments, a tumor sample contains cells from a malignant tumor, e.g., cancerous cells.
[0129] Examples of tumors include, but are not limited to, adenoma, fibroma, hemangioma, lipoma, cervical dysplasia, pulmonary metaplasia, leukoplakia, carcinoma, sarcoma, germ cell tumor, and blastoma.
[0130] A blood sample, in some embodiments, refers to a sample comprising cells, e.g., cells from a blood sample. In some embodiments, the blood sample comprises non-cancerous cells. In some embodiments, the blood sample comprises pre-cancerous cells. In some embodiments, the blood sample comprises cancerous cells. In some embodiments, the blood sample comprises blood cells. In some embodiments, the blood sample comprises red blood cells. In some embodiments, the blood sample comprises or consists of white blood cells, "WBCs," or peripheral blood mononuclear cells, "PBMCs." In some embodiments, the blood sample comprises platelets. Examples of cancerous blood cells include, but are not limited to, leukemia, lymphoma, and myeloma. In some embodiments, a blood sample is taken to obtain cell-free nucleic acids (e.g., cell-free DNA, cell-free RNA, etc.) in the blood.
[0131] The blood sample may be a whole blood sample or a fractionated blood sample. In some embodiments, the blood sample comprises whole blood. In some embodiments, the whole blood sample comprises an anticoagulant. In some embodiments, the blood sample comprises fractionated blood. In some embodiments, the blood sample comprises buffy coat. In some embodiments, the blood sample comprises serum. In some embodiments, the blood sample comprises plasma. In some embodiments, the blood sample comprises a blood clot.
[0132] A tissue sample, in some embodiments, refers to a sample containing cells from the tissue. In some embodiments, a tumor sample contains non-cancerous cells from the tissue. In some embodiments, a tumor sample contains pre-cancerous cells from the tissue.
[0133] Any biological sample described herein may be obtained from a subject using any known technique. See, for example, the following publications regarding collection, processing, and storage of biological samples, each of which is incorporated herein by reference in its entirety: Biospecimens and biorepositories: from afterthought to science by Vaught et al. (Cancer Epidemiol Biomarkers Prev. 2012 Feb; 21(2):253-5), and Biological sample collection, processing, storage, and information management by Vaught and Henderson (IARC Sci Publ. 2011; (163):23-42).
[0134] Any biological sample from a subject described herein may be stored using any method that maintains the stability of the biological sample. In some embodiments, maintaining the stability of a biological sample means preventing the degradation of the components of the biological sample (e.g., DNA, RNA, protein, or tissue structure or morphology) until they are measured, so that the measurements correspond to the state of the sample at the time of obtaining it from the subject. In some embodiments, the biological sample is stored in a composition that is permeable to the biological sample and can protect the components of the biological sample (e.g., DNA, RNA, protein, or tissue structure or morphology) from degradation. As used herein, degradation refers to the change of a component from one form to another, such that the original form is no longer detectable at the same level as before degradation.
[0135] In some embodiments, biological samples are preserved using a frozen storage method. Non-limiting examples of frozen storage methods include, but are not limited to, step-down freezing, blast freezing, direct plunge freezing, snap freezing, slow freezing using a programmable freezer, and vitrification. In some embodiments, biological samples are preserved using freeze-drying. In some embodiments, after collection of the biological sample from a subject, the biological sample is placed in a container that already contains a protectant (e.g., RNALater to protect RNA) and then frozen (e.g., by snap freezing). In some embodiments, such storage in a frozen state occurs immediately after collection of the biological sample. In some embodiments, the biological sample may be kept in a protectant or a buffer solution without a protectant at either room temperature or 4°C for a period of time (e.g., up to 1 hour, up to 8 hours, or up to 1 day, or several days) before freezing.
[0136] Non-limiting examples of protectants include formalin solution, formaldehyde solution, RNALater or other equivalent solution, TriZol or other equivalent solution, DNA / RNA Shield or equivalent solution, EDTA (e.g., Buffer AE (10 mM Tris·Cl; 0.5 mM EDTA, pH 9.0)) and other coagulants, and citrate dextrose (e.g., for blood specimens). In some embodiments, specialized containers may be used to collect and / or store biological samples. For example, vacutainers may be used to store blood. In some embodiments, vacutainers may contain protectants (e.g., coagulants or anticoagulants). In some embodiments, the container in which the biological sample is stored may be placed in a secondary container for better storage or to prevent contamination.
[0137] Any biological sample from a subject described herein may be stored under any conditions that maintain the stability of the biological sample. In some embodiments, the biological sample is stored at a temperature that maintains the stability of the biological sample. In some embodiments, the sample is stored at room temperature (e.g., 25°C). In some embodiments, the sample is stored under refrigeration (e.g., 4°C). In some embodiments, the sample is stored under frozen conditions (e.g., -20°C). In some embodiments, the sample is stored under ultra-low temperature conditions (e.g., -50°C to -800°C). In some embodiments, the sample is stored under liquid nitrogen (e.g., -1700°C). In some embodiments, biological samples are stored at -60°C to -80°C (e.g., -70°C) for up to 5 years (e.g., up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, or up to 5 years). In some embodiments, biological samples are stored for up to 20 years (e.g., up to 5 years, up to 10 years, up to 15 years, or up to 20 years) as described by any of the methods described herein.
[0138] The methods of the present disclosure involve obtaining one or more biological samples from a subject for analysis. In some embodiments, one biological sample is collected from a subject for analysis. In some embodiments, two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) biological samples are collected from a subject for analysis. In some embodiments, one biological sample from a subject will be analyzed. In some embodiments, two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) biological samples may be analyzed. If more than one biological sample from a subject is analyzed, the biological samples may be obtained at the same time (e.g., two or more biological samples may be taken during the same procedure), or the biological samples may be taken at different time points (e.g., between different procedures, including at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 days after the initial procedure; 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 weeks; 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 months, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 years, or 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 years).
[0139] The second or subsequent biological sample may be taken or obtained from the subject after one or more treatments, and may be taken from the same or a different area. As a non-limiting example, the second or subsequent biological sample may be useful in determining whether the cancer in each biological sample has different characteristics (e.g., when biological samples are taken from two physically separate tumors in a patient) or whether the cancer has responded to one or more treatments (e.g., in the case of two or more biological samples taken before and after treatment from the same tumor or different tumors). In some embodiments, each of the at least one biological sample is a bodily fluid sample, a cell sample, or a tissue biopsy sample.
[0140] Flow cytometry Aspects of the present disclosure relate to processing cytometry data to generate cellular composition percentages. In some embodiments, the cytometry data may include flow cytometry data. In some embodiments, a flow cytometry platform may be used to perform a flow cytometry interrogation of a flowing (e.g., blood) sample. The flow sample may include target particles having specific particle attributes. The flow cytometry interrogation of the flowing sample may provide a flow cytometry result for the flowing sample.
[0141] In some embodiments, the flow sample may be exposed to a stain or dye that, when exposed to interrogating excitation radiation, provides a response radiation that can be measured by the radiation detection system of the flow cytometry platform. In some embodiments, the flow cytometry platform includes multiple photodetectors. As a particle passes through the laser beam, time-correlated pulses will occur on the forward scatter (FSC) and side scatter (SSC) detectors, and possibly also on the fluorescence detector. This may be considered an "event," and for each event, the magnitude of the detector output for each detector, FSC, SSC, and fluorescence detector, is stored. The data obtained includes measured signals for each of the light scatter parameters and fluorescence emission.
[0142] The flow cytometry platform may further include components for storing detector output and analyzing data. For example, data storage and analysis may be performed using a computer connected to the detection electronics. For example, data may be stored logically in a table format, where each row corresponds to data for one particle (or one event) and columns correspond to each measured parameter. Using a standard file format, such as the "FCS" file format, for storing data from the flow cytometer facilitates analysis of the data using a separate program and / or machine. In some embodiments, data may be displayed in a two-dimensional (2D) plot for ease of visualization, although other methods may be used to visualize multidimensional data.
[0143] In some embodiments, parameters measured using a flow cytometer may include FSC, which refers to excitation light scattered generally along the forward direction by a particle; SSC, which refers to excitation light scattered generally to the side by a particle; and light emitted by fluorescent molecules in one or more channels (frequency bands) of the spectrum, designated FL1, FL2, etc., or by the name of the fluorochrome that is primarily emitted in that channel.
[0144] Both flow cytometers and scanning cytometers are commercially available, for example, from BD Biosciences (San Jose, Calif.). Flow cytometry is described, for example, in Landy et al. (eds.), Clinical Flow Cytometry, Annals of the New York Academy of Sciences Volume 677 (1993); Bauer et al. (eds.), Clinical Flow Cytometry: Principles and Applications, Williams & Wilkins (1993); Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford University Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); and Practical Flow Cytometry, 4th ed., Wiley-Liss (2003) (all of which are incorporated herein by reference). Fluorescence imaging microscopy is described, for example, in Pawley (ed.), Handbook of Biological Confocal Microscopy, 2nd Edition, Plenum Press (1989), incorporated herein by reference.
[0145] In some embodiments, the cytometry data includes cytometry measurements obtained during each cytometry event. As described herein, a cytometry event corresponds to an object (e.g., a cell, debris, bead, doublet, or indeterminate object) under measurement by a cytometry platform (e.g., a flow cytometry platform or a mass cytometry platform). In some embodiments, a cytometry event includes an event subset corresponding to cells in a biological sample under measurement by the cytometry platform. For example, an event subset may include one, some, or all of the cytometry events. The number of cells measured using the cytometry platform may include any suitable number of cells, as aspects of the technology described herein are not limited in this respect. For example, the number of cells measured by the cytometry platform can include at least 5,000 cells, at least 10,000 cells, at least 20,000 cells, at least 50,000 cells, at least 100,000 cells, at least 500,000 cells, at least 600,000 cells, at least 900,000 cells, between 500 cells and 1 million cells, between 5,000 cells and 900,000 cells, or between 20,000 cells and 700,000 cells. In some embodiments, flow cytometry is performed using a panel of antibodies set forth in Table 12.
[0146] Mass cytometry Aspects of the present disclosure relate to processing cytometry data to generate cellular composition percentages. In some embodiments, the cytometry data may include mass cytometry data. In some embodiments, a mass cytometry platform may be used to perform a mass cytometry interrogation of a flowing (e.g., blood) sample. The flowing sample may include target particles having specific particle attributes. The mass cytometry interrogation of the flowing sample may provide a mass cytometry result for the flowing sample.
[0147] In some embodiments, the flow sample may be exposed to a target-specific antibody labeled with a metal isotope. In some embodiments, the conjugated antibody is detected using elemental mass spectrometry (e.g., inductively coupled plasma mass spectrometry (ICP-MS) and time-of-flight mass spectrometry (TOF-MS)). For example, elemental mass spectrometry can distinguish between isotopes of different atomic masses and measure the electrical signal for the isotope associated with each particle or cell. Data obtained for a single cell or particle is considered an "event." In some embodiments, mass cytometry is also referred to as "time-of-flight cytometry" or "CyTOF." CyTOF techniques are described, for example, in Shiskova et al. "Deep immune profiling by mass cytometry revealed an association between the state of the immune system before treatment and response to checkpoint inhibitor therapy in clear cell renal cell carcinoma," Cancer Res (2022) 82 (12_Supplement): 2061.
[0148] The mass cytometry platform may further include components for storing the detector output and analyzing the data. For example, data storage and analysis may be performed using a computer connected to the detection elements. Storage of data from the mass cytometry platform may use a standard file format, such as the "FCS" file format, to facilitate analysis of the data using separate programs and / or machines.
[0149] Mass cytometry platforms are commercially available, for example, from Fluidigm (San Francisco, CA). Mass cytometry is described, for example, in Bendall et al., "A deep profiler's guide to cytometry," Trends in Immunology, 33(7), 323-332 (2012) and Spitzer et al., "Mass Cytometry: Single Cells, Many Features," Cell, 165(4), 780-791 (2016), both of which are incorporated herein by reference in their entireties.
[0150] In some embodiments, the cytometry data includes cytometry measurements obtained during each cytometry event. As described herein, a cytometry event corresponds to an object (e.g., a cell, debris, bead, doublet, or indeterminate object) under measurement by a cytometry platform (e.g., a flow cytometry platform or a mass cytometry platform). In some embodiments, a cytometry event includes an event subset corresponding to cells in a biological sample under measurement by the cytometry platform. For example, an event subset may include one, some, or all of the cytometry events. The number of cells measured using the cytometry platform may include any suitable number of cells, as aspects of the technology described herein are not limited in this respect. For example, the number of cells measured by the cytometry platform can include at least 5,000 cells, at least 10,000 cells, at least 20,000 cells, at least 50,000 cells, at least 100,000 cells, at least 500,000 cells, at least 600,000 cells, at least 900,000 cells, between 500 cells and 1 million cells, between 5,000 cells and 900,000 cells, or between 20,000 cells and 700,000 cells.
[0151] RNA expression data Aspects of the present disclosure relate to methods for determining a subject's leukocyte immune profile type using sequencing data or RNA expression data obtained from a biological sample of the subject. The RNA expression data used in the methods described herein is typically derived from sequencing data obtained from the biological sample.
[0152] Sequencing data can be obtained from biological samples using any suitable sequencing technique and / or device. In some embodiments, the sequencing device used to sequence the biological sample can be selected from any suitable sequencing device known in the art, including, but not limited to, Illumina™, SOLid™, Ion Torrent™, PacBio™, nanopore-based sequencing device, Sanger sequencing device, or 454™ sequencing device. In some embodiments, the sequencing device used to sequence the biological sample is an Illumina sequencing (e.g., NovaSeq™, NextSeq™, HiSeq™, MiSeq™, or MiniSeq™) device.
[0153] After obtaining sequencing data, data is processed to obtain RNA expression data.RNA expression data can be obtained by any method known in the art, including but not limited to whole transcriptome sequencing, whole exome sequencing, whole RNA sequencing, mRNA sequencing, target RNA sequencing, RNA exome capture sequencing, next-generation sequencing and / or deep RNA sequencing.In some embodiments, RNA expression data can be obtained by using microarray assay.
[0154] In some embodiments, sequencing data is processed to generate RNA expression data. In some embodiments, to generate expression data, RNA sequence data is processed by one or more bioinformatics methods or software tools, such as RNA sequence quantification tools (e.g., Kallisto) and genome annotation tools (e.g., Gencode v23). Kallisto software is described in Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, Near-optimal probabilistic RNA-seq quantification, Nature Biotechnology 34, 525-527 (2016), doi: 10.1038 / nbt.3519 (which is incorporated herein by reference in its entirety).
[0155] In some embodiments, the microarray expression data is processed using a bioinformatics R package, such as "affy" or "limma," to generate expression data. The "affy" software is described in Bioinformatics. 2004 Feb 12; 20(3):307-15. Doi: 10.1093 / bioinformatics / btg405. "affy-analysis of Affymetrix GeneChip data at the probe level" by Laurent Gautier 1, Leslie Cope, Benjamin M Bolstad, Rafael A Irizarry PMID: 14960456 DOI: 10.1093 / bioinformatics / btg405, which is incorporated herein by reference in its entirety. The "limma" software is described in Ritchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, Smyth GK "limma powers differential expression analyses for RNA-sequencing and microarray studies." Nucleic Acids Res. 2015 Apr 20; 43(7):e47. 20. Doi.org / 10.1093 / nar / gkv007 PMID: 25605792, PMCID: PMC4402510, which is incorporated herein by reference in its entirety.
[0156] In some embodiments, the sequencing data and / or expression data comprises more than 5 kilobases (kb). In some embodiments, the size of the obtained RNA data is at least 10 kb. In some embodiments, the size of the obtained RNA sequencing data is at least 100 kb. In some embodiments, the size of the obtained RNA sequencing data is at least 500 kb. In some embodiments, the size of the obtained RNA sequencing data is at least 1 megabase (Mb). In some embodiments, the size of the obtained RNA sequencing data is at least 10 Mb. In some embodiments, the size of the obtained RNA sequencing data is at least 100 Mb. In some embodiments, the size of the obtained RNA sequencing data is at least 500 Mb. In some embodiments, the size of the obtained RNA sequencing data is at least 1 gigabase (Gb). In some embodiments, the size of the obtained RNA sequencing data is at least 10 Gb. In some embodiments, the size of the obtained RNA sequencing data is at least 100 Gb. In some embodiments, the size of the RNA sequencing data obtained is at least 500 Gb.
[0157] In some embodiments, expression data is obtained through bulk RNA sequencing. Bulk RNA sequencing can include obtaining the expression level of each gene in RNA extracted from a large input cell population (for example, a mixture of different cell types). In some embodiments, expression data is obtained through single-cell sequencing (for example, scRNA-seq). Single-cell sequencing can include sequencing individual cells.
[0158] In some embodiments, the bulk sequencing data comprises at least 1 million reads, at least 5 million reads, at least 10 million reads, at least 20 million reads, at least 50 million reads, or at least 100 million reads. In some embodiments, the bulk sequencing data comprises between 1 million and 5 million reads, between 3 million and 10 million reads, between 5 million and 20 million reads, between 10 million and 50 million reads, between 30 million and 100 million reads, or between 1 million and 100 million reads (or any number of reads inclusive and in between).
[0159] In some embodiments, the expression data comprises next generation sequencing (NGS) data. In some embodiments, the expression data comprises microarray data.
[0160] Any of the methods or compositions described herein may use expression data (e.g., indicating RNA expression levels) for multiple genes. The number of genes that may be examined may be less than or equal to the total number of genes of interest. In some embodiments, expression levels may be determined for all genes of interest. As non-limiting examples, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, 29 or more, 30 or more, 35 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 225 or more, 250 or more, 275 or more, or 300 or more genes may be used in any assessment described herein.
[0161] In some embodiments, the RNA expression data is obtained by accessing the RNA expression data from at least one computer storage medium on which the RNA expression data is stored. Additionally or alternatively, in some embodiments, the RNA expression data may be received from one or more sources via any suitable type of communications network. For example, in some embodiments, the RNA expression data may be received from a server (e.g., an SFTP server or Illumina BaseSpace).
[0162] The obtained RNA expression data may be in any suitable format, as aspects of the technology described herein are not limited in this regard. For example, in some embodiments, the RNA expression data may be obtained in a text-based file (e.g., FASTQ, FASTA, BAM, or SAM format). In some embodiments, the file storing the sequencing data may include a quality score of the sequencing data. In some embodiments, the file storing the sequencing data may include sequence identification information.
[0163] Expression data, in some embodiments, includes gene expression levels. Gene expression levels may be detected by detecting products of gene expression, such as mRNA and / or protein. In some embodiments, gene expression levels are determined by detecting mRNA levels in a sample. As used herein, the terms "determining" or "detecting" may include assessing the presence, absence, quantity and / or amount (which may be an effective amount) of a substance within a sample, including deriving a qualitative or quantitative concentration level of such substance, or otherwise assessing the value and / or classification of such substance in a sample from a subject.
[0164] In some embodiments, the sequencing data is obtained from a biological sample obtained from a subject. The sequencing data is obtained by any suitable method, for example, using any of the methods described herein, including those in the section entitled "Biological Samples." In some embodiments, the sequencing data comprises RNA-seq data. In some embodiments, the biological sample comprises blood or tissue. In some embodiments, the biological sample comprises one or more tumor cells and / or one or more immune cells (e.g., PBMCs).
[0165] In some embodiments, the obtained sequencing data is normalized to transcript kilobases per million (TPM). Normalization may be performed using any suitable software and in any suitable manner. For example, in some embodiments, TPM normalization may be performed according to the techniques described in Wagner et al. (Theory Biosci. (2012) 131:281-285), which is incorporated herein by reference in its entirety. In some embodiments, TPM normalization may be performed using a software package, such as the gcrma package. Aspects of the gcrma package are described in Wu J, Gentry RIwcfJMJ (2021). "gcrma: Background Adjustment Using Sequence Information. R package version 2.66.0.", which is incorporated herein by reference in its entirety. In some embodiments, the RNA expression level in TPM for a particular gene may be calculated according to the following formula:
number
[0166] In some embodiments, RNA expression levels normalized to TPM units may be logarithmically transformed. In some embodiments, RNA expression levels may not be normalized to transcripts per million units, but may instead be converted to another type of unit (e.g., reads per kilobase per million (RPKM) or fragments per kilobase per million (FPKM) or any other suitable unit). Additionally, or alternatively, in some embodiments, the logarithmic transformation may be omitted. Alternatively, in some embodiments, no transformation may be applied, or one or more other transformations may be applied instead of the logarithmic transformation.
[0167] In some embodiments, RNA expression data may include sequence data generated by a sequencing protocol (e.g., a series of nucleotides in a nucleic acid molecule identified by next generation sequencing, Sanger sequencing, etc.), as well as information contained therein that may be considered information that can be inferred or determined from the sequence data (e.g., information that indicates source, tissue type, population of cell types, etc.). In some embodiments, expression data may include information contained in a FASTA file, descriptions and / or quality scores contained in a FASTQ file, aligned positions contained in a BAM file, and / or any other suitable information obtained from any suitable file.
[0168] Techniques for correlating leukocyte signatures with leukocyte immune profile types As discussed herein, a subject may be determined to have a particular leukocyte immune profile type. To this end, a leukocyte signature for the subject may be determined (e.g., using cytometry (e.g., flow cytometry) or RNA sequencing), and the leukocyte signature for the subject may be associated with one particular leukocyte signature cluster in a collection of leukocyte signature clusters (each corresponding to a respective leukocyte immune profile type (e.g., naive, primed, advanced, chronic, or suppressed)). The leukocyte signature for the subject may be associated with one of the signature clusters in the collection in various ways.
[0169] For example, in some embodiments, a subject's leukocyte signature may be associated with a particular one of a plurality of leukocyte signature clusters by using a distance-based comparison or any other suitable metric, and based on the results of the comparison, the leukocyte signature may be associated with the closest leukocyte signature cluster ("closest" in the sense that a distance-based comparison is performed or whatever distance metric or measure is used). Examples of this are described herein, including with respect to FIG. 4.
[0170] For example, in some embodiments, a subject's leukocyte signature may be associated with a particular one of a plurality of leukocyte signature clusters by using a trained classifier. The trained classifier may be a multi-class classifier. The trained classifier may process the leukocyte signature to obtain an output indicative of a particular one of the plurality of leukocyte signature clusters. To this end, the leukocyte signature may be provided as an input to the trained classifier (optionally with suitable preprocessing, e.g., normalization) to obtain an output indicative of a particular one of the plurality of leukocyte signature clusters. For example, the output may indicate a numerical value (e.g., score, likelihood, and / or probability) for each of the signature clusters, and these numerical values may be used to select the signature cluster with which the subject's leukocyte signature is associated (e.g., selecting the cluster with the largest value (e.g., when the numerical value is a probability) or the smallest score (e.g., when the score is a log-likelihood)).
[0171] For example, in some embodiments, a leukocyte signature may include cell percentages for several cell populations (e.g., respective cell percentages for each of at least some (e.g., all) of the cell populations listed in Table 1, Table 2, Table 3, and / or Table 4). The cell percentages in the leukocyte signature may be normalized and then provided as input features to a trained classifier, thereby generating an output for selecting a signature cluster associated with the leukocyte signature. For example, the cell percentages may be renormalized as percentages from the PBMC fraction of a PBMC population or as percentages from the WBC fraction of a granulocyte population. Additionally or alternatively, the cell percentages may be recalculated as a minimax normalization that creates signature clusters using only the 2 and 98 percentiles of the cohort (e.g., out of 850 samples). Thus, in some embodiments, the input to the trained classifier may be a one-dimensional vector of normalized cell percentages (e.g., a 30x1 or 34x1 vector of numbers in the range [0,1]), and the output may be the probability (or other numerical value indicating likelihood) of belonging to each of five leukocyte signature clusters (these clusters correspond to each immune profile type), and thus may be a 5x1 vector of numbers in the range [0,1] that would sum to 1. In this example, the signature cluster with the highest predicted probability of the five may be the signature cluster to which the subject's leukocyte signature is assigned. As a specific, non-limiting example, if the output of the trained classifier (for an input of a subject's normalized cell percentages) is [0.8, 0.1, 0.07, 0.0, 0.03], then the leukocyte signature for the subject may be assigned to the first cluster (e.g., the "naive" or "G1" signature cluster).
[0172] Any of a number of types of classifiers may be used to associate a subject's leukocyte signature with a particular one of multiple leukocyte signature clusters, for example, a k-nearest neighbor (KNN) classifier, a decision tree classifier, a gradient boosting decision tree classifier, a Bayesian classifier, or a neural network classifier.
[0173] As an example, a neural network classifier may be used. For example, in some embodiments, a tabular pre-fitted network transformer (TabPFN) classifier may be used. For example, a TabPFN classifier having an architecture, for example, as described in N. Hollmann, S. Muller, K. Eggensperger, and F. Hutter, "TABPFN: A transformer that solves small tabular classification problems in a second," The Eleventh International Conference on Learning Representations (ICLR) 2023 (which is incorporated herein by reference in its entirety), and trained using the method described therein may be used. As an example, a TabPFN classifier may be trained for a cohort having N (e.g., 850) samples with cluster labels and M (e.g., 34) normalized cell population percentages using leave-one-out cross-validation by taking N-1 samples, evaluating the accuracy of the prediction on the remaining samples, and repeating this method N times (e.g., 850 times) to estimate the error.
[0174] As another example, a decision tree classifier may be used. Any suitable type of decision tree classifier may be used, and any suitable supervised decision tree learning method may be used to train it. For example, the decision tree classifier may be trained by iterative bisection (e.g., the ID3 algorithm, as described, for example, in Quinlan, JR 1986. Induction of Decision Trees. Mach. Learn. 1, 1 (Mar. 1986), 81-106), C4.5 (e.g., as described, for example, in Quinlan, JR C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers, 1993), or classification and regression tree (CART) (e.g., as described, for example, in Breiman, Leo; Friedman, JH; Olshen, RA; Stone, CJ (1984). Classification and regression trees. Monterey, CA: Wadsworth & Brooks / Cole Advanced Books & Software). It should be understood that any other suitable training method may be used to train the decision tree classifier, as aspects of the technology described herein are not limited in this respect.
[0175] As another example, a gradient boosting decision tree classifier may be used.
[0176] A gradient-boosted decision tree classifier can be an ensemble of multiple decision tree classifiers (sometimes referred to as "weak learners"). The prediction (e.g., classification) generated by the gradient-boosted decision tree classifier is based on the predictions generated by the multiple decision tree portions of the ensemble. The ensemble may be trained using an iterative optimization method that involves calculating the gradient of a loss function (hence the name "gradient" boosting). For example, the gradient-boosted decision tree classifier may be trained using any suitable supervised training algorithm, including any of the algorithms described in Hastie, T.; Tibshirani, R.; Friedman, JH (2009). "v10. Boosting and Additive Trees." The Elements of Statistical Learning (2nd ed.). New York: Springer. pp. 337-384. In some embodiments, the gradient boosting decision tree classifier may be implemented using any suitable publicly available gradient boosting framework, such as XGBoost (e.g., as described in Chen, T., & Guestrin, C. (2016). XGBoost: AScalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). New York, NY, USA: ACM. The XGBoost software may be available, for example, from http: / / xgboost.ai).
[0177] Another example of a framework that may be utilized is LightGBM (e.g., as described in, e.g., Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., … Liu, T.-Y. (2017). Lightgbm: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146-3154.). The LightGBM software may be available, for example, from https: / / lightgbm.readthedocs.io / ).
[0178] It should be understood that while some embodiments may use a multi-class classifier to associate leukocyte signatures with clusters, other embodiments may use multiple classifiers (e.g., multiple binary classifiers). For example, each cluster may be associated with a respective binary classifier trained to generate a numerical score to which the leukocyte signature belongs in that cluster. The outputs from the multiple classifiers may then be compared to potentially identify the signature cluster to which the leukocyte signature for the subject should be associated.
[0179] Yet another approach to associating a subject's leukocyte signature with a signature cluster involves determining, for each particular one of a plurality of leukocyte signature clusters, a score indicating whether the subject's leukocyte signature is associated with that particular cluster, wherein determining the score for a particular cluster includes applying a linear regression model associated with the particular cluster to the cell composition percentages in the leukocyte signature. The linear regression model may be trained using a regularization method, for example, a model whose coefficients are determined using elastic net linear regression.
[0180] Cellular composition percentage and application Aspects of the present disclosure relate to generating a leukocyte signature for a subject by processing cytometry data and / or RNA expression data to obtain a cellular composition percentage. As used herein, "cellular composition percentage" refers to the percentage of a particular cell type in a plurality of cells. For example, if 100 cells are identified as CD4 T cells in a total cell population of 500 cells, the cellular composition percentage of CD4 T cells in the population is 20%.
[0181] 3 is a flowchart of a method 300 that may be used to perform process 108 (and thus is an exemplary implementation of process 108) for determining cellular composition percentages using cytometry data. Method 300 may be performed partially or fully using a laptop computer, a desktop computer, one or more servers, a computing device as described herein in connection with FIG. 15, in a cloud computing environment, or any other suitable computing device or devices, as aspects of the technology described herein are not limited in this respect.
[0182] Process 300 begins with operation 302, in which cytometry data is obtained for a biological sample from a subject, the biological sample including a plurality of cells. In some embodiments, operation 302 may be performed in any suitable manner as described herein. For example, cytometry (e.g., flow cytometry) may be performed on the biological sample (e.g., using any suitable flow cytometry device or platform) to obtain the cytometry data.
[0183] Next, in process 304, each of at least a portion of the plurality of cells is identified as a respective type based on the cytometry data obtained in process 302. In some embodiments, process 304 may be performed at least according to the techniques described herein, including with respect to FIG. 1 for identifying the type of cells in a biological sample.
[0184] Next, in process 306, a cell count is determined for each of the plurality of cell types identified in process 304. In some embodiments, this includes determining the number of cells, i.e., cell count, of each type of cell for which cytometry measurements are obtained in process 302. The cell count, in some embodiments, may be used to determine the number of cells of each type of cell included in at least a cell type hierarchy. A cell type hierarchy may indicate the relationship between different cell types. For example, a cell type hierarchy may include a parent cell type and cell types that are children, or subtypes, of the parent cell type. In some embodiments, process 306 receives as input data indicating a cell type hierarchy. Such data may be provided in any suitable format, as aspects of the technology described herein are not limited in this respect.
[0185] In some embodiments, process 306 may also receive data indicating the type identified (in process 304) for each of a plurality of objects (e.g., cells, debris, beads, unidentified objects, etc.) in the biological sample. For example, the input may include a file of tab-delimited values having a number of rows corresponding to the number of objects. At least some of the rows may each include an indication of the type determined for that object. In some embodiments, at least some of the cell types indicated for the objects are included in a cell type hierarchy. In some embodiments, the cell type hierarchy includes one or more cell types. For example, the identified cell types may include a type for "double," which is a combination of two different cell types (e.g., "monocytes and neutrophils"). As another example, the identified cell types may include one or more custom cell types (e.g., "dead neutrophils") that one or more of the machine learning models were trained to predict.
[0186] In some embodiments, a "pre-processing" cell count is determined for each unique cell type listed in the type-indicating data identified for the subsample, for example, this includes determining counts for types included in a hierarchy of cell types and types not included in a hierarchy of cell types.
[0187] In some embodiments, the determined cell count is then updated according to the cell types included in the cell type hierarchy. For example, this may include attributing a cell count determined for an identified cell type not included in the hierarchy to a cell type included in the hierarchy. For example, a cell count determined for an identified cell type "dead neutrophil," which is not included in the hierarchy, may be attributed to the cell type "neutrophil," which is included in the hierarchy. For example, the cell count may be added to the cell count for neutrophils. Thus, in some embodiments, the cell count for "dead neutrophil" may be discarded because this cell count is considered to be the "neutrophil" cell type. In some embodiments, in updating the determined cell count according to the cell types included in the cell type hierarchy, the "double" may also be split into two different cell types, and the cell count for each cell type may be updated accordingly. For example, a count of "monocytes and neutrophils" may be split into a count of monocytes and a count of neutrophils. Thus, in some embodiments, any existing cell counts for monocytes and neutrophils may be updated to include said counts, which may be considered by the "monocyte" and "neutrophil" cell types, and the "monocyte and neutrophil" cell counts may be discarded.
[0188] In some embodiments, the cell count of a parent cell type in a cell type hierarchy is determined as the sum of the cell counts of its descendants (e.g., subtypes). For example, because "classical monocytes" are a subtype of "monocytes," a cell identified as a "classical monocyte" is also a "monocyte." Thus, in some embodiments, the cell count of a parent cell type in a cell type hierarchy may be updated based on the cell counts of its descendants. For example, the cell count of the descendants may be added to the existing parent cell count, or may be added from zero if there is no existing cell count for the parent cell type. In some embodiments, the technique of updating the cell count of a parent cell type may be performed sequentially from the bottom of the cell type hierarchy to the top of the cell type hierarchy.
[0189] Next, in process 308, a cellular composition percentage is determined for each of at least some of the identified cell types. In some embodiments, determining the cellular composition percentage for a particular cell type includes determining the ratio of the number of cells of a particular type to the total number of cells determined for the biological sample. In some embodiments, determining the cellular composition percentage for a particular cell type includes determining the ratio of the number of cells of a particular type to the total number of immune cells determined for the biological sample. In some embodiments, determining the cellular composition percentage for a particular cell type includes determining the percentage of the particular cell type relative to a cell type class associated with the particular cell type in the biological sample. For example, determining the percentage of naive T cells relative to the total number of T cells identified in the biological sample. For example, the total number of cells may be determined as the number of white blood cells determined for the biological sample.
[0190] In some embodiments, the determined cellular composition percentages for specific cell types are used to determine the cell concentrations of those cell types in a biological sample. For example, the normalized cellular composition percentages may be multiplied by respective coefficients that convert the cellular composition percentages to cell concentrations. Aspects of machine learning models are described herein, including at least those in the section "Cytometry-Based Cell Deconvolution."
[0191] In other embodiments of the methods described herein, the RNA expression data is processed using cellular deconvolution techniques to generate cellular composition percentages for some (or all) of the cell types listed in Table 3. Generation of cellular composition percentages using cellular deconvolution techniques, such as the BostonGene Kassandra technique, is described, for example, in PCT / US2021 / 022155, published September 16, 2021 as WO 2021 / 183917; and PCT / US2022 / 027088, published November 3, 2022 as WO 2022 / 232615, the entire contents of each of which are incorporated herein by reference. Aspects of the machine learning model are described herein, including at least those in the section "RNA-Based Cellular Deconvolution."
[0192] Other cellular deconvolution techniques may also be used in the methods described by the present disclosure, such as, for example, Cibersort (e.g., as described by Newman et al. Nature Methods volume 12, pages 453-457 (2015)) or CibersortX (e.g., as described by Newman et al. Nature Biotechnology volume 37, pages 773-782 (2019)). In some embodiments, two or more cellular deconvolution methods are used, and then a consensus from the two or more cellular deconvolution methods is used to determine the cellular deconvolution.
[0193] Cytometry-based cell deconvolution The cytometry data may be processed using any suitable technique to identify the cellular composition percentages, and any one of a number of techniques may be used for this purpose, as aspects of the technology described herein are not limited in this respect.
[0194] For example, in some embodiments, processing the cytometry data to determine cellular composition percentages may include plotting the cytometry data as a series of two-dimensional plots and identifying distinctive cell populations based on shared marker expression, commonly referred to as "gating." Gating the cytometry data may include manually gating the cytometry data to isolate distinctive cell populations. Additionally or alternatively, gating may be performed using any suitable gating technique, such as by using FlowJo™ (FlowJo™ Software. Ashland, OR: Beckton, Dickinson and Company; 2021). In some embodiments, the number of cells contained in the identified cell populations may be used to determine the corresponding cellular composition percentages.
[0195] Additionally or alternatively, processing the cytometry data to determine the cellular composition percentages may include clustering the cytometry data to identify characteristic cell populations. In some embodiments, clustering the cytometry data may include calculating a two-dimensional t-SNE plot for the sample and calculating a FlowSOM for the sample. FlowSOM is described by Van Gassen et al. (“FlowSOM: Using self-organizing maps for visualization and interpretation of cytometry data,” in Journal of Quantitative Cell Science, vol. 87, no. 7, pp. 636-645, 2015), which is incorporated herein by reference in its entirety. The resulting clusters may correspond to characteristic cell populations. In some embodiments, the number of cells contained in the identified cell populations may be used to determine the corresponding cellular composition percentages.
[0196] Additionally or alternatively, processing the cytometry data to determine the cellular composition percentages can include processing the cytometry data using machine learning methods. For example, the cytometry data may be processed using the machine learning methods described by PCT / US2023 / 012003, published on August 3, 2023 as WO 2023 / 147177, which is incorporated herein by reference in its entirety.
[0197] For example, a machine learning method may include processing cytometry data using one or more machine learning models to identify the type of cells present in a biological sample. In some embodiments, the multiple machine learning models used to process the cytometry data include a first machine learning model and a second machine learning model different from the first machine learning model. In some embodiments, cytometry measurements corresponding to a particular event are processed using the first machine learning model to determine an event type for the particular event. An "event" corresponds to an object (e.g., a cell, debris, bead, doublet, or unidentified object) in the biological sample being measured by a cytometry platform (e.g., a flow cytometry platform or a mass cytometry platform). For example, an event may correspond to a cell in the biological sample being measured by the cytometry platform, and measurements obtained during the event may be included in the cytometry data. The determined event type (e.g., predicted by the first machine learning model) indicates whether the event corresponds to a cell being measured by the cytometry platform, debris being measured by the cytometry platform, or beads being measured by the cytometry platform. For example, the first machine learning model can include a multi-class classifier that is trained to distinguish between at least some event types.In some embodiments, when the determined event type indicates that a particular event corresponds to the cell that is being measured by the cytometry platform, the second machine learning model is used to process the cytometry measurement value corresponding to the particular event, and determine the type of cell for the particular event.For example, the second machine learning model can include a multi-class classifier that is trained to distinguish between at least some event cell types.
[0198] In some embodiments, the machine learning method includes processing cytometry data for cells (e.g., a type of event) in the biological sample using a hierarchy of machine learning models corresponding to a cell type hierarchy. The machine learning models in the hierarchy of machine learning models may be trained to predict a specific type of cell using the cytometry data corresponding to the cells. Additionally or alternatively, the machine learning models in the hierarchy of machine learning models may include a multi-class classifier trained to distinguish between at least some cell types at a particular level in the hierarchy. Different levels of the hierarchy of machine learning models may be used to predict a type of cell with different levels of specificity (e.g., a general cell type or a specific subtype). In some embodiments, the cell types determined for the cells in the biological sample are then used to determine the cellular composition percentage for the cell type.
[0199] 30 depicts an exemplary technique 3080 for processing cytometry data using multiple machine learning models to determine respective types for one or more events (e.g., cells or particles). In some embodiments, the exemplary technique 3080 includes providing first cytometry data 3032-1 of a first event as input to a hierarchy of machine learning models, which are used to determine one or more event types 3084, 3090 of the first event.
[0200] The cytometry data processed using the exemplary technique 3080 may include cytometry data for each of a plurality of cells and particles processed using the cytometry platform. For example, the cytometry data may include first cytometry data 3032-1 for a first event. FIG. 30 illustrates processing the first cytometry data 3032-1 to determine one or more types for the first event. However, it should be understood that the exemplary technique 3080 may be used to process cytometry data for any suitable number of events, such as second cytometry data for a second event, as aspects of the technology described herein are not limited to processing cytometry data for any particular number of events.
[0201] In some embodiments, the technique 3080 includes processing the first cytometry data 3032-1 through a machine learning model hierarchy. Figure 30 shows an exemplary machine learning model hierarchy, which includes machine learning models 3082a-c, 3086a-b, and 3088.
[0202] In some embodiments, a machine learning model may be trained to determine whether a first event is of a particular type based on the first cytometry data 3032-1. In some embodiments, this may include determining a probability that the first event is of that particular type. For example, the first event may correspond to a cell, and at least some (e.g., all) of the machine learning models in the hierarchy may each be trained to predict whether the cell is of a particular cell type. The cell types may include any suitable cell types, as aspects of the technology described herein are not limited in this respect. For example, the cell types may include any of the cell types listed in Table 1. In the example shown in FIG. 30 , event type B 3084 and event type 3090 may be cell types identified for a cell using the machine learning models in the machine learning model hierarchy. Event type B 3084 and event type 3090 may each be any suitable cell type, such as, for example, any of the cell types listed in Table 1.
[0203] As an example, machine learning model 3082a may be trained to determine whether a first event is of type A. As another example, machine learning model 3086b may be trained to determine whether a first event is of type E. Additionally or alternatively, the machine learning model may include a multi-class classifier trained to determine whether a first event is one of a plurality of different event types based on first cytometry data 3032-1. For example, machine learning model A 3082a may be trained to determine whether a first event is of type A1, type A2, or type A3. For example, the machine learning model may output the most likely type (e.g., among type A1, type A2, or type A3) for the first event. Such a machine learning model may output a type and / or a probability that the event is of the identified type. For example, machine learning model A 3082a may identify that an event is more likely to be of type A2 than type A1 or type A3, along with the probability that the event is of type A2. In some embodiments, the machine learning model may include a decision tree classifier, a gradient boosting decision tree classifier, a neural network, a support vector machine classifier, or any other suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect. In some embodiments, the machine learning model may include an ensemble of machine learning models of any suitable type (a machine learning model that is part of an ensemble may be referred to as a "weak learner"). For example, the machine learning model may include an ensemble of decision tree classifiers. Aspects of machine learning models are described herein, including at least those in the section "Machine Learning."
[0204] In some embodiments, different levels of the machine learning model hierarchy may be used to determine the event type with different levels of specificity. For example, machine learning models 3082a-c may be used to determine that a first event is of type B 3084, while machine learning models 3086a-b may be used to determine that the first event is of type E 3090, which is a subtype of type B 3084.
[0205] In some embodiments, the output of machine learning models 3082a-c is used to provide information about which one or more machine learning models in the hierarchy will subsequently be used to process the first cytometry data 3032-1. For example, the output of machine learning models 3082a-c may indicate which event type, of the event types associated with each of the models, is the most likely event type for the first event. As shown in this example, the output of machine learning models 3082a-c indicates that the first event is of type B 3084. Based on this output, technique 3080 may continue by determining whether the first event is a subtype of type B 3084. Thus, in some embodiments, the first cytometry data 3032-1 may be processed using machine learning models 3086a-b that are trained to determine whether an event is a subtype of type B 3084.
[0206] In some embodiments, a level in the machine learning model hierarchy may not indicate any type for the first event. For example, the level of the hierarchy that includes machine learning model 3088 does not indicate a type for the first event. In some embodiments, this may indicate that no machine learning model at that level in the hierarchy predicted the particular event type associated with the machine learning model for the first event (e.g., the machine learning model is trained to determine about). For example, machine learning model 3088 predicted that the first event is not type F. In some embodiments, if a level in the hierarchy does not indicate an event type, then the event type indicated at a previous level in the hierarchy may be determined to be the type of the first event. For example, type E 3090 may be determined to be the type for the first event. In this case, type E 3090 corresponds to the most specific type for the first event because type E 3090 is a subtype of type B 3084.
[0207] In some embodiments, the cell types identified for cells in a biological sample can be used to determine the number of cells of each type in the sample (e.g., cell count). The cell count can be used to determine the cellular composition percentage for different cell types. Exemplary techniques for determining cellular composition percentage based on cell count are described herein, including at least those related to Figure 3.
[0208] RNA-based cellular deconvolution Using a suitable cellular deconvolution technique, RNA expression data can be processed to identify cellular composition percentages. To this end, any one of several cellular deconvolution techniques can be used, as aspects of the technology described herein are not limited in this respect. Non-limiting examples of cellular deconvolution techniques include Kassandra, CIBERSORT, CIBERSORTx, QuanTIseq, FARDEEP, Xcell, ABIS, EPIC, MCP-counter, Scaden, and MuSiC. Kassandra is described in PCT / US2021 / 022155, published on September 16, 2021 as WO 2021 / 183917; and PCT / US2022 / 027088, published on November 3, 2022 as WO 2022 / 232615 (the entire contents of each of which are incorporated herein by reference). CIBERSORTx is described by Newman et al. ("Robust enumeration of cell subsets from tissue expression profiles." Nat. Methods 12, 453-457 (2015)), which is incorporated herein by reference in its entirety. CIBERSORTx is described by Newman, A., et al. ("Determining cell type abundance and expression from bulk tissues with digital cytometry." Nature biotechnology 37.7 (2019): 773-782), which is incorporated herein by reference in its entirety.QuanTIseq is described by Finotello, F., et al. ("Molecular and pharmacological modulators of the tumor immune contexture revealed by deconvolution of RNA-seq data." Genome medicine 11.1 (2019): 1-20), which is incorporated herein by reference in its entirety. FARDEEP is described by Hao, Yuning, et al. ("Fast and robust deconvolution of tumor infiltrating lymphocyte from expression profiles using least trimmed squares." PLoS computational biology 15.5 (2019): e1006976), which is incorporated herein by reference in its entirety. Xcell is described by Aran, D., et al. ("xCell: digitally portraying the tissue cellular heterogeneity landscape." Genome biology 18 (2017): 1-14), which is incorporated herein by reference in its entirety. Abis is described by Monaco, G, et al. (“RNA-Seq signatures normalized by mRNA abundance allow absolute deconvolution of human immune cell types.” Cell reports 26.6 (2019): 1627-1640), which is incorporated herein by reference in its entirety.EPIC is described by Racle, J. and Gfeller, D. ("EPIC: a tool to estimate the proportions of different cell types from bulk gene expression data." Bioinformatics for Cancer Immunotherapy: Methods and Protocols (2020): 233-248), which is incorporated herein by reference in its entirety. MCP-counter is described by Becht, E., et al. ("Estimating the population abundance of tissue-infiltrating immune and stromal cell populations using gene expression." Genome biology 17.1 (2016): 1-20), which is incorporated herein by reference in its entirety. Menden, K., et al. ("Deep learning-based cell composition analysis from tissue expression profiles." Science advances 6.30 (2020): eaba2619), which is incorporated herein by reference in its entirety. MuSiC is described by Wang, X., et al. (“Bulk tissue cell type deconvolution with multi-subject single-cell expression reference.” Nature communications 10.1 (2019): 380), which is incorporated herein by reference in its entirety.
[0209] In some embodiments, the cellular composition percentages are identified by processing the RNA expression data using the Kassandra cell deconvolution technique. In some embodiments, the Kassandra deconvolution technique includes processing the RNA expression data using one or more machine learning models to determine the cellular composition percentages for one or more cell types. For example, determining the cellular composition percentages for a particular cell type may include obtaining RNA expression data for a set of genes associated with the cell type (e.g., one or more marker genes, which may be specific or semi-specific genes for a particular cell type), and processing the RNA expression data with at least one machine learning model to determine the cellular composition percentages for the particular cell type. According to some embodiments, this method may be repeated or performed in parallel for each of multiple cell types to achieve deconvolution across multiple cell types.
[0210] In some embodiments, determining cellular composition percentages using the Kassandra deconvolution technique for a particular cell type includes estimating RNA percentages for the particular cell type and using the estimated RNA percentages to determine the cellular composition percentages. For example, estimating RNA percentages may include processing RNA expression data obtained for the cell type using at least one machine learning model trained to predict RNA percentages for the cell type. FIG. 29A depicts an illustrative example of determining RNA percentages based on RNA expression data using machine learning methods. In this example, RNA expression data from a biological sample 2902 is processed using one or more machine learning models to obtain RNA percentages 2906 for cell type A, cell type B, and cell type C. While this FIG. 29A example shows three cell types, the technique may be applied to determining RNA percentages for any suitable number of cell types, as aspects of the technology described herein are not limited in this respect. Additionally, the cell type may be any suitable cell type, such as, for example, any of the cell types listed in Table 18 and / or any of the cell types listed in WO 2021 / 183917.
[0211] In this example shown in Figure 29A, RNA expression data is obtained for biological sample 2902. The RNA expression data may be obtained using any suitable technique, such as those described herein.
[0212] Regardless of how the RNA expression data is obtained from the biological sample 2902, the RNA expression data may be processed using one or more machine learning models 2904. The one or more machine learning models 2904 may include any suitable one or more machine learning models, such as, for example, a nonlinear regression model (e.g., a logistic regression model), a neural network model, a support vector machine, a Gaussian mixture model, a random forest model, a decision tree model, or any other suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect. For example, the one or more machine learning models 2904 may be one or more nonlinear regression models. The one or more nonlinear regression models may be implemented using gradient boosting techniques (e.g., as implemented in XGBoost). Aspects of machine learning models are described herein, including at least those in the section "Machine Learning."
[0213] In some embodiments, the one or more machine learning models 2904 may include a separate machine learning model for each of a plurality of cell types. In this example shown in FIG. 29A , the machine learning models 2904 include a machine learning model for cell type A, a machine learning model for cell type B, and a machine learning model for cell type C. As shown, additional machine learning models for one or more additional cell types and / or subtypes may be provided in some embodiments. For example, one or more machine learning models may be provided for one or more cell types listed in Table 18 and / or any of the cell types listed in WO 2021 / 183917.
[0214] In some embodiments, the input to each of the machine learning models 2904 may include a subset of the RNA expression data 2902 selection. For example, the input to a machine learning model for a particular cell type may include RNA expression data of genes specific and / or semi-specific for that cell type. Table 18 and WO 2021 / 183917 list examples of genes specific and / or semi-specific for different cell types. In some embodiments, expression data obtained for genes listed for a particular cell type in Table 18 and / or WO 2021 / 183917 may be provided as input to a machine learning model trained to predict RNA percentages for the particular cell type. As a non-limiting example, expression data obtained for genes listed for basophils in Table 18 may be provided as input to a machine learning model trained to predict RNA percentages for basophils in a biological sample. In some embodiments, other information about the RNA expression data (e.g., the median of the RNA expression data, or any other suitable statistic) may additionally or alternatively be provided as input to the machine learning model.
[0215] In some embodiments, the output of the machine learning model 2904 may be RNA percentages 2906 for each cell type and / or subtype. For example, a machine learning model for cell type A may produce as its output the predicted percentage of RNA from cells of type A in the input RNA expression data. Similarly, a machine learning model for cell type B may produce as its output the predicted percentage of RNA from cells of type B, and a machine learning model for cell type C may produce as its output the predicted percentage of RNA from cells of type C. As described herein, the predicted percentages of RNA may be used to calculate corresponding cellular composition percentages for some or all of the cell types and / or subtypes under analysis.
[0216] FIG. 29B is a diagram depicting the use of machine learning models 2920, 2922, 2924 including first sub-models 2926, 2928, 2930 and second sub-models 2938, 2940, 2942 to determine RNA percentages based on RNA expression data.
[0217] As shown in Figure 29B, different machine learning models 2920, 2922, 2924 are used to process gene expression data 2914, 2916, 2918 associated with each cell type: cell type A 2908, cell type B 2910, and cell type C 2912. In some embodiments, each exemplary machine learning model includes a first sub-model 2926, 2928, 2930 for generating a first value 2932, 2934, 2936 for the estimated percentage of RNA from each cell type, and a second sub-model 2938, 2940, 2942 for generating a second value 2944, 2946, 2948 for the estimated percentage of RNA from each cell type.
[0218] As a non-limiting example of the use of a machine learning model including one or more sub-models, consider machine learning model 2922 trained to estimate RNA percentage for cell type B 2910. In some embodiments, expression data 2916 may be obtained from a set of genes associated with cell type B 2910 and may be used as input to machine learning model 2922. For example, cell type B 2910 may include basophils, and expression data 2916 may include expression data for at least some of the genes from the gene set associated with basophils listed in Table 18. In some embodiments, at least a portion of expression data 2916 (e.g., expression data associated with a subset of genes, expression data associated with all genes, etc.) is used as input to first sub-model 2928. For example, a subset of expression data 2916 including expression data for a subset of genes from the gene set associated with basophils may be used as input. The first sub-model may then process the input expression data to determine a first value 2934 of the estimated percentage of RNA from cell type B 2910 .
[0219] In some embodiments, the exemplary machine learning model 2922 may include a second sub-model 2940 that generates a second value 2946 of the estimated RNA percentage from cell type B 2910. In some embodiments, the second sub-model 2940 may use one or more inputs to generate the second value 2940. For example, in some embodiments, at least a portion of the expression data 2916 may be used as an input. In some embodiments, the expression data may include the same expression data input into the first sub-model 2928. In some embodiments, the expression data may include the same expression data input into the first sub-model, as well as additional expression data. In some embodiments, the expression data may include expression data that differs from the expression data input into the first sub-model.
[0220] Additionally or alternatively, in some embodiments, the second sub-model 2940 may take as input the estimated RNA percentages output for other cell types 2908, 2912 by the first sub-models 2926, 2930 of the machine learning models 2920, 2924. As shown, the second sub-model 2940 for cell type B 2910 takes as input a first value 2932 for the estimated RNA percentage from cell type A 2908 and a first value 2936 for the estimated RNA percentage from cell type C 2912. This type of input can be informative when attempting to determine the RNA percentages from cell types associated with the same genes or the same set of genes as another cell type or types. For example, if cell type B 2910 is associated with the same gene as cell type C 2912, then expression data obtained for that gene may not be very informative regarding which of these two cell types is present in the biological sample, as it may be unclear which cell type generated the expression data. However, consider a scenario in which first sub-model 2930 outputs 0% as the first value 2936 for the estimated RNA percentage determined for cell type C, indicating that there are no cells of cell type C 2912 in the biological sample. As a result, any expression data obtained for the shared gene should be expressed by cell type B 2910. In some embodiments, second sub-model 2940 can make such inferences using first values 2932, 2936.
[0221] In some embodiments, the output of the second sub-model 2940 is a second value 2946 for the estimated RNA percentage from cell type B 2910.
[0222] In some embodiments, the estimated RNA percentages can be processed to determine the cellular composition percentages for each of the cell types. The RNA percentages can be processed using any technique suitable for obtaining cellular composition percentages, as aspects of the technology described herein are not limited in this respect.
[0223] As a non-limiting example, determining the cellular composition percentages based on the RNA percentages can include applying Equation 1 to the RNA percentages:
number
number
[0224] RNA factor A per cell cell may correspond to the RNA concentration per cell. The RNA coefficient per cell may be used to convert RNA percentages to corresponding cellular composition percentages. In some embodiments, the RNA coefficient per cell A cell may be determined as part of the model training process (e.g., from simulated or artificial data where the percentages of different cell types are known). In some embodiments, the RNA coefficient A per each cell cellmay be experimentally determined for some or all cell types. For example, the RNA coefficient per cell may be obtained by accessing data related to RNA expression for each cell type (e.g., from available scientific literature, such as PMID:29130882, PMID:30726743, etc., or estimated from single-cell data using average or nonlinearly transformed UMI counts per cell type), and that data may be used to determine the corresponding RNA coefficient per cell for each cell type (e.g., by analyzing purity and / or histological TCGA lymphocyte data). In some embodiments, the RNA coefficient per cell may be tissue-specific and may vary based on the disease under analysis (e.g., between cancers). In some embodiments, the RNA coefficient per cell may be tissue-independent and may not vary based on the disease under analysis (e.g., because non-malignant microenvironment cells may be represented by the same or substantially similar cellular phenotypes even between different cancers, tissues, or diseases). In the latter case, data from multiple types of cancers, tissues, diseases, etc. may be combined to calculate the RNA coefficient per cell. For example, in some embodiments, more than 10,000 different TCGA cancer tissue samples were analyzed as part of determining the RNA coefficients per cell for each cell type. The inventors recognized and understood that the non-malignant cell composition percentage may correspond to tumor cellularity as defined by histology and WES analysis. Therefore, in some embodiments, determining the RNA coefficients per cell may include aligning the non-malignant cell composition percentages obtained from RNA with the cell composition percentages obtained from DNA to determine the coefficients for each RNA type per cell.
[0225] In some embodiments, Equation 1 may be applied to each RNA percentage independently (e.g., sequentially), or in some embodiments, to some or all of the RNA percentages together (e.g., in parallel). In some embodiments, Equation 1 may initially be applied to RNA percentages for cell types that are not subtypes of each other. In some embodiments, Equation 1 may subsequently be applied to RNA percentages for cell types that are subtypes of one or more of the originally used cell types. In some embodiments, the calculation of cellular composition percentages for cell subtypes may be modified based on the originally calculated cellular composition percentages. For example, in some embodiments, subsequently calculated cellular composition percentages for cell subtypes may be normalized or otherwise adjusted to be the sum of the cellular composition percentages for all cell types (i.e., for the originally calculated cell type they are subtypes of).
[0226] Generation of leukocyte signatures and identification of leukocyte immune profile types In some embodiments, leukocyte immune profile types may be generated by (1) obtaining leukocyte signatures (using techniques described herein) for a plurality of subjects; and (2) clustering the leukocyte signatures so obtained into a plurality of leukocyte immune profile types. Any suitable clustering technique may be used for this purpose, including, but not limited to, density clustering, spectral clustering, k-means clustering, hierarchical clustering, and / or agglomerative clustering algorithms.
[0227] In some embodiments, the leukocyte immune profile type may be identified by clustering multiple leukocyte signatures for multiple subjects. In some embodiments, the leukocyte immune profile type is identified using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm. In some embodiments, the clustering is performed using a spectral clustering algorithm.
[0228] For example, the similarity between samples may be calculated using Pearson correlation. The distance matrix may be transformed into a graph where each sample forms a node and two nodes form an edge with a weight equal to their Pearson correlation coefficient. Edges with a weight less than a specified threshold may be removed. The Louvain community detection algorithm may be applied to calculate the graph's partition into clusters. The minimum Davies-Bouldin, maximum Calinski-Harabasz, and Silhouette methods may be used to mathematically determine the optimal weight threshold for the observed clusters. Separations in clusters with low populations (less than 5% of the samples) may be eliminated.
[0229] Thus, in some embodiments, generating a leukocyte immune profile type involves (A) obtaining a plurality of sets of data (e.g., cytometry data, RNA expression data, etc.) from biological samples obtained from a plurality of respective subjects, wherein each of the plurality of data sets includes information indicative of the presence, absence, and / or respective amounts of a certain cell type, such as WBCs or PBMCs (e.g., some or all of the cell types listed in Table 1, Table 2, Table 3, or Table 4), in the subject's biological sample; (B) generating a plurality of leukocyte signatures from the plurality of data sets, wherein each of the plurality of leukocyte signatures includes a cellular composition percentage for a respective cell type (e.g., some or all of the cell types listed in Table 1, Table 2, Table 3, or Table 4) in the respective subject's biological sample; and (C) clustering the plurality of leukocyte signatures to obtain a plurality of leukocyte immune profile types.
[0230] The resulting leukocyte immune profile types may each comprise any suitable number of leukocyte signatures, e.g., at least 10, at least 100, at least 500, at least 500, at least 1000, at least 5000, 100-10,000, 500-20,000, or any other suitable range within these ranges, as aspects of the technology described herein are not limited in this respect.
[0231] The number of leukocyte immune profile types in this example is 5. An important aspect of the present disclosure is the inventors' discovery that certain unique types of leukocyte immune profiles can be characterized into five types based on the generation of leukocyte signatures using the methods described herein.
[0232] 4, a subject's leukocyte signature 400 may be associated with one of five leukocyte immune profile types: 402, 404, 406, 408, and 410. Each of clusters 402, 404, 406, 408, and 410 may be associated with a respective leukocyte immune profile type. In this example, leukocyte signature 400 is compared to each cluster (e.g., using a distance-based comparison or any other suitable metric), and based on the results of the comparison, leukocyte signature 400 is associated with the closest leukocyte signature cluster ("closest" in the sense that a distance-based comparison is performed, or whatever distance metric or measure is used). In this example, white blood cell signature 400 is associated with white blood cell signature cluster 4 406 (as indicated by the matching shading) because the measured distance D4 between white blood cell signature 400 and cluster 406 (e.g., its centroid or other point representative thereof) is less than the measured distances D1, D2, D3, and D5 between white blood cell signature 400 and clusters 402, 404, 408, and 410, respectively (e.g., their centroids or other points representative thereof).
[0233] In some embodiments, a subject's leukocyte signature may be associated with one of five leukocyte immune profile types by using a machine learning method (such as, for example, k-nearest neighbor (KNN) or any other suitable classifier) to assign the leukocyte signature to one of the five leukocyte immune profile types. The machine learning method may be trained to assign the leukocyte signature to a meta-cohort represented by the signature in a cluster. Aspects of the machine learning model are described herein, including at least those in the section "Techniques for Associating Leukocyte Signatures with Leukocyte Immune Profile Types."
[0234] In some embodiments, the subject's leukocyte signature may be associated with one of five leukocyte immune profile types using a linear regression model, e.g., elastic net linear regression. In some embodiments, associating the subject's leukocyte signature with a particular one of the plurality of leukocyte immune profile types comprises determining, for each particular one of the plurality of leukocyte immune profile types, a score indicating whether the subject's leukocyte signature is associated with that particular cluster. In some embodiments, determining the score for a particular cluster comprises applying a linear regression model associated with the particular cluster to the cellular composition percentages in the leukocyte signature.
[0235] In some embodiments, leukocyte immune profile types include naive type (e.g., G1), primed type (e.g., G2), progressive type (e.g., G3), chronic type (e.g., G4), and suppressed type (e.g., G5). The leukocyte immune profile types described herein may be described by qualitative characteristics, such as, for example, a high cellular composition percentage for certain cell types or a low signal cellular composition percentage for certain other cell types. In some embodiments, a high cellular composition percentage refers to a higher cellular composition percentage of the same cell type in a subject under analysis compared to a subject with a different type of cancer or a healthy subject. In some embodiments, a low cellular composition percentage refers to a lower cellular composition percentage of the same cell type in a subject under analysis compared to a subject with a different type of cancer or a healthy subject. In some embodiments, a "high" signal refers to a cellular composition percentage signal that is at least 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 50-fold, 100-fold, 1000-fold, or more increased compared to the cellular composition percentage of the same cell type in a subject with a different type of cancer or a healthy subject. In some embodiments, a "low" signal refers to a cellular composition percentage that is at least 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 50-fold, 100-fold, 1000-fold, or more decreased compared to the cellular composition percentage of the same cell type in a subject with a different type of cancer or a healthy subject.
[0236] In some embodiments, a suppressed leukocyte immune profile type (e.g., G5) is characterized by an increased number of myeloid cell populations, including classical monocytes and neutrophils, compared to other leukocyte immune profile types.
[0237] In some embodiments, a chronic leukocyte immune profile type (e.g., G4) is characterized by increased numbers of CD8 memory and effector cells and NKT cell populations compared to other leukocyte immune profile types.
[0238] In some embodiments, a primed leukocyte immune profile type (e.g., G2) is characterized by an increased number of T helper memory cells, including CD4 central memory, compared to other leukocyte immune profile types.
[0239] In some embodiments, advanced cell memory leukocyte immune profile types (e.g., G3) are characterized by increased numbers of CD4 and CD8 memory cells and a large increase in CD8 transitional memory cells compared to other leukocyte immune profile types.
[0240] In some embodiments, a naive leukocyte immune profile type (eg, G1) is characterized by increased numbers of naive CD4, CD8, and B cells relative to other leukocyte immune profile types.
[0241] In some embodiments, leukocyte immune profile types may be characterized by gene expression profiles that indicate the biological processes underlying a particular leukocyte immune profile type. For example, in some embodiments, leukocyte immune profile types are characterized according to Molecular Signature Database (MSigDB; described by Liberzon et al. Cell Syst. 2015 Dec 23; 1(6): 417-425) signatures. In some embodiments, the MSigDB signature is selected from TCF LEF CTNNB1 binding to target promoters, WNT beta-catenin signaling, T cell receptor and costimulatory signaling, CTLA4 pathway, NK cell-mediated cytotoxicity, antigen processing and presentation, graft-versus-host disease, ILC family development and heterogeneity, CD8 TCR downstream pathway, NFAT TF pathway, cancer immunotherapy with PD1 blockade, CTL pathway, allograft rejection, IL12 2 pathway, neutrophil degranulation, innate immune system, IL1 family signaling, signaling by GPCRs, signaling by receptor tyrosine kinases, elevated KRAS signaling, negative regulation of the PI3K AKT network, VEGFR1 2 pathway, naive vs. PD-1 high CD8 T cells, naive vs. activated CD8 T cells, PD-1 signaling, and cancer immunotherapy signatures by PD1 blockade.
[0242] In some embodiments, the "TCF LEF CTNNB1 binding to target promoter" signature includes gene expression scores (e.g., GSEA scores) for the following genes: AXIN2, CTNNB, LEF1, MYC, RUNX3, TCF7, TCF7L1, and TCF7L2.
[0243] In some embodiments, the "WNT beta-catenin signaling" signature includes gene expression scores (e.g., GSEA scores) for the following genes: ADAM17, AXIN1, AXIN2, CCND2, CSNK1E, CTNNB1, CUL1, DKK1, DKK4, DLL1, DVL2, FRAT1, FZD1, FZD8, GNAI1, HDAC11, HDAC2, HDAC5, HEY1, HEY2, JAG1, JAG2, KAT2A, LEF1, MAML1, MYC, NCOR2, NCSTN, NKD1, NOTCH1, NOTCH4, NUMB, PPARD, PSEN2, PTCH1, RBPJ, SKP2, TCF7, TP53, WNT1, WNT5B, and WNT6.
[0244] In some embodiments, the "T cell receptor and costimulatory signaling" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: PDK1, NFKB1, NFKBIA, NFATC2, IL2, CSNK1A1, PLCG1, ZAP70, PRKCA, PTPN6, DYRK1A, LCK, PDCD1, PPP3CA, DYRK2, CTLA4, PTEN, GSK3B, GSK3A, FYN, CD28, AKT1, CALM1, RASA1, RASGRP1, CD8A, CALM2, CD8B, and ITK.
[0245] In some embodiments, the "CTLA4 pathway" signature includes gene expression scores (e.g., GSEA scores) for the following genes: CD247, CD28, CD3D, CD3E, CD3G, CD80, CD86, CTLA4, GRB2, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, ICOS, ICOSLG, IL2, ITK, LCK, PIK3CA, PIK3R1, and PTPN11.
[0246] In some embodiments, the "NK cell-mediated cytotoxicity" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: ARAF, BID, BRAF, CASP3, CD244, CD247, CD48, CHP1, CHP2, CSF2, FAS, FASLG, FCER1G, FCGR3A, FCGR3B, FYN, GRB2, GZMB, HCST, HLA-A, HLA-B, HLA-C, HLA-E, HLA-G, HRAS, ICAM1, ICAM2, IFNA1, IFNA10, IF NA13, IFNA14, IFNA16, IFNA17, IFNA2, IFNA21, IFNA4, IFNA5, IFNA6, IFNA7, IFNA8, IFNAR1, IFNAR2, IFNB1, IFNG, IFNGR1, IFNGR2, ITGAL, I TGB2, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DL5A, KIR2DS1, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1, KIR3DL2, KLRC1, KLRC2, KLRC3, KLRD 1, KLRK1, KRAS, LAT, LCK, LCP2, MAP2K1, MAP2K2, MAPK1, MAPK3, MICA, MICB, NCR1, NCR2, NCR3, NFAT5, NFATC1, NFATC2, NFATC3, NFATC4, NRAS , PAK1, PIK3CA, PIK3CB, PIK3CD, PIK3CG, PIK3R1, PIK3R2, PIK3R3, PIK3R5, PLCG1, PLCG2, PPP3CA, PPP3CB, PPP3CC, PPP3R1, PPP3R2, PRF1, PR KCA, PRKCB, PRKCG, PTK2B, PTPN11, PTPN6, RAC1, RAC2, RAC3, RAET1E, RAET1G, RAET1L, RAF1, SH2D1A, SH2D1B, SH3BP2, SHC1, SHC2, SHC3, SHC 4, SOS1, SOS2, SYK, TNF, TNFRSF10A, TNFRSF10B, TNFRSF10C, TNFRSF10D, TNFSF10, TYROBP, ULBP1, ULBP2, ULBP3, VAV1, VAV2, VAV3, and ZAP70.
[0247] In some embodiments, the "antigen processing and presentation" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: B2M, CALR, CANX, CD4, CD74, CD8A, CD8B, CIITA, CREB1, CTSB, CTSL, CTSS, HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPAl, HLA-DPBl, HLA-DQA1, HLA-DQA2, HLA-DQBl, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, HSP90AA1, HSP90AB1, HSPA1A, HSPA1B, HSPA1L, HSPA 2, HSPA4, HSPA5, HSPA6, HSPA8, IFI30, IFNA1, IFNA10, IFNA13, IFNA14, IFNA16, IFNA17, IFNA2, IFNA21, IFNA4, IFNA5, IFNA6, IFNA7, IFNA8, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DL5A, KIR2DS1, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1, KIR3DL2, KIR3DL3, KLRC1, KLRC2, KLRC3, KLRC4, KLRD1, LGMN, LTA, NFYA, NFYB, NFYC, PDIA3, PSME1, PSME2, PSME3, RFX5, RFXANK, RFXAP, TAP1, TAP2, and TAPBP.
[0248] In some embodiments, the "graft-versus-host disease" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: CD28, CD80, CD86, FAS, FASLG, GZMB, HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA -DQA2, HLA-DQB1, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, IFNG, IL1A , IL1B, IL2, IL6, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL5A, KIR3DL1, KIR3DL2, KLRC1, KLRD1, PRF1, and TNF.
[0249] In some embodiments, the "ILC family development and heterogeneity" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: AHR, AREG, BCL11B, EOMES, GATA3, GFI1, HNF1A, ID2, IFNG, IL12A, IL12B, IL13, IL15, IL17A, IL18, IL1B, IL22, IL23A, IL25, IL33, IL4, IL5, IL6, IL7, IL9, NFIL3, RORA, TBX21, TNF, TOX, TSLP, and ZBTB16.
[0250] In some embodiments, the "CD8 TCR downstream pathway" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: B2M, BRAF, CD247, CD3D, CD3E, CD3G, CD8A, CD8B, EGR1, EGR4, ELK1, EOMES, FASLG, FOS, FOSL1, GZMB, HLA-A, HRAS, IFNA1, IFNA10, IFNA14, IFNA16, IFNA17, IFNA2, IFNA21, IFNA4, IFNA5, IFNA6, IFNA7, IFNA 8, IFNAR1, IFNAR2, IFNG, IL2, IL2RA, IL2RB, IL2RG, JUN, JUNB, KRAS, MAP2K1, MAP2K2, MAPK1, MAPK3, MAPK8, MAPK9, NFATC1, NFATC2, NFATC3, NRAS, PPP3CA, PPP3CB, PPP3R1, PRF1, PRKCA, PRKCB, PRKCE, PRKCQ, PTPN7, RAF1, STAT4, TNF, TNFRSF18, TNFRSF4, and TNFRSF9.
[0251] In some embodiments, the "NFAT TF pathway" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: BATF3, CASP3, CBLB, CD40LG, CDK4, CSF2, CTLA4, CXCL8, DGKA, E2F1, EGR1, EGR2, EGR3, EGR4, FASLG, FOS, FOSL1, FOXP3, GATA3, GBP3, IFNG, IKZF1, IL2, IL2RA, IL3, IL4, IL5, IRF4, ITCH, JUN, JUNB, MAF, NFATC1, NFATC2, NFATC3, POU2F1, PPARG, PRKCQ, PTGS2, PTPN1, PTPRK, RNF128, SLC3A2, TBX21, and TNF.
[0252] In some embodiments, the "cancer immunotherapy with PD1 blockade" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: NFKB1, PTPN11, PDCD1, NFATC1, STAT3, NFATC2, HLA-DRB1, BATF, NFAT5, IFNG, HLA-A, CD274, PDCD1LG2, ZAP70, NFATC3, NFATC4, CD8A, CD3D, LCK, CD8B, CD3E, JUN, and CD3G.
[0253] In some embodiments, the "CTL pathway" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: B2M, CD247, CD3D, CD3E, CD3G, FAS, FASLG, GZMB, HLA-A, ICAM1, ITGAL, ITGB2, and PRF1.
[0254] In some embodiments, the "allograft rejection" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: AARS1, ABCE1, ABI1, ACHE, ACVR2A, AKT1, APBB1, B2M, BCAT1, BCL10, BCL3, BRCA1, C2, CAPG, CARTPT, CCL11, CCL13, CCL19, CCL2, CCL22, CCL4, CCL5, CCL7, CCND2, CCND3, CCR1, CCR2, CCR5, CD1D, CD2, CD247, CD28, CD3D, CD3E, CD3G , CD4, CD40, CD40LG, CD47, CD7, CD74, CD79A, CD80, CD86, CD8A, CD8B, CD96, CDKN2A, CFP, CRTAM, CSF1, CSK, CTSS, CXCL13, CXCL9, CXCR3, DARS1, DEGS1, D YRK3, EGFR, EIF3A, EIF3D, EIF3J, EIF4G3, EIF5A, ELANE, ELF4, EREG, ETS1, F2, F2R, FAS, FASLG, FCGR2B, FGR, FLNA, FYB1, GALNT1, GBP2, GCNT1, GLMN, GP R65, GZMA, GZMB, HCLS1, HDAC9, HIF1A, HLA-A, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DQA1, HLA-DRA, HLA-E, HLA-G, ICAM1, ICOSLG, IFNAR2, IFNG, I FNGR1, IFNGR2, IGSF6, IKBKB, IL10, IL11, IL12A, IL12B, IL12RB1, IL13, IL15, IL16, IL18, IL18RAP, IL1B, IL2, IL27RA, IL2RA, IL2RB, IL2RG, IL4, IL4R , IL6, IL7, IL9, INHBA, INHBB, IRF4, IRF7, IRF8, ITGAL, ITGB2, ITK, JAK2, KLRD1, KRT1, LCK, LCP2, LIF, LTB, LY75, LY86, LYN, MAP3K7, MAP4K1, MBL2, MMP 9, MRPL3, MTIF2, NCF4, NCK1, NCR1, NLRP3, NME1, NOS2, NPM1, PF4, PRF1, PRKCB, PRKCG, PSMB10, PTPN6, PTPRC, RARS1, RIPK2, RPL39, RPL3L, RPL9, RPS19,RPS3A, RPS9, SIT1, SOCS1, SOCS5, SPI1, SRGN, ST8SIA4, STAB1, STAT1, STAT4, TAP1, TAP2, TAPBP, TGFB1, TG FB2, THY1, TIMP1, TLR1, TLR2, TLR3, TLR6, TNF, TPD52, TRAF2, TRAT1, UBE2D1, UBE2N, WARS1, WAS, and ZAP70. ,
[0255] In some embodiments, the "IL12 2 pathway" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: ATF2, B2M, CCL3, CCL4, CCR5, CD247, CD3D, CD3E, CD3G, CD4, CD8A, CD8B, EOMES, FASLG, FOS, GADD45B, GADD45G, GZMA, GZMB, HLA-A, HLA-DRA, HLX, IFNG, IL12A, IL12B, IL12RB1, IL12RB2, IL18 , IL18R1, IL18RAP, IL1B, IL1R1, IL2, IL2RA, IL2RB, IL2RG, IL4, JAK2, LCK, MAP2K3, MAP2K6, MAPK14, MTOR, NFKB1, NFKB2, NO S2, PPP3CA, PPP3CB, PPP3R1, RAB7A, RELA, RELB, RIPK2, SOCS1, SPHK2, STAT1, STAT3, STAT4, STAT5A, STAT6, TBX21, and TYK2.
[0256] In some embodiments, the "neutrophil degranulation" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: A1BG, ABCA13, ACAA1, ACLY, ACP3, ACTR10, ACTR1B, ACTR2, ADA2, ADAM10, ADAM8, ADGRE3, ADGRE5, ADGRG3, AGA, AGL, AGPAT2, AHSG, ALAD, ALDH3B1, ALDOA, ALDOC, ALOX5, AMPD3, ANO6, ANPEP, ANXA2, AOC1, AP1M1, AP2A2, APAF1, APEH, APRT, ARG1, ARHGAP45, ARHGAP9, ARL8A, ARMC8, ARPC5, ARSA, ARSB, ASAH1, ATAD3B, ATG7, ATP11A, ATP11B, ATP6AP2, ATP6V0A1, ATP6V0C, ATP6V1D, ATP8A 1, ATP8B4, AZU1, B2M, B4GALT1, BIN2, BPI, BRI3, BST1, BST2, C1orf35, C3, C3AR1, C5AR1, C6orf120, CAB39, CALML5, CAMP, CAND1, CANT1, CAP1, CAPN1, CA T, CCT2, CCT8, CD14, CD177, CD300A, CD33, CD36, CD44, CD47, CD53, CD55, CD58, CD59, CD63, CD68, CD93, CDA, CDK13, CEACAM1, CEACAM3, CEACAM6, CEACAM 8, CEP290, CFD, CFP, CHI3L1, CHIT1, CHRNB4, CKAP4, CLEC12A, CLEC4C, CLEC4D, CLEC5A, CMTM6, CNN2, COMMD3, COMMD9, COPB1, COTL1, CPNE1, CPNE3, CPPE D1, CR1, CRACR2A, CREG1, CRISP3, CRISPLD2, CSNK2B, CST3, CSTB, CTSA, CTSB, CTSC, CTSD, CTSG, CTSH, CTSS, CTSZ, CXCL1, CXCR1, CXCR2, CYB5R3, CYBA, C YBB, CYFIP1, CYSTM1, DBNL, DDOST, DDX3X, DEFA1, DEFA1B, DEFA4, DEGS1, DERA, DGAT1, DIAPH1, DNAJC13, DNAJC3, DNAJC5, DNASE1L1, DOCK2, DOK3, DPP7,DSC1、DSG1、DSN1、DSP、DYNC1H1、DYNC1LI1、DYNLL1、DYNLT1、EEF1A1、EEF2、 ELANE、ENPP4、EPX、ERP44、FABP5、FAF2、FCAR、FCER1G、FCGR2A、FCGR3B、FCN 1、FGL2、FGR、FLG2、FOLR3、FPR1、FPR2、FRK、FRMPD3、FTH1、FTL、FUCA1、FUCA 2、GAA、GALNS、GCA、GDI2、GGH、GHDC、GLA、GLB1、GLIPR1、GM2A、GMFG、GNS、GOL GA7、GPI、GPR84、GRN、GSDMD、GSN、GSTP1、GUSB、GYG1、HBB、HEBP2、HEXB、HGS NAT、HK3、HLA-B、HLA-C、HMGB1、HMOX2、HP、HPSE、HRNR、HSP90AA1、HSP90AB1、 HSPA1A, HSPA1B, HSPA6, HSPA8, HUWE1, HVCN1, IDH1, IGF2R, ILF2, IMPDH1, IMPDH2, IQGAP1, IQGAP2, IRAG2, IST1, ITGAL, ITGAM, ITGAV, ITGAX, ITGB2, JU P、KCMF1、KCNAB2、KPNB1、KRT1、LAIR1、LAMP1、LAMP2、LAMTOR1、LAMTOR2、LA MTOR3、LCN2、LGALS3、LILRA3、LILRB2、LILRB3、LPCAT1、LRG1、LRRC7、LTA4H LTF, LYZ, MAGT1, MAN2B1, MANBA, MAPK1, MAPK14, MCEMP1, METTL7A, MGAM, MGST1, MIF, MLEC, MME, MMP25, MMP8, MMP9, MNDA, MOSPD2, MPO, MS4A3, MVP, NAP RT、NBEAL2、NCKAP1L、NCSTN、NDUFC2、NEU1、NFAM1、NFASC、NFKB1、NHLRC3、N IT2、NME2、NPC2、NRAS、OLFM4、OLR1、ORM1、ORM2、ORMDL3、OSCAR、OSTF1、P2RX 1、PA2G4、PADI2、PAFAH1B2、PDAP1、PDXK、PECAM1、PFKL、PGAM1、PGLYRP1、PG M1、PGM2、PGRMC1、PIGR、PKM、PKP1、PLAC8、PLAU、PLAUR、PLD1、PLEKHO2、PNP、PPBP, PPIA, PPIE, PRCP, PRDX4, PRDX6, PRG2, PRG3, PRKCD, PRSS2, PRSS3, PRTN3, PSAP, PSEN1, PSMA2, PSMA5, PSMB1, PSMB7, PSMC2, PSMC3, PSMD1, PSMD11, PSMD12, PSMD13, PSMD14, PSMD2, PSMD3, PSMD6, PSMD7, PTAFR, PTGES2, PTPN6, PTPRB, PTPRC, PTPRJ, PTPRN2, PTX3, PYCARD, PYGB, PYGL, QPCT, QSOX1, RAB10, RAB14, RAB18, RAB24, RAB27A, RAB31, RAB37, RAB3A, RAB3D, RAB44, RAB4B, RAB5B, RAB5C, RAB6A, RAB7A, RAB9B, RAC1, RAP1A, RAP1B, RAP2B, RAP2C, RETN, RHOA, RHOF, RHOG, RNASE2, RNASE3, RNASET2, ROCK1, S100A11, S100A12, S100A7, S100A8, S100A9, S100P, SCAMP1, SDCBP, SELL, SERPINA1, SERPINA3, SERPINB1, SERPINB10, SERPINB12, SERPINB3, SERPINB6, SIGLEC14, SIGLEC5, SIGLEC9, SIRPA, SIRPB1, SLC11A1, SLC15A4, SLC27A2, SLC2A3, SLC2A5, SLC44A2, SLCO4C1, SLPI, SNAP23, SNAP25, SNAP29, SPTAN1, SRP14, STBD1, STING1, STK10, STK11IP, STOM, SURF4, SVIP, SYNGR1, TARM1, TBC1D10C, TCIRG1, TCN1, TICAM2, TIMP2, TLR2, TMBIM1, TMC6, TMEM179B, TMEM30A, TMEM63A, TNFAIP6, TNFRSF1B, TOLLIP, TOM1, TRAPPC1, TRPM2, TSPAN14, TTR, TUBB, TUBB4B, TXNDC5, TYROBP, UBR4, UNC13D, VAMP8, VAPA, VAT1, VCL, VCP, VNN1, VPS35L, XRCC5, XRCC6, and YPEL5.,
[0257] In some embodiments, the "innate immune system" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: A1BG, AAMP, ABCA13, ABI1, ABI2, ABL1, ACAA1, ACLY, ACP3, ACTB, ACTG1, ACTR10, ACTR1B, ACTR2, ACTR3, ADA2, ADAM10, ADAM8, ADGRE3, ADGRE5, ADGRG3, AGA, AGER, AGL, AGPAT2, AHCYL1, AHSG, AIM2, ALAD, ALDH3B1, ALDOA, ALDOC, ALO X5, ALPK1, AMPD3, ANO6, ANPEP, ANXA2, AOC1, AP1M1, AP2A2, APAF1, APEH, APOB, APP, APRT, ARG1, ARHGAP45, ARHGAP9, ARL8A, ARMC8, ARPC1A, ARPC1B, ARP C2, ARPC3, ARPC4, ARPC5, ARSA, ARSB, ART1, ASAH1, ATAD3B, ATF1, ATF2, ATG12, ATG5, ATG7, ATOX1, ATP11A, ATP11B, ATP6AP2, ATP6V0A1, ATP6V0A2, ATP6 V0A4, ATP6V0B, ATP6V0C, ATP6V0D1, ATP6V0D2, ATP6V0E1, ATP6V0E2, ATP6V1A, ATP6V1B1, ATP6V1B2, ATP6V1C1, ATP6V1C2, ATP6V1D, ATP6V1E1, ATP6V1 E2, ATP6V1F, ATP6V1G1, ATP6V1G2, ATP6V1G3, ATP6V1H, ATP7A, ATP8A1, ATP8B4, AZU1, B2M, B4GALT1, BAIAP2, BCL10, BCL2, BCL2L1, BIN2, BIRC2, BIRC3, BPI, BPIFA1, BPIFA2, BPIFB1, BPIFB2, BPIFB4, BPIFB6, BRI3, BRK1, BST1, BST2, BTK, BTRC, C1orf35, C1QA, C1QB, C1QC, C1R, C1S, C2, C3, C3AR1, C4A, C4B , C4B_2, C4BPA, C4BPB, C5, C5AR1, C5AR2, C6, C6orf120, C7, C8A, C8B, C8G, C9, CAB39, CALM1, CALML5, CAMP, CAND1, CANT1, CAP1, CAPN1, CAPZA1, CAPZA2,CARD11、CARD9、CASP1、CASP10、CASP2、CASP4、CASP8、CASP9、CAT、CCL17、CCL22、CCR2、CCR6、CCT2、CCT8、CD14、CD177、CD180、CD19、CD209、CD247、CD300A、CD300E、CD300LB、CD33、CD36、CD3G、CD4、CD44、CD46、CD47、CD53、CD55、CD58、CD59、CD63、CD68、CD81、CD93、CDA、CDC34、CDC42、CDK13、CEACAM1、CEACAM3、CEACAM6、CEACAM8、CEP290、CFB、CFD、CFH、CFHR1、CFHR2、CFHR3、CFHR4、CFHR5、CFI、CFL1、CFP、CGAS、CHGA、CHI3L1、CHIT1、CHRNB4、CHUK、CKAP4、CLEC10A、CLEC12A、CLEC4A、CLEC4C、CLEC4D、CLEC4E、CLEC5A、CLEC6A、CLEC7A、CLU、CMTM6、CNN2、CNPY3、COLEC10、COLEC11、COMMD3、COMMD9、COPB1、COTL1、CPB2、CPN1、CPN2、CPNE1、CPNE3、CPPED1、CR1、CR2、CRACR2A、CRCP、CREB1、CREBBP、CREG1、CRISP3、CRISPLD2、CRK、CRP、CSNK2B、CST3、CSTB、CTNNB1、CTSA、CTSB、CTSC、CTSD、CTSG、CTSH、CTSK、CTSL、CTSS、CTSV、CTSZ、CUL1、CXCL1、CXCR1、CXCR2、CYB5R3、CYBA、CYBB、CYFIP1、CYFIP2、CYLD、CYSTM1、DBNL、DCD、DDOST、DDX3X、DDX41、DEFA1、DEFA1B、DEFA3、DEFA4、DEFA5、DEFA6、DEFB1、DEFB103A、DEFB103B、DEFB104A、DEFB104B、DEFB105A、DEFB105B、DEFB106A、DEFB106B、DEFB107A、DEFB107B、DEFB108B、DEFB109B、DEFB110、DEFB112、DEFB113、DEFB114、DEFB115、DEFB116、DEFB118、DEFB119、DEFB121、DEFB123、DEFB124、DEFB125、DEFB126、DEFB127、DEFB128、DEFB129、DEFB13 0A、DEFB130B、DEFB131A、DEFB132、DEFB133、DEFB134、DEFB135、DEFB136、D EFB4A, DEFB4B, DEGS1, DERA, DGAT1, DHX36, DHX58, DHX9, DIAPH1, DNAJC13, DNAJC3, DNAJC5, DNASE1L1, DNM1, DNM2, DNM3, DOCK1, DOCK2, DOK3, DPP7, DSC 1、DSG1、DSN1、DSP、DTX4、DUSP3、DUSP4、DUSP6、DUSP7、DYNC1H1、DYNC1LI1、 DYNLL1、DYNLT1、ECSIT、EEA1、EEF1A1、EEF2、ELANE、ELK1、ELMO1、ELMO2、ENP P4、ENSG00000284958、EP300、EPPIN、EPPIN-WFDC6、EPX、ERP44、F2、FABP5、 FADD、FAF2、FBXW11、FCAR、FCER1A、FCER1G、FCGR1A、FCGR2A、FCGR3A、FCGR3B 、FCN1、FCN2、FCN3、FGA、FGB、FGG、FGL2、FGR、FLG2、FOLR3、FOS、FPR1、FPR2、 FRK、FRMPD3、FTH1、FTL、FUCA1、FUCA2、FYN、GAA、GAB2、GALNS、GCA、GDI2、GGH 、GHDC、GLA、GLB1、GLIPR1、GM2A、GMFG、GNLY、GNS、GOLGA7、GPI、GPR84、GRAP 2、GRB2、GRN、GSDMD、GSDME、GSN、GSTP1、GUSB、GYG1、GZMM、HBB、HCK、HEBP2、H ERC5、HEXB、HGSNAT、HK3、HLA-B、HLA-C、HLA-E、HMGB1、HMOX1、HMOX2、HP、HP SE、HRAS、HRNR、HSP90AA1、HSP90AB1、HSP90B1、HSPA1A、HSPA1B、HSPA6、HSPA 8、HTN1、HTN3、HUWE1、HVCN1、ICAM2、ICAM3、IDH1、IFI16、IFIH1、IFNA1、IFN A10、IFNA13、IFNA14、IFNA16、IFNA17、IFNA2、IFNA21、IFNA4、IFNA5、IFNA6、IFNA7、IFNA8、IFNB1、IGF2R、IGHE、IGHG1、IGHG2、IGHG4、IGHV1-2、IGHV1-4 6、IGHV1-69、IGHV2-5、IGHV2-70、IGHV3-11、IGHV3-13、IGHV3-23、IGHV3-3 0、IGHV3-33、IGHV3-48、IGHV3-53、IGHV3-7、IGHV4-34、IGHV4-39、IGHV4-5 9、IGKV1-12、IGKV1-16、IGKV1-17、IGKV1-33、IGKV1-39、IGKV1-5、IGKV1-1D-1 2、IGKV1D-16、IGKV1D-33、IGKV1D-39、IGKV2-28、IGKV2-30、IGKV2D-28、IG KV2D-30、IGKV2D-40、IGKV3-11、IGKV3-15、IGKV3-20、IGKV3D-20、IGKV4-1、 IGKV5-2, IGLC2, IGLC3, IGLV1-40, IGLV1-44, IGLV1-47, IGLV1-51, IGLV2- 11、IGLV2-14、IGLV2-23、IGLV2-8、IGLV3-1、IGLV3-19、IGLV3-21、IGLV3-25 、IGLV3-27、IGLV6-57、IGLV7-43、IKBIP、IKBKB、IKBKE、IKBKG、IL1B、ILF2、 IMPDH1, IMPDH2, IQGAP1, IQGAP2, IRAG2, IRAK1, IRAK2, IRAK3, IRAK4, IRF3 、IRF7、ISG15、IST1、ITCH、ITGAL、ITGAM、ITGAV、ITGAX、ITGB2、ITK、ITLN1、 ITPR1、ITPR2、ITPR3、JUN、JUP、KCMF1、KCNAB2、KIR2DL5A、KIR2DS1、KIR2DS2 、KIR2DS4、KIR2DS5、KIR3DS1、KLRC2、KLRD1、KLRK1、KPNB1、KRAS、KRT1、LAI R1、LAMP1、LAMP2、LAMTOR1、LAMTOR2、LAMTOR3、LAT、LAT2、LBP、LCK、LCN2、LC P2、LEAP2、LGALS3、LGMN、LILRA3、LILRB2、LILRB3、LIMK1、LPCAT1、LPO、LRG 1、LRRC14、LRRC7、LRRFIP1、LTA4H、LTF、LY86、LY96、LYN、LYZ、MAGT1、MALT1、MAN2B1、MANBA、MAP2K1、MAP2K3、MAP2K4、MAP2K6、MAP2K7、MAP3K1、MAP3K14、MAP3K7、MAP3K8、MAPK1、MAPK10、MAPK11、MAPK12、MAPK13、MAPK14、MAPK3、M APK7, MAPK8, MAPK9, MAPKAPK2, MAPKAPK3, MASP1, MASP2, MAVS, MBL2, MCEMP1, MEF2A, MEF2C, MEFV, METTL7A, MGAM, MGST1, MIF, MLEC, MME, MMP25, MMP8, M MP9, MNDA, MOSPD2, MPO, MRE11, MS4A2, MS4A3, MUC1, MUC12, MUC13, MUC15, MUC16, MUC17, MUC20, MUC21, MUC3A, MUC4, MUC5AC, MUC5B, MUC6, MUC7, MUCL1, MVP, MYD88, MYH2, MYH9, MYO10, MYO1C, MYO5A, MYO9B, N4BP1, NAPRT, NBEAL2, NCF1, NCF2, NCF4, NCK1, NCKAP1, NCKAP1L, NCKIPSD, NCR2, NCSTN, NDUFC2, N EU1, NF2, NFAM1, NFASC, NFATC1, NFATC2, NFATC3, NFKB1, NFKB2, NFKBIA, NFKBIB, NHLRC3, NIT2, NKIRAS1, NKIRAS2, NLRC3, NLRC4, NLRC5, NLRP1, NLRP3, NLRP4, NLRX1, NME2, NOD1, NOD2, NOS1, NOS2, NOS3, NPC2, NRAS, OLFM4, OLR1, ORM1, ORM2, ORMDL3, OSCAR, OSTF1, OTUD5, P2RX1, P2RX7, PA2G4, PADI2, PAF AH1B2, PAK1, PAK2, PAK3, PANX1, PCBP2, PDAP1, PDPK1, PDXK, PDZD11, PECAM1, PELI1, PELI2, PELI3, PFKL, PGAM1, PGLYRP1, PGLYRP2, PGLYRP3, PGLYRP4, PGM1, PGM2, PGRMC1, PI3, PIGR, PIK3C3, PIK3CA, PIK3CB, PIK3R1, PIK3R2, PIK3R4, PIN1, PKM, PKP1, PLA2G2A, PLA2G6, PLAC8, PLAU, PLAUR, PLCG1, PLCG2,PLD1, PLD2, PLD3, PLD4, PLEKHO2, PLPP4, PLPP5, PNP, POLR1C, POLR1D, POLR2E, POLR2F, POLR2H, POLR2K, POLR2L, POLR3A, POLR3B, POLR3C, POLR3D, POLR3E, PO, LR3F、POLR3G、POLR3GL、POLR3H、POLR3K、PPBP、PPIA、PPIE、PPP2CA、PPP2CB、PPP2R1A、PPP2R1B、PPP2R5D、PPP3CA、PPP3CB、PPP3R1、PRCP、PRDX4、PRDX6、PRG2、PRG3、PRKACA、PRKACB、PRKACG、PRKCD、PRKCE、PRKCQ、PRKCSH、PRKDC、PROS1、PRSS2、PRSS3、PRTN3、PSAP、PSEN1、PSMA1、PSMA2、PSMA3、PSMA4、PSMA5、PSMA6、PSMA7、PSMA8、PSMB1、PSMB10、PSMB11、PSMB2、PSMB3、PSMB4、PSMB5、PSMB6、PSMB7、PSMB8、PSMB9、PSMC1、PSMC2、PSMC3、PSMC4、PSMC5、PSMC6、PSMD1、PSMD10、PSMD11、PSMD12、PSMD13、PSMD14、PSMD2、PSMD3、PSMD4、PSMD5、PSMD6、PSMD7、PSMD8、PSMD9、PSME1、PSME2、PSME3、PSME4、PSMF1、PSTPIP1、PTAFR、PTGES2、PTK2、PTPN11、PTPN4、PTPN6、PTPRB、PTPRC、PTPRJ、PTPRN2、PTX3、PYCARD、PYGB、PYGL、QPCT、QSOX1、RAB10、RAB14、RAB18、RAB24、RAB27A、RAB31、RAB37、RAB3A、RAB3D、RAB44、RAB4B、RAB5B、RAB5C、RAB6A、RAB7A、RAB9B、RAC1、RAC2、RAF1、RAP1A、RAP1B、RAP2B、RAP2C、RASGRP1、RASGRP2、RASGRP4、RBSN、REG3A、REG3G、RELA、RELB、RETN、RHOA、RHOF、RHOG、RIGI、RIPK1、RIPK2、RIPK3、RNASE2、RNASE3、RNASE6、RNASE7、RNASE8、RNASET2、RNF125、RNF135、RNF216、ROCK1、RPS27A、RPS6KA1、RPS6KA2、RPS6KA3、RPS6KA5、S100A1、S100A11、S100A12、S100A7、S100A7A、S100A8、S100A9、S100B、S100P、SAA1、SARM1、SCAMP1、SDCBP、SELL、SEM1、SEMG1、SERPINA1、SERPINA3、SERPINB1、SERPINB10、SERPINB12、SERPINB3、SERPINB6、SERPING1、SFTPA1、SFT PA2、SFTPD、SHC1、SIGIRR、SIGLEC14、SIGLEC15、SIGLEC5、SIGLEC9、SIKE1、 SIRPA、SIRPB1、SKP1、SLC11A1、SLC15A4、SLC27A2、SLC2A3、SLC2A5、SLC44A 2、SLCO4C1、SLPI、SNAP23、SNAP25、SNAP29、SOCS1、SOS1、SPTAN1、SRC、SRP1 4、STAT6、STBD1、STING1、STK10、STK11IP、STOM、SUGT1、SURF4、SVIP、SYK、S YNGR1, TAB1, TAB2, TAB3, TANK, TARM1, TAX1BP1, TBC1D10C, TBK1, TCIRG1, TCN1, TEC, TICAM1, TICAM2, TIFA, TIMP2, TIRAP, TKFC, TLR1, TLR10, TLR2, TLR 3、TLR4、TLR5、TLR6、TLR7、TLR8、TLR9、TMBIM1、TMC6、TMEM179B、TMEM30A、T MEM63A、TNFAIP3、TNFAIP6、TNFRSF1B、TNIP2、TOLLIP、TOM1、TOMM70、TP53、 TRAF2、TRAF3、TRAF6、TRAPPC1、TREM1、TREM2、TREX1、TRIM21、TRIM25、TRIM 32、TRIM4、TRIM56、TRPM2、TSPAN14、TTR、TUBB、TUBB4B、TXK、TXN、TXNDC5、T XNIP, TYROBP, UBA3, UBA52, UBA7, UBB, UBC, UBE2D1, UBE2D2, UBE2D3, UBE2 、UBE2L6、UBE2M、UBE2N、UBE2V1、UBR4、UNC13D、UNC93B1、USP14、USP18、VAM P8、VAPA、VAT1、VAV1、VAV2、VAV3、VCL、VCP、VNN1、VPS35L、VRK3、VTN、WAS、W ASF1、WASF2、WASF3、WASL、WIPF1、WIPF2、WIPF3、XRCC5、XRCC6、YES1、YPEL5、and ZBP1.
[0258] In some embodiments, the "IL1 family signaling" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: AGER, ALOX5, ALPK1, APP, BTRC, CASP1, CASP8, CHUK, CTSG, CUL1, FBXW11, GSDMD, HMGB1, IKBIP, IKBKB, IKBKG, IL13, IL18, IL18BP, IL18R1, IL18RAP, IL1A, IL1B, IL1F10, IL1R1, IL1R2, IL1RAP, IL1RAPL1, IL1RL1, IL1RL2. , IL1RN, IL33, IL36A, IL36B, IL36G, IL36RN, IL37, IL4, IRAK1, IRAK2, IRAK3, IRAK4, LRRC14, MAP2K1, MAP2K4, MAP2K6, MAP3K3, MAP3K7, MAP3K8, MAPK8, MYD88, N4BP1, NFKB1, NFKB2, NFKBIA, NFKBIB, NKIRAS1, NKIRAS2, NLRC5, NLRX1, NOD1, NOD2, PELI1, PELI2, PELI3, PSMA1, PSMA2, PSMA3, PSMA4, PSMA5, P SMA6, PSMA7, PSMA8, PSMB1, PSMB10, PSMB11, PSMB2, PSMB3, PSMB4, PSMB5, PSMB6, PSMB7, PSMB8, PSMB9, PSMC1, PSMC2, PSMC3, PSMC4, PSMC5, PSMC6, PSM D1, PSMD10, PSMD11, PSMD12, PSMD13, PSMD14, PSMD2, PSMD3, PSMD4, PSMD5, PSMD6, PSMD7, PSMD8, PSMD9, PSME1, PSME2, PSME3, PSME4, PSMF1, PTPN11, PT PN12, PTPN13, PTPN14, PTPN18, PTPN2, PTPN20, PTPN23, PTPN4, PTPN5, PTPN6, PTPN7, PTPN9, RBX1, RELA, RIPK2, RPS27A, S100A12, S100B, SAA1, SEM1, SI GIRR, SKP1, SMAD3, SQSTM1, STAT3, TAB1, TAB2, TAB3, TBK1, TIFA, TNIP2, TOLLIP, TP53, TRAF2, TRAF6, UBA52, UBB, UBC, UBE2N, UBE2V1, USP14, and USP18.
[0259] In some embodiments, the "GPCR signaling" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: ABHD12, ABHD6, ABR, ACKR1, ACKR2, ACKR3, ACKR4, ADCY1, ADCY2, ADCY3, ADCY4, ADCY5, ADCY6, ADCY7, ADCY8, ADCY9, ADCYAP1, ADCYAP1R1, ADGRE1, ADGRE2, ADGRE3, ADGRE5, ADM, ADM2, ADORA1, ADORA2A, ADORA2B, ADORA3, A DRA1A, ADRA1B, ADRA1D, ADRA2A, ADRA2B, ADRA2C, ADRB1, ADRB2, ADRB3, AGT, AGTR1, AGTR2, AHCYL1, AKAP13, AKT1, AKT2, AKT3, ANXA1, APLN, APLNR, APP, ARHGEF1, ARHGEF10, ARHGEF10L, ARHGEF11, ARHGEF12, ARHGEF15, ARHGEF16, ARHGEF17, ARHGEF18, ARHGEF19, ARHGEF2, ARHGEF25, ARHGEF26, ARHGEF3, A RHGEF33, ARHGEF35, ARHGEF37, ARHGEF38, ARHGEF39, ARHGEF4, ARHGEF40, ARHGEF5, ARHGEF6, ARHGEF7, ARHGEF9, ARRB1, ARRB2, AVP, AVPR1A, AVPR1B, AV PR2, BDKRB1, BDKRB2, BRS3, BTK, C3, C3AR1, C5, C5AR1, C5AR2, CALCA, CALCB, CALCR, CALCRL, CALM1, CAMK2A, CAMK2B, CAMK2D, CAMK2G, CAMK4, CAMKK1, CA MKK2, CASR, CCK, CCKAR, CCKBR, CCL1, CCL11, CCL13, CCL16, CCL17, CCL19, CCL2, CCL20, CCL21, CCL22, CCL23, CCL25, CCL27, CCL28, CCL3, CCL3L1, CCL3L 3, CCL4, CCL4L2, CCL5, CCL7, CCR1, CCR10, CCR2, CCR3, CCR4, CCR5, CCR6, CCR7, CCR8, CCR9, CCRL2, CD55, CDC42, CDK5, CGA, CHRM1, CHRM2, CHRM3, CHRM4,CHRM5、CMKLR1、CNR1、CNR2、CORT、CREB1、CRH、CRHBP、CRHR1、CRHR2、CX3CL1、CX3CR1、CXCL1、CXCL10、CXCL11、CXCL12、CXCL13、CXCL16、CXCL2、CXCL3、CXCL5、CXCL6、CXCL8、CXCL9、CXCR1、CXCR2、CXCR3、CXCR4、CXCR5、CXCR6、CYSLTR1、CYSLTR2、DAGLA、DAGLB、DGKA、DGKB、DGKD、DGKE、DGKG、DGKH、DGKI、DGKK、DGKQ、DGKZ、DHH、DRD1、DRD2、DRD3、DRD4、DRD5、ECE1、ECE2、ECT2、EDN1、EDN2、EDN3、EDNRA、EDNRB、EGFR、F2、F2R、F2RL1、F2RL2、F2RL3、FFAR1、FFAR2、FFAR3、FFAR4、FGD1、FGD2、FGD3、FGD4、FN1、FPR1、FPR2、FPR3、FSHB、FSHR、FZD1、FZD10、FZD2、FZD3、FZD4、FZD5、FZD6、FZD7、FZD8、FZD9、GABBR1、GABBR2、GAL、GALR1、GALR2、GALR3、GAST、GCG、GCGR、GHRH、GHRHR、GHRL、GHSR、GIP、GIPR、GLP1R、GLP2R、GNA11、GNA12、GNA13、GNA14、GNA15、GNAI1、GNAI2、GNAI3、GNAL、GNAQ、GNAS、GNAT1、GNAT2、GNAT3、GNAZ、GNB1、GNB2、GNB3、GNB4、GNB5、GNG10、GNG11、GNG12、GNG13、GNG2、GNG3、GNG4、GNG5、GNG7、GNG8、GNGT1、GNGT2、GNRH1、GNRH2、GNRHR、GPBAR1、GPER1、GPHA2、GPHB5、GPR132、GPR143、GPR15、GPR150、GPR17、GPR176、GPR18、GPR183、GPR20、GPR25、GPR27、GPR31、GPR32、GPR35、GPR37、GPR37L1、GPR39、GPR4、GPR45、GPR55、GPR65、GPR68、GPR83、GPR84、GPRC6A、GPSM1、GPSM2、GPSM3、GRB2、GRK2、GRK3、GRK5、GRK6、GRM1、GRM2、GRM3、GRM4、GRM5、GRM6、GRM7、GRM8、GRP、GRPR、HBEGF、HCAR1、HCAR2、HCAR3、HCRT、HCRTR1、HCRTR2、HEBP1、HRAS、HRH1、HRH2、HRH3、HRH4、HTR1A、HTR1B、HTR1D、HTR1E、HTR1F、HTR2A、HTR2B、HTR2C、HTR4、HTR5A、HTR6、HTR7、IAPP、IHH、INSL3、INSL5、ITGA5、ITGB1、ITPR1、ITPR2、ITPR3、ITSN1、KALRN、KEL、KISS1、KISS1R、KNG1、KPNA2、KRAS、LHB、LHCGR、LPAR1、LPAR2、LPAR3、LPAR4、LPAR5、LPAR6、LTB4R、LTB4R2、MAPK1、MAPK3、MAPK7、MC1R、MC2R、MC3R、MC4R、MC5R、MCF2、MCF2L、MCHR1、MCHR2、MGLL、MLN、MLNR、MMP3、MTNR1A、MTNR1B、NBEA、NET1、NGEF、NLN、NMB、NMBR、NMS、NMU、NMUR1、NMUR2、NPB、NPBWR1、NPBWR2、NPFF、NPFFR1、NPFFR2、NPS、NPSR1、NPW、NPY、NPY1R、NPY2R、NPY4R、NPY5R、NRAS、NTS、NTSR1、NTSR2、OBSCN、OPN1LW、OPN1MW、OPN1SW、OPN3、OPN4、OPN5、OPRD1、OPRK1、OPRL1、OPRM1、OXER1、OXGR1、OXT、OXTR、P2RY1、P2RY10、P2RY11、P2RY12、P2RY13、P2RY14、P2RY2、P2RY4、P2RY6、PAK1、PCP2、PDE10A、PDE11A、PDE1A、PDE1B、PDE1C、PDE2A、PDE3A、PDE3B、PDE4A、PDE4B、PDE4C、PDE4D、PDE7A、PDE7B、PDE8A、PDE8B、PDPK1、PDYN、PENK、PF4、PIK3CA、PIK3CG、PIK3R1、PIK3R2、PIK3R3、PIK3R5、PIK3R6、PLA2G4A、PLCB1、PLCB2、PLCB3、PLCB4、PLEKHG2、PLEKHG5、PLPPR1、PLPPR2、PLPPR3、PLPPR4、PLPPR5、PLXNB1、PMCH、PNOC、POMC、PPBP、PPP1CA、PPP1R1B、PPP2CA、PPP2CB、PPP2R1A、PPP2R1B、PPP2R5D、PPP3CA、PPP3CB、PPP3CC、PPP3R1、PPY、PREX1、PRKACA、PRKACB、PRKACG、PRKAR1A、PRKAR1B、PRKAR2A、PRKAR2B、PRKCA、PRKCB、PRKCD、PRKCE、PRKCG、PRKCH、PRKCQ、PRKX、PRLH、PRLHR、PROK1、PROK2、PROKR1、PROKR2、PSAP、PTAFR、PTCH1、PTCH2、PTGDR、PTGDR2、PTGER1、PTGER2、PTGER3、PTGER4、PTGFR、PTGIR、PTH、PTH1R、PTH2、PTH2R、PTHLH、PYY、QRFP、QRFPR、RAMP1、RAMP2、RAMP3、RASGRF2、RASGRP1、RASGRP2、RGR、RGS1、RGS10、RGS11、RGS12、RGS13、RGS14、RGS16、RGS17、RGS18、RGS19、RGS2、RGS20、RGS21、RGS22、RGS3、RGS4、RGS5、RGS6、RGS7、RGS8、RGS9、RGSL1、RHO、RHOA、RHOB、RHOC、RLN2、RLN3、ROCK1、ROCK2、RPS6KA1、RPS6KA2、RPS6KA3、RRH、RXFP1、RXFP2、RXFP3、RXFP4、S1PR1、S1PR2、S1PR3、S1PR4、S1PR5、SAA1、SCT、SCTR、SHC1、SHH、SMO、SOS1、SOS2、SRC、SST、SSTR1、SSTR2、SSTR3、SSTR4、SSTR5、SUCNR1、TAAR1、TAAR2、TAAR5、TAAR6、TAAR8、TAAR9、TAC1、TAC3、TACR1、TACR2、TACR3、TAS1R1、TAS1R2、TAS1R3、TAS2R1、TAS2R10、TAS2R13、TAS2R14、TAS2R16、TAS2R19、TAS2R20、TAS2R3、TAS2R30、TAS2R31、TAS2R38、TAS2R39、TAS2R4、TAS2R40、TAS2R41、TAS2R42、TAS2R43、TAS2R46、TAS2R5、TAS2R50、TAS2R60、TAS2R7, TAS2R8, TAS2R9, TBXA2R, TIAM1, TIAM2, TRH, TRHR, TRIO, TRPC3, TRPC6, TRPC7, TSHB, TSHR, UCN, UCN2, UCN3, UTS2, UTS2B, UTS2R, VAV1, VAV2, VAV3, VIP, VIPR1, VIPR2, WNT1, WNT10A, WNT10B, WNT11, WNT16, WNT2, WNT2B, WNT3, WNT3A, WNT4, WNT5A, WNT6, WNT7A, WNT7B, WNT8A, WNT8B, WNT9A, WNT9B, XCL1, XCL2, XCR1, and XK.
[0260] In some embodiments, the "receptor tyrosine kinase signaling" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: AAMP, ABI1, ABI2, ACTB, ACTG1, ADAM10, ADAM12, ADAM17, ADAP1, ADCYAP1, ADCYAP1R1, ADORA2A, AHCYL1, AKT1, AKT2, AKT3, ALK, ANOS1, AP2A1, AP2A2, AP2B1, AP2M1, AP2S1, APH1A, APH1B, APOE, ARC, AREG, ARF6, AR HGEF7, ASCL1, ATF1, ATF2, ATP6AP1, ATP6V0A1, ATP6V0A2, ATP6V0A4, ATP6V0B, ATP6V0C, ATP6V0D1, ATP6V0D2, ATP6V0E1, ATP6V0E2, ATP6V1A, ATP6V1B1 , ATP6V1B2, ATP6V1C1, ATP6V1C2, ATP6V1D, ATP6V1E1, ATP6V1E2, ATP6V1F, ATP6V1G1, ATP6V1G2, ATP6V1G3, ATP6V1H, AXL, BAIAP2, BAX, BCAR1, BDNF, BR AF, BRK1, BTC, CALM1, CAV1, CBL, CD274, CDC37, CDC42, CDH5, CDK5, CDK5R1, CDK5R2, CHD4, CHEK1, CILP, CLTA, CLTC, CMA1, COL11A1, COL11A2, COL1A1, CO L1A2, COL24A1, COL27A1, COL2A1, COL3A1, COL4A1, COL4A2, COL4A3, COL4A4, COL4A5, COL5A1, COL5A2, COL5A3, COL6A1, COL6A2, COL6A3, COL6A5, COL6A6 , COL9A1, COL9A2, COL9A3, CREB1, CRK, CRKL, CSK, CSN2, CTNNA1, CTNNB1, CTNND1, CTSD, CUL5, CXCL12, CYBA, CYBB, CYFIP1, CYFIP2, DIAPH1, DLG4, DNAL4 , DNM1, DNM2, DNM3, DOCK1, DOCK3, DOCK7, DUSP3, DUSP4, DUSP6, DUSP7, EGF, EGFR, EGR1, EGR2, EGR3, EGR4, ELK1, ELMO1, ELMO2, EP300, EPGN, EPN1, EPS15,EPS15L1、ERBB2、ERBB3、ERBB4、ERBIN、EREG、ESR1、ESRP1、ESRP2、F3、FER、F ES、FGF1、FGF10、FGF16、FGF17、FGF18、FGF19、FGF2、FGF20、FGF22、FGF23、FG F3、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFBP1、FGFBP2、FGFBP3、FGFR1、FG FR2、FGFR3、FGFR4、FGFRL1、FLRT1、FLRT2、FLRT3、FLT1、FLT3、FLT3LG、FLT4、 FN1、FOS、FOSB、FOSL1、FRS2、FRS3、フューリン、FYN、GAB1、GAB2、GABRA1、GABRB1 、GABRB2、GABRB3、GABRG2、GABRG3、GABRQ、GALNT3、GFAP、GGA3、GIPC1、GRAP、 GRAP2、GRB10、GRB2、GRB7、GRIN2B、GTF2F1、GTF2F2、HBEGF、HDAC1、HDAC2、H DAC3、HGF、HGFAC、HGS、HIF1A、HNRNPA1、HNRNPF、HNRNPH1、HNRNPM、HPN、HRAS 、HSP90AA1、HSPB1、ID1、ID2、ID3、ID4、IDE、IGF1、IGF1R、IGF2、IL2RG、INS、 INSR、IRS1、IRS2、IRS4、ITCH、ITGA2、ITGA3、ITGAV、ITGB1、ITGB3、ITPR1、IT PR2, ITPR3, JAK2, JAK3, JUNB, JUND, JUP, KDR, KIDINS220, KIT, KITLG, KL, KLB, KRAS, LAMA1, LAMA2, LAMA3, LAMA4, LAMA5, LAMB1, LAMB2, LAMB3, LAMC1, L AMC2、LAMC3、LCK、LRIG1、LYL1、LYN、MAP2K1、MAP2K2、MAP2K5、MAPK1、MAPK1 1、MAPK12、MAPK13、MAPK14、MAPK3、MAPK7、MAPKAP1、MAPKAPK2、MAPKAPK3、MA TK、MDK、MEF2A、MEF2C、MEF2D、MEMO1、MET、MKNK1、MLST8、MMP9、MST1、MST1R 、MTOR、MUC20、MXD4、MYC、MYCN、NAB1、NAB2、NCBP1、NCBP2、NCF1、NCF2、NCF4、NCK1、NCK2、NCKAP1、NCKAP1L、NCOR1、NCSTN、NEDD4、NELFB、NGF、NOS3、NRAS、NRG1、NRG2、NRG3、NRG4、NRP1、NRP2、NTF3、NTF4、NTRK1、NTRK2、NTRK3、PAG1、PAK1、PAK2、PAK3、PCSK5、PCSK6、PDE3B、PDGFA、PDGFB、PDGFC、PDGFD、PDGFRA、PDGFRB、PDPK1、PGF、PGR、PIK3C3、PIK3CA、PIK3CB、PIK3R1、PIK3R2、PIK3R3、PIK3R4、PLAT、PLCG1、PLG、POLR2A、POLR2B、POLR2C、POLR2D、POLR2E、POLR2F、POLR2G、POLR2H、POLR2I、POLR2J、POLR2K、POLR2L、PPP2CA、PPP2CB、PPP2R1A、PPP2R1B、PPP2R5D、PRDM1、PRKACA、PRKACB、PRKACG、PRKCA、PRKCB、PRKCD、PRKCE、PRKCZ、PRR5、PSEN1、PSEN2、PSENEN、PTBP1、PTK2、PTK2B、PTK6、PTN、PTPN1、PTPN11、PTPN12、PTPN18、PTPN2、PTPN3、PTPN6、PTPRF、PTPRJ、PTPRK、PTPRO、PTPRS、PTPRU、PTPRZ1、PXN、RAB4A、RAB4B、RAC1、RALA、RALB、RALGDS、RANBP10、RANBP9、RAP1A、RAP1B、RAPGEF1、RASA1、RBFOX2、REST、RHOA、RICTOR、RIT1、RIT2、RNF41、ROCK1、ROCK2、RPS27A、RPS6KA1、RPS6KA2、RPS6KA3、RPS6KA5、RRAD、S100B、SGK1、SH2B2、SH2B3、SH2D2A、SH3GL1、SH3GL2、SH3GL3、SH3KBP1、SHB、SHC1、SHC2、SHC3、SIN3A、SOCS1、SOCS6、SOS1、SPARC、SPHK1、SPINT1、SPINT2、SPP1、SPRED1、SPRED2、SPRY1、SPRY2、SRC、SRF、STAM、STAM2、STAT1、STAT3、STAT5A、STAT5B、STAT6、STMN1、STUB1、TAB2、TCF12、TCIRG1, TEC, TGFA, TGFBR3, THBS1, THBS2, THBS3, THBS4, THEM4, TIA1, TIAL1, TIAM1, TLR9, TNS3, TNS4, TPH1, TRIB1, TRIB3, UBA52, UBB, UBC, USP8, VAV1, VAV2, VAV3, VEGFA, VEGFB, VEGFC, VEGFD, VGF, VRK3, WASF1, WASF2, WASF3, WWOX, WW1, YAP1, YES1, and YWHAB.
[0261] In some embodiments, the "elevated KRAS signaling" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: ABCB1, ACE, ADAM17, ADAM8, ADAMDEC1, ADGRA2, ADGRL4, AKAP12, AKT2, ALDH1A2, ALDH1A3, AMMECR1, ANGPTL4, ANKH, ANO1, ANXA10, APOD, ARG1, ATG10, AVL9, BIRC3, BMP2, BPGM, BTBD3, BTC, C3AR1, CA2, CAB39L, CBL, CBR4, CB X8, CCL20, CCND2, CCSER2, CD37, CDADC1, CFB, CFH, CFHR2, CIDEA, CLEC4A, CMKLR1, CPE, CROT, CSF2, CSF2RA, CTSS, CXCL10, CXCR4, DCBLD2, DNMBP, DOCK 2, DUSP6, EMP1, ENG, EPB41L3, EPHB2, EREG, ERO1A, ETS1, ETV1, ETV4, ETV5, EVI5, F13A1, F2RL1, FBXO4, FCER1G, FGF9, FLT4, FUCA1, G0S2, GABRA3, GADD4 5G, GALNT3, GFPT2, GLRX, GNG11, GPNMB, GPRC5B, GUCY1A1, GYPC, H2BC3, HBEGF, HDAC9, HKDC1, HOXD11, HSD11B1, ID2, IGF2, IGFBP3, IKZF1, IL10RA, IL1 B, IL1RL2, IL2RG, IL33, IL7R, INHBA, IRF8, ITGA2, ITGB2, ITGBL1, JUP, KCNN4, KIF5C, KLF4, LAPTM5, LAT2, LCP1, LIF, LY96, MAFB, MALL, MAP3K1, MAP4K1 , MAP7, MMD, MMP10, MMP11, MMP9, MPZL2, MTMR10, MYCN, NAP1L2, NGF, NIN, NR0B2, NR1H4, NRP1, PCP4, PCSK1N, PDCD1LG2, PECAM1, PEG3, PIGR, PLAT, PLAU , PLAUR, PLEK2, PLVAP, PPBP, PPP1R15A, PRDM1, PRELID3B, PRKG2, PRRX1, PSMB8, PTBP2, PTCD2, PTGS2, PTPRR, RABGAP1L, RBM4, RBP4, RELN, RETN, RGS16,SATB1, SCG3, SCG5, SCN1B, SDCCAG8, SEMA3B, SERPINA3, SLPI, SNAP25, SNAP91, SOX9, SPARCL1, SPON1, SPP1, SPRY2, ST6GAL1, STRN, TFPI, TLR8, TMEM100, TMEM158, TMEM176A, TMEM176B, TNFAIP3, TNFRSF1B, TNNT2, TOR1AIP2, TPH1, TRAF1, TRIB1, TRIB2, TSPAN1, TSPAN13, TSPAN7, USH1C, USP12, VWA5A, WDR33, WNT7A, YRDC, ZNF277, and ZNF639.
[0262] In some embodiments, the "negative regulation of the PI3K AKT network" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: AKT1, AKT2, AKT3, AREG, BTC, CD19, CD28, CD80, CD86, EGF, EGFR, EPGN, ERBB2, ERBB3, ERBB4, EREG, ESR1, ESR2, FGF1, FGF10, FGF16, FGF17, FGF18, FGF19, FGF2, FGF20, FGF22, FGF23, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGFR1, FGFR2, FGFR3, FGFR4, F LT3, FLT3LG, FRS2, FYN, GAB1, GAB2, GRB2, HBEGF, HGF, ICOS, IER3, IL1RAP, IL1RL1, IL33, INS, INSR, IRAK1, I RAK4, IRS1, IRS2, KIT, KITLG, KL, KLB, LCK, MAPK1, MAPK3, MET, MYD88, NRG1, NRG2, NRG3, NRG4, PDGFA, PDGFB , PDGFRA, PDGFRB, PHLPP1, PHLPP2, PIK3AP1, PIK3CA, PIK3CB, PIK3CD, PIK3R1, PIK3R2, PIK3R3, PIP4K2A, PIP 4K2B, PIP4K2C, PIP5K1A, PIP5K1B, PIP5K1C, PPP2CA, PPP2CB, PPP2R1A, PPP2R1B, PPP2R5A, PPP2R5B, PPP2R5 C, PPP2R5D, PPP2R5E, PTEN, PTPN11, RAC1, RAC2, RHOG, SRC, STRN, TGFA, THEM4, TRAF6, TRAT1, TRIB3, and VAV1.
[0263] In some embodiments, the "VEGFR1 2 pathway" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: AKAP1, AKT1, ARF1, BRAF, CAMKK2, CAV1, CBL, CDC42, CDH5, CTNNA1, CTNNB1, DNM2, FBXW11, FES, FLT1, FYN, GAB1, GRB10, GRB2, HGS, HSP90AA1, HSP90AB1, IQGAP1, ITGAV, ITGB3, KDR, MAP2K1, MAP2K2, MAP2K3, MAP2K6, MAPK 1, MAPK11, MAPK14, MAPK3, MAPKAPK2, MYOF, NCK1, NCK2, NEDD4, NOS3, PAK2, PDPK1, PIK3CA, PIK3R1, PLCG1, PRKAA1, PRKAA2, PRKAB1, PRKAC A, PRKAG1, PRKCA, PRKCB, PRKCD, PTK2, PTK2B, PTPN11, PTPN2, PTPN6, PTPRJ, PXN, RAF1, RHOA, ROCK1, SH2D2A, SHB, SRC, VCL, VEGFA, and VTN.
[0264] In some embodiments, the "naive CD8 T cells vs. PD-1 high CD8 T cells" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: ACSS2, ACTN1, ADGRA3, ADPRM, AEBP1, AGBL2, AGMAT, AK5, AMIGO1, APBB1, ARHGEF4, ARMCX1, ATG9B, ATP6V0E2-AS1, AZIN2, BDH1, BEND5, BEX3, BPHL, C17orf67, C19orf18, C3orf18, CA6, CAPN5, CARS1, CATSPERE, CBR3, CCR7, CD248, CD55, CE P170, CEP41, CHCHD7, CHMP7, CLEC11A, CLN5, CLTRN, CNKSR1, CNKSR2, CT75, CYP2J2, DCHS1, DENND5A, DSC1, ECRG4, EDAR, EFHC2, EFHD1, EFNA1, EIF2 D, ENSG00000280119, EPB41L2, EPHA1, EPHA1-AS1, FAM117B, FAM184A, FAM216A, FBLN2, FBP1, FBXO15, FLNB, FOXO1, FOXP1, GAL3ST4, GIPC3, GNG7, G P5, GPRASP2, HAPLN3, HPCAL4, HSBP1L1, IGF1R, IL6R, IL6ST, IPCEF1, IRS1, ITGA6, KLF7, KLHL6, KRTCAP3, LDLRAP1, LEF1, LEF1-AS1, LINS1, LMF1, L RRN3, MAL, MAML2, MAN1C1, MCF2L-AS1, MDS2, MEST, MICU3, MMEL1, MRRF, MYB, NAA16, NAT9, NDFIP1, NELL2, NEXMIF, NOG, NR3C2, NRCAM, NREP, NT5E, N UDT9P1, OBSCN, OVGP1, OXNAD1, PABPC3, PASK, PCSK5, PDCD4-AS1, PDE9A, PDK1, PIK3IP1, PKIA, PKIG, PLAG1, PLEKHG4, PLPP1, PRKCA, PRKCQ-AS1, PR RT1, PRXL2A, RAB43, RBM26-AS1, REG4, RETREG1, RFX2, RHPN2, RNF157, RNF175, ROBO3, SALL2, SARAF, SCML1, SCML2, SCOC-AS1, SELL, SERP1, SFXN4,SFXN5, SH3RF3, SH3YL1, SLC16A10, SLC22A17, SLC7A3, SNED1, SNHG32, SOX8, SPART, SPEG, SPEN-AS1, SPINK2, SPINT2, SREBF1, STRADB, STXBP1, SULT1B1, SUSD3, TAF4B, TBXA2R, TCEAL3, TCF3, TECPR1, THEM4, TKTL1, TMEM220, TMEM272, TNFRSF10D, TOP1MT, TP73-AS1, TPST1, TRABD2A, TSEN2, TXNRD3, UBE2E2, UBIAD1, UBQLN2, USP51, USP6NL, VIPR1, YPEL2, ZBTB10, ZBTB18, ZNF285, ZNF436-AS1, ZNF496, ZNF662, ZNF667-AS1, and ZNF93.,
[0265] In some embodiments, the "naive vs. activated CD8 T cell" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: ACADL, ACOT7, ACP1, ACTG1, AFG3L1P, AIF1, AKIP1, ALCAM, ALYREF, ANAPC15, ANXA1, ANXA4, APOBEC2, ASF1B, AURKA, BEX3, BSPRY, BUB3, C4orf3, CAB39L, CALM1, CALM3, CARHSP1, CCDC34, CD44, CD48, CD80, CDK2AP1, CDKN1A, CDKN2C, CENPP, C HST11, CISD1, CLIC1, COMMD3, COPS4, COPS5, CRELD2, CTLA4, CXCL10, CXCR3, DAPK2, DBI, DCK, DDOST, DDX39A, DEPDC1, DLAT, DPAGT1, DSCC1, DUSP5, E2F7, EME1, EMP1, EPAS1, ERG28, ERH, ETFB, FAM136A, FBXO5, FCGRT, FIGNL1, FKBP2, FLNB, GABARAPL1, GCNT1, GEM, GGH, GLRX, GPR160, GRB7, GSAP, H1 -1, H2BC4, H4C8, HIRIP3, HMGN2, HOPX, HPRT1, ID2, IDI1, IFITM1, IFNGR1, IL1B, IL1R2, INSL6, IRAK3, ITSN1, KIF22, KLF11, KLRC2, LAG3, LAIR1, LA MP2, LSM12, LSM2, LSM3, MDH2, MICOS10, MIS18BP1, MPHOSPH6, MRPL18, MRPL42, MXD3, MYADM, MYL4, NCALD, NDUFAF2, NDUFS6, NME1, NRP1, NUDT1, NUP3 7, NUP43, NUP54, ORC6, PANX1, PBK, PGAM1, PHF11, PLAC8, PMAIP1, PMM1, PNP, POLR3K, PPA1, PRELID1, PRF1, PRIM2, PSMA1, PSMA5, PSMB2, PSMC3IP, PS MD8, PYCARD, RAD18, RAD51AP1, RAN, RANBP1, RBBP7, RBM47, RFC3, RGS1, RPA2, SAMSN1, SAR1B, SCRN3, SELENOS, SEPHS2, SERPINB9, SERPINE2, SF3B6,SIVA1, SMIM3, SNRPA1, SNX10, SPDL1, SURF4, SYCE2, SYPL1, SYTL3, TAF12, TBCB, TCEAL9, TEX15, TEX30, TFDP1, TIMM17A, TIMM2 3, TMBIM4, TMED10, TMEM163, TROAP, TTC39B, TTC9C, TUBB4B, TXNDC17, UBE2N, UBE2S, UCK2, UFC1, VDAC3, VIM, YBX3, and ZBTB32. ,
[0266] In some embodiments, the "PD-1 signaling" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: CD247, CD274, CD3D, CD3E, CD3G, CD4, CSK, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, LCK, PDCD1, PDCD1LG2, PTPN11, PTPN6, TRAV19, TRAV29DV5, TRAV8-4, TRBV12-3, and TRBV7-9.
[0267] In some embodiments, the "cancer immunotherapy with PD1 blockade" signature comprises gene expression scores (e.g., GSEA scores) for the following genes: NFKB1, PTPN11, PDCD1, NFATC1, STAT3, NFATC2, HLA-DRB1, BATF, NFAT5, IFNG, HLA-A, CD274, PDCD1LG2, ZAP70, NFATC3, NFATC4, CD8A, CD3D, LCK, CD8B, CD3E, JUN, and CD3G.
[0268] In some embodiments, leukocyte immune profile types are characterized according to cytokine expression profiles. Table 12 of this example further describes additional characteristics of leukocyte immune profile types, for example, by cell type enrichment, functional significance, and / or T cell receptor (TCR) repertoire.
[0269] In some embodiments, the present disclosure provides methods for identifying a subject having, suspected of having, or at risk of having cancer as having, or likely to have, a favorable prognosis (e.g., as measured by overall survival (OS) or progression-free survival (PFS)). A favorable prognosis may refer to subjects with a first immune profile type associated with a decreased risk of cancer progression, an increased chance of responding to a therapeutic agent, and / or an increased life expectancy compared to subjects with a different leukocyte immune profile type. For example, in some embodiments, subjects with primed-type HNSCC are predicted to respond better to immunotherapy (e.g., a PD1 inhibitor) compared to subjects with HNSCC of a different immune profile type.
[0270] In some embodiments, the methods include determining the subject's leukocyte immune profile type as described herein.
[0271] In some embodiments, the methods include identifying a subject as having a reduced risk of cancer progression relative to subjects with a different leukocyte immune profile type. In some embodiments, a "reduced risk of cancer progression" may indicate that the subject has a more favorable cancer prognosis or a reduced likelihood of having advanced disease. In some embodiments, a "reduced risk of cancer progression" may indicate that a subject with cancer is expected to have a higher response rate to certain treatments. For example, a "reduced risk of cancer progression" may indicate that the subject is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% more likely to experience a progression-free survival event (e.g., recurrence, retreatment, or death) relative to another cancer patient or population of cancer patients (e.g., patients with cancer, but who do not have the same cancer leukocyte immune profile type as the subject).
[0272] In some embodiments, the method further includes identifying the subject as having an increased risk of cancer progression relative to other leukocyte immune profile types. In some embodiments, an "increased risk of cancer progression" may indicate that the subject is increased likelihood of having a less positive cancer prognosis or advanced disease. In some embodiments, an "increased risk of cancer progression" may indicate that a subject with cancer is less likely to respond to or is not responsive to certain treatments, and shows only modest or no improvement in disease symptoms. For example, an "increased risk of cancer progression" may indicate that the subject is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% more likely to experience a progression-free survival event (e.g., recurrence, retreatment, or death) relative to another cancer patient or cancer patient population (e.g., patients with cancer, but not with the same leukocyte immune profile type as the subject).
[0273] In some embodiments, the methods described herein involve performing the determination using at least one computer hardware processor.
[0274] In some embodiments, the present disclosure provides a method of providing a prognosis, predicting survival, or stratifying patient risk in a subject suspected of having or at risk of having cancer, hi some embodiments, the method comprises determining the subject's leukocyte immune profile type as described herein.
[0275] Updated leukocyte immune profile types based on new data Described herein are techniques for generating a leukocyte immune profile type. It should be understood that clusters may be updated as additional signatures are calculated for a patient. In some embodiments, the subject's leukocyte signature is one of a threshold number of leukocyte signatures for a threshold number of subjects. In some embodiments, once the threshold number of leukocyte signatures are generated, the leukocyte immune profile type is updated. For example, once a threshold number of new leukocyte signatures are obtained (e.g., 1 new signature, 10 new signatures, 100 new signatures, 500 new signatures, or any suitable threshold number of signatures ranging from 10 to 1,000 signatures), the new signatures may be combined with the leukocyte signatures previously used to generate the leukocyte immune profile type, and the combined set of old and new leukocyte signatures may be re-clustered (e.g., using any of the clustering algorithms described herein or any other suitable clustering algorithm) to obtain an updated set of leukocyte immune profile types.
[0276] In this way, data obtained from future patients can be analyzed in a way that utilizes information learned from patients whose leukocyte signatures have been calculated prior to the future patient. In this sense, the machine learning techniques described herein (e.g., unsupervised clustering machine learning methods) are adaptive and learn as new patient data accumulates. This can facilitate improved characterization of the types of leukocyte immune profiles future patients may have, which can improve treatment selection for those patients.
[0277] Treatment indications Aspects of the present disclosure relate to methods of identifying or selecting a therapeutic agent for a subject based on determining the subject's leukocyte immune profile type. The present disclosure is based, in part, on the recognition that subjects with certain leukocyte immune profile types (e.g., naive immune profile type, primed immune profile type) have an increased likelihood of responding to certain therapies (e.g., immunotherapeutic agents) compared to subjects with other leukocyte immune profile types (e.g., suppressed). In some embodiments, subjects with a suppressed leukocyte immune profile type are not selected for immunotherapy. In some embodiments, subjects with a suppressed leukocyte immune profile type are administered a treatment that is not an immunotherapy.
[0278] In some embodiments, the therapeutic agent is an immuno-oncology (IO) agent. The IO agent can be a small molecule, a peptide, a protein (e.g., an antibody such as a monoclonal antibody), an interfering nucleic acid, or a combination of any of the foregoing. In some embodiments, the IO agent comprises a PD1 inhibitor, a PD-L1 inhibitor, or a PD-L2 inhibitor. Examples of IO agents include, but are not limited to, cemiplimab, nivolumab, pembrolizumab, avelumab, durvalumab, atezolizumab, BMS1166, BMS202, and the like. In some embodiments, the IO agent comprises a combination of atezolizumab and albumin-bound paclitaxel, pembrolizumab and albumin-bound paclitaxel, pembrolizumab and paclitaxel, or pembrolizumab, gemcitabine, and carboplatin.
[0279] In some embodiments, the methods described by the present disclosure further include administering one or more therapeutic agents to the subject based on the determination of the subject's leukocyte immune profile type, hi some embodiments, the subject is administered one or more (e.g., 1, 2, 3, 4, 5, or more) IO agents.
[0280] Aspects of the present disclosure relate to methods of treating a subject having (or suspected of or at risk of having) cancer based on determining the subject's leukocyte immune profile type. In some embodiments, the method includes administering one or more (e.g., 1, 2, 3, 4, 5, or more) therapeutic agents to the subject. In some embodiments, the therapeutic agent(s) administered to the subject is selected from small molecules, peptides, nucleic acids, radioisotopes, cells (e.g., CAR T cells, etc.), and combinations thereof. Examples of therapeutic agents include chemotherapy (e.g., cytotoxic agents, etc.), immunotherapy (e.g., immune checkpoint inhibitors such as PD-1 inhibitors, PD-L1 inhibitors, etc.), antibodies (e.g., anti-HER2 antibodies), cell therapy (e.g., CAR T cell therapy), gene silencing therapy (e.g., interfering RNA, CRISPR, etc.), antibody-drug conjugates (ADCs), and combinations thereof.
[0281] In some embodiments, the present disclosure relates to methods of treating a subject having (or suspected of or at risk of having) head and neck squamous cell carcinoma (HNSCC) based on determining the subject's leukocyte immune profile type. For example, subjects with HNSCC and a primed leukocyte immune profile type may exhibit a higher response rate to immunotherapy (e.g., an immune checkpoint inhibitor, e.g., a PD-1 blocking antibody such as nivolumab) compared to subjects with HNSCC and a different immune profile type (e.g., naive, advanced, chronic, or suppressed). In some embodiments, subjects with HNSCC and a chronic or suppressed immune type may exhibit a lower response rate to immunotherapy compared to subjects with HNSCC and a primed immune type.
[0282] In some embodiments, a subject is administered an effective amount of a therapeutic agent. As used herein, "effective amount" refers to the amount of each active agent, either alone or in combination with one or more other active agents, required to confer a therapeutic effect on a subject. As will be recognized by those skilled in the art, the effective amount will vary depending on the specific condition being treated, the severity of the condition, the individual patient's parameters, including age, health, size, sex, and weight, the duration of treatment, the nature of concomitant therapy (if any), the specific route of administration, and similar factors within the knowledge and expertise of a healthcare practitioner. These factors are well known to those skilled in the art and can be addressed with no more than routine experimentation. Generally, it is preferable to use the maximum dose of each component or combination thereof, i.e., the highest safe dose according to sound medical judgment. However, those skilled in the art will understand that a patient may require a lower or tolerable dose for medical, psychological, or virtually any other reason.
[0283] Empirical considerations, such as the half-life of the therapeutic compound, generally contribute to determining the dosage. For example, to prolong the half-life of an antibody and to prevent the antibody from being attacked by the host's immune system, an antibody compatible with the human immune system, such as a humanized antibody or a fully human antibody, may be used. The frequency of administration may be determined and adjusted during the course of therapy, generally (but not necessarily) based on the treatment and / or suppression and / or improvement and / or delay of cancer. Alternatively, a continuous sustained-release formulation of the anti-cancer therapeutic agent may be appropriate. Various formulations and devices for achieving sustained release are known in the art.
[0284] In some embodiments, dosage may be empirically determined in individuals who have received one or more doses of the anti-cancer therapeutic agent. Individuals may be administered increasing dosages of the anti-cancer therapeutic agent. To determine the effectiveness of the administered anti-cancer therapeutic agent, one or more aspects of the cancer (e.g., leukocyte immune profile type, tumor microenvironment, tumor formation, tumor growth, etc.) may be analyzed.
[0285] Generally, the initial dosage candidate for any administration of the anti-cancer antibodies described herein can be about 2 mg / kg. For purposes of this disclosure, a typical daily dosage can range anywhere from about 0.1 μg / kg to 3 μg / kg to 30 μg / kg to 300 μg / kg to 3 mg / kg, to 30 mg / kg to 100 mg / kg, or more, depending on the factors discussed above. For repeated administration over several days or more, treatment is continued until a desired suppression or improvement of symptoms occurs, or until a therapeutic level sufficient to alleviate the cancer or one or more symptoms thereof is achieved, depending on the condition. An exemplary titration regimen involves administering an initial dose of about 2 mg / kg, followed by a maintenance dose of about 1 mg / kg of antibody weekly, or a maintenance dose of about 1 mg / kg every other week. However, other dosage regimens may be useful depending on the pharmacokinetic decay pattern the practitioner (e.g., physician) desires to achieve. For example, titration from one to four times weekly is contemplated. In some embodiments, a dose ranging from about 3 μg / mg to about 2 mg / kg (e.g., about 3 μg / mg, about 10 μg / mg, about 30 μg / mg, about 100 μg / mg, about 300 μg / mg, about 1 mg / kg, and about 2 mg / kg) may be used. In some embodiments, the dose titration frequency is once every 1, 2, 4, 5, 6, 7, 8, 9, or 10 weeks; or once every 1, 2, or 3 months, or more. The progress of this therapy may be monitored by conventional techniques and assays and / or by monitoring leukocyte immune profile types as described herein. The dose titration regimen (including the treatments used) may vary over time.
[0286] Dosing of immuno-oncology agents is well known, as described, for example, by Louedec et al. Vaccines (Basel). 2020 Dec; 8(4): 632. For example, dosages for pembrolizumab include 200 mg administered every 3 weeks or 400 mg every 6 weeks, e.g., by 30-minute infusion.
[0287] When the anti-cancer therapeutic agent is not an antibody, it may be administered at a rate of about 0.1 to 300 mg / kg of patient body weight in 1 to 3 divided doses, or as disclosed herein. In some embodiments, for a normal weight adult patient, a dose ranging from about 0.3 to 5.00 mg / kg may be administered. The particular dosage regimen, e.g., dose, timing, and / or repetition, will depend on the specific subject and their individual medical history, as well as the characteristics of the individual agent (such as the half-life of the agent and other considerations well known in the art).
[0288] For purposes of this disclosure, appropriate dosages of anti-cancer therapeutics will depend on the particular anti-cancer therapeutic(s) (or compositions thereof) utilized, the type and severity of the cancer, whether the anti-cancer therapeutic is administered for prophylactic or therapeutic purposes, previous therapy, the patient's medical history and response to the anti-cancer therapeutic, and the discretion of the attending physician. Typically, a clinician will administer an anti-cancer therapeutic, such as an antibody, until a dosage is reached that achieves the desired result.
[0289] Administration of anti-cancer therapeutic agents can be continuous or intermittent, depending, for example, on the physiological condition of the recipient, whether the purpose of administration is therapeutic or prophylactic, and other factors known to those skilled in the art. Administration of anti-cancer therapeutic agents (e.g., anti-cancer antibodies) can be essentially continuous over a preselected period of time, or can be in a series of spaced doses, for example, either before, during, or after the onset of cancer.
[0290] As used herein, the term "treating" refers to the application or administration of a composition containing one or more active agents to a subject having cancer, a symptom of cancer, or a predisposition to cancer, with the intent to eradicate, cure, alleviate, palliate, alter, repair, ameliorate, reverse, or affect the cancer or one or more symptoms of cancer, or the predisposition to cancer.
[0291] Alleviating cancer includes delaying the onset or progression of the disease or reducing the severity of the disease. Alleviating the disease does not necessarily require a curative result. As used herein, "delaying" the onset of a disease (e.g., cancer) means delaying, preventing, slowing, slowing, stabilizing, and / or postponing the progression of the disease. This delay may be of varying lengths of time, depending on the course of the disease and / or individual being treated. A method that "delays" or alleviates the onset of a disease, or a method that delays the occurrence of a disease, is a method that reduces the likelihood of one or more symptoms of the disease developing in a given time frame and / or reduces the severity of symptoms in a given time frame, compared to the absence of the method. Such comparisons are typically based on clinical trials using a sufficient number of subjects to produce statistically significant results.
[0292] "Onset" or "progression" of a disease refers to the first symptoms and / or subsequent progression of the disease. Onset of a disease can be detected and determined using clinical techniques known in the art. Alternatively, or in addition to clinical techniques known in the art, onset of a disease can be detectable and determined based on other criteria. However, onset also refers to progression, which may be undetectable. For purposes of this disclosure, onset or progression refers to the biological course of a symptom. "Onset" includes onset, recurrence, and onset. As used herein, "onset" or "onset" of cancer includes initial onset and / or recurrence.
[0293] Examples of immunotherapies include, but are not limited to, PD-1 or PD-L1 inhibitors, CTLA-4 inhibitors, adoptive cell transfer, therapeutic cancer vaccines, oncolytic virotherapy, T-cell therapy, and immune checkpoint inhibitors.
[0294] In some embodiments, the present disclosure provides methods of treating cancer, the method comprising administering one or more therapeutic agents (e.g., one or more anti-cancer agents, such as one or more immunotherapeutic agents) to a subject identified as having a particular leukocyte immune profile type, wherein the subject's leukocyte immune profile type was identified by a method as described by the present disclosure. [Example]
[0295] Example Recent advances in immunotherapy demonstrate the need for a better understanding of the immune system of individual cancer patients and how it impacts cancer treatment response. In one representative example, we describe the development of an immune profiling platform that uses peripheral immune cell heterogeneity to stratify patients into different categories or immune types by assessing characteristics present in the blood of cancer patients and monitor disease progression and treatment response. Therefore, we used multiparameter flow cytometry to develop a unique diagnostic immune profiling assay and analytical framework based on the analysis of leukocytes in peripheral blood.
[0296] Supervised manual gating analysis of flow cytometry data from a cohort of 50 healthy donors identified 415 cell types and immune activation states, which were used to train and subsequently independently validate a machine learning model to automatically identify immune cell subsets from raw cytometry data. Similarly, a cohort of 650 patients was also analyzed by flow cytometry. Applying this tool to peripheral blood (e.g., WBC) samples from a mixed cohort of 299 healthy donors and 323 cancer patients, we developed a machine learning classification model capable of distinguishing between these two groups with 91% accuracy (ROC-AUC). Further refinement of this model using spectral clustering with bootstrapping revealed five clusters characterized by specific physiological immune profiles, namely immune types: (1) naive T and B lymphocytes, (2) Tregs and various CD4+ T helper cell subsets, (3) mature NK, CD8+ transitional memory and PD1+ TIGIT+ CD8+ T cells, (4) terminally differentiated effector memory and TEMRA CD4 and CD8+ T cells, and (5) myeloid cells such as monocytes and neutrophils.
[0297] Few healthy donors were assigned to cluster 1 or 5. These profiles were further validated using the cellular deconvolution algorithm Kassandra with corresponding RNA-seq, and differential gene expression analysis revealed immune type-specific signatures consistent with potential immune responses. Patients in the terminally differentiated CD8+ T cell cluster had a narrower range of HLA types than other clusters, and TCR repertoire analysis indicated significantly increased clonality and decreased clonotype diversity. Within this cluster, there was a high degree of overlap between peripheral blood TCR sequences and tumors, suggesting a relationship between peripheral blood immune types and tumor infiltration.
[0298] Example 1 The immune system plays an important role in protecting organisms from various diseases, including cancer. However, sometimes immunity fails to prevent tumor development. Furthermore, immune cells can even support malignant growths that are part of the tumor microenvironment. Most immune cell populations are also present in the blood and can be analyzed after collection as readily available biopsies. Blood sampling is a largely non-invasive procedure that can obtain human immune cells. This representative example outlines analyses performed on blood samples collected from both cancer patients and healthy donors.
[0299] A total of 621 blood samples were collected: 299 from healthy donors, 221 from patients with epithelial cancer, and 101 from patients with sarcoma. A second cohort was also analyzed—a total of 850 blood samples were collected: 408 from healthy donors, 309 from patients with epithelial cancer, and 133 from patients with sarcoma. Samples were subjected to crosslinking multipanel flow cytometry (FC) analysis and a hematology analyzer. RNA sequencing was also performed on many of the samples (Figures 5A-5B). As a result, cohorts with multiple blood cell population percentages (e.g., cell types listed in Table 1) were generated. For many of the blood samples from cancer patients, corresponding RNA-seq data from tumor biopsies was available. For the RNA-seq data, expression values calculated in TPM format were available for approximately 20,000 genes.
[0300] First, we analyzed the flow cytometry data using classical dimensionality reduction methods such as PCA, tSNE, and uMAP. Spectral clustering analysis was then performed on the data. The best cluster stability was observed when the number of clusters in the spectral clustering algorithm was equal to 5. Uneven distribution of healthy donor and cancer patient samples was observed among these clusters (see Figures 6A-6B). Some of the identified clusters consisted of blood samples with different cell population ratios, as described below:
[0301] Cluster 1 (myeloid-derived suppressor / NK cell cluster; also referred to in this example as "monocyte" or "G1" or "suppressor"): This cluster is characterized by an increased number of myeloid cell populations, including classical monocytes and neutrophils, compared to the other clusters.
[0302] Cluster 2 (terminally differentiated CD8+ T cell cluster; also referred to in this example as "CD8 T cells" or "G2" or "chronic"): This cluster is characterized by increased numbers of CD8 memory and effector cells and NKT cell populations compared to the other clusters.
[0303] Cluster 3 (mixed CD4+ T helper cell cluster; also referred to in this example as "CD4 T cells" or "G3" or "progression"): This cluster is characterized by an increased number of T helper memory cells, including CD4 central memory, compared to the other clusters.
[0304] Cluster 4 (CD4+ Th1 and CD8+ T cell memory cluster; also referred to in this example as "CD4 / CD8 T cells" or "G4" or "primed"): This cluster is characterized by an increased number of CD4 and CD8 memory cells and a large increase in CD8 transitional memory cells compared to the other clusters.
[0305] Cluster 5 (naive T and B lymphocyte cluster; also referred to in this example as "G5" or "naive"): This cluster is characterized by an increased number of naive CD4, CD8 and B cells compared to the other clusters.
[0306] These clusters can also be described statistically as shown in Tables 5-7 below, which show the 25%, 50% (median), and 75% quantiles for each of the five clusters for each cell type.
[0307] [Table 5-1]
[0308] [Table 5-2]
[0309] [Table 6-1]
[0310] [Table 6-2]
[0311] [Table 7-1]
[0312] [Table 7-2]
[0313] To verify these observations, we analyzed the corresponding RNA-seq data of blood samples belonging to different clusters (for samples with RNA-seq data). The RNA-seq data were processed using the BostonGene Kassandra Cell Deconvolution Tool (e.g., as described in PCT / US2021 / 022155, published September 16, 2021, as International Publication No. 2021 / 183917; and PCT / US2022 / 027088, published November 3, 2022, as International Publication No. 2022 / 232615, the entire contents of each of which are incorporated herein by reference). The data indicated that the flow cytometry analysis results were consistent with the RNA-seq-based cellular composition (Figures 7A-7B). Gene signature and differential cytokine expression analysis were also performed (Figure 8).
[0314] T cell receptor and B cell receptor (TCR / BCR) analysis was also performed on the blood RNA-seq sample data. As expected, cluster 2, described above, which is enriched for effector CD8 and CD4 cells, had lower TCR diversity than the other clusters. Interestingly, the overlap between TCR clonotypes between tumor and blood samples was also high in this cluster (Figures 9A-B).
[0315] Example 2 Figure 10A shows a schematic depicting a representative exemplary immune profiling pipeline. Peripheral blood samples from 442 cancer patients with a variety of different diagnoses and 408 healthy donors were collected from multiple centers. White blood cells (WBCs) were isolated, stained with a custom antibody panel in 96-well plates, and processed by multiparameter flow cytometry (n=850). Each panel was manually labeled, and then the percentages of cell populations (e.g., cell types listed in Table 2) were determined. A machine learning model was developed to classify healthy and cancer groups and refined to stratify immune profiles. Figure 10B shows representative data for a cytometry panel of cell populations shown as a heatmap of normalized signal intensity and tSNE of immune cell populations.
[0316] Supervised manual gating analysis of flow cytometry data from a cohort of 50 healthy donors identified 415 cell types. Analysis of additional cancer samples led to the identification of 650 cell types and immune activation states, which were used to train and independently validate a machine learning (ML) model to automatically identify immune cell subsets from the raw cytometry data. Using the maximum relevance-minimum redundancy (MRMR) algorithm with stepwise leave-one-out cross-validation, the most significantly different cell populations between healthy donors and cancer patients were identified. From this flow cytometry data, 20 significant features that distinguished between healthy donors and cancer patients were selected. In another analysis, we used the Boruta feature selection algorithm (see, e.g., M Kursa and W. Rudnicki, “Feature Selection with the Boruta Package”, Journal of Statistical Software, vol. 36, issue 11, 2010) to select 78 significant features that distinguish between healthy donors and cancer patients, and further refined the random forest model using spectral clustering with bootstrapping to identify immune profiles, and measured cluster stability with the Jaccard index metric.
[0317] The developed machine learning classification model can distinguish between healthy individuals and cancer patients from flow cytometry analysis of peripheral blood samples (Figures 11A-11D).
[0318] We then analyzed the flow cytometry data using spectral clustering to classify immune cell heterogeneity in an individual's peripheral blood into five leukocyte immune profile types, each characterized by a specific physiological immune program, and supported this with transcriptome analysis. A brief description of the clusters is as follows:
[0319] Cluster 1 (naive T and B lymphocyte cluster or "naive" cluster; also referred to in this example as "G1"): This cluster is characterized by an increased number of naive CD4, CD8 and B cells compared to the other clusters.
[0320] Cluster 2 (CD4+ T cell cluster or "primed" cluster; also referred to in this example as "G2"): This cluster is characterized by an increased number of T helper memory cells, including CD4 central memory cells, compared to the other clusters.
[0321] Cluster 3 (CD4+ CD8+ T cell cluster or "progression" cluster; also referred to in this example as "G3"): This cluster is characterized by increased numbers of CD4 and CD8 memory cells, increased dendritic cells, increased NK cells, and a large increase in CD8 transitional memory cells compared to the other clusters.
[0322] Cluster 4 (CD8+ T cell cluster or "chronic" cluster; also referred to in this example as "G4"): This cluster is characterized by increased numbers of CD8 memory, CD45RA-reexpressing effector memory cells (TEMRA), and increased numbers of effector and NKT cell populations compared to the other clusters.
[0323] Cluster 5 (myeloid-derived suppressor / NK cell cluster or "suppressor" cluster; also referred to in this example as "G5"): This cluster is characterized by increased numbers of myeloid cell populations, including classical monocytes and neutrophils, compared to the other clusters.
[0324] These clusters can also be described statistically as shown in Tables 8-10 below, which show the 25%, 50% (median), and 75% quantiles for each of the five clusters for each cell type.
[0325] Table 8-1
[0326] Table 8-2
[0327] Table 9-1
[0328] Table 9-2
[0329] Table 10-1
[0330] Table 10-2
[0331] The first cluster, G1, was enriched for naive B and T cell populations; G2 contained CD4 T helper memory subsets and CD4 Tregs; G3 contained CD8 transitional memory T cells, dendritic cells, TIGIT- and PD1-positive CD8 T cells; G4 contained CD4 / CD8 effector and TEMRA cells; and G5 was highly enriched for classical / non-classical monocytes, HLA-DR-low monocytes, and neutrophils. The healthy-to-cancer ratio was lowest in the G1 cluster and highest in G5, supporting its relevance as a signature of an individual's immune status (Figures 12A-12D). Figures 13A-13C show representative data supporting differential expression of cytokine pathways across leukocyte immune profile types. Figures 13A-13B show representative heatmaps demonstrating correlations between functional gene signatures for cytokine-related pathways from the MSigDB database across leukocyte immune profile types G1-G5. The health status of each patient is also shown. Figure 13C shows representative data showing a comparison of differential gene expression levels of the cytokine and chemokine genes FLT3LG, CCL4, CXCL16, CCR7, TGFBR3, and IL1R1 for five leukocyte immune profile types.
[0332] Assessment of T cell receptor (TCR) and B cell receptor (BCR) content of leukocyte immune profile types was also performed. Figures 14A-14C show representative data for TCR and BCR analysis of PBC immune profile types. Figure 14A shows a representative TCR analysis (for both α and β chains) landscape stratified by PMBC immune profile types (G1-G5). Figure 14B shows representative data for TCR β chain clonality and Chao1 index. Figure 14C shows a representative comparison of TCR analysis for blood and tumor RNA-seq data. Analysis of the distribution of HLA alleles of MHC I classes indicated that the G2 (primed) cluster was enriched for CD8+ T cells and showed lower abundance of HLA B compared to other leukocyte immune profile types. Analysis of the TCR landscape indicated that the G2 (primed) cluster contained more samples with a higher proportion of dominant clonotypes for both the α and β chains of the TCR compared to other leukocyte immune profile types.
[0333] Example 3 Recent advances in immune-based cancer therapy demonstrate the need for a greater understanding of the molecular and cellular characteristics of each individual cancer patient's immune system. The lack of a comprehensive diagnostic capable of describing a patient's immune system status represents a major barrier to predicting and monitoring response to immunotherapy. Here, we developed a clinical immune profiling platform to characterize immune cell heterogeneity in the peripheral blood of healthy donors and patients with solid tumors. We selected robust cell populations differentially represented in these two groups to train a machine learning (ML)-based classifier and used unsupervised clustering to identify groups or immune types with putative functional significance. Using flow cytometry, we identified five immune types corresponding to immune response states characterized by predominant cellular differentiation patterns. These observations were cross-validated using bulk RNA sequencing and T cell repertoire analysis, revealing conserved physiological states that can be easily interrogated from a single blood draw.
[0334] Human populations are genetically and developmentally diverse, with immune systems shaped by a unique set of immunological challenges, including microbial exposure, metabolic changes, chronic diseases such as cancer, and aging. With regard to cancer, each patient's immune system poses ongoing challenges to their responses and can critically inform how cancer patients will respond to various therapies, including immunotherapy. The success of immune checkpoint blockade (ICB) in cancer is complicated by the failure of many patients to respond. These treatments also result in immune-related adverse events (irAEs), which often lead to serious, lifelong complications. Existing biomarkers derived from tumor biopsy evaluation, such as PD-L1 expression by immunohistochemistry, microsatellite instability (MSI), DNA mismatch repair alterations (dMMR), and tumor mutation burden (TMB), have only modestly improved response rates. Detailed analysis of the tumor microenvironment (TME), including analysis of tumor-infiltrating T cells, inflammatory and immunosuppressive cell types, and a variety of different tissue microdomains, has improved positive predictive value over these consensus biomarkers. However, beyond the TME, the contribution of the patient's immune system is not factored into response prediction.
[0335] Surprisingly, there is no consensus on how to assess immune status. Many techniques applied to this problem favor reductionist approaches that consider individual cell populations one at a time, but such methods often involve technical and natural variability that introduces confounding and bias. However, unbiased techniques such as single-cell RNA sequencing are difficult to apply to large cohorts or individual patients in clinical settings due to their low throughput, high coefficient of variation, and high cost per sample.
[0336] This example describes a pan-cancer framework developed for patient stratification using a flow cytometry-based immune profiling assay, using real-world samples from 408 healthy donors and 442 patients with solid tumors. Results demonstrate that comprehensive characterization of the immune system in peripheral blood can be achieved with flow cytometry. Machine learning (ML) techniques were used to develop a classification model capable of discriminating between healthy donors and patients with solid tumors with high accuracy. Unsupervised clustering was used to identify five distinctive immune types, each characterized by a distinct distribution of immune cell types and activation states, as supported by paired bulk RNA-seq analysis. Analysis of over 18,000 transcriptomes from PBMCs demonstrated that these clusters were highly conserved across different patient populations and diseases. These signatures were validated in a cohort of head and neck squamous cell carcinoma (HNSCC) patients treated with the PD-1 inhibitor nivolumab. In this cohort, objective responses were associated with immune types enriched in central and transitional memory CD4+ T cells, demonstrating the functional significance of this classification. Importantly, these features represent functional meta-signatures that can be targeted to select more effective immunotherapies. Identification of tumor immune profiles, or portraits, through an inexpensive blood test is highly correlated with immunotherapy response and holds promise for effective patient stratification in clinical trials and treatment selection in the clinical setting.
[0337] Briefly, we developed an immune profiling assay to evaluate the immune signature of cancer patients in a comprehensive pan-cancer analysis using conventional flow cytometry on red blood cell (RBC)-dependent white blood cell (WBC) samples (Figure 16A). The cohort represented patients with a broad cross-section of solid tumors and healthy donors (n=850). We established a clinical-grade, end-to-end process, including an absolute quantification step of blood cells using a standard hematology analyzer for complete blood count (CBC) analysis of whole blood (Figure 16A). Consistent with previously published studies, significant differences were observed between healthy donors and cancer patients in terms of absolute numbers of RBCs, platelets, neutrophils, and lymphocytes, while absolute numbers of monocytes were similar (Figure 25).
[0338] To ensure broad coverage of immune cell subpopulations across different immune cell lineages, we developed a collection of 10 overlapping antibody panels linked through a lineage backbone panel for quantification of total CD45+ cells in peripheral blood (Panel CP10-General, Figures 16A-B). Multicolor flow cytometry was performed on isolated WBCs stained with the custom panel. The combination and alignment of these overlapping antibody panels enabled comprehensive cell typing, assigning combinations of cell surface markers to distinct immune groups: NK cells, dendritic cells, monocytes, CD4 and CD8 T cells, nonconventional T cells, and B cells (Figure 16B). Collectively, 650 cell types and immune activation states were distinguished based on surface marker combinations (Figure 16B).
[0339] Manual analysis of cell typing defined extensive cell type hierarchies and subpopulations, which were then used to supervise the training of machine learning (ML) gradient boosting event-type models to identify reproducible immune cell subsets for each panel (Figure 16A). Each panel model accurately detected various immune cell populations and activation states from peripheral blood by reproducing the manual supervised gating analysis in validation tests (F1 score: 0.74-0.95, P4-metric: 0.84-0.97, 10 panels).
[0340] Preliminary comparison of peripheral blood samples from healthy donors and cancer patients revealed several differences in the distribution of immune cell subsets, with notable differences in the frequencies of monocytes, naive, central memory, and terminally differentiated CD4+ and CD8+ T cells (Figure 16C). This preliminary analysis is consistent with previously published reports and indicates that differences in the overall immune cell composition between these two groups may be diagnostically detectable in patients with solid tumors. Within this immune profiling flow cytometry-based platform, we developed an ML-based classifier to distinguish between healthy and cancer groups and refined the model to stratify immune profiles (Figures 16A, 17, and 18).
[0341] To thoroughly analyze the immune status of cancer patients and distinguish features specifically associated with tumor development, independent of patient age, solid tumor type, and administered therapy, we collected peripheral blood samples from 408 healthy donors and 442 cancer patients aged 16–98 years with 84 different solid tumor diagnoses within seven major therapy groups (total n = 850; Figures 17A and 17B, supplemental cohort). Significant differences have previously been demonstrated between peripheral blood samples from healthy donors and cancer patients, as well as their immune cell composition analyzed by flow cytometry (Figures 25 and 16C). To accurately distinguish the immune cell content in peripheral blood samples from healthy donors from cancer patients, we developed an ML-based classifier.
[0342] We first assessed the distribution of patients within the cohort with respect to donor age, diagnosis, and therapy using uniform manifold approximation projection (UMAP) for dimensionality reduction of all 650 cytometry populations (Figure 17C). Clusters defined by the presence or absence of cancer in patients (which also corresponded to age, as healthy donors are generally younger than patients with solid tumors) revealed the greatest separation, while clusters based on diagnosis or treatment type were indistinguishable (Figure 17C). Findings indicate that immune features differ between healthy donors and cancer patients, independent of the specific cancer or treatment type.
[0343] Using the maximum relevance-minimum redundancy (MRMR) algorithm with stepwise leave-one-out cross-validation, we selected 20 cell populations that were significantly different between healthy donors and cancer patients (Figure 17D, populations showing differential distributions). Interestingly, these over- and under-represented cell populations included naive CD4+ and CD8+ T cells, naive and memory B cells, and CD8+ TemRA and CD14+ classical monocytes, respectively (Figure 17E), which were highly significant between these two groups. Using the 20 selected populations for a set of 503 samples as features, we trained a TabPFN-based (Hollman et al. 2022) classifier model. The labels predicted by the classifier model (healthy or cancer) corresponded to the true labels, as observed in the slope of the UMAP healthy-cancer distribution for the selected populations (Figure 17F).
[0344] The classifier demonstrated high performance in classifying healthy and cancer classes, as assessed using leave-one-out cross-validation on the training dataset (area under the receiver operating characteristic curve (AUC-ROC) = 0.91). The classifier model outperformed a simpler "standard" model featuring common high-level populations derived from a standard clinical cytometry panel (BD Multitest™ 6-color TBNK; e.g., as described by Omana-Zapata et al. PLoS One 2019 Jan 28;14(1):e0211207) combined with major populations identified using CBC (basophils, eosinophils, neutrophils, monocytes, NK cells, NKT cells, B cells, CD4 T cells, and CD8 T cells) (AUC-ROC = 0.81, Figure 17G, left). Similarly, the healthy / cancer classifier model performed better than the "standard" panel in classifying healthy donors and cancer patients in a validation subset of 347 patient samples (AUC = 0.84 and 0.77, respectively, Figure 17G, right). Overall, cellular immune profiling can discriminate between healthy donors and patients with solid tumors based solely on the composition of white blood cell populations in peripheral blood. Figure 27A additionally shows an exemplary training workflow schema for the healthy / cancer classifier model. Figure 27B shows the high correlation between model cell type predictions and manual markup.
[0345] Next, we focused on uncovering the most robust immune features characterizing the heterogeneity of the cohort and, consequently, identifying functional immune signatures that reflect physiological states associated with overall disease response, rather than transient features corresponding to specific diagnoses or treatments. Unsupervised spectral clustering was applied to the normalized frequencies of selected cell types from 34 cases obtained by flow cytometry to reveal immunologically distinct phenotypes. Immune cell types were selected from a hierarchical tree consistent with the Kassandra algorithm's cell deconvolution (Figure 23B) and from cell populations previously identified as immunotherapy response biomarkers in studies across a variety of cancer types: TIGIT+ PD1+ CD8 T cells, Vδ2+ γ-δ T cells, CD39+ Tregs, and HLA-DR-low monocytes.
[0346] We identified five distinct functional immune types, G1–G5. G1-naive was characterized by high frequencies of naive CD4+, naive CD8+, and naive B cells. G2-primed showed higher proportions of differentiated CD4+ central and transitional memory T cells and CD39+ regulatory T cells (Tregs). G3-progressive included increased frequencies of mature NK cells, CD8 transitional memory, and PD1+ TIGIT+ CD8+ T cells. G4-chronic was enriched for NKT and terminally differentiated effector memory CD45RA+ (TemRA) and CD45RA- (TemRA) T cells, both CD4+ and CD8+ T cells. Finally, the G5-suppressive cluster was highly enriched for classical monocytes, HLA-DR-low monocytes, and neutrophils, with a reduced content of lymphoid cell populations (Figure 18A). Consistent with the UMAP analysis of the entire cohort (Fig. 17C), the presence or absence of cancer diagnosis (healthy or cancer) and patient age were unequally distributed within the immune type clusters (Fig. 18A). Importantly, clusters enriched for terminally differentiated CD8 T cells (G4) and classical monocytes (G5) contained few healthy donors and a high proportion of cancer patients, whereas the G1 group, with the highest proportion of naive T and B lymphocytes, contained the highest proportion of healthy donors (Fig. 18A, top). Additionally, analyses were performed to determine the relationship between gene expression and TME clusters (Fig. 27C).
[0347] To analytically validate the immune groups, we compared RNA-seq-based cellular deconvolution with G1–G5 immune types clustered from flow cytometry data. Consistent with published results by Zaitsev et al. (PMID: 35944503), Kassandra algorithm cellular deconvolution, quantifying cell population frequencies derived from bulk RNA-seq of paired samples (n = 797, supplemental cohort), showed high concordance with frequencies obtained by flow cytometry (Figures 18A and 18B). Figure 27D shows a representative Kassandra algorithm deconvolution heatmap with labels based on flow cytometry clustering. Figure 27E further demonstrates the agreement between cytometry-based quantification of cell population frequency distributions among immune types and the corresponding predictions made by the deconvolution algorithm.
[0348] To demonstrate immune type-associated gene expression profiles, we selected the 200 most differentially expressed genes from each cluster and performed gene set enrichment analysis (GSEA) using curated functional gene signatures from MsigDB for immunologically relevant pathways. G1 and G2 were significantly enriched for signatures of TCF and LEFCTNNB1 transcriptional regulation, TCR, and WNT / β-catenin signaling. G4 was enriched for genes associated with cytotoxic effector T cell responses, and G5 contained multiple pathways associated with innate and myeloid cells (Figure 18C). Comparison of individual gene expression levels of cytokine and chemokine signaling-related genes across immune types revealed expression patterns: FLT3LG and CCR7 expression were highest in G1–G4 compared with low expression in G5; CCL4 and TGFBR3 expression were high in G4, and CXCL16 and IL1R1 were higher in the G5 group (Figure 18D).
[0349] Characterizing the developmental relationships between different T cell lineages allowed us to infer the trajectories of immune type evolution. Using peripheral immune cell composition data obtained from cytometry, we performed pseudotime analysis to establish a developmental hierarchy of these response states (Figure 18E). Patients in G1 and G2, which contain the highest frequencies of naive and central memory CD4+ T cells, clustered most frequently at the origin and terminated at the end of a branch where groups G4 and G5 were most abundant. This analysis indicates that these immune types represent a continuum of functional response states and, therefore, potential responses to subsequent immunological challenges.
[0350] Interestingly, we observed overlapping patterns of gene signatures between the different clusters, consistent with immune cell population distribution by flow cytometry and consistent with pseudotime analysis. For example, G3 appears to be a transitional state containing elements of both G4 and G5. These findings further indicate that these distinct immune type groups are shaped by convergent responses to environmental and immunological stimuli. Accordingly, we associated each immune type group with a functional and developmental state: G1-naive, G2-primed, G3-progressive, G4-chronic, and G5-suppressed. These response characteristics, combined with the fact that these clusters are present in both healthy donors and cancer patients, indicate that these immune types may represent conserved features of immune physiology across different patient populations.
[0351] To further validate our immune typing, we used a multiclass classifier trained on cohort RNA-seq data to stratify cell population percentages from the Kassandra algorithm's cellular deconvolution of bulk RNA-seq samples (n = 18,712, see Open Source Dataset List) collected from the open-source GEO and ArrayExpress databases (Barrett et al., 2012). The open-source validation dataset consisted of whole blood samples from healthy donors and patients with a variety of different diagnoses (>90 types), grouped based on shared features. Using the multiclass classifier, we clustered the samples into five characteristically distinct immune profiles, G1–G5, as seen in three-dimensional PCA projections. Conservation of these immune categories across diverse diseases is noted (Figure 19A).
[0352] Each data set was divided into subgroups primarily based on disease etiology, such that samples from patients with persistent Mycobacterium tuberculosis or Leishmania spp. infections were assigned to the "Intracellular Bacterial and Parasitic Infections" group, and patients with influenza or coronavirus were assigned to the "Acute Respiratory Viral Infections" group (Figure 19B). Consistent with previous findings (Figure 18A), the most frequent immune types in healthy donors were G1 and G2 (Figure 19B). Patients with autoimmune diseases were most commonly classified as naive (G1) immune types. The primed immune type (G2), enriched for central and transitional memory CD4+ T cells, was the most frequent immune type overall across all cohorts and was particularly enriched in patients with "Intracellular Bacterial and Parasitic Infections," consistent with helper T cell responses being essential for phagosomal pathogen clearance. The progressive or transitional immune type (G3) was enriched in samples from patients with viral infections and was most commonly associated with intestinal tissues. Compared with healthy donors and patients with other viral infections, HIV patients were most frequently assigned to the chronic immune type (G4). Interestingly, diseases known to be associated with high systemic inflammation, such as bacterial sepsis, were most frequently classified as the suppressed (G5) subtype, pointing to inflammation-dependent recruitment of monocytes and neutrophils into the blood in these patients (Figure 19B). Overall, a correlation was observed between patients assigned to the naive (G1) and suppressed (G5) immune types (Figure 19C). Analysis of open-source data using deconvolution of bulk RNA-seq data demonstrated that immune types are associated with a conserved immune response state and can be used to classify patient responses based on immune cell heterogeneity in peripheral blood, independent of specific diagnoses.
[0353] Clonal expansion of antigen-specific T cells is a fundamental feature of effective immune responses. This analysis of functional immune type groups points to an association between the predominant cell phenotype and T cell receptor (TCR) repertoire composition in each cluster; therefore, bulk RNA-seq data generated for many patients was used to assess the T cell repertoire. Coverage of CDR3 sequences from the TCR β chain was consistent across cohorts and reflected the overall frequency of T cells in each sample (Figure 20A, top). While rare across the cohort, dominant clones representing >10% of all CDR3s in each patient (Figure 20A, bottom) were enriched in the chronic (G4) subtype. TCR β repertoire clonality in individuals in G4 was threefold higher than in all other immune type groups (Figure 20B), consistent with the abundance of terminally differentiated T cells in this group (Figure 20B). Conversely, naive, primed, and advanced immune types (G1–G3) had significantly higher TCRβ diversity than immune types G4 and G5, with the Chao1 richness index decreasing from G1 to G3 (Figure 20C), consistent with the frequency of naive, central, and transitional memory T cells in each group. These differences in TCRβ chain clonality and diversity were also observed for TCRα CDR3 sequences (Figures 24A–H and 24M).
[0354] Next, we analyzed the distribution of MHC class I HLA alleles A, B, and C (Figure 20D, HLA-B) within the different functional immune type groups, as HLA type skew within different immune type groups may lead to the observed changes in repertoire diversity. The distribution of HLA alleles from this cohort was heterogeneous (Figure 20D). Patients with the HLA-B07:02 allele were found to be less numerous compared to all other groups (Figure 20D, right), whereas other HLA alleles were not significantly different. This analysis provides further evidence that functional immune type categories are shaped by convergent responses to immunological stimuli. Overall, this supports the validity of these categories when assessing the efficacy of immune checkpoint blockade (ICB).
[0355] To further explore this, we extended the GSEA analysis to assess functional immune type groups using annotated gene signatures corresponding to T cell differentiation state, repertoire diversity, and gene expression patterns consistent with PD-1-expressing T cells targeted by immune checkpoint blockade (ICB). First, we observed an enrichment pattern for a common T cell differentiation signature, including genes differentially expressed between naive and activated CD8+ T cells, similar to TCRβ repertoire diversity (Figure 20D). In addition, individual gene expression levels of the prominent transcription factors TCF-7, LEF1, and ID3 associated with naive and self-renewing memory T cells were highest in the G1-naive and G2-primed groups, reflecting the T cell differentiation signature score (Figure 20E) (PMID: 21383243, 19204323). Consistent with the TCRβ clonality and frequency of terminally differentiated T cells in the cluster, the chronic (G4) group had the highest enrichment score for the PD-1-high CD8+ T cell signature (Figure 20F). The transcription factors TBX21, EOMES, and TOX, which are critical regulators of effector T cell differentiation and exhaustion, also had the highest expression in the G4 group (Figure 20E). Collectively, this analysis suggests that patients with the chronic (G4) immune type are characterized by prolonged antigen exposure consistent with the evolution of antitumor responses and are associated with ICB responses in the peripheral blood and TME.
[0356] BCR repertoire diversity analysis showed a similar trend to the TCR repertoire (Figures 24A-M). The G1-naive cluster had the highest chao1 diversity scores for both BCR heavy chains and lambda and kappa light chains compared to the other clusters, and was significantly different from the G3, G4, and G5 clusters. BCR clonality did not show significant differences between the clusters, and did not reach significance (Figures 24I-L).
[0357] Patients in this cohort with different cancer diagnoses undergoing treatment with ICB alone or in combination (n = 72) were hypothesized to most frequently belong to the chronic (G4) immune type. Although there was an increased frequency of patients in clusters G3 and G4, this distribution was not significantly different from patients with cancer diagnoses as a whole. PDCD1 expression levels by RNA-seq did not differ significantly between immune groups (Figure 20G). Patients in the G1, G2, and G4 groups had the highest expression of annotated genes associated with PD1 signaling and cancer immunotherapy with PD1 blockade (Figure 20F).
[0358] Importantly, this immune profiling platform was tested in a clinical cohort of 36 patients with advanced head and neck squamous cell carcinoma (HNSCC) treated with the PD-1-blocking antibody nivolumab. Cryopreserved peripheral blood samples were obtained for each patient for retrospective analysis before nivolumab infusion (baseline or pretreatment) and after treatment (Figure 21A). According to objective radiological and pathological assessment, 22 / 36 patients (64%) responded to therapy (responders). The HNSCC clinical cohort also showed a similar distribution of UMAP relative to the internal cohort (Figure 21B).
[0359] A multiclass immune typing model was applied to the clinical cohort samples to assign patients to immune type groups (Figure 21C). Consistent with previous analyses (Figures 17-18), the largest proportion of patient samples was classified as G2-primed immune type (Figure 22C). Response to anti-PD-1 therapy was significantly associated with baseline (pre-treatment) G2 immune type. Non-responders were observed within each group, but the naive immune type group (G1) contained predominantly non-responders at baseline (Figure 21D). Dynamic analysis of post-treatment samples showed that 39% (14 / 36) of patients remained within the same functional immune type group. Patients with a baseline G2-primed immune type who remained in this group after treatment administration had a high response rate to therapy (Figure 21E). Patients with a G2-primed immune type at the time point of treatment had a significantly higher proportion of responders compared to the overall cohort, whereas chronic (G4) and suppressed (G5) immune types were dominated by non-responders (Figure 21F). Collectively, this analysis demonstrates that functional immune type assignment can effectively stratify patients who will respond to ICB based on the composition of immune cell types in the peripheral blood.
[0360] The responding G2-primed immune type was enriched for central and transitional memory CD4+ T cells in the peripheral blood of the HNSCC cohort, which led us to further evaluate all immune cell populations between responders and non-responders to validate these findings. Population analysis of baseline sample differences between responders and non-responders revealed significant increases in 10 cell populations in the peripheral blood of responders, nine of which belong to the CD4+ T cell lineage (Figures 22A and 22B). Th2, Th17, central memory T helper, and effector CD4+ T helper cell populations were significantly enriched in the blood of responders compared with non-responders (Figure 22C), which contrasted with the overall proportions of naive and total CD4+ T cells. Overall, these results further support the relationship between enrichment of memory CD4+ T cells in the peripheral blood and ICB efficacy in HNSCC.
[0361] The distribution of the internal cohort based on the frequencies of different immune cell subpopulations demonstrated a continuous gradient of characteristically distinct immune type clusters during UMAP. However, patient assignment to distinctive clusters is limited in capturing the dynamic transitions of immune responses across immune type clusters (Figure 22D, left). It was reasonable that creating a linear continuous score reflecting sample similarity within and across its individual immune type clusters would increase the utility of this framework as a predictive biomarker. By using machine learning to identify the most differentially enriched populations in each cluster (Table 9), we generated population frequency-based coefficients for each subpopulation, which we used to calculate a functional immune type signature score (0–10) based on the degree of similarity to the specific clusters phenotyped in the internal cohort (Figure 22D, right). The G2-primed signature score was significantly higher in ICB responders compared with non-responders in the HNSCC cohort at both baseline and post-treatment administration (p values = 0.017 and 0.014, respectively, Mann-Whitney U test) (Figure 22E). Next, we used the G2-primed signature score to train a binary classifier to predict ICB response in HNSCC patients, revealing a positive predictive value of 74% for pre-treatment and 76% for on-treatment samples (Figure 22F). Thus, the newly developed G2-primed signature has promising clinical utility.
[0362] [Table 11-1]
[0363] [Table 11-2]
[0364] [Table 11-3]
[0365] [Table 12]
[0366] method Internal cohort description Peripheral blood samples from cancer patients were collected at multiple medical facilities across the United States and sent to Boston Gene Laboratory. Blood from healthy donors was purchased from multiple collection sites in the area: Research Blood Components (Watertown, MA), STEMCELL Technologies (Vancouver, BC, Canada), and Discovery Life Sciences (Huntsville, AL). All patients provided written consent under an IRB-approved protocol. Initially, 960 blood samples were collected for flow cytometry analysis, including 470 patients with various cancer types (145 sarcoma cancer subtypes and 325 epithelial-origin cancers) and 449 healthy donor samples. 145 patients had sarcoma cancer subtypes and 325 had epithelial-origin cancers. After excluding samples based on insufficient quality, a total of 850 flow cytometry samples were analyzed in this study. White blood cell analysis was performed for all patients using a proprietary flow cytometry method (Figure 16A).
[0367] The median age of the cohort was 47 years for healthy donors and 61.5 years for cancer patients. Only patients with sarcomas and carcinomas were included, and the most common diagnoses of epithelial origin were pancreatic cancer (n = 37), breast neoplasms (n = 65), non-small cell lung cancer (n = 32), colorectal neoplasms (n = 41), melanoma (n = 19), and prostate (n = 18). Treatment information was available for 417 patients (417 / 442, 94.3%). Previous treatment, including chemotherapy, radiation therapy, ICI, or systemic therapy classified as other, was administered within 1 year of blood collection in 211 patients (211 / 417, 50.6%). 234 patients (234 / 417, 56.1%) were currently undergoing therapy at the time of specimen collection. Based on the data provided, 44 patients (44 / 417, 10.55%) had no evidence of therapy administration after cancer diagnosis. In addition, 797 RNA samples from both healthy and cancer donors were analyzed. This diverse cohort was used for multiscale analysis of the relationship between cancer and peripheral blood immunity.
[0368] Head and neck squamous cell carcinoma cohort To further explore the relationship between the newly discovered immune clusters and cancer immunotherapy, we applied this flow cytometry analytical framework to a cohort of 36 head and neck squamous cell carcinoma (HNSCC) patients. The HNSCC cohort was part of a prospective phase II trial conducted at Thomas Jefferson University Hospital. During this trial, patients received either anti-PD1 monoclonal antibody treatment (nivolumab) or nivolumab in combination with a specific IDO inhibitor (BMS986205). Pre- and post-treatment cryopreserved PBMCs were thawed and subjected to multicolor flow cytometry staining. A total of 70 samples were analyzed; two patients only had pre-treatment samples due to poor quality of post-treatment PBMCs.
[0369] Blood sample processing and white blood cell (WBC) isolation Upon receipt, all fresh peripheral blood samples were subjected to a complete blood count using a DxH 500 hematology analyzer (Beckman Coulter, Brea, CA). Samples received within 24 hours of collection were subjected to red blood cell (RBC) lysis of 3 ml whole blood, and white blood cells (WBC) were isolated using 42 ml nuclease-free HyPure water mixed with 5 ml 10x RBC lysis buffer (eBioscience). Samples were lysed for 10 minutes at room temperature with continuous mixing on a tube rotator. Cells were then centrifuged at 300 x g for 5 minutes and washed with sorter buffer (2% NBCS in PBS + 1 mM EDTA).
[0370] Thawing cryopreserved PBMCs Cryopreserved peripheral blood mononuclear cell (PBMC) samples were stored in a liquid nitrogen tank and thawed at 37°C using pre-made thawing medium (500 mL RPMI 1640 medium + 10 mL HEPES + 10 mL PENSTREP + 10 mL MEMNEAA + 10 mL NAHEP + 5 mL 20% NBCS in GlutaMAX). Prior to thawing, 15 mL aliquots of thawing medium were pre-warmed to 37°C in a water bath and supplemented with 75 μL DNase (20 mg / mL) and 75 μL glutathione (200 mM). Samples were removed from the liquid nitrogen tank and immediately immersed in a 37°C water bath without submerging the caps. Thawing was monitored visually, and samples were agitated in the water bath for approximately 1 minute until only a few ice crystals remained. Using a wide-bore 1 mL pipette, each sample was transferred to an empty 15 mL tube. Prewarmed, supplemented thawing medium was slowly pipetted into the tube, gently layering the medium over the sample. After a 3-4 mL layer was formed, the warmed medium was slowly pipetted directly into the sample, simultaneously stirring until the sample was homogenous. Once homogenous, the sample was topped up while warming and supplemented with thawing medium to a final volume of 15 mL. The PBMC sample was then centrifuged at 300 × g for 8 minutes, washed with thawing medium at 300 × g for 8 minutes, and then stained.
[0371] Cell staining and flow cytometry Isolated WBCs or PBMCs were centrifuged at 300 x g for 5 minutes, resuspended, and blocked with blocking buffer (IMDM + 10% NBCS + DNase I (1:200) + human TrueStain FcX (1:50) + monocyte blocker (1:50) + unlabeled normal mouse IgG (1:200)) for 10 minutes at room temperature. After blocking, each sample was aliquoted into 10 unique wells of a 96-well plate, centrifuged at 300 x g for 3 minutes, and the supernatant was removed. Each well was stained with Ghost Dye Violet 510 Viability Dye (1:400, Tonbo) in PBS for 10 minutes at room temperature. After staining with the viability dye, 200 μL of sorter buffer was added to each well and centrifuged at 300 x g for 3 minutes, followed by removal of the supernatant. Samples were stained with 10 custom flow cytometry panels (Table 13) for 20 minutes at room temperature. After staining, 200 μL of sorter buffer was added to each well and centrifuged at 300 × g for 3 minutes, followed by removal of the supernatant. Cells were then fixed overnight at 4 °C in 1% paraformaldehyde solution (Cytofix / Cytoperm, BD Biosciences). The fixative was then washed with sorter buffer and resuspended in acquisition buffer (PBS + 0.5% (w / v) BSA + 0.75% (w / v) glycine + 5 mM EDTA + Tween-20 (1:2000) + sodium azide (1:100)).
[0372] Stained and fixed cells were acquired on a BD FACSCelesta flow cytometer. CS&T Research beads (BD Biosciences) were used to validate the performance of the BD FACSCelesta before each acquisition. A compensation matrix was generated by calculating the spectral overlap from single-stain controls using FACSDiva software. Single-stain controls were prepared in-house by staining a set of 13 samples on Ultracomp eBeads compensation beads (Thermofisher) with a unique antibody in each channel.
[0373] [Table 13-1]
[0374] [Table 13-2]
[0375] [Table 13-3]
[0376] [Table 13-4]
[0377] [Table 13-5]
[0378] [Table 13-6]
[0379] [Table 13-7]
[0380] [Table 13-8]
[0381] RNA isolation Isolated WBCs for RNA sequencing were centrifuged at 300 × g for 5 minutes, with a maximum of 1e6 cells per vial. The supernatant was removed, and cells were resuspended in cold homogenization buffer (2% 1-thioglycerol, Promega). Samples were then frozen at -80°C until extraction. RNA extraction from frozen samples was performed using a stationary automated Maxwell RSC instrument (Promega) according to the Maxwell RSC simplyRNA Cells Kit (Promega).
[0382] Library preparation and sequencing of samples Libraries were prepared using Illumina TruSeq® standard mRNA library preparation (Poly-A mRNA; standard). Libraries were sequenced on a NovaSeq 6000 with paired-end reads (2 x 150) and a target coverage of 50 ml reads.
[0383] Flow cytometry data processing Flow cytometry data underwent several quality control steps to ensure consistent analytical input and overall high quality. All selected patient samples contained 10,000 or more cells per panel. Files with poor compensation or random PMT failure were excluded. Flow cytometry data were exported in fcs 3.0 file format and analyzed by applying compensation matrices as Pandas DataFrames (v1.1.4). Data processing and analysis were performed using FlowKit (v.0.5.0, https: / / github.com / malcommac / FlowKit / releases). All fluorochrome marker channel values were transformed using the arcsinh transformation: x = ln(x + √((x^2+1))) divided by a factor of 190. Forward and side scatter values (FCS-A / H / W and SSC-A / H / W) were divided by 105 to match the order of the arcsinh-transformed data.
[0384] Manual Data Analysis We developed a framework for the accurate manual analysis of cell populations in 2D scatter plots, combining classical gating and clustering steps. Each panel was analyzed separately according to its own specific strategy. Each strategy consisted of implementing several successive steps of the following cell selection / labeling method: Clustering method. Events were clustered using FlowSOM (v0.1.1, https: / / pypi.org / project / FlowSom / ). Data were visualized using the tSNE algorithm (openTSNE, v0.6.2, https: / / pypi.org / project / openTSNE / ), and color-coded by both clustering results and overall marker intensity to visualize marker intensity combinations in specific clusters. Each cluster was manually matched to a cell population based on the marker intensity combination in that cluster.
[0385] Processing the cytometry data before clustering can include noise transformation. Noise transformation adjusts the marker intensities to reduce the impact of noise on the clustering results, which involves lowering the marker intensities below a certain threshold. The noise threshold for a marker is manually defined based on a two-dimensional plot of the marker's intensity against the intensity of other markers in the panel. The boundary between the marker's noise and positive signal is selected at the visually observed minimum point in the marker distribution. The following equation describes the marker intensity after noise transformation:
number
[0386] Population selection by two-dimensional plots shows pairwise projections of data distribution histograms, color-coded by event distribution density (as done in classical gating methods). The boundary between the positive and negative populations is manually selected at the visually observed minimum in the marker distribution. To facilitate visual observation of the distribution minimum, a kernel density estimation plot is used on top of the density plot.
[0387] The end result of manual data labeling was a cell population label for every event in the fcs file.
[0388] Model training for flow cytometry data analysis LightGBM decision tree boosting machine learning models were trained using manually labeled data (using default parameters, https: / / lightgbm.readthedocs.io / en / latest / Parameters.html). These models were trained to predict the label of each cytometry event. Approximately 200–300 labeled FCS samples were used for the models of each cytometry panel. Forward scatter, side scatter, and corrected fluorescence channel signal values, normalized to the maximum value and with different quantiles selected for each panel, were used as inputs (Table 14). A voting model was trained for each panel. The base voting model consisted of two submodels, each represented by a LightGBM decision tree boosting classifier (LightGBM, v3.3.2, https: / / pypi.org / project / lightgbm / 3.3.2 / ). The first type predicted the "top" population, such as leukocytes in the general panel and CD8 T cells in the CD8 T cell panel. The second type classified the target population into subtypes. The Supplementary Model Figure shows the overall model training method.
[0389] The performance of the model was confirmed on a validation sample set (approximately 30 samples per panel) that was not used for model training. Predictions were generated for each of the validation samples. These predictions were then compared to the manual labels of these samples based on the f1 score and p4 score metrics (see Cytometry_supplement.xlsx, listing "models_quality," which shows the average f1 score and p4 score per panel across the entire population used in this paper).
[0390] [Table 14]
[0391] Quality control of predicted labels All predicted labels generated by the model were subjected to a manual quality control procedure. The quality of the predicted labels was determined using a panel-specific set of two-dimensional plots of the intensity of one marker against the intensity of another. The major populations of the panel were plotted in different colors on these plots to verify the accuracy of population selection and separation from each other. If the predicted labels of a file were incorrect, the gating of that population was manually corrected.
[0392] Determination of cell percentage To calculate the final population percentages from the labeled data, the results from the different cytometry panels were combined together using the common panel (CP10). The cell count values of the corresponding populations from the other panels were multiplied by a normalization factor to fit the results from the linear panel. The normalization factor was obtained by dividing the number of cells in the reference population in the linear panel by the number of cells in the reference population in the other panel (monocytes for the monocyte panel, T cells for the CD4 T cell panel, etc.). Table 15 contains a complete list of the refer...
Claims
1. 1. A method for determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, comprising: Using at least one computer hardware processor, obtaining cytometry data or RNA expression data from a biological sample obtained from said subject; processing said cytometry data or said RNA expression data to determine cellular composition percentages for at least 20 cell types listed in Table 4; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, wherein the leukocyte signature comprises the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types. To carry out A method comprising:
2. the cytometry data comprises flow cytometry data; processing the cytometry data or the RNA expression data comprises processing the flow cytometry data to determine the cellular composition percentages for at least the 20 cell types listed in Table 4; The method of claim 1.
3. 3. The method of claim 2, wherein the flow cytometry data is obtained from a biological sample consisting of white blood cells.
4. 4. The method of claim 3, wherein processing the flow cytometry data comprises determining a cellular composition percentage for each cell type listed in Table 1.
5. 3. The method of claim 2, wherein the flow cytometry data is obtained from a biological sample consisting of peripheral blood mononuclear cells (PBMCs).
6. 6. The method of claim 5, wherein processing the flow cytometry data comprises determining a cellular composition percentage for each cell type listed in Table 2.
7. 7. The method of any one of claims 1-6, wherein processing the cytometry data comprises applying one or more machine learning models to the cytometry data to obtain cellular composition percentages for the at least 20 cell types listed in Table 4.
8. obtaining RNA expression data includes obtaining sequencing data that was previously obtained by sequencing the biological sample obtained from the subject; Including, processing the cytometry data or the RNA expression data comprises processing the RNA expression data to determine the cellular composition percentages for at least the 20 cell types listed in Table 4; The method of claim 1.
9. 9. The method of claim 8, wherein the sequencing data comprises at least 1 million reads, at least 5 million reads, at least 10 million reads, at least 20 million reads, at least 50 million reads, or at least 100 million reads.
10. 10. The method of claim 8 or 9, further comprising normalizing the RNA expression data to transcripts per million (TPM) units prior to processing the RNA expression data to determine the cellular composition percentages.
11. 11. The method of any one of claims 8 to 10, wherein processing the RNA expression data comprises determining a cellular composition percentage for each cell type listed in Table 3.
12. 12. The method of any one of claims 8-11, wherein processing the RNA expression data comprises applying a cellular deconvolution algorithm to the RNA expression data to obtain cellular composition percentages for the at least 20 cell types listed in Table 3.
13. the plurality of leukocyte immune profile types being associated with a respective plurality of leukocyte immune profile types; using the leukocyte signature and identifying the leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types; Associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types; identifying the leukocyte immune profile type for the subject as the leukocyte immune profile type corresponding to the particular one of the plurality of leukocyte immune profile types with which the leukocyte signature of the subject is associated; Including, The method according to any one of claims 1 to 12.
14. Associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types, processing the leukocyte signature with a trained classifier to obtain an output indicative of the particular one of the plurality of leukocyte immune profile types; 14. The method of claim 13, comprising:
15. 15. The method of claim 14, wherein the trained classifier comprises a trained neural network classifier, optionally a Tabular Pre-Data Fitted Network Transformer (TabPFN) classifier.
16. said associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types; determining a score for each particular one of the plurality of leukocyte immune profile types indicating whether the leukocyte signature of the subject is associated with that particular leukocyte immune profile type; Including, 14. The method of claim 13, wherein determining the score for a particular leukocyte immune profile type comprises applying a linear regression model associated with the particular leukocyte immune profile type to the cellular composition percentages in the leukocyte signature.
17. 17. The method of any one of claims 1 to 16, further comprising generating the plurality of leukocyte immune profile types, said generating comprising: obtaining a plurality of cytometry data or RNA expression datasets from biological samples obtained from a plurality of respective subjects, wherein each of the plurality of cytometry data or RNA expression datasets indicates a cellular composition percentage for at least 20 cell types listed in Table 4; generating a plurality of leukocyte signatures from the plurality of cytometry data or RNA expression data sets, each of the plurality of leukocyte signatures comprising a cellular composition percentage for the at least 20 cell types listed in Table 4, wherein said generating comprises, for each particular one of the plurality of leukocyte signatures: determining the leukocyte signature by using the cytometry data or RNA expression data to determine the percentage composition of cells in the particular cytometry data or RNA expression data set from which the particular leukocyte signature is generated; and clustering the plurality of leukocyte signatures to obtain the plurality of leukocyte immune profile types. A method comprising:
18. updating the plurality of leukocyte immune profile types using the leukocyte signature of the subject, wherein the leukocyte signature of the subject is one of a threshold number of leukocyte signatures for a threshold number of subjects, and the leukocyte immune profile type is updated when the threshold number of leukocyte signatures are generated. The method of any one of claims 15 to 17, further comprising: The method, wherein the threshold number of leukocyte signatures is at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, or at least 5000 leukocyte signatures.
19. 20. The method of claim 18, wherein the updating is performed using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and an agglomerative clustering algorithm.
20. determining a leukocyte immune profile type of a second subject, wherein the leukocyte immune profile type of the second subject is identified using the updated leukocyte immune profile type.
20. The method of claim 18 or 19, wherein said identifying further comprises: determining a leukocyte signature of the second subject from cytometry data or RNA expression data from a biological sample obtained from the second subject; Associating the leukocyte signature of the second subject with a particular one of the plurality of updated leukocyte immune profile types; and identifying the leukocyte immune profile type for the second subject as the leukocyte immune profile type that corresponds to the particular one of the plurality of updated leukocyte immune profile types with which the leukocyte signature of the second subject is associated; A method comprising:
21. 20. The method of any one of claims 17 to 19, wherein the clustering is performed using a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm.
22. 22. The method of claim 21, wherein the clustering is performed using a spectral clustering algorithm.
23. 23. The method of any one of claims 1 to 22, wherein the plurality of leukocyte immune profile types comprises a naive type, a primed type, an advanced type, a chronic type, and a suppressed type.
24. 24. The method of any one of claims 1 to 23, further comprising identifying the subject as a candidate for immunotherapeutic treatment based on said identifying the leukocyte immune profile type for the subject.
25. The method of any one of claims 1 to 24, further comprising identifying the subject as a candidate for treatment with immunotherapy when the subject is identified as having the primed type.
26. 26. The method of any one of claims 1 to 25, further comprising administering a therapeutic agent to the subject based on the identification of the subject's leukocyte immune profile type.
27. The method of any one of claims 1 to 26, further comprising administering immunotherapy to the subject when the subject is identified as having the primed type.
28. 1. A method for determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, comprising: Using at least one computer hardware processor, obtaining flow cytometry data for white blood cells (WBCs) isolated from a biological sample obtained from said subject; processing said flow cytometry data to determine cellular composition percentages for at least 20 cell types listed in Table 1; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, wherein the leukocyte signature comprises the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types. To carry out A method comprising:
29. 29. The method of claim 28, wherein the WBCs consist of granulocytic and agranulocytic cells.
30. 30. The method of claim 28 or 29, wherein processing the flow cytometry data comprises applying one or more machine learning models to the flow cytometry data to obtain cellular composition percentages for the at least 20 cell types listed in Table 1.
31. 31. The method of any one of claims 28-30, wherein processing the flow cytometry data comprises determining the cellular composition percentages for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD4+ T cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
32. Processing the flow cytometry data can include identifying naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, Vδ2+ γδ T cells, post-class-switched memory B cells, central memory CD8+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD39+ CD4+ Tregs, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, effector memory CD4+ T cells, NKT cells, CD8+ TEMRA, effector memory CD8+ 32. The method of any one of claims 28 to 31, comprising determining the cellular composition percentages for T cells, CD4+ TEMRA, neutrophils, granulocytes, classical monocytes, non-classical monocytes, and HLA-DR low monocytes.
33. 33. The method of any one of claims 28 to 32, wherein the leukocyte signature comprises a cellular composition percentage for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD4+ T cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
34. The leukocyte signature may be selected from the group consisting of naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, Vδ2+ γδ T cells, post-class-switched memory B cells, central memory CD8+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD39+ CD4+ Tregs, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, effector memory CD4+ T cells, NKT cells, CD8+ TEMRA, and effector memory CD8+ 34. The method of any one of claims 28 to 33, comprising the cell composition percentages for T cells, CD4+ TEMRA, neutrophils, granulocytes, classical monocytes, non-classical monocytes, and optionally HLA-DR low monocytes.
35. the plurality of leukocyte immune profile types being associated with a respective plurality of leukocyte immune profile types; using the leukocyte signature and identifying the leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types; Associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types; and identifying the leukocyte immune profile type for the subject as the leukocyte immune profile type corresponding to the particular one of the plurality of leukocyte immune profile types with which the leukocyte signature of the subject is associated; The method of any one of claims 28 to 34, comprising:
36. Associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types, processing the leukocyte signature with a trained classifier to obtain an output indicative of the particular one of the plurality of leukocyte immune profile types; 36. The method of claim 35, comprising:
37. 37. The method of claim 36, wherein the trained classifier comprises a trained neural network classifier, optionally a Tabular Pre-Data Fitted Network Transformer (TabPFN) classifier.
38. said associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types; determining a score for each particular one of the plurality of leukocyte immune profile types indicating whether the leukocyte signature of the subject is associated with that particular leukocyte immune profile type; Including, 36. The method of claim 35, wherein determining the score for a particular leukocyte immune profile type comprises applying a linear regression model associated with the particular leukocyte immune profile type to the cellular composition percentages in the leukocyte signature.
39. 39. The method of any one of claims 28 to 38, further comprising generating the plurality of leukocyte immune profile types, said generating comprising: obtaining a plurality of flow cytometry data sets from white blood cells (WBCs) isolated from a plurality of biological samples obtained from each of the subjects, each of the plurality of flow cytometry data sets indicating a cellular composition percentage for at least 20 cell types listed in Table 1; generating a plurality of leukocyte signatures from the plurality of flow cytometry data sets, each of the plurality of leukocyte signatures comprising a cellular composition percentage for at least 20 cell types listed in Table 1, said generating including, for each particular one of the plurality of leukocyte signatures: determining the leukocyte signature by determining the cellular composition percentages using the flow cytometry data in the particular flow cytometry data set from which the particular leukocyte signature was generated; and clustering the plurality of leukocyte signatures to obtain the plurality of leukocyte immune profile types. A method comprising:
40. updating the plurality of leukocyte immune profile types using the leukocyte signature of the subject, wherein the leukocyte signature of the subject is one of a threshold number of leukocyte signatures for a threshold number of subjects, and the leukocyte immune profile type is updated when the threshold number of leukocyte signatures are generated.
40. The method of any one of claims 35 to 39, further comprising: The method, wherein the threshold number of leukocyte signatures is at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, or at least 5000 leukocyte signatures.
41. 41. The method of claim 39 or 40, wherein the updating is performed using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and an agglomerative clustering algorithm.
42. Determining the leukocyte immune profile type of the second subject 42. The method of claim 40 or 41, wherein the leukocyte immune profile type of the second subject is identified using the updated leukocyte immune profile type, and wherein said identifying further comprises: determining a leukocyte signature of the second subject from flow cytometry data from leukocyte cells isolated from a biological sample obtained from the second subject; Associating the leukocyte signature of the second subject with a particular one of the plurality of updated leukocyte immune profile types; and identifying the leukocyte immune profile type for the second subject as the leukocyte immune profile type that corresponds to the particular one of the plurality of updated leukocyte immune profile types with which the leukocyte signature of the second subject is associated; A method comprising:
43. 43. The method of claim 42, wherein the clustering is performed using a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm.
44. 44. The method of claim 43, wherein the clustering is performed using a spectral clustering algorithm.
45. 45. The method of any one of claims 28 to 44, wherein the plurality of leukocyte immune profile types comprises a naive type, a primed type, an advanced type, a chronic type, and a suppressed type.
46. 46. The method of any one of claims 28-45, further comprising identifying the subject as a candidate for immunotherapeutic treatment based on said identifying the leukocyte immune profile type for the subject.
47. 47. The method of any one of claims 28 to 46, further comprising identifying the subject as a candidate for treatment with immunotherapy when the subject is identified as having the primed type.
48. 48. The method of any one of claims 28 to 47, further comprising administering a therapeutic agent to the subject based on the identification of the subject's leukocyte immune profile type.
49. 49. The method of any one of claims 28 to 48, further comprising administering immunotherapy to the subject when the subject is identified as having the primed type.
50. 50. The method of any one of claims 1 to 49, wherein the subject has head and neck squamous cell carcinoma (HNSCC).
51. 1. A method for determining a leukocyte immune profile type of a subject having, suspected of having, or at risk of having cancer, comprising: Using at least one computer hardware processor, obtaining RNA expression data for peripheral blood mononuclear cells (PBMCs) isolated from a biological sample obtained from said subject; processing said RNA expression data to determine cellular composition percentages for at least 20 cell types listed in Table 3; generating a leukocyte signature for the subject using the determined cellular composition percentages for the at least 20 cell types, wherein the leukocyte signature comprises the cellular composition percentages for the at least 20 cell types; and using the leukocyte signature and identifying a leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types. The method includes performing the following.
52. 52. The method of Claim 51, wherein said RNA expression data comprises bulk RNA expression data.
53. 53. The method of Claim 52, wherein processing said RNA expression data comprises applying a cellular deconvolution technique comprising one or more machine learning models to obtain said cellular composition percentages.
54. 54. The method of any one of claims 51 to 53, wherein the cellular composition percentages comprise cellular composition percentages for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
55. 55. The method of any one of claims 51 to 54, wherein the cellular composition percentages comprise cellular composition percentages for naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, class-switched memory B cells, central memory, CD4+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD4+ T cells, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, CD8+ TEMRA, effector memory CD8+ T cells, neutrophils, granulocytes, classical monocytes, and non-classical monocytes.
56. 56. The method of any one of claims 51 to 55, wherein the leukocyte signature comprises a cellular composition percentage for naive CD4+ Tregs, naive CD4+ T cells, naive CD8+ T cells, naive B cells, effector memory CD8+ T cells, classical monocytes, and non-classical monocytes.
57. 57. The method of any one of claims 51 to 56, wherein the leukocyte signature comprises cellular composition percentages for naive CD4+ T cells, naive CD8+ T cells, naive B cells, non-transformed memory IgM B cells, class-switched memory B cells, central memory, CD4+ T cells, CD4+ Tregs, transitional memory CD4+ T cells, central memory CD4+ T cells, total memory CD4+ T cells, CD4+ T cells, CD4+ T cells, eosinophils, basophils, plasmacytoid dendritic cells, dendritic cells, PD1 high CD8+ T cells, transitional memory CD8+ T cells, cytotoxic NK cells, regulatory NK cells, total memory CD8+ T cells, CD8+ T cells, CD8+ TEMRA, effector memory CD8+ T cells, neutrophils, granulocytes, classical monocytes, and non-classical monocytes.
58. 58. The method of any one of claims 51 to 57, wherein the plurality of leukocyte immune profile types is associated with a respective plurality of leukocyte immune profile types, using the leukocyte signature and identifying the leukocyte immune profile type for the subject from among a plurality of leukocyte immune profile types; Associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types; and identifying the leukocyte immune profile type for the subject as the leukocyte immune profile type corresponding to the particular one of the plurality of leukocyte immune profile types with which the leukocyte signature of the subject is associated; A method comprising:
59. Associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types, processing the leukocyte signature with a trained classifier to obtain an output indicative of the particular one of the plurality of leukocyte immune profile types; 59. The method of claim 58, comprising:
60. 60. The method of claim 59, wherein the trained classifier comprises a trained neural network classifier, optionally a Tabular Pre-Data Fitted Network Transformer (TabPFN) classifier.
61. said associating the leukocyte signature of the subject with a particular one of the plurality of leukocyte immune profile types; determining a score for each particular one of the plurality of leukocyte immune profile types indicating whether the leukocyte signature of the subject is associated with that particular leukocyte immune profile type; 59. The method of claim 58, comprising: The method, wherein determining the score for a particular leukocyte immune profile type comprises applying a linear regression model associated with the particular leukocyte immune profile type to the cell composition percentages in the leukocyte signature.
62. 62. The method of any one of claims 51 to 61, further comprising generating the plurality of leukocyte immune profile types, said generating comprising: obtaining a plurality of RNA expression datasets from white blood cells (WBCs) isolated from biological samples obtained from a plurality of respective subjects, wherein each of the plurality of RNA expression datasets indicates a cellular composition percentage for at least 20 cell types listed in Table 3; generating a plurality of leukocyte signatures from said plurality of RNA expression datasets; wherein each of the plurality of leukocyte signatures comprises a cellular composition percentage for at least 20 cell types listed in Table 3, and wherein said generating comprises, for each particular one of the plurality of leukocyte signatures: determining the leukocyte signature by determining the cellular composition percentages using the RNA expression data in the particular RNA expression dataset from which the particular leukocyte signature was generated; and clustering the plurality of leukocyte signatures to obtain the plurality of leukocyte immune profile types. A method comprising:
63. updating the plurality of leukocyte immune profile types using the leukocyte signature of the subject, wherein the leukocyte signature of the subject is one of a threshold number of leukocyte signatures for a threshold number of subjects, and the leukocyte immune profile type is updated when the threshold number of leukocyte signatures are generated.
63. The method of any one of claims 51 to 62, further comprising: The method, wherein the threshold number of leukocyte signatures is at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, or at least 5000 leukocyte signatures.
64. 64. The method of claim 63, wherein said updating is performed using a clustering algorithm selected from the group consisting of a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and an agglomerative clustering algorithm.
65. Determining the leukocyte immune profile type of the second subject 65. The method of claim 63 or 64, wherein the leukocyte immune profile type of the second subject is identified using the updated leukocyte immune profile type, said identifying further comprising: determining a leukocyte signature of the second subject from RNA expression data from leukocyte cells isolated from a biological sample obtained from the second subject; Associating the leukocyte signature of the second subject with a particular one of the plurality of updated leukocyte immune profile types; and identifying the leukocyte immune profile type for the second subject as the leukocyte immune profile type that corresponds to the particular one of the plurality of updated leukocyte immune profile types with which the leukocyte signature of the second subject is associated; A method comprising:
66. 66. The method of claim 65, wherein the clustering is performed using a density clustering algorithm, a spectral clustering algorithm, a k-means clustering algorithm, a hierarchical clustering algorithm, and / or an agglomerative clustering algorithm.
67. 67. The method of claim 66, wherein the clustering is performed using a spectral clustering algorithm.
68. 68. The method of any one of claims 51 to 67, wherein the plurality of leukocyte immune profile types comprises a naive type (G1), a primed type (G2), a progressive type (G3), a chronic type (G4), and a suppressed type (G5).
69. 69. The method of any one of claims 51-68, further comprising identifying the subject as a candidate for immunotherapeutic treatment based on said identifying the leukocyte immune profile type for the subject.
70. 70. The method of any one of claims 51 to 69, further comprising identifying the subject as a candidate for treatment with immunotherapy when the subject is identified as having the primed type.
71. 71. The method of any one of claims 51 to 70, further comprising administering a therapeutic agent to the subject based on the identification of the subject's leukocyte immune profile type.
72. 72. The method of any one of claims 51 to 71, further comprising administering immunotherapy to said subject when said subject is identified as having the primed type.
73. 73. The method of any one of claims 51 to 72, wherein the subject has head and neck squamous cell carcinoma (HNSCC).
74. at least one computer hardware processor; and At least one computer readable storage medium storing processor executable instructions which, when executed by said at least one computer hardware processor, cause said at least one computer hardware processor to perform the method of any one of claims 1 to 73. A system including:
75. At least one computer readable storage medium storing processor executable instructions that, when executed by at least one computer hardware processor, cause said at least one computer hardware processor to perform the method of any one of claims 1 to 73.