Method for quantifying molecular activity in cancer cells of human tumour
Patent Information
- Application Number
- JP2024196543
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-10-18
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-05
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to Singapore Provisional Patent Application No. 10201809232S, filed on October 18, 2018, the entire contents of which are incorporated herein by reference for all purposes.
[0002] FIELD OF THEINVENTION The present invention relates generally to the field of bioinformatics. In particular, the present invention relates to identifying biomarkers for use in the detection and diagnosis of cancer. [Background technology]
[0003] 2. Background of the Invention Tumors are heterogeneous masses of malignantly mutated cancer cells, non-malignant (stromal and immune) cells, and cell-cell junctional structures. Collectively, these components form the tumor microenvironment (TME), a multifaceted cellular milieu that is both suppressive and supportive of developing tumors. Understanding how cancer cells interact with their environment inside human tumors has been a long-standing challenge. Importantly, cancer cells typically constitute <60% of all cells in the composite tumor mass. When molecular activity (i.e., mRNA expression) is profiled in bulk tumor samples, it is not possible to determine whether a given factor is expressed primarily in cancer cells or in non-cancer cells. Any molecular readout is the sum of signals from cancer cells and numerous non-cancer cells in the TME.
[0004] Experimental models can simulate and measure crosstalk in the tumor microenvironment, but such models are generally limited by how quickly tumor cells adapt to physiology outside of their natural environment. Immunohistochemistry (IHC) can directly measure selected proteins in tumor tissue, but is not suitable for large-scale, unbiased discovery. It can be performed on single tumors, but it is labor-intensive, biased (as it can only be applied to selected markers), and not quantitative (based on the percentage of cells expressing the markers). Also, current sequencing of bulk tumor transcriptomes does not give information about cancer cells specifically. Instead, transcriptome-wide profiles of cancer and stromal cells may be generated using microdissection or single-cell profiling of tumor tissue, but these approaches are difficult to apply to tumor biopsies, and dissociation may confound cell physiology and gene expression profiles to some extent. Moreover, the above methods cannot be applied retrospectively to existing large-scale cancer genomics bulk tumor data, which represents a vast and largely unexplored resource for studying crosstalk in the tumor microenvironment.
[0005] One major branch of oncology drug discovery focuses on developing antibodies (or drugs conjugated with antibodies) that specifically target antigens / proteins inside or on the surface of cancer cells. Therefore, early in drug discovery, it is important to have access to accurate molecular profiles of cancer cells. While experimental models (cell lines and animal models) can provide an approximation, such models are generally limited by how quickly cancer cells adapt to physiology outside their natural environment. For example, expression of EGFR (and copies of the EGFR gene) in glioblastoma cancer cells is greatly reduced immediately after the cancer cells are cultured in vitro.
[0006] Currently, cancer cell gene expression can also be estimated by single cell profiling or laser microdissection. However, these approaches have limitations in that the molecular profile is biased after cell dissociation, the techniques are labor-intensive and expensive, cells cannot be easily separated, e.g., non-malignant cells from malignant (cancer) epithelial cells, and they cannot be easily applied to standard frozen tumor samples or formalin-fixed paraffin-embedded (FFPE) tumor samples, and these methods are not scalable.
[0007] Thus, there is a need for techniques that enable high-throughput profiling of cancer cells ex vivo. Furthermore, other desirable features and characteristics will become apparent from the subsequent detailed description and appended claims taken in conjunction with the accompanying drawings and this background of this disclosure. Summary of the Invention
[0008] In one aspect, the present invention refers to a method for predicting the expression profiles of cancerous and non-cancerous cells based on a set of expression profiles, where each set of expression profiles is obtained from a tumor-derived sample comprising a mixture of cancerous and non-cancerous cells of one tumor type, the method comprising the steps of: a. determining a tumor purity value for one or more tumor-derived samples; b. providing different sets of expression profiles, where the different sets of expression profiles comprise composite expression data for a plurality or all molecules expressed by cancerous and non-cancerous cells contained in the one or more tumor-derived samples; c. deconvoluting each composite expression data referred to in b. by extrapolating the expression profiles of a plurality or all molecules expressed in the different tumor samples having different tumor purity values to a tumor purity value at least substantially equal to 1 or 0, thereby predicting the expression profiles of cancerous and non-cancerous cells, respectively, from the set of expression profiles. [The present invention 1001] A method for predicting expression profiles of cancerous cells and non-cancerous cells, respectively, based on a plurality of sets of expression profiles, comprising: wherein each set of the plurality of sets of expression profiles is obtained from a tumor-derived sample comprising a mixture of cancerous and non-cancerous cells of a tumor type, the method comprising the steps of: a. determining a tumor purity value for one or more tumor-derived samples; b. providing distinct sets of expression profiles, wherein the distinct sets of expression profiles comprise combined expression data for a plurality or all of the molecules expressed by cancerous and non-cancerous cells contained in the one or more tumor-derived samples; c. Deconvoluting each of the composite expression data referred to in b. by extrapolating the expression profiles of the plurality or all of the molecules expressed in different tumor samples having different tumor purity values to a tumor purity value at least substantially equal to 1 or 0, thereby predicting the expression profiles of the cancerous and non-cancerous cells, respectively, from the plurality of sets of expression profiles. [The present invention 1002] The method of claim 10, wherein said tumor-derived sample is obtained from a single subject. [The present invention 1003] The method of claim 1002, wherein the tumor-derived sample is divided into two or more sections and one set of expression profiles is generated for each section. [The present invention 1004] The method of any of claims 1001 to 1003, wherein the step of providing a plurality of different sets of expression profiles comprises the use of an existing dataset of expression profiles. [The present invention 1005] The method of claim 1004, wherein said existing dataset of expression profiles is derived from the TCGA and ICGC databases. [The present invention 1006] Any of the methods of the invention, wherein the tumor type is selected from the group consisting of BLCA (bladder urothelial carcinoma), BRCA (invasive breast carcinoma), CESC (cervical squamous cell carcinoma), CRC (colorectal adenocarcinoma) (COAD (colon adenocarcinoma) and READ (rectal adenocarcinoma) combined), ESCA (esophageal carcinoma), GBM (glioblastoma multiforme), HNSC (head and neck squamous cell carcinoma), KIRC (renal clear cell carcinoma), KIRP (renal papillary cell carcinoma), LGG (brain low grade glioma), LIHC (hepatocellular carcinoma), LUAD (lung adenocarcinoma), LUSC (lung squamous cell carcinoma), OV (ovarian serous cystadenocarcinoma), PAAD (pancreatic adenocarcinoma), PRAD (prostate adenocarcinoma), SKCM (cutaneous melanoma), STAD (gastric adenocarcinoma), THCA (thyroid carcinoma), and UCEC (uterine endometrial carcinoma). [The present invention 1007] Any of the methods of the invention, wherein the expression profile is selected from the group consisting of gene expression, RNA expression, epigenetic expression, protein expression, proteome expression, and combinations thereof, such as RNA expression and epigenetic expression, and RNA expression and protein expression. [The present invention 1008] Any of the methods of the above invention, wherein said method for determining tumor purity is selected from the group consisting of a distribution of somatic DNA variant allele frequency, somatic DNA copy number change amplitude, germline B allele frequency, gene expression signature or pattern, protein expression signature or pattern, and DNA methylation signature or pattern, and a distribution of combinations thereof. [The present invention 1009] The method of the present invention 1008, wherein at least two, or at least three, or at least four, or at least five, or two, or three, or four, or five, or all of the methods of the present invention 1008 are used together to determine the average tumor purity. [The present invention 1010] The method of any one of claims 1001 to 1008, wherein the tumor purity value is a mean tumor purity value. [The present invention 1011] The method of the present invention 1001 further comprises a step of scoring the molecules of step c. based on the level of up-regulation or down-regulation in the cancer tissue compared to the stromal tissue; and / or a step of scoring the molecules of step c. based on the level of up-regulation or down-regulation in the cancer tissue compared to the healthy tissue. [The present invention 1012] The method of the present invention 1011 further comprising the steps of assigning the up-regulated and down-regulated molecules to gene or transcript isoforms for a known dataset of membrane-bound proteins or membrane-bound receptors; and / or assigning the up-regulated and down-regulated molecules to gene or transcript isoforms for a known dataset of HLA-binding peptides and T-cell antigen-binding peptides. [The present invention 1013] The method of claim 1012, wherein said known dataset for assigning genes or transcript isoforms is derived from Gene Ontology and / or TANTIGEN. [The present invention 1014] The method of any one of claims 1012 to 1013, further comprising the step of selecting a gene or transcript isoform for antibody-based therapy and / or T cell-based therapy. [The present invention 1015] Any of the aforementioned methods of the invention, wherein said gene or transcript isoform is a membrane bound protein, a membrane bound receptor, an antigenic peptide, a target protein, a peptide, and / or is targetable by an antibody. [The present invention 1016] Any of the aforementioned methods of the invention, wherein the molecule is selected from the group of a gene, a DNA, an RNA, or a protein molecule, or a combination thereof. [Brief description of the drawings]
[0009] The invention will be better understood by reference to the detailed description when considered in conjunction with the non-limiting examples and the accompanying drawings, in which: [Figure 1] 1 shows a diagram comparing conventional clinical sequencing with TUMERIC-solo sequencing according to the present embodiment. [Diagram 2] 1 shows a schematic diagram of the TUMERIC sequencing process according to the present embodiment. [Diagram 3] 1 shows a flow diagram of the overall TUMERIC-solo process according to the present embodiment. [Figure 4] 4 shows a flow diagram of the TUMERIC-solo tumor purity estimation process of FIG. 3 according to a present embodiment. [Diagram 5] 4 shows a flow diagram of TUMERIC-solo transcriptome deconvolution in FIG. 3 according to the present embodiment. [Figure 6]Examples of tumor transcriptome deconvolution according to this embodiment are shown: FIG. 6a shows estimated tumor purity values for approximately 8000 bulk tumor samples across 20 solid tumor types; FIG. 6b shows genes specifically expressed in cancer and stromal cells across cancer types; as expected, only genes specific to cancer cells are affected by DNA copy number changes in the corresponding tumors; FIG. 6c shows inferred cancer and stromal compartment expression levels for 280 known stroma-specific genes; FIG. 6d shows inferred cancer and stromal compartment expression levels in melanoma (cutaneous melanoma - SKCM) and cancer and stromal compartment expression levels previously identified by melanoma single cell RNA sequencing (scRNA sequencing). Figure 6e shows genes and pathways ordered by inferred expression differences between the cancer and stromal compartments in each tumor type; Figure 6f shows protein expression inferred for the cancer and stromal compartments in (OV) and breast (BRCA) cancer cohorts using iTRAQ protein quantification data and compared to RNA sequencing data from the same tumors; Figure 6g shows genes with highly variable cancer mRNA expression differences relative to stromal mRNA expression across cancer types where identified, immunohistochemistry (IHC) staining data was compared to deconvoluted RNA sequencing data for the gene with the highest mRNA abundance (S100A6). [Figure 7]7a shows the results of inferring crosstalk between cancer and stromal cells according to the present embodiment, where FIG. 7b shows the Relative Crosstalk (Relative Crosstalk) function estimating the relative flow of signaling in four possible directions between the cancer and stromal cell compartments, including the bulk (non-deconvoluted) normal tissue signaling estimates. Figure 7b shows the median RC scores across 20 solid tumor types estimated and plotted per signaling direction; Figures 7c and 7d show the five ligand-receptor pairs with the highest median autocrine cancer signaling scores across tumor types (Figure 7c) and the highest median paracrine signaling scores from stroma to cancer (Figure 7d), RC scores for individual pairs and by cancer type; Figure 7e shows the RC scores for canonical EGF family ligand-receptor pairs across breast cancer subtypes; Figures 7f and 7g show the estimated expression of EGF family receptors (f) and ligands (g) in cancer and stromal cell compartments across breast cancer subtypes, with undeconvoluted expression of normal tissue included for comparison. [Figure 8] An example query is shown to illustrate the process of identifying membrane protein drug targets in glioblastoma using TUMERIC. In this query, the user specifies the tumor type (glioblastoma) and further specifies the genetic / molecular subtype of the tumor to be analyzed (in this case tumors without IDH1 mutations). Known membrane proteins are then ranked by total bulk tumor expression (x-axis) and the degree to which they are specifically expressed in cancer cells (y-axis) as inferred by TUMERIC. The predicted toxicity of each target, derived from gene expression in healthy vital organs such as brain / heart / kidney, can be visualized simultaneously to aid the target selection process. [Figure 9]A diagram showing an overview of the tumor transcriptome deconvolution methodology and platform according to the present embodiment is shown, and Figure 9A shows the concept of the algorithm for tumor transcriptome deconvolution according to the present embodiment, which is utilized to infer cancer cell-specific drug targets. Figure 9B shows an overview of the components required for such a platform, such as: a large data warehouse of bulk patient tumor samples with genomic and transcriptomic data, a fast algorithm (online transcriptome deconvolution) and visualization to facilitate the discovery and identification of drug targets and biomarkers, and a query example to illustrate the process of identifying drug targets in glioblastoma. [Figure 10] 1 shows data showing that TUMERIC-Solo can estimate cancer and stromal cell expression of PD-L1 in an individual lung cancer patient (A014). TUMERIC-Solo gene expression deconvolution of PD-L1 in a single lung cancer patient (A014) according to the present embodiment compared to PD-L1 expression data inferred from a cohort of patients according to the present embodiment (total, approximately 60 lung cancer patients with TUMERIC applied); measured bulk tumor gene expression was included for comparison. [Figure 11] Shown is data from TUMERIC-Solo applied to a single lung cancer patient (A014) according to the present embodiment compared to data from a cohort of patients (overall, TUMERIC applied to approximately 60 lung cancer patients) according to the present embodiment. Deconvoluted cancer and stromal cell gene expression of 4 genes shows concordance of TUMERIC-solo of single patients and TUMERIC (overall) of multiple patients; measured bulk tumor gene expression was included for comparison. [Figure 12] Detailed data from TUMERIC-Solo applied to a sector of a single lung cancer patient's tumor (A014) according to the present embodiment, compared to data from a cohort of patients according to the present embodiment (TUMERIC applied to approximately 60 lung cancer patients). The plot shows the association between measured bulk gene expression (y-axis) and estimated tumor purity (x-axis) for three selected genes. [Figure 13] Figure 1 shows TUMERIC-Solo applied to a set of published biomarker genes associated with response to pembrolizumab treatment response. Expression of six genes in cancer and stromal cells of a single lung cancer patient (A014) was measured with TUMERIC-solo and compared with data from a cohort of patients (~60 lung cancer patients applied with TUMERIC). Measured bulk tumor gene expression was included for comparison. [Figure 14] Relative change in gene expression (signal to noise ratio) of six pembrolizumab biomarker genes as assessed using bulk or TUMERIC-solo deconvoluted gene expression in lung cancer patients (A014) is shown. Changes in gene expression are measured relative to bulk, cancer, and stromal cell expression as determined using data from a cohort of patients (TUMERIC was applied to approximately 60 lung cancer patients). Expression of PD-L1 / CD274 is compared for cancer cells, while expression of the other five biomarkers is compared for stromal cells. [Figure 15] A graph showing a patient-specific therapeutic antibody recommendation by TUMERIC-Solo according to the present embodiment (left) compared to a similar recommendation based on measured bulk gene expression (right). The graph shows the absolute (x-axis) and relative (y-axis, compared to normal lung tissue) expression of known membrane proteins in a lung cancer patient (A014). Based on the data shown, a CLDN6 antibody treatment (antibody or antibody-drug conjugate) is designated for this lung cancer patient. [Figure 16] Figure 1 shows that TUMERIC was used to identify biomarkers associated with response to pembrolizumab treatment in gastric cancer. TUMERIC analysis identified genes with strong gene expression dysregulation in cancer or stromal cells in responding (R) tumors compared to non-responding (PD) tumors. The signal-to-noise ratio (R compared to PD) measured by TUMERIC (y-axis) is shown together with the signal-to-noise ratio measured by naive bulk gene expression profiling (x-axis). [Figure 17]Data are shown for biglycan (BGN) expression in patients with different responses to pembrolizumab treatment (response, R; stable disease, SD; progressive disease, PD). Bulk tumor gene expression of BGN (left) compared to TUMERIC deconvoluted gene expression of BGN in cancer cells (right): BGN is highly overexpressed in non-responder (PD) cancer cells, with only moderate changes in bulk gene expression measurable. [Figure 20] Detailed data from TUMERIC applied to biglycan (BGN) expression in patients with different responses to pembrolizumab treatment (response, R; stable disease, SD; progressive disease, PD). The plot shows the association between measured BGN bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Figure 19] 1 shows a schematic diagram of the TUMERIC-solo sequencing process according to the present embodiment. [Figure 20] We present the landscape of available technologies for high-throughput profiling of the tumor transcriptome. Existing technologies either offer high resolution (single-cell RNA-seq, sc-RNAseq) or high scalability (e.g., immunohistochemistry IHC and bulk tumor profiling). Tumeric-Solo offers higher resolution than bulk tumor profiling (it profiles cancer and stromal cells separately) and Tumeric-Solo is easier to scale than sc-RNAseq (it can analyze FFPE samples). [Figure 21] The mathematical model underlying TUMERIC and TUMERIC-Solo is shown. The measured bulk tumor mRNA abundance (sector / part for TUMERIC-Solo) in a sample is determined by the sum of the mRNA molecules from cancer and non-cancer cells in that sample. Tumor purity can be estimated from DNA sequence data obtained from the same tumor sample / sector. [Figure 22]Breakdown of 8000 bulk tumors in the Cancer Genome Atlas (TCGA) used for the TUMERIC validation analysis. All tumors have DNA (exome sequencing) and RNA (RNA sequencing) data. [Figure 23] We show the process by which TUMERIC and TUMERIC-Solo estimate the cancer / stroma compartment ratio (tumor purity). Mutational data (DNA), copy number count data (aCGH), and / or mRNA expression data from the same tumor (sectors for TUMERIC-solo) are used to generate a consensus tumor purity estimate. Purity estimates from different methods are normalized, missing data are imputed, and estimates are averaged per sample / sector. [Figure 24] IFNG: Upregulated in the stroma of MSI and ICI-responsive tumors. Figure 24A shows expression of IFNG as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 24B shows data from TUMERIC applied to IFNG expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 25]FASLG: Upregulated in the stroma of MSI and ICI-responsive tumors. Figure 25A shows the expression of FASLG as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 25B shows data from TUMERIC applied to FALSG expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Figure 26] CXCL13: Upregulated in the stroma of MSI and ICI-responsive tumors. Figure 26A shows the expression of CXCL13 as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 26B shows data from TUMERIC applied to CXCL13 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Figure 27]ZNF683: Upregulated in the stroma of MSI and ICI-responsive tumors. Figure 27 shows the expression of ZNF683 as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 27B shows data from TUMERIC applied to ZNF683 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Figure 28] IL2RA: Upregulated in the stroma of MSI and ICI-responsive tumors. Figure 28A shows the expression of IL2RA as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 28B shows data from TUMERIC applied to IL2RA expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Figure 29]CD274 / PD-L1: Upregulated in the stroma of MSI and ICI-responsive tumors. Figure 29A shows the expression of CD274 / PD-L1 as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 29B shows data from TUMERIC applied to CD274 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 30] CPNE1: Downregulated in cancer cells of MSI and ICI-responsive tumors. Figure 30A shows the expression of CPNE1 as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 30B shows data from TUMERIC applied to CPNE1 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 31]TTC19: Upregulated in cancer cells of MSI and ICI-responsive tumors. Figure 31A shows the expression of TTC19 as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 31B shows data from TUMERIC applied to TTC19 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 32] OXCT1: Upregulated in cancer cells of MSI and ICI-responsive tumors. Figure 32A shows the expression of OXCT1 as a function of tumor purity in microsatellite unstable (MSI, dark gray dots) and stable (MSS, light gray dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 32B shows data from TUMERIC applied to OXCT1 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 33]ALDH6A1: Upregulated in cancer cells of MSI and ICI-responsive tumors. Figure 33A shows the expression of ALDH6A1 as a function of tumor purity in microsatellite unstable (MSI, dark gray dots) and stable (MSS, light gray dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 33B shows data from TUMERIC applied to ALDH6A1 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 34] COX15: Upregulated in cancer cells of MSI and ICI-responsive tumors. Figure 34A shows the expression of COX15 as a function of tumor purity in microsatellite unstable (MSI, dark grey dots) and stable (MSS, light grey dots) tumors of colorectal cancer (CRC, left), gastric cancer (STAD, center), and endometrial cancer (UCEC, right). Regression lines show TUMERIC-inferred cancer and stromal cell gene expression in each cancer type and MSI / MSS subtype. Figure 34B shows data from TUMERIC applied to COX15 expression in patients with different responses (response; stable; progression) to pembrolizumab treatment. Plots show the association between measured bulk gene expression (y-axis) and estimated tumor sample purity (x-axis) for the three treatment response groups. [Diagram 35]First, box plots are shown showing that tumor purity was estimated by different methods for approximately 8000 TCGA tumors across 20 cancer types. The median estimated purity for a given method and cancer type is plotted. Tumeric is the normalized average of AbsCN-seq, ASTAC, ESTIMATE, and PurBayes (see Methods). CPE is a consensus purity estimate from previously published TCGA samples, which was included for comparison. To explore the agreement of the different purity estimation methods, we clustered the methods based on Pearson correlation (1-r) and Ward linkage, whose data are provided in the second part of Figure 35. CPE is primarily based on purity estimates from ESTIMATE, and therefore these two methods cluster tightly together (r=0.83) as expected. In the third part of Figure 35, box plots show the estimated tumor purity values for approximately 8000 bulk tumor samples across 20 solid tumor types. Consistent with previous observations, pancreatic adenocarcinoma (PAAD) tumors had a very low mean purity (approximately 39%). Glioblastoma (GBM) and ovarian cancer (OV) samples had the highest purity estimates, likely due to tumor selection bias in phase 1 of the TCGA project. [Diagram 36] Deconvolution of fibroblast activation protein alpha (FAP) gene expression across different cancer types. Estimated cancer (C) and stromal (S) cell gene expression (log2 FPKM+1) for each cancer type is shown. [Figure 37] Deconvolution of T cell surface glycoprotein CD3 delta chain (CD3D) gene expression across different cancer types. Estimated cancer (C) and stromal (S) cell gene expression (log2 FPKM+1) for each cancer type is listed. [Figure 38] Deconvolution of CD4 gene expression across different cancer types. Estimated cancer (C) and stromal (S) cell gene expression (log2 FPKM+1) for each cancer type is listed. [Figure 39]Deconvolution of colony-stimulating factor 1 receptor (CSF1R) gene expression across different cancer types. Estimated cancer (C) and stromal (S) cell gene expression (log2 FPKM+1) for each cancer type is listed. [Diagram 40] Deconvolution of epithelial cell adhesion molecule (EPCAM) gene expression across different cancer types. Estimated cancer (C) and stromal (S) cell gene expression (log2 FPKM+1) for each cancer type is shown. [Diagram 41] Heatmap of normalized enrichment scores (NES) of MSigDB Hallmark Gene Sets obtained by GSEA prerank analysis of log2((cancer_FPKM+1) / (stroma_FPKM+1)) after deconvolution. Immune system related pathways such as inflammatory response, interferon alpha / gamma response are upregulated in stroma, while known cancer cell specific pathways such as MYC targets, G2M checkpoint, DNA repair are upregulated. Red / blue cells have FDR<=0.25 and white cells have FDR>0.25. [Diagram 42] a) We identified genes with highly variable mRNA expression differentials in cancer compared to stroma across cancer types, and b) we compared immunohistochemistry (IHC) staining data with RNA-seq data for the highest (S100A6) and second highest (LDHB) abundant genes. [Diagram 43]Deconvolution of gene expression for estrogen receptor 1 (ESR1) in invasive ductal carcinoma (IDC) luminal A (IDC_LumA), luminal B (IDC_LumB), basal (IDC_Basal) and HER2 (IDC_Her2) breast cancer subtypes (first graph). As expected, ESR1-negative HER2 and basal subtypes express ESR1 lowly. Similarly, ESR1-positive subtypes LumA and LumB express ESR1 very highly in cancer cells (fpkm ~387 for LumA and ~221 for LumB). Deconvolution of ERBB2 / HER2 expression in basal (left) and HER2+ (right) subtypes (second and third graphs). Deconvolution of ERBB2 expression in HER2 tumors also took into account frequent HER2 amplification events (see Methods). [Diagram 44] Gene set enrichment analysis (GSEA) comparing gene expression of the cancer compartment (i.e., gene expression of cancer cells) in luminal A (luma), luminal B (lumb), and HER2 (her2) tumors with basal-type tumors. Heatmap shows GSEA normalized enrichment scores (NES) for individual gene sets for three different subtype comparisons, with blue (negative NES) reflecting gene sets that are upregulated in basal-type compared to other cancer types. Light / dark grey cells have FDR<=0.25, whereas white cells have FDR>0.25. [Diagram 45] Comparison of deconvolution using linear and log-transformed RNA-seq gene expression (FPKM+1). Plot shows the top 5% coefficient of determination (R2) of purity versus gene expression obtained for each transformation. Across all cancer types, tumor purity is overall more linearly correlated with log-transformed RNA-seq gene expression data. [Figure 46] Shown is an IHC image of a BLCA (bladder urothelial carcinoma) tumor sample stained with S100A6. [Figure 47] Shown is an IHC image of a BLCA (bladder urothelial carcinoma) tumor sample stained with S100A6. [Figure 48] Shown is an IHC image of a LIHC (hepatocellular carcinoma) tumor sample stained with S100A6. [Figure 49] Shown is an IHC image of a LIHC (hepatocellular carcinoma) tumor sample stained with S100A6. [Figure 50] Shown is an IHC image of a PAAD (pancreatic adenocarcinoma) tumor sample stained with S100A6. [Figure 51] Shown is an IHC image of a PAAD (pancreatic adenocarcinoma) tumor sample stained with S100A6. [Figure 52] Shown is an IHC image of a PRAD (prostate adenocarcinoma) tumor sample stained with S100A6. [Figure 53] Shown is an IHC image of a PRAD (prostate adenocarcinoma) tumor sample stained with S100A6. [Figure 54] Shown is an IHC image of a PRAD (prostate adenocarcinoma) tumor sample stained with LDHB. [Figure 55] Shown is an IHC image of a PRAD (prostate adenocarcinoma) tumor sample stained with LDHB. [Figure 56] Shown is an IHC image of a PAAD (pancreatic adenocarcinoma) tumor sample stained with LDHB. [Figure 57] Shown is an IHC image of a PAAD (pancreatic adenocarcinoma) tumor sample stained with LDHB. [Figure 58] Shown is an IHC image of an OV (ovarian serous cystadenocarcinoma) tumor sample stained with LDHB. [Figure 59] Shown is an IHC image of an OV (ovarian serous cystadenocarcinoma) tumor sample stained with LDHB. [Figure 60] Shown is an IHC image of a LIHC (hepatocellular carcinoma) tumor sample stained with LDHB. [Figure 61] Shown is an IHC image of a LIHC (hepatocellular carcinoma) tumor sample stained with LDHB. [Figure 62] Shown is an IHC image of a HNSC (head and neck squamous cell carcinoma) tumor sample stained with LDHB. [Figure 63] Shown is an IHC image of a HNSC (head and neck squamous cell carcinoma) tumor sample stained with LDHB. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] definition As used herein, the term "tumor type" refers to tumors selected by their anatomical form, such as breast cancer or lung cancer; tumors selected by cancer type, such as carcinoma or melanoma; tumor subtypes of the same cancer type; or tumors treated with the same treatment type. Examples of such treatments are, but are not limited to, gefitinib, erlotinib and afatinib for the treatment of cancers related to EGFR; OSI-906 (linsitinib) for the treatment of cancers related to IGF1R; everolimus (also known as RAD001) and sirolimus for the treatment of cancers related to mTOR; BKM120 (buparisib) and BYL719 (alpelisib) for the treatment of cancers related to PIK3CB and PIK3R3; idelalisib for the treatment of cancers related to PIK3CD and dacomatinib and lapatinib for the treatment of cancers related to ERBB4, or combinations thereof. In one example, the anti-cancer drug used to treat EGFR-associated cancer is, but is not limited to, gefitinib, erlotinib, afatinib, or a combination thereof. In another example, the anti-cancer drug used to treat mTOR-associated cancer is, but is not limited to, everolimus (RAD001), sirolimus, or a combination thereof. In another example, the anti-cancer drug used to treat IGF1R-associated cancer is, but is not limited to, linsitinib. In another example, the anti-cancer drug used to treat PIK3CB and PIK3R3-associated cancer is, but is not limited to, BKM120 (buparisib), BYL719 (alpelisib), or a combination thereof. In another example, the anti-cancer drug used to treat PIK3CD-associated cancer is, but is not limited to, idelalisib. In another example, the anticancer drug used to treat ERBB4-related cancer is, but not limited to, dacomitinib, lapatinib, or a combination thereof.In one example, the anticancer drug is a tyrosine kinase inhibitor.In another example, the tyrosine kinase inhibitor is an EGFR inhibitor.In yet another example, the tyrosine kinase inhibitor is, but is not limited to, gefitinib, erlotinib, erlotinib HCl, lapatinib, dacomitinib, TAE684, afatinib, dasatinib, saracatinib, veratinib, AEE788, WZ4002, icotinib, osimertinib, BI1482694, ASP8273, EGF816, AZD3759, cetuximab, necitumumab, pannitumumab, nimotuzumab, and combinations thereof.In a further example, the tyrosine kinase inhibitor is, but is not limited to, gefitinib, erlotinib, lapatinib, and combinations thereof. In one example, the tumor type can be, but is not limited to, BLCA, BRCA, CESC, CRC (COAD and READ combined), ESCA, GBM, HNSC, KIRC, KIRP, LGG, LIHC, LUAD, LUSC, OV, PAAD, PRAD, SKCM, STAD, THCA and UCEC as described in the TCGA database.
[0011] As used herein, the term "scoring" refers to the process of ranking genes, biomarkers, or therapeutic targets.As used in this application, the term "scoring" can also be used synonymously with the term "ranking".For example, in a cohort of cancer patients (TUMERIC) or individual cancer patients (TUMERIC-solo), all genes can be scored or ranked according to their predicted expression in cancer cells to identify the top-ranked therapeutic target candidates.
[0012] As used herein, the term "tumor purity value" refers to the estimated proportion of cancerous cells among all cells present in a tumor. In the context of this disclosure, the terms "cancer cells" and "malignant cells" are used interchangeably. The tumor purity value of a given tumor can be estimated, for example, from the somatic mutation variant allele frequency (VAF) measured in a given sample. For example, if a known (clonal) cancer driver mutation is measured in gene X with a variant allele frequency (VAF) of 0.2 (20%), and gene X is unchanged by somatic copy number changes in the given sample (gene X has two alleles / chromosomes in cancer cells), this variant allele frequency (VAF) can be explained by a tumor that contains 40% cancer cells (one mutant allele and one wild type allele) and 60% non-cancer cells (two wild type alleles). Since many genes are mutated in tumors, the purity value is given by the consensus value that best fits all observed variant allele frequencies (VAFs).
[0013] As used herein, the term "variant allele frequency (VAF)" refers to the relative frequency of an allele (variant of a gene) at a particular locus in a population, expressed as a proportion or percentage relative to the entire population. In other words, variant allele frequency (VAF) represents the proportion of people who carry a specific allele for all chromosomes in a population.
[0014] As used herein, the terms "robust" and "accurate" can be used interchangeably.
[0015] As used herein, the term "TANTIGEN" refers to the tumor T cell antigen database developed and maintained by the Bioinformatics Core of the Cancer Vaccine Center, Dana-Farber Cancer Institute, and described in Cancer Immunol Immunother. 2017 Jun; 66(6):731-735. (doi: 10.1007 / s00262-017-1978-y. Epub 2017 Mar 9). The tumor T cell antigen database is a data source and analysis platform for cancer vaccine target discovery focused on human tumor antigens containing HLA ligands and T cell epitopes. It catalogs over 1000 tumor peptides from 292 different proteins. The database also provides information on T cell epitopes and HLA ligands with full references, gene expression profiles, antigen isoforms, and mutations. Predicted binding peptides of 15 HLA class I and class II alleles are also included in the database.
[0016] As used herein, the term "Gene Ontology" refers to the Gene Ontology Resource database, a source of information about gene function and maintained by the Open Biological Ontologies Foundry.
[0017] As used herein, the term "TCGA" refers to the Cancer Genome Atlas Program operated and maintained by the National Cancer Institute, BG 9609 MSC 9760, 9609 Medical Center Drive, Bethesda, MD 20892-9760, USA.
[0018] As used herein, the term "Human Protein Atlas" refers to a Sweden-based program launched in 2003 with the goal of mapping all human proteins in cells, tissues, and organs using the integration of various omics technologies, including antibody-based imaging, mass spectrometry-based proteomics, transcriptomics, and systems biology. All data in this intellectual resource is available online and open access, allowing scientists from both academia and industry to freely access the data for exploring the human proteome. The Human Protein Atlas consists of six separate divisions, each focusing on a specific aspect of genome-wide analysis of human proteins: the Tissue Atlas, which shows the distribution of proteins across all major tissues and organs of the human body; the Cell Atlas, which shows the subcellular localization of proteins in single cells; the Pathology Atlas, which shows the impact of protein levels on cancer patient survival; the Blood Atlas, the Brain Atlas, and the Metabolic Atlas.
[0019] As used herein, the term "cBioPortal" refers to an online portal for cancer genomics. The cBioPortal for Cancer Genomics was originally developed at Memorial Sloan Kettering Cancer Center. The public site of cBioPortal is hosted by the Center for Molecular Oncology at Memorial Sloan Kettering Cancer Center. The cBioPortal software is currently available under an open source license through GitHub. The software is currently developed and maintained by a multi-institutional team consisting of Memorial Sloan Kettering Cancer Center, Dana Farber Cancer Institute, Princess Margaret Cancer Centre in Toronto, Children's Hospital of Philadelphia, Hyve in the Netherlands, and Bilkent University in Ankara, Turkey.
[0020] As used herein, the term "Genomic Data Commons" refers to a research program of the National Cancer Institute (NCI; NCI Center for Cancer Genomics (CCG), 31 Center Drive, Bldg. 31, Suite 3A20, Bethesda, MD 20892).
[0021] As used herein, the term "cancer compartment" refers to cancer cells.For example, as used herein, Tumeric-solo is used to estimate / predict the expression of genes in cancer cells / compartments.Gene is ranked / ordered from high to low based on this estimated cancer expression level.
[0022] The embodiments exemplified and described herein may suitably be practiced in the absence of any element or elements, or any limitation or limitations not specifically disclosed herein. Thus, for example, terms such as "comprising", "including", "containing" and the like are to be read expansively and without limit. Additionally, the terms and expressions used herein are used as terms of description and not of limitation, and in the use of such terms and expressions, there is no intention to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the claimed invention. Thus, although the present invention has been specifically disclosed by the present embodiments and any features, it should be understood that modifications and changes of the embodiments embodied therein disclosed herein may be exercised by those skilled in the art, and such modifications and changes are considered to be within the scope of the present invention.
[0023] As used in this application, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "genetic marker" includes a plurality of genetic markers, including mixtures and combinations thereof.
[0024] As used herein, the term "about" in reference to a concentration of a formulation component typically means ±5% of the stated value, more typically ±4% of the stated value, more typically ±3% of the stated value, more typically ±2% of the stated value, even more typically ±1% of the stated value, and even more typically ±0.5% of the stated value.
[0025] Throughout this disclosure, certain embodiments may be disclosed in a range format. Descriptions in range format should be understood merely for convenience and brevity, and should not be interpreted as rigid limitations on the disclosed range range. Thus, descriptions of ranges should be considered to have specifically disclosed all possible subranges and individual numerical values within that range. For example, descriptions of ranges such as 1-6 should be considered to have specifically disclosed subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0026] Certain embodiments may be described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also forms part of the disclosure. This includes the generic description of embodiments with a condition or negative limitation that removes any matter from the genus, regardless of whether the excised material is specifically described herein.
[0027] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also forms part of the invention. This includes the generic description of the invention with a condition or negative limitation that removes any matter from the genus, regardless of whether the excised material is specifically described herein.
[0028] Detailed Description of the Invention Described herein is an approach for quantifying genome-wide, high-throughput molecular activity (e.g., mRNA, DNA methylation, or protein expression) in cancer cells and non-cancer cells of individual patient tumors, with particular application for discovering new biomarkers and treating individual patients based on abnormal molecular activity. Signaling between cancer cells and non-malignant (e.g., stromal) cells in the tumor microenvironment is difficult to study in patient tumors. Thus, disclosed herein is a data-driven method for deconvolution of cancer and stromal cell transcriptomes in bulk tumor tissue and inference of cell-cell signaling crosstalk. Using this approach, common crosstalk across different solid tumor types and inferred modes of EGF family crosstalk in breast cancer subtypes are advantageously identified in bulk tumor tissue. This method is further demonstrated to be advantageous for recommending new drug targets, recommending patient-specific treatments, and identifying and / or quantifying biomarkers for immune checkpoint inhibition anti-cancer therapy.
[0029] According to the present embodiment, a combined experimental-computational method / algorithm (hereinafter referred to as "TUMERIC-solo") for inferring cancer and non-cancer molecular activity in individual bulk tumor samples is disclosed. The combined experimental-computational method / algorithm according to the present embodiment can be applied to any type of molecular data (e.g., mRNA expression (RNA sequencing), expression of mRNA transcript isoforms, protein expression (using iTRAQ), or epigenetic profiling) co-extracted from different physical sections / sectors of a bulk tumor sample. The combined experimental-computational method / algorithm according to the present embodiment requires as input both DNA data and molecular data from N sectors of a single bulk tumor sample and outputs estimates of molecular activity / expression in cancer and non-cancer cells of that tumor sample. The data disclosed herein below validates the use of the combined experimental-computational method / algorithm according to the present embodiment for RNA sequencing and protein using a cohort of bulk tumor samples from different patients.
[0030] The combined experimental and computational method / algorithm according to this embodiment also encompasses a method for treating a patient's tumor based on specific molecular signals in cancer or non-cancer cells of individual tumors.For example, a patient's tumor sample can be analyzed by TUMERIC-solo, and the patient can be treated according to the measured molecular activity in cancer cells (e.g., tamoxifen for ESR1-positive breast tumors, PDL1-positive for checkpoint blockade immunotherapy) or non-cancer cells (e.g., PDL1-positive for checkpoint blockade immunotherapy in gastrointestinal tumors).For example, the latter may be relevant for future immunotherapy.
[0031] The inventors are not aware of any method in the art that allows deconvolution of cancer cell mRNA expression in a single patient. The combined experimental and computational method / algorithm according to this embodiment requires the tumor sample to be physically sectioned into N parts or sectors. It is understood that the accuracy of the method according to this embodiment increases with the number N of parts or sectors of the tumor sample (e.g., for N greater than 5 to 10). However, it is also understood that some tumor samples may possibly be too small / fragile for such sectioning.
[0032] Figure 1 shows a diagram 100 comparing a conventional clinical sequencing operation 102 with the present embodiment's TUMERIC-solo sequencing operation 104. As an alternative to deconvolution using transcriptional signatures, the proportion of cancer cells (tumor purity) is first estimated from the tumor's mutant allele frequency and copy number profile, and averaged to form a consensus tumor purity value. Importantly, the present embodiment avoids making assumptions about the transcriptional profiles of cancer and stromal cells found in a given tumor (see also Figure 23).
[0033] Examples of procedures for estimating tumor purity from DNA and CNA data can be found, for example, in the following publications: Bao, L., Pu, M., and Messer, K. AbsCN-seq: a statistical method to estimate tumor purity, ploidy and absolute copy numbers from next-generation sequencing data. Bioinformatics 30, 18 1056-1063; Larson, N., and Fridley, B. PurBayes: estimating tumor cellularity and subclonality in next-generation sequencing data. Bioinformatics 29, 1888-1889. Estimation of purity from gene expression data is presented in the following publication: Yoshihara, K., Shahmoradgoli, M., Martinez, E., Vegesna, R., Kim, H., Torres-Garcia, W., Trevino, V., Shen, H., Laird, PW, Levine, DA, et al. (2013). Inferring tumour purity and stromal and immune cell admixture from expression data. Nature Communications 4, 2612.
[0034] Thus, in one example, the method disclosed herein predicts the expression profile of cancerous cells and non-cancerous cells, respectively, based on a set of expression profiles, where each set of expression profiles is obtained from a tumor-derived sample that comprises a mixture of cancerous cells and non-cancerous cells of one tumor type.In another example, the method disclosed herein includes the steps of: determining a tumor purity value for one or more tumor-derived samples; providing different sets of expression profiles, where the different sets of expression profiles comprise composite expression data for a plurality or all of the molecules expressed by cancerous cells and non-cancerous cells contained in one or more tumor-derived samples; and deconvoluting each composite expression data obtained by the method disclosed herein by extrapolating the expression profiles of a plurality or all of the molecules expressed in the different tumor samples with different tumor purity values to a tumor purity value at least substantially equal to 1 or 0, thereby predicting the expression profile of cancerous cells and non-cancerous cells, respectively, from the set of expression profiles.
[0035] In one example, the molecule can be, but is not limited to, a gene, a DNA, an RNA, or a protein molecule, or a combination thereof.
[0036] In another example, the methods disclosed herein may further comprise scoring the molecules disclosed herein based on their level of up- or down-regulation in the cancer tissue compared to the stromal tissue; and / or scoring the molecules disclosed herein based on their level of up- or down-regulation in the cancer tissue compared to the healthy tissue.
[0037] In another example, the methods disclosed herein include assigning the up-regulated and down-regulated molecules to gene or transcript isoforms for a known dataset of membrane-bound proteins or receptors; and / or assigning the up-regulated and down-regulated molecules to gene or transcript isoforms for a known dataset of HLA-binding peptides and T-cell antigen-binding peptides.
[0038] In one example, known datasets for assigning genes or transcript isoforms are derived from, for example, but not limited to, Gene Ontology, Human Protein Atlas, and / or TANTIGEN.
[0039] In another example, a gene or transcript isoform disclosed herein can be, without limitation, a membrane-bound protein, a membrane-bound receptor, an antigenic peptide, a target protein, a peptide, and / or can be targeted by an antibody.
[0040] Sequencing according to this embodiment, when combined with large-scale genomic and molecular data from human tumors (e.g., from TCGA or clinical trials), allows for the estimation of cancer-specific molecular profiles (mRNA, epigenetic, or protein abundance) for target and biomarker discovery using bulk human tumor tissue.
[0041] In one example, providing the different sets of expression profiles includes using an existing dataset of expression profiles. In such cases, the existing dataset of expression profiles is derived from a database such as, but not limited to, TCGA, Genomic Data Commons, cBioPortal, and / or ICGC database.
[0042] Tumor molecular profiles have been deconvoluted into cancer and stromal cell components using a constrained linear regression approach as described in TUMERIC-solo sequencing104 and in more detail below. Inferred cancer and stromal compartment expression profiles are combined with a curated database of ligand-receptor interactions to infer autocrine and paracrine signaling crosstalk between these two compartments in the tumor microenvironment (TME).
[0043] While new computational methods allow the inference of cell type proportions from bulk tumor mRNA profiles using knowledge of primary cell type transcriptional signatures, conventional implementations of these methods generally focus on deconvolution of specific immune cell types and do not provide estimates of gene expression in individual cell types. Previous approaches to estimate cancer and stromal cell gene expression profiles in tumor tissues are either heavily customized to individual tumor types or assume that tumors are a mixture of cancer cells and healthy tissue. The customization of individual tumor cells limits the use of such methods, and assuming that tumors are a mixture of cancer cells and healthy tissue ignores unique stromal cell types and biological processes of the tumor microenvironment, which can strongly confound the inferred gene expression profile.
[0044] Few laboratory techniques exist that allow for the discrimination of signals from cancer cells and non-cancerous cells in the tumor microenvironment. Immunohistochemistry (IHC) can directly measure selected proteins in tumor tissue, but is generally not quantitative and is not suitable for large-scale, unbiased profiling or discovery. Furthermore, IHC is labor intensive and requires a skilled pathologist to help interpret the data.
[0045] Transcriptome-wide profiles of cancer and stromal cells may be generated using microdissection or single-cell profiling of tumor tissue, but these approaches are difficult to apply to tumor biopsies, and dissociation may perturb cell physiology and gene expression profiles to some extent. Moreover, these methods require special handling and processing of tissue, which makes them less suitable as standard data-generating assays in precision oncology.
[0046] Targeted exome sequencing is becoming a routine diagnostic assay, with companies offering clinical sequencing as a service. See, for example, Figure 20. As the cost of sequencing continues to fall, companies are now also offering whole exome sequencing and RNA sequencing as clinical diagnostic services. Importantly, these services are scalable, since they only require frozen or formalin-fixed paraffin-embedded (FFPE) tumor tissue and next-generation sequencing (NGS). However, bulk exome sequencing and RNA sequencing do not allow for direct measurement of cancer cell populations in tumors. This is important, for example, to determine breast cancer patients with estrogen-positive tumors (due to tamoxifen treatment) or tumors with increased PDL1 expression in the cancer cells (PD1 / PDL1 checkpoint inhibition).
[0047] TUMERIC is a method to estimate the cancer and stromal (including any non-cancer cells) compartment molecular profiles for a set of tumors, and the crosstalk signaling between average representative cells in these two compartments. With reference to FIG. 2, a schematic diagram 200 of the TUMERIC sequencing process according to the present embodiment begins with the estimation of tumor purity 210. The purity (proportion of cancer cells) of each bulk tumor sample is estimated 210 from DNA data (exome sequencing), copy number data (aCGH), and mRNA expression data (RNA sequencing) using a consensus approach. Next, a deconvolution 220 of the mRNA expression levels in the "average" cancer and stromal cells is inferred for a given gene and tumor set (e.g., representing a tumor type) using non-negative least squares regression. Finally, the resulting mRNA expression profile is used 230 to infer candidate autocrine and paracrine signaling pathways between cancer and stromal cells using a database of curated receptor-ligand signaling interactions.
[0048] Thus, in one example, the method disclosed herein can include, but is not limited to, determining a tumor purity value based on the distribution of somatic DNA variant allele frequency, somatic DNA copy number change amplitude, germline B allele frequency, gene expression signature or pattern, protein expression signature or pattern, and DNA methylation signature or pattern, and the distribution of their combinations.In one example, the tumor purity value is based on gene expression signature (or gene expression profile).In another example, the tumor purity value is based on allele frequency, for example, somatic DNA variant allele frequency and / or germline B allele frequency.In another example, the tumor purity value is based on methylation signature.
[0049] In one example, at least two, or at least three, or at least four, or at least five, or two, or three, or four, or five, or all of the methods disclosed herein are used together to determine the average tumor purity.
[0050] In another example, the tumor purity value is the average tumor purity value.
[0051] In one example, the tumor types referred to herein can be, but are not limited to, BLCA, BRCA, CESC, CRC (a composite of COAD and READ), ESCA, GBM, HNSC, KIRC, KIRP, LGG, LIHC, LUAD, LUSC, OV, PAAD, PRAD, SKCM, STAD, THCA, and UCEC as described in the TCGA database.
[0052] Referring to FIG. 3, flowchart 300 discloses the TUMERIC-solo method according to this aspect. First, for example, using a microtome, cryosectioning, or a frozen tumor array, a frozen tumor sample is divided into N sectors at 302 (e.g., the value of N is greater than 5 but less than 20 (5 < N < 20)). DNA data and RNA data are simultaneously extracted from each sector, barcoded, pooled, and profiled by next-generation sequencing. The obtained next-generation DNA sequencing data 304 and the obtained next-generation RNA sequencing data 306 are demultiplexed, and the DNA sequencing data 304 and the RNA sequencing data 306 are used to estimate 308 the proportion of cancer cells (tumor purity) 310 for each sector using the mutant allele frequency and copy number profile and transcriptional signature of each sector. As described below and as shown in Equation (1) (see also FIG. 21), the purity data 310 (p i ) and the RNA sequencing data 306 (E 腫瘍,i ) at the sector level are deconvolved at 312 to obtain the cancer cells 314 (E がん ) and non-cancer cells 316 (E間質 ) to predict the molecular profile of TIFF2025026905000002.tif5128
[0053] The cancer cell 314 and non-cancer cell 316 profiles can be used to provide recommendations 318 for immune checkpoint inhibitors. Additionally, recommendations 322 for antibody-based targeting of cancer cells can be determined and prioritized from the cancer cell profile 314 using cross-referencing a database 320 of known membrane proteins and antigens and intercellular signaling.
[0054] 4, a flow diagram 400 illustrates a TUMERIC-solo tumor purity estimation process 308 according to the present embodiment. The purity of each bulk tumor sample (i.e., the percentage of cancer cells in each sample) is estimated by first estimating purity from DNA sequencing data 304 and RNA sequencing data 306 using three methodologies: purity is estimated from DNA sequencing data 304 using somatic variant allele frequencies 402 and DNA copy number alterations and B allele frequencies 404, and purity is estimated from RNA sequencing data 306 using gene expression signatures of epithelial and immune / stromal infiltrating cells 406. If none of the estimation methodologies 402, 404, 406 converge or yield estimates that are too high or too low, the purity value estimates are imputed 408 using a statistical method for imputation (e.g., average, regression, or k-nearest neighbors). If the estimate from one of the three methodologies 402, 404, 406 is very high (e.g., >98%) but the other estimates from the methodologies 402, 404, 406 are not very high (e.g., <95%), the estimate is considered to be too high. Similarly, if the estimate from one of the three methodologies 402, 404, 406 is very low (e.g., <10%) but the other estimates from the methodologies 402, 404, 406 are not very low (e.g., >20%), the estimate is considered to be too low. The final step in tumor purity estimation 308 to infer 310 an average tumor purity estimate for each of the N tumor sectors is normalization 410 of the purity distribution. Normalization 410 aligns the different estimated purity distributions. This may be done by using quantile normalization or other normalization techniques and / or by weighting each estimate by its correlation with the average consensus estimate, such that estimates that are more highly correlated with the average consensus estimate are weighted more heavily during normalization. Normalization 410 may also filter out purity estimate distributions that deviate too much from the average consensus estimate.
[0055] Thus, in one example, tumor-derived sample is obtained from a single subject.In another example, tumor-derived sample is divided into two or more sections.In yet another example, tumor-derived sample is divided into two or more sections, and one set of expression profiles is generated for each section.
[0056] 5, a flow diagram 500 illustrates TUMERIC-solo transcriptome deconvolution 312 according to the present embodiment. To infer the molecular profiles of cancer cells 314 and non-cancerous cells 316 in the original tumor sample, the purity data 310 (i.e., tumor purity estimates for each of the N tumor sectors) and the sector-wise RNA sequencing data 306 are deconvoluted 312. The deconvolution 312 includes transcriptome deconvolution 502 of the tumor purity estimates 310 and the RNA sequencing data 306, whose expression has been summarized 404 at the level of genes, transcript isoforms, or exons.
[0057] In one example, expression profile can be, but is not limited to, gene expression, RNA expression, epigenetic expression, protein expression, proteome expression, and combinations thereof, for example, RNA expression and epigenetic expression, and RNA expression and protein expression.In another example, expression profile is gene expression profile.In another example, expression profile is RNA expression profile.
[0058] Using the tumor purity estimate (p) 310 and the RNA sequencing data 306, the transcriptome deconvolution 402 advantageously uses generalized linear model (GLM) regression to infer cancer (E_cancer) compartment expression 314 and stromal (E_stromal) compartment expression 316 at the gene level, transcript isoform level, or exon level from the measured bulk RNA data (E_obs) 306, where the RNA data is expressed as Equation 2: 404 summarized as shown in TIFF2025026905000003.tif4128.
[0059] If the expression data 314, 316 is summarized as fragments / reads per kilobase of transcript per million mapped reads (FPKM / RPKM), a normal distribution link function may be used in the generalized linear model (GLM) according to the present embodiment, and the observed data may be in linear or log scale. If the expression data 314, 316 is summarized as read counts, a Poisson distribution, negative binomial distribution, or other overdispersed exponential family of distributions may be used as the link function in the generalized linear model (GLM) according to the present embodiment.
[0060] FIG. 6 shows an example validating tumor transcriptome deconvolution according to the present embodiment, and as shown in FIG. 6a, consensus tumor purity estimates were obtained for approximately 8000 samples across 20 solid tumor types in the Cancer Genome Atlas (TCGA), showing that the majority of tumor samples had 40-70% purity. Pancreatic adenocarcinoma (PAAD) tumors were shown to have very low purity (average purity approximately 39%), consistent with previous observations. Glioblastoma (GBM) and ovarian cancer (OV) cohorts had the highest purity estimates, likely due to tumor selection bias in the first phase of the Cancer Genome Atlas project. It was found that mRNA expression-derived tumor purity estimates, as well as previously published consensus tumor purity estimates, correlated well with TUMERIC consensus purity estimates, but may have systematically overestimated purity by 20-50% compared to mutation and copy number-based methods (FIG. 35).
[0061] Figure 6b shows the genes inferred for each tumor type that are differentially expressed in cancer and stromal cells. Correlation between mRNA expression at each locus and somatic copy number alterations (CNAs) was assessed (top column). Tumor types were ranked by the mean difference in correlation of cancer and stromal genes, and the fraction of the genome altered by CNAs was determined for each tumor sample (bottom column). Multiple analyses were performed to evaluate the accuracy of TUMERIC in deconvolving the transcriptomes of the cancer and stromal cell compartments. First, since somatic copy number alterations (CNAs) are a hallmark of the cancer cell genome, without being bound by theory, we reasoned that the expression of genes derived exclusively from stromal cells should not be affected by such alterations. Indeed, using TUMERIC to infer the top cancer- and stromal cell-specific genes in each tumor type, strong correlations were found between tumor copy number alterations and the expression of cancer-specific genes, but not between tumor copy number alterations and the expression of stromal-specific genes. Variation in correlations between tumor types can also be explained by the overall prevalence of copy number alterations in a given tumor type. Similarly, we found that TUMERIC consistently inferred that previously obtained stromal and immune cell specific genes were significantly more highly expressed in the stromal compartment of all tumor types, as shown in Figure 6c. In Figure 6c, inferred cancer and stromal compartment expression levels for 280 known stromal specific genes are shown.
[0062] To test the concordance of TUMERIC with tumor single-cell RNA sequencing (scRNA-seq) profiling, we compared TUMERIC expression estimates for cancer- and stromal cell-specific genes identified by single-cell RNA sequencing of melanoma. Figure 6d shows inferred cancer compartment and stromal compartment expression levels in melanoma (cutaneous melanoma - SKCM), as well as bulk tumor measurements, for cancer- and stromal-specific genes previously identified in melanoma single-cell RNA sequencing (scRNA-seq). TUMERIC inferred significantly higher stromal compartment expression for stromal cell-specific genes (P=2e-55, Mann Whitney, two-tailed), and significantly higher cancer compartment expression for cancer cell-specific genes (P=3.6e-4).
[0063] The possible biological functions of genes with cancer- or stroma-specific expression across tumor types were evaluated using gene set enrichment analysis (see Methods section). Gene sets consistently upregulated in the cancer compartment across tumor types were associated with known features of cancer cells, such as activation of cell cycle, MYC signaling, metabolism, and DNA repair. Figure 6e shows genes ranked by inferred expression difference between the cancer and stromal compartments in each tumor type. Gene set enrichment analysis (GSEA) was used to identify gene sets enriched in cancer and stroma. Non-significant associations (false discovery rate (FDR) > 0.25) are displayed in white. In contrast, gene sets consistently upregulated in the stromal compartment across all cancer types included genes related to angiogenesis, immune response, and mesenchymal cell state.
[0064] To assess the extent to which the deconvoluted mRNA profiles represent an accurate surrogate for protein levels in cancer and stromal cells, TUMERIC was applied to deconvolute protein expression data from TCGA tumors. Figure 6f shows protein expression inferred for the cancer and stromal compartments in the (OV) and breast cancer (BRCA) cohorts using iTRAQ protein quantification data and compared to RNA sequencing data, and it was found that the mRNA expression estimates were generally consistent with the relative levels of protein abundance in cancer and stroma.
[0065] Finally, Figure 6g shows genes identified as having highly variable mRNA expression differences in cancer compared to stroma across cancer types, comparing immunohistochemistry (IHC) staining data with RNA sequencing data for the gene with the highest mRNA abundance (S100A6) to confirm that the expression pattern of one such gene was indeed variable across tumor types (Figure 6g).
[0066] Referring to FIG. 7, the results of inferring crosstalk between cancer cells and stromal cells according to the present embodiment are shown. To infer and distinguish between the types of autocrine (signaling within the same compartment) and paracrine (signaling between cancer cell and stromal cell compartments) ligand-receptor (LR) crosstalk in tumors, a metric, the relative crosstalk (RC) score, was developed. This relative crosstalk (RC) score estimates the relative flow of signaling in four possible directions between cancer cell and stromal cell compartments, including bulk (non-deconvoluted) normal tissue signaling estimates, and quantifies the directionality of relative signaling for a given ligand-receptor pair, as shown in FIG. 7a, and makes several simplifying assumptions about cell-cell signaling (e.g., ignoring local competition and saturation effects). Nevertheless, the relative crosstalk (RC) score is a reasonable approximation for determining the directionality of relative signaling in tumors.
[0067] First, the extent to which some ligand-receptor pairs exhibited a consistent mode of crosstalk across tumor types was assessed, with differences in relative crosstalk scores found between the cancer and stromal compartments. As shown in Figure 7b, 264 ligand-receptor pairs were found to have high autocrine stromal signaling scores, whereas only three ligand-receptor pairs showed evidence of strong autocrine cancer signaling across tumor types (median cancer-to-cancer RC score >40%). This suggests that for solid tumors, autocrine cancer signaling tends to be tumor type specific and may be determined by the cancer cell of origin, whereas stromal autocrine signaling is generally independent of tumor type and site of origin. Interestingly, the paracrine signaling interface between the cancer and stromal cell compartments also had a large number of recurrent interactions (26 and 40 interactions with median RC score >40% for cancer-to-stroma and stromal-to-cancer signaling, respectively), highlighting the importance of the tumor environment on the biology of cancer cells. As shown in Figure 7c, recurrent autocrine cancer signaling predicted across tumor types included signaling through FGFR8, LRP6, and MST1R. Of note, MST1R (RON) has been found to be a prognostic marker and is currently being evaluated as a therapeutic target in a range of tumor types. Signaling through ACVR2B, in particular, was included in both the top cancer autocrine signaling interactions predicted across tumor types and stroma-to-cancer signaling interactions (Figures 7c and 7d).
[0068] As another example, about 130 lung adenocarcinoma tumor samples were analyzed using the methods disclosed herein. All samples had exome (DNA) and RNA sequencing data. A patient tumor sample (A014) that was divided into 8 independent sectors and then subjected to the TUMERIC-solo analysis workflow was also analyzed. The methodology according to the present embodiment was further used to study the role of EGF family signaling across breast cancer subtypes, as shown in FIG. 7e. As shown in FIG. 7f, a 30-fold increased expression of ERBB2 was predicted in cancer cells of HER2-positive tumors. Focusing on the canonical EGF family LR interaction and the predicted signaling through this receptor, it was found that both cancer cell and stromal cell EGFR expression was generally reduced in tumors compared to normal breast tumors (FIG. 7e and 7f). EGFR expression was predicted to be expressed in cancer cells of basal and HER2-positive tumors, but was nearly absent in cancer cells of luminal A and B tumor subtypes (FIG. 7e and 7f). Amphiregulin (AREG) appeared to be the main source of EGFR ligand (Fig. 7g). It is noteworthy that AREG was speculated to be expressed mainly by stromal cells in both luminal subtypes, whereas AREG was expressed almost exclusively by cancer cells in basal and HER2-positive tumors (Fig. 7g). This data supports the existence of a cancer cell autocrine feedback loop between AREG and EGFR that is unique to HER2-positive and basal breast tumors and demonstrates how this approach can be applied to study cell-cell crosstalk associated with specific molecular or genetic subtypes of tumors.
[0069] In summary, a data-driven method is provided herein for deconvoluting cancer and stromal cell transcriptomes and inferring cell-cell crosstalk in the tumor microenvironment using only bulk genomic and transcriptomic data from a set of tumors. The methods disclosed herein are not limited by transcriptomic data and can be advantageously used with other types of bulk tumor molecular data, such as, but not limited to, epigenetic or proteomic profiles.
[0070] Validation of the TUMERIC-solo approach First, the ability of TUMERIC and TUMERIC-solo to quantify cancer and stromal expression for known marker genes was evaluated. Referring to FIG. 8, this figure shows an example query to illustrate the process of identifying membrane protein drug targets in glioblastoma using TUMERIC. In this query, the user specifies the tumor type (glioblastoma) and further specifies the genetic / molecular subtype of the tumor to be analyzed (here, tumors without IDH1 mutations). Known membrane proteins are then ranked by their overall bulk tumor expression (x-axis) and the extent to which TUMERIC predicts that they are specifically expressed in cancer cells (y-axis). The predicted toxicity of each target, derived from gene expression in healthy vital organs such as the brain / heart / kidney, can be visualized simultaneously to aid the target selection process.
[0071] Referring to Figure 9A, diagram 910 represents an overview of the tumor transcriptome (or proteome) deconvolution methodology and platform according to the present embodiment. Figure 9B shows work packages WP1 920, WP2 930, and WP3 940 together with an overview 950 of the methodology according to the present embodiment.
[0072] Referring to FIG. 11, we show bar graph data from TUMERIC-Solo when applied to a single lung cancer patient (A014) compared to data from a cohort of patients (TUMERIC was applied to approximately 60 lung cancer patients). First, the data demonstrates that TUMERIC-Solo can reliably identify known stromal factors (such as CD3D, CD68, which are overexpressed in stroma compared to cancer) and epithelial / cancer markers (EGFR, EPCAM). Second, using TUMERIC-Solo, we showed that while expression of PDL1 (CD274) is commonly expressed in the stroma, in patient A014 PDL1 was overexpressed by more than six-fold (>6-fold) in cancer cells, while stromal expression remained unchanged (FIG. 10). This identifies blockade of the PD1 / PDL1 checkpoint as a potential target in this particular patient. If the same analysis were performed using bulk tumor profiling, an approximately two-fold upregulation of PDL1 expression would be observed, although it would be unclear whether the increased expression levels were due to stromal cells or cancer cells overexpressing PDL1.
[0073] This demonstrates how TUMERIC and TUMERIC-solo can provide concordant results, even though TUMERIC uses data obtained from different patient tumors and TUMERIC-solo uses data from different sections of one individual patient tumor. To further illustrate this concept and concordance, we illustrated these two deconvolution approaches by plotting the measured (bulk) gene expression of CD68, CD74, and EPCAM as a function of estimated sample / sector tumor purity for TUMERIC (N=130 samples) and TUMERIC-solo (N=8 sectors of patient tumor A014), respectively (Figure 12). While the analyzed and inferred gene expression levels are overall concordant between TUMERIC and TUMERIC-solo, the analysis of CD74 demonstrates how well TUMERIC-solo can infer patient-specific changes in gene expression.
[0074] Prediction of patient-specific PDL1 expression using TUMERIC-solo Tumor PDL1 (CD274) expression is a biomarker of immune checkpoint inhibition treatment response in lung cancer. However, PDL1 checkpoint inhibition only works in a subset of patients (<20%), and it is debated whether cancer or stromal cells predominantly overexpress PDL1 in patients who benefit from the treatment. TUMERIC-solo analysis of A014 tumors identified that PDL1 is highly upregulated in cancer cells but not in stromal cells. Of note, PD-L1 upregulation is a phenomenon specific to A014 patients and was not observed in TUMERIC analysis of 130 patient tumors, highlighting the added value of TUMERIC-solo. Taken together, this indicates that PD1 / PDL1 immune checkpoint inhibition may be an effective treatment for patient A014. Furthermore, the signal-to-noise ratio (SNR) for TUMERIC-solo (background / total 1 vs. cancer 6) was much higher than PDL1 upregulation measurements in naive bulk tumors (background / total 1.7 vs. bulk 3.9) (Figure 10).
[0075] Improved quantification of immune checkpoint biomarker signatures using TUMERIC-solo It has been previously reported that bulk tumor 6-gene biomarkers were responsive for response to pembrolizumab (PD1 / PDL1 inhibition) treatment. These 6 genes are IDO1 / CD274, CXCL10, CXCL9, HLA-DRA, STAT1, and IFNG. TUMERIC-solo was used to infer the activity of these genes in patient A014. This analysis demonstrated that one gene (CD274 / PDL1) was strongly upregulated in cancer cells, whereas the other 4 genes (CXCL10, HLA-DRA, IFNG, STAT1) were strongly upregulated in stroma (Figure 13). The signal-to-noise ratio for the combination of these 6 markers was compared for TUMERIC-solo and naive bulk tumor approaches. It was found that TUMERIC-solo provided a significant improvement in the signal-to-noise ratio for these markers by virtue of its ability to discriminate between cancer and stromal expression (Figure 14). Thus, TUMERIC-solo may provide a more accurate aggregate biomarker activity score for pembrolizumab treatment recommendations.
[0076] Using TUMERIC and TUMERIC-Solo to guide therapeutic and target discovery As can be seen from Figures 8 and 9 above, TUMERIC and TUMERIC-solo can be applied to a set of patient tumors or to one individual tumor to identify and / or recommend drug targets and treatments. A methodology according to this embodiment is outlined herein, which uses at least the following steps: 1. Applying TUMERIC / TUMERIC-solo to a set of samples / sectors; 2. Ranking genes or transcript isoforms by predicted cancer compartment expression; 3. Score genes or transcript isoforms by upregulation level in cancer compartment compared to stromal compartment (identifying cancer cell specific factors); 4. Score genes or transcript isoforms by upregulation level in cancer compared to healthy / normal tissue (identifying cancer cell specific factors); 5. Subsetting genes or transcript isoforms to known membrane-bound proteins or membrane-bound receptors (e.g., using known resources / databases). This results in a list of candidates for targets for antibody-based (e.g., antibody drug conjugate) therapy. 6. Subsetting (e.g., using known resources / databases) gene or transcript isoforms of proteins that produce known HLA binding and T cell antigen peptides. This results in a candidate list of tumor associated antigens (TAA) that are specifically associated with and overexpressed by cancer cells of the tumor, and recommends candidates for engineered T cell based therapies (e.g., but not limited to, CAR-T).
[0077] Thus, in one example, a method is disclosed for analyzing a single patient tumor.The method disclosed herein can also identify the transcripts that are aberrantly expressed in the cancer cells of a single patient.The method disclosed also requires only a minimal number of (mathematical) assumptions, allowing for unbiased analysis.
[0078] Patient-specific recommendations for therapeutic antibodies using TUMERIC-solo The extent to which the method disclosed herein (TUMERIC-solo) can be used to make recommendations regarding treatment with specific antibodies targeting membrane proteins of cancer cells was analyzed in subject A014. Approximately 4000 known annotated membrane proteins were analyzed for specificity of expression (log fold change >3, cancer compared to normal lung) and abundance (expression >50 FPKM) in cancer cells of A014 tumors, as these are important parameters for therapeutic antibody targeting. The top target from this approach using TUMERIC-solo was CLDN6, which is currently under evaluation as a therapeutic antibody target elsewhere (Figure 15). Thus, the results of TUMERIC-solo indicate that targeting CLDN6 is recommended in patient A014, for example, by using a therapeutic antibody against CLDN6. A similar naive bulk target recommendation approach only highlighted a single target (COL1A1) and failed to report the CLDN6 antibody target.
[0079] TUMERIC-solo identifies biomarkers of response to PD-L1 blockade in gastric cancer We further tested whether TUMERIC or TUMERIC-solo could also reveal previously untargeted biomarkers of PD-L1 inhibition treatment response by estimating gene expression more specifically in cancer or stromal / immune cells (compared to bulk tumor tissue). In this regard, TUMERIC was used to identify robust biomarkers across a cohort of treated patients, and then TUMERIC-Solo was applied as a biomarker testing assay (companion diagnostic) in the context of treating individual patients. Data from a recent cohort of approximately 50 metastatic gastric cancer patients treated with a PD-L1 inhibitor (pembrolizumab) were used. Patients were grouped based on their treatment response (complete / partial response (R); stable disease (SD); progressive disease (PD)) and TUMERIC was applied to each group of patients.
[0080] First, this analysis revealed a large set of genes with robust cancer or stromal cell gene expression dysregulation between responders (R) and non-responders (PD). The signal-to-noise ratio (predictive power) for these genes was much stronger with TUMERIC than measured by bulk tumor profiling (see Figure 16), suggesting that many of these biomarkers are only useful when combined with TUMERIC-solo. For example, biglycan (BGN) shows very high expression levels in cancer cells of non-responders (PD), but almost zero expression levels in responders (R+SD). Because this difference is much less clear and more variable in bulk tissue gene expression profiling (see Figure 17), tests based on bulk BGN expression alone may have insufficient prognostic power to be considered as a biomarker.
[0081] Data was obtained from a multi-patient gastric cancer cohort to test / simulate what TUMERIC-solo data for biglycan would likely look like in individual patients with putative metastatic gastric cancer with different pembrolizumab treatment outcomes (Figure 18). This shows that biglycan cancer / stroma expression levels can be inferred by TUMERIC-solo for a given patient to determine whether pembrolizumab would be effective in the metastatic gastric cancer setting.
[0082] The identification of biomarkers predictive of response to PD-L1 inhibition is illustrated as a further example showing a joint TUMERIC analysis of clinical trial cohorts and treatment-naïve microsatellite unstable (MSI) / microsatellite stable (MSS) tumors.
[0083] The discovery of robust predictive biomarkers of response to immune checkpoint inhibition (ICI) therapy is challenged by the paucity of available transcriptomic data from tumors responding and non-responding to ICI treatment. Because microsatellite unstable (MSI) tumors often have robust clinical responses to ICI therapy, we performed a joint TUMERIC analysis of immune checkpoint inhibition (ICI) clinical trial cohorts and a large cohort of treatment-naïve microsatellite unstable (MSI) / microsatellite stable (MSS) tumors. This joint analysis yielded five cancer and six stromal compartment gene expression biomarkers that were robustly associated with both ICI response and MSI status across three different tumor types.
[0084] Microsatellite instability is frequently found in colorectal, gastric, and endometrial cancers. A cohort of approximately 1000 treatment-naïve tumors was collected from these three tumor types in TCGA. Using TUMERIC, we identified cancer and stromal cell gene expression differences between microsatellite unstable (MSI) and microsatellite stable (MSS) tumors that were present in all three tumor types. Next, TUMERIC was used to analyze transcriptome data from a clinical trial of metastatic gastric cancer patients treated with a PD-L1 inhibitor (pembrolizumab; Nature Medicine. 2018, DOI: 10.1038 / s41591-018-0101-z; information disclosed in this study can also be found in the European Nucleotide Archive [ENA; part of the ELIXIR infrastructure at EMBL-EBI, Wellcome Genome Campus, Hinxton, Cambridgeshire, CB10 1SD, UK] under trial number PRJEB25780). Briefly, patients were grouped based on their treatment response (complete / partial response (R); stable disease (SD); progressive disease (PD)) and TUMERIC was applied within each group of patients. Significant cancer and stromal cell gene expression differences were then identified between the complete / partial response (R) and progressive disease (PD) groups. Finally, biomarkers from MSI / MSS and clinical trial data analysis were intersect, which resulted in a final list of six stromal cell-associated biomarkers (IFNG, FASLG, CXCL13, ZNF683, IL2RA, and CD274 / PD-L1) and five cancer cell-associated biomarkers (CPNE1, TTC19, OXCT1, ALDH6A1, and COX15). By applying TUMERIC-Solo, compartment-specific gene expression changes of these biomarkers can be measured in individual patient tumors, and the compartment-specific changes can then be used to predict response to ICI treatment. Data for the identified biomarker genes are summarized in Figures 24-34.
[0085] Treatment concepts within the scope of this disclosure include, but are not limited to, cancer cell-targeting antibodies (eg, ADCs), e.g., therapeutic antibodies against cell surface receptors, and chemotherapeutic agents.
[0086] In another example, the methods disclosed herein further comprise selecting genes or transcript isoforms for antibody-based therapy and / or T cell-based therapy.
[0087] Advantages of the methods disclosed herein include that these methods are applicable to both frozen and formalin-fixed paraffin-embedded (FFPE) tissue samples, meaning that immunohistochemical staining and the like can still be performed after analysis. Also, as illustrated in the data provided herein, the disclosed methods can distinguish between cancer and stromal (any non-cancer) cell types, providing more information than bulk / average profiling. Also, while the methods disclosed herein focus on transcriptome profiling, the same types of "omics" can be adapted to other types (e.g., without limitation, epigenomics, proteomics, etc.). As disclosed herein, the methods can also be performed with data derived from parallel DNA sequencing and sectored RNA data alone (e.g., by purity estimation based on RNA expression alone).
[0088] The method can also be applied in a complementary approach in the study of tumor microenvironment cell biology and antibody drug discovery in situations where bulk tumor biopsy data are already abundant or are the only viable source of data. Furthermore, insights gained from the method can be used to design in vitro assays and co-culture models that more accurately mimic the biological properties of the human tumor microenvironment.
[0089] Thus, it can be seen that the disclosed method has the potential to revolutionize the molecular data that can be extracted from individual bulk tumor samples. Using the methodology according to this embodiment is envisioned to create a near future where the cost of sequencing drops by >10-fold ($100 genome), meaning that the additional sequencing cost associated with the approach disclosed herein (about 5-fold higher) is negligible compared to the overall administrative and handling overhead associated with sequencing as a service for bulk tumor samples. The ability to directly and unbiasedly profile cancer cells from bulk tumor samples should be of immediate interest to companies selling clinical sequencing as a service, oncology precision computing in cancer hospitals, and large pharmaceutical companies interested in developing companion biomarkers. The methodology according to this embodiment can be used for any molecular activity (mRNA, epigenetic, protein expression) that can be co-extracted from individual sections and is ideally suited to the analysis of mRNA expression, because DNA and RNA can be easily co-extracted and analyzed by next generation sequencing.
[0090] Other embodiments are within the scope of the following claims and non-limiting examples. Additionally, when features or aspects of the invention are described in the form of a Markush group, those skilled in the art will recognize that the invention is also described thereby in terms of any individual members or subgroups of members of the Markush group. EXAMPLES
[0091] Experimental Section method Tumor data sources Twenty solid tumor types were analyzed. These solid tumor types have the Cancer Genome Atlas (TCGA) acronyms BLCA (bladder urothelial carcinoma), BRCA (invasive breast cancer), CESC (cervical squamous cell carcinoma), CRC (colorectal adenocarcinoma) (COAD (colon adenocarcinoma) and READ (rectal adenocarcinoma) combined), ESCA (esophageal carcinoma), GBM (glioblastoma multiforme), HNSC (head and neck squamous cell carcinoma), KIRC (renal clear cell carcinoma), KIRP (renal papillary cell carcinoma), LGG (brain low-grade glioma), LIHC (hepatocellular carcinoma), LUAD (lung adenocarcinoma), LUSC (lung squamous cell carcinoma), OV (ovarian serous cystadenocarcinoma), PAAD (pancreatic adenocarcinoma), PRAD (prostate adenocarcinoma), SKCM (cutaneous melanoma), STAD (gastric adenocarcinoma), THCA (thyroid carcinoma) and UCEC (uterine endometrial carcinoma). Somatic mutation (SNV) and copy number variation (CNV) data for 20 tumor types were obtained from the Broad Institute Firehose website (see Data Deposition section below). Uniformly processed Cancer Genome Atlas RNA sequencing (FPKM) data were obtained from the UCSC Xena server.
[0092] Tumor purity estimation Four different published methods for consensus tumor purity estimation were used. These are AbsCNseq, PurBayes, Ascat and ESTIMATE. AbsCNseq uses segmentation of copy number alterations and variant allele frequency (VAF) data of single nucleotide variants (SNVs) for individual tumors. PurBayes utilizes SNV VAF data (inferred from copy number alteration data) for diploid genes. Ascat purity estimation is based on copy number alteration (single nucleotide polymorphism (SNP) array) data, where tumor ploidy and purity are simultaneously estimated to identify allele-specific copy number alterations. Pre-calculated Ascat tumor purity estimates for the Cancer Genome Atlas cohort were obtained from the COSMIC website (see Data Deposit section below). ESTIMATE uses mRNA expression signatures of known immune and stromal gene signatures to infer tumor purity, and tumor purity values were obtained by applying ESTIMATE to Cancer Genome Atlas RNA sequencing (log2 FPKM [fragments per kilobase]) data. To obtain consensus tumor purity estimates, imputation of missing data was performed, followed by quantile normalization separately for each cancer type. Some tumor purity values were missing because these algorithms failed in certain input data instances. Additionally, several instances of very high (>98%) or low (<10%) purity estimates were observed, but such cases were typically found by only a single method for a given tumor, so these were also assigned as missing data. Missing data were then imputed using sequential principal component analysis of the matrix of incomplete algorithms and sample tumor purity (using the missMDA R package).
[0093] Quantile normalization was used to further standardize the tumor purity distributions of the different algorithms. Briefly, we sorted the tumor purity values for each algorithm and calculated the mean value for each rank in these distributions. We then assigned these mean values to the original positions of the individual purity distributions. As ESTIMATE produced purity estimates with large biases compared to the other three methods (typically 30-50% higher), we used only ESTIMATE purity values for the ranking step. We obtained the final TUMERIC consensus tumor purity estimate as the mean of these normalized purity values.
[0094] Cancer-stroma gene expression deconvolution Tumors were assumed to be composed of cancer and stromal (any non-cancerous) cells. The measured bulk tumor mRNA abundance was then determined by the sum of the mRNA molecules derived from these two compartments. The measured mRNA expression for a given gene in sample i was then calculated using Equation 3: It can be expressed as shown in TIFF2025026905000004.tif5128. In the formula, p i means the ratio of cancer cells (tumor purity), TIFF2025026905000005.tif5128 are the average expression levels for genes in the cancer and stromal compartments, respectively. See also FIG. 21, which shows the underlying mathematical model, disclosed herein. A simplifying assumption was made that these (non-negative) average compartment expression levels were constant across a set of tumors, and these expression levels were estimated using non-negative least squares regression (SciPy library). RNA sequencing fragments per kilobase (FPKM) data were log-transformed prior to deconvolution, log2(X+1). It has been discussed whether gene expression deconvolution should be performed using linear or log-transformed gene expression values. First, it was observed that the relationship between tumor purity and bulk tumor gene expression was often heteroscedastic. Second, both transformations were evaluated, and it was found that while the results were similar overall, the log transformation provided improved separation between inferred cancer and stromal compartment gene expression for known stromal genes (data not shown). It was also found that the above formula (Equation 1) tended to overestimate stromal gene expression for genes with somatic copy number alterations (CNAs) that affect gene expression in a subset of samples (e.g., ERBB2 in HER2-positive breast tumors). Therefore, a correction approach was used for such genes. Genes were identified using the correlation between copy number alterations (CNAs) and mRNA expression in a given set of tumors (Mann-Whitney U test P<1e-6 to account for multiple testing comparing expression for samples with diploid and non-diploid copy number alterations), and then a two-step approach was used to estimate gene expression in the cancer and stromal compartments. The above approach was used to first infer mRNA expression in the stromal compartment using only samples with diploid copy numbers for the gene of interest. The inferred mean stromal compartment expression, the measured mean tumor expression, and the mean purity of the tumor samples were then used to calculate the mean cancer compartment expression using the above formula.
[0095] Deconvolution of iTRAQ tumor protein expression data iTRAQ data was obtained for BRCA (breast cancer) and ovarian cancer (OV) tumor types using CPTAC Consortium data available on cBioPortal (www.cbioportal.org). Data was deconvoluted into cancer and stromal compartment expression similar to the RNA sequencing data above.
[0096] Ligand-receptor relative crosstalk (RC) score To estimate the relative flow of signaling between the cancer and stromal cell compartments, we developed a relative crosstalk (RC) score. The product of gene expression inferred for a given compartment is used to estimate the ligand-receptor (LR) complex activity (linear scale). The RC score, calculated with Equation 4, then estimates the relative complex activity, taking into account all four possible directions of signaling, for example in the case of cancer-cancer (CC) signaling, as well as the state of normal tissue. TIFF2025026905000006.tif16155
[0097] A normal term is included in the denominator to account for the combined activity in normal tissues, and this term is calculated directly from the observed gene expression levels in the matched normal tissue samples available for each tumor type in TCGA. It is noted that the relative crosstalk (RC) score is based on several simplifying assumptions, such as no competition or saturation effects for individual ligand-receptor complexes, mRNA expression is a reasonable surrogate for ligand and receptor concentrations at the site of formation of the ligand-receptor complex, cancer cells and stromal cells are homogeneously mixed in the tumor, and all cancer and stromal cells have the same properties and gene expression profiles.
[0098] Gene set enrichment (GSEA) analysis To study genes differentially expressed between cancer and stromal cells, gene set enrichment (GSEA) analysis was performed on a pre-ranked analysis of genes sorted by differential expression (log fragments per kilobase) in the cancer and stromal compartments. All signature gene signatures were analyzed to determine differentially enriched gene sets using a false discovery rate (FDR) cutoff of 0.25.
[0099] Immunohistochemistry (IHC) quantitative analysis To quantify cancer and stromal cell expression of genes, color deconvolution of IHC images obtained from the Human Protein Atlas (proteinatlas.org) was performed using the ImageJ software package and standard protocols. After manual selection and segmentation of cancer and stromal cells (without knowledge of antibody staining), color intensity was measured in ImageJ and DAB (target), hematoxylin (cell), and complementary color components were estimated. The average antibody intensity was then estimated for the cancer and stromal compartments of a given slide. In summary, IHC images of various human tumor samples stained with antibodies against S100A6 and LDHB were obtained from the Human Protein Atlas and analyzed with ImageJ software. Color deconvolution of DAB and hematoxylin was performed using the protocol described by Ruifrok et al. First, two high-quality images in which cancer and stromal cells were clearly visualized were randomly selected. Next, stromal and cancer cells in each IHC image were manually detected and segmented (using ROI manager) into stromal and cancer regions based on pathological features (cancer type, size, shape, arrangement of cells and cell nuclei) [3]. Pixel intensities were then calculated for the identified cancer and stromal regions based on DAB vectors (antibodies). The percentage of each cancer / stromal region with DAB staining was estimated, and the average cancer / stromal staining score was calculated for the whole slide by Equation 5 (shown below). log2((mean cancer staining rate + 1%) / (mean stromal staining rate + 1%) (5)
[0100] A 1% false count was added to the numerator and denominator to account for cases where cancer / stroma staining was zero.
[0101] (Table 1) TCGA cancer types and samples used. See also Figure 22. TIFF2025026905000007.tif128170
Claims
1. A method for predicting expression profiles of cancerous cells and non-cancerous cells, respectively, based on a plurality of sets of expression profiles, comprising: wherein each set of the plurality of sets of expression profiles is obtained from a tumor-derived sample comprising a mixture of cancerous and non-cancerous cells of a tumor type, the method comprising the steps of: a. determining a tumor purity value for one or more tumor-derived samples; b. providing distinct sets of expression profiles, wherein the distinct sets of expression profiles comprise combined expression data for a plurality or all of the molecules expressed by the cancerous and non-cancerous cells in the one or more tumor-derived samples, wherein the molecules are genes, the expression profiles are mRNA expression profiles and gene expression profiles, and wherein the expression profiles are log-transformed expression profiles; c. Deconvoluting each of the combined expression data referred to in b. by extrapolating the expression profiles of the plurality or all of the molecules expressed in different tumor samples having different tumor purity values to a tumor purity value of about 1 or about 0, wherein the expression profiles of the cancerous and non-cancerous cells are expressed according to the following formula: are obtained using non-negative least squares regression of where p i means the tumor purity value obtained in step a, and are the average expression levels for genes in the cancer and stromal compartments, respectively; This stage.
2. The method of claim 1, wherein the tumor-derived sample is obtained from a single subject.
3. The method described in claim 2, wherein the tumor-derived sample is divided into two or more sections and one set of expression profiles is generated for each section.
4. A method described in any one of claims 1 to 3, wherein the step of providing multiple sets of different expression profiles includes the use of an existing dataset of expression profiles.
5. The method described in claim 4, wherein the existing dataset of expression profiles is derived from the TCGA and ICGC databases.
6. The method of claim 1, wherein the tumor purity value is a mean tumor purity value.
7. The method described in claim 6, wherein the tumor purity value is obtained using a method selected from the group consisting of AbsCNseq, PurBayes, Ascat, ESTIMATE, and combinations thereof.
8. The method of claim 1, wherein the tumor type is selected from the group consisting of BLCA (bladder urothelial carcinoma), BRCA (invasive breast carcinoma), CESC (cervical squamous cell carcinoma), CRC (colorectal adenocarcinoma) (COAD (colon adenocarcinoma) and READ (rectal adenocarcinoma) combined), ESCA (esophageal carcinoma), GBM (glioblastoma multiforme), HNSC (head and neck squamous cell carcinoma), KIRC (renal clear cell carcinoma), KIRP (renal papillary cell carcinoma), LGG (brain low-grade glioma), LIHC (hepatocellular carcinoma), LUAD (lung adenocarcinoma), LUSC (lung squamous cell carcinoma), OV (ovarian serous cystadenocarcinoma), PAAD (pancreatic adenocarcinoma), PRAD (prostate adenocarcinoma), SKCM (cutaneous melanoma), STAD (gastric adenocarcinoma), THCA (thyroid carcinoma), and UCEC (uterine endometrial carcinoma).
9. The method of claim 1, further comprising the steps of: scoring the molecules of step c. based on their level of up-regulation or down-regulation in the cancer tissue compared to stromal tissue; and / or scoring the molecules of step c. based on their level of up-regulation or down-regulation in the cancer tissue compared to healthy tissue.
10. The method of claim 9, further comprising the steps of: assigning the up-regulated and down-regulated molecules to gene or transcript isoforms for a known dataset of membrane-bound proteins or membrane-bound receptors; and / or assigning the up-regulated and down-regulated molecules to gene or transcript isoforms for a known dataset of HLA-binding peptides and T-cell antigen-binding peptides.
11. The method of claim 10, wherein the known dataset for assigning genes or transcript isoforms is derived from Gene Ontology and / or TANTIGEN.
12. The method of claim 10 or 11, further comprising a step of selecting a gene or transcript isoform for antibody-based therapy and / or T cell-based therapy.
13. The method of any one of claims 10 to 12, wherein the gene or transcript isoform is a membrane-bound protein, a membrane-bound receptor, an antigenic peptide, a target protein, a peptide, and / or is targetable by an antibody.