Cellular ancestry signatures for subtyping cancers priority
By assaying biological samples for specific biomarkers and comparing expression levels to reference values, the method offers a stable breast cancer classification system that improves prognostic performance and treatment guidance.
Patent Information
- Application Number
- PCT/US2024/059083
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-12-08
- Publication Date
- 2025-06-12
AI Technical Summary
Current breast cancer classification systems are unstable and rely on predictive or prognostic clusters, which change with treatment and survival modality, lacking a firm anchor in anatomic, developmental, and biological categories.
A method involving the assay of biological samples for expression levels of specific biomarkers such as ET-60, THR-50, THR-70, CA-110, and CA-130, compared to reference values, to identify altered expression levels that provide a prognosis for longer overall survival and recurrence-free survival, and optionally guiding cancer treatment decisions.
The described method provides a robust foundation for breast cancer classification, optimizing the layering of prognostic and predictive biomarkers, and achieving significant prognostic performance across different clinical and molecular subtypes, surpassing existing prognostic signatures.
Smart Images

Figure IMGF000062_0001 
Figure IMGF000030_0001 
Figure IMGF000031_0001
Abstract
Description
[0001] CELLULAR ANCESTRY SIGNATURES FOR SUBTYPING CANCERS PRIORITY
[0002] This application claims the benefit of the filing date of U.S. application No. 63 / 607,694, filed December 8, 2023, the disclosure of which is incorporated by reference herein. BACKGROUND
[0003] The term 'classification' is often used to denote subgroups of breast cancer with different outcomes or treatment. However, utilizing predictive or prognostic clusters as a basis of classification has inherent instability, wherein group definition changes with both treatment and survival modality. In contrast, a genuine disease classification system should be firmly anchored in fundamental anatomic, developmental, and biological categories, independent of predictive or prognostic groups. A classification based on these stable groups can establish a robust foundation, optimizing the subsequent layering of prognostic and predictive biomarkers onto the classification categories, which further stratifies patients.
[0004] Following this paradigm, the hematopoietic malignancies have been successfully classified using normal cell-types as robust references. Since there is a well-established consensus about the normal hematopoietic cell types, this approach provides a stable foundation for lymphoma and leukemia classification. In this approach, both morphological features and markers that distinguish normal lymphoid and hematopoietic cell subtypes are used to classify their malignant counterparts. Next, prognostic / predictive molecular markers are used to refine each cell-type-based category further. As a result, the World Health Organization classification defines over 100 lymphoma and leukemia subtypes within cell-type categories including T-cell, B-cell, or natural killer-cell lymphomas, and lymphocytic, myelocytic, erythroid, or monocytic leukemias (1,2). This approach has resulted in a comprehensive cell-type-based lymphoma and leukemia classification that has withstood the test of time over the past four decades or more, reinforced by extensive molecular studies that largely upheld these cell-type-based categories.
[0005] SUMMARY
[0006] Provided herein is a method comprising: (a) assaying a biological sample from a subject for expression of (1) ET-60 biomarkers combined with THR-50 biomarkers recited in Table 1 as CA-110, or (2) ET-60 biomarkers combined with THR-70 biomarkers recited in Table 1 as CA-130, to determine expression levels for the ET-60 and THR-50 biomarkers or the ET-60 and THR-70; (b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’s biological sample, wherein altered expression levels of the CA-110 (ET-60 and THR-50) biomarkers or the CA-130 (ET- 60 and THR-70) biomarkers in the biological sample relative to the reference value provides a prognosis of longer overall survival and / or longer recurrence-free survival; and c) optionally administering or withholding one or more cancer treatments to a subject determined to have a cancer with a prediction of low overall survival and / or low progression-free survival. In some aspects, the sample is a breast cancer, colon cancer, lung cancer, prostate cancer, kidney cancer, cervix cancer, melanoma, ovarian cancer, liver cancer, pancreatic cancer, head & neck cancer, bladder cancer, brain tumors such as glioma, or hematopoietic neoplasms such as multiple myeloma sample. In some aspects, the CA-110 (ET-60 and THR-50) biomarkers are prognostic for breast cancer, colon cancer, lung cancer, prostate cancer, or pancreatic cancer. In some aspects, the CA-130 (ET-60 and THR-70) biomarkers are prognostic for breast cancer, kidney cancer, cervix cancer, melanoma, ovarian cancer, liver cancer, colon cancer, pancreas cancer, lung cancer, prostate cancer, head & neck cancer, bladder cancer, brain tumors such as glioma, or hematopoietic neoplasms such as multiple myeloma. In some aspects, RNA expression is assayed. In some aspects, nucleic acid amplification is employed prior to assaying. In some aspects, protein expression is assayed.
[0007] Provided herein is a method comprising: (a) assaying a biological sample from a subject for expression of THR-50 biomarkers recited in Table 1, or THR-70 biomarkers recited in Table 1 to determine expression levels for the THR-50 or THR-70 biomarkers; (b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’s biological sample, wherein a low mean expression of the THR-50 or THR-70 genes correlates to shorter overall survival (OS), or shorter recurrence-free survival (RFS), or metastasis-free survival (DMFS) of breast cancer; and (c) optionally administering or withholding one or more cancer treatments to a subject determined to have a breast cancer with a prediction of shorter OS, shorter RFS and shorter DMFS. In some aspects, expression of the THR-50 or THR-70 genes is correlated with survival in each of ER-neg, LN+, LN-, AR+, Grade 2, Grade 3, Lum-A, Lum-B, HER2+, HER2-, HER2-like and Basal-like breast cancer subtypes. In some aspects, expression of the THR-50 or THR-70 genes is combined with cellular ancestry information, triple hormone receptor protein expression status, gene mutation information, immune infiltrate information, or a combination thereof, to provide OS, RFS, or DMFS time. In some aspects, expression levels for THR-6E biomarkers are determined. In some aspects, expression of the THR-6E biomarkers is correlated with survival in ERp-HER2n-LNn breast cancers. In some aspects, expression of the THR-6E biomarkers is correlated with survival in each of LN-positive and LN-negative patients, untreated, hormone- treated, chemotherapy-treated, and ER-negative cancers. In some aspects, expression of the THR-6E biomarkers is correlated with response to endocrine therapy in ER-positive / LN- negative cancers and chemotherapy in ER-positive / LN-positive breast cancers. In some aspects, expression levels for THR-4E biomarkers are determined. In some aspects, expression of the THR-4E biomarkers is correlated with survival in ER-positive breast cancer. In some aspects, expression of the THR-4E biomarkers is correlated, with response to endocrine therapy in ER-positive / LN-negative breast cancers, and chemotherapy in ER-positive / LN-positive breast cancers. In some aspects, THR70 biomarkers are combined with i20 immune biomarkers recited in Table 1, and +HER2 signature to identify subgroups with shorter and longer survival. In some aspects, expression of the THR70 biomarkers combined with the i20 immune biomarkers is correlated with survival in each of Lum-A, ER+LP, Lum-B, ER+HP, Basal-like, TNBC, and HER-2-like tumors. In some aspects, RNA expression is assayed. In some aspects, nucleic acid amplification is employed prior to assaying. In some aspects, protein expression is assayed.
[0008] Provided herein is a method comprising: (a) assaying a biological sample from a subject for expression of i20 biomarkers recited in Table 1 to determine expression levels for the i20 biomarkers; (b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’s biological sample, wherein altered expression levels correlate to overall survival (OS) or relapse-free survival (RFS) of ERnegative breast cancer; and (c) optionally administering or withholding one or more cancer treatments to a subject determined to have a breast cancer with a prediction of shorter OS or RFS. In some aspects, expression levels of THR-i8 or THR-i3 biomarkers are determined. In some aspects, expression levels of the THR-i8 or the THR-i3 biomarkers are correlated with the outcome of double-negative (ER / HER2-negative), triple-negative (ER / PR / HER2- negative), quadruple-negative (ER / AR / HER2 / AR or ER / AR / HER2 / VDR-negative) and pentaplex-negative (ER / AR / HER2 / VDR / AR-negative) breast cancers. In some aspects, expression levels of the THR-i8 or THR-i3 biomarkers are correlated with survival of a gynecologic tumor, kidney cancer, or sarcoma. In some aspects, the gynecologic tumor is cervix, ovary, or uterus adenocarcinomas into survival outcome groups. In some aspects, expression levels of the THR-i8 biomarkers are correlated with overall survival (OS) or relapse-free (RFS) outcomes in the cervix, ovary, uterus, kidney carcinomas, and sarcomas. In some aspects, RNA expression is assayed. In some aspects, nucleic acid amplification is employed prior to assaying. In some aspects, protein expression is assayed.
[0009] Provided herein is a method comprising: (a) assaying a biological sample from a subject for expression of ET- 12H, ET- 13T biomarkers recited in Table 1 to determine expression levels for the ET-12H, ET-13T biomarkers; (b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’s biological sample, wherein altered expression levels correlate to overall survival (OS) or relapse-free survival (RFS) of ER-negative breast cancer; and (c) optionally administering or withholding one or more cancer treatments to a subject determined to have a breast cancer with a prediction of shorter OS or RFS. In some aspects, expression levels of the ET-12H biomarkers are correlated with survival of HER2-positive breast cancer, and the ET-13T biomarker is correlated with survival of ER / HER2-negative breast cancer. In some aspects, expression levels of the ET-13T biomarker is correlated with response to breast cancer chemotherapy and anti- HER2 therapy.
[0010] BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIGS. 1A-1C show the breast cancer classification based on triple-hormone receptor (THR) expression. A) Immunohistochemical (IHC) staining with ER, AR, and VDR of tissue microarrays (TMA) from breast cancer patients (top) identifies four distinct subtypes (bottom). Hormone receptor-positive tumors were identified as those with >1% protein expression. B) Kaplan-Meier survival curves showing the difference in survival between the four THR subtypes. C) Table showing the hazard ratio and 95% confidence interval (CI) for each of the THR in multivariate COX proportional hazards model. ER: estrogen receptor, AR: androgen receptor, VDR: vitamin D receptor. THR: triple hormone receptor. P: logrank p-value. Y-axis = percent survival, X-axis = time (months).
[0012] FIGS. 2A-2D show the development and validation of the THR-50 signature. A) Heatmaps showing the expression of the top 300 (left) and top 50 (right) differentially expressed genes between THR0&1 and THR2&3 cell lines in the Cancer Cell Line Encyclopedia (CCLE) dataset. B) Heatmap showing the expression of the THR-50 genes in human samples from the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC) dataset. The heatmap shows the sample annotation including ER status measured by immunohistochemistry, clinical subtypes (based on ER, HER2, and MIB-1 IHC) classification subtypes, histological grade, and PAM-50 groups. C-D) Kaplan-Meier survival plots showing the difference in recurrence-free survival (RFS) (C) and overall survival (OS) (D) between METABRIC samples predicted as 0 (better survival) and 1 (worse survival) using the THR-50 signature. Survival time is in months. The hazard ratios and 95% confidence intervals are shown. THR: triple-hormone receptors, HR: hazard ratio, CI: confidence interval. Y-axis = percent survival, X-axis = time (months). FIGS. 3A and 3B show the THR-50 signature is associated with recurrence-free survival across different breast cancer subtypes in the METABRIC dataset. A-B) Kaplan-Meier survival curves showing the difference in survival between samples with the highest probability of recurrence (Q4) and those with the lowest probability (QI) in different subtypes based on ER, HER2, and MIB-1 immunohistochemistry (IHC) (A), and in different PAM-50 groups (B). The hazard ratios (HR) and 95% confidence intervals (CI) are shown. Y-axis = percent survival, X-axis = time (months).
[0013] FIGS. 4A-4C show the THR-50 signature demonstrates significant association with recurrence-free survival (RFS) across different breast cancer clinical groups, outperforming established tests. A) Kaplan-Meier (KM) survival plots comparing the prognostic accuracy of the THR-50 signature and other existing tests including MammaPrint, Oncotype, and PAM-50 at predicting RFS. B-C) KM plots comparing the prognostic accuracy of the THR-50 signature (B) to the PAM-50 classification system (C) at capturing RFS in different breast cancer clinical groups. The analysis uses an independent validation cohort (KMP dataset) comprising 2,032 samples from 50 gene expression datasets. The reported p-values are derived from the log-rank test, assessing the statistical difference in survival distributions between the groups. The hazard ratios (HR) along with their corresponding 95% confidence intervals (CI) are shown. Y-axis = percent survival, X-axis = time (months).
[0014] FIGS. 5A-5C show the THR-50 signature outperforms established tests at capturing recurrence-free survival (RFS) across different PAM-50 groups. Kaplan-Meier survival plots comparing the association with RFS across different PAM-50 groups using the THR-50 signature (A), Oncotype (B), and MammaPrint (C) tests. The analysis uses an independent validation cohort (KMP dataset) comprising 2,032 samples from 50 gene expression datasets. The reported p-values are derived from the log-rank test, assessing the statistical difference in survival distributions between the groups. The hazard ratios (HR) along with their corresponding 95% confidence intervals (CI) are shown. Y-axis = percent survival, X-axis = time (months).
[0015] FIGS. 6A-6D show the development and validation of the THR-70 signature. A) Pie charts showing the percentages of PAM-50 groups in the different triple-hormone receptor (THR) subtypes in the BC-855 dataset. B) Venn diagram showing the genes in common between the top differentially expressed genes (DEGs) between the THR0&1 and THR2&3 groups in the CCLE and BC-855 datasets. The THR-70 signature comprises the top 70 DEGs in common between both datasets. C-D) Kaplan-Meier survival plots showing the association of the THR-70 signature top and bottom quartiles with overall survival (OS) across different breast cancer groups derived from PAM-50 (C) or immunohistochemical staining (D). Survival times are in months. Hazard ratios (HR) and 95% confidence intervals (CI) are shown. Statistically significant HRs are highlighted in red.
[0016] FIGS. 7A-7F show the THR-70 signature is prognostic in an independent patient cohort comprising samples from 10 different datasets. The THR-70 signature demonstrates a significant association with recurrence-free survival (RFS) (A) and distant metastasis -free survival (DMFS) (B). THR-70 is also prognostic in lymph node (LN) positive (C) and negative (D) disease, and in patients who received endocrine therapy following surgery (E), and in those who did not receive neoadjuvant treatment (F). Y-axis = percent survival, X-axis = time (years).
[0017] FIGS. 8A-8B show the triple hormone receptor signatures are enriched in cell-of-origin and immune gene set signatures in gene set variation analysis (GSVA). A) Heatmap showing the spearman correlation coefficients between different breast cancer signatures (rows) and key cancer-associated pathways and biological processes (columns). B) Heatmap of the enrichment scores of immune cell types (rows) in different breast cancer signatures (columns). * indicates a p-value <0.05 and # indicates a false discovery rate (FDR) <0.05.
[0018] FIG. 9 shows unsupervised clustering based on THR-70 uncovers five distinct breast cancer groups. Heatmap showing the expression of THR-70 genes in the METABRIC dataset. Five distinct groups were identified: El, E2a, E2b, E3, and PQNBC.
[0019] FIG. 10 shows the triple-hormone receptor signatures outperform existing classifiers in identifying subgroups with different survival outcomes. Kaplan-Meier survival plots comparing the 20-year recurrence-free survival rates between different breast cancer groups identified by immunohistochemical staining or ER, HER2, and MIB-1 (clinical), PAM-50, THR-50, and THR-70 signatures. THR: triple-hormone receptor. Pi+: pentaplex- (ER, PR, AR, VDR, and HER2) negative, immune-positive tumors. Pi-: pentaplex-negative, immunenegative tumors. Survival times are in months. Y-axis = percent survival, X-axis = time (months).
[0020] FIGS. 11A-11B show that THR-70 demonstrates a significant improvement in identifying subgroups with different survival rates compared to immunohistochemistry and PAM-50. A) Kaplan-Meier (KM) survival plots comparing the 20-year recurrence-free survival (RFS) in ER-negative groups identified by immunohistochemistry (IHC) (HER2+ and TNBC), PAM-50 (basal and claudin-low), and THR-70i (PQNBC_i- and PQNBC_i+). B) KM plots comparing the 20-year RFS in ER+ groups identified by IHC (ER+ high and low proliferation), PAM-50 (luminal A and B), and THR-70i (El, E2, and E3). Survival times are in months. The hazard ratios (HR) and 95% confidence intervals (CI) are shown. Y-axis - percent survival, X-axis = time (months).
[0021] FIGS. 12 -12C show the survival difference between the triple-hormone receptor- derived groups is not confounded by HER2 status. A-B) Kaplan-Meier (KM) survival plots comparing the 20-year recurrence-free survival (RFS) between HER2+ and HER2- samples within the PQNBC (A) and E clusters (B) derived from THR-70. C) KM survival plot comparing the 20-year RFS between HER2+ samples from the bottom THR-50 quartile (QI) and HER2+ samples from the top three quartiles (Q2:Q4). Hazard ratios (HR) and 95% confidence intervals (CI) are shown. Y-axis = percent survival, X-axis = time (months).
[0022] FIGS. 13A-13B show THR70+i20+HER2 signature outperforms PAM-50 (Mammaprint) at identifying subgroups with significantly shorter and longer survival. Kaplan- Meier survival plots comparing the 240 months of recurrence-free survival rates between different breast cancer groups identified by THR-70-i20-HER2 (A) and PAM-50 (B) signatures in the METABRIC dataset. THR: triple-hormone receptor. Survival times are in months. THR- 70+i20+HER2 signature identifies six prognostic groups with significant differences in survival hazard ratio: PQNBC i- (HR=1), El (HR=1.6), E2 (HR=2.5), E3 (HR=3.0), HER2 (HR=3.7), and PQNBC_i+ (HR=5.8). PAM-50 hazard ratios were not calculated because the survival curves cross each other.
[0023] FIG. 14 shows that the ET-60 and THR-50 signatures are additive in breast cancer. Kaplan-Meier survival plots comparing the recurrence-free survival rates of low-risk (black) and high-risk (red) groups. The ET-60 (HR=3.5) and THR-50 (HR=2.6) signatures demonstrate a significant association with recurrence-free survival (RFS) individually. However, their combination (CA-110) increases their prognostic power nearly two-fold (HR= 8.54, 95% CI 4.66-15.7). In contrast, the combination of commercially available standard prognostic signatures MammaPrint (MAM70, HR=3.4) and Prosignia (PAM-50, HR=2.3) is not additive; MAM70+PAM-50 (HR=3.7). Y-axis = percent survival, X-axis = time (years), HR=hazard ratio.
[0024] FIG. 15 shows that the CA- 110, combination of ET-60 and THR-50, is highly prognostic in multiple cancers. Kaplan-Meier survival plots comparing the recurrence-free survival rates of low-risk (black) and high-risk (red) groups in colon cancer (HR=16.4), lung cancer (HR=25.1), prostate cancer (HR=24.1) and pancreatic cancer (HR=17xlO9). Y-axis = percent survival, X-axis = time (years), HR=hazard ratio.
[0025] FIG. 16 shows that the ET-60 and THR-70 signatures are additive in breast cancer. Top row: Kaplan-Meier survival plots comparing the low-risk (green) and high-risk (red) groups (bottom row). The ET-60 (HR=3.5) and THR-70 (HR=4.3) signatures demonstrate a significant association with overall survival rates individually. Importantly, their combination (CA-130) increases their prognostic power nearly two-fold (HR= 7.3). Y-axis = percent survival, X-axis = time (years), HR=hazard ratio. Bottom row: The same analysis was carried out using a different software with similar results (bottom row); survival plots comparing the low-risk (green) and high-risk (red) groups. The ET-60 (HR=6.9) and THR-50 (HR=10.7) signatures demonstrate a significant association with overall survival rates individually. However, their combination (CA-130) increases their prognostic power nearly two-fold (HR=18.3). Y-axis = percent survival, X-axis = time (days), HR=hazard ratio.
[0026] FIG. 17 shows that the CA-130, the combination of ET-60 and THR-70 signatures, is a powerful prognostic biomarker for breast cancer. Kaplan-Meier survival plots comparing the low-risk (blue), intermediate-risk (green), and high-risk (red) groups in two datasets: TCGA (left), Y-axis = percent survival, X-axis = time (days); and Breast Cancer Metabase (right), Y axis = percent survival, X-axis = time (years), HR=hazard ratio.
[0027] FIG. 18 shows the heatmap of the CA-130 signature in the TCGA PanCancer dataset. The bar above indicates organ sites, in left to right order: Breast cancer (orange), Ovarian cancer (yellow), Endometrial cancer (green), Bladder cancer (blue), Prostate cancer (purple), Breast cancer (orange), Pancreatic cancer (dark green), Endometrial cancer (light green), Colon cancer (dark pink), Lung cancer (light green), Kidney (renal) cancer (light blue), Thyroid cancer (red), Melanoma (pink), Sarcoma (pink), Glioma (light brown), Liver cancer (dark blue), Lymphoma (dark brown). High expression (red), low expression (blue). cBioPortal PanCancer dataset; 2,583 patients, 2,922 samples.
[0028] FIG. 19 shows that the CA- 130, combination of ET-60 and THR-70, is highly prognostic in multiple cancers. Kaplan-Meier survival plots comparing the recurrence-free survival rates of low-risk (green) and high-risk (red) groups in kidney cancer (HR=6.8), cervix cancer (HR=2.68xlO9) (HR=25.1), melanoma (HR=6.2), ovarian cancer (HR=4.1), liver cancer (HR=9.0) and brain cancer glioma (HR=13.9), colon cancer (HR=19.9), pancreas cancer (HR=19.4), bladder cancer (HR=4.8), lung cancer (HR=2.2), head & neck cancer (HR=4.7), multiple myeloma (HR=7.2), Y axis = percent survival, X-axis = time (days), HR=hazard ratio.
[0029] FIGS. 20A-20B show that THR-50 is significantly associated with overall survival (OS) in the KMP breast cancer dataset. Kaplan-Meier (KM) survival plots showing the association between THR-50 and OS in different breast cancer subtypes and clinical groups using the best cut-off (a) and using quartiles (QI versus Q4) (b) to divide the expression into low and high. FIGS. 21A-21C show that THR-70 is associated with overall survival (OS) and recurrence-free survival (RFS) across different PAM-50 groups in the METABRIC dataset. A- B) Kaplan-Meier survival plots comparing OS (A) and RFS (B) between quartiles (Q1:Q4) of THR-70-derived probability scores across PAM-50 groups. C) The OS and RFS hazard ratios (HR) and 95% confidence intervals (CI) of quartiles Q2:Q4 compared to QI across the five PAM-50 groups. Significant HRs (p-value <0.05) are highlighted in red.
[0030] FIGS. 22A-22C show that THR-70 is associated with overall survival (OS) and recurrence-free survival (RFS) across different breast cancer groups based on immunohistochemical (IHC) staining in the METABRIC dataset. A-B) Kaplan-Meier survival plots comparing OS (A) and RFS (B) between quartiles (Q1:Q4) of THR-70-derived probability scores across the IHC groups. C) The OS and RFS hazard ratios (HR) and 95% confidence intervals (CI) of quartiles Q2:Q4 compared to QI across the four IHC groups. The groups are based on IHC staining with ER, HER2, and MIB-1. Significant HRs (p-value <0.05) are highlighted in red.
[0031] FIGS. 23A-23B show that THR-70 is significantly associated with recurrence-free survival (RFS) in an independent cohort comprising samples from 10 different datasets. Kaplan-Meier (KM) survival plots showing the association between THR-70 signature and RFS in the META-10 cohort overall (a) and in each dataset separately (b).
[0032] FIGS. 24A-24D show that THR-70 is significantly associated with overall survival in four different gene expression datasets.
[0033] FIGS. 25A-25B show unsupervised clustering based on THR-70 identifies distinct groups with different survival rates. A) Kaplan-Meier survival plots comparing 20-year recurrence-free survival between the five distinct clusters identified by THR-70. B) Same analysis after merging E2a and E2b in a single cluster (E2).
[0034] FIG. 26 shows the percentage of PAM-50 and THR-50 genetic mutations in samples from the TCGA and METABRIC datasets.
[0035] FIG. 27 shows the recurrence-free survival (RFS) rates in the THR-70-derived predominantly ER+ clusters in the METABRIC dataset. Kaplan-Meier survival plots comparing the 20-year RFS rates between the predominantly ER+ clusters overall (top left) and by HER2 status in each cluster separately. P: log-rank p-value.
[0036] FIG. 28 shows the overlap between THR-70 derived clusters and PAM-50 groups. The hierarchical plot at the top shows the METABRIC samples ordered by the THR-70-derived clusters together with other important annotations including histological grade, ER, PR, and HER2 status, immunohistochemistry-derived subtypes, and PAM-50 groups. The pie charts show the overlap between each THR-derived cluster and PAM-50 groups.
[0037] FIGS. 29A-29G. THR-6E outperforms random controls, Oncotype DX, and PAM-50 in predicting the survival of ER-positive breast cancer patients in the KMP dataset, (a) Heatmap of the expression of THR-70 in ER+ breast cancer in the METABRIC dataset. Four ER+ clusters were identified using hierarchical clustering of THR-70 expression in patient samples. THR-70 was divided into five gene groups with distinct expression across ER+ clusters. THROE refers to the fourth block of genes with the highest variability of expression across ER+ clusters. Analysis was carried out using the cBioPortal online platform https: / / www.cbioportal.org / . (b-c) Kaplan-Meier survival plots show the difference in relapse- free survival between patients with high and low average expression of THR-70 (b) and THROE (c) in the Kaplan-Meier Plotter (KMP) cohort. Categorization into low and high expression is based on using the median average expression as a threshold. P-values are derived from the log-rank test. Analysis was carried out using the KMP online platform https: / / kmplot.com / analysis / index.php?p=service&cancer=breast. (d) Kaplan-Meier survival plot showing the difference in relapse-free survival between ER+ / HER2- / LN breast cancer patients with high and low average expression of THR-OE. Categorization into high and low expression is based on the best cutoff of average THR-6E expression. The P-value is derived from the log-rank test. Analysis was carried out using the KMP online platform https : / / kmplot.com / analysis / index.php?p=service&cancer=breast. (e-g) Scatter plot comparing the distribution of relapse-free survival p-values between THR-6E (e), Oncotype DX (f), and PAM-50 (g) (red square) and random signatures with the same number of genes (circles). Signatures are shown on the x-axis while p-values are displayed on the y-axis in descending order. The p-values are derived from a log-rank test comparing the difference in relapse-free survival between patients with low and high average expression of each signature. Random signatures with a significant difference in survival (p<0.05) are shown in yellow while nonsignificant ones (p>=0.05) are shown in blue.
[0038] FIGS. 30A-30C. Expression of THR-6E genes is tissue-specific, (a) Expression of breast-specific hormone receptors (ER, AR, and VDR) and THR-6E genes Kinesin Family Member 4A and 2C (KIF4A, and KIF2C), Cell Division Cycle Protein 20 (CDC20), Microtubule Nucleation Factor (TPX2 / P100) in different normal human tissues. Analysis was carried out using the ProteomicsDB online platform https: / / www.proteomicsdb.org / . (b) The UMAP (Uniform Manifold Approximation and Projection) scatter plot of CDC20 expression in different cell types in normal human breast tissue Human Protein Atlas (HP A) single-cell mRNA dataset. The table summarizes the expression of all the THR-6E genes in different glandular breast epithelial (GBE) cell subsets. Analysis was carried out using the HPA online platform https: / / www.proteinatlas.org / . (c) The bar plot shows varying levels of CDC20 expression across different GBE subsets, as an example. All the other THR-6E genes were analyzed similarly.
[0039] FIG. 31. THR-6E genes are not mutated in breast cancer. Mutation rates of THR-6E, PAM-50, and MammaPrint signatures in breast cancer samples from the TCGA and METABRIC datasets. THR-6E genes show a low mutation rate (1.1% on average) across 3,593 patients, while PAM-50 and MammaPrint signatures contain multiple potential cancer driver genes amplified in 15-20% of samples, supporting the role of THR-6E as a stable cell-of-origin signature in breast cancer. Analysis was carried out using the cBioPortal online platform https: / / www.cbioportal.org / .
[0040] FIGS. 32A-32D. Expression of THR-6E genes (KIF4A, KIF2C, CDC20, TPX2, FAM64A, and LMNB2) is upregulated in primary and metastatic breast cancer compared to normal breast tissue, (a) The Density Plot shows RNA-Seq-based log2 expression of individual THR-6E genes in normal (pink), tumor (green), and metastasis (blue) breast tissue, (b) The Targetgram plots provide an overview of individual THR-6E genes in normal (green), tumor (orange) and metastasis (yellow) breast tissue, using gene chip-based data, the size of the segments represent the mean values, length of the dashed lines represent the median values of each type, (c) The Violin plots show the combined THR-6E signature expression in normal (green), tumor (orange), and metastasis (yellow) breast tissue, (d) The pan-cancer heatmap analysis displays log2 FC values of tumor / normal RNA-Seq data, the red color represents higher expression in tumors, while the blue color indicates higher expression in normal tissues. The sizes of the circles are inversely proportional to the adjusted P values. Analysis was carried out using the TNM online platform https: / / tnmplot.com / analysis / .
[0041] FIGS. 33A-33B. THR-6E genes form an interactome. (a) Protein-protein interaction network analysis of THR-6E and related hormone receptors. Each node represents a protein, with colors distinguishing between query proteins and the first shell of interactors (colored nodes). The edges between nodes represent protein-protein associations, including both known interactions from curated databases and experimental data, as well as predicted interactions based on gene neighborhood, gene fusions, gene co-occurrence, text mining, co-expression, and protein homology. The network highlights the functional connections between THR-6E components and key hormone receptors such as ER, AR, and VDR. Analysis was carried out using the String online platform https: / / string-db.org / . (b) Expanded network visualization of the regulatory interactions between THR-6E genes and their associated cancer hallmark phenotypes, the nodes represent proteins, with the query proteins shown in yellow and their first neighbors shown in green. White circles and blue clover leaves represent protein families and complexes, respectively. Direct interactions are displayed as solid lines and indirect ones as dashed lines, while edge color and arrow shape represent the effect (up- or down-regulation): up-regulations are represented as blue arrows, while down-regulations as ‘T-shaped’ red arrows. Analysis was carried out using the Cancer Gene Net online platform https: / / signor.uniroma2.it / CancerGeneNet / .
[0042] FIGS. 34A-34H. THR-6E outperforms random controls, Oncotype DX, PAM-50, and MammaPrint in predicting the survival of ER-positive breast cancer patients in the METABRIC dataset, (a-d) Kaplan-Meier survival curves for relapse-free survival (RFS) stratified by THR-6E (a), Oncotype DX (b), PAM-50 (c), and MammaPrint (d) signatures. Each panel shows the survival probability over time for low-risk (blue) and high-risk (yellow) patient groups, with the hazard ratio (HR) and corresponding 95% confidence interval (CI) and ward-test p- value displayed. The number of patients at risk at various time points is indicated below each curve. For each patient, a risk score was computed using each signature’s genes, and stratification into low- and high-risk was performed based on the median risk score, (e-h) Comparison of THR-6E (e), Oncotype DX (f), PAM-50 (g), and MammaPrint (h) signatures with random signatures of the same size in predicting RFS. Each scatter plot shows the distribution of p-values for the signatures, ranked by significance. THR-6E, Oncotype DX, PAM-50, and MammaPrint are shown in green, while randomly generated signatures are shown in red. The dashed line represents the significance threshold (p<0.05). The arrows indicate the position of the tested signatures relative to the random distribution, demonstrating the statistical significance of the THR-6 signature compared to the others in predicting RFS. P-values are derived from the Ward test comparing the difference in survival between low- and high-risk patients defined by each signature.
[0043] FIGS. 35A-35B. THR-6E outperforms Oncotype DX, PAM-50, and MammaPrint in predicting response to endocrine treatment and chemotherapy in ER-positive breast cancer patients, (a) The performance of THR-6E, Oncotype DX, PAM-50, and MammaPrint at predicting the response to endocrine therapy (ET) in patients with ER+ / HER2- and lymph node-negative (LN-) breast cancer. The boxplots (top) show the difference in gene expression levels in responders versus non-responders for each of the analyzed signatures. The Receiver Operating Characteristic (ROC) curves (bottom) show the predictive power, measured by the Area Under the Curve (AUC), of each signature for distinguishing between responders and non-responders. (b) The performance of THR-6E, Oncotype DX, PAM-50, and MammaPrint at predicting the response to chemotherapy in patients with ER+ / HER2- and lymph nodepositive (LN+) breast cancer. The boxplots (top) show the difference in gene expression levels in responders versus non-responders for each of the analyzed signatures. The ROC curves (bottom) show the predictive power, measured by the AUC, of each signature for distinguishing between responders and non-responders. Analysis was carried out using the ROC online platform https : / / www.rocplot.com / site / treatment.
[0044] FIGS. 36A-36C. THR-6E is specifically prognostic for different clinical groups of breast cancer, (a) Kaplan-Meier survival curves for lymph node (LN)-negative, LN-positive, and estrogen receptor (ER)-negative breast cancer patients, stratified by high (red) and low (black) gene average THR-6E expression levels. Each panel includes the hazard ratio (HR) with 95% confidence intervals (CI) and the log -rank p-value. The number of patients at risk at various time points is shown below each curve, (b) KM survival curves for untreated breast cancer patients and those receiving hormone therapy or chemotherapy, (c) KM survival curves for other cancer types, including ovarian cancer, small cell carcinoma of the lung (Lung SCCC), and leukemia, stratified by THR-6E expression levels. The analysis shows no significant difference in survival between high and low expression groups for Lung SCCC, ovarian cancer, and leukaemia which highlights the specificity of THR-6E to breast cancer. Analysis was carried out using the KMP online platform https : / / kmplot.com / analysis / index.php?p=service&cancer=breast.
[0045] FIGS. 37 -37H. THR-4E is a prognostic and predictive biomarker for ER-positive breast cancer, (a-b) Kaplan-Meier (KM) survival curves for THR-4E in estrogen receptor (ER)- positive (a); estrogen receptor (ER)-negative (b) breast cancer patients, stratified by high (red) and low (black) gene average THR-4E expression levels. Each panel includes the hazard ratio (HR) with 95% confidence intervals (CI) and the log-rank p-value. The number of patients at risk at various time points is shown below each curve. Categorization into low and high expression is based on using the median average expression as a threshold. P-values are derived from the log-rank test. Analysis was carried out using the KMP online platform https: / / kmplot.com / analysis / index.php?p=service&cancer=breast. (c) KM survival plots show the difference in relapse-free survival (RFS) between patients with high and low risk based on ET-4E in the METABRIC cohort. Categorization into high and low expression is based on the median risk. P-values are derived from the log-rank test, (d) Scatter plots comparing the distribution of RFS p-values between ET-4E and 1000 random signatures of the same size (red circle). Signatures are shown on the x-axis while p-values are displayed on the y-axis in descending order. The horizontal dashed line corresponds to a p-value of 0.05. P- values are derived from the Ward test, (e-h) The performance of THR-4E at predicting the response to endocrine therapy (ET) in patients with ER-positive breast cancer (e-f); and response to chemotherapy in patients with ER-positive / HER-2-negative and lymph-node positive breast cancer (g-h). The boxplots (top) show the difference in gene expression levels in responders versus non-responders for each of the analyzed signatures. The Receiver Operating Characteristic (ROC) curves (bottom) show the predictive power, measured by the Area Under the Curve (AUC), of each signature for distinguishing between responders and non-responders. Analysis was carried out using the ROC online platform https: / / www.rocplot.com / site / treatment.
[0046] FIGS. 38A-38E. THR-4E outperforms 1,000 random controls and four of the most commonly used commercialized signatures in ER-positive and HER2-negative breast cancer. Scatter plots comparing the distribution of RFS p-values between test signature (green dot) and 1,000 random signatures of the same size (red line), analyzed and ranked by machine learning. Signatures are shown on the x-axis while p-values are displayed on the y-axis in descending order. The horizontal dashed line corresponds to a p-value of 0.05. P-values are derived from the Ward test, (a) THR-4E, (b) Endopredict, (c) Oncotype DX, (d) PAM-50, and (e) MammaPrint.
[0047] FIGS. 39A-39F. Comparison of THR-i20, -i8, and -i3 in KMP and TCGA datasets, (a- c) Kaplan-Meier (KM) survival curves for THR-i20, -i8, and -i3E in estrogen receptor (ER) negative and HER2-negative breast cancer patients in the KMP dataset. THR-i high (red) and low (black). Analysis was carried out using the KMP online platform https: / / kmplot.com / analysis / index.php?p=service&cancer=breast. (d-e) Kaplan-Meier (KM) survival curves for THR-i20, -i8, and -i3E in estrogen receptor, HER2, and progestin-negative (TNBC) breast cancer patients in the TCGA dataset. THR-i high (black) and low (red). Analysis was carried out using the Tumor online Prognostic analyses Platform (ToPP) http: / / www.biostatistics.online / topp / survival.php#mutpng.Each panel includes the hazard ratio (HR) with 95% confidence intervals (CI) and the log-rank p-value. The number of patients at risk at various time points is shown below each curve.
[0048] FIGS. 40A-40H. Comparison of THR-i20, -i8, and -i3 in TNBC and PQNBC in METABRIC, (a-d) Kaplan-Meier (KM) survival curves for THR-i20, -i8, and -i3E in estrogen receptor, HER2, and progestin-negative (TNBC) breast cancer patients in the METABRIC dataset. THR-i high (yellow) and low (blue), (e-h) Kaplan-Meier (KM) survival curves for THR-i20, -i8, and -i3E in estrogen receptor, HER2, progestin, androgen receptor, and / or vitamin-D negative (PQNBC) breast cancer patients in the METABRIC dataset. THR-i high (yellow) and low (blue).
[0049] FIG. 41. Outcome stratification of gynecological cancers by THR-i3. Kaplan-Meier (KM) survival curves for THR-i3 stratify gynecologic tumors such as cervix, ovary, and uterus adenocarcinomas into overall survival (OS) and relapse-free survival (RFS) outcome groups in the Pan-cancer RNA-seq dataset. Analysis was carried out using the KMP plot online platform https : / / kmplot.com / analysis / index.php?p=service&cancer=pancancer_rnaseq#.
[0050] FIGS. 42A-42H. Outcome stratification of gynecological cancers, kidney cancers, and sarcomas that are enriched in immune cells by THR-i8. Kaplan-Meier (KM) survival curves for THR-i8 stratify overall survival (OS) outcome groups in multiple tumor types including cervix, ovary, uterus, and kidney clear cell carcinomas (CCC) that are enriched for natural killer cells (NK), kidney papillary carcinoma enriched for CD8+ T-cells and sarcomas enriched for type 1 T-helper cells in the Pan-cancer RNA-seq dataset. Analysis was carried out using the KMP platform https : / / kmplot.com / analysis / index.php?p=service&cancer=pancancer_rnaseq#.
[0051] FIG. 43. Outcome stratification of gynecological cancers, kidney cancers, and sarcomas cancers regardless of immune cell enrichment by THR-i8.
[0052] FIGS. 44A-44I. THR-12H is a prognostic biomarker for HER2 -positive breast cancer, (a) Kaplan-Meier (KM) survival curves in the KMP dataset show the difference in relapse-free survival (RFS) between HER-2-positive breast cancer with high (red) and low (black) THR- 12H average gene expression levels. Each panel includes the hazard ratio (HR) with 95% confidence intervals (CI) and the log-rank p-value. The number of patients at risk at various time points is shown below each curve. Categorization into low and high expression is based on using the median average expression as a threshold. P-values are derived from the log-rank test. Analysis was carried out using the KMP online platform https: / / kmplot.com / analysis / index.php?p=service&cancer=breast. (b) Kaplan-Meier (KM) survival curves in the METABRIC dataset show the difference in relapse-free survival (RFS) between HER-2-positive breast cancer with high (yellow) and low (blue) THR-12H average gene expression levels, (c) Kaplan-Meier (KM) survival curves in the TCGA dataset show the difference in disease-specific survival (DSS) between patients with HER-2-positive breast cancer with high (black and low (red) THR-12H average gene expression levels, (d) Table showing the rank of each signature compared to 1,000 random signatures of equal size, (e-i) Scatter plots comparing the distribution of RFS p-values between test signature (blue dot) and 1,000 random signatures of the same size (red line), analyzed and ranked by machine learning. Signatures are shown on the x-axis while p-values are displayed on the y-axis in descending order. The horizontal dashed line corresponds to a p-value of 0.05. P-values are derived from the Ward test; (e) THR-12H, (f) Endopredict, (g) Oncotype DX, (h) PAM-50, and (i) MammaPrint.
[0053] FIGS. 45A-45L THR-13T is a prognostic biomarker for ER / HER2-negative breast cancer, (a) Kaplan-Meier (KM) survival curves in the KMP dataset show the difference in relapse-free survival (RFS) between ER / HER-2-negative breast cancer with high (red) and low (black) THR-13T average gene expression levels. Each panel includes the hazard ratio (HR) with 95% confidence intervals (CI) and the log -rank p-value. The number of patients at risk at various time points is shown below each curve. Categorization into low and high expression is based on using the median average expression as a threshold. P-values are derived from the log-rank test. Analysis was carried out using the KMP online platform https: / / kmplot.com / analysis / index.php?p=service&cancer=breast. (b) Kaplan-Meier (KM) survival curves in the METABRIC dataset show the difference in relapse-free survival (RFS) between ER / HER-2-negative breast cancer with high (yellow) and low (blue) THR-13T average gene expression levels, (c) Kaplan-Meier (KM) survival curves in the TCGA dataset show the difference in disease-specific survival (DSS) between patients with HER-2-positive breast cancer with high (black and low (red) THR-13T average gene expression levels, (c-d) The performance of ET-13T at predicting the response to chemotherapy in patients with ER / HER2-negative breast cancer (c); and response to anti-HER2-therapy in patients with HER2 -positive breast cancer (d). The boxplots (top) show the difference in gene expression levels in responders versus non-responders for each of the analyzed signatures. The Receiver Operating Characteristic (ROC) curves (bottom) show the predictive power, measured by the Area Under the Curve (AUC), of each signature for distinguishing between responders and non-responders. Analysis was carried out using the ROC online platform https: / / www.rocplot.com / site / treatment, (e-i) Scatter plots comparing the distribution of RFS p-values between test signature (blue dot) and 1 ,000 random signatures of the same size (red line), analyzed and ranked by machine learning. Signatures are shown on the x-axis while p- values are displayed on the y-axis in descending order. The horizontal dashed line corresponds to a p-value of 0.05. P-values are derived from the Ward test; (e) THR-13T, (f) Endopredict, (g) Oncotype DX, (h) PAM-50, and (i) MammaPrint.
[0054] DESCRIPTION
[0055] As illustrated herein, ET-60, THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, and THR-i3 are markers useful for detecting, diagnosing, and determining the prognosis of cancer, including breast cancer. Methods for detecting, diagnosing, and determining the prognosis of cancer, including breast cancer, are also described herein.
[0056] The methods generally involve obtaining a sample from a subject and comparing gene expression levels in the sample with one or more reference values, where the expression levels of the following genes are compared: THR-50 genes, THR-70 genes, CA-110 genes, CA-130 genes, THR-6E genes, THR-4E genes, i20 genes, THR-i8 genes, THR-i3 genes, ET-12H, ET- 13T genes or a combination of those genes. The method can also include classifying the subject from whom the sample was obtained as having cancer (i.e., being a cancer patient) or not having cancer. The method can also include classifying a cancer patient as having a poor prognosis based upon the expression levels of the THR-50 genes, THR-70 genes, CA-110 genes, CA-130 genes, THR-6E genes, THR-4E genes, i20 genes, THR-i8 genes, THR-i3 genes, ET-12H, ET-13T genes, or a combination of those genes in the patient’s sample. In some cases, the subject is a breast cancer patient.
[0057] For example, a method for classifying a breast cancer patient according to prognosis can include: (a) comparing the respective levels of expression of THR-50 genes, THR-70 genes, CA-110 genes, CA-130 genes, THR-6E genes, THR-4E genes, i20 genes, THR-i8 genes, THR-i3 genes, or a combination of the genes in a sample taken from a breast cancer patient to respective reference values of expression of the genes; and (b) classifying the breast cancer patient according to prognosis of his or her breast cancer based on altered expression levels of the THR-50 genes, THR-70 genes, CA-110 genes, CA-130 genes, THR-6E genes, THR-4E genes, i20 genes, THR-i8 genes, THR-i3 genes, ET-12H, ET-13T genes, or a combination thereof.
[0058] Samples
[0059] Breast cancer can be assessed through the evaluation of expression patterns, or profiles, of the THR-50 genes, THR-70 genes, CA-110 genes, CA-130 genes, THR-6E genes, THR-4E genes, 120 genes, THR-i8 genes, ET-12H, ET-13T genes, or THR-i3 genes in one or more subject samples. The term subject, or subject sample, refers to an individual regardless of health and / or disease status. A subject can be a subject, a study participant, a control subject, a screening subject, or any other class of individual from whom a sample is obtained and assessed using the markers and / or methods described herein. Accordingly, a subject can be diagnosed with breast cancer, can present with one or more symptoms of breast cancer, or a predisposing factor, such as a family (genetic) or medical history (medical) factor, for breast cancer, can be undergoing treatment or therapy for breast cancer, or the like. Alternatively, a subject can be
[0060] Y1 healthy with respect to any of the aforementioned factors or criteria. It will be appreciated that the term “healthy” as used herein, is relative to breast cancer status, as the term “healthy” cannot be defined to correspond to any absolute evaluation or status. Thus, an individual defined as healthy with reference to any specified disease or disease criterion, can in fact be diagnosed with any other one or more diseases, or exhibit any other one or more disease criterion, including one or more cancers other than breast cancer. However, the healthy controls are preferably free of any cancer.
[0061] In some cases, the methods for detecting, predicting, and / or assessing the prognosis of breast cancer include collecting a biological sample comprising a cell or tissue, such as a breast tissue sample or a primary breast tumor tissue sample. By “biological sample” is intended for any sampling of cells, tissues, or bodily fluids in which expression of THR-50 genes, THR-70 genes, CA-110 genes, CA-130 genes, THR-6E genes, THR-4E genes, i20 genes, THR-i8 genes, ET-12H, ET-13T genes, or THR-i3 genes can be detected. Examples of such biological samples include but are not limited to, biopsies and smears. Bodily fluids useful in the present invention include blood, lymph, urine, saliva, nipple aspirates, gynecological fluids, or any other bodily secretion or derivative thereof. Blood can include whole blood, plasma, serum, or any derivative of blood. In some embodiments, the biological sample includes breast cells, particularly breast tissue from a biopsy, such as a breast tumor tissue sample. Biological samples may be obtained from a subject by a variety of techniques including, for example, by scraping or swabbing an area, by using a needle to aspirate cells or bodily fluids, or by removing a tissue sample (i.e., biopsy). In some embodiments, a breast tissue sample is obtained by, for example, fine needle aspiration biopsy, core needle biopsy, or excisional biopsy.
[0062] The samples can be stabilized for evaluating and / or quantifying THR-50, THR-70, CA- 110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 expression levels.
[0063] In some cases, fixative and staining solutions may be applied to some of the cells or tissues for preserving the specimen and for facilitating examination. Biological samples, particularly breast tissue samples, may be transferred to a glass slide for viewing under magnification. In one embodiment, the biological sample is a formalin-fixed, paraffin- embedded breast tissue sample, particularly a primary breast tumor sample.
[0064] Gene Expression
[0065] Various methods can be used for evaluating and / or quantifying THR-50, THR-70, CA- 110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 expression levels. By “evaluating and / or quantifying” is intended determining the quantity or presence of an RNA transcript or its expression product of THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 genes.
[0066] Methods for detecting expression of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 genes, including gene expression profiling, can involve methods based on hybridization analysis of polynucleotides, methods based on sequencing of polynucleotides, immunohistochemistry methods, and proteomicsbased methods. The methods generally involve detecting expression products (e.g., mRNA or proteins) encoded by the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR- i8, ET-12H, ET-13T or THR-i3 genes. In some cases, PCR-based methods, which can include reverse transcription PCR (RT-PCR) (Weis et al., TIG 8:263-64, 1992), array-based methods such as microarray (Schena et al., Science 270:467-70, 1995), or combinations thereof are used. By “microarray” is intended an ordered arrangement of hybridizable array elements, such as, for example, polynucleotide probes, on a substrate. The term “probe” refers to any molecule that is capable of selectively binding to a specifically intended target biomolecule, for example, a nucleotide transcript or a protein encoded by or corresponding to THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 genes. Probes can be synthesized or obtained from THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8 or THR-i3 nucleic acids or they can be derived from appropriate biological preparations. Probes may be specifically designed to be labeled. Examples of molecules that can be utilized as probes include, but are not limited to, RNA, DNA, proteins, antibodies, and organic molecules.
[0067] Many expression detection methods use isolated RNA. The starting material is typically total RNA isolated from a biological sample, such as a cell or tissue sample, a tumor or tumor cell line, a corresponding normal tissue or cell line, or a combination thereof. If the source of RNA is a sample from a subject, RNA (e.g., mRNA) can be extracted, for example, from stabilized, frozen, or archived paraffin-embedded, or fixed (e.g., formalin-fixed) tissue samples (e.g., pathologist-guided tissue core samples).
[0068] General methods for RNA extraction are available and are disclosed in standard textbooks of molecular biology, including Ausubel et al., ed., Current Protocols in Molecular Biology, lohn Wiley & Sons, New York 1987-1999. Methods for RNA extraction from paraffin-embedded tissues are disclosed, for example, in Rupp and Locker (Lab Invest. 56: A67, 1987) and De Andres et al. (Biotechniques 18:42-44, 1995). In some cases, RNA isolation can be performed using a purification kit, a buffer set, and protease from commercial manufacturers, such as Qiagen (Valencia, Calif.), according to the manufacturer's instructions. For example, total RNA from cells can be isolated using Qiagen RNeasy mini columns. Other commercially available RNA isolation kits include MASTERPURE™ Complete DNA and RNA Purification Kit (Epicentre, Madison, Wis.) and the Paraffin Block RNA Isolation Kit (Ambion, Austin, Tex.). Total RNA from tissue samples can be isolated, for example, using RNA Stat-60 (Tel-Test, Friendswood, Tex.). RNA prepared from tissue or cell samples (e.g. tumors) can be isolated, for example, by cesium chloride density gradient centrifugation. Additionally, large numbers of tissue samples can readily be processed using available techniques, such as, for example, the single-step RNA isolation process of Chomczynski (U.S. Pat. No. 4,843,155).
[0069] Isolated RNA can be used in hybridization or amplification assays that include, but are not limited to, PCR analyses and probe arrays. One method for the detection of RNA levels involves contacting the isolated RNA with a nucleic acid molecule (probe) that can hybridize to the mRNA encoded by the gene being detected. The nucleic acid probe can be, for example, a full-length cDNA, or a portion thereof, such as an oligonucleotide of at least 7, 15, 30, 60, 100, 250, or 500 nucleotides in length and sufficient to specifically hybridize under stringent conditions to any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, THR-i3 genes, or any derivative DNA or RNA. Hybridization of an mRNA with the probe indicates that the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET- 12H, ET-13T, or THR-i3 genes in question are being expressed.
[0070] In some cases, the mRNA from the sample is immobilized on a solid surface and contacted with a probe, for example by running the isolated mRNA on an agarose gel and transferring the mRNA from the gel to a membrane, such as nitrocellulose. In other cases, the probes are immobilized on a solid surface and the mRNA is contacted with the probes, for example, in an Agilent gene chip array. A skilled artisan can readily adapt available mRNA detection methods for use in detecting the level of expression of the THR-50, THR-70, CA- 110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T or THR-i3 genes.
[0071] An alternative method for determining the level of THR-50, THR-70, CA-110, CALIO, THR-6E, THR-4E, i20, THR-18, ET-12H, ET-13T, or THR-i3 gene expression in a sample involves the process of nucleic acid amplification of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 mRNA (or cDNA thereof), for example, by RT-PCR (U.S. Pat. No. 4,683,202), ligase chain reaction (Barany, Proc. Natl. Acad. Sci. USA 88: 189-93, 1991), self-sustained sequence replication (Guatelli et al., Proc. Natl. Acad. Sci. USA 87: 1874-78, 1990), transcriptional amplification system (Kwoh et al., Proc. Natl. Acad. Sci. USA 86:1173-77, 1989), Q-Beta Replicase (Lizardi et al., Bio / Technology 6: 1197, 1988), rolling circle replication (U.S. Pat. No. 5,854,033), or any other nucleic acid amplification method, followed by the detection of the amplified molecules using available techniques. These detection schemes are especially useful for the detection of nucleic acid molecules if such molecules are present in very low numbers.
[0072] In some cases, THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 gene expression is assessed by quantitative RT-PCR. Numerous different PCR or QPCR protocols are available and can be directly applied or adapted for use using the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET- 13T, or THR-i3 genes. Generally, in PCR, a target polynucleotide sequence is amplified by reaction with at least one oligonucleotide primer or pair of oligonucleotide primers. The primer(s) hybridize to a complementary region of the target nucleic acid and a DNA polymerase extends the primer(s) to amplify the target sequence. Under conditions sufficient to provide polymerase-based nucleic acid amplification products, a nucleic acid fragment of one size dominates the reaction products (the target polynucleotide sequence which is the amplification product). The amplification cycle is repeated to increase the concentration of the single target polynucleotide sequence. The reaction can be performed in any thermocycler commonly used for PCR. However, preferred are cyclers with real-time fluorescence measurement capabilities, for example, SMARTCYCLER® (Cepheid, Sunnyvale, Calif.), ABI PRISM 7700® (Applied Biosystems, Foster City, Calif.), ROTOR-GENE™ (Corbett Research, Sydney, Australia), LIGHTCYCLER® (Roche Diagnostics Corp, Indianapolis, Ind.), ICYCLER® (Biorad Laboratories, Hercules, Calif.) and MX4000® (Stratagene, La Jolla, Calif.).
[0073] Quantitative PCR (QPCR) (also referred to as real-time PCR) is preferred under some circumstances because it provides not only a quantitative measurement but also reduces time and contamination. In some instances, the availability of full gene expression profiling techniques is limited due to requirements for fresh frozen tissue and specialized laboratory equipment, making the routine use of such technologies difficult in a clinical setting. However, QPCR gene measurement can be applied to standard formalin-fixed paraffin-embedded clinical tumor blocks, such as those used in archival tissue banks and routine surgical pathology specimens (Cronin et al. (2007) Clin Chem 53:1084-91) (Mullins 2007) (Paik 2004). As used herein, “quantitative PCR (or “real-time QPCR”) refers to the direct monitoring of the progress of PCR amplification as it is occurring without the need for repeated sampling of the reaction products. In quantitative PCR, the reaction products may be monitored via a signaling mechanism (e.g., fluorescence) as they are generated and are tracked after the signal rises above a background level but before the reaction reaches a plateau. The number of cycles required to achieve a detectable or “threshold’- level of fluorescence varies directly with the concentration of amplifiable targets at the beginning of the PCR process, enabling a measure of signal intensity to provide a measure of the amount of target nucleic acid in a sample in real-time.
[0074] In some cases, microarrays are used for expression profiling. Microarrays are particularly well suited for this purpose because of the reproducibility between different experiments. DNA microarrays provide one method for the simultaneous measurement of the expression levels of large numbers of genes. Each array consists of a reproducible pattern of capture probes attached to a solid support. Labeled RNA or DNA is hybridized to complementary probes on the array and then detected by laser scanning. Hybridization intensities for each probe on the array are determined and converted to a quantitative value representing relative gene expression levels. See, for example, U.S. Pat. Nos. 6,040,138, 5,800,992 and 6,020,135, 6,033,860, and 6,344,316. High-density oligonucleotide arrays are particularly useful for determining the gene expression profile for a large number of RNAs in a sample. Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, for example, U.S. Pat. No. 5,384,261. Although a planar array surface can be used, the array can be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays can be nucleic acids (or peptides) on beads, gels, polymeric surfaces, fibers (such as fiber optics), glass, or any other appropriate substrate. See, for example, U.S. Pat. Nos. 5,770,358, 5,789,162, 5,708,153, 6,040,193 and 5,800,992. Arrays can be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device. See, for example, U.S. Pat. Nos. 5,856,174 and 5,922,591.
[0075] When using microarray techniques, PCR-amplified inserts of cDNA clones can be applied to a substrate in a dense array. The microarrayed genes immobilized on the microchip, are suitable for hybridization under stringent conditions. Fluorescently labeled cDNA probes can be generated through the incorporation of fluorescent nucleotides by reverse transcription of RNA extracted from tissues of interest. Labeled cDNA probes applied to the chip hybridize with specificity to each spot of DNA on the array. After stringent washing to remove non- specifically bound probes, the chip is scanned by confocal laser microscopy or by another detection method, such as a CCD camera. Quantitation of hybridization of each arrayed element allows for assessment of corresponding mRNA abundance.
[0076] With dual color fluorescence, separately labeled cDNA probes generated from two sources of RNA can be hybridized pairwise to the array. The relative abundance of the transcripts from the two sources corresponding to each specified gene is thus determined simultaneously. A miniaturized scale can be used for hybridization, which provides convenient and rapid evaluation of the expression pattern for large numbers of genes. Such methods have been shown to have the sensitivity required to detect rare transcripts, which are expressed at a few copies per cell, and to reproducibly detect at least approximately two-fold differences in the expression levels (Schena et al., Proc. Natl. Acad. Sci. USA 93:106-49, 1996). Microarray analysis can be performed by commercially available equipment, following manufacturer's protocols, such as by using the Affymetrix GenChip technology, or Agilent inkjet microarray technology. The development of microarray methods for large-scale analysis of gene expression makes it possible to search systematically for molecular markers of cancer classification and outcome prediction in a variety of tumor types.
[0077] As used herein “level”, refers to a measure of the amount of, or a concentration of a transcription product, for instance, an mRNA, or a translation product, for instance, a protein or polypeptide.
[0078] As used herein “activity” refers to a measure of the ability of a transcription product or a translation product to produce a biological effect or to a measure of a level of biologically active molecules.
[0079] As used herein “expression level” further refers to gene expression levels or gene activity. Gene expression can be defined as the utilization of the information contained in a gene by transcription and translation leading to the production of a gene product.
[0080] The terms “increased,” or “increase” in connection with the expression of the biomarkers described herein generally mean an increase by a statically significant amount. For the avoidance of any doubt, the terms “increased” or “increase” means an increase of at least 10% as compared to a reference value, for example, an increase of at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90% or up to and including a 100% increase or any increase between 10-100% as compared to a reference value or level, or at least about a 1.5-fold, at least about a 1.6-fold, at least about a 1.7-fold, at least about a 1.8-fold, at least about a 1.9-fold, at least about a 2-fold, at least about a 3-fold, or at least about a 4-fold, or at least about a 5-fold, at least about a 10-fold increase, any increase between 2-fold and 10-fold, at least about a 25-fold increase, or greater as compared to a reference level. In some embodiments, an increase is at least about a 1.8-fold increase over a reference value.
[0081] Similarly, the terms “decrease,” or “reduced,” or “reduction,” or “inhibit” in connection with expression of the biomarkers described herein generally refer to a decrease by a statistically significant amount. However, for the avoidance of doubt, “reduced”, “reduction” or “decrease” or “inhibit” means a decrease by at least 10% as compared to a reference level, for example, a decrease by at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90% or up to and including a 100% decrease (e.g. absent level or non-detectable level as compared to a reference sample), or any decrease between 10-100% as compared to a reference level.
[0082] A “reference value” is a predetermined reference level, such as an average or median of expression levels of each of THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 biomarkers in, for example, biological samples from a population of healthy subjects. The reference value can be an average or median of expression levels of each of THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-18, ET- 12H, ET-13T, or THR-i3 biomarkers in a chronological age group matched with the chronological age of the tested subject. In some embodiments, the reference biological samples can also be gender matched. In some embodiments, the reference biological samples can also be cancer-containing tissue from a specific subgroup of patients, such as stage 1, stage 2, stage 3, or grade 1, grade 2, grade3 cancers, non-metastatic cancers, untreated cancers, hormone treatment-resistant cancers, HER2 amplified cancers, triple negative cancers, estrogen negative cancers, or other relevant biological or prognostic subsets. For example, as explained herein, malignancy-associated response signature expression levels in a sample can be assessed relative to normal breast tissue from the same subject, or from a sample from another subject, or from a repository of normal subject samples. If the expression level of a biomarker is greater or less than that of the reference or the average expression level, the biomarker expression is said to be “increased” or “decreased,” respectively, as those terms are defined herein. Exemplary analytical methods for classifying the expression of a biomarker, determining a malignancy-associated response signature status, and scoring a sample for expression of a malignancy-associated response signature biomarker are explained in detail herein.
[0083] Treatment
[0084] Methods are described herein for treating cancer. Such methods can involve administering therapeutic agents that can treat cancers with poor prognosis. Examples of such therapeutic agents can include one or more histone deacetylase inhibitors, histone demethylase inhibitors, mTOR inhibitors, polo-like kinase (PLK) inhibitors, heat shock factor inhibitors, and / or inhibitors of any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET- 13T, or THR-i3 breast cancer cell-origin associated signature biomarkers described herein. In some cases, the cancer includes breast cancer, ovarian cancer, colon cancer, brain cancer, pancreatic cancer, prostate cancer, lung cancer, or melanoma. In some embodiments, the cancer includes leukemia, myeloma, or lymphoma. In some aspects, the cancer includes breast cancer, kidney cancer, cervix cancer, melanoma, ovarian cancer, liver cancer, colon cancer, pancreas cancer, lung cancer, prostate cancer, head & neck cancer, bladder cancer, brain tumors such as glioma, or hematopoietic neoplasms such as multiple myeloma.
[0085] The methods can include downregulating expression of one or more of the following: histone deacetylase, histone demethylase, mTOR, polo-like kinase, proteins with heat shock factors, any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET- 12H, ET-13T, or THR-i3 biomarkers, or a combination thereof. Suitable methods for downregulating such expression can include inhibiting transcription of mRNA; degrading mRNA by methods including, but not limited to, the use of interfering RNA (RNAi); blocking translation of mRNA by methods including, but not limited to, the use of antisense nucleic acids or ribozymes, or the like. In some embodiments, a suitable method for downregulating expression may include providing to the cancer a small interfering RNA (siRNA) targeted to histone deacetylase, histone demethylase, mTOR, polo-like kinase, proteins with heat shock factors, any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, THR- i3, ET-12H, ET-13T biomarkers or a combination.
[0086] Suitable methods for down-regulating the function or activity of histone deacetylase, histone demethylase, mTOR, polo-like kinase, proteins with heat shock factors, any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, THR-i3, ET-12H, ET- 13T or a combination thereof may include administering a small molecule inhibitor that inhibits the function or activity of any of these markers or factors.
[0087] In some cases, one or more histone deacetylase inhibitors can be administered to treat cancers with poor prognosis, such as cancers identified by measuring and / or monitoring any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 biomarkers described herein. In some cases, histone deacetylase inhibitors are not administered to treat cancers with poor prognosis, such as cancers identified by measuring any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 markers described herein. As used herein a “Histone Deacetylase inhibitor” or “HD AC inhibitor” refers to inhibitors of Histone Deacetylase 1 (HDAC1), Histone Deacetylase 7 (HDAC7), and / or phosphorylated HDAC7, including agents that inhibit the level and / or activity of HDAC1 and / or HDAC7 and / or phosphorylated HDAC7, as well as agents that inhibit the phosphorylation of HDAC7 e.g., inhibitors of EMK protein kinase, C-TAK1 protein kinase, and / or CAMK protein kinase, and agents that activate or increase the level and / or activity of phosphatase activity to remove phosphoryl groups from HDAC7, e.g., activators of PP2A phosphatase and / or myosin phosphatase. In some cases, HD AC inhibitors include molecules that bind directly to a functional region of HDAC1 and / or HDAC7 and / or phosphorylated HDAC7 in a manner that interferes with the enzymatic activity of HDAC1 and / or HDAC7 and / or phosphorylated HDAC7 e.g., agents that interfere with substrate binding to HDAC1 and / or HDAC7 and / or phosphorylated HDAC7. In some embodiments, HDAC inhibitors include molecules that bind directly to HDAC7 in a manner that prevents the phosphorylation of HDAC7. HDAC inhibitors include agents that inhibit the activity of peptides, polypeptides, or proteins that modulate the activity of HDAC1 and / or HDAC7 e.g., inhibitors of EMK protein kinase, C-TAK1 kinase, CAMK protein kinase inhibitors of C- TAK1 protein kinase. Examples of suitable inhibitors include but are not limited to antisense oligonucleotides, oligopeptides, interfering RNA e.g., small interfering RNA (siRNA), small hairpin RNA (shRNA), aptamers, ribozymes, small molecule inhibitors, or antibodies or fragments thereof, and combinations thereof.
[0088] In some cases, HDAC inhibitors are specific inhibitors or specifically inhibit the level and / or activity of HDAC1 and / or HDAC7 and / or phosphorylated HDAC7. As used herein, “specific inhibitor(s)” refers to inhibitors characterized by their ability to bind to with high affinity and high specificity to HDAC1 and / or HDAC7 and / or phosphorylated HDAC7 proteins or domains, motifs, or fragments thereof, or variants thereof, and preferably have little or no binding affinity for non-HDACl and / or non-HDAC7 and / or non-phosphorylated HDAC7 proteins. As used herein, “specifically inhibit(s)” refers to the ability of an HDAC inhibitor of the present invention to inhibit the level and / or activity of a target polypeptide, e.g., HDAC1, and / or HDAC7, and / or phosphorylated HDAC7, and / or EMK protein kinase, and / or C-TAK1 protein kinase and / or CAMK protein kinase and preferably have little or no inhibitory effect on non-target polypeptides. As used herein, “specifically activate(s)” and “specifically increase(s)” refers to the ability of an HDAC inhibitor of the present invention to stimulate (e.g., activate or increase) the level and / or activity of a target polypeptide, e.g., PP2A phosphatase and / or myosin phosphatase and preferably to have little or no stimulatory effect on non-target polypeptides.
[0089] Examples of HDAC inhibitors include Vorinostat (SAHA), Entinostat (MS-275), Panobinostat (LBH589), Trichostatin A (TSA), Mocetinostat (MGCD0103), 4-Phenylbutyric acid (4-PBA), ACY-775, Belinostat (PXD101), Romidepsin (FK228, Depsipeptide), MC1568, Tubastatin A HC1, Givinostat (ITF2357), Dacinostat (LAQ824), CUDC-101, Quisinostat (JNJ- 26481585) 2HC1, Pracinostat (SB939), PCI-34051, Droxinostat, Abexinostat (PCI-24781), RGFP966, AR-42, Ricolinostat (ACY-1215), Valproic Acid (NSC 93819) sodium salt, Tacedinaline (CI994), Fimepinostat (CUDC-907), Sodium butyrate, Curcumin, M344, Tubacin, RG2833 (RGFP109), Resminostat, Divalproex Sodium, Scriptaid, Sodium Phenylbutyrate, Tubastatin A, Tubastatin A TFA, Sinapinic Acid, TMP269, Santacruzamate A (CAY10683), TMP195, Valproic acid (VPA), UF010, Tasquinimod, SKLB-23bb, Isoguanosine, NKL22, Sulforaphane, BRD73954, BG45, Domatinostat (4SC-202), Citarinostat (ACY-241), Suberohydroxamic acid, BRD33O8, Splitomicin, HPOB, LMK-235, Biphenyl-4-sulfonyl chloride, Nexturastat A, BML-210 (CAY10433), TC-H 106, SR-4370, TH34, Tucidinostat (Chidamide), SIS17, (-)-Parthenolide, WT161, CAY10603, ACY-738, Raddeanin A, GSK3117391, Tinostamustine(EDO-SlOl), or combinations thereof. Such HDAC inhibitors are available from Selleckchem.com.
[0090] In some cases, one or more histone demethylase inhibitors can be administered to treat cancers with poor prognosis, such as cancers identified by measuring and / or monitoring any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 biomarkers described herein. Examples of histone demethylase inhibitors include GSK-I4, 2,4-Pyridinedicarboxylic Acid, AS8351, Clorgyline hydrochloride, CPI-455, Daminozide, GSK-2879552, GSK-J1, GSK-J2, GSK-J5, GSK-LSD1, IOX1, IOX2, JIB-04, ML-324, NCGC00244536, OG-L002, ORY-lOOl, SP-2509, TC-E 5002, UNC-926, 0- Lapachone, or combinations thereof. Such inhibitors are available, e.g., from Selleckchem.com.
[0091] In some cases, one or more mTOR inhibitors can be administered to treat cancers with poor prognosis, such as cancers identified by measuring and / or monitoring any THR-50, THR- 70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 biomarkers described herein. Examples of mTOR inhibitors include Rapamycin (AY-22989), Everolimus (RAD001), AZD8O55, Temsirolimus (CCI-779), PI-103, NU7441 (KU-57788), KU-0063794, Torkinib (PP242), Ridaforolimus (Deforolimus, MK-8669), Sapanisertib (MLN0128), Voxtalisib (XL765) Analogue, Torin 1, Omipalisib (GSK2126458), OSI-027, PF- 04691502, Apitolisib (GDC-0980), GSK1059615, WYE-354, Gedatolisib (PKI-587), Vistusertib (AZD2014), Torin 2, WYE-125132 (WYE-132), BGT226 (NVP-BGT226) maleate, Palomid 529 (P529), PP121, WYE-687, Clemastine (HS-592) fumarate, Nitazoxanide (NSC 697855), WAY-600, ETP-46464, GDC-0349, PI3K / Akt Inhibitor Library, 4EGI-1, XL388, MHY1485, 3 -Hydroxy anthranilic acid, Bimiralisib (PQR309), Samotolisib (LY3023414), Lanatoside C, Rotundic acid, L-Leucine, Chrysophanic Acid, Voxtalisib (XL765), GNE-477, CZ415, Astragaloside IV, CC-115, Salidroside, Compound 401, 3BDO, Zotarolimus (ABT-578), GNE-493, Paxalisib (GDC-0084), Onatasertib (CC 223), ABTL- 0812, PQR620, SF2523, Niclosamide, or combinations thereof. Such HDAC inhibitors are available from Selleckchem.com.
[0092] In some cases, one or more Polo-Like Kinase (PLK) inhibitors can be administered to treat cancers with poor prognosis, such as cancers identified by measuring and / or monitoring any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET- 13T, or THR-i3 biomarkers described herein. Examples of PLK inhibitors include BI 2536, Volasertib (BI 6727), Wortmannin (KY 12420), Rigosertib (ON-01910), GSK461364, HMN- 214, MLN0905, Ro3280, SBE 13 HC1, Centrinone (LCR-263), CFI-400945, HMN-176, Onvansertib (NMS-P937), or combinations thereof.
[0093] In some cases, one or more heat shock factor inhibitors can be administered to treat cancers with poor prognosis, such as cancers identified by measuring and / or monitoring any of the THR-50, THR-70, CA-110, CA-130, THR-6E, THR-4E, i20, THR-i8, ET-12H, ET-13T, or THR-i3 biomarkers described herein. Examples of heat shock factor inhibitors include one or more of the following Tanespimycin (17-AAG), Pimitespib (TAS-116, Luminespib (NVP- AUY922), Alvespimycin (17-DMAG) HC1, Ganetespib (STA-9090), Onalespib (AT13387), Geldanamycin (NSC 122750), SNX-2112 (PF-04928473), PF-04929113 (SNX-5422), KW- 2478, Cucurbitacin D, VER155008, VER-50589, CH5138303, VER-49009, NMS-E973, Zelavespib (PU-H71), HSP990 (NVP-HSP990), XL888 NVP-BEP800, BIIB021or a combination thereof. Such heat shock factor inhibitors can be obtained from Tocris.com.
[0094] In some cases, chemotherapy (e.g., cyclophosphamide, methotrexate, 5-fluorouracil, vinorelbie, doxorubicin, docetaxel, bleomycin, vinblastine, dacarbazine, mustine, vincristine, procarbazine, prednisolone, Bleomycin, etoposide, cisplatin, epirubicin, capecitabine, methotrexate, folinic acid, oxaliplatin, gemcitabine, ifosfamide, and others or combinations thereof), endocrine therapy (e.g., hormone therapy, such as androgen deprivation or estrogen deprivation or high dose estrogen therapy) and / or immunotherapy can be administered to treat cancers.
[0095] As used herein, “solid tumor’’ is intended to include, but not be limited to, the following sarcomas and carcinomas: fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, testicular tumor, lung carcinoma, small cell lung carcinoma, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, and retinoblastoma. Solid tumor is also intended to encompass epithelial cancers.
[0096] The THR-50, THR-70, i20, ET-60, CA-110, and CA-130 signature genes are listed below in Table 1.
[0097] TABLE 1
[0098]
[0099]
[0100] THR-6E, THR-4E, THR-i8, THR-i3, and ET-12H signatures are prognostic for relapse- free and overall survival in different subsets of breast cancer, and predictive of anti-estrogen and anti-HER2 targeted treatments, as well as chemotherapy in breast cancer (Table 4).
[0101] Isoforms and variants of the THR-50, THR-70, CA-1 10, CA-130, THR-6E, THR-4E, i20, THR-i8, THR-i3, and ET-12T genes and gene products can be present in subjects and can be detected, measured, evaluated, and the subjects with such isoforms and variants can be treated by the methods and compositions described herein. Such isoforms and variants can have sequences with between 65-100% sequence identity to a reference sequence, for example with at least 65%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% sequence, at least 98%, at least 99%, or at least 99.5% identity to a sequence described herein or a reference sequence (such as one described in the NCBI or Uniprot databases) over a specified comparison window. Optimal alignment may be ascertained or conducted using the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol. 48:443-53 (1970).
[0102] Definitions
[0103] The “absolute amplitude” of correlation expressions means the distance, either positive or negative, from a zero value; i.e., both correlation coefficients -0.35 and 0.35 have an absolute amplitude of 0.35.
[0104] “Status’" means a state of gene expression of a set of genetic markers whose expression is strongly correlated with a particular phenotype.
[0105] “Good prognosis” means that a patient is expected to have longer overall survival (OS), or progression-free survival (PFS), or disease-specific survival (DSS), or recurrence-free survival (RFS) compared to “poor prognosis” patients. These metrics are typically described by National Cancer Institute (NCI) as overall survival (OS), or progression-free survival (PFS) which is the length of time during and after the treatment of cancer, that a patient lives with the disease but it does not get worse, or disease-specific survival (DSS) that is the percentage of people in a treatment group who have not died from their cancer in a defined period of time, or recurrence-free survival (RFS) that is length of time after primary treatment for a cancer ends that the patient survives without any signs or symptoms of that cancer, also called as disease- free survival (DFS), or relapse-free survival (see website at cancer.gov / publications / dictionaries / cancer-terms / def / rfs).
[0106] “Poor prognosis” means that a patient is expected to have a shorter overall survival (OS), or progression -free survival (PFS), or disease-specific survival (DSS) or recurrence-free survival (RFS) compared to “good prognosis” patients.
[0107] “Marker” means an entire gene, mRNA, EST, or a protein product derived from that gene, where the expression or level of expression changes under different conditions, where the expression of the gene (or combination of genes) correlates with a certain condition, the gene or combination of genes is a marker for that condition.
[0108] “Marker-derived polynucleotides” means the RNA transcribed from a marker gene, any cDNA, or cRNA produced therefrom, and any nucleic acid derived therefrom, such as synthetic nucleic acid having a sequence derived from the gene corresponding to the marker gene.
[0109] A “similarity value” is a number that represents the degree of similarity between two things being compared. For example, a similarity value may be a number that indicates the overall similarity between a patient's expression profile using specific phenotype-related markers and a control specific to that phenotype (for instance, the similarity to a “good prognosis” template, where the phenotype is a good prognosis). The similarity value may be expressed as a similarity metric, such as a correlation coefficient, or may simply be expressed as the expression level difference, or the aggregate of the expression level differences, between a patient sample and a template.
[0110] The present description is further illustrated by the following examples, which should not be construed as limiting in any way. EXAMPLES
[0111] Example I
[0112] Introduction
[0113] Motivated by the lymphoma / leukemia classification paradigm, provided herein is the development of a similar approach for breast cancer, and other gynecological cancers.
[0114] Tumor phenotype is shaped both by transforming genomic alterations and the normal cell-origin. While it was shown that the normal cell-of-origin signatures dominate the molecular tumor patterns, unlike genetic alterations, it has been difficult to translate this genome-wide signature into actionable information. Provided herein is the development of a breast cellular ancestry signature guided by cell-type specific co-expression of hormone receptors for androgen (AR), estrogen (ER), and vitamin D (VDR). This triple hormone receptor (THR) pattern was used to identify mRNA signatures (THR-50 and -70) that differentiate between breast cancers with low and high triple hormone receptor co-expression. It was determined that both THR signatures predict overall survival and progression-free survival in breast cancer patients in multiple datasets and it was demonstrated that they provide robust prognostic performance across different clinical and molecular breast cancer subtypes, surpassing existing prognostic signatures.
[0115] Four differentiation types were identified in normal breast luminal epithelial (NBLE) cells in human breast lobules based on the co-expression of estrogen, androgen, and vitamin- D receptors; ER, AR, and VDR (3,4). It was found that some NBLE cells, defined as triple hormone receptor-positive (THR3), express all three receptors, while others may express two, one, or none of these receptors, respectively defined as THR2, THR1, and THR-zero (THRO) (3,4). Given the association of ER, AR, and VDR with differentiation, it was hypothesized that the THR state may represent the differentiation stages of NBLE cells.
[0116] Numerous studies propose that DNA methylation plays a role in cellular differentiation programs. Consistent with this observation, it was discovered that each THR differentiation state corresponds to a unique DNA methylation signature (5). This suggested that the combined expression of ER, AR, and VDR can serve as a surrogate index for differentiation. Interestingly, it was discovered that both the DNA methylation signature and the THR states of NBLE cells are preserved in human breast tumors (6). While NBLE cells exhibit a heterogeneous phenotype with all THR states present within the same lobule, human breast cancers predominantly exhibit one of these differentiation states (7). Thus, each breast cancer can be categorized as THRO, THR1, THR2, or THR3 phenotype. This suggests that either human breast tumors primarily preserve the initial THR state of their singular normal cell-of- origin, or a differentiation block due to acquired mutations restricts them to one dominant THR state, akin to hematopoietic malignancies (4). Alternatively, a lineage-specific expression profile may emerge due to the unique interaction of the cell-origin with oncogenic alterations, as has been shown previously (Ince et al., Cancer Cell, 2007). The term ‘cellular ancestry’ is used to encompass all of these different possibilities.
[0117] While the estrogen receptor (ER) has been a key prognostic and predictive marker in breast cancer (8), the prognostic role of VDR and AR has been more unsettled. As a single marker, it is well-established that VDR protein expression does not correlate with breast cancer overall survival outcome in multivariate analysis (9-12). While some studies found a correlation between AR protein expression and breast cancer outcome, others failed to confirm this conclusion (13). For example, AR was identified as an independent overall survival (OS) prognostic marker < 50% of reviewed studies (10 / 22), in a meta-analysis of unselected breast cancer. In ER-positive breast cancer, AR was identified as a prognostic marker in five out of eight by multivariate analysis (14). In ER-negative breast cancer, AR had no statistically significant impact on OS in a meta-analysis of thirteen studies with 2,826 patients (15). Variable results were also observed in other studies (3), while AR mRNA level did not correlate with breast cancer (p = 0.13) (3), it was found that AR protein expression is associated with improved prognosis in ER+ and worse prognosis in ER- breast cancer (16,17). However, in another study, AR expression was not associated with breast cancer-free interval (18). In summary, the role of AR as a prognostic marker was controversial at best and VDR did not appear to be prognostic in most studies. Therefore, the initial motivation behind combining AR and VDR with ER was purely based on the ability of this triple marker panel to identify distinct normal breast cell types. Based on the reviewed literature, an additive prognostic power by combining AR and VDR with ER was not expected. Nevertheless, it was found that the combination of these three markers can be unexpectedly powerful as a prognostic panel (3,4,19).
[0118] Several prognostic and predictive biomarkers currently inform breast cancer management in the clinical setting. Since the 1990s, breast tumors have typically been categorized into three groups based on the presence of ER, PR, and the human epidermal growth factor receptor 2 (HER2+). These proteins have functioned as both prognostic and predictive biomarkers, aiding in the selection of patients for anti-estrogen or anti-HER2 targeted therapies. The estrogen receptor (ER) is a crucial biomarker, with approximately 70% of breast cancers being ER-positive (8). The activation of ER by estrogen triggers the growth of cancer cells, rendering ER-positive tumors responsive to hormonal therapies such as tamoxifen or aromatase inhibitors (20,21). HER2 is amplified in 10-15% of breast cancers that are treated with anti-HER2 monoclonal antibodies such as Herceptin or small molecules that inhibit kinase activity and signaling such as lapatinib. The remainder of tumors that are negative for ER, PR, and HER2 are called triple-negative breast cancers (TNBC), that have no targeted therapies.
[0119] In the last two decades, several mRNA expression-based prognostic signatures such as MammaPrint (22,23), Oncotype Dx (24), and PAM-50 (25) have been developed to move beyond the ER / PR / HER2 paradigm. MammaPrint measures the expression levels of 70 genes using microarrays to predict recurrence probability in breast cancer patients (22,23). Similarly, Oncotype DX uses RT-PCR to assess the expression of 21 genes to calculate a recurrence score which is then used to divide patients into low-, intermediate-, and high-risk groups (24). However, it's important to note that the application and prognostic value of these signatures are limited to early-stage, node-negative, small, ER-positive, and HER2-negative tumors (26,27). As such they are only useful in less than half of the patient population (26). Moreover, these tests were developed by identifying differentially expressed genes associated with survival. As such these signatures include an amalgam of signals from the tumor cells, stroma, normal cells, inflammatory cells, etc. While these complex signatures are useful for prognostic clustering, they are difficult to implement for a traditional cell-type -based classification paradigm.
[0120] Results
[0121] Classification of Breast Tumors by Triple Hormone Receptor Protein Expression
[0122] The triple-hormone receptor (THR) categorization is based on protein co-expression of ER, AR, and VDR, assessed by immunohistochemical (IHC) staining of formalin-fixed paraffin-embedded (FFPE) breast cancer tissue microarrays (TMA) (3,4). The individual IHC stains are scored as positive or negative as described before, and breast cancers are organized in four subgroups: THRO, THR1, THR2, and THR3, representing 7%, 11%, 28%, and 54% of breast cancers respectively (Figure 1A-B). Kaplan-Meier survival analysis shows that breast cancers with fewer number hormone receptors were associated with shorter overall survival (Figure IB) with a statistically significant (p< 0.002) hazard ratio (HR) in multivariate analysis (Figure 1C): THRO = 2.6, 95%CI: 1.6-4. 1, THR1 = 2.5, 95%CI: 1.8-3.5, THR2 = 1.61 95%CI 1.19- 2.19, and THR3 = 1.0).
[0123] Derivation of a Triple Hormone Receptor mRNA Signature: THR-50
[0124] Next, it was sought to validate the THR index in other datasets. However, IHC-based THR results of breast cancer TMAs with outcome data were not publicly available. Thus, an mRNA signature was developed that would serve as a surrogate for THR to allow one to examine publicly available breast cancer gene expression data sets.
[0125] To identify a gene list that can separate breast cancers with THR0&1 phenotype from THR2&3, gene expression data was used from the Expression data from the Cancer Cell Line Encyclopedia (CCLE) dataset (28). First, the expression profiles were compared between THR0&1 cell lines (BT-20, HCC1187, HCC1937, HCC1143, and MDA-MB-231) and THR2&3 cell lines (MCF7, T47D, CAMA-1, YMB-1, ZR-75-1), which identified 300 differentially expressed genes (DEGs) between both groups. The top 50 most significant genes (lowest p-value) were selected to study further, thereafter, called the THR-50 signature (see Figure 2A and Table 1).
[0126] Validation of the THR-50 Signature in Human Breast Cancer
[0127] The THR-50 signature was examined in the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC) dataset (29) (1904 samples), and it was found that it split the human tumors into two major clusters similar to the CCLE dataset (Figure 2B). As expected, the THR2&3 associated genes in CCLE (Figure 2A) are highly expressed in ER+ human breast cancer in the METABRIC dataset (Figure 2B). Additionally, THR-50 was significantly associated with recurrence-free survival (RFS) (HR= 1.5, 95%CI: 1.3-1.7) and overall survival (OS) (HR= 1.7, 95%CI: 1.5. -1.9) (see Figure 2C-D). These results suggested that THR-50 can serve as a mRNA surrogate for IHC-based THR categories.
[0128] Analysis of THR-50 Signature Across Breast Cancer Subtypes
[0129] Having observed that THR-50 signature correlates with outcome in breast cancer globally, breast cancer molecular subtypes were analyzed based on clinical clusters based on ER, HER2, and Mibl IHC and PAM-50 groups. The THR-50 prediction probability scores were used to divide the METABRIC samples into four equal-sized quartiles, with quartile 1 (QI) and quartile 4 (Q4) samples having the lowest and highest probabilities of survival, respectively. Subsequently, the difference in survival between QI and Q4 samples was examined in different breast cancer subtypes, based on the three IHC markers and PAM-50 groups.
[0130] Notably, Q4 samples had a significantly less favorable RFS compared to QI samples in all clinical subtypes, except in HER2+ tumors where they had a more favorable RFS (Figure 3A). Across the PAM-50 groups, Q4 samples had a significantly less favorable RFS compared to QI samples in the basal, luminal A, and B subtypes, while no significant difference was observed in the Claudin-low, and HER2 subtypes (Figure 3B). Similarly, in comparison to QI, Q4 samples had a favorable OS in all clinical subtypes except HER2+, and in the PAM-50 basal, luminal A, and B subtypes (Figure 20). These results indicated that THR cellular ancestry signatures can be used as a prognostic index to divide breast cancer patients into good vs. poor outcome groups, which we examined in further detail next.
[0131] The THR-50 Signature Surpasses Usual Prognostic Biomarker Tests
[0132] Multigene biomarker tests provide useful prognostic information for a selected group of breast cancer patients (30). For early-stage, ER-positive, HER2-negative, lymph nodenegative breast tumors, the Oncotype, Prosignia (PAM-50), MammaPrint (MAM-70), and Endopredict signatures are now recommended by ASCO (31,32). Furthermore, the MammaPrint signature is suggested for breast tumors with up to three metastatic lymph nodes (Nl) (33,34). However, these tests are not recommended by ASCO for ER-negative, HER2- positive, lymph-node metastatic (>N1), or treated breast cancers (35,36).
[0133] The performance of the THR-50 signature was compared to other known breast cancer prognostic signatures for recurrence-free survival (RFS) in another Kaplan-Maier Plot (KMP) patient cohort comprising 2,032 samples from 50 gene expression datasets. Patients were divided into two groups (high versus low risk) based on the expression of genes in each subtyping system and then the survival of the two groups was compared using Kaplan-Meier survival estimate.
[0134] Notably, the THR-50 signature performed better at describing two distinct RFS cohorts (p<1016), compared to the MammaPrint (MAM-70, NS, p=0.11), Oncotype Dx, (ONC-21, NS, p=0.09), and PAM-50 (p=4xl0-5) systems. In all breast cancer samples (n=2,032), patients with a low mean expression of the THR-50 genes had shorter RFS compared to those with a high mean expression (HR= 2.04, 95%CI: 1.8-2.4), while this difference in survival was not significant using the MammaPrint, Oncotype DX, and less notable with PAM-50 (HR= 1.37, 95%CI: 1.2-1.6), (Figure 4A). The THR-50 signature was even capable of significantly predicting RFS in lymph node-positive (HR=2.7, 95%CI: 1.9-3.8), AR-positive (HR=4.6, 95%CI: 2.0-11.0), grade 2 (HR=2.1, 95%CI: 1. 1-4.0), and grade 3 (HR=1.6, 95%CI: 1.1-2.5) breast cancer subgroups (Figures 4B and SI). PAM-50 was not associated with RFS in AR+, grade 2, and grade 3 breast cancers (Figure 4C).
[0135] Additionally, the THR-50 signature split all PAM-50 groups into patients with shorter and longer RFS (Figure 5A). The PAM-50 groups with high mean expression of the THR-50 genes had a significantly shorter RFS (Lum-A HR=2.2, Lum-B HR=2.0, Her2-like HR=3.3, basal-like HR=3.2), indicating that THR50 provides additional information about the tumor biology distinct from PAM-50. As expected, Oncotype DX was not significantly associated with RFS in any of the PAM-50 groups (Figure 5B), while MammaPrint was significant only in patients with HER2+ breast cancer (Figure 5C). In summary, THR-50 is significantly associated with survival in all eight clinically relevant breast cancer subtypes examined (LN+, AR+, Grade2, Grade3, Lum-A, Lum-B, HER2-like, and Basal-like), whereas the other signatures are significant in one subtype or none (37).
[0136] Derivation and Validation of the THR-70 Signature
[0137] In addition to the THR-50 signature which was developed using expression profiles from breast cancer cell lines, human breast cancer datasets were also used to develop a surrogate mRNA signature for THR. Specifically, 855 breast cancer cases were compiled from four datasets (22,38-40), hereafter referred to as BC855. Subsequently, the samples were categorized based on the co-expression of ESRI, AR, and VDR into four groups: THRO, THR1, THR2, and THR3 (Figure 6A).
[0138] PAM-50 subtyping in BC855 shows that the THR categories are not a mere relabeling of existing groups. Indeed, each THR group includes all six PAM-50 subtypes with fluctuating proportions. Interestingly, THR1 that has the worst outcome includes 15.6 % Luminal-A, 18.8 % Luminal-B, 18.4 % HER2-enriched, 13.9 % Claudin-low, 7.7% Normal like and 24.3 % Basal-like PAM-50 subtypes. The nearly balanced distribution of these subtypes suggests that the poor outcome is not due to the dominance of this THR group by one PAM-50 category (Figure 6A). The signatures derived in CCLE and BC855 were examined and found that 350 genes are associated with THR both in cell lines (CCLE) and human breast cancers (BC-855) (Figure 6B). The top 70 genes based on SAM fold change were selected and referred to as the THR-70 signature (see Table 1).
[0139] Using a similar approach to the one used in the THR-50 signature, the association of the THR-70 signature with RFS and OS in the METABRIC dataset across different breast cancer subtypes was examined. In all subtypes except HER2+, patients with the highest probability of OS events (Q4) had the least favorable RFS and OS compared to those with the lowest probability (QI) (Figures 6C-D, 21, and 22). These results show that the THR-70 signature is also capable of capturing prognostic information across different breast cancer subtypes which further highlights the utility of classification systems based on the THR status in breast cancer.
[0140] To validate the prognostic performance, the THR-70 signature was further tested in the Meta-10 cohort which includes samples from 10 different gene expression datasets. In all samples (n=1888), the signature was significantly associated with RFS (HR= 2.54, 95%CI: 2.03-3.18) (Figure 7A) and distant metastasis-free survival (DMFS) (HR= 3.77, 95%CI: 2.7- 5.26) (Figure 7B). This performance was also maintained in different patient subgroups including those with LN+ (n=240) and LN- (n=1047) disease (Figure 7C-D), together with patients who received endocrine therapy following surgery (n=492) (Figure 7E) and those who did not receive neoadjuvant treatment (n=1791) (Figure 7F). Cumulatively, these results are striking since it is unusual for a standard prognostic signature to correlate with such varied aspects of tumor biology and in multiple datasets (see Figures 23 and 24). However, this may not be unusual for a cellular ancestry signature that would be expected to influence all aspects of the tumor phenotype.
[0141] THR-50 and -70 are Associated with Cellular Ancestry and Immune Gene Set Signatures
[0142] To gain further biological insights about the THR-50 and THR-70 signatures, a gene set variation analysis (GSVA) was performed to delineate their functional roles. Consistent with their cell-of-origin derivation, both THR signatures were enriched for AR and ER pathways, which are negatively correlated with PAM-50, MAM-70, and PCNA-131 signatures. Interestingly, THR-70 is also correlated with the epithelial-mesenchymal transition (EMT) signature (Figure 8A).
[0143] Proliferation-related genes were found to be over-represented in 22 of 24 breast prognostic signatures (41), and removing proliferation-associated genes from 47 published breast cancer prognostic signatures, significantly reduced or eliminated their association with outcome (42), indicating that most breast cancer prognostic tests may essentially function as surrogate proliferation markers (43). Consistent with this, in GSVA it was found that the most significant gene set associated with PAM-50 and MAM-70 are cell cycle and apoptosis (Figure 8A), further underscoring the novelty of the THR signatures disclosed herein at capturing novel dimensions of breast cancer biology.
[0144] Furthermore, the potential correlation of these signatures with immune response was explored by measuring their enrichment in gene signatures associated with different immune cell types. Notably, both THR signatures were highly enriched in tumors with an immune infiltrate containing higher central memory, gamma delta (yS), CD4+ T, Thl7, T follicular helper (Tfh), NK, and MAIT cell subsets (Figure 8B). In contrast, PAM-50, MammaPrint, and PCNA-131 classifiers were more enriched in signatures of myeloid lineages like neutrophils, dendritic cells, and monocytes (Figure 8B). While the y5 T cells are involved in anti-tumor cytotoxicity suppressing various anti-tumor responses and increasing angiogenesis. The dominant intra-tumoral Thl CD4+subset in early-stage cancers switches to Treg and Thl7 in the late stages of progression, associated with poor breast cancer patient outcomes. The T follicular helper cells (Tfh), another subset of CD4+T cells, play a role in helping B cells produce antibodies, and CD8+ T cells are associated with favorable clinical outcomes (44). Thus, the infiltrates associated with THR signatures contain anti-tumor and pro-tumor growth components. There is a remarkable association between cell-origin and immune response; the THR markers correlate with immune infiltrates that differ significantly from MAM-70 and PAM-50.
[0145] Unsupervised Clustering of Breast Cancer Samples Using the THR Signatures Reveals Distinct Subtypes
[0146] Having shown that THR signatures can be utilized as a prognostic index to identify breast cancers with different outcomes, the utility of THR in developing a cell type-based breast cancer classification was explored. To achieve this, the expression levels of the THR signatures were used to perform unsupervised clustering of the METABRIC dataset to identify different subgroups and examine their specific survival rates.
[0147] It was found that the THR-70 signature divides the METABRIC dataset into 5 distinct groups named El, E2a, E2b, E3, and PQNBC, with the E groups being predominantly ER+ compared to the PQNBC group which comprised mainly ER- samples (see Figure 9). The E2a and E2b were combined into a single prognostic group E2 given their overlap (Figure 25). Similarly, the THR-50 signature was able to identify 5 distinct groups in the METABRIC dataset: El, E2, E3, E4, and PQNBC (data not shown). While there is some overlap between Lum-A, Lum-B, lower grade, and THR-E, these clusters also contained HER2+, claudin-low, and high-grade tumors (see Figures 9 and 28). The PQNBC cluster included tumors that are pentaplex-negative for five markers ER, PR, AR, VDR, and HER2 (PNBC), as well tumors quadruple-negative for four markers ER, PR, AR, and VDR (QNBC). The PQNBC cluster predominantly contains basal-like (49.7%), claudin-low (34.5%), and HER2-like (14.7%) PAM-50 groups (Figure 28). An immune signature consisting of 20 genes (hereafter called i20) was used to further divide the PQNBC cluster into two subgroups (PQNBCi-i- and PQNBCi-) with PQNBCi-i- characterized by high immune infiltration compared to the PQNBCi- subgroup. This combined signature is referred to as THR-70i (Table 1).
[0148] Next, the RFS rates in the THR + i20-derived groups were examined using Kaplan- Meier survival estimates. Interestingly, compared to existing classifiers like the clinical three marker IHC and PAM-50, the THR signatures were able to identify distinct subtypes of breast cancer with significantly fewer cross-over in survival curves. For example, while TNBC and basal-like tumors are generally mentioned as worse outcome groups, however, their survival curve crosses over both ER-positive clusters in clinical and PAM-50 groups. Moreover, the claudin-low cluster cross-over with Lum-A and HER2+ cluster cross-over with Lum-B (Figure 10). It is worth pointing out that these multiple crossovers in PAM-50 survival curves are seen in all other studies and multiple datasets (25,45-50). Therefore, while a cluster may have a worse outcome during the first five years, it may have a better outcome at ten years, which makes their clinical utilization difficult. In contrast, it was observed that the THR signature separated PQNBC clusters by a wide- margin, from E clusters, and particularly the THR-70i signature, eliminated all but one cross-over between two good outcome clusters E3 and Pi-i- (Figure 10).
[0149] Furthermore, the THR-70i signature significantly increased the dynamic range of prognostic clusters. For example, it was observed that Lum-A, ER+LP, Lum-B, ER+HP, Basal- like, TNBC, and HER-2-like tumors have a survival probability between 45% to 65% at 20 years. In contrast, the survival probability of PQNBCi-i- is 75 to 80% at 20 years, compared to 25% to 30% for PQNBCi+. In experimental systems, it is shown that some of the early driver mutations may become redundant or non-essential during the later progression of the same tumor. Since tumors are constantly evolving by accumulating new mutations and genetic alterations, it may not be surprising that a prognostic signature that is applied to the primary tumor tissue may lose some of its predictive power after 10 or 20 years. Since the genes in the PAM-50 signature are mutated in 35% of breast cancers, it is possible some of the crossovers in the PAM-50 survival curves may be due to such genetic evolution. In contrast, only 1.2% of THR-50 genes are mutated in breast cancers, consistent with a cell-of-origin signature that may be more durable over the lifespan of the tumor, since it is not dependent on mutations (see Figure 26).
[0150] Next, the difference in RFS in breast cancer subgroups was examined. While there is an approximately 1.5 -fold survival probability difference between the ER-negative groups defined by IHC (TNBC vs. HER2) or PAM-50 (basal-like vs. claudin-low), the hazard ratio between THR70+i20 clusters (PQNBCi-i- vs. PQNBCi-) is over 15-fold. Thus, THR signatures provide a significant improvement in identifying ER-negative breast tumors with extremely different outcomes (Figure 11A).
[0151] While IHC and PAM-50 identify two ER-positive subtypes with a 1.5-to-l.8-fold difference in hazard rations, THR-70 was able to identify three significant ER+ clusters (El, E2, and E3) with a range of 2. 1 -fold difference in survival, and a further gradation of estrogen- driven tumors (Figure 11B).
[0152] Finally, the effect of HER2 status on the survival estimates of the THR-derived clusters was examined. In the PQNBC cluster, HER2 status did not affect RFS (p-vahie=0.28) (Figure 12A), further underscoring the importance of the i20 gene signature which effectively separated this cluster into two subgroups with significantly different survival rates (HR=15.7), as shown previously in Figure 11A. In contrast, HER2 status was significantly associated with RFS in the E clusters overall (p-value<0.001) (Figure 12B) and also in El (p-value=0.012) and E2 (p- value<0.001) separately, but not in E3 (Figure 27). Interestingly, it seems that in the highest THR signature group (E3) it seems that hormone receptors dominate the outcome; thus, the E3+ / HER2+ patients have a similar survival with E3+ / HER2- patients. Whereas, in E2 / E1 clusters with a lower THR expression it seems that HER2 dominates hormone receptors, suggesting that overlaying this genetic profile over the cellular ancestry signature may further improve the stratification of breast cancer patients. Likewise, it was found that THR-50 signature can be used to stratify HER2+ patients into two outcome groups with significantly different survival (HR=2.2 CI: 1.4-3.7) (Figure 12C).
[0153] In brief, THR cell-of-origin signature reveals a new dimension of breast cancer biology, that is distinct from the existing clustering schemes (Figure 28), which can provide a robust and durable foundation on which other prognostic biomarkers can be overlayed to produce an improved breast cancer stratification.
[0154] Next, it was found that initially grouping human breast cancers based on the cell ancestry and immune THR-70+i20 signature, followed by separating the genetic category of HER2-amplified breast cancers, helps in identifying six prognostic groups. These groups exhibit a significant variation in hazard ratios: PQNBC_i- (HR=1), El (HR=1.6), E2 (HR=2.5), E3 (HR=3.0), HER2 (HR=3.7), and PQNBC_i+ (HR=5.8). (Figure 13A). Notably, the PAM- 50 cluster KM curves cross each other four times within the first years. Thus, it is not possible to carry out the Cox -proportional hazard analysis to calculate hazard ratios with the PAM-50 signature (Figure 13B).
[0155] Combination of THR and metastatic cellular ancestry signatures identify prognostic groups in multiple tumor types.
[0156] Previously, a metastatic breast cancer cellular ancestry signature (ET60) that also predicts breast cancer outcome (57) and drug response (19) was identified. The ET60 signature was disclosed as CORNELL 10059-02-PC, and a U.S. application No. 63 / 292,943, was filed December 22, 2021, which was published on June 29, 2023, as WO2023122758A1, incorporated herein by reference. The ET60 signature is reproduced herein above.
[0157] Interestingly, it was found that the combination of ET60 is additive with THR signatures. For example, while the hazard ratio of ET60 and THR50 are 3.5 and 2.6 respectively, their combination, hereafter referred to as a cellular ancestry signature CA-110 (Table 1), has a more than two-fold increase in the hazard ratio (HR=8.5) in the TCGA breast cancer database. In contrast, the combination of commercially available standard prognostic signatures Mammaprint (MAM70, HR=3.4) and Prosigna (PAM-50, HR=2.3) is not additive; MAM70+PAM-50 HR=3.7 (Figure 14). Importantly, the CA-110 signature is highly prognostic in colon cancer (HR=16.4), Lung cancer (HR=25.7), prostate cancer (HR=24.1), and pancreatic cancer (HR=1.7xlO9) (Figure 15).
[0158] It was also found that the combination of ET60 (HR=6.9) and THR70 (HR=10.7), hereafter referred to as CA-130 (Table 1), is also additive (HR=18.3) in breast cancer (Figure 16). The CA-130 signature is also highly prognostic in breast cancer (Figure 16), and able to cluster the TGCA PanCancer dataset according to tissue origin (Figure 17). Remarkably, CALIO signature is highly prognostic in multiple other cancers including kidney cancer (HR=4.0), cervix cancer (HR=2.6x 109), melanoma (HR=6.2), ovarian cancer (HR=4.2), liver cancer (HR=9.0), glioma (HR=13.4), colon cancer (HR=19.9), pancreas cancer (HR=19.4), bladder cancer (HR=4.8), lung cancer (HR=2.2), head & neck cancer (HR=4.7), multiple myeloma (HR=7.2) (Figure 19). In some aspects, the prognostic signatures produce a hazard ratio spread of 2-fold to 3-fold. Thus, the range of hazard ratios that are achieved by CA-110 and CA-130 are truly remarkable.
[0159] Discussion
[0160] Species taxonomy uses one feature for each hierarchical level to place each organism in a taxon in an ancestral branch. Suggested herein is classifying a tumor initially into organ (phylum), tissue (class), cell-type (order), and differentiation state (genus) taxons. After which genetic and epigenetic alterations describe the species within that genus. The microenvironment and immunity define the ecosystem of the tumor species. In this approach, one may classify an invasive ductal carcinoma based on breast (phylum), lobule (class), duct (class) epithelium (order), and triple-hormone receptor (genus) taxons. Once the tumor is placed in this ancestral hierarchy, one can use genetic alterations such as HER2 and immune signature to define the species and its ecosystem, which is what was done herein by using the THR70 signature first, followed by i20 immune signature and HER2 amplification information. The modular approach makes it possible to add other tumor hallmarks such as angiogenesis, proliferation, and invasion signatures in this classification.
[0161] In contrast with a standard classification paradigm, breast tumors are currently categorized by hybrid taxons, in which some tumors are defined by cellular ancestry (ER, PR) and others by mutations (HER2) (51). Likewise, molecular signatures Oncotype DX (24) and Prosignia (PAM-50) (25) are a hybrid taxons; a blend of cell-type markers (ER, PR, KRT5, KRT14 and KRT17) and mutations (MYC, EGRF, GRB7, MDM2, FGFR, and HER2). In addition, the prognostic signatures such as MammaPrint (22,52) and PAM-50 contain signals from the entire tumor including vascular, blood, nerve, inflammatory, adipose, and fibroblastic cells. These composite signatures are useful in clinical practice but challenging to deconstruct retroactively and attach biological meaning (53), and they have not evolved with the increasing understanding of breast cancer pathogenesis (54).
[0162] Notably, none of these clinical, prognostic, or experimental subtypes fully incorporate the cellular ancestry information which is critical to the clinical course of many cancer types including breast cancer (55,56). It was found that breast carcinomas can be divided into four groups based on the co-expression of three cell type-specific hormone receptors AR, ER, and VDR. As shown herein, the patients with the THR3 subtype (AR+ER+VDR+) have significantly better overall and relapse-free survival compared to those with the THRO subtype (AR-ER-VDR-) (3,4). It was confirmed that breast cancer cellular ancestry signature correlates with breast cancer metastasis (57) and drug response (19).
[0163] One of the key strengths of this discovery is the utilization of large and well- characterized cohorts of breast cancer patients, which allowed one to validate the prognostic performance of the THR signatures across various clinical and molecular subgroups. Furthermore, by comparing the performance of these signatures with existing classifiers, one could demonstrate their significant value in breast cancer prognostication. The results show that the THR signatures could be integrated into current clinical decision-making algorithms to refine risk stratification and guide personalized treatment strategies for breast cancer patients.
[0164] A recent analysis titled “Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer” convincingly illustrated the importance of cell- of-origin (58). However, since these cell-specific patterns can involve over thirty percent of the entire epigenome (59,60), these observations have been difficult to translate into mechanistic insights (61-70). This study provides an example where these genome-wide patterns can be translated into practical signatures that predict overall survival and recurrence- free survival in breast cancer patients from different cohorts. By incorporating cellular ancestry information and triple hormone receptor status, the THR-50 and 70 signatures provide a more comprehensive and clinically meaningful prognostic model for breast cancer.
[0165] Materials and Methods
[0166] Data Collection and Inclusion Criteria
[0167] Cell lines: To develop the gene signature, gene expression profiles were used from breast cancer cell lines sourced from the Cancer Cell Line Encyclopedia (CCLE) (28) including those positive for a single hormone receptor or none with THR-0&1 phenotype (BT-20, HCC1187, HCC1937, HCC1143, and MDA-MB-231) together with cell lines positive for two or three receptors with a THR-2&3 phenotype (MCF7, T47D, CAMA-1, YMB-1, and ZR-75- 1).
[0168] Patient Data: Gene expression data was incorporated from multiple cohorts comprising samples from breast cancer patients with available survival information. These cohorts included the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC) dataset (29) which consists of samples from 1,904 patients with a mean age of 61.1 years at diagnosis. Additional cohorts were sourced from the KM plotter database (71) encompassing 2,032 samples from 50 different studies, the Meta- 10 cohort which includes 1,888 samples from 10 studies (72), and the BC855 cohort including 855 samples from four datasets (22,38-40).
[0169] Development of the Triple-Hormone Receptors Gene Signatures
[0170] The CellExpress platform (73) was used to perform differential expression analysis between the THR0&1 and THR2&3 breast cancer cell lines in the CCLE dataset (28). Specifically, the “Gene Signature Explore” function was used with the following parameters: platform: Affymetrix U 133plus2 platform, use all the genes with a p-value cutoff of 0.01 , group 1 : CCLE (GSE36133) with breast cell lines HCC1937, MDA-MB231, HCC1187, HCC1143, group 2: CCLE (GSE36133) with breast cell lines YMB-1, MCF7, ZR-75-1, T47D, CAMA-1. The differential expression analysis was performed using a t-test, and the top 300 differentially expressed genes (DEGs) were ranked based on their p-value. The THR-50 signature comprised the top 50 most significant (lowest p-value) DEGs between the THR0&1 and THR2&3 cell lines.
[0171] Furthermore, the gene expression profiles of 855 human breast cancer samples from the BC855 cohort (22,38-40) was used to derive a second THR-based gene signature. Specifically, samples were ranked based on the expression of ESRI, AR, and VDR. Samples with low expression for each of ESRI , AR, and VDR were classified as THRO, those with high expression for one receptor as THR1, those with high expression for two receptors as THR2, and those with high expression for ER, AR, and VDR were labeled THR3. Similar to the analysis performed on the CCLE dataset, we identified the top DEGs between THR0&1 and THR2&3. The top 70 DEGs common between the CCLE dataset and the human cohort were used to develop the THR-70 signature.
[0172] Derivation of 20 gene immune signature (i-20)
[0173] A differential expression analysis was conducted comparing patients within the PQNBC cluster from the THR-70 dataset in the METABRIC cohort. This analysis focused on contrasting patients with an overall survival (OS) time of 50 months or more against those with an OS of less than 50 months. Next, the top 20 differentially expressed genes (DEGs) were identified, which were then utilized to develop a logistic regression model, with the binary OS event (0 and 1) as the response variable and the expression levels of these 20 genes as covariates. The prediction probabilities were further characterized from this model into two groups, termed PQNBC-i+ and PQNBC-i-, based on a threshold derived from the Receiver Operating Characteristics (ROC) Curve analysis. To explore the potential biological function of these 20 genes, Gene Set Enrichment Analysis (GSEA) was employed using the Molecular Signatures Database (MSigDB). By analyzing 18,542 gene sets in MsigDB, we found that 16 of the top 20 gene sets that strongly correlated with our 20-gene signature are linked to immune cell signatures. Based on this, we named this signature as 'immune-20', or i20 for brevity. The specific genes constituting i20 are detailed in Table 1.
[0174] Survival Analysis
[0175] To develop a robust predictive and prognostic signature using patient data, a list of genes in THR-50 and THR-70 signatures were used in the METABRIC dataset to fit a logistic regression model with overall survival (OS) as the outcome variable. The OS probability scores for every patient, as predicted by the THR models, were then categorized into binary labels of 0 and 1 using the optimal threshold from the Receiver Operating Characteristics (ROC) Curve.
[0176] Survival analysis was employed to investigate the association between the THR signatures and survival across different patient cohorts. Within the METABRIC dataset, the probability scores were computed to divide the samples into two groups (median cutoff) and four groups (quartiles): Q1:Q4 where QI and Q4 samples had the lowest and highest probability of death, respectively. Then the Kaplan-Meier (KM) survival estimate was used with a logrank test to compare OS and RFS between the different groups. In the KM plotter breast cancer dataset (74), survival analysis was performed by computing the average expression of the THR signatures genes, then using median and quartiles cutoffs to divide the patients into low- and high-risk groups. The same approach was employed for comparing the THR-50 signature to the PAM-50, MammaPrint, and Oncotype DX assays.
[0177] The Surv Express analysis was carried out by selecting as previously described (57); with two maximized risk groups optimized using an algorithm that decides where the partitions should be made to maximize the statistical significance of the separation of risk groups. The SurvExpress algorithm tests different cut-off points to partition risk groups with minimum p- value. The relative hazard is computed using Cox proportional hazard regression analysis and the p values are computed using Log Rank test (72). The Kaplan-Meier Plotter analysis was carried out as described before (57); by various approaches to define comparison cohorts by trichotomizing (QI vs Q4) which involves assigning the data into four cohorts and omitting the middle two cohorts, or by using the best available cut-off value where each possible cutoff value is tested between QI and Q4. The Benjamini -Hochberg method is used to correct for multiple hypothesis testing (74).
[0178] Gene Set Enrichment Analysis
[0179] The immune module of Gene Set Cancer Analysis (GSCA) online platform (75) was used to determine the correlation between immune cell infiltrates and GSVA enrichment score, using ImmuCellAI (Immune Cell Abundance Identifier) that estimates the abundance of 24 immune cells from gene expression dataset including RNA-Seq and microarray data, comprised of 18 T-cell subtypes and 6 other immune cells: B cell, NK cell, Monocyte cell, Macrophage cell, Neutrophil cell, and DC cell (76). The expression module of GSCA was used to summarize the association between GSVA score and activity of cancer-related pathways in breast cancer. The GSCA analyzes reverse phase protein array data to calculate the pathway activity score of 10 cancer-related pathways in 7876 samples (TCGA database) including TSC / mTOR, RTK, RAS / MAPK, PI3K / AKT, Hormone ER, Hormone AR, EMT, DNA Damage Response, Cell Cycle, and Apoptosis pathways. The GSVA score represents the variation of gene set activity over a sample population in an unsupervised manner and is represented in a heatmap and table with Spearman cor., P value, and FDR (75).
[0180] Unsupervised Clustering of Patient Samples
[0181] To explore the potential of the THR signatures in identifying distinct breast cancer subtypes, unsupervised hierarchical clustering of the samples within the METABRIC dataset was performed. Hierarchical clustering was performed on all samples using the ward minimum variance method (77), then the hierarchical tree was cut into 5 groups. The optimal number of groups was determined based on the observed patterns in the data and the clinical interpretability of the results. The overall and recurrence-free survival probabilities of the 5 groups were compared using Kaplan-Meier survival curves and logrank test.
[0182] Software and Statistical Analysis
[0183] All statistical analyses were performed using R software (version 4.0.3). The survival and survminer R packages were used for generating Kaplan-Meier survival curves and the COX proportional hazards models (78,79). The stats package was used for hierarchical clustering while the glmnet package was used to fit the logistic regression models (80). The significance level (p- value and false discovery rate) was set at 0.05 for all statistical tests.
[0184] Bibliography 1. Khoury JD, Solary E, Abla O, Akkari Y, Alaggio R, Apperley JF, et al. The 5th edition of the World Health Organization Classification of Haematolymphoid Tumours: Myeloid and Histiocytic / Dendritic Neoplasms. Leukemia. 2022 Jul;36(7): 1703-19.
[0185] 2. Alaggio R, Amador C, Anagnostopoulos I, Attygalle AD, Araujo IB de O, Berti E, et al. The 5th edition of the World Health Organization Classification of Haematolymphoid Tumours: Lymphoid Neoplasms. Leukemia. 2022 Jul;36(7): 1720-48.
[0186] 3. Santagata S, Thakkar A, Ergonul A, Wang B, Woo T, Hu R, et al. Taxonomy of breast cancer based on normal cell phenotype predicts outcome. J Clin Invest. 2014 Feb 3;124(2):859— 70.
[0187] 4. Santagata S, Ince TA. Normal cell phenotypes of breast epithelial cells provide the foundation of a breast cancer taxonomy. Expert Review of Anticancer Therapy. 2014 Dec 1 ;14(12): 1385— 9.
[0188] 5. Houseman EA, Kile ML, Christiani DC, Ince TA, Kelsey KT, Marsit CJ. Reference- free deconvolution of DNA methylation data and mediation by cell composition effects. BMC Bioinformatics. 2016 Jun 29; 17(1):259.
[0189] 6. Houseman EA, Ince TA. Normal Cell-Type Epigenetics and Breast Cancer Classification: A Case Study of Cell Mixture- Adjusted Analysis of DNA Methylation Data from Tumors. Cancer Inform. 2014 Dec 9;13(Suppl 4):53— 64.
[0190] 7. Dontu G, Ince TA. Of Mice and Women: A Comparative Tissue Biology Perspective of Breast Stem Cells and Differentiation. J Mammary Gland Biol Neoplasia. 2015 ;20( 1— 2):51— 62.
[0191] 8. Allison KH, Hammond MEH, Dowsett M, Me Kernin SE, Carey LA, Fitzgibbons PL, et al. Estrogen and Progesterone Receptor Testing in Breast Cancer: ASCO / CAP Guideline Update. JCO. 2020 Apr 20;38( 12): 1346-66.
[0192] 9. Freake HC, Abeyasekera G, Iwasaki J, Marcocci C, MacIntyre I, McClelland RA, et al. Measurement of 1,25 -dihydroxy vitamin D3 receptors in breast cancer and their relationship to biochemical and clinical indices. Cancer Res. 1984 Apr;44(4): 1677-81.
[0193] 10. Al-Azhri J, Zhang Y, Bshara W, Zirpoli G, McCann SE, Khoury T, et al. Tumor Expression of Vitamin D Receptor and Breast Cancer Histopathological Characteristics and Prognosis. Clin Cancer Res. 2017 Jan 1 ;23(1):97- 103.
[0194] 11. Huss L, Butt ST, Borgquist S, Elebro K, Sandsveden M, Rosendahl A, et al. Vitamin D receptor expression in invasive breast tumors and breast cancer survival. Breast Cancer Research. 2019 Jul 29;21(1):84. 12. Huss L, Butt ST, Borgquist S, Elebro K, Sandsveden M, Manjer J, et al. Levels of Vitamin D and Expression of the Vitamin D Receptor in Relation to Breast Cancer Risk and Survival. Nutrients. 2022 Aug 16; 14(16):3353.
[0195] 13. Narayanan R, Dalton JT. Androgen Receptor: A Complex Therapeutic Target for Breast Cancer. Cancers (Basel). 2016 Dec 2;8(12): 108.
[0196] 14. Ricciardelli C, Bianco-Miotto T, Jindal S, Butler LM, Leung S, McNeil CM, et al. The Magnitude of Androgen Receptor Positivity in Breast Cancer Is Critical for Reliable Prediction of Disease Outcome. Clinical Cancer Research. 2018 May 14;24(10):2328-41.
[0197] 15. Wang C, Pan B, Zhu H, Zhou Y, Mao F, Lin Y, et al. Prognostic value of androgen receptor in triple negative breast cancer: A meta-analysis. Oncotarget. 2016 Jun 21;7(29):46482-91.
[0198] 16. Kensler KH, Poole EM, Heng YJ, Collins LC, Glass B, Beck AH, et al. Androgen Receptor Expression and Breast Cancer Survival: Results From the Nurses’ Health Studies. JNCI: Journal of the National Cancer Institute. 2019 Jul 1 ; 111(7) :700— 8.
[0199] 17. Hu R, Dawood S, Holmes MD, Collins LC, Schnitt SJ, Cole K, et al. Androgen receptor expression and breast cancer survival in postmenopausal women. Clin Cancer Res. 2011 Apr 1; 17(7): 1867-74.
[0200] 18. Kensler KH, Regan MM, Heng YJ, Baker GM, Pyle ME, Schnitt SJ, et al. Prognostic and predictive value of androgen receptor expression in postmenopausal women with estrogen receptor-positive breast cancer: results from the Breast International Group Trial 1- 98. Breast Cancer Research. 2019 Feb 22;21(l):30.
[0201] 19. Bhattacharya U, Kamran M, Manai M, Cristofanilli M, Ince TA. CelLof-Origin Targeted Drug Repurposing for Triple-Negative and Inflammatory Breast Carcinoma with HDAC and HSP90 Inhibitors Combined with Niclosamide. Cancers (Basel). 2023 Jan 4;15(2):332.
[0202] 20. Cardoso F, Paluch-Shimon S, Senkus E, Curigliano G, Aapro MS, Andre F, et al. 5th ESO-ESMO international consensus guidelines for advanced breast cancer (ABC 5). Ann Oncol. 2020 Dec;31(12):1623-49.
[0203] 21. Robertson JohnFR, Paridaens RJ, Lichfield J, Bradbury I, Campbell C. Meta-analyses of phase 3 randomised controlled trials of third generation aromatase inhibitors versus tamoxifen as first-line endocrine therapy in postmenopausal women with hormone receptorpositive advanced breast cancer. European Journal of Cancer. 2021 Mar 1; 145: 19— 28. 22. van de Vijver MJ, He YD, van ’t Veer LJ, Dai H, Hart AAM, Voskuil DW, et al. A Gene-Expression Signature as a Predictor of Survival in Breast Cancer. New England Journal of Medicine. 2002 Dec 19;347(25): 1999-2009.
[0204] 23. Buyse M, Loi S, van’t Veer L, Viale G, Delorenzi M, Gias AM, et al. Validation and Clinical Utility of a 70-Gene Prognostic Signature for Women With Node-Negative Breast Cancer. JNCI: Journal of the National Cancer Institute. 2006 Sep 6;98( 17): 1183— 92.
[0205] 24. Paik S, Shak S, Tang G, Kim C, Baker J, Cronin M, et al. A Multigene Assay to Predict Recurrence of Tamoxifen-Treated, Node-Negative Breast Cancer. N Engl J Med. 2004 Dec 30;351(27):2817-26.
[0206] 25. Parker JS, Mullins M, Cheang MCU, Leung S, Voduc D, Vickery T, et al. Supervised Risk Predictor of Breast Cancer Based on Intrinsic Subtypes. J Clin Oncol. 2009 Mar 10;27(8): 1160-7.
[0207] 26. Harris LN, Ismaila N, McShane LM, Andre F, Collyar DE, Gonzalez-Angulo AM, et al. Use of Biomarkers to Guide Decisions on Adjuvant Systemic Therapy for Women With Early-Stage Invasive Breast Cancer: American Society of Clinical Oncology Clinical Practice Guideline. J Clin Oncol. 2016 Apr 1;34( 10): 1134-50.
[0208] 27. Duffy MJ, Harbeck N, Nap M, Molina R, Nicolini A, Senkus E, et al. Clinical use of biomarkers in breast cancer: Updated guidelines from the European Group on Tumor Markers (EGTM). European Journal of Cancer. 2017 Apr l;75:284-98.
[0209] 28. Barretina J, Caponigro G, Stransky N, Venkatesan K, Margolin AA, Kim S, et al. The Cancer Cell Line Encyclopedia enables predictive modelling of anticancer drug sensitivity. Nature. 2012 Mar;483(7391):603-7.
[0210] 29. Curtis C, Shah SP, Chin SF, Turashvili G, Rueda OM, Dunning MJ, et al. The genomic and transcriptomic architecture of 2,000 breast tumours reveals novel subgroups. Nature. 2012 Jun;486(7403):346-52.
[0211] 30. The Way of the Future: Personalizing Treatment Plans Through Technology I American Society of Clinical Oncology Educational Book [Internet], [cited 2023 Jun 27]. Available from: https: / / ascopubs.org / doi / 10.1200 / EDBK_320593?url_ver=Z39.88- 2003&rfr_id=ori:rid:crossref.org&rfr_dat=cr_pub%20%200pubmed
[0212] 31. Vieira AF, Schmitt F. An Update on Breast Cancer Multigene Prognostic Tests — Emergent Clinical Biomarkers. Frontiers in Medicine [Internet]. 2018 [cited 2023 Apr 5];5. Available from : https : / / w w w. frontiersin. org / articles / 10.3389 / fmed.2018.00248 32. Barbi M, Makower D, Sparano JA. The clinical utility of gene expression assays in breast cancer patients with 0-3 involved lymph nodes. Ther Adv Med Oncol. 2021 Aug 14; 13: 17588359211038468.
[0213] 33. Jacob L, Witteveen A, Beumer I, Delahaye L, Wehkamp D, van den Akker J, et al. Controlling technical variation amongst 6693 patient microarrays of the randomized MIND ACT trial. Commun Biol. 2020 Jul 27 ;3:397.
[0214] 34. Piccart M, Veer LJ van ’t, Poncet C, Cardozo JMNL, Delaloge S, Pierga JY, et al. 70- gene signature as an aid for treatment decisions in early breast cancer: updated results of the phase 3 randomised MIND ACT trial with an exploratory analysis by age. The Lancet Oncology. 2021 Apr l;22(4):476-88.
[0215] 35. Varga Z, Sinn P, Fritzsche F, Hochstetter A von, Noske A, Schraml P, et al. Comparison of EndoPredict and Oncotype DX Test Results in Hormone Receptor Positive Invasive Breast Cancer. PLOS ONE. 2013 Mar 7;8(3):e58483.
[0216] 36. Bosl A, Spitzmiiller A, Jasarevic Z, Rauch S, Jager S, Offner F. MammaPrint versus EndoPredict: Poor correlation in disease recurrence risk classification of hormone receptor positive breast cancer. PLOS ONE. 2017 Aug 29;12(8):e0183458.
[0217] 37. Sestak I, Buus R, Cuzick J, Dubsky P, Kronenwett R, Denkert C, et al. Comparison of the Performance of 6 Prognostic Signatures for Estrogen Receptor-Positive Breast Cancer: A Secondary Analysis of a Randomized Clinical Trial. JAMA Oncology. 2018 Apr 1;4(4):545— 53.
[0218] 38. Wang Y, Klijn JG, Zhang Y, Sieuwerts AM, Look MP, Yang F, et al. Geneexpression profiles to predict distant metastasis of lymph-node-negative primary breast cancer. The Lancet. 2005 Feb 19;365(9460):671-9.
[0219] 39. Bos PD, Zhang XHF, Nadal C, Shu W, Gomis RR, Nguyen DX, et al. Genes that mediate breast cancer metastasis to the brain. Nature. 2009 Jun;459(7249): 1005-9.
[0220] 40. Minn AJ, Gupta GP, Siegel PM, Bos PD, Shu W, Giri DD, et al. Genes that mediate breast cancer metastasis to lung. Nature. 2005 Jul 28;436(7050):518-24.
[0221] 41. Sole X, Bonifaci N, Lopez-Bigas N, Berenguer A, Hernandez P, Reina O, et al. Biological Convergence of Cancer Signatures. PLOS ONE. 2009 Feb 20;4(2):e4544.
[0222] 42. Venet D, Dumont JE, Detours V. Most Random Gene Expression Signatures Are Significantly Associated with Breast Cancer Outcome. PLOS Computational Biology. 2011 Oct 20;7(10):el002240. 43. Nunes AT, Collyar DE, Harris LN. Gene Expression Assays for Early-Stage Hormone Receptor-Positive Breast Cancer: Understanding the Differences. JNCI Cancer Spectr. 2017 Dec l l;l(l):pkx008.
[0223] 44. Byrne A, Savas P, Sant S, Li R, Virassamy B, Luen SJ, et al. Tissue-resident memory T cells in breast cancer control and immunotherapy responses. Nat Rev Clin Oncol. 2020 Jun;17(6):341-8.
[0224] 45. Sprlie T, Perou CM, Tibshirani R, Aas T, Geisler S, Johnsen H, et al. Gene expression patterns of breast carcinomas distinguish tumor subclasses with clinical implications. Proceedings of the National Academy of Sciences. 2001 Sep 11 ;98(19): 10869-74.
[0225] 46. Perou CM. Molecular Stratification of Triple-Negative Breast Cancers. The Oncologist. 2010 Nov 1 ; 15(S5):39— 48.
[0226] 47. Prat A, Parker JS, Karginova O, Fan C, Livasy C, Herschkowitz JI, et al. Phenotypic and molecular characterization of the claudin-low intrinsic subtype of breast cancer. Breast Cancer Res. 2010;12(5):R68.
[0227] 48. Milioli HH, Vimieiro R, Tishchenko I, Riveros C, Berretta R, Moscato P. Iteratively refining breast cancer intrinsic subtypes in the METABRIC dataset. BioData Mining. 2016 Jan 13;9(1):2.
[0228] 49. Jiang YZ, Ma D, Suo C, Shi J, Xue M, Hu X, et al. Genomic and Transcriptomic Landscape of Triple-Negative Breast Cancers: Subtypes and Treatment Strategies. Cancer Cell. 2019 Mar 18;35(3):428-440.e5.
[0229] 50. Vallon-Christersson J, Hiikkinen J, Hegardt C, Saal LH, Larsson C, Ehinger A, et al. Cross comparison and prognostic assessment of breast cancer multigene signatures in a large population-based contemporary clinical series. Sci Rep. 2019 Aug 21;9(1): 12184.
[0230] 51. Johansson ALV, Trewin CB, Fredriksson I, Reinertsen KV, Russnes H, Ursin G. In modern times, how important are breast cancer stage, grade and receptor subtype for survival: a population-based cohort study. Breast Cancer Res. 2021;23:17.
[0231] 52. van ’t Veer LJ, Dai H, van de Vijver MJ, He YD, Hart AAM, Mao M, et al. Gene expression profiling predicts clinical outcome of breast cancer. Nature. 2002 Jan;415(6871):530-6.
[0232] 53. Emmert-Streib F, Manjang K, Dehmer M, Yli-Harja O, Auvinen A. Are There Limits in Explainability of Prognostic Biomarkers? Scrutinizing Biological Utility of Established Signatures. Cancers. 2021 Jan;13(20):5087. 54. Manjang K, Tripathi S, Yli-Harja O, Dehmer M, Glazko G, Emmert-Streib F. Prognostic gene expression signatures of breast cancer are lacking a sensible biological meaning. Sci Rep. 2021 Jan 8; 11(1): 156.
[0233] 55. Rycaj K, Tang DG. Cell-of-Origin of Cancer versus Cancer Stem Cells: Assays and Interpretations. Cancer Research. 2015 Sep 30;75(19):4003-l l.
[0234] 56. Skibinski A, Kuperwasser C. The origin of breast tumor heterogeneity. Oncogene. 2015 Oct 16;34(42):5309— 16.
[0235] 57. Kamran M, Bhattacharya U, Omar M, Marchionni L, Ince TA. ZNF92, an unexplored transcription factor with remarkably distinct breast cancer over-expression associated with prognosis and cell-of-origin. npj Breast Cancer. 2022 Aug 29;8(1): 1— 11.
[0236] 58. Hoadley KA, Yau C, Hinoue T, Wolf DM, Lazar AJ, Drill E, et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell. 2018 Apr 5;173(2):291-304.e6.
[0237] 59. Mancarella D, Plass C. Epigenetic signatures in cancer: proper controls, current challenges and the potential for clinical translation. Genome Medicine. 2021 Feb 10; 13( 1) :23.
[0238] 60. Hawkins RD, Hon GC, Ren B. Next-generation genomics: an integrative approach. Nat Rev Genet. 2010 Jul;l l(7):476-86.
[0239] 61. Xin L. Cells of origin for cancer: an updated view from prostate cancer. Oncogene. 2013 Aug;32(32):3655-63.
[0240] 62. Merritt MA, Bentink S, Schwede M, Iwanicki MP, Quackenbush J, Woo T, et al. Gene Expression Signature of Normal Cell-of-Origin Predicts Ovarian Tumor Outcomes. PLOS ONE. 2013 Nov 26;8(1 l):e80314.
[0241] 63. Bhagirath D, Zhao X, West WW, Qiu F, Band H, Band V. Cell type of origin as well as genetic alterations contribute to breast cancer phenotypes. Oncotarget. 2015 ;6(11):9018— 30.
[0242] 64. Bu W, Liu Z, Jiang W, Nagi C, Huang S, Edwards DP, et al. Mammary Precancerous Stem and Non-Stem Cells Evolve into Cancers of Distinct Subtypes. Cancer Research. 2019 Jan 2;79(1):61-71.
[0243] 65. Kwon S, Kim SS, Nebeck HE, Ahn EH. Immortalization of Different Breast Epithelial Cell Types Results in Distinct Mitochondrial Mutagenesis. International Journal of Molecular Sciences. 2019 Jan;20(l l):2813.
[0244] 66. Ferone G, Lee MC, Sage J, Berns A. Cells of origin of lung cancers: lessons from mouse studies. Genes Dev. 2020 Jan 8;34( 15— 16): 1017— 32. 67. Kim HJ, Park JW, Lee JH. Genetic Architectures and Cell-of-Origin in Glioblastoma. Frontiers in Oncology [Internet]. 2021 [cited 2023 Jun 27]; 10. Available from: https : / / www. frontiersin. org / articles / 10.3389 / fonc .2020.615400
[0245] 68. Flowers BM, Xu H, Mulligan AS, Hanson KJ, Seoane JA, Vogel H, et al. Cell of Origin Influences Pancreatic Cancer Subtype. Cancer Discovery. 2021 Mar 2;l l(3):660-77.
[0246] 69. Geboes K, Hoorens A. The cell of origin for Barrett’s esophagus. Science. 2021 Aug 13;373(6556):737— 8.
[0247] 70. Moeini A, Haber PK, Sia D. Cell of origin in biliary tract cancers and clinical implications. JHEPReport [Internet]. 2021 Apr 1 [cited 2023 Jun 27];3(2). Available from: https : / / www.jhep-reports.eu / article / S2589-5559(21)00002-l / fulltext
[0248] 71. Gyorffy B. Survival analysis across the entire transcriptome identifies biomarkers with the highest prognostic power in breast cancer. Computational and Structural Biotechnology Journal. 2021 Jan 1 ; 19:4101— 9.
[0249] 72. Aguirre-Gamboa R, Gomez-Rueda H, Martinez-Ledesma E, Martmez-Torteya A, Chacolla-Huaringa R, Rodriguez-Barrientos A, et al. SurvExpress: An Online Biomarker Validation Tool and Database for Cancer Gene Expression Data Using Survival Analysis. PLoS One. 2013 Sep 16;8(9):e74250.
[0250] 73. Lee YF, Lee CY, Lai LC, Tsai MH, Lu TP, Chuang EY. CellExpress: a comprehensive microarray -based cancer cell line and clinical sample gene expression analysis online system. Database. 2018 Jan l;2018:baxl01.
[0251] 74. Lanczky A, Gyorffy B. Web-Based Survival Analysis Tool Tailored for Medical Research (KMplot): Development and Implementation. J Med Internet Res. 2021 Jul 26;23(7):e27633.
[0252] 75. Liu CJ, Hu FF, Xia MX, Han L, Zhang Q, Guo AY. GSCALite: a web server for gene set cancer analysis. Bioinformatics. 2018 Nov 1;34(21):3771— 2.
[0253] 76. Miao YR, Zhang Q, Lei Q, Luo M, Xie GY, Wang H, et al. ImmuCellAI: A Unique Method for Comprehensive T-Cell Subsets Abundance Prediction and its Application in Cancer Immunotherapy. Advanced Science. 2020;7(7): 1902880.
[0254] 77. Ward JH. Hierarchical Grouping to Optimize an Objective Function. Journal of the American Statistical Association. 1963 Mar l;58(301):236- 44.
[0255] 78. Therneau TM, Grambsch PM. Modeling Survival Data: Extending the Cox Model [Internet]. New York, NY: Springer; 2000 [cited 2023 Feb 20]. (Dietz K, Gail M, Krickeberg K, Samet J, Tsiatis A, editors. Statistics for Biology and Health). Available from: http : / / link, springer, com / 10.1007 / 978-1-4757-3294-8 79. Cox DR. Regression Models and Life-Tables. Journal of the Royal Statistical Society Series B (Methodological). 1972;34(2): 187-220.
[0256] 80. Friedman JH, Hastie T, Tibshirani R. Regularization Paths for Generalized Linear Models via Coordinate Descent. Journal of Statistical Software. 2010 Feb 2;33:l-22.
[0257] Example II
[0258] Random Genes Can Substitute for Prognostic Oncogene Signatures, but a Not Concise Cell- of-Origin Signature
[0259] Introduction
[0260] Predicting the behavior of complex systems, whether physical or biological, requires an understanding of their initial conditions. For instance, ecological studies have demonstrated that population dynamics are highly sensitive to these initial conditions, where even small changes can lead to significantly different population trajectories over time (1). It was hypothesized that the initial conditions of cancer are bounded by the normal cell-of-origin, which could be equally important in determining the trajectory and cellular population dynamics of a tumor (2). It has been also recognized that using too many parameters to define complex systems introduces excess uncertainty (3). Based on these first principles, breast cancer cell-of-origin was defined with the fewest possible parameters. This hypothesis-driven approach differs from genome- wide exploratory methods that are increasingly prevalent in biomarker studies, striving to integrate a growing number of parameters, rather than reducing them, often without a specific hypothesis.
[0261] It was previously reported that breast cancer cell-of-origin determines metastatic potential and drug response of breast cancer in experimental models (4-6). Next, it was found that breast- specific triple-hormone receptor (THR) co-expression of estrogen (ER), androgen (AR) and vitamin-D (VDR) receptors define putative cells-of-origin with distinct epigenetic DNA methylation profiles (7,8). Finally, the THR cell-of-origin signature comprised of thousands of genes was reduced to a seventy-mRNA signature (THR-70), that proved to be prognostic in all breast cancer subtypes (Example I).
[0262] In a translational study, the THR-70 signature was validated through outcome analysis of 6,679 patients across 65 publicly available datasets. It was also shown that THR-70 is enriched in normal breast epithelium using a single-cell transcriptomics analysis. While this seventy-gene signature performed well as a prognostic biomarker, subsequent investigations suggested that well-established breast cancer prognostic signatures exceeding twenty transcripts in size become increasingly indistinguishable from random signatures. This discovery further motivated the identification of a more concise signature for predicting breast cancer prognosis. A small, unmutated cell-of-origin signature, composed of just six transcripts was identified, that outperforms larger prognostic signatures, even those with potent driver mutations.
[0263] THR-6E Signature
[0264] While the THR-70 signature performs well in identifying breast cancer subgroups with different outcomes, identifying subsets of this signature with fewer genes can help reduce this invention to practice. To reduce the size of the THR-70 cell-of-origin signature, unsupervised hierarchical clustering was used, which revealed a six-transcript signature (THR-6E) with variable expression across the three ER-positive breast cancer clusters El, E2, and E3 compared to other THR-70 sub-signatures (Fig. 29A). Signature size reduction does not decrease prognostic power in this instance; the THR-6E relative hazard ratio (RHR) of relapse- free survival (RFS) is greater than THR-70, RHR =1.8, CI 1.7-2. 1, p<le-16 and RHR =1.7, CI 1.4-2.0, p=2e- 11 respectively, among all breast cancer regardless of subtype in the KMP dataset (Fig. 29B-C).
[0265] Several mRNA transcript-based prognostic signatures are utilized to guide the treatment of breast cancer, such as the 21-gene Oncotype DX, 50-gene PAM-50, and 70-gene MammaPrint. The commercial versions of these tests are used in ER-positive (ERp), HER2- negative (HER2n), and lymph node-negative (LNn) breast cancers. They are not recommended in ER-negative (ERn) and HER2 -positive (HER2p) breast cancers (9). It was found that THROE is highly prognostic in ERp-HER2n-LNn breast cancers (RHR =2.2, CI 1.7-2.7, p=4.8e-l 1) (Fig. 29D).
[0266] Comparison of THR-6E to well-established signatures in ERp-HER2n-LNn breast cancers shows that while all appear to be significant based on the standard significance threshold a=0.05, there is a significant difference compared to random sequences of equal size (Fig. 29E-G). Among the twenty random sequences with six genes (Random-6, or R-6), only four achieved a p-value less than 0.05. However, it is worth mentioning that there is an eightorder magnitude difference in the p-value of these four random signatures (4.0e-2 to 3.7e-3) versus THR-6E (p=2. le-11). In contrast, the p values of eight R-21 and eleven R-50 random sequences are lower than Oncotype DX and PAM-50 respectively, and below the 0.05 threshold (Fig. 29E-G). Thus, in this comparison, Oncotype DX (21-genes) and PAM-50 (50- genes) are indistinguishable from nearly half of the random signatures of equal size, R-21, and R-50 respectively. Consistently with this, others have found that the MammaPrint classifier based on the top 1-70 genes is no better than lower-ranked classifiers such as 71-140, 141-210, and so on until 701-770 (10). These results suggest that the standard significance threshold of a=0.05 may need to be adjusted with increasing signature size.
[0267] Having confirmed THR-6E’s prognostic significance in breast cancer compared to random controls, as well as established signatures like Oncotype DX and PAM-50, a more in- depth analysis of its characteristics was undertaken. To verify the breast specificity of the THROE signature, the ProteomicsDB database was examined. The findings reveal that while THROE genes are expressed individually in various tissues, they are co-expressed only in breast tissue, particularly in conjunction with ER, AR, and VDR co-expression (Table 3, Fig. 30A).
[0268] It was previously shown herein, using a single cell transcriptomics analysis, that THR- 70 is enriched in normal breast epithelium (Breast Cancer Res. 2024 Sep 13;20(l):132. doi: 10.1180 / sl3058-024-01870-9. PMID: 39272208). Analysis of published clusters of glandular breast epithelial (GBE) cells in the Human Protein Atlas (HPA) single-cell mRNA dataset unveiled varying levels of THR-0E expression across different GBE subsets. While some normal GBE cells express all six THR-0E transcripts, others express one, two, or three transcripts, suggesting a potential for a variable gradient of expression of THR-6E in breast cancers (Fig. 30B, Table 3).
[0269] As would be expected from a true cell-of-origin signature, it was found that the THROE genes are rarely mutated in breast cancer, on average only 1.1%, in the combined TCGA and METABRIC dataset with 3,593 patients. In contrast, Oncotype DX, PAM-50, and MammaPrint signatures contain multiple potential cancer driver genes that are amplified in 15- 20% of the breast samples (Fig. 31).
[0270] It is shown that both the individual THR-6E genes and the average expression of the entire signature are upregulated in primary and metastatic breast cancers (Fig. 32A-C). Pancancer heatmap analysis demonstrates that THR-6E, ER, AR, and VDR co-expression is highest in breast cancer, which is maintained in metastatic tumors as well (Fig. 32D). These results cumulatively demonstrate that THR-6E remains unmutated in breast cancer, and its coexpression pattern is unique to normal breast glandular epithelial cells, which is maintained in breast cancer, validating it as a putative breast cancer cell-of-origin signature.
[0271] An interactome analysis shows multiple interactions among THR-6E proteins and hormone receptors, indicating activation of Kinesin Family Member 4A and 2C (KIF4A and KIF2C) and Cell Division Cycle Protein 20 (CDC20) by ER and VDR respectively (11,12), activation of AR by KIF4A (13) and Microtubule Nucleation Factor (TPX2 / P 100) (14), and reciprocal activation of Family With Sequence Similarity 64, Member A FAM64A / PIMREG) and AR and ER (15), and interaction of Lamin B2 (EMNB2) with ER (Fig. 33A) (16). Expanding this interactome beyond direct interactions suggests a THR-6E gene network associated with proliferation, differentiation, survival, spindle assembly, microtubule organization, intracellular transport, and Golgi-ER membrane trafficking (Fig. 33B). These findings suggest that THR-6E is likely not a random collection of genes identified by chance, but rather may be a component of the breast specific network associated with hormone regulation.
[0272] Next, the prognostic function of THR-6E was validated in a second dataset. Kaplan- Meier overall survival analysis charts reveal that the high-risk THR-6E group correlates with poor survival in the METABRIC cohort (RHR=1.7, CI: 1.3-2.2, p=9.7E-05). Initially, Oncotype DX, PAM-50, and MammaPrint also seemed to correlate with the outcome in this dataset (Fig. 34A-D). However, when compared to 1,000 random sequences comprising six genes (R-6), once again THR-6E emerges as the sole biomarker that outperforms random controls, based on their p-value ranking (Fig. 34E). Thus, the null hypothesis (HO) that THROE and R-6 exhibit no significant difference in their impact on the outcome variable was rejected. In contrast, Oncotype DX, PAM-50, and MammaPrint were once again indistinguishable from random control sequences matched according to their size, R-21, R-50, and R-70, respectively (Fig. 34F-H). Indeed, Oncotype DX, PAM-50, and MammaPrint ranked approximately 450th, 150th, and 525threspectively, among 1,000 random control signatures with an equal number of genes to each specific signature (Fig. 34F-H). In contrast, THR-6E ranked 2ndout of 1,000 random controls (Fig. 34E). Importantly, the single random signature that outperformed THR-6E in the METABRIC cohort did not correlate with the outcome in another dataset (KMP, not shown). Therefore, it was determined that THR-6E that was developed through hypothesis-based natural-learning, outperforms 1,000 signatures fitted into a survival model by machine-learning.
[0273] To explore whether THR-6E could function as a predictive biomarker, in addition to its role as a prognostic signature, a Receiver Operating Characteristics (ROC) analysis of breast cancer endocrine treatment response was conducted. The findings reveal that high THR-6E expression is significantly associated with a lack of response to endocrine therapy in ERp- HER2n-LNn breast cancers (p=0.035, AUC= 0.611) (Fig. 35A). Among established signatures, only Oncotype DX correlates with hormone treatment response (p=0.041), unlike PAM-50 (p=0.14) and MammaPrint (p=0.4) (Fig. 35 A). This result is in accordance with the American Society of Clinical Oncology (ASCO) guideline that recognizes Oncotype DX as the only test with strong evidence for predicting therapy benefit, whereas PAM-50 and MammaPrint are classified as having intermediate or moderate evidence (9). It was found that THR-6E can stratify lymph node-positive (LNp) breast cancers for chemotherapy sensitivity, measured by the pathological response (p=0.000023, AUC=0.714) (Fig. 35B).
[0274] Additionally, THR-6E demonstrates prognostic value in both LN-positive and LN- negative patients, as well as in untreated, hormone-treated, chemotherapy-treated, and ERnegative cancers (Fig. 36A-B). These findings hold significance because established signatures are not recommended by the American Society of Clinical Oncology (ASCO) for ER-negative cancers or ER-positive cancers with more than three positive nodes (9). Consistent with this recommendation, it was observed that Oncotype DX (p=0.071), PAM-50 (p=0.83), and MammaPrint (p=0.59) do not correlate with chemotherapy response in lymph node-positive patients (Fig. 35B).
[0275] Taken together, the THR-6E signature stratifies breast cancer patients to guide decisions for adjuvant endocrine and chemotherapy. Notably, in a different control, it was observed that THR-6E does not exhibit prognostic value in other cancer types such as ovarian adenocarcinoma, lung squamous cell carcinoma, or leukemia (Fig. 36C). This observation is consistent with the derivation of THR-6E as a breast- specific cell-of-origin signature. It was also observed that the THR and ET signatures are particular to different breast cancer subtypes in both KMP and METABRIC datasets.
[0276] A. Tissue ER-a AR VDR KTF2C KTF4A CDC20 FAM64ATPX2 LMNB2 The above Table shows tissue and cell-specific expression of THR-6E genes. (A) Expression of breast-specific hormone receptors (ER, AR, and VDR) and THR-6E genes of Kinesin Family Member 4A and 2C (KIF4A, and KIF2C), Cell Division Cycle Protein 20 (CDC20), Microtubule Nucleation Factor (TPX2 / P100), Family with Sequence Similarity 64, Member A (FAM64A / PIMREG) and Lamin B2 (LMNB2) in different normal human tissues. THR-4E Signature
[0277] Next, it was determined if one could further reduce the size of this signature to only four genes without compromising the prognostic power. This four-gene signature is referred as THR-4E, which is prognostic only ER-positive breast cancers (Fig. 37). The THR-4E relative hazard ratio (RHR) of relapse- free survival (RFS) of ER-positive breast cancer is 2.8, CI 1.8-2.6, p<le-16, which is greater than THR-70 and THR-6E (Fig. 29 A). Moreover, THR- 4E is not prognostic in ER-negative breast cancer (p=0.19), demonstrating its specific utility in ER-positive but not in ER-negative breast cancer (Fig. 37A-B). To explore whether THR- 4E could function as a predictive biomarker, in addition to its role as a prognostic signature, a Receiver Operating Characteristics (ROC) analysis of breast cancer endocrine treatment response was conducted. The findings reveal that high THR-4E expression is significantly associated with a lack of response to endocrine therapy in ER-positive breast cancers (p=0.00002, AUC=0.615) (Fig. 37E-F), and a lack of response to chemotherapy in ER- positive / HER2-negative breast cancers with lymph node metastasis (p=1.7e-06, AUC=0.767).
[0278] When compared to 1,000 random sequences comprising four genes (R-6), once again THR-4E emerged as the sole biomarker that outperforms all the random controls, based on their p-value ranking (Fig. 38A). In contrast, Endopredict, Oncotype DX, PAM-50, and MammaPrint were once again indistinguishable from random control sequences matched according to their size (Fig. 38B-E). Indeed, Endopredict, Oncotype DX, PAM-50, and MammaPrint ranked l l111, 403rd, 194111’ and 183rdrespectively, among 1,000 random control signatures with an equal number of genes to each specific signature (Fig. 38B-E).
[0279] THR-i8 and i3 signatures
[0280] As described herein (Example I), a twenty-gene immune infiltration signature (THR- i20), can be used to stratify ER-negative breast cancers into outcome groups with significant survival differences. To further reduce this signature into practice, an eight-gene subset (THR- i8) and a three-gene subset (THR-i3) of THR-i20 can be used to stratify double-negative (ERHER2-negative), triple-negative (ER / PR / HER2-negative), quadruple-negative (ER / AR / HER2 / AR or ER / AR / HER2 / VDR-negative) and pentaplex-negative (ER / AR / HER2 / VDR / AR-negative) breast cancers into different outcome subgroups. Kaplan-Meier overall survival analysis charts reveal that the high-risk THR-immune subgroups correlate with poor survival in the KMP ER / HER2-negative cohorts (Fig. 39); THR- i20 (RHR=1.5, CI: 1.2-2.2, p=0.003), THR-i8 (RHR=2.1, CI: 1.6-3.0, p=8e-07) and THR-i3 (RHR=2.0, CI: 1.5-2.8, p=6.6e-06). In the TCGA triple-negative breast cancer group they also correlate with poor survival; THR-i20 (RHR=4.3, CI: 1.8-10.7, p=0.00039), THR-i8 (RHR=2.5, CI: 1.2-5.4, p=0.012) and THR-i3 (RHR=3.4, CI: 1.4-8.2, p=0.027). Figure 40 shows that THR-immune signatures also stratify poor outcome triple-negative breast cancer (TNBC) and Pentaplex / quadruple negative (PQNBC) breast cancer in the METABRIC dataset. While in the KMP dataset, the signatures with fewer genes (THR-i8 and -3) appear to outperform THR-i20, it appears that in the TCGA and METABRIC datasets, THR-i20 performs better. Lastly, in addition to breast cancer, we show that THR-i3 signature can stratify other gynecologic tumors such as cervix, ovary, and uterus adenocarcinomas into survival outcome groups in the Pan-cancer RNA-seq dataset (Fig. 41). THR-i8 performed similarly to THR-i3 in the same analysis (data not shown). Additionally, THR-i8 stratifies overall survival (OS) and relapse-free (RFS) outcome groups in multiple tumor types enriched for specific immune cell types; including cervix, ovary, uterus, and kidney clear cell carcinomas (CCC) that are enriched for natural killer cells (NK), kidney papillary carcinoma enriched for CD8+ T-cells and sarcomas enriched for type-1 T-helper cells in the Pan-cancer RNA-seq dataset (Fig. 42).
[0281] ET-12H and ET-13T signatures
[0282] We determined that the expression of a 12-gene signature (ET-12H) composed of ABTB1, CYP2S1, DEF6, FCHO1, SHKBP1, CCDC69, CCND2, CSK, HSA011916, MGAT1, PREXland WNT10A genes correlates with relapse-free (RFS) and disease-specific (DSS) survival of human epidermal growth factor receptor 2-positive (HER2+) breast cancer in three different cohorts; KMP, METABRIC and TCGA (Fig. 44A-C). When compared to 1,000 random sequences THR-12 outperforms 999 random controls, based on their p-value ranking (Fig. 44E). In contrast, Endopredict, Oncotype DX, PAM-50, and MammaPrint were once again indistinguishable from random control sequences matched according to their size (Fig. 44F-I), ranking 525th, 443rd, 992nd’ and 155* among 1,000 random controls (Fig. 44D).
[0283] Lastly, we determined that the expression of a 13-gene signature (ET-13T) composed of CACNG4, DKK3, FGF1, ADGRG1, ID3, PCDH1, THBS1, BAD, CLIC1, FOSL2, GRIP2, PLB1, and SNHG32 genes correlates with relapse-free (RFS) ER and HER-2-double-negative breast cancer in KMP and METABRIC datasets, chemotherapy response (Fig. 44A_C), and anti-HER2 treatment response in HER2-positive breast cancer (Fig. 45D) . When compared to 1,000 random sequences ET-13H outperforms 997 random controls, based on their p-value ranking (Fig. 45E). In contrast, Endopredict, Oncotype DX, PAM-50, and MammaPrint were once again indistinguishable from random control sequences matched according to their size (Fig. 44F-I), ranking 291st, 895th, 967th’ and 883rdamong 1,000 random controls (Fig. 45C-D).
[0284] Discussion
[0285] This example explores various fundamental questions in cancer research, concerning the interplay between initial conditions, represented by the cell-of-origin, and complex dynamical systems, represented by prognostic gene expression networks.
[0286] Is cancer biology primarily driven by acquired mutations? While both pre-existing conditions, (nature) and acquired factors (nurture) contribute to shaping complex traits, cancer research has frequently emphasized acquired mutations as the primary tumor drivers. Yet, a recent study titled “Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer”, suggests that cancer may be similar to other complex systems (17). Translating these genome-wide patterns, which involve hundreds of genes, into clinical applications has been challenging. This study addresses this by distilling the complexity into a practical signature of just six genes. This demonstrates that the normal cell- of-origin can serve as a powerful biomarker for breast cancer survival, outperforming established signatures that include driver mutations.
[0287] Are more parameters better in cancer modeling? The complexity of multi-omics data, often known as the curse of dimensionality, complicates the identification of clinically significant features. Research shows that too many parameters can introduce uncertainty into machine learning models (3). The findings support this, demonstrating that increasing the number of genes in a signature often leads to poorer performance compared to random controls of the same size. It was found that signatures larger than 20 genes tend to perform similarly to random gene sets. The larger the signature, the less effective it can become.
[0288] How can we determine the optimal size for a prognostic gene signature? It was found that the significance of a prognostic signature is inversely correlated with its size. Thus, the size of the THR-70 signature was reduced. However, testing all possible combinations was impractical due to the number of potential gene sets — ranging from 1 to 70 genes — totaling over one quintillion (1.18xl021), which created significant computational and multiple-testing challenges. Hence, the approach of using unsupervised heatmaps to explore gene clusters has proven unexpectedly effective in identifying optimal candidate signatures, such as THR-6Ee.
[0289] What is the most effective approach for multiple testing? This study raises questions about whether standard statistical significance thresholds are adequate for multi-gene signatures (18,19). The ongoing debate about statistical testing for omics studies often involves assumptions about the 'independence' of tests (20). Many omics studies, including those supporting approval of Oncotype DX, PAM-50, and MammaPrint were not corrected for multiple comparisons (21). In contrast, the use of random gene controls avoids any biases by providing an empirical, assumption-free alternative. Random sequences have proven useful as controls in other fields, such as siRNA gene knock-down studies, protein-DNA binding evaluations, and functional genomic screens, where they help distinguish true biological effects from non-specific results. Without such controls, it would be challenging to validate these studies.
[0290] Summary
[0291] This study combines two understudied yet powerful concepts for the discovery and evaluation of candidate biomarkers: the cancer cell-of-origin and random controls.
[0292] Understanding the initial conditions of complex dynamical systems is necessary for predicting their future behavior. Defining these conditions with a minimal number of parameters is also required given that unnecessary parameters can compromise predictive power. Based on these first principles, it was hypothesized that the initial conditions of cancer, defined by the normal cell-of-origin, can be succinctly described using a minimal number of unmutated genes. Such a signature would be less likely to be prognostic by chance alone.
[0293] It was found that prognostic signatures exceeding twenty gene transcripts become increasingly indistinguishable from random signatures of equal size. Therefore, concise, unmutated cell-of-origin signatures were identified including THR-6E, -4E, -i8, -i3, ET-12H, and ET-13T composed of just three to twelve transcripts, that outperform random controls and larger breast cancer prognostic signatures that are FDA cleared and commercialized, even those that include potent driver mutations.
[0294] Nearly two decades ago, a study in the Lancet found that many transcriptomic signatures failed to outperform chance in classifying patients. The authors recommended random sampling as a control, for validation of omics studies. However, despite additional studies with similar results, random controls have not become a standard practice. It seems disconcerting that the results presented herein reaffirm that even well-established and commercialized prognostic signatures such as Oncotype DX, PAM-50, and MammaPrint do not outperform random signatures of equal size. Two decades later, this is no longer just a theoretical debate; these signatures have received regulatory approval and are now used in oncology practice. Methods
[0295] Survival Analysis
[0296] Unsupervised Clustering of Patient Samples
[0297] Hierarchical clustering was performed on all samples using the ward minimum variance method, then the hierarchical tree was cut into five groups. The optimal number of groups was determined based on the observed patterns in the data and the clinical interpretability of the results.
[0298] Software and Statistical Analysis
[0299] All statistical analyses were performed using R software (version 4.0.3). The survival and survminer R packages were used for generating Kaplan-Meier survival curves and the COX proportional hazards models. The stats package was used for hierarchical clustering, while the glmnet package was used to fit the logistic regression models. The significance level (p-value and FDR) was set at 0.05 for all statistical tests.
[0300] Survival analysis was carried out to investigate the association between the THR and ET signatures and survival across different patient cohorts using the Kaplan-Meier (KM) survival estimate with a log-rank test. In the METABRIC cohort, a logistic regression model was used to compute a risk-score for each patient using the expression levels of the signature genes. The risk scores were then categorized into binary labels of 0 (low-risk) and 1 (high-risk) using the optimal threshold from the Receiver Operating Characteristics (ROC) Curve. Additionally, these scores were categorized into lowest and highest risk using the median cutoff or optimized cut-offs. In the KMP BrCa dataset, survival analysis was performed by a median cut-off. The Benjamini-Hochberg method was used to correct for multiple hypothesis testing.
[0301] Comparison with Other Signatures
[0302] The prognostic performance of the THR signature was compared with that of other established BrCa signatures, including PAM-50, Oncotype DX, and MammaPrint. This comparison was conducted within the same patient cohort (KMP or METABRIC cohorts), applying the same patient stratification methodology. For comparisons performed within the KMP cohort, an expression-based approach was used, wherein the average expression of the signature genes was computed across the patient cohort and then categorized patients into low- and high-expression groups using the optimum cutoff. For the comparisons with the METABRIC cohort, a risk-score-based approach was followed, wherein for each signature, a risk-score for each patient was computed using a logistic regression model, then categorized patients into low- and high-risk groups using the optimum cutoff. Notably, not all genes / probes in the studied signatures were identified in the patient cohorts used in the study. For instance, in the KMP cohort, four genes four genes of MammaPrint (AA555029_RC, LOC100131053, LOC100288906, and hCG 1980668) were not identified.
[0303] Bibliography
[0304] 1. Rogers TL, Johnson BJ, Munch SB. Chaos is not rare in natural ecosystems. Nat Ecol Evol 2022;6(8): 1105-11 doi 10.1038 / s41559-022-01787-y.
[0305] 2. Santagata S, Ince TA. Normal cell phenotypes of breast epithelial cells provide the foundation of a breast cancer taxonomy. Expert Rev Anticancer Ther 2014;14(12):1385-9 doi 10.1586 / 14737140.2014.956096.
[0306] 3. Torres R, Judson-Torres RL. Research Techniques Made Simple: Feature Selection for Biomarker Discovery. J Invest Dermatol 2019; 139(101:2068-74 el doi 10.1016 / j.jid.2019.07.682.
[0307] 4. Ince TA, Richardson AL, Bell GW, Saitoh M, Godar S, Karnoub AE, et al. Transformation of different human breast epithelial cell types leads to distinct tumor phenotypes. Cancer Cell 2007; 12(2): 160-70 doi 10.1016 / j.ccr.2007.06.013.
[0308] 5. Bhattacharya U, Kamran M, Manai M, Cristofanilli M, Ince TA. Cell-of-Origin Targeted Drug Repurposing for Triple-Negative and Inflammatory Breast Carcinoma with HDAC and HSP90 Inhibitors Combined with Niclosamide. Cancers (Basel) 2023; 15(2) doi 10.3390 / cancersl5020332.
[0309] 6. Kamran M, Bhattacharya U, Omar M, Marchionni L, Ince TA. ZNF92, an unexplored transcription factor with remarkably distinct breast cancer over-expression associated with prognosis and cell-of-origin. NPJ Breast Cancer 2022;8(l):99 doi 10. 1038 / s41523-022 -00474- 2.
[0310] 7. Santagata S, Thakkar A, Ergonul A, Wang B, Woo T, Hu R, et al. Taxonomy of breast cancer based on normal cell phenotype predicts outcome. J Clin Invest 2014;124(2):859-70 doi 10.1172 / JCI70941.
[0311] 8. Houseman EA, Ince TA. Normal cell-type epigenetics and breast cancer classification: a case study of cell mixture-adjusted analysis of DNA methylation data from tumors. Cancer Inform 2014;13(Suppl 4):53-64 doi 10.4137 / CIN.S 13980.
[0312] 9. Andre F, Ismaila N, Allison KH, Barlow WE, Collyar DE, Damodaran S, et al. Biomarkers for Adjuvant Endocrine and Chemotherapy in Early-Stage Breast Cancer: ASCO Guideline Update. J Clin Oncol 2022;40(16):1816-37 doi 10. 1200 / JC0.22.00069. 10. Domany E. Using high-throughput transcriptomic data for prognosis: a critical overview and perspectives. Cancer Res 2014;74(17):4612-21 doi 10.1158 / 0008-5472.CAN- 13-3338.
[0313] 11. Zou JX, Duan Z, Wang J, Sokolov A, Xu J, Chen CZ, et al. Kinesin family deregulation coordinated by bromodomain protein ANCCA and histone methyltransferase MLL for breast cancer cell growth, survival, and tamoxifen resistance. Mol Cancer Res 2014;12(4):539-49 doi 10.1158 / 1541-7786.MCR-13-0459.
[0314] 12. Huang W, Ray P, Ji W, Wang Z, Nancarrow D, Chen G, et al. The cytochrome P450 enzyme CYP24A1 increases proliferation of mutant KRAS-dependent lung adenocarcinoma independent of its catalytic activity. J Biol Chem 2020;295(18):5906-17 doi 10.1074 / jbc.RAl 19.011869.
[0315] 13. Cao Q, Song Z, Ruan H, Wang C, Yang X, Bao L, et al. Targeting the KIF4A / AR Axis to Reverse Endocrine Therapy Resistance in Castration-resistant Prostate Cancer. Clin Cancer Res 2020;26(6): 1516-28 doi 10.1158 / 1078-0432.CCR-19-0396.
[0316] 14. Sun B, Long Y, Xiao L, Wang J, Yi Q, Tong D, et al. Target Protein for Xklp2 Functions as Coactivator of Androgen Receptor and Promotes the Proliferation of Prostate Carcinoma Cells. J Oncol 2022;2022:6085948 doi 10.1155 / 2022 / 6085948.
[0317] 15. Zhou Y, Ou L, Xu J, Yuan H, Luo J, Shi B, et al. FAM64A is an androgen receptor- regulated feedback tumor promoter in prostate cancer. Cell Death Dis 2021;12(7):668 doi 10.1038 / s41419-021-03933-z.
[0318] 16. Novaro V, Roskelley CD, Bissell MJ. Collagen-IV and laminin- 1 regulate estrogen receptor alpha expression and function in mouse mammary epithelial cells. J Cell Sci 2003;116(Pt 14):2975-86 doi 10.1242 / jcs.00523.
[0319] 17. Hoadley KA, Yau C, Hinoue T, Wolf DM, Lazar AJ, Drill E, et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell 2018;173(2):291-304 e6 doi 10.1016 / j.cell.2018.03.022.
[0320] 18. loannidis JP. Microarrays and molecular research: noise discovery? Lancet 2005;365(9458):454-5 doi 10.1016 / S0140-6736(05)17878-7.
[0321] 19. loannidis JPA. The Proposal to Lower P Value Thresholds to .005. JAMA 2018:319(14): 1429-30 doi 10.1001 / jama.2018.1536.
[0322] 20. Streiner DL. Best (but oft-forgotten) practices: the multiple problems of multiplicity- whether and how to correct for many statistical tests. Am J Clin Nutr 2015;102(4):721-8 doi 10.3945 / ajcn. 115. 113548. 21. Sparano JA, Gray RJ, Ravdin PM, Makower DF, Pritchard KI, Albain KS, et al. Clinical and Genomic Risk to Guide the Use of Adjuvant Therapy for Breast Cancer. N Engl J Med 2019;380(25):2395-405 doi 10.1056 / NEJMoal904819.
[0323] 22. Michiels S, Koscielny S, Hill C. Prediction of cancer outcome with microarrays: a multiple random validation strategy. Lancet 2005;365(9458):488-92 doi 10.1016 / S0140- 6736(05)17866-0.
[0324] 23. Venet D, Dumont JE, Detours V. Most random gene expression signatures are significantly associated with breast cancer outcome. PLoS Comput Biol 2011;7(10):el002240 doi 10. 1371 / journal.pcbi. 1002240.
[0325] 24. Boyle EA, Li YI, Pritchard JK. An Expanded View of Complex Traits: From Polygenic to Omnigenic. Cell 2017; 169(7): 1177-86 doi 10.1016 / j.cell.2017.05.038.
[0326] 25. Michaelis AC, Brunner AD, Zwiebel M, Meier F, Strauss MT, Bludau I, et al. The social and structural architecture of the yeast protein interactome. Nature 2023; 624(7990): 192- 200 doi 10.1038 / s41586-023-06739-5.
[0327] 26. Kumar S, Vindal V. Architecture and topologies of gene regulatory networks associated with breast cancer, adjacent normal, and normal tissues. Funct Integr Genomics 2023;23(4):324 doi 10.1007 / sl0142-023-01251-5.
[0328] 27. Embar V, Handen A, Ganapathiraju MK. Is the average shortest path length of gene set a reflection of their biological relatedness? J Bioinform Comput Biol 2016; 14(6): 1660002 doi 10. 1142 / S0219720016600027.
[0329] 28. Tschodu D, Lippoldt J, Gottheil P, Wegscheider AS, Kas JA, Niendorf A. Re- evaluation of publicly available gene-expression databases using machine-learning yields a maximum prognostic power in breast cancer. Sci Rep 2023;13(l):16402 doi 10. 1038 / s41598- 023-41090-9.
[0330] 29. Bartlett JMS, Bayani J, Kornaga E, Xu K, Pond GR, Piper T, et al. Comparative survival analysis of multiparametric tests-when molecular tests disagree-A TEAM Pathology study. NPJ Breast Cancer 2021;7(l):90 doi 10.1038 / s41523-021-00297-7.
[0331] 30. Varga Z, Sinn P, Seidman AD. Summary of head-to-head comparisons of patient risk classifications by the 21 -gene Recurrence Score(R) (RS) assay and other genomic assays for early breast cancer. Int J Cancer 2019;145(4):882-93 doi 10. 1002 / ijc.32139.
[0332] 31. Bartlett JM, Bayani J, Marshall A, Dunn JA, Campbell A, Cunningham C, et al. Comparing Breast Cancer Multiparameter Tests in the OPTIMA Prelim Trial: No Test Is More Equal Than the Others. J Natl Cancer Inst 2016; 108(9) doi 10.1093 / jnci / djw050. All patents and publications referenced or mentioned herein are indicative of the levels of skill of those skilled in the art to which the invention pertains, and each such referenced patent or publication is hereby specifically incorporated by reference to the same extent as if it had been incorporated by reference in its entirety individually or set forth herein in its entirety. Applicants reserve the right to physically incorporate into this specification any and all materials and information from any such cited patents or publications.
[0333] The specific methods, devices, and compositions described herein are representative of preferred embodiments and are exemplary and not intended as limitations on the scope of the invention. Other objects, aspects, and embodiments will occur to those skilled in the art upon consideration of this specification and are encompassed within the spirit of the invention as defined by the scope of the claims. It will be readily apparent to one skilled in the art that varying substitutions and modifications may be made to the invention disclosed herein without departing from the scope and spirit of the invention.
[0334] The invention illustratively described herein suitably may be practiced in the absence of any element or elements, or limitation or limitations, which is not specifically disclosed herein as essential. The methods and processes illustratively described herein suitably may be practiced in differing orders of steps, and the methods and processes are not necessarily restricted to the orders of steps indicated herein or in the claims.
[0335] Under no circumstances may the patent be interpreted to be limited to the specific examples or embodiments or methods specifically disclosed herein. Under no circumstances may the patent be interpreted to be limited by any statement made by any Examiner or any other official or employee of the Patent and Trademark Office unless such statement is specifically and without qualification or reservation expressly adopted in a responsive writing by Applicants.
[0336] The terms and expressions that have been employed are used as terms of description and not of limitation, and there is no intent in the use of such terms and expressions to exclude any equivalent of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, it will be understood that although the present invention has been specifically disclosed by preferred embodiments and optional features, modification, and variation of the concepts herein disclosed may be resorted to by those skilled in the art and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims and statements of the invention. The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also forms part of the invention. This includes the generic description of the invention with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein. In addition, where features or aspects of the invention are described in terms of Markush groups, those skilled in the art will recognize that the invention is also thereby described in terms of any individual member or subgroup of members of the Markush group.
Claims
WHAT IS CLAIMED IS:
1. A method comprising:(a) assaying a biological sample from a subject for expression of (1) ET-60 biomarkers combined with THR-50 biomarkers recited in Table 1 as CA-110, or (2) ET-60 biomarkers combined with THR-70 biomarkers recited in Table 1 as CA-130, to determine expression levels for the ET-60 and THR-50 biomarkers or the ET-60 and THR-70;(b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’ s biological sample, wherein altered expression levels of the CA-110 (ET-60 and THR-50) biomarkers or the CA- 130 (ET-60 and THR-70) biomarkers in the biological sample relative to the reference value provides a prognosis of longer overall survival and / or longer recurrence-free survival; and(c) optionally administering or withholding one or more cancer treatments to a subject determined to have a cancer with a prediction of low overall survival and / or low progression-free survival.
2. The method of claim 1, wherein the sample is a breast cancer, colon cancer, lung cancer, prostate cancer, kidney cancer, cervix cancer, melanoma, ovarian cancer, liver cancer, pancreatic cancer, head & neck cancer, bladder cancer, brain tumors such as glioma, or hematopoietic neoplasms such as multiple myeloma sample.
3. The method of claim 1, wherein the CA-110 (ET-60 and THR-50) biomarkers are prognostic for breast cancer, colon cancer, lung cancer, prostate cancer, or pancreatic cancer.
4. The method of claim 1, wherein the CA-130 (ET-60 and THR-70) biomarkers are prognostic for breast cancer, kidney cancer, cervix cancer, melanoma, ovarian cancer, liver cancer, colon cancer, pancreas cancer, lung cancer, prostate cancer, head & neck cancer, bladder cancer, brain tumors such as glioma, or hematopoietic neoplasms such as multiple myeloma.
5. The method of claim 1, wherein RNA expression is assayed.
6. The method of claim 1, wherein nucleic acid amplification is employed prior to assaying.
7. The method of claim 1, wherein protein expression is assayed.
8. A method comprising:(a) assaying a biological sample from a subject for expression of THR-50 biomarkers recited in Table 1, or THR-70 biomarkers recited in Table 1 to determine expression levels for the THR-50 or THR-70 biomarkers;(b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’ s biological sample, wherein a low mean expression of the THR-50 or THR-70 genes correlates to shorter overall survival (OS), or shorter recurrence-free survival (RFS), or metastasis-free survival (DMFS) of breast cancer; and(c) optionally administering or withholding one or more cancer treatments to a subject determined to have a breast cancer with a prediction of shorter OS, shorter RFS, and shorter DMFS.
9. The method of claim 8, wherein expression of the THR-50 or THR-70 genes is correlated with survival in each of ER-neg, LN+, LN-, AR+, Grade 2, Grade 3, Lum- A, Lum- B, HER2+, HER2-, HER2-like and Basal-like breast cancer subtypes.
10. The method of claim 8, wherein the expression of the THR-50 or THR-70 genes are combined with cellular ancestry information, triple hormone receptor protein expression status, gene mutation information, immune infiltrate information, or a combination thereof, to provide OS, RFS or DMFS time.
11. The method of claim 8, wherein expression levels of THR-6E biomarkers are determined.
12. The method of claim 11, wherein the expression of the THR-6E bio markers is correlated with survival in ERp-HER2n-LNn breast cancers.
13. The method of claim 11, wherein the expression of the THR-6E bio markers iscorrelated with survival in each of LN-positive and LN-negative patients, untreated, hormone-treated, chemotherapy-treated, and ER-negative cancers.
14. The method of claim 11, wherein the expression of the THR-6E biomarkers is correlated with response to endocrine therapy in ER-positive / LN-negative cancers, and chemotherapy in ER-positive / LN-positive breast cancers.
15. The method of claim 8, wherein expression levels of THR-4E biomarkers are determined.
16. The method of claim 15, wherein the expression of the THR-4E bio markers is correlated with survival in ER-positive breast cancer.
17. The method of claim 11, wherein the expression of the THR-4E biomarkers is correlated, with response to endocrine therapy in ER-positive / LN-negative breast cancers, and chemotherapy in ER-positive / LN-positive breast cancers.
18. The method of claim 8, wherein the THR70 biomarkers are combined with i20 immune biomarkers recited in Table 1, and +HER2 signature to identify subgroups with shorter and longer survival.
19. The method of claim 18, wherein the expression of the THR70 biomarkers combined with the i20 immune biomarkers is correlated with survival in each of Lum-A, ER+LP, Lum- B, ER+HP, Basal-like, TNBC and HER-2-like tumors.
20. The method of claim 8, wherein RNA expression is assayed.
21. The method of claim 8, wherein nucleic acid amplification is employed prior to assaying.
22. The method of claim 8, wherein protein expression is assayed.
23. A method comprising:(a) assaying a biological sample from a subject for expression of i20 biomarkersrecited in Table 1 to determine expression levels for the i20 biomarkers;(b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’s biological sample, wherein altered expression levels correlate to overall survival (OS) or relapse-free survival (RFS) of ERnegative breast cancer; and(c) optionally administering or withholding one or more cancer treatments to a subject determined to have a breast cancer with a prediction of shorter OS or RFS.
24. The method of claim 23, wherein expression levels of THR-i8 or THR-i3 biomarkers are determined.
25. The method of claim 24, wherein expression levels of the THR-i8 or the THR-i3 biomarkers are correlated with outcome of double-negative (ER / HER2-negative), triplenegative (ER / PR / HER2-negative), quadruple-negative (ER / AR / HER2 / AR or ER / AR / HER2 / VDR-negative) and pentaplex-negative (ER / AR / HER2 / VDR / AR-negative) breast cancers.
26. The method of claim 24, wherein expression levels of the THR-i8 or THR-i3 biomarkers are correlated with survival of a gynecologic tumor, kidney cancer, or sarcoma.
27. The method of claim 26, wherein the gynecologic tumor is cervix, ovary, or uterus adenocarcinomas into survival outcome groups.
28. The method of claim 24, wherein expression levels of the THR-i8 biomarkers are correlated with overall survival (OS) or relapse-free (RFS) outcome in the cervix, ovary, uterus, kidney carcinomas, and sarcomas.
29. The method of claim 23, wherein RNA expression is assayed.
30. The method of claim 23, wherein nucleic acid amplification is employed prior to assaying.
31. The method of claim 23, wherein protein expression is assayed.
32. A method comprising:(a) assaying a biological sample from a subject for expression of ET-12H, or ET-13T biomarkers recited in Table 1 to determine expression levels for the ET-12H, or ET-13T biomarkers;(b) comparing the determined expression levels with one or more reference values to identify any altered expression levels in the subject’s biological sample, wherein altered expression levels correlate to overall survival (OS) or relapse-free survival (RFS) of ERnegative breast cancer; and(c) optionally administering or withholding one or more cancer treatments to a subject determined to have a breast cancer with a prediction of shorter OS or RFS.
33. The method of claim 32, wherein expression levels of the ET-12H biomarkers are correlated with survival of HER2-positive breast cancer, and ET-13T biomarker is correlated with survival of ER / HER2-negative breast cancer.
34. The method of claim 32, wherein expression levels of the ET-13T biomarker are correlated with response to breast cancer chemotherapy and anti-HER2 therapy.
Citation Information
Patent Citations
Prognostic / predictive epigenetic breast cancer signature
WO2023122758A1