Methods, kits and systems for determining the er status of cancer and methods for treating cancer based on same

WO2026206885A1PCT designated stage Publication Date: 2026-10-01PRECEDE BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020450
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-10-10
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure US2026020450_01102026_PF_FP_ABST
    Figure US2026020450_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure includes, among other things, methods, kits, and systems for determining ER status of cancer, e.g., a breast cancer. In various embodiments, the present disclosure relates to the use of one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation that that are characteristic of ER status of cancer. In some embodiments, differential modifications and / or differential accessibility are detected and quantified at one or more genomic loci of a biological sample, e.g., in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from a subject with cancer. In various embodiments a determined ER status is useful, e.g., in selecting treatment for and / or treating a cancer, e.g., a breast cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket: 2014191-0051METHODS, KTTS AND SYSTEMS FOR DETERMINING THE ER STATUS OF CANCER AND METHODS FOR TREATING CANCER BASED ON SAMEBACKGROUND

[0001] It has long been recognized that some human breast cancers are hormone dependent. Estrogen regulates the differentiation and proliferation of breast epithelial cells and interacts with the estrogen receptor (ER) in the nucleus. Prolonged exposure of estrogen is an important risk factor for cancer. Progesterone receptor (PR) expression in normal breast epithelium is regulated by ER (Jensen, Cancer (1980) 46:2759-2761). Presence of ER, PR and human epidermal growth factor receptor-2 (HER2) status in invasive breast carcinoma is now routinely estimated as these markers are considered to be important prognostic factors. ER and PR status has been used for many years to determine a patient’s suitability for treatment with endocrine therapy (e.g., tamoxifen).

[0002] To determine if a cancer is ER-positive, medical practitioners currently order testing that is conducted on a tissue sample using immunohistochemistry (IHC). Samples are reviewed by a pathologist and typically reported as (a) the word positive or negative, (b) a percentage that tells you how many cells out of 100 stained positive for hormone receptors, i.e., a number between 0% (none have receptors) and 100% (all have receptors), and / or (c) an Allred score between 0 and 8. The Allred scoring system looks at what percentage of cells test positive for hormone receptors, along with how well the receptors show up after staining, called intensity (Allred et al., Breast Cancer Res (2004) 6:240-245). This information is then combined to score the sample on a scale from 0 to 8 where, the higher the score, the more receptors were found and the easier they were to see in the sample.

[0003] ER-positive cancers can be treated with ER-targeted agents that lower estrogen levels or block estrogen receptors. Conversely, treatment with ER-targeted agents is not helpful for ER-negative cancers. These cancers may instead be treated with one or more of surgery and / or radiation, HER2 -targeted therapy (if HER2 -positive), chemotherapy and immunotherapy.

[0004] Current methods for assessing ER status are invasive and rely on a single tissue biopsy to characterize the ER status in metastatic breast cancer. Such methods focus only on a small region at a single tumor site at a given time and therefore do not accurately capture tumor heterogeneity or receptor evolution and therefore only partially characterize the relevant patient population.13403741vl Page 1 of 173Attorney Docket: 2014191-0051SUMMARY

[0005] There is a need in the art for more comprehensive, less invasive, and more precise diagnostic methods for determining ER status, including methods that are independent of IHC testing. Improved diagnostic methods would improve patient outcomes and also better support future clinical trials that seek to identify subpopulations of patients that respond to ER-targeted agents. They would also expand our understanding of the underlying biology of ERpositive cancer and help identify new treatments.

[0006] The present disclosure is based, at least in part, on the demonstration that the ER status of a cancer in a subject can be determined by detecting and quantifying the presence of histone modifications and / or DNA methylation at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample, e.g, a plasma sample obtained or derived from the subject. The present disclosure also encompasses methods where chromatin accessibility and / or binding of one or more transcription factors are detected at the one or more genomic loci instead of (or in addition to) histone modifications and / or DNA methylation. The present disclosure is also based, at least in part, on the demonstration that genomic loci at which the amount of one or more epigenetic biomarkers (e.g., histone methylation marks such as H3K4me3 and histone acetylation marks such as H3K27ac and / or DNA methylation) is correlated with target gene expression (e.g., ESRI Nor ESRI related gene expression) can be combined into multimodal classifiers to determine ER status. These new monomodal and multimodal classifiers provide minimally invasive ways of determining ER status that are more accurate, objective, and comprehensive than the current tissue-based approaches. No liquid biopsy platform to date has been able to provide actionable resolution on a transcriptionally regulated phenotype relevant for therapy such as ER status.

[0007] The present disclosure includes, among other things, technologies for the determination of ER status and for the detection, monitoring, and / or treatment of cancer (including, e.g., breast, ovarian, or endometrial cancer) based on ER status. In various embodiments, the present disclosure relates to the measurement of histone modifications in a sample obtained or derived from a subject to detect and / or treat cancer (including, e.g., breast, ovarian, or endometrial cancer) based on ER status. The present disclosure includes, among other things, histone modification measurements in cell-free DNA (cfDNA) that are characteristic of cancer, and which in various embodiments are useful, e.g., for detecting, monitoring, selecting13403741vl Page 2 of 173Attorney Docket: 2014191-0051treatment for, and / or treating cancer (including, e.g., breast, ovarian, or endometrial cancer) based on ER status. The present disclosure includes, among other things, histone modification measurements in cfDNAthat are characteristic of ER-positive and ER-negative cancers, which in various embodiments are useful, e.g., in detecting, monitoring, selecting treatment for, and / or treating an ER-positive and ER-negative cancers. In some embodiments, histone modification measurements in cfDNA can be used to detect or determine resistance of a cancer (e.g., breast, ovarian, or endometrial cancer) to a therapy or transformation of a cancer from one subtype to another. In various embodiments, the present disclosure includes exemplary genomic loci that are differentially modified in ER-positive vs. ER-negative cancer, e.g., breast, ovarian, or endometrial cancer. In various embodiments, genomic loci differentially modified in cfDNA are or include one or more enhancers. In various embodiments, genomic loci differentially modified in cfDNA are or include one or more promoters.

[0008] In various embodiments, a genomic locus is differentially modified if it is characterized by increased or decreased histone modification as compared to a reference (e.g., a sample from an ER-negative or healthy subject). Increased or decreased histone modification can be or include, e.g., increased or decreased histone methylation (hypermethylation or hypomethylation, respectively) of one or more particular methylation marks, or a combination thereof; increased or decreased pan-methylation; increased or decreased histone acetylation (hyperacetylation or hypoacetylation, respectively) of one or more particular acetylation marks, or a combination thereof; and / or increased or decreased pan-acetylation (e.g., pan-H3 acetylation). In various embodiments, histone methylation can be or include histone methylation marks selected from H3K4mel, H3K4me2, H3K4me3, or a combination thereof. In various embodiments, histone methylation can be or include H3K4me3. In various embodiments, histone acetylation can be or include histone acetylation marks selected from H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, or a combination thereof. In various embodiments, histone acetylation can be or include H3K27ac.

[0009] In various embodiments, the present disclosure relates to the measurement of DNA methylation in a sample obtained or derived from a subject to detect and / or treat cancer (including, e.g., breast, ovarian, or endometrial cancer) based on ER status. The present disclosure includes, among other things, DNA methylation measurements in cell-free DNA (cfDNA) that are characteristic of cancer, and which in various embodiments are useful, e.g., for13403741vl Page 3 of 173Attorney Docket: 2014191-0051detecting, monitoring, selecting treatment for, and / or treating cancer (including, e.g., breast, ovarian, or endometrial cancer) based on ER status. In some embodiments, DNA methylation measurements in cfDNA can be used to detect or determine resistance of a cancer (e., breast, ovarian, or endometrial cancer) to a therapy or transformation of a cancer from one subtype to another. In various embodiments, the present disclosure includes exemplary genomic loci that are differentially DNA methylated in ER-positive vs. ER-negative cancer, e.g., breast, ovarian, or endometrial cancer. In various embodiments, a genomic locus is differentially modified if it is characterized by increased or decreased DNA methylation as compared to a reference (e.g., a sample from an ER-negative or healthy subject). In various embodiments, genomic loci differentially modified in cfDNA are or include one or more enhancers. In various embodiments, genomic loci differentially modified in cfDNA are or include one or more promoters.

[0010] The present disclosure further relates, in various embodiments, to the measurement of chromatin accessibility in cell-free DNA (cfDNA) to determine ER status. The present disclosure includes, among other things, chromatin accessibility measurements in cfDNA that are characteristic of ER-positive cancers, which in various embodiments are useful, e.g., in detecting, monitoring, selecting treatment for, and / or treating an ER-positive cancer. In some embodiments, chromatin accessibility measurements in cfDNA can be used to detect or determine resistance of a cancer (e.g., breast, ovarian, or endometrial cancer) to a therapy or transformation of a cancer from one subtype to another. In various embodiments, the present disclosure includes genomic loci that are differentially accessible in ER-positive vs. ER-negative cancers. In various embodiments, genomic loci differentially accessible in cfDNA are or include one or more enhancers. In various embodiments, genomic loci differentially accessible in cfDNA are or include one or more promoters.

[0011] In various embodiments, without wishing to be bound by any particular scientific theory, histone methylation (e.g., H3K4me3) corresponds and / or is correlated with chromatin accessibility. In various embodiments, without wishing to be bound by any particular scientific theory, histone acetylation (e.g., H3K27ac) corresponds and / or is correlated with chromatin accessibility. In various embodiments, without wishing to be bound by any particular scientific theory, DNA methylation corresponds and / or is correlated with chromatin accessibility.

[0012] In various embodiments, a genomic locus is differentially accessible if it is characterized by increased or decreased chromatin accessibility as compared to a reference (e.g.,13403741vl Page 4 of 173Attorney Docket: 2014191-0051a sample from an ER-negative or healthy subject). Increased or decreased histone modification can be or include, e.g., increased or decreased accessibility as determined by various chromatin accessibility assays known in the art.

[0013] The present disclosure further relates, in various embodiments, to the measurement of transcription factor binding in cell-free DNA (cfDNA) to determine ER status. The present disclosure includes, among other things, transcription factor binding measurements in cfDNA that are characteristic of ER-positive cancers, which in various embodiments are useful, e.g., in detecting, monitoring, selecting treatment for, and / or treating an ER-positive cancer. In some embodiments, transcription factor binding measurements in cfDNA can be used to detect or determine resistance of a cancer e.g., breast, ovarian, or endometrial cancer) to a therapy or transformation of a cancer from one subtype to another. In various embodiments, the present disclosure includes genomic loci that are differentially bound by transcription factors in ER-positive vs. ER-negative cancers. In various embodiments, genomic loci that are differentially bound by transcription factors in cfDNA are or include one or more enhancers. In various embodiments, genomic loci that are differentially bound by transcription factors in cfDNA are or include one or more promoters.

[0014] In various embodiments, without wishing to be bound by any particular scientific theory, histone methylation (e.g., H3K4me3) corresponds and / or is correlated with transcription factor binding. In various embodiments, without wishing to be bound by any particular scientific theory, histone acetylation (e.g., H3K27ac) corresponds and / or is correlated with transcription factor binding. In various embodiments, without wishing to be bound by any particular scientific theory, DNA methylation corresponds and / or is correlated with transcription factor binding.

[0015] In various embodiments, a genomic locus is differentially bound by transcription factors if it is characterized by increased or decreased transcription factor binding as compared to a reference (e.g., a sample from an ER-negative or healthy subject). Increased or decreased transcription factor binding can be or include, e.g., increased or decreased transcription factor binding as determined by various transcription factor binding assays known in the art.

[0016] In some aspects, the present disclosure is directed to a method of prediction of ER status that comprises predicting ESRI expression level by quantifying one or more epigenetic biomarkers in a liquid biopsy sample from a subject. The method may include providing an expression level prediction model that has been produced from digital samples for an indication13403741vl Page 5 of 173Attorney Docket: 2014191-0051each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication (e.g., breast cancer) and (ii) an expression level for ESRI, wherein the digital samples have been generated using data derived from cell samples specific to the indication and healthy volunteers. The method may include providing input data including signal for the one or more epigenetic biomarkers derived from a liquid biopsy sample for a subject. The method may include predicting expression level of ESRI for the subject from the input data using the model.

[0017] In some embodiments, one or more epigenetic biomarkers described herein comprise (i) one or more histone modifications, (ii) chromatin accessibility, (iii) binding of one or more transcription factors, and / or (iv) DNA methylation, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject. In some embodiments, one or more genomic loci comprise one or more expression-level correlated loci for ESRI. In some embodiments, one or more genomic loci comprise one or more expressionlevel correlated loci for ESRI and / or one or more ESRI related genes.

[0018] In some embodiments, one or more genomic loci described herein comprise one or more genomic loci that are within + / - 200 kB of ESRI, and optionally wherein the one or more genomic loci include one or more of the ESRI associated loci provided in Table 1.

[0019] In some embodiments, one or more ESRI related genes include ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.

[0020] In some embodiments, one or more genomic loci described herein comprise: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESR1. In some embodiments, one or more genomic loci comprise: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for one or more of ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof. In some embodiments, one or more genomic loci comprise: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRP, and one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for13403741vl Page 6 of 173Attorney Docket: 2014191-0051EN01, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, CCDC170, or any combination thereof.

[0021] In some embodiments, one or more expression-level correlated loci for ESRI, EN01, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, or CCDC170 are proximal to (e.g., within + / - 200 kB of the transcription start site (TSS) of) ESRI, EN01, YBX1, GATA3, FOXAl, HAPLN3, ENI, P1M1, or CCDC170, respectively.

[0022] In some embodiments, one or more expression-level correlated loci for ESRI, EN01, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, CCDC170, or any combination thereof include one or more promoter regions for one or more of ESRI, EN01, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, CCDC170, or any combination thereof. In some embodiments, one or more expression-level correlated loci for ESRI, ENO1, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, CCDC170, or any combination thereof include one or more enhancer regions for one or more of ESRI, ENO1, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, CCDC170, or any combination thereof. In some embodiments, one or more expression-level correlated loci for ESRI, EN01, YBX1, GATA3, FOXAl, HAPLN3, ENI, PIM1, CCDC170, or any combination thereof, are genomic regions at which signal of one or more epigenetic biomarkers (i) has been shown to be correlated with ESR1 expression, and / or (ii) has been shown to be correlated with ENO1, YBX1, GATA3, FOXAl, HAP LN 3, ENI, P1M1, or CCDC170 expression, or any combination thereof. In some embodiments, one or more genomic loci described herein comprise: (i) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expressionlevel correlated loci that are proximal to ESRI and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, or 75 or more expression-level correlated loci that are proximal to ENO1 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, or 15 or more expression-level correlated loci that are proximal to YBX1 and that are provided in Table 1; (iv) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, or 90 or more expression-level correlated loci that are proximal to GATA3 and that are provided in Table 1; (v) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci that are proximal to FOXAl and that are provided in Table 1; (vi) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 55 or13403741vl Page 7 of 173Attorney Docket: 2014191-0051more expression-level correlated loci that are proximal to HAPLN3 and that are provided in Table 1; (vii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, or 125 or more expression-level correlated loci that are proximal to ENJ and that are provided in Table 1; (viii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more expression-level correlated loci that are proximal to PIM1 and that are provided in Table 1; (ix) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 115 or more expression-level correlated loci that are proximal to CCDC170 and that are provided in Table 1; or (x) any combination of (i)-(ix).

[0023] In some embodiments, one or more expression-level correlated loci for ESRI include: (i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with ESRI and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 60 or more “H3K27ac” analyte loci that are associated with ESRI and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ESRI and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0024] In some embodiments, one or more expression-level correlated loci for ENO1 include: (i) one or more, 5 or more, or 10 or more “H3K4me3” analyte loci that are associated with ENO1 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “H3K27ac” analyte loci that are associated with ENO1 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ENO I and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0025] In some embodiments, one or more expression-level correlated loci for YBX1 include: (i) one or more “H3K27ac” analyte loci that are associated with YBX1 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, or 15 or more “MBD” analyte loci that are associated with YBX1 and that are provided in Table 1; or (iii) a combination of (i) and (ii).

[0026] In some embodiments, one or more expression-level correlated loci for GATA3 include: (i) one or more, 5 or more, 10 or more, or 15 or more “H3K4me3” analyte loci that are13403741vl Page 8 of 173Attorney Docket: 2014191-0051associated with GATA3 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 “H3K27ac” analyte loci that are associated with GATA3 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “MBD” analyte loci that are associated with GATA3 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0027] In some embodiments, one or more expression-level correlated loci for FOXA1 include: (i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with FOXA1 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with FOXA1 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with FOXA1 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0028] In some embodiments, one or more expression-level correlated loci for HAPLN3 include: (i) one or more, 2 or more, 3 or more, 4 or more, 5 or more, 5 or more, 7 or more, 8 or more, 9 or more, or 10 or more “H3K4me3” analyte loci that are associated with HAPLN3 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K27ac” analyte loci that are associated with HAPLN3 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “MBD” analyte loci that are associated with HAPLN3 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0029] In some embodiments, one or more expression-level correlated loci for ENJ include: (i) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “H3K4me3” analyte loci that are associated with EN1 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with EN1 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with EN1 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0030] In some embodiments, one or more expression-level correlated loci for PIM1 include: (i) one or more, 2 or more, or 3 or more “H3K4me3” analyte loci that are associated with PIM1 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 15 or more,13403741vl Page 9 of 173Attorney Docket: 2014191-005120 or more, 25 or more, 30 or more, or 35 or more “H3K27ac” analyte loci that are associated with PIM1 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with PIM1 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0031] In some embodiments, one or more expression-level correlated loci for CCDC170 include: (i) one or more, 5 or more, 10 or more, or 20 or more “H3K4me3” analyte loci that are associated with CCDC170 and that are provided in Table 1; (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more “H3K27ac” analyte loci that are associated with CCDC170 and that are provided in Table 1; (iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with CCDC170 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

[0032] In various embodiments of the present disclosure, correlation between a level of one or more epigenetic biomarkers disclosed herein and ESRI expression is determined via Spearman correlation. In some embodiments, a Spearman correlation coefficient of at least 0.2 shows that a level of one or more epigenetic biomarkers is correlated with ESRI expression. In some embodiments, a level of the one or more epigenetic biomarkers at the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci have been shown to be correlated with ESRI expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8). In some embodiments, a level of the one or more epigenetic biomarkers at the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci have been shown to be correlated with ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, ENJ, PIMl, or CCDC170 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8).

[0033] In some embodiments, one or more expression-level correlated loci have been determined to exhibit low or no signal of the one or more epigenetic biomarkers in one or more liquid biopsy samples obtained from one or more healthy subjects (e.g., low or no signal as determined using one or more assays described herein). In some embodiments, one or more13403741vl Page 10 of 173Attorney Docket: 2014191-0051expression-level correlated loci have been determined to exhibit low or no signal of the one or more epigenetic biomarkers in 10% or more of healthy subjects.

[0034] In some embodiments, one or more histone modifications of the present disclosure are quantified using a histone modification assay that measures one or more of H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, H3K4me3, and panacetylation. In some embodiments, a histone modification assay detects H3K4me3 modifications. In some embodiments, a histone modification assay detects H3K27ac modifications. In some embodiments, a histone modification assay is selected from ChlP-seq (Chromatin ImmunoPrecipitation sequencing), CUT& RUN (Cleavage Under Targets and Release Using Nuclease) sequencing, and CUT& Tag (Cleavage Under Targets and Tagmentation) sequencing.

[0035] In some embodiments, chromatin accessibility is quantified using an ATAC-seq (Assay of Transpose Accessible Chromatin sequencing) assay, a NOMe-seq (Nucleosome Occupancy and Methylome sequencing) assay, a FAIRE-seq (Formaldehyde-Assisted Isolation of Regulatory Elements sequencing) assay, an MNase-seq (Micrococcal Nuclease digestion with sequencing) assay, a DNase hypersensitivity assay, or a fragmentomics assay.

[0036] In some embodiments, binding of one or more transcription factors is quantified using a transcription factor binding assay that detects binding of one or more of p300, mediator complex, cohesin complex, RNA pol II, FOXA1, ESRI, PR, MYC, EN1, FOXM1, KLF4, AP-2, RARa, or RUNX1. In some embodiments, a transcription factor binding assay is ChlP-seq (Chromatin ImmunoPrecipitation sequencing), CUT& RUN (Cleavage Under Targets and Release Using Nuclease) sequencing, or CUT& Tag (Cleavage Under Targets and Tagmentation) sequencing.

[0037] In some embodiments, DNA methylation is quantified using Bisulfite sequencing (BS-Seq), Whole Genome Bisulfite Sequencing (WGBS), Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq), or Methyl-CpG-Binding Domain sequencing (MBD-seq).

[0038] In some embodiments, two or more of the epigenetic biomarkers are quantified at one or more genomic loci. In some embodiments, two or more histone modifications are quantified. In some embodiments, H3K4me3 and H3K27ac modifications are quantified. In some embodiments, one or more histone modifications and DNA methylation are quantified. In13403741vl Page 11 of 173Attorney Docket: 2014191-0051some embodiments, H3K4me3 and / or H3K27ac modifications and DNA methylation are quantified. In some embodiments, H3K4me3 modifications, H3K27ac modifications, and DNA methylation are quantified.

[0039] In some embodiments, a method of the present disclosure comprises (i) quantifying H3K4me3 modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “H3K4me3” analyte loci provided in Table 1; (ii) quantifying H3K27ac modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “H3K27ac” analyte loci provided in Table 1; (iii) quantifying DNA methylation at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “MBD” analyte loci provided in Table 1; or (iv) any combination of (i)-(iii).

[0040] In some embodiments, a method of the present disclosure provides comprises enrichment of cfDNA comprising certain histone modifications describe din the present disclosure. In some embodiments, quantifying H3K4me3 modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the genomic loci using an assay that comprises enriching for cfDNA comprising one or more H3K4me3 modifications and sequencing the cfDNA enriched for H3K4me3 modifications (e.g., using a cfChlP-seq assay). In some embodiments, quantifying H3K27ac modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the genomic loci using an assay that comprises enriching for cfDNA comprising one or more H3K27ac modifications and sequencing the cfDNA enriched for H3K27ac modifications (e.g., using a cfChlP-seq assay). In some embodiments, quantifying DNA methylation at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) using an assay that comprises enriching for methylated cfDNA and sequencing the cfDNA enriched for methylated cfDNA (e.g., using a MBD-seq assay).

[0041] In some embodiments, cfDNA comprising H3K4me3 modifications is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody) that binds H3K4me3 modifications. In some embodiments, cfDNA comprising H3K27ac modifications is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody) that binds H3K27ac modifications. In some embodiments, methylated cfDNA is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody or a methyl13403741vl Page 12 of 173Attorney Docket: 2014191-0051binding domain) that binds methylated DNA. In some embodiments, an agent that binds H3K4me3 modifications, an agent that binds H3K27ac modifications, and / or an agent that binds methylated DNA is attached (e.g., via a covalent or noncovalent bond) to a physical support (e.g., a bead, a magnetic bead, an agarose bead, or a magnetic epoxy bead) prior to incubating with the sample. In some embodiments, two or more of (i) an agent that binds H3K4me3 modifications, (ii) the agent that binds H3K27ac modifications, and (iii) an agent that binds methylated DNA are incubated. In some embodiments, a sample is incubated with two or more agents in sequence or in parallel (e.g., wherein the sample is divided into fractions and each fraction is incubated with a different agent).

[0042] In some embodiments, quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from a subject, one or more epigenetic biomarkers comprises sequencing. In some embodiments, sequencing is performed using a next generation sequencing method.

[0043] In some embodiments, DNA sequencing adapters are attached (e.g., covalently attached) to cfDNA obtained from a subject. In some embodiments, DNA sequencing adapters are attached after cfDNA has been enriched for cfDNA comprising one or more H3K4me3 modifications, cfDNA comprising one or more H3K27ac modifications, methylated cfDNA, or any combination thereof. In some embodiments, cfDNA attached to the DNA sequencing adapters is amplified.

[0044] In some embodiments, quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from a subject, one or more epigenetic biomarkers comprises mapping sequence reads to a reference genome. In some embodiments, non-uniquely mapped and redundant sequence reads are discarded.

[0045] In some embodiments, quantifying H3K4me3 modifications, H3K27ac modifications, methylated DNA, or any combination thereof at each of one or more genomic loci comprises summing the number of sequence reads having at least one nucleotide overlap with each of one or more genomic loci. In some embodiments, number of sequence reads at each of the one or more genomic loci is adjusted on the basis of sequencing depth (e.g., quantile normalizing sequence reads to a common reference distribution) and / or ChIP quality, and wherein the adjusting is done prior to or subsequent to summing the number of sequence reads having at least one nucleotide overlap with each of the one or more genomic loci. In some13403741vl Page 13 of 173Attorney Docket: 2014191-0051embodiments, an estimate of local background signal is subtracted from the sequence reads at each genomic loci prior to summing.

[0046] In some embodiments, quantifying H3K4me3 modifications, H3K27ac modifications, methylated DNA, or any combination thereof at each of one or more genomic loci comprises calculating sequence read density at one or more genomic loci. In some embodiments, sequence read density is calculated using a method that comprises: (i) summing background adjusted sequence counts at each of the one or more genomic loci, and (ii) dividing the sum of the background adjusted sequence counts by the combined sum of the length (e.g., the number of nucleotides) of the one or more genomic loci. In some embodiments, sequence read density is calculated using a method that comprises: (i) for each genomic loci, dividing the background adjusted fragment count by the length (e.g., number of nucleotides) of the genomic loci, and (ii) summing the resulting value of (i) for each of the one or more genomic loci. In some embodiments, sequence reads are normalized to aggregate counts in a given sample across a set of regions (e.g., 10,000 regions) previously determined to have DNAse hypersensitivity in most cell types.

[0047] In some embodiments, a method of the present disclosure comprises determining: (i) a point estimate of H3K27ac modifications, (ii) a point estimate of H3K4me3 modifications, (iii) a point estimate of methylated DNA; or (iv) any combination of (i)-(iii). In some embodiments, a point estimate is a mean. In some embodiments, a point estimate is a geometric mean.

[0048] In some embodiments, a method of the present disclosure comprises inputting values obtained by quantifying one or more epigenetic biomarkers into a model. In some embodiments, a model is produced based on (i) a signal of one or more of epigenetic markers at one or more of genomic loci provided in Table 1 measured in one or more cancer cell lines, and / or (ii) a signal of one or more of epigenetic markers at one or more of the genomic loci provided in Table 1 measured in one or more samples obtained from one of more subjects having the cancer.

[0049] In some embodiments, a model is produced based on the signal of one or more of epigenetic markers at one or more of genomic loci provided in Table 1 measured in (i) one or more ER-positive cancer cell lines and (i) one or more ER-negative cell lines (e.g., ER-negative cancer cell lines). In some embodiments, an ER-positive cancer cell line and / or an ER-negative13403741V 1 Page 14 of 173Attorney Docket: 2014191-0051cancer cell line are each breast cancer cell lines.

[0050] In some embodiments, a model is produced based on a signal of one or more of epigenetic markers at one or more of genomic loci provided in Table 1 measured in one or more samples obtained from one of more subjects having an ER-positive cancer and one or more subjects having a ER-negative cancer or one or more healthy subjects.

[0051] In some embodiments, model inputs are point estimates of one or more epigenetic biomarkers. In some embodiments, point estimates are means or geometric means of one or more epigenetic biomarkers.

[0052] In some embodiments, a model is produced by a method that comprises regressing values obtained by quantifying one or more epigenetic biomarkers against measured ER expression. In some embodiments, regressing is performed using an ordinary least squares (OLS) regression method.

[0053] In some embodiments, a quantified amount of the one or more epigenetic biomarkers is compared to a reference.

[0054] In some embodiments, a reference is a predetermined threshold, a measurement from a liquid biopsy sample, a measurement from liquid biopsy samples obtained from a cohort of subjects, and / or a normalized value. In some embodiments, a predetermined threshold and / or the normalized value distinguish ER-positive and ER-negative cancers with an AUC of 0.5 or greater (e.g., 0.7 or greater, 0.8 or greater, 0.9 or greater, 0.95 or greater, or 0.99 or greater).

[0055] In some embodiments, a reference is a measurement from a liquid biopsy sample obtained from a cohort of subjects who have previously been determined to have a ER-positive or a ER-negative cancer.

[0056] In some embodiments, a reference is a measurement from a liquid biopsy sample obtained from a cohort of healthy subjects.

[0057] In some embodiments, a method of the present disclosure determines (i) whether the cancer is ER-positive (ER+) or ER-negative (ER-), (ii) the percentage of cells in the cancer that express ER (e.g., that would stain positive for ER expression using an IHC and / or ISH assay), and / or (iii) an Allred score between 0 and 8 for the cancer.

[0058] In some embodiments, method of the present disclosure comprises obtaining a liquid biopsy sample from a subject. In some embodiments, a liquid biopsy sample is a plasma sample, serum sample, or urine sample. In some embodiments, DNA (e.g., cfDNA) is purified13403741vl Page 15 of 173Attorney Docket: 2014191-0051from about 1 mL, about 2 mL, about 3 mL, about 4 mL, or about 5 mL of a liquid biopsy sample (e g., a plasma sample).

[0059] In some embodiments, a method of the present disclosure comprises treating a subject having a cancer, the method comprising: administering a cancer therapy to a subject based on the ER status of the cancer, wherein the ER status of the cancer has been determined by quantifying one or more epigenetic biomarkers comprising (i) one or more histone modifications, (ii) chromatin accessibility, (iii) binding of one or more transcription factors, and / or (iv) DNA methylation, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from a subject. In some embodiments, if a cancer is determined to be ER-positive, a method comprises administering a therapy for treating a ERpositive cancer, and (ii) if a cancer is determined to be ER-negative, a method comprises administering a therapy for treating an ER-negative cancer. In some embodiments, a therapy for treating an ER-positive cancer comprises administering an ER-targeted agent.

[0060] In some embodiments, a method of the present disclosure comprises monitoring the ER status of a cancer in a subject. In some embodiments, monitoring the ER status of a cancer further comprises treating the cancer. In some embodiments, monitoring the ER status of a cancer comprises determining ER status of the cancer by quantifying the presence of histone modifications and / or DNA methylation at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample, c.., a plasma sample obtained or derived from the subject at a first and a second time point. In some embodiments, a subject has been administered a cancer therapy before the first time point or at the first time point, or wherein the subject has been administered a cancer therapy after the first time point and before the second time point.

[0061] In some embodiments, monitoring the ER status of a cancer in a subject further comprises administering a cancer therapy to a subject based on an ER status of the cancer at a second time point and / or a change in ER status between a first time point and a second time point. In some embodiments, dose and / or frequency of administration of a cancer therapy is adjusted based on an ER status of the cancer at a second time point and / or the change in ER status between a first time point and a second time point. In some embodiments, if a cancer is ER-positive at a second time point, a method of the present disclosure comprises administering an ER-targeted therapeutic to a subject, and (ii) if a cancer is ER-negative at a second time point, a method does not comprise administering an ER-targeted therapeutic to a subject.13403741vl Page 16 of 173Attorney Docket: 2014191-0051

[0062] In one aspect, the present disclosure provides a kit comprising reagents for quantifying one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci, wherein one or more genomic loci are selected from Table 1.

[0063] In some embodiments, a kit comprises reagents for quantifying: (i) H3K4me3 modifications, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 or more genomic loci in Table 1; (ii) H3K27ac modifications, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1; (iii) DNA methylation, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1; or (iv) any combination of (i)-(iii);. In some embodiments, reagents for quantifying comprise one or more reagents for enriching for one or more genomic loci for which H3K4me3 modifications, H3K27ac modifications, and / or DNA methylation is being quantified (e g., reagents for selectively amplifying (e.g., PCR primers) or preferentially binding (e.g., using nucleic acid bait molecules) the one or more genomic loci).

[0064] In some embodiments, a kit comprises one or more antibodies for use in ChlP-seq. In some embodiments, a one or more antibodies specifically bind H3K4me3- or H3K27ac-modified histones. In some embodiments, a kit comprises one or more methyl-binding domains for use in MBD-seq. In some embodiments, a kit comprises one or more antibodies that bind methylated DNA for use in MeDIP-seq. In some embodiments, a kit comprises reagents for isolation of cell-free DNA (cfDNA) from a liquid biopsy sample. In some embodiments, a kit comprises reagents for library preparation for sequencing. In some embodiments, a kit comprises reagents for sequencing. In some embodiments, a kit comprises instructions for determining if a subject has an ER-positive cancer.

[0065] In one aspect, the present disclosure provides a non-transitory computer readable storage medium encoded with a computer program, wherein the program comprises instructions that when executed by one or more processors cause the one or more processors to perform operations to perform a method of the present disclosure, e.g., quantifying of one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci, wherein one or more genomic loci are selected from Table 1.

[0066] In one aspect, the present disclosure provides a computer system comprising a13403741vl Page 17 of 173Attorney Docket: 2014191-0051memory and one or more processors coupled to the memory, wherein the one or more processors are configured to perform operations to perform a method of the present disclosure, e.g., quantifying of one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci, wherein one or more genomic loci are selected from Table 1.

[0067] In one aspect, the present disclosure provides a system for determining the ER status of a cancer in a subject. In some embodiments, a system comprises a sequencer configured to generate a sequencing dataset from a sample; and a non-transitory computer readable storage medium and / or a computer system. In some embodiments, a system comprises a sample preparation device configured to prepare the sample for sequencing from a biological sample, optionally a liquid biopsy sample. In some embodiments, a sequencer is configured to generate a Whole Genome Sequencing (WGS) dataset from the sample.BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The present teachings described herein will be more fully understood from the following description of various illustrative embodiments, when read together with the accompanying drawings. It should be understood that the drawing described below is for illustration purposes only and is not intended to limit the scope of the present teachings in any way. The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and may be better understood by referring to the following description taken in conjunction with the accompanying drawings.

[0069] FIG. 1 illustrates a method of making a preliminary prediction of expression level for a target gene (e.g., ESRI) for an indication (e.g., breast cancer), according to illustrative embodiments of the present disclosure.

[0070] FIG. 2 is a block diagram of an example network environment for use in the methods and systems described herein, according to illustrative embodiments of the present disclosure.

[0071] FIG. 3 is a block diagram of an example computing device and an example mobile computing device, for use in illustrative embodiments of the present disclosure.

[0072] FIG. 4 is a series of graphs showing performance of a model by refining Locus Expression Models Improves Accuracy and Lowers ctDNA Detection Limits. FIG. 4A shows13403741vl Page 18 of 173Attorney Docket: 2014191-0051AUC curves for the classification of ER status using plasma samples. FIG. 4B shows AUC values at different ctDNA fractions for the classification of ER status. “Multigene” refers to an ER classifier that was improved through incorporation of loci from additional genes (i.e., in addition to ESRl). “ISP” stands for in silico plasma and refers to simulated plasma samples generated by diluting in silico plasma sequencing data from cancer patients with plasma sequencing data from healthy patients.

[0073] FIG. 5 A depicts comparison of performance of ER status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation) based on genomic loci near ESRI (a single gene model). Comparison was made by plotting area under the curve (AUC) for each model based on specificity and sensitivity of models.

[0074] FIG. 5B depicts comparison of performance of ER status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation) based on genomic loci near multiple genes associated with expression of ESRI (a multigene model). Comparison was made by plotting area under the curve (AUC) for each model based on specificity and sensitivity of models.

[0075] FIG. 6 depicts comparison of performance of ER status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation) based on genomic loci near ESRI (a single gene model) or genomic loci near multiple genes associated with expression of ESRI (a multigene model). Model performance was evaluated at ctDNA concentrations ranging from 10% to 0.5% in diluted in silico plasma samples. Comparison was made by determining an area under the curve (AUC) for each model based on specificity and sensitivity of models.DETAILED DESCRIPTION

[0076] The present disclosure is based, at least in part, on the demonstration that the ER status of a cancer in a subject can be determined by detecting and quantifying the presence of13403741vl Page 19 of 173Attorney Docket: 2014191-0051histone modifications and / or DNA methylation at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample, e.g., a plasma sample obtained or derived from the subject. The present disclosure also encompasses methods where chromatin accessibility and / or binding of one or more transcription factors are detected at the one or more genomic loci instead of (or in addition to) histone modifications and / or DNA methylation. The present disclosure is also based, at least in part, on the demonstration that measurements of one or more epigenetic biomarkers (e.g., histone methylation marks such as H3K4me3 and histone acetylation marks such as H3K27ac and / or DNA methylation) at one or more expression correlated loci (e.g., genomic loci at which the amount of one or more epigenetic biomarkers is correlated with ESRI expression or / / / / / / / -related gene expression) can be combined into multimodal classifiers to determine ER status. These new monomodal and multimodal classifiers provide minimally invasive ways of determining ER status that are more accurate, objective, and comprehensive than the current tissue-based approaches. No liquid biopsy platform to date has been able to provide actionable resolution on a transcriptionally regulated phenotype relevant for therapy such as ER status.ER status and cancer

[0077] Estrogens are steroidal hormones that function as the primary female sex hormone. There are three major forms of estrogen, namely estrone (El), estradiol (E2) and estriol (E3). Estradiol (E2) is the predominant estrogen in nonpregnant females, while estrone (El) and estriol (E3) are primarily produced during pregnancy and following the onset of menopause, respectively. All estrogens are produced from androgens through actions of enzymes such as aromatase. Follicle-stimulating hormone and luteinizing hormone stimulate the synthesis of estrogen in the ovaries. However, some estrogens are also produced in smaller amounts by other tissues such as the liver, adrenal glands, and mammary gland. Studies have shown that estrogen is associated with mammary turn ori genesis, ovarian and endometrial carcinogenesis (Folkerd and Dowsett, J Clin Oncol (2010) 28:4038-4044). Also, mounting evidence suggests that estrogen and its target gene encoding progesterone receptor (PR) play critical roles in regulating breast cancer progression (Knutson et al., J Hematol Oncol (2017) 10:89).

[0078] The biological effects of estrogen are mostly mediated by its binding and activation of ERa and ERp, which are members of the nuclear receptor superfamily of13403741vl Page 20 of 173Attorney Docket: 2014191-0051transcription factors that are characterized by highly conserved DNA- and ligand-binding domains (Wang etal., JHematol Oncol (2017) 10:168). The DNA binding domain, which is extremely well conserved between ERa and ER0 (97% homology), contains two functionally distinct zinc finger motifs that are responsible for specific DNA binding, as well as mediating receptor dimerization (Hewitt and Korach, Endocr Rev (2018) 39(5):664-675). The unliganded ER has been shown to be present in a cytosolic complex with hsp90 and associated proteins, with ligand binding allowing dissociation from the hsp90 complex, receptor dimerization, nuclear localization and binding to estrogen response elements (EREs) in promoters of estrogen-regulated genes (Pratt and Toft, Endocr Rev (1997) 18:306-360). Genome-wide chromatin immunoprecipitation studies have confirmed that the majority of ER-binding sites in estrogen responsive genes conform well to this consensus sequence (Welboren et al., EMBO J (2009) 28: 1418-1428). While ERa and ERp can bind to most ERE identically, the differences in ERa and ER may lead to tethering differential transcription factors and then modulating different target genes. Thus, the activation of ERa or ERP can produce both unique and overlapping effects.

[0079] ERa has also been shown to modulate gene transcription through heterodimerizing with other transcription factors such as activating protein 1 (API) and nuclear factor kappa-light-chain-enhancer of activated B cells (NF-kB). There is a large profile of estrogen-responsive genes, including pS2, cathepsin D, c-fos, c-jun, c-myc, TGF-a, retinoic acid receptor al, efp, progesterone receptor (PR), insulin-like growth factor 1 (IGF1) (Ikeda et al., Acta Pharmacol Sin (2015) 36:24-31). Many of these ER-regulated genes, including IGF1, cyclin DI, c-myc, and efp, are important for cell proliferation and survival. C-myc is a bona-fide oncogene that is amplified or overexpressed in a variety of human tumors. Efp is an ubiquitin ligase that promotes proteasomal degradation of 14-3-3 sigma thereby stimulating cellular proliferation. While PR is an estrogen-responsive gene, it may antagonize ERa action to inhibit tumor growth, particularly through interacting with RNA polymerase III and inhibiting tRNA transcription.

[0080] A pool of ERa are located in the plasma membrane and cytoplasm (Adlanmerini et al., Proc Natl Acad Sci USA (2014) 111: E283-290), where it binds to diverse membrane or cytoplasmic signaling molecules such as the p85 regulatory subunit of class I phosphoinositide 3-kinase, mitogen-activated protein kinase (MAPK) and Src (Omarjee et al., Oncogene (2017)13403741vl Page 21 of 173Attorney Docket: 2014191-005136:2503-2514). Activation of these signal transduction pathways by estrogen initiates cell survival and proliferation signals. Additionally, these signaling molecules are able to phosphorylate the ERa and its co-regulators to augment nuclear ERa signaling (Arnal et al., Physiol Rev (2017) 97:1045-1087). The genomic and non-genomic actions of ERa play a crucial role in breast epithelial cell proliferation and survival, as well as mammary tumorigenesis

[0028] , The purpose of this review is to decipher the complex mechanisms underlying the aberrant expression of ERa and ERp in human cancer.

[0081] Based on the ER status, breast tumors can be classified as ER-positive and ERnegative. About 75% of breast cancer cases are ERa positive at diagnosis (Allred et al., Breast Cancer Res (2004) 6:240-245). To determine if a cancer is ER-positive, medical practitioners currently order testing that is conducted on a tissue sample using immunohistochemistry (IHC). Samples are reviewed by a pathologist and typically reported as (a) the word positive or negative, (b) a percentage that tells you how many cells out of 100 stained positive for hormone receptors, i.e., a number between 0% (none have receptors) and 100% (all have receptors), and / or (c) an Allred score between 0 and 8. The Allred scoring system looks at what percentage of cells test positive for hormone receptors, along with how well the receptors show up after staining, called intensity (Allred et al., Breast Cancer Res (2004) 6:240-245). This information is then combined to score the sample on a scale from 0 to 8 where, the higher the score, the more receptors were found and the easier they were to see in the sample. The terms “ER-positive” and “ER-negative” as used herein can correspond to any of these tissue based approaches for determining ER status.

[0082] ER-positive cancers can be treated with ER-targeted agents that lower estrogen levels or block estrogen receptors. ER-positive cancers tend to grow more slowly than those that are ER-negative. Women with hormone receptor-positive breast cancers tend to have a better outlook in the short-term, but these cancers can sometimes come back many years after treatment.

[0083] Treatment with ER-targeted agents is not helpful for ER-negative cancers. These cancers may instead be treated with one or more of surgery and / or radiation, HER2 -targeted therapy (if HER2-positive), chemotherapy and immunotherapy. These cancers tend to grow faster than ER-positive cancers. If they come back after treatment, it is often in the first few years. ER-negative breast cancers are more common in women who have not yet gone through13403741vl Page 22 of 173Attorney Docket: 2014191-0051menopause.ER-targeted agents

[0084] The introduction of ER-targeted agents has dramatically influenced the outcome of patients with ER-positive breast cancers. ER-targeted agents block or degrade estrogen receptors or lower estrogen levels. Many ER-targeted agents have already been approved and others are in development or being tested in clinical trials for ER-positive breast cancer and other ER-positive cancers.Agents that block or degrade estrogen receptors (ER)

[0085] In some embodiments, an ER-targeted agent is an agent that blocks or degrades ER. These agents stop estrogen from fueling breast cancer cells to grow. These agents work by preventing estrogen from activating estrogen receptors. They do this by blocking estrogen from binding to estrogen receptors or by degrading estrogen receptors. The former are called Selective Estrogen Receptor Modulators (SERMs) while the latter are called Selective Estrogen Receptor Degraders (SERDs).Selective Estrogen Receptor Modulators (SERMs)

[0086] In some embodiments, an ER-targeted agent is a SERM. SERMs bind estrogen receptors and block them from binding to estrogen. These agents are pills, taken orally.Tamoxifen

[0087] Tamoxifen is a SERM that can be used to treat women with breast cancer who have or have not gone through menopause. This agent can be used in several ways. In women at high risk of breast cancer, tamoxifen can be used to help lower the risk of developing breast cancer.

[0088] For women who have been treated with breast-conserving surgery for ductal carcinoma in situ (DCIS) that is ER-positive, taking tamoxifen for 5 years lowers the chance of the DCIS coming back in the same breast. It also lowers the chance of getting an invasive breast cancer or another DCIS in both breasts.

[0089] For women with ER-positive invasive breast cancer treated with surgery, tamoxifen can help lower the chances of the cancer coming back and improve the chances of13403741v 1 Page 23 of 173Attorney Docket: 2014191-0051living longer. It can also lower the risk of a new cancer developing in the other breast. Tamoxifen can be started either after (adjuvant) or before (neoadjuvant) surgery. When given after surgery, it is usually taken for 5 to 10 years. This drug is used mainly for women with early-stage breast cancer who have not yet gone through menopause. If the subject has gone through menopause, aromatase inhibitors (see below) are often used instead.

[0090] For women with ER-positive breast cancer that has spread to other parts of the body, tamoxifen can often help slow or stop the growth of the cancer and might even shrink some tumors.Toremifene

[0091] Toremifene is a SERM that works in a similar way to tamoxifen, but it is used less often and is only approved to treat post-menopausal women with metastatic breast cancer. It is not likely to work if tamoxifen has already been used and has stopped working.Selective estrogen receptor degraders (SERDs)

[0092] In some embodiments, an ER targeted agent is a SERD. Like SERMs, these agents bind estrogen receptors but do so in a manner that causes them to be degraded. SERDs are used most often in post-menopausal women. When given to pre-menopausal women, they need to be combined with a luteinizing-hormone releasing hormone (LHRH) agonist to turn off the ovaries.Fulvestrant

[0093] In some embodiments, an ER-targeted agent is fulvestrate. In some embodiments, fulvestrant can be used (i) alone to treat advanced breast cancer that has not been treated with other hormone therapy, (ii) alone to treat advanced breast cancer after other hormone drugs (like tamoxifen and often an aromatase inhibitor) have stopped working, or (iii) in combination with a CDK 4 / 6 inhibitor or PI3K inhibitor to treat metastatic breast cancer as initial hormone therapy or after other hormone treatments have been tried. It is given as two injections into the buttocks (bottom). For the first month, the two shots are given two weeks apart. After that, they are given once a month.13403741vl Page 24 of 173Attorney Docket: 2014191-0051Elacestrant

[0094] In some embodiments, an ER targeted agent is Elacestrant. Elacestrant can be used to treat advanced, ER-positive, HER2 -negative breast cancer when the cancer cells have an ESRI gene mutation, and the cancer has grown after at least one other type of hormone therapy. Elacestrant is taken daily as pills, orally.Drugs that lower estrogen levels

[0095] In some embodiments, an ER targeted agent is a drug that can low estrogen levels. Because estrogen stimulates ER-positive cancers to grow, lowering the estrogen level can help slow the cancer’s growth or help prevent it from coming back.Aromatase inhibitors (AIs)

[0096] Aromatase inhibitors (AIs) are drugs that stop most estrogen production in the body. Before menopause, most estrogen is made by the ovaries. But in women whose ovaries are not working, either because they have gone through menopause or because of certain treatments, estrogen is still made in body fat by an enzyme called aromatase. AIs work by preventing aromatase from making estrogen.

[0097] These drugs are useful for women who have gone through menopause, although they can also be used in pre-menopausal women when they are combined with ovarian suppression. These AIs are pills taken orally every day to treat breast cancer and include letrozole, anastrozole, and exemestane.Other ER-targeted agents and other cancers

[0098] While the sections above focus on FDA approved ER-targeted agents, many other ER-targeted agents are being developed and / or assessed in clinical trials. It is to be understood that these other ER-targeted agents can also be used in treatment methods of the present disclosure. In addition, while the sections above focus on the treatment of ER-positive breast cancer, many of these ER-targeted agents can also be used to treat other ER-positive cancers, e.g., ovarian or endometrial ER-positive cancers.ESRI Related Genes13403741V 1 Page 25 of 173Attorney Docket: 2014191-0051

[0099] Tn some embodiments, ER expression status can be determined by a method that comprises quantifying one or more epigenetic markers at one or more ESRI expression correlated loci that are proximal to ESRI. In some embodiments, ER expression status can be determined by a method that comprises quantifying one or more epigenetic biomarkers at one or more expression correlated genomic loci that are proximal to one or more ESRI related genes. Expression-correlated genomic loci that are proximal to one or more ESRI related genes are regions at which signal of one or more epigenetic biomarkers is correlated with ESRJ expression and / or expression of the proximal ESRI related genes.

[0100] In some embodiments, ER expression status can be determined by a method that comprises quantifying one or more epigenetic biomarkers at one or more expression-correlated loci that are proximal to 1 or more (e.g., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more 12 or more, 15 or more, or 20 or more) ESRI related genes.

[0101] ESRI related genes may be selected (e.g., identified) in a number of ways. In some embodiments, ESRI related genes include one or more genes whose expression has been shown to be correlated with that of ESRJ. In some embodiments, ESRI related genes can be identified using publicly accessible databases such as The Cancer Genome Atlas (TCGA). In some embodiments, ESRI related genes may be selected based on, for example, genes being in a transcriptional complex with ESRJ, being master regulators for an indication (e.g., breast cancer), being related to a particular pathway relevant for an indication (e.g., breast cancer), or a combination thereof. As an example, FOXA1 and GATA3 may be selected as related genes for ESRJ for being in a same transcriptional complex. As another example or addition to the prior example, ENJ and PIMJ may be selected as related genes to ESRJ because they are master regulators of ER- breast cancer. A number of related genes may be selected, for example at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15. Related genes may be selected by selecting a certain number of genes having highest rank according to a measure of correlation (e.g., as ranked by a data source (e.g., TCGA)), for example, though not only the highest-ranking set of related genes need be used.

[0102] In some embodiments, one or more ESRI related genes include one or more of ENO1 (Enolase 1, ANOL1, MBP-1, or PPH), YBX1 (YB-1, DBPB, MDR-NF1, NSEP-1, CSDA2, BP-8, or CSDB), GATA3 (HDR or HDRS), FOXA1 (HNF3A or TCF3A), HAPLN313403741vl Page 26 of 173Attorney Docket: 2014191-0051(HsT 19883 or EXLD1), EN1 (HME-1, ENDOVESLB, or Hu-En-1), PIM1 (PIM or EC 2.7.11.1), CCDC170 (C6orf97, FLJ23305, or BA282P11.1), or any combination thereof. In some embodiments, a method comprises quantifying one or more epigenetic biomarkers at one or more expression correlated loci that are proximal to ESRI and one or more expression correlated loci that are proximal to one or more of EN01, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.Subjects and Samples

[0103] A sample used in methods and systems provided herein can be derived from any biological sample including any processed sample that includes circulating tumor DNA (ctDNA) derived from a biological sample. In various embodiments, a sample analyzed using methods and systems provided herein can be derived from a sample obtained from a mammalian subject. In various embodiments, a sample analyzed using methods and systems provided herein can be derived from a sample obtained from a human subject.

[0104] In various instances, a human subject is a subject diagnosed or seeking diagnosis as having, diagnosed as, or seeking diagnosis as at risk of having, and / or diagnosed as or seeking diagnosis as at immediate risk of having cancer, e.g., breast cancer, small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC), etc. In various instances, a human subject is a subject identified as needing cancer therapy. In certain instances, a human subject is a subject identified as needing ER status screening by a medical practitioner.

[0105] The subject may not have undergone previous treatments for cancer, such as the treatments recited in this disclosure. In other embodiments, the subject has undergone previous treatments for cancer, such as the treatments recited in this disclosure.

[0106] In various embodiments a subject has one or more biomarkers and / or risk factors for cancer, e.g., breast cancer, small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC), etc. In certain embodiments, a human subject is identified as in need of ER status screening based on an initial cancer diagnosis, e.g., a breast cancer, small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC), etc. diagnosis. In various instances, a human subject is a subject not yet diagnosed as having, not at risk of having, not at immediate risk of having, not diagnosed as having, and / or not seeking diagnosis for a cancer. Genetic factors may also contribute to cancer risk, as evidenced by individuals with a family history of cancer.13403741vl Page 27 of 173Attorney Docket: 2014191-0051

[0107] Tn various embodiments, a sample from a subject, e.g, a human can be obtained from a liquid biopsy. In certain embodiments, a sample and / or reference is obtained from serum, plasma, or urine. In certain embodiments, the sample is serum. In certain embodiments, a sample includes circulating tumor DNA (ctDNA). In certain embodiments, a sample is derived from about 1 mb of blood obtained from the subject. In certain embodiments, a sample is derived from about 0.5-2 mb of blood obtained from the subject, e.g, about 0.5 to 1.75 mb, about 0.5 to 1.5 mb, about 0.75 to 1.25 mb or about 0.9 to 1.1 mb of blood.

[0108] In various embodiments, a sample is a sample of cell-free DNA (cfDNA). cfDNA is typically found in human biofluids (e.g, plasma, serum, or urine) in short, double-stranded fragments. The concentration of cfDNA is typically low, but can significantly increase under particular conditions, including without limitation pregnancy, autoimmune disorders, myocardial infarction, and cancer. Circulating tumor DNA (ctDNA) is the component of cell-free DNA specifically derived from cancer cells. ctDNA can be present in human biofluids bound to leukocytes and erythrocytes or not bound to leukocytes and erythrocytes. Various tests for detection of tumor-derived ctDNA are based on detection of genetic or epigenetic modifications that are characteristic of cancer (e.g, of a relevant cancer). Genetic or epigenetic factors characteristic of cancer can include, without limitation, oncogenic or cancer-associated mutations in tumor-suppressor genes, activated oncogenes, chromosomal disorders, histone modifications (e.g., histone methylation and / or histone acetylation), chromatin accessibility, binding of one or more transcription factors and / or DNA methylation.

[0109] In various embodiments, ctDNA includes less than 30%, less than 20%, or less than 10% of the cfDNA in the liquid biopsy sample obtained from the subject, e.g., less than 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or less than 1% of the cfDNA in the sample. In some embodiments, the percentage of ctDNA in the liquid biopsy sample is assessed using ichorCNA which estimates the percentage of ctDNA in a sample probabilistically (see Adalsteinsson et al., Nat Commun (2017) 8(1): 1324 the entire contents of which are incorporated herein by reference).

[0110] cfDNA and ctDNA can provide a real-time or nearly real time metric of status of a source tissue. cfDNA and ctDNA demonstrate a half-life in blood of about 2 hours, such that a sample taken at a given time provides a relatively timely reflection of the status of a source tissue.13403741vl Page 28 of 173Attorney Docket: 2014191-0051

[0111] Various methods of isolating nucleic acids from a sample (e.g., of isolating cfDNAfrom blood or plasma) are known in the art. Nucleic acids can be isolated using, without limitation, standard DNA purification techniques, by direct gene capture (e.g., by clarification of a sample to remove assay-inhibiting agents and capturing a target nucleic acid, if present, from the clarified sample with a capture agent to produce a capture complex and isolating the capture complex to recover the target nucleic acid).

[0112] Reagents and protocols for obtaining and analyzing cfDNA and ctDNA, such as circulating in blood or other tissue, are commercially available as described in the Examples and well-known in the art (see, for example, Anker et al., Cancer and Metastasis Rev (1999) 18:65-73; Wua et al., Clin Chim Acta (2002) 321:77-87; Fiegl et al., Cancer Res (2005) 15:1141-1145; Pathak et al., Clin Chem (2006) 52:1833-1842; Schwarzenbach et al., Clin Cancer Res (2009) 15:1032-1038; Schwarzenbach et al., Nat Rev Cancer (2011) 11:426-437) the contents of each of which is separately incorporated herein by reference in their entirety).

[0113] In various embodiments, samples can be collected from individuals repeatedly over a period of time (e.g., once daily, weekly, monthly, annually, biannually, etc.). In various embodiments, such samples can be used to verify results from earlier detections and / or to identify an alteration in biological pattern because of, for example, disease progression, resistance to therapy, treatment, remission, and the like. For example, subject samples can be taken and monitored every month, every two months, or combinations of one, two, or three-month intervals according to the present disclosure. In various embodiments, samples can be collected for monitoring over time beginning at or at certain clinically determined stages, such as at resistance to a therapy, before radiographic progression, after radiographic progression, and / or at tissue biopsy. In addition, results from samples obtained at different points in time can be conveniently compared with each other, as well as with those of normal controls during the monitoring period, thereby providing the subject’s own values, as an internal, or personal, control for long-term monitoring.

[0114] Samples include materials prepared by processes including, without limitation, steps such as concentration, dilution, adjustment of pH, removal of high abundance polypeptides (e.g., albumin, gamma globulin, and transferrin, etc.), addition of preservatives, addition of calibrants, addition of protease inhibitors, addition of denaturants, desalting, concentration and / or extraction of sample nucleic acids, and / or amplification of sample nucleic acids e.g., by PCR or13403741v1 Page 29 of 173Attorney Docket: 2014191-0051other nucleic acid amplification techniques). Samples also include materials prepared by techniques that isolate, e.g., nucleosomes or transcription factors and / or nucleic acids associated with nucleosomes or transcription factors.

[0115] Removal from a sample of proteins that are not desirable for a relevant purpose or context (e.g, high abundance, uninformative, or undetectable proteins) can be achieved using high affinity reagents, high molecular weight filters, ultracentrifugation and / or electrodialysis. High affinity reagents include antibodies or other reagents (e.g., aptamers) that selectively bind to high abundance proteins. Sample preparation can also include ion exchange chromatography, metal ion affinity chromatography, gel filtration, hydrophobic chromatography, chromatofocusing, adsorption chromatography, isoelectric focusing and related techniques. Molecular weight filters include membranes that separate molecules based on size and molecular weight. Such filters may further employ reverse osmosis, nanofiltration, ultrafiltration and microfiltration. Ultracentrifugation is the centrifugation of a sample at about 15,000-60,000 rpm while monitoring with an optical system the sedimentation (or lack thereof) of particles.Electrodialysis is a procedure which uses an electromembrane or semipermeable membrane in a process in which ions are transported through semi-permeable membranes from one solution to another under the influence of a potential gradient. Since the membranes used in electrodialysis may have the ability to selectively transport ions having positive or negative charge, reject ions of the opposite charge, or to allow species to migrate through a semipermeable membrane based on size and charge, it renders electrodialysis useful for concentration, removal, or separation of electrolytes.

[0116] Separation and purification in the present disclosure may include any procedure known in the art, such as capillary electrophoresis (e.g., in capillary or on-chip) or chromatography (e.g, in capillary, column or on a chip). Electrophoresis is a method that can be used to separate ionic molecules under the influence of an electric field. Electrophoresis can be conducted in a gel, capillary, or in a microchannel on a chip. Examples of gels used for electrophoresis include starch, acrylamide, polyethylene oxides, agarose, or combinations thereof. A gel can be modified by its cross-linking, addition of detergents, or denaturants, immobilization of enzymes or antibodies (affinity electrophoresis) or substrates (zymography) and incorporation of a pH gradient. Examples of capillaries used for electrophoresis include capillaries that interface with an electrospray.13403741V 1 Page 30 of 173Attorney Docket: 2014191-0051

[0117] Capillary electrophoresis (CE) is preferred for separating complex hydrophilic molecules and highly charged solutes. CE technology can also be implemented on microfluidic chips. Depending on the types of capillary and buffers used, CE can be further segmented into separation techniques such as capillary zone electrophoresis (CZE), capillary isoelectric focusing (CIEF), capillary isotachophoresis (CITP) and capillary electrochromatography (CEC). An embodiment to couple CE techniques to electrospray ionization involves the use of volatile solutions, for example, aqueous mixtures containing a volatile acid and / or base and an organic such as an alcohol or acetonitrile.

[0118] Capillary isotachophoresis (CITP) is a technique in which the analytes move through the capillary at a constant speed but are nevertheless separated by their respective mobilities. Capillary zone electrophoresis (CZE), also known as free-solution CE (FSCE), is based on differences in the electrophoretic mobility of the analytes, determined by the charge on the analytes, and the frictional resistance the analytes encounter during migration, which is often directly proportional to the size of the analytes. Capillary isoelectric focusing (CIEF) allows weakly-ionizable amphoteric molecules, to be separated by electrophoresis in a pH gradient. CEC is a hybrid technique between traditional high performance liquid chromatography (HPLC) and CE.

[0119] Separation and purification techniques used in the present disclosure can include any chromatography procedures known in the art. Chromatography can be based on the differential adsorption and elution of certain analytes or partitioning of analytes between mobile and stationary phases. Different examples of chromatography include, but not limited to, liquid chromatography (LC), gas chromatography (GC), high performance liquid chromatography (HPLC), etc.

[0120] In some embodiments, whole blood is collected from a subject, and a plasma layer is separated by centrifugation. cfDNA may be then extracted from the plasma using methods known in the art.Histone Modifications, Chromatin Accessibility and Transcription Factor Binding

[0121] Histone methylation is understood to increase or decrease expression of associated coding sequences, depending on which histone residue is methylated. Histone methylation is an essential modification that can cause monomethylation (mel), dimethylation13403741vl Page 31 of 173Attorney Docket: 2014191-0051(me2), and trimethylation (me3) of several amino acids, thus directly affecting heterochromatin formation, gene imprinting, X chromosome inactivation, and gene transcriptional regulation. Histone methyltransferases promote monomethylation, dimethylation, or trimethylation of histones while histone demethylases promote demethylation of histones. In general, lysine (Lys or K), arginine (Arg or R), and rarely histidine (His or H) are the most common histone methyl acceptors. Histone methylation only occurs at specific lysine and arginine sites of histone H3 and H4. In histone H3, lysine 4, 9, 26, 27, 36, 56, and 79 and arginine 2, 8, and 17 can be methylated. By comparison, histone H4 has fewer methylation sites, in which only lysine 5, 12, and 20 and arginine 3 can be methylated. Histone methylation is often associated with transcriptional activation or inhibition of downstream genes. The methylation of histone H3K4, R8, R17, K26, K36, K79, H4R3, and K 12 can activate gene transcription. However, the methylation of histone H3K9, K27, K56, H4K5, and K20 can inhibit gene transcription. Thus, for example, H3K4 methylation generally activates gene expression, while H3K27 methylation generally represses gene expression.

[0122] Histone acetylation occurs predominantly at lysine residues and is generally understood to increase expression of associated coding sequences. Without wishing to be bound by any theory, acetylation of lysine residues is thought to neutralize lysine’s positive charge and thereby cause histones to drift away from DNA, which has a negative charge. The released structure facilitates access to transcriptional machinery such as transcription factors and RNA polymerase II. Histone acetylation and deacetylation are generally catalyzed by histone acetyltransferases (HATs) and HDACs, respectively. Acetyl-CoA can be a source and co-factor of acetylation. In regulatory regions, HATs can acetylate histones and recruit HAT-containing complexes to activate the transcriptional process. For instance, H3K9ac and H3K27ac levels can be associated with promoter and enhancer activities. Furthermore, H3K27ac enhances not only the kinetics of transcriptional activation, but also accelerates the transition of RNA polymerase II from the initiation state to the elongation state.

[0123] Differential modification of a genomic locus (e.g., differential histone methylation and / or differential histone acetylation) can refer to, or be determined by or detected as, a comparative difference or change in modification status of one or more genomic loci between a first sample, condition, disease, or state and a second or reference sample, condition, disease, or state. Those of skill in the art will appreciate that a reference is typically produced by13403741vl Page 32 of 173Attorney Docket: 2014191-0051measurement using a methodology identical, similar, or comparable to that by which a compared non-reference measurement was taken.

[0124] Chromatin accessibility can refer to the degree to which nuclear macromolecules are able to physically contact DNA and is determined in part by the occupancy and modification status of nucleosomes. Modified histones can regulate chromatin accessibility through a variety of mechanisms, such as altering transcription factor (TF) binding through steric hindrance and modulating nucleosome affinity for active chromatin remodelers. The topological organization of nucleosomes across the genome is non-uniform: while histones can be densely arranged within facultative and constitutive heterochromatin, histones can be depleted at regulatory loci, including within enhancers, insulators and transcribed gene bodies. Active regulatory elements of the genome are generally accessible.

[0125] Differential accessibility of a genomic locus can refer to, or be determined by or detected as, a comparative difference or change in modification status of one or more genomic loci between a first sample, condition, disease, or state and a second or reference sample, condition, disease, or state. Those of skill in the art will appreciate that a reference is typically produced by measurement using a methodology identical, similar, or comparable to that by which a compared non-reference measurement was taken.

[0126] A reference can be a value or set of values that are predetermined or derived from a sample or set of samples. A reference can be a sample or set of samples. A reference value can be a predetermined threshold value, a value that varies in accordance with circumstances (e.g., according to patient subpopulation, age, weight, or other variables), or a ratio. Reference ratios can be ratios relating to the modification and / or accessibility of multiple loci within individual samples and / or references, or across or between samples and / or references. In various embodiments, a reference can have or represent a normal, non-diseased state. In some embodiments, such as for staging of disease or for evaluating the efficacy of treatment, a reference can have or represent a diseased state, e.g., a cancer, stage of cancer, or subtype of cancer. In some embodiments, a reference can represent a particular level of target gene (e.g., ER) expression based on IHC testing, e.g., ER-positive or ER-negative cancer.

[0127] In certain instances, a reference is a non-contemporaneous sample from the same source, e.g., a prior sample from the same source, e.g., from the same subject. In certain instances, a reference for the modification status of one or more genomic loci (e.g., one or more13403741vl Page 33 of 173Attorney Docket: 2014191-0051expression correlated genomic loci) can be the modification status of the one or more genomic loci (e.g., one or more expression correlated genomic loci) in a sample (e.g., a sample from a subject), or a plurality of samples, known to represent a particular state (e.g., ER-positive or ERnegative cancer). In certain instances, a reference for the accessibility status of one or more genomic loci (e.g., one or more differentially accessible genomic loci) can be the accessibility status of the one or more genomic loci (e.g., one or more differentially accessible genomic loci) in a sample (e.g., a sample from a subject), or a plurality of samples, known to represent a particular state (e.g., Er-positive or ER-negative cancer).

[0128] In some illustrative but non-limiting embodiments of the present disclosure differential modification or differential accessibility can refer to a differential (e.g., between a sample and a reference) with an absolute log2(fold-change) that is greater than or equal to 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0 or more, or any range in between, inclusive, e.g., as measured according to an assay provided herein.

[0129] Enhancers are genomic loci that can be differentially modified or differentially accessible in and / or between conditions, diseases, and other states. Enhancers are cis-acting DNA regulatory regions that are thought to bind trans-acting proteins that contribute to expression patterns of associated genes. Chromatin ImmunoPrecipitation sequencing (ChlP-seq) of histone modifications (e.g., acetylation) have identified millions of enhancers in mammalian genomes. The number of active enhancers in any given cell type is estimated to be in the tens of thousands. Certain transcription factors (TFs), sometimes referred to as “master” transcription factors, associate with active enhancers with important impacts on gene expression and cell function. Certain such transcription factors preferentially associate with enhancers that regulate genes required for establishing cell identity and function, including enhancer domains known as “super-enhancers”. Moreover, master TFs can participate in inter-connected auto-regulatory circuitries or “cliques” that are self-reinforcing, show marked cell selectivity, and function to maintain cell state and / or cell survival.Techniques for Detecting and Quantifying Histone Modifications and Transcription Factor Binding

[0130] Various techniques of molecular biology are well known in the art and / or disclosed in the present application for detecting and quantifying histone modifications and / or13403741vl Page 34 of 173Attorney Docket: 2014191-0051transcription factor binding. In some embodiments, the methods, kits and systems of present disclosure involve the detection and quantification of histone modifications and / or transcription factor binding in samples, e.g., in liquid biopsy samples including cfDNA such as plasma samples including cfDNA. Chromatin ImmunoPrecipitation (ChIP) is one technique of molecular biology useful in detecting and quantifying histone modifications and transcription factor binding in samples. CUT& RUN or CUT& Tag are other more recent techniques that can also be used to detect and quantify histone modifications and transcription factor binding sites. ChlP-chip, ChlP-exo, ChIP Re-ChIP, and ChlPmentation are other alternative techniques that could be used.

[0131] ChIP can involve various steps including one or more of fixation, sonication, immunoprecipitation, and analysis of the immunoprecipitated DNA. ChIP has become a very widely used tissue-based technique for determining the in vivo location of binding sites of various transcription factors and histones. Because the proteins are captured at the sites of their binding with DNA, ChIP helps to detect DNA-protein interactions that take place in living cells. More importantly, ChIP can be coupled to many commonly used molecular biology techniques such as PCR and real-time PCR, PCR with single-stranded conformational polymorphism, Southern blot analysis, Western blot analysis, cloning, and microarray. The resulting versatility has increased the potential of this technique.

[0132] ChIP of tissue samples usually involves cross-linking of the chromatin-bound proteins by formaldehyde, followed by sonication or nuclease treatment to obtain small DNA fragments. Immunoprecipitation can be then carried out using specific antibodies to the DNA-binding protein of interest. The DNA can be then released from the proteins and analyzed using various methods. ChIP has also been used to study RNA-protein interactions. X-ChIP methods utilize fixed chromatin fragmented by sonication, while the N-ChIP methods utilize native chromatin, which can be unfixed and nuclease digested.

[0133] The first step of the technique can be the cross-linking of DNA and proteins. Formaldehyde is one of the most used cross-linking agents. One advantage of using formaldehyde can be the ease of reversibility of the cross-links and its ability to form bonds that span approximately 2 angstroms. This means that formaldehyde can bind molecules in close association with each other. Generally, formaldehyde can be added to the medium in the cell culture flask or plate. It enters the cells through the cell membrane and cross-links the proteins to13403741vl Page 35 of 173Attorney Docket: 2014191-0051the chromatin. Formaldehyde fixation of tumor tissues has also been done. Other cross-linking agents that have been used include chemicals such as methylene blue and acridine orange, cisplatin, dimethylarsinic acid, potassium chromate, and ultraviolet (UV) light and lasers.

[0134] Harvested chromatin can be sonicated in one or more sonication cycles. DNA can be typically broken into to 100-500 bp fragments to pinpoint the location of the DNA sequence of interest. An alternative to sonication can be nuclease digestion of the chromatin, e.g., in N-ChlP methods. Purification of chromatin can be achieved using a cesium chloride (CsCl) gradient centrifugation.

[0135] Chromatin can be immunoprecipitated using one or more antibodies that bind a target epitope. For example, an antibody used in ChIP can selectively bind a particular transcription factor or one or more particular histone modifications, such as one or more particular histone acetylation modifications or histone methylation modifications. In some embodiments, an antibody used to bind a target epitope can be a “pan” antibody (e.g., a panacetylation antibody, a pan-methylation antibody, an antibody that binds a group of histone modifications associated with increased transcription activation, and / or an antibody that binds a group of histone modifications associated with increased transcription repression). The antibody against the protein of interest is allowed to bind to the protein-DNA complex, and the complex can be then precipitated. Immunosorbants commonly used to separate the antigen-antibody complex from the lysate include salmon sperm DNA-protein A-Sepharose®, protein G, magnetic beads, and other engineered immunoprecipitation systems known to those of skill in the art.

[0136] Immunoprecipitated DNA can be eluted. Once the DNA of interest is isolated, many detection and quantification methods can be used to study the isolated gene fragments. Commonly utilized methods include PCR, real-time PCR, slot blot hybridization, microarray techniques, and deep or next-generation sequencing. ChlP-seq combines chromatin immunoprecipitation (ChIP) with massively parallel DNA sequencing to identify the binding sites of DNA-associated proteins. ChlP-seq can be used to map DNA-binding proteins, e.g., transcription factor binding sites and histone modifications in a genome-wide manner.

[0137] Cell-free Chromatin ImmunoPrecipitation sequencing (cfChlP-seq) involves applying ChlP-seq to samples that include cell-free DNA, e.g., liquid biopsy samples including cfDNA such as plasma samples including cfDNA (e.g., see Sadeh et al., Nat Biotechnol (2021) 39: 586-598 and Jang et al., Life Sci Alliance (2023) 6(12):e202302003 the entire contents of13403741vl Page 36 of 173Attorney Docket: 2014191-0051each of which are incorporated herein by reference). In some embodiments, cfChlP-seq uses antibodies or antibody fragments that bind specific histone modifications (e.g., H3K4me3 and / or H3K27ac) and / or transcription factors that are coupled (covalently or non-covalently) to beads, e.g., magnetic beads such as Dynabeads® magnetic beads and incubated with a volume, e.g., about 1 mL of thawed plasma obtained from a subject. Without limitation, exemplary antibodies that bind H3K4me3 include PA5-27029 (available from Thermo Fisher Scientific in Waltham, MA) and C15410003 (available from Diagenode in Denville, NJ) and exemplary antibodies that bind H3K27ac include ab21623 or ab4729 (both available from Abeam in Cambridge, UK) and Cl 5210016 (available from Diagenode in Denville, NJ).

[0138] In some embodiments, the antibodies or antibody fragments can be covalently coupled to beads, e.g., epoxy beads. In some embodiments, the antibodies or antibody fragments can be non-covalently coupled to beads, e.g., Protein A or Protein G beads such as Dynabeads® Protein A or Dynabeads® Protein G beads. After washing, a cfDNA library is then typically prepared from the captured cfDNA. Library preparation can be done on-bead or after releasing the captured cfDNA by digestion of bound histones, e.g., using proteinase K. The cfDNA library is then sequenced to generate reads of captured cfDNA sequences, e.g., by next-generation sequencing (NGS) as is known in the art. The reads are then analyzed, e.g., aligned and counted using standard bioinformatic techniques as is known in the art. A cfChlP-seq bioinformatic pipeline can include, e.g., alignment of sequence reads to a reference genome with BWA or Bowtie2. Aligned reads can be used to call and quantify peaks as compared to a reference.

[0139] CUT& Tag involves antibody -based binding of a target protein, e.g., transcription factor or histone modification of interest, where antibody incubation is directly followed by the shearing of the chromatin and library preparation (see Kaya-Okur et al., Nat Comm (2019) 10: 1930). CUT& Tag assays take advantage of a Tn5 transposase that is fused with Protein Ato direct the enzyme to the antibody bound to its target on chromatin. Tn5 transposase is pre-loaded with sequencing adapters (generating the assembled pA-Tn5 adapter transposome) to carry out antibody-targeted tagmentation. In a typical CUT& Tag assay samples are incubated with an antibody immobilized on Concanavalin A-coated magnetic beads to facilitate subsequent washing steps. Cells can be incubated with a primary antibody specific for the target protein of interest followed by incubation with a secondary antibody. Samples can then be incubated with assembled transposomes, which consist of Protein A fused to the Tn5 transposase enzyme that is13403741vl Page 37 of 173Attorney Docket: 2014191-0051conjugated to NGS adapters. After incubation, unbound transposome can be washed away using stringent conditions. Tn5 is a Mg2+-dependent enzyme so Mg2+can be added to activate the reaction, which results in the chromatin being cut close to the protein binding site and simultaneous addition of the NGS adapter DNA sequences. Chromatin cleavage and library preparation can be achieved in one single step.

[0140] CUT& RUN is an epigenetic profiling strategy in which antibody-targeted controlled cleavage by micrococcal nuclease releases specific protein-DNA complexes into the supernatant for paired-end DNA sequencing (see Skene and Henikoff, Elife (2017) 6:1-35, Skene et al., Nat Protoc (2018) 13:1006-1019). As only targeted fragments enter into solution, and the vast majority of DNA is left behind, CUT& RUN has low background levels. In an example CUT& RUN assay, a sample is incubated with an antibody or antibody fragment that binds the target protein, e.g., transcription factor or histone modification of interest. The sample is then incubated with Protein-A-MNase after which CaCl₂ can be added to initiate the calcium dependent nuclease activity of MNase to cleave the DNA around the target protein. The protein-A-MNase reaction can be quenched by adding chelating agents (EDTA and EGTA). Cleaved DNA fragments are then liberated, extracted, and used to construct a sequencing library.Techniques for Detecting and Quantifying DNA Methylation

[0141] DNA from a sample may be sequenced in order to obtain sequencing data for the sample. DNA from a sample may be sequenced in order to determine methylation status. Various methylation sequencing techniques may be used to produce methylation sequencing data. In some embodiments, DNA extracted from plasma is enriched for densely methylated fragments as part of a sequencing method, such as, for example, Methyl -CpG-Binding Domain Sequencing (MBD-seq). In some embodiments, after sequencing, FASTQ files are processed to produce a table of unique fragments that align to a genome for a subject from which a sample is derived (e.g., the human genome for human subjects / samples). Sequencing data used to produce a cTF estimation model and / or used to estimate cTF for a sample may be preprocessed to get unique fragments from aligned sequencing reads. In some embodiments, sequencing data received and / or obtained may be initially deduplicated.

[0142] Various techniques of molecular biology are well known in the art and / or disclosed in the present application for detecting and quantifying DNA methylation, for example13403741vl Page 38 of 173Attorney Docket: 2014191-0051of cfDNA (e.g., ctDNA) in a sample. In some embodiments, the methods and systems of the present disclosure involve the detection and quantification of DNA methylation in samples, e.g., in liquid biopsy samples including cfDNA (e.g., ctDNA) such as plasma samples including cfDNA. Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq) and Methyl-CpG-Binding Domain sequencing (MBD-seq) are exemplary techniques of molecular biology useful in detecting and quantifying DNA methylation in samples, though others, such as Bisulfite sequencing (BS-Seq) and Whole Genome Bisulfite Sequencing (WGBS) exist. In some embodiments, an enrichment method, such as MBD-seq, is used to obtain sequencing data. In some embodiments, sequencing data from an enrichment method, such as MBD-seq, is received and processed.

[0143] DNA methylation typically refers to the methylation of the 5’ position of cytosine (mC) by DNA methyltransferases (DNMT). It is a major epigenetic modification in humans and many other species. In mammals, most DNA methylations occur within the context of CpG dinucleotides. DNA methylation is thought to be a repressive chromatin modification. Aberrant methylation can lead to many diseases including cancers (Robertson, Nat Rev Genet (2005) 6:597-610 and Bergman and Cedar, Nat Struct Mol Biol (2013) 20:274-281).

[0144] MeDIP-seq was first reported by Weber et al., Nat Genet (2005) 37:853-862. In a typical MeDIP-seq protocol, antibody or antibody-fragment that binds 5-methylcytidine (5mC) is used to enrich methylated DNA fragments, then these fragments are sequenced and analyzed. If using 5mC-specific antibodies or antibody fragments, methylated DNA is isolated from genomic DNA via immunoprecipitation. Anti-5mC antibodies are incubated with fragmented genomic DNA and precipitated, followed by DNA purification and sequencing.

[0145] Methyl-CpG-Binding Domain sequencing (MBD-seq) is similar to MeDIP-seq except that it uses methyl binding domain (MBD) proteins instead of antibodies or antibody fragments to bind methylated DNA. In a typical MBD-seq protocol, genomic DNA is first sonicated and incubated with tagged MBD proteins that can bind methylated cytosines. The protein-DNA complex is then precipitated with antibody-conjugated beads that are specific to the MBD protein tag, followed by DNA purification and sequencing.Techniques for Detecting and Quantifying Chromatin Accessibility

[0146] Various techniques of molecular biology are well known in the art and / or13403741vl Page 39 of 173Attorney Docket: 2014191-0051disclosed in the present application for detecting and quantifying chromatin accessibility. In some embodiments, the methods, kits and systems of the present disclosure involve the detection and quantification of chromatin accessibility in samples, e.g., in liquid biopsy samples including cfDNA such as plasma samples including cfDNA. ATAC-seq (Assay of Transpose Accessible Chromatin sequencing), NOMe-seq (Nucleosome Occupancy and Methylome sequencing), FAIRE-seq (Formaldehyde-Assisted Isolation of Regulatory Elements sequencing), MNase-seq (Micrococcal Nuclease digestion with sequencing), and DNase hypersensitivity assays are exemplary techniques of molecular biology useful in detecting and quantifying chromatin accessibility in samples. Sono-Seq is another alternative method that could be used (see Auerbach et al., Proc Natl Acad USA (2009) 106(35): 14926-14931). Fragmentomics-based methods are yet another method that can be used to assess chromatin accessibility (see Ding, Spencer C., and YM Dennis Lo. " Cell-free DNA fragmentomics in liquid biopsy." Diagnostics 12.4 (2022): 978).

[0147] DNase hypersensitivity assays can use the non-specific DNA endonuclease Deoxyribonuclease I (DNase I), which selectively digests accessible DNA regions. DNase I hypersensitivity sites (DHS) identified by DNase-seq include open chromatin regulatory regions. A typical DNase hypersensitivity assay can include a first step in which nuclei are isolated from cells using lysis buffer, and nuclei are digested using DNase I. DNA fragment sizes are measured to identify optimal digestion using gel electrophoresis. Biotinylated linkers can be ligated to the ends of digested DNA after polishing to make blunt ends, and the DNA can then be isolated. DNA with biotinylated linker can be digested by restriction endonuclease Mmel and captured by streptavidin coated Dynabeads® to generate short tags to which a second sequencing adaptor can be ligated. A second linker can be ligated and amplified to generate a library for sequencing. A DNase-seq bioinformatic pipeline can include, e.g., alignment of sequence reads to a reference genome with BWA or Bowtie2. Aligned reads can be used to call and quantify peaks as compared to a reference.

[0148] MNase-seq determines chromatin accessibility with micrococcal nuclease (MNase) that preferentially digests nucleosome-free, protein-unbound DNA. Atypical MNase-seq assay can include a first step in which nuclei are isolated from either native or crosslinked chromatin and digested using MNase with titration. In vivo formaldehyde crosslinking step that is designed to capture the interaction between proteins and DNA. This crosslinking allows bound13403741vl Page 40 of 173Attorney Docket: 2014191-0051proteins to shield their associated DNA from digestion by MNase. Following crosslinking, samples are digested with MNase, which can be specifically activated by addition of Ca2+ to the buffer. Digestion can be halted by chelating the reaction, at which point the samples are RNase treated, crosslinks are reversed, and proteins are digested away from the chromatin. DNA can then be isolated via a phenol-chloroform extraction. Uncut DNA is purified and mononucleosome bands are isolated and excised through gel electrophoresis. Isolated DNA can be amplified by adding adapters to generate a library, and sequenced. MNase-seq primarily sequences regions of DNA bound by histones or other proteins. Therefore, it indirectly determines which regions of DNA are accessible by directly determining which regions are bound to nucleosomes or proteins.

[0149] FAIRE-seq is a method in which nucleosome-depleted regions of DNA (NDRs) are isolated from chromatin. Atypical FAIRE-seq assay can include a first step in which cells are fixed using formaldehyde so that histones are crosslinked to interacting DNA. Crosslinked chromatin can then be sheared by sonication that generates protein-free DNA and protein-crosslinked DNA fragments. Protein-free DNA can be isolated using a phenol-chloroform extraction: DNA crosslinked with protein stays in organic phase, while protein-free DNA stays in aqueous phase. Highly crosslinked DNA remains in the organic phase and the non-crosslinked DNA is pulled to the aqueous phase. Non-crosslinked DNA from the aqueous phase can then be amplified and sequenced. Reads enriched in the sequencing pool tend to have lower nucleosome and transcription factor binding and are therefore inferred to come from accessible regions.

[0150] NOMe-seq is a method to identify nucleosome-depleted regions of DNA (NDRs) with M. CviPI methyltransferase that methylates cytosine in GpC dinucleotides not protected by nucleosomes or other proteins. Unlike CmpG, GpCmin the human genome does not occur naturally in most cell types. GpCmlevels at open chromatin regions can be compared to background signals and used to detect and quantify NDRs. Atypical NOMe-seq protocol can include a step in which samples are treated with M. CviPI and S-adenosylhomocysteine (SAM) to methylate accessible GpC sites. M. CviPI treated DNA can be sheared using a sonicator, so that DNA fragments can be sequenced. DNA is treated with bisulfite, which converts unmethylated cytosine to uracil using sodium bisulfite, while methylated cytosine is unaffected. A library is generated using adapters and sequenced. Accessible chromatin is expected to have high levels of GpCmbut low levels of CmpG. Therefore, NOMe-seq identifies NDRs using the two separate13403741vl Page 41 of 173Attorney Docket: 2014191-0051methylation analyses that serve as independent (but opposite) measures, providing matched chromatin designations for each regulatory element.ATAC-seq uses hyperactive Tn5 transposase that preferentially cuts accessible chromatin regions and simultaneously inserts adapters to the fragmented region (Buenrostro et al., Nat Methods (2013) 10(12): 1213-1218 the entirety of which is incorporated herein by reference). Atypical ATAC-seq assay can include a first step in which samples are incubated with Tn5 transposase. DNA can then be isolated and purified. DNA fragmented and tagged by Tn5 transposase can be purified and then amplified to generate a library and sequenced for analysis.

[0151] ‘Fragmentomics” or a “fragmentomics assay” refers to methods that use certain size and sequence characteristics of cfDNAto gain insight into the epigenetic state of cells at the time their genomic DNA was released into the extracellular environment. Without wishing to be bound by theory, upon release of genomic DNA from a cell into the extracellular environment, nucleases rapidly cleave the genomic DNA into short fragments. The cleavage pattern and sequences of the fragments reflect the positioning of nucleosomes genome- wide at the point of cell death, and by finding nucleosomes that are consistently genomically positioned across cancer cells (i.e. many of the circulating tumor DNA fragments that map to that small region of the genome have the same start and end positions or similar fragment length characteristics) fragmentomics attempts to infer the location of stably positioned nucleosomes at regulatory sites, and thus to infer where the active regulatory sites are in a given cell type. Accordingly, analysis of cfDNA fragmentation patterns can be used to infer characteristics of the cells at the time they released genomic DNA. Examples of metrics commonly measured in fragmentomics include fragment size, preferred ends, end motifs, single-stranded jagged ends, and nucleosomal footprints. Approaches for measuring fragmentomics metrics include, e.g., qPCR, electron microscopy, single molecule sequencing, and next-generation sequencing. A relationship between fragmentomic metrics and histone modifications (h3K4me3 and H3K27ac) has been established. See Bai, Jinyue, et al. " Histone modifications of circulating nucleosomes are associated with changes in cell-free DNA fragmentation patterns." Proceedings of the National Academy of Sciences 121.42 (2024): e2404058121; and Wang, Yadong, et al. " Cell-free epigenomes enhanced fragmentomics-based model for early detection of lung cancer." Clinical and Translational Medicine 15.2 (2025): e70225.13403741vl Page 42 of 173Attorney Docket: 2014191-0051Exemplary Genomic Loci

[0152] Among other things, the present disclosure identifies exemplary genomic loci that (i) are differentially modified and / or differentially accessible depending on ER expression status, and shows that differential modifications and / or accessibility of said exemplary genomic loci can be used to accurately predict (e.g., determine) ER expression status. See Table 1, which shows the chromosomal coordinates of each genomic locus.

[0153] In some embodiments, genomic loci described herein have been determined to have certain characteristics that make them particularly advantageous for predicting or determining ER expression status.

[0154] Among other things, genomic loci described herein have been determined to exhibit statistically meaningful signal that can be predictive of expression level of a gene (e.g., ESRJ or one or more ESRI related genes). In general, not every portion of a genomic region comprising a gene (e.g., a genomic region comprising a gene and / or regions proximal to the gene) will produce signal of one or more epigenetic biomarkers that exhibits a statistically meaningful relationship with expression level of the gene (e.g., for at least one of a set of one or more epigenetic biomarkers being considered) and those regions can be ignored (e.g., at least with respect to the at least one of the set of one or more epigenetic biomarkers) when producing a model to predict expression level of the gene. In some embodiments, considering genomic regions or subregions that are too large in predicting expression can result in technical noise or other detrimental effects that can reduce predictive power. Thus, use of subregions of a genomic region that exhibit a statistically meaningful relation with expression level can reduce technical noise or other detrimental effects and therefore improve predictive power (e.g., as compared to use of genomic regions that include regions that do not exhibit a statistically meaningful relationship with expression level).

[0155] In some embodiments, expression-correlated loci described herein exhibit minimal or no signal of one of more epigenetic biomarkers in one or more samples (e.g., liquid biopsy samples) obtained from healthy subjects. In some embodiments, expression-correlated loci described herein have minimal or no signal of one of more epigenetic biomarkers in less than 10% of samples obtained from healthy subjects. Regions that exhibit minimal or no signal of one or more epigenetic biomarkers in one or more samples from one or more healthy subjects13403741V 1 Page 43 of 173Attorney Docket: 2014191-0051can exhibit highly cancer-cell-specific signal for the one or more epigenetic biomarkers.

[0156] The present disclosure is not limited to methods that use the exact same chromosomal coordinates that are recited in Table 1. The present disclosure encompasses methods that use any of the genomic loci in Table 1 and also subregions thereof, i.e., references herein to methods that involve detecting and / or quantifying one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci of Table 1 encompasses methods that detect these marks anywhere within these genomic loci including within any subregions. For example, where Table 1 references chrl:8720924-8721417 as a genomic locus for detecting and / or quantifying H3K27ac modification, this encompasses methods that detect and / or quantify H3K27ac modification at any position or sub-region of chrl:8720924-8721417, e.g., methods that detect and / or quantify H3K27ac modification within chrl:8721024-8721317, etc. In some embodiments, a subregion may span at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500 or at least 3000 contiguous base pairs that are located between the lower and upper coordinates of a genomic locus recited in Table 1. In some embodiments, a subregion may span less than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500 or at least 3000 contiguous base pairs that are located between the lower and upper coordinates of a genomic locus recited in Table 1. In some embodiments, a subregion may have the same central coordinate as a genomic locus recited in Table 1. In some embodiments, a subregion may have a different central coordinate as a genomic locus recited in Table 1. It is also to be understood that the lower / upper coordinates of the genomic loci in Table 1 are approximate and that the present disclosure encompasses methods where any one or more of the genomic loci are expanded by increasing the size of the genomic locus by 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40% or up to 50% in one or both directions.

[0157] In some embodiments a classifier is generated using a set of differentially modified and / or differentially accessible genomic loci that are correlated with increased ER expression. Sequence reads that fall into each selected genomic locus are analyzed and counted, e.g., as described herein, including in the Examples. In some embodiments, counts from genomic loci that are correlated with increased ER expression are aggregated. Other ways of using the genomic loci and related sequencing data to generate and apply a classifier to determine ER expression status are described herein and known in the art, e g., without limitation, methods that13403741vl Page 44 of 173Attorney Docket: 2014191-0051use a learning statistical classifier system or a combination of learning statistical classifier systems.

[0158] In some embodiments, exemplary genomic loci from Table 1 are used in a monomodal ER expression status classifier, e.g., an expression status classifier that uses a single histone modification (e.g., H3K4me3 or H3K27ac) or DNA methylation at one or more genomic loci for purposes of determining ER expression status. In some embodiments, exemplary genomic loci from Table 1 are used in combination in a multimodal classifier, e.g., a ER expression status classifier that uses more than one histone modification (e.g., H3K4me3 and H3K27ac) or one or more histone modifications (e.g., H3K4me3 and / or H3K27ac) and DNA methylation at one or more genomic loci for purposes of determining ERexpression status.

[0159] In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor at one or more loci provided in Table 1. In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor at one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more loci listed in Table 1.

[0160] In some embodiments, one or more genomic loci comprise: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRI (e.g., ESRI associated loci provided in Table 1). In some embodiments, one or more genomic loci comprise: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for one or more of ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, (e.g., ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170 associated loci provided in Table 1) or any combination thereof. In some embodiments, one or more genomic loci comprise: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRI (e.g., ESRI associated loci provided in Table 1); and one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level13403741vl Page 45 of 173Attorney Docket: 2014191-0051correlated loci for one or more of EN01, YBXI, GATA3, F0XA1, HAPI. N3, EN1, PJM1, CCDC170, (e.g., ENOJ, YBXI, GATA3, F0XA1, HAPLN3, ENJ, PIM1, CCDC170 associated loci provided in Table 1) or any combination thereof.

[0161] In some embodiments, one or more expression-level correlated loci for ESRI, ENO J, YBXI, GATA3, FOXA1, HAPLN3, EN1, P1M1, or CCDC170 are proximal to (e.g., within + / - 200 kB of the transcription start site (TSS) o ) ESRI, ENOJ, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170, respectively. In some embodiments, one or more expression-level correlated loci for ESRI, EN01, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof include one or more promoter regions for one or more of ESRI, EN01, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof. In some embodiments, one or more expression-level correlated loci for ESRI, EN01, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof include one or more enhancer regions for one or more of ESRI, EN01, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.

[0162] In some embodiments, one or more expression-level correlated loci for ESRI, EN01, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof, are genomic regions at which signal of one or more epigenetic biomarkers (i) has been shown to be correlated with ESRI expression, and / or (ii) has been shown to be correlated with EN01, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170 expression, or any combination thereof. In some embodiments, (i) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRI include one or more ESRI associated loci provided in Table 1; (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, or 75 or more expression-level correlated loci for ENO1 include one or more ENO1 associated loci provided in Table 1; (iii) one or more, 5 or more, 10 or more, or 15 or more expression-level correlated loci for YBXI include one or more YBXI associated loci provided in Table 1; (iv) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, or 90 or more expressionlevel correlated loci for GATA3 include one or more GATA3 associated loci provided in Table 1; (v) none or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-13403741vl Page 46 of 173Attorney Docket: 2014191-0051level correlated loci for F0XA1 include one or more F0XA1 associated loci provided in Table 1; (vi) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 55 or more expression-level correlated loci for HAPLN3 include one or more HAPLN3 associated loci provided in Table 1; (vii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, or 125 or more expression-level correlated loci for EN1 include one or more EN1 associated loci provided in Table 1; (viii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more expression-level correlated loci for PIM1 include one or more PIM1 associated loci provided in Table 1; (ix) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 115 or more expression-level correlated loci for CCDC170 include one or more CCDC170 associated loci provided in Table 1; or (x) any combination of (i)-(ix).

[0163] In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor at each of the loci provided in Table 1. In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor for at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, 40%, 50%, 75%, or 100% of loci identified in Table 1. In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor for at least a percent of loci identified in Table 1 having a lower bound selected from 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 1%, 2%, 3%, 4%, 5%, or 10%, and an upper bound selected from 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, 40%, 50%, 75%, or 100%.Differential H3K4me3 modification

[0164] Exemplary genomic loci whose H3K4 methylation state (in particular, H3K4 trimethylation or H3K4me3 state) is associated with ER expression status (e.g., ER+ status) are provided in Table 1 (see “H3K4me3” loci).

[0165] A person of skill in the art will recognize that the methods disclosed herein do not require that every H3K4me3 analyte genomic locus listed in Table 1 be assessed for H3K4me313403741vl Page 47 of 173Attorney Docket: 2014191-0051modification. Instead, a subset of H3K4me3 analyte loci may be assessed for H3K4me3 modification. Subsets of the H3K4me3 analyte genomic loci of Table 1 can be selected (e.g., for use in determining ER expression status) based on various performance criteria, e.g., to select genomic loci that demonstrate differential modification with a particular level of statistical significance and / or a particular threshold of differential between relevant states (e.g., a measured log2(fold-change)). Subsets of the genomic loci may also be selected based on an algorithm, e.g., during the process of obtaining a classifier. Those of skill in the art will appreciate that such subsets of loci of Table 1, and loci included in such subsets, are together, individually, and / or in randomly selected subsets, at least as informative (e.g., as statistically significant and / or reliable) for uses disclosed herein, e.g., for determining ER expression status.

[0166] In various embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 or more H3K4me3 loci identified in Table 1 are differentially H3K4me3 modified (e.g., (a) a number of loci identified in Table 1 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, including, e.g., about 1 to about 120, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100 as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER expression or a cohort of subjects with aberrant ER expression (e.g., a subject or cohort of subjects with a ER+ cancer)).

[0167] In various embodiments, a sample or subject from which the sample is derived, is determined to have an ER+ cancer if one or more promoter regions of one or more genes (e.g., one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more) in Table 1 are differentially H3K4me3 modified as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER expression or a cohort of subjects with aberrant ER expression (e.g., a subject or cohort of subjects with a ER+ cancer)).

[0168] In some embodiments, one or more expression correlated loci for ESRI include: one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are13403741vl Page 48 of 173Attorney Docket: 2014191-0051associated with ESRI and that are provided in Table 1. In some embodiments, one or more expression correlated loci for EN01 include: one or more, 5 or more, or 10 or more “H3K4me3” analyte loci that are associated with EN01 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for GATA3 include: (i) one or more, 5 or more, 10 or more, or 15 or more “H3K4me3” analyte loci that are associated with GATA3 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for FOXA1 include: (i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with FOXA1 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for HAPLN3 include: one or more, 2 or more, 3 or more, 4 or more, 5 or more, 5 or more, 7 or more, 8 or more, 9 or more, or 10 or more “H3K4me3” analyte loci that are associated with HAPLN3 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for ENl include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “H3K4me3” analyte loci that are associated with ENl and that are provided in Table 1. In some embodiments, one or more expression correlated loci for PIM1 include: (i) one or more, 2 or more, or 3 or more “H3K4me3” analyte loci that are associated with PIM1 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for CGDC170 include: (i) one or more, 5 or more, 10 or more, or 20 or more “H3K4me3” analyte loci that are associated with CCDC170 and that are provided in Table 1.

[0169] In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with ESRI and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 5 or more, or 10 or more “H3K4me3” analyte loci that are associated with ENO1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 5 or more, 10 or more, or 15 or more “H3K4me3” analyte loci that are associated with GATA3 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with FOXA1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 2 or more, 3 or more, 4 or more, 5 or more, 5 or more, 7 or more, 8 or more, 9 or more, or 10 or more13403741vl Page 49 of 173Attorney Docket: 2014191-0051“H3K4me3” analyte loci that are associated with HAPLN3 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “H3K4me3” analyte loci that are associated with EN1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 2 or more, or 3 or more “H3K4me3” analyte loci that are associated with / 7A77 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K4me3 modifications at one or more, 5 or more, 10 or more, or 20 or more “H3K4me3” analyte loci that are associated with CCDC170 and that are provided in Table 1.

[0170] In various embodiments, a promoter region refers to a region a certain number of nucleotides upstream of a gene (e.g., 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, or 1,000 nucleotides upstream of a gene). In some embodiments, a promoter region refers to an H3K4me3 locus identified in Table 1.

[0171] In various embodiments, differentially H3K4me3 modified refers to a methylation status characterized by an increase or decrease in a value measuring methylation (e.g., of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring methylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase or decrease of 0.1-fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-13403741vl Page 50 of 173Attorney Docket: 2014191-0051fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold. 1 2-fold to 4.0-fold. 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.Differential H3K27ac modification

[0172] Exemplary genomic loci whose H3K27ac state is associated with ER expression status (e.g., ER+ status) are provided in Table 1 (see “H3K27ac” loci).

[0173] A person of skill in the art will recognize that the methods disclosed herein do not require that every H3K27ac genomic locus listed in Table 1 be assessed for H3K27ac modifications. Instead, a subset of H3K27ac loci may be assessed for H3K27ac modification. Subsets of the H3K27ac genomic loci of Table 1 can be selected (e.g., for use in determining ER expression status) based on various performance criteria, e.g., to select genomic loci that demonstrate differential modification with a particular level of statistical significance and / or a particular threshold of differential between relevant states (e.g., a measured log2(fold-change)). Subsets of the genomic loci may also be selected based on an algorithm, e.g., during the process of obtaining a classifier. Those of skill in the art will appreciate that such subsets of loci of Table 1, and loci included in such subsets, are together, individually, and / or in randomly selected subsets, at least as informative (e.g., as statistically significant and / or reliable) for uses disclosed herein, e.g., for determining ER expression status.

[0174] In various embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, or about 25 or more H3K27ac loci identified in Table 1 (e.g., (a) a number of loci identified in Table 1 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 120, about 150, about 200, about 250, about 300, about 350, about 400, including, e.g., about 1 to about 400, about 1 to about 300, about 1 to about 200, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100 are differentially H3K27ac modified as compared to a13403741vl Page 51 of 173Attorney Docket: 2014191-0051reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER expression or a cohort of subjects with aberrant ER expression (e.g., a subject or cohort of subjects with a ER+ cancer)).

[0175] In various embodiments, a sample or subject from which the sample is derived, is determined to have a particular ER expression status if one or more enhancer regions of one or more (e.g., 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more) H3K27ac loci in Table 1 are differentially H3K27ac modified as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER expression status or a cohort of subjects with aberrant an ER expression status (e.g., a subject or cohort of subjects with an ER+ cancer)).

[0176] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if H3K27ac modifications for 1, 2, 3, 4, 5, 10, 15, 20, or 25, or more of the H3K27ac loci identified in Table 1 are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER expression status or a cohort of subjects with an aberrant ER expression status (e.g., a subject or cohort of subjects with an ER+ cancer)).

[0177] In some embodiments, one or more expression correlated loci for ESRI include: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 60 or more “H3K27ac” analyte loci that are associated with ESRi and that are provided in Table 1. In some embodiments, one or more expression correlated loci for EN01 include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “H3K27ac” analyte loci that are associated with EN01 and that are provided in Table 1. In some embodiments, one or more expression-level correlated loci for YBX1 include: one or more “H3K27ac” analyte loci that are associated with YBX1 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for GATA3 include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 “H3K27ac” analyte loci that are associated with GATA3 and that are provided in Table 1. In some embodiments, one or more one or more expression correlated loci for FOXA1 include: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with FOXA1 and that are provided in Table 1. In some embodiments, one or more expression-level correlated loci for HAPLN3 include: one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K27ac” analyte loci that are associated with HAPLN3 and that are13403741vl Page 52 of 173Attorney Docket: 2014191-0051provided in Table 1. In some embodiments, one or more expression correlated loci for EN1 include: one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with EN1 and that are provided in Table 1. In some embodiments, one or more expression correlated loci for PIM1 include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “H3K27ac” analyte loci that are associated with PIM1 and that are provided in Table 1. In some embodiments, one or more one or more expression correlated loci for CCDC170 include: (ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more “H3K27ac” analyte loci that are associated with CCDC170 and that are provided in Table 1.

[0178] In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 60 or more “H3K27ac” analyte loci that are associated with ESRI and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “H3K27ac” analyte loci that are associated with ENO1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more “H3K27ac” analyte loci that are associated with YBX1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 “H3K27ac” analyte loci that are associated with GATA3 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with FOXA1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K27ac” analyte loci that are associated with HAPLN3 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with EN1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “H3K27ac”13403741vl Page 53 of 173Attorney Docket: 2014191-0051analyte loci that are associated with PIM1 and that are provided in Table 1. In some embodiments, a method comprises quantifying H3K27ac modifications for one or more 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more “H3K27ac” analyte loci that are associated with CCDC170 and that are provided in Table 1.

[0179] In various embodiments, differentially H3K27ac modified refers to an acetylation status characterized by an increase or decrease in a value measuring acetylation (e.g., of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring acetylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase or decrease of 0.1-fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold. 1.2-fold to 4.0-fold. 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.

[0180] In some embodiments, one or more enhancer regions of a recited gene are provided in Table 1. In some embodiments, one or more enhancer regions of a recited gene corresponds to: (i) one or more loci with increased or decreased H3K27ac modifications as compared to a reference (e.g., a sample from a healthy subject) within a certain number of13403741vl Page 54 of 173Attorney Docket: 2014191-0051nucleotides (e.g., 50,000 nucleotides) of the recited gene; and / or (ii) one or more loci with increased or decreased H3K27ac modifications as compared to a reference (e.g., a sample from a healthy subject) that are closest to the recited gene in the genome.Differential DNA methylation

[0181] Exemplary genomic loci whose DNA methylated state is associated with ER expression status (e.g., ER+ status) are provided in Table 1 (see " MBD” loci).

[0182] A person of skill in the art will recognize that the methods disclosed herein do not require that every MBD locus listed in Table 1 be assessed for DNA methylation. Instead, a subset of MBD loci may be assessed for DNA methylation. Subsets of the MBD loci of Table 1 can be selected (e.g., for use in determining ER expression status) based on various performance criteria, e.g., to select genomic loci that demonstrate differential modification with a particular level of statistical significance and / or a particular threshold of differential between relevant states (e.g., a measured log2(fold-change)). Subsets of the genomic loci may also be selected based on an algorithm, e.g., during the process of obtaining a classifier. Those of skill in the art will appreciate that such subsets of loci of Table 1, and loci included in such subsets, are together, individually, and / or in randomly selected subsets, at least as informative (e.g., as statistically significant and / or reliable) for uses disclosed herein, e.g., for determining ER expression status.

[0183] In various embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, or about 25 or more MBD loci identified in Table 1 (e.g., (a) a number of loci identified in Table 1 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 105, about 110, about 115, about 125, about 150, about 200, about 250, including, e.g., about 1 to about 200, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about lOOare differentially DNA methylated as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER status or a cohort of subjects with an aberrant ER expression status (e.g., a subject or cohort of subjects with an ER+ cancer)).13403741vl Page 55 of 173Attorney Docket: 2014191-0051

[0184] In various embodiments, a sample or subject from which the sample is derived, is determined to have a particular ER expression status if one or more MBD loci (e.g., 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more loci) in Table 1 are differentially DNA methylated as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER expression status or a cohort of subjects with an aberrant ER expression status (e.g., a subject or cohort of subjects with an ER+ cancer)).

[0185] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if DNA methylation for 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more MBD loci that are identified in Table 1 as having a positive association with ER expression status are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER expression status or a cohort of subjects with an aberrant ER expression status (e.g., a subject or cohort of subjects with an ER+ cancer)).

[0186] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Table 1 as having a positive association with ER expression are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER expression or a cohort of subjects with aberrant ER expression (e g., a subject or cohort of subjects with a ER+ cancer)). In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Table 1 as having a negative association with ER expression are decreased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER expression or a cohort of subjects with aberrant ER expression (e.g., a subject or cohort of subjects with a ER+ cancer)).

[0187] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER expression status if:(a) DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Table 1 as having a positive association with ER expression are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject13403741vl Page 56 of 173Attorney Docket: 2014191-0051with aberrant ER expression or a cohort of subjects with aberrant ER expression (e.g., a subject or cohort of subjects with a ER+ cancer)); and(b) DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Table 1 as having a negative association with ER expression are decreased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER expression or a cohort of subjects with aberrant ER expression (e.g., a subject or cohort of subjects with a ER+ cancer)).

[0188] In various embodiments, “differentially DNA methylated” refers to a methylation status characterized by an increase or decrease in a value measuring methylation (e.g., of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring methylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase of 0.1 -fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold. 1.2-fold to 4.0-fold. 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.

[0189] In some embodiments, one or more ESRI expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD”13403741vl Page 57 of 173Attorney Docket: 2014191-0051analyte loci that are associated with ESRI and that are provided in Table 1. In some embodiments, one or more EN01 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ENO1 and that are provided in Table 1. In some embodiments, one or more YBX1 expression correlated loci include: one or more, 5 or more, 10 or more, or 15 or more “MBD” analyte loci that are associated with YBX1 and that are provided in Table 1. In some embodiments, one or more GATA3 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “MBD” analyte loci that are associated with GATA3 and that are provided in Table 1. In some embodiments, one or more FOXA1 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with FOXA1 and that are provided in Table 1. In some embodiments, one or more HAP LN 3 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, or 20 or more “MBD” analyte loci that are associated with HAPLN3 and that are provided in Table 1. In some embodiments, one or more EN1 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with EN1 and that are provided in Table 1. In some embodiments, one or more PIM1 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with PIM1 and that are provided in Table 1. In some embodiments, one or more CCDC170 expression correlated loci include: one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with CCDC170 and that are provided in Table 1.Differential chromatin accessibility or transcription factor binding

[0190] Genomic loci provided in Table 1 can also demonstrate differential chromatin accessibility or transcription factor binding in different ER expression states.

[0191] In various embodiments, without wishing to be bound by any particular scientific theory, histone methylation (e.g., H3K4me3) corresponds and / or is correlated with chromatin accessibility. In various embodiments, without wishing to be bound by any particular scientific theory, histone acetylation (e.g., H3K27ac) corresponds and / or is correlated with chromatin accessibility. In various embodiments, without wishing to be bound by any particular scientific13403741vl Page 58 of 173Attorney Docket: 2014191-0051theory, DNA methylation corresponds and / or is correlated with chromatin accessibility.

[0192] In some embodiments, without wishing to be limited to any particular scientific theory, chromatin accessibility corresponds and / or is correlated with H3K4me3 modifications. As a result, in some embodiments, ER expression status may be determined by detecting and quantifying chromatin accessibility at one or more genomic loci in Table 1 in accordance with the section above discussing exemplary genomic loci with differential H3K4me3 modifications.

[0193] In some embodiments, without wishing to be limited to any particular scientific theory, chromatin accessibility corresponds and / or is correlated with H3K27ac modifications. As a result, in some embodiments, ER expression status may be determined by detecting and quantifying chromatin accessibility at one or more genomic loci in Table 1 in accordance with the section above discussing exemplary genomic loci with differential H3K27ac modifications.

[0194] In some embodiments, without wishing to be limited to any particular scientific theory, chromatin accessibility corresponds and / or is correlated with DNA methylation. As a result, in some embodiments, ER expression status can be measured by detecting and quantifying chromatin accessibility at one or more genomic loci in Table 1 in accordance with the section above discussing exemplary genomic loci with differential DNA methylation.

[0195] In various embodiments, without wishing to be bound by any particular scientific theory, histone methylation (e.g., H3K4me3) corresponds and / or is correlated with transcription factor binding. In various embodiments, without wishing to be bound by any particular scientific theory, histone acetylation (e.g., H3K27ac) corresponds and / or is correlated with transcription factor binding. In various embodiments, without wishing to be bound by any particular scientific theory, DNA methylation corresponds and / or is correlated with transcription factor binding.

[0196] In some embodiments, without wishing to be limited to any particular scientific theory, binding of RNA pol II corresponds and / or is correlated with H3K4me3 modifications. As a result, in some embodiments, ER expression status may be determined by detecting and quantifying binding of RNA pol II at one or more genomic loci in Table 1 in accordance with the section above discussing exemplary genomic loci with differential H3K4me3 modifications.

[0197] In some embodiments, without wishing to be limited to any particular scientific theory, binding of p300, mediator complex, cohesin complex or RNA pol II corresponds and / or is correlated with H3K27ac modifications. As a result, in some embodiments, ER expression status may be determined by detecting and quantifying binding of p300, mediator complex,13403741vl Page 59 of 173Attorney Docket: 2014191-0051cohesin complex or RNA pol II at one or more genomic loci in Table 1 in accordance with the section above discussing exemplary genomic loci with differential H3K27ac modifications.Using Classifiers

[0198] Models produced according to methods disclosed herein may be used to predict ER status of a cancer in a subject based on one or more samples from the subject, for example liquid biopsy samples. Such predictions can be made for different samples at different times. Such predictions can be used to diagnose and / or monitor subjects. In some embodiments, a model is used to quantify expression level of one or more ESRI related genes (e.g., ESRJ, EN01, YBX1, GATA3, FOXAJ, HAPLN3, ENJ, PIMJ, CCDC170) for a subject that has or is suspected of having an indication, such as, for example, a cancer. Different samples, such as liquid biopsy (e g., plasma) samples, may be taken from a subject and used to monitor the subject and / or diagnose the subject. Models disclosed herein may be used to track response to treatment with a therapy. For example, a subject may be treated with an ER-targeted agent (e.g., an ER degrader) and a model may be used to monitor ER status or reduction in expression of ER or one or more ER related genes for the subject using data derived from liquid biopsy samples for the subject.

[0199] In some embodiments, a method is for predicting ESRJ expression level in a subject and / or an ER status of a cancer in a subject. Such a method may include providing sample data for a subject, wherein the sample data include signal for one or more epigenetic biomarkers for ESRI (e.g., and optionally one or more related genes). Such a method may further include providing a model that has been produced using a method disclosed herein. Such a method may further include predicting an ESRJ expression level for the subject and / or ER status for a cancer in the subject using the sample data with the model. In some embodiments, the sample data have been derived from a liquid biopsy sample (e.g., from plasma) for the subject.

[0200] In some embodiments, a method is for monitoring a subject. Such a method may include providing (e g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time. Such a method may further include predicting a first expression level for ESRJ and / or a first ER status for the first sample and a second ESRJ expression level and / or a second ER status using the signal for the second sample using a model13403741vl Page 60 of 173Attorney Docket: 2014191-0051that has been produced using a method disclosed herein, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g., an ER-targeted agent). Such a method may further include determining a difference in ESRI expression level and / or ER status between the first sample and the second sample. In some embodiments, the change is a reduction in ESRI expression level and / or a change in ER status and the therapy is an ER-targeted agent.

[0201] In some embodiments, a method is for characterizing cancer recurrence and / or progression. Such a method may include predicting an ESRI expression level for a subject having a cancer and / or an ER status of a cancer in a subject based on a first sample from the subject taken at a first time point using a model that has been produced using a method disclosed herein. Such a method may further include predicting an ESRI expression level for the subject and / or an ER status for a cancer in the subject based on a second sample for the subject taken at a second time point after the first time point using a method disclosed herein. Such a method may further include determining a difference in the ESRI expression level and / or ER status at the second time point and at the first time point.

[0202] In some embodiments, a method is for monitoring cancer in a subject. Such a method may include predicting an ESRI expression level and / or ER status based on a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein. Such a method may further include determining whether there is a difference in ESRJ expression level and / or ER status for the subject over time.

[0203] In some embodiments, a method is for determining effectiveness of a therapeutic agent in a subject having cancer. Such a method may include predicting an ESRI expression level and / or ER status based on a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein. Such a method may further include determining whether there is a difference in ESRI expression level and / or ER status for the subject over time.

[0204] In some embodiments, a method is for monitoring response of a subject having an indication to a therapeutic agent for the indication. Such a method may include predicting an ESRI expression level and / or ER status based on a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein. Such a method may further include determining whether there is a difference in ESRI expression level and / or ER status for13403741V1 Page 61 of 173Attorney Docket: 2014191-0051the subject over time.

[0205] In some embodiments, a method includes administering a therapy including a therapeutic agent to a subject between when two or more samples were obtained. In some embodiments, a method includes administering a therapy to the subject when a difference in expression level of a target gene (e g., ESRI or an ESRI related gene) is determined to be at least as large as a threshold difference. In some embodiments, a method includes altering administration of a therapy to a subject when a difference in expression level over time is determined to be at least as large as a threshold difference. In some embodiments, altering administration includes increasing a dosage and / or frequency of administration.

[0206] In some embodiments, a method is for prognosing cancer. Such a method may include predicting an ESRI expression level and / or an ER status for a subject having a cancer based on a sample for the subject using a model that has been produced using a method disclosed herein. Such a method may further include prognosing cancer in the subject based on the determined expression level. In some embodiments, a method includes administering a therapy based on a prognosis.

[0207] In some embodiments, a method is for diagnosing cancer in a subject. Such a method may include predicting an ESRI expression level and / or ER status for a subject having a cancer based on a sample for the subject using a model that has been produced using a method disclosed herein. Such a method may further include determining that the expression level exceeds a threshold. In some embodiments, a method includes initiating administration of a therapy based on the expression level. In some embodiments, a method includes selecting a dosing regimen for a therapy based on the expression level.

[0208] In some embodiments, a method is for determining whether a cancer has been removed from a subject, the method including, after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, predicting ESRI expression level of and / or ER status for the subject based on a sample from the subject using a model that has been produced using a method disclosed herein. In some embodiments, a method includes continuing administration of a therapy based on the expression level. In some embodiments, a method includes ceasing administration of a therapy based on the expression level.

[0209] In some embodiments, the present disclosure provides methods for obtaining a classifier, e.g., a validated classifier that can be used to determine ER status. In some13403741vl Page 62 of 173Attorney Docket: 2014191-0051embodiments, a subject is determined to have a validated epigenetic profde indicative of an ERpositive cancer based on analysis of a biological sample, optionally of cell-free DNA (cfDNA) from a liquid biopsy sample, obtained or derived from the subject, wherein the presence of the validated epigenetic profile has been determined using a validated classifier.Applications

[0210] Methods, kits and systems of the present disclosure include analysis of differentially modified and / or differentially accessible genomic loci to determine the ER status of a cancer. Methods, kits and systems of the present disclosure can be used in any of a variety of applications. For example, methods, kits and systems of the present disclosure can be used in detecting and / or treating cancers based on ER status. Methods, kits and systems of the present disclosure can also be used to detect or determine resistance of a cancer, e.g., breast, ovarian, or endometrial cancer to a therapy or transformation from one cancer subtype to another.

[0211] In various embodiments, methods, kits and systems of the present disclosure can be applied to an asymptomatic human subject. As used herein, a subject can be referred to as “asymptomatic” if the subject does not report, and / or demonstrate by non-invasively observable indicia (e., without one, several, or all of device-based probing, tissue sample analysis, bodily fluid analysis, surgery, or cancer screening), sufficient characteristics of cancer to support a medically reasonable suspicion that the subject is likely suffering from cancer, e.g., breast, ovarian, or endometrial cancer. Detection of early-stage cancer can be achieved using methods, kits and systems of the present disclosure, with attendant medical benefits including potential for early treatment and attendant improvement in therapeutic outcomes.

[0212] In various embodiments, methods, kits, and systems of the present disclosure can be applied to a human subject that has increased susceptibility for cancer (including breast cancer).

[0213] In various embodiments, methods, kits and systems of the present disclosure can be applied to a symptomatic human subject. As used herein, a subject can be referred to as “symptomatic” if the subject report, and / or demonstrates by non-invasively observable indicia (e.g., without one, several, or all of device-based probing, tissue sample analysis, bodily fluid analysis, surgery, or cancer screening), sufficient characteristics of cancer to support a medically reasonable suspicion that the subject is likely suffering from cancer, e.g., breast, ovarian, or13403741vl Page 63 of 173Attorney Docket: 2014191-0051endometrial cancer. For example, in various embodiments a sample from a subject, optionally where the subject has a cancer that is of unknown ER status, can be assayed according to one or more embodiments of the present disclosure to determine ER status (e.g., to determine if the cancer is ER-positive or ER-negative). In various embodiments a sample from a subject, where the subject has a cancer that is known or suspected of being ER-positive (or ER-negative), can be assayed according to one or more embodiments of the present disclosure to determine the ER status of the cancer (e.g., if the cancer is ER-positive or ER-negative).

[0214] In some embodiments, methods, kits and systems of the present disclosure can be used to determine that a subject has an ER-positive cancer, optionally an ER-positive cancer that correlates with an ER+ Allred score of 3, 4, 5, 6, 7 or 8 based on IHC testing. In some embodiments, methods, kits and systems of the present disclosure can be used to determine that a subject has an ER-positive cancer, optionally an ER-positive cancer that correlates with an ER+ Allred score of at least 3, at least 4, at least 5, at least 6, at least 7 or 8 based on IHC testing. In some embodiments, methods, kits and systems of the present disclosure can be used to determine that a subject has an ER-negative cancer, optionally an ER-negative cancer that correlates with ER- Allred score of 0, 1, or 2 based on IHC testing.

[0215] In some embodiments, methods, kits and systems of the present disclosure can be used to validate or confirm a prior determination that a subject has an ER-positive cancer, optionally an ER-positive cancer that correlates with an ER+ Allred score of 3, 4, 5, 6, 7 or 8 based on IHC testing. In some embodiments, methods, kits and systems of the present disclosure can be used to validate or confirm a prior determination that a subject has an ER-positive cancer, optionally an ER-positive cancer that correlates with an ER+ Allred score of at least 3, at least 4, at least 5, at least 6, at least 7 or 8 based on IHC testing. In some embodiments, methods, kits and systems of the present disclosure can be used to validate or confirm a prior determination that a subject has an ER-negative cancer, optionally an ER-negative cancer that correlates with ER- Allred score of 0, 1, or 2 based on IHC testing.

[0216] In some embodiments, methods, kits and systems of the present disclosure are used to identify and detect new ER related categories that are independent of IHC or ISH scoring. For example, instead of training the classifier on samples from cohorts that were defined based on ER IHC or ISH testing, classifiers are trained on samples from cohorts that are defined based on whether they respond or do not respond to a particular ER-targeted agent. The resulting13403741vl Page 64 of 173Attorney Docket: 2014191-0051classifiers are then used to identify subjects that are more likely to respond to the particular ER-targeted agent independent of any IHC or ISH scoring. It is therefore to be understood that the term “ER status” as used herein is not limited to ER-positive and ER-negative or the traditional ER scoring based on IHC or ISH testing but can encompass any ER related categories including whether a subject will or will not respond to a particular ER-targeted agent.

[0217] Those of skill in the art will appreciate that regular, preventative, and / or prophylactic screening to determine ER status improves diagnosis of cancer, including and / or particularly early-stage cancer. Thus, the present disclosure provides, among other things, methods, kits and systems particularly useful for the diagnosis and treatment of early-stage cancer. Generally, and particularly in embodiments in which ER-positive cancer detection in accordance with the present disclosure is carried out annually, and / or in which a subject is asymptomatic at time of detecting, methods, kits and systems of the present disclosure are especially likely to detect early-stage ER-positive cancer. In various embodiments, detecting in accordance with methods, kits and systems of the present disclosure reduces cancer mortality, e.g., by early cancer diagnosis.

[0218] In various embodiments ER status determination in accordance with the present disclosure is performed once for a given subject or multiple times for a given subject. In various embodiments, ER status determination in accordance with the present disclosure is performed on a regular basis, e.g., every six months, annually, every two years, every three years, every four years, every five years, or every ten years.

[0219] In various embodiments, methods, kits and systems disclosed herein provide a determination of ER status. In other instances, methods, kits and systems disclosed herein will be indicative of ER status but not definitive for ER status. In various instances in which methods, kits and systems of the present disclosure are used to determine ER status, the same can be followed by a further confirmatory assay, which further assay can confirm, support, undermine, or reject a determination resulting from a prior determination, c.., a determination in accordance with the present disclosure. As used herein, a confirmatory assay can be an ER test that is currently recognized by medical practitioners, e.g., ER scoring based on IHC or ISH testing.

[0220] In various embodiments, ER status determination according to one or more methods, kits and / or systems disclosed herein is followed by treatment of cancer. In various embodiments, treatment of cancer includes administration of a therapeutic regimen including one13403741v 1 Page 65 of 173Attorney Docket: 2014191-0051or more cancer therapies provided herein, including without limitation one or more of ER targeted therapy, surgery, radiation, endocrine therapy, chemotherapy, and / or immunotherapy. In various embodiments, treatment of cancer includes administration of a therapeutic regimen including one or more treatments provided herein as available, appropriate, and / or preferred for a particular ER status.

[0221] In various embodiments, methods, kits and systems can be used to determine whether a particular subject and / or cancer is likely to be and / or is characterized as responsive to ER targeted therapy. In some such embodiments, methods, kits and systems can be followed by treatment of the subject with an ER targeted therapy.

[0222] In various embodiments, methods, kits and systems can be used to determine whether a particular subject and / or cancer is likely to be and / or is characterized as resistant to, non-responsive to, or not recommended treatment with to ER targeted therapy. In some such embodiments, methods, kits and systems can be followed by treatment with one or more of surgery and / or radiation, a HER2-targeted agent (if HER2-positive), chemotherapy and immunotherapy instead of ER targeted therapy.

[0223] Responsiveness can refer to the ability or likelihood of a therapy to cause a reduction in tumor size or inhibit tumor growth or metastasis. Responsiveness can refer to improvement in prognosis (e.g., increased time to cancer recurrence or increased life expectancy, e.g., overall survival, recurrence-free survival, metastasis-free survival, or disease-free survival). Responsiveness can refer to achievement of a treatment benefit, including e.g., improvement in one or more symptoms of cancer, e.g., breast, ovarian, or endometrial cancer. Responsiveness can be measured quantitatively (e.g., as in the case of tumor size; as in the case of measurement of histone modification, chromatin accessibility, transcription factor binding, or DNA methylation at one or more genomic loci; or as in the calculation of clinical benefit (CBR)), or qualitatively (e.g., by measures such as “pathological complete response” (pCR), “clinical complete remission” (cCR), “clinical partial remission” (cPR), “clinical stable disease” (cSD), “clinical progressive disease” (cPD), or other qualitative criteria). Resistance can refer to the inability or unlikelihood of a therapy to achieve a desired therapeutic effect (e.g., a reduction in tumor size, improvement in prognosis, or other treatment benefit such as, e.g., improvement in one or more symptoms of cancer) in a subject and / or cancer. Resistance includes both acquired and natural resistance. In certain embodiments, resistance includes the extent to which one or13403741vl Page 66 of 173Attorney Docket: 2014191-0051more desired therapeutic benefits results from administration of a therapy to a subject and / or cancer is less than that expected and / or achieved in a reference (e.g., less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of benefit achieved in a reference).

[0224] In various embodiments, methods, kits and systems can be used to detect the clinical efficacy of a course of therapy for cancer, e.g., breast, ovarian, or endometrial cancer. For example, methods and / or compositions of the present disclosure could be used to determine the presence, absence, or ER status of a cancer in a subject over the course of treatment. Methods and / or compositions of the present disclosure could be used in conjunction with, or confirmed by, other means of determining the presence, absence, or ER status of a cancer including, for example measurements of tumor size or character by techniques such as CT, PET, mammogram, ultrasound, palpation, histology, caliper measurement after biopsy or surgical resection, or by various qualitative, quantitative, or semi quantitative scoring systems including without limitation based on IHC or ISH testing, residual cancer burden (Symmans et al., J Clin Oncol (2007) 25:4414-4422, incorporated by reference herein in its entirety) or Miller-Payne score (Ogston et al., Breast (2003) 12:320-327, incorporated by reference herein in its entirety) in a qualitative fashion like “pathological complete response” (pCR), “clinical complete remission” (cCR), “clinical partial remission” (cPR), “clinical stable disease” (cSD), or “clinical progressive disease” (cPD).

[0225] In some embodiments, methods, kits, and systems described herein can be used to monitor progression of disease in a subject. In some embodiments, monitoring progression entails obtaining and characterizing samples from a subject at at least a first and a second time point. In some embodiments, at the first time point, a subject has already been diagnosed with lung cancer (e.g., an ER-negative cancer or ER-positive cancer). In some embodiments, at a first time point, a subject has been determined to have cancer and therapy is administered before or close to (e.g., the same day as) the first time point or between the first time point and the second time point; in such embodiments, determination of ER status at at least the first and the second time points can be used to monitor treatment efficacy and / or determine when a change in therapy should be made. For example, in some embodiments, a subject has previously been diagnosed with an ER-positive at the first time point, an ER-positive therapy is being or will be administered to the subject, and disease status can be monitored, which can be useful, e.g., for determining whether a change in therapy should be made. In some embodiments, treatment13403741vl Page 67 of 173Attorney Docket: 2014191-0051efficacy can be monitored, e.g., by using a method described herein to determine a decrease or increase in disease state signal, which can be useful, e.g., for determining whether an administered therapy is effective and / or whether a change in therapy should be made. In some embodiments, at the first time point, a cancer has gone into remission for a subject (e.g., the subject has minimal residual disease). In embodiments where a cancer has gone into remission, methods, kits, and systems described herein can be useful, e.g., for detecting reoccurrence of cancer, and can be faster, less expensive, and / or less invasive than, e.g., approaches that rely on tissue biopsies and / or imaging techniques.

[0226] In some embodiments, methods, kits and systems for ER status determination provided herein can inform treatment and / or payment (e.g., reimbursement for or reduction of cost of medical care, such as detecting or treatment) decisions and / or actions, e.g., by individuals, healthcare facilities, healthcare practitioners, health insurance providers, governmental bodies, or other parties interested in healthcare cost.

[0227] In some embodiments, methods, kits and systems for ER status determination provided herein can inform decision making relating to whether health insurance providers reimburse a healthcare cost payer or recipient (or not), e.g., for (1) ER status determination itself (e.g., reimbursement for detecting otherwise unavailable, available only for periodic / regular detecting, or available only for temporally- and / or incidentally- motivated detecting); and / or for (2) treatment, including initiating, maintaining, and / or altering therapy, e.g., based on the determined ER status. For example, in some embodiments, methods, kits and systems for ER status determination provided herein are used as the basis for, to contribute to, or support a determination as to whether a reimbursement or cost reduction will be provided to a healthcare cost payer or recipient. In some instances, a party seeking reimbursement or cost reduction can provide results of ER status determination conducted in accordance with the present disclosure together with a request for such reimbursement or reduction of a healthcare cost. In some instances, a party making a determination as to whether or not to provide a reimbursement or reduction of a healthcare cost will reach a determination based in whole or in part upon receipt and / or review of results of ER status determination conducted in accordance with the present disclosure.

[0228] In various embodiments, ER status determination using methods, kits and systems disclosed herein can be used in classifying subjects, samples, and / or tumors (e.g., breast13403741vl Page 68 of 173Attorney Docket: 2014191-0051cancer subjects, samples, and / or tumors). In various embodiments, methods, kits and systems disclosed herein can be used to generate a set of subjects, samples, and / or tumors identified according to the present methods, kits and systems each classified as corresponding to a particular ER status, and optionally using two or more of such classified subjects, samples, and / or tumors to identify biomarkers that distinguish the classes (i.e., distinguish the subjects, samples, and / or tumors according to their class, e.g, according to their ER status).

[0229] For illustration purposes and without limitation, in an exemplary assay of the present disclosure, one or more samples obtained from a subject (e.g, a liquid biopsy sample including cfDNA, e.g, a plasma sample including cfDNA) are analyzed by a method comprising enriching for cfDNA comprising a particular histone modification, wherein enriching is performed by a method that comprises incubating the sample with a reagent that specifically binds the histone modification being enriched for, and sequencing the enriched cfDNA. One example of such an assay is ChlP-seq for a histone modification (e.g., H3K4me3 and / or H3K27ac). Sequence reads (e.g., ChlP-seq sequence reads) can be aligned to human genome build hg19, e.g., using the Burrows-Wheeler Aligner (BWA). Non-uniquely mapping and redundant reads are optionally discarded.

[0230] To provide one example of peak calling, MACS v2.1.1.20140616 can be used for sequence (e.g., ChlP-seq) peak calling with a q-value (FDR) threshold of 0.01. Sequence (e.g., ChlP-seq) data quality can optionally be evaluated by any of one or more of a variety of measures, including total peak number, FRiP (fraction of reads in peak) score, number of high-confidence peaks (e.g, enriched > ten-fold over background), and percent of peak overlap with “blacklist” DHS peaks derived from the ENCODE project (Amemiya et al., Sci Rep (2019) 9( 1 ): 9354). If the sequence (e.g., ChlP-seq) data quality is below a particular threshold, the data may be discarded and the assay repeated. The number of reads overlapping the selected genomic loci for the relevant histone modification can be summed. In some embodiments, the average number of reads in the local background of each ChlP-seq peak is subtracted to improve signal to noise. In some embodiments, a sequence read density for one or more histone modifications can be calculated by a method that comprises (1) summing background adjusted sequence counts at at one or more genomic loci and dividing the resulting sum by the total number of kilobases of the one or more genomic loci, or (2) for each genomic loci, determining the ratio of background adjusted fragment counts to the number of kilobases of the genomic loci, and then summing the13403741vl Page 69 of 173Attorney Docket: 2014191-0051ratios for each loci.

[0231] Data can be log2 -transformed and quantile normalized to match the distribution of the data used to train a classifier. Normalized data can be used as input into a classifier that was trained using the same histone modification(s) and selected genomic loci. The classifier can then use inputted data to determine ER status of a subject’s cancer. It will be appreciated that this or similar approaches can be applied to assays of the present disclosure that quantify chromatin accessibility, transcription factor binding and / or DNA methylation.

[0232] Expression-level correlated loci refers to genomic regions where there has been determined to be a correlation between target gene expression (e.g., ESRI expression and / or expression of an ESRI related gene) and signal for one or more epigenetic biomarkers. In some embodiments, ESRI expression-level correlated loci are proximal to ESRI (e.g., within + / - 200 kB of ESR ). In some embodiments, ESRJ expression-level correlated loci are not proximal to ESRJ. In some embodiments, ESRI expression-level correlated loci are proximal to ESRJ-r elated genes. In some embodiments, ESRI expression-level correlated loci are proximal to (e.g., within + / - 200 kB of) ESRI, EN01, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170, or any combination thereof. In some embodiments, ESRI expression-level correlated loci are also expression level correlated with ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170, or any combination thereof.

[0233] FIG. 1 illustrates a method 100 according to the present disclosure that may be used to identify expression-correlated genomic loci and determine ESRI expression level and / or ER status. In step 102, digital samples for an indication that include signal for one or more epigenetic biomarkers and an expression level for a target gene (e.g., ESRI and / or one or more ESRI related genes) are received. In some embodiments, method 100 includes generating the digital samples. In step 104, the samples are tiled into tiles. In step 106, a loop is performed for at least one subset of the samples received in step 102 and tiled in step 104, for example a cross validated loop where each fold uses a different subset of the samples and rotates which exactly one of the samples is held out between the folds. The loop may be performed with a subset of samples all corresponding to (e.g., having) a particular ctDNA fraction. The loop may be performed a number of iterations for different subsets of samples all corresponding to (e.g., having) a particular ctDNA fraction. In some embodiments, the loop is performed a number of iterations using subsets of samples all corresponding to the same initial (e.g., relatively high)13403741vl Page 70 of 173Attorney Docket: 2014191-0051ctDNA fraction (e.g., 10% or 5%) and then, if the predictions from those iterations are within one or more predefined criteria, performed for a second number of iterations using subsets of samples all corresponding to a second, lower ctDNA fraction (e.g., 5% or 3%).

[0234] The loop includes performing steps 106a- 106c. In step 106a, one or more expression-level correlated tiles are determined from the tiles for each sample of the subset of samples used in the loop. The one or more expression-level correlated tiles may be determined by determining for each of the tiles whether there is correlation in the tile between expression level for a sample and signal for one or more epigenetic biomarkers for the sample (e.g., each of one or more epigenetic biomarkers) across the subset of samples. In some cases, adjacent overlapping tiles may both be correlated and collapsed together into one expression-level correlated tile. In step 106b, a model is produced using the expression-level correlated tiles (e.g., based on signal for the one or more epigenetic biomarkers corresponding to the expression-level correlated tiles and the expression level for the subset of the samples). Producing the model may include producing one or more point estimates (e.g., mean(s), e.g., geometric mean(s)) for signal across the expression-level correlated tiles. Producing the model may include producing a regression-based model, such as an OLS multiple regression using point estimates determined from signal for one or more epigenetic biomarkers for the subset of the samples in the expression-level correlated tiles determined in step 106a and the expression level for each of the subset of the samples. In step 106c, expression level (e.g., mRNAor protein expression level) is predicted for a sample received in step 102 and not included in the subset used to produce the model (e.g., a held out sample for that cross validation fold).

[0235] If all loops have been completed, then the method proceeds to step 108. If not, then another loop of step 106 (including steps 106a- 106c) proceeds. In step 108, once all appropriate loops are completed, it is determined whether the predictions fall within one or more predefined criteria by comparing the predicted expression level to the actual expression level for each sample used for prediction. The actual expression level may be assumed from a measured expression level corresponding to the cell line used to generate the sample. Predefined criteria that may be used include an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%) and / or an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9). Successful prediction within one or more predefined criteria may lead to tuning a model to perform at lower ctDNA fraction (e.g., clinically relevant fractions of 1-3%) and / or exhibit better13403741vl Page 71 of 173Attorney Docket: 2014191-0051performance (e.g., based on the one or more predefined criteria) and / or may lead to repeating the method 100 with samples corresponding to (e.g., having) a lower ctDNA fraction. Method 100 may be used to test whether a model for a target gene (e.g., ESRI or an ESRl-related gene) provides expression level prediction that correlate with expression for an indication.

[0236] In some embodiments, expression-correlated tiles and expression-correlated genomic regions may be determined by testing tiles for samples (e.g., a subset of samples). Signal for one or more epigenetic biomarkers for samples in different tiles spanning a genomic region corresponding to a target gene (e.g., ESRI or one or more ESRI related genes) may be compared to expression level for the samples to determine whether there is correlation. Such processes may identify the particular subregions in a genomic region where there is statistically meaningful signal that may be predictive of expression level. In general, not every portion of a genomic region corresponding to a target gene will produce signal (e.g., sequencing counts) that exhibits a statistically meaningful relationship with expression level (e.g., for at least one of a set of one or more epigenetic biomarkers being considered) and those regions can be ignored (e.g., at least with respect to the at least one of the set of one or more epigenetic biomarkers) when producing a model to predict expression level. Furthermore, by considering genomic regions or subregions that are too large, technical noise or other detrimental effects can reduce predictive power. By considering genomic regions (e.g., only genomic regions) that do not exhibit signal or that exhibit low signal in healthy volunteer samples (e.g., plasma samples), the signal associated with a target gene can be highly cancer-cell-specific. Tiling samples and then predicting an expression-level correlated tiles therefrom can mitigate technical noise and / or other detrimental effects that can reduce predictive power. Tiling approaches as disclosed herein may helpfully reduce dimensionality for further model tuning (e.g., feature engineering). A set of expressionlevel correlated tiles may be determined as tiles that each exhibit correlation between each of one or more epigenetic biomarkers and expression level or may be determined as tiles that each exhibit correlation between at least one of one or more epigenetic biomarkers and expression level.

[0237] In some embodiments, determining a set of expression-level correlated tiles includes determining, for one or more tiles, that a correlation between signal for at least one epigenetic biomarker with expression level exceeds a threshold for a correlation measure.Different correlation measures may be used in various embodiments. In some embodiments, a13403741vl Page 72 of 173Attorney Docket: 2014191-0051correlation measure is a Spearman correlation. In some embodiments, a threshold for a Spearman correlation used as a correlation measure is at least 0.1, at least 0.2, at least 0.3, at least 0.4, or at least 0.5. In some embodiments, a correlation measure is a Pearson correlation. In some embodiments, a fixed threshold may be used (e.g., of at least 0.1, at least 0.2, at least 0.3, at least 0.4, or at least 0.5). In some embodiments, a dynamic threshold is used, for example where the dynamic threshold is defined as a point estimate (e.g., median or mean) of correlations determined for a particular biomarker and gene. For example, a correlation measure (e.g., Pearson correlation) may be calculated for epigenetic biomarker signal and expression across a set of tiles and expression-level correlated tiles may be determined based on (e.g., as) those tile(s) for which the correlation measure is above the median. In some embodiments, a combination threshold is used that is a combination of a fixed threshold and a dynamic threshold, For example, a combination threshold may be a combination of a fixed threshold of 0.1 and a dynamic threshold of a median correlation measure between biomarker signal and expression level for a set of tiles (i.e., both thresholds must be exceeded to determine a tile as an expressionlevel correlated tile).

[0238] In some embodiments, one or more tiles are determined to be in a set of expression-level correlated tiles further based on signal for at least one epigenetic biomarker in healthy volunteer samples (e.g., used to generate digital samples being used to determine the set of expression-level correlated tiles) being below a threshold. Such a criterion can be used to eliminate tiles where there is an apparent correlation between signal for an epigenetic biomarker and expression level but the correlation is not caused by the indication but rather is a latent correlation for the genome (e.g., human genome). In some embodiments, healthy volunteer sample signal being below a threshold is determined by identifying peaks in epigenetic biomarker signal for healthy volunteer samples and excluding any tile with a peak in at least a threshold amount of healthy volunteer samples, for example with a peak in at least 10% of health volunteer samples. Such peaks may be identified using tools known in the art, such as, for example MACS2. In some embodiments, healthy volunteer sample signal being below a threshold is determined using pebbling with a p-value cutoff. Such pebbling may estimate likelihood that signal (e.g., sequencing counts) in a tile are above background.

[0239] In some embodiments, determining a set of expression-level correlated tiles includes determining signal for at least one epigenetic biomarker for one or more of a set of tiles13403741vl Page 73 of 173Attorney Docket: 2014191-0051that is uncorrelated with expression level across a subset of samples and excluding the one or more of the tiles from the set of expression-level correlated tiles [e.g., for the at least one of the one or more epigenetic biomarkers (e.g., excluding the one or more of the tiles from the set of expression-level correlated tiles entirely)]. Such exclusion may exclude that tile for only each epigenetic biomarker for which the tile is determined to be uncorrelated. In some embodiments, a set of expression-level correlated tiles includes only tiles for which there is correlation between epigenetic biomarker signal and expression level across each of a set of epigenetic biomarkers being considered or includes tiles for which there is correlation between epigenetic biomarker signal and expression level for at least one of a set of epigenetic biomarkers being considered; in general, the latter approach is preferred.

[0240] Tiles where there is correlation between expression level and epigenetic biomarker signal may be collapsed together where ones of the tiles overlap. Such collapsing may be used to produce a set of expression-level correlated tiles that are mutually exclusive. There may be correlation between signal for one or more epigenetic biomarkers and expression level for overlapping tiles where the correlation results from a common region shared by the overlapping tiles, where a correlative region spans over two or more tiles, or where there are a plurality of small nearly-spaced correlative regions, for example. In some embodiments, determining a set of expression-level correlated tiles includes (i) determining overlapping tiles have a correlation between signal corresponding to the tile for each of one or more epigenetic biomarkers and the expression level and (ii) collapsing the overlapping tiles such that the set of expression-level correlated tiles are mutually non-overlapping. Collapsing overlapping tiles may result in a set of expression-level correlated tiles includes tiles having non-uniform size.

[0241] A model for predicting ESRI expression level and / or ER status may be produced based on raw signal (e.g., normalized counts) or data derived from raw signal, for example one or more point estimates and expression level. A point estimate may be a mean, such as, for example, a geometric mean. Only one point estimate may be used per epigenetic biomarker or multiple point estimates may be used per epigenetic biomarker. For example, DNA methylation may be positively associated or negatively associated and different point estimates may be used for expression-level correlation tiles corresponding to positively associated DNA methylation and expression-level correlation tiles corresponding to negatively associated DNA methylation. In some embodiments, producing a model includes, for each of one or more epigenetic13403741vl Page 74 of 173Attorney Docket: 2014191-0051biomarkers, determining a point estimate for signal for the epigenetic biomarker across a set of expression-level correlated tiles and producing the model based on the point estimate. In some embodiments, the point estimate is a mean. In some embodiments, the mean is a geometric mean. In some embodiments, a model is produced based on a respective point estimate of signal for each of one or more epigenetic biomarkers for a set of expression-level correlation tiles. In some embodiments, producing a model includes determining a plurality of point estimates for signal for at least epigenetic biomarker across a set of expression-level correlated tiles and the model is produced based on the point estimate. In some embodiments, the plurality of point estimates includes a first point estimate for positively associated tiles for an epigenetic biomarker and a second point estimate for negatively associated tiles for the epigenetic biomarker.

[0242] One or more predefined criteria may be used to characterize a relationship (e.g., correlation) between predicted expression level from a model and actual expression level (e.g., taken from a cell line used to make samples). A predefined criterion may be, for example, an R2being above a threshold (e.g., at least 50%, at least 60%, at least 70%, or at least 80%) or an area under curve (AUC) above a threshold (e.g., at least 0.6, at least 0.7, at least 0.8, or at least 0.9); a combination of these criteria may be used. In some embodiments, a method includes determining whether a relationship (e.g., correlation) between an expression level for a target gene (e.g., ESRI or ES 7 -related gene) for a sample predicted from a model and an expression level for the sample across a plurality predicted samples are within one or more predefined criteria, for example an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%) and / or an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9). Such a determination may be made at a particular ctDNA fraction or for each of a plurality of ctDNA fractions (e.g., of no more than 10%, no more than 8%, no more than 6%, or no more than 5%).

[0243] A model may be tuned upon determining that the model is performant within one or more predefined criteria. A model may be tuned upon determining that predictions for expression level for a target gene (e.g., ESRI or one or more ESRI related genes) are within one or more predefined criteria. A method may be iterated at a lower ctDNA fraction upon determining that a model is performant within one or more predefined criteria. A method may be iterated at a lower ctDNA fraction upon determining that predictions for expression level for a target gene are within one or more predefined criteria. In some embodiments, a method includes determining that a relationship between expression level predicted with a model and actual13403741vl Page 75 of 173Attorney Docket: 2014191-0051expression level is within the one or more predefined criteria and, responsive to that determination, (e.g., manually) tuning (e.g., feature engineering) the model. Tuning a model may include selecting a subset of one or more expression-level correlated tiles previously determined and tuning the model using the subset of the one or more expression-level correlated tiles. In some embodiments, a model is for an initial ctDNA fraction and tuning the model includes tuning the model to produce predictions at a ctDNA fraction lower than the initial ctDNA fraction. In some embodiments, a method includes determining that a relationship between expression level predicted with a model and actual expression level is within one or more predefined criteria and, responsive to that determination, performing a loop with each of at least one subset of samples that correspond to a lower ctDNA fraction.

[0244] Samples may be tiled over a genomic region corresponding to a target gene (e.g., ESRI or one or more ESRi related genes). A genomic region may include a region corresponding to a transcript-encoding region of a target gene. A genomic region may correspond to an exon for a target gene. A genomic region may include a buffer around the target gene or a subregion thereof (e.g., around an exon-encoding region or a transcript-encoding region). For example, a genomic region may include initial and ending buffer regions. For example, a buffer of ±100 kb, ±200 kb, ±300 kb, ±400 kb, or ±500 kb may be used (e.g., around an exon-encoding region, a transcript-encoding region, or a transcription start site (TSS)).

[0245] Samples may be tiled into tiles to determine regions of a target gene (e.g., ESRI or one or more £iS7?7-related genes) where one or more epigenetic biomarkers are correlated with an expression level. In some embodiments, tiles are overlapping. In some embodiments, no more than two tiles are mutually overlapping. In some embodiments, adjacent tiles overlap by at least 10% (e.g., at least 20%, at least 30% of a length of the tiles) and no more than 70% (e.g., no more than 60% or no more than 50% of the length of the tiles. In some embodiments, adjacent tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp). In some embodiments, tiles have a length in a range of from 100 to 1000 bp (e.g., from 250 to 750 bp).

[0246] Signal for an epigenetic biomarker included in a sample may be normalized and / or pebbled, for example to allow different samples to be appropriately compared to each other and / or to account for background signal. Signal may be normalized and / or pebbled before performing one or more loops to determine correlation between expression level and epigenetic biomarker signal for sample tiles, produce a model, and use the model to predict expression13403741vl Page 76 of 173Attorney Docket: 2014191-0051level. Normalization may include a quantile normalization. In some embodiments, a method includes normalizing signal for each of one or more epigenetic biomarkers in samples prior to performing a loop. In some embodiments, a method includes pebbling signal for each of one or more epigenetic biomarkers in samples prior to performing the loop such that the loop is performed using the pebbled signal. In some embodiments, pebbling a signal for an epigenetic biomarker includes determining a background signal for the epigenetic biomarker for a genomic region based on signal (e.g., a number of fragments) in a sample in a pebbling region and subtracting the background signal from the signal. Normalization may be ctDNA fraction dependent (e.g., different normalizations applied to samples corresponding to (e.g., having) different ctDNA fractions). A pebbling region may be larger than a genomic region corresponding to a target gene (e.g., ESRI or an ESKY-related gene). A pebbling region may be at least 1 mega-base-pairs (Mbp), at least 2 Mbp, at least 3 Mbp, at least 4 Mbp, or at least 5 Mbp large. A pebbling region may include or overlap with a genomic region corresponding to a target gene, preferably include.

[0247] In some embodiments, one or more additional related (e.g., correlated) genes are considered in addition to ESRI when producing a model. For example, an ensemble model may be produced from constituent models each corresponding to a different gene or a single model may be produced that simultaneously considers signal across a larger genomic region corresponding to multiple genes. Such approaches may benefit from using predictive epigenetic signal to genomic regions outside ESRI that are related to ESRI.

[0248] Related genes may be selected (e.g., identified) in a number of ways, including using publicly accessible databases such as The Cancer Genome Atlas (TCGA). Related genes may be selected based on, for example, genes being in a transcriptional complex with ESRI being master regulators for an indication, being related to a particular pathway relevant for an indication, or a combination thereof. As an example, FOXA1 and GATA3 may be selected as related genes for ESRI for being in a same transcriptional complex. As another example or addition to the prior example, EN1 and PIM1 may be selected as related genes to ESRI because they are master regulators of ER- breast cancer. A number of related genes may be selected, for example at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15. Related genes may be selected by selecting a certain number of genes having highest rank according to a measure of correlation (e.g., as ranked by a data source (e.g., TCGA)), for example, though not only the13403741vl Page 77 of 173Attorney Docket: 2014191-0051highest ranking set of related genes need be used.

[0249] A set of expression-level correlated tiles for one or more epigenetic biomarkers may be determined for each of the related genes in addition to ESRI from digital samples that include signal for the one or more epigenetic biomarkers for the related genes in addition to ESRI in a manner disclosed herein with respect to ESRI. A model may be produced based on signal for one or more epigenetic biomarkers for a set of expression-level correlated tiles for each related gene and ESRI together (e.g., using a ridge regression) or constituent models for each related gene and ESRI may be produced separately from individual sets for each of one or more epigenetic biomarkers and then an ensemble model may be produced therefrom (e.g., based on an ordinary least squares regression or a ridge regression of the predictions from each constituent model). The latter approach may simplify the model and may be less prone to overfitting. The former approach may benefit where certain epigenetic biomarkers and / or regions might be more correlated for genes that are under tight regulatory control. For example, in certain such cases, enhancers are more correlated than promoters or DNA methylation. Because models can be produced for a large numbers of target genes quickly using methods disclosed herein, a model may have already been produced and / or a set of expression-level correlated tiles may have already been determined for a related gene to a target gene and therefore multigene models can be rapidly produced as well.

[0250] In some embodiments, a method of producing a multigene expression level prediction model may include producing a first constituent model for ESRI using a method disclosed herein; selecting one or more related genes related to ESRR, and producing a respective second constituent model for each of the one or more related genes using a method disclosed herein; and producing an ensemble model based on the first constituent model and the respective second constituent model(s). In some embodiments, the ensemble model uses a ridge regression of the first constituent model and the second constituent model.

[0251] In some embodiments, a method includes producing a model according to a method disclosed herein and selecting one or more related genes related to ESRI, wherein the digital samples include signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes such that set of expression-level correlated tiles includes at least one tile corresponding to each of the one or more related genes.13403741vl Page 78 of 173Attorney Docket: 2014191-0051

[0252] Tn some embodiments, a method includes producing a model according to a method disclosed herein and selecting one or more related genes related to ESRI, wherein the digital samples include signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes, and determining the set of expression-level correlated tiles includes testing each of the tiles corresponding to the one or more related genes for correlation between the signal in the sample corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level for the sample across the subset of the samples.

[0253] In some embodiments, a method of producing a model that predicts expression level of ESRI includes receiving digital samples for an indication (e.g., breast cancer) each including (i) signal for each of one or more epigenetic biomarkers for / correspond! ng to the indication and one or more selected related genes related to ESRi and (ii) an expression level (e g., mRNA expression level) for ES J producing a set of constituent models based on the digital samples, wherein the set includes one constituent model for each of ESRI and the one or more related genes; and producing an ensemble model based on a combination (e.g., using a ridge regression) of the constituent models in the set.

[0254] In some embodiments, a method of producing a model that predicts expression level of ESRI includes receiving digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for ESRI corresponding to the indication and one or more selected related genes related to ESRI and (ii) an expression level (e.g., mRNA expression level) for ESRR determining genomic regions corresponding to ESRI and the one or more related genes where the signal for at least one of the one or more epigenetic biomarkers correlates with the expression level for the samples (e.g., determine expression-level correlated tiles corresponding to ESRI and each of the one or more related genes); producing a model based on the signal for the one or more epigenetic biomarkers for the genomic regions (e.g., for the expression-level correlated tiles) and the expression level.

[0255] In some embodiments, signal for digital samples and for a subject is sequencing counts. In some embodiments, selecting one or more related genes includes selecting one or more genes that are in a transcriptional complex with ESRI. In some embodiments, selecting one or more related genes includes selecting one or more genes that are master regulators for an indication. In some embodiments, selecting one or more related genes includes selecting one or13403741vl Page 79 of 173Attorney Docket: 2014191-0051more genes that are related to a pathway relevant for an indication. In some embodiments, one or more related genes includes at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes. In some embodiments, selecting one or more related genes includes selecting a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].

[0256] In some embodiments, digital samples have been generated using data derived from cell samples specific to an indication and healthy volunteers. In some embodiments, cell samples include tissue samples. In some embodiments, cell samples have been derived from one or more cell lines, one or more patient-derived xenografts or a biopsy therefrom, one or more organoids, or a combination thereof. In some embodiments, a liquid biopsy sample is a plasma sample. In some embodiments, one or more epigenetic biomarkers include one or more histone modifications and / or DNA methylation. In some embodiments, one or more epigenetic biomarkers includes H3K27ac modification, H3K4me3 modification, and DNA methylation.

[0257] In some embodiments, multiple epigenetic biomarkers (e.g., one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation) can be quantified for a single sample. In such embodiments, two or more assays for assessing the epigenetic biomarkers can be performed in sequence (meaning a single sample can be probed for each modification in sequence) or in parallel (meaning that a single sample can be divided into multiple fractions, and then each fraction analyzed to quantifying an epigenetic biomarker). In some embodiments, H3K4me3 and H3K27ac histone modifications; H3K4me3 modifications and DNA methylation; H3K27ac modifications and DNA methylation; or H3K4me3 modifications, H3K27ac histone modifications, and DNA methylation are quantified in a single sample.

[0258] For the avoidance of any doubt, those of skill in the art will appreciate from the present disclosure that methods, kits and systems for ER status determination of the present disclosure are at least for in vitro use. Accordingly, all aspects and embodiments of the present disclosure can be performed and / or used at least in vitro.

[0259] Those of skill in the art will also appreciate that, in certain embodiments, methods of the present disclosure can be implemented on and / or in conjunction with a computer program and computer system. In some embodiments, methods of the present disclosure can be implemented on and / or in conjunction with a non-transitory computer readable storage medium13403741vl Page 80 of 173Attorney Docket: 2014191-0051encoded with the computer program, wherein the program comprises instructions that when executed by one or more processors cause the one or more processors to perform operations to perform the method. A computer system can also store and manipulate data generated by methods of the present disclosure that comprise a plurality of genomic locus modification status and / or accessibility status changes / profiles, which data can be used by a computer system in implementing methods disclosed herein. In certain embodiments, a computer system (i) receives modification status and / or accessibility status data; (ii) stores the data; and (iii) compares the data in any number of ways described herein (e.g., analysis relative to appropriate references), e.g., to determine ER status. In certain embodiments, a computer system (i) compares the genomic locus modification and / or accessibility status to a reference; and (ii) outputs an indication of whether the modification status and / or accessibility status of the genomic locus is significantly different from the reference and / or provides a determination regarding ER status.

[0260] Numerous types of computer systems can be used to implement methods of the present disclosure according to knowledge possessed by a skilled artisan in the bioinformatics and / or computer arts. Several software components can be loaded into memory during operation of such a computer system. The software components can comprise both software components that are standard in the art and components that are special to the present disclosure (e.g., dCHIP software described in Lin et al., Bioinformatics (2004) 20:1233-1240, incorporated herein by reference in its entirety; radial basis machine learning algorithms (RBM) known in the art). Methods of the present disclosure can also be programmed or modeled in mathematical software packages that allow symbolic entry of equations and high-level specification of processing, including specific algorithms to be used, thereby freeing a user of the need to procedurally program individual equations and algorithms. Such packages include, e.g., Matlab from Mathworks (Natick, MA), Mathematica from Wolfram Research (Champaign, IL), S-Plus from MathSoft (Seattle, WA), R from R Foundation for Statistical Computing (Vienna, Austria), Python from Python Software Foundation (Wilmington, DE), or Perl from Perl Foundation (Holland, MI). In certain embodiments, a computer system comprises a database for storage of genomic locus modification status and / or accessibility status data. Such stored profiles can be accessed and used to perform comparisons of interest at a later point in time. In addition to the exemplary program structures and computer systems described herein, other, alternative program structures and computer systems will be readily apparent to the skilled artisan.13403741vl Page 81 of 173Attorney Docket: 2014191-0051

[0261] Various algorithms can be applied to the comparison, between samples and references, of the modification status and / or accessibility status of genomic loci. In various embodiments, an algorithm can be a single learning statistical classifier system. Other suitable statistical algorithms are well known to those of skill in the art. For example, learning statistical classifier systems include a machine learning algorithmic technique capable of adapting to complex datasets (e.g., a panel of genomic loci of interest) and making decisions based upon such datasets. In some embodiments, a single learning statistical classifier system such as a classification tree (e.g., random forest) is used. In other embodiments, a combination of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more learning statistical classifier systems are used, preferably in tandem. Examples of learning statistical classifier systems include, but are not limited to, those described in the Examples and also those using inductive learning (e.g, decision / classification trees such as random forests, classification and regression trees (C& RT), boosted trees, etc ), Probably Approximately Correct (PAC) learning, connectionist learning (e.g., neural networks (NN), artificial neural networks (ANN), neuro fuzzy networks (NFN), network structures, perceptrons such as multi-layer perceptrons, multi-layer feed-forward networks, applications of neural networks, Bayesian learning in belief networks, etc ), reinforcement learning (e.g., passive learning in a known environment such as naive learning, adaptive dynamic learning, and temporal difference learning, passive learning in an unknown environment, active learning in an unknown environment, learning action-value functions, applications of reinforcement learning, etc ), and genetic algorithms and evolutionary programming. Other learning statistical classifier systems include support vector machines (e.g., Kernel methods), multivariate adaptive regression splines (MARS), Levenberg-Marquardt algorithms, Gauss-Newton algorithms, mixtures of Gaussians, gradient descent algorithms, and learning vector quantization (LVQ). In certain embodiments, methods of the present disclosure can include sending classification results to a medical practitioner, e.g., an oncologist.

[0262] In various embodiments, the area under the receiver operating characteristic (AUROC) for determining if a subject has a particular ER cancer status (e.g., an ER-positive cancer vs. an ER-negative cancer) is greater than 0.5 (e.g., greater than 0.55, greater than 0.6, greater than 0.65, greater than 0.7, greater than 0.75, greater than 0.8, greater than 0.85, greater than 0.9, or greater than 0.95).13403741vl Page 82 of 173Attorney Docket: 2014191-0051Formulation and Administration of Therapeutic Agents

[0263] The present disclosure includes methods where a therapeutic agent or regimen is administered to a subject based on the ER status of a cancer (e.., breast cancer, ovarian cancer, or endometrial cancer). In general, the therapeutic agent or regimen provided herein will be available, appropriate, and / or preferred for the determined ER status. Those of skill in the art will be aware of recommended and / or governmentally approved formulations and / or dosages for various therapeutic agents provided herein.

[0264] The present disclosure includes pharmaceutical compositions for delivery of one or more therapeutic agents to a subject. As disclosed herein, a pharmaceutical composition may be in any form known in the art, including formulations for administration according to any route known in the art. A suitable means of administration can be selected based on the age and condition of a subject.

[0265] Pharmaceutical composition forms of the present disclosure can include, e.g., liquid, semi-solid and solid dosage forms. Pharmaceutical composition forms of the present disclosure can include, e.g., liquid solutions (e.g., injectable and infusible solutions), dispersions or suspensions, tablets, pills, powders, and liposomes. Selection or use of any particular form may depend, in part, on the intended mode of administration and therapeutic application.Accordingly, the compositions can be formulated for administration by a parenteral mode (e.g., intravenous, subcutaneous, intraperitoneal, or intramuscular injection) or a non-parenteral mode. As used herein, parenteral administration refers to modes of administration other than enteral and topical administration, usually by injection or infusion.

[0266] In some embodiments, the compositions provided herein are present in unit dosage form, which unit dosage form can be suitable for self-administration. Such a unit dosage form may be provided within a container, e.g., a pill, vial, cartridge, prefilled syringe, or disposable pen.

[0267] A pharmaceutical composition of the present disclosure can be in an injectable or infusible form. For example, the present disclosure includes sterile formulations for injection or infusion, which can be formulated in accordance with conventional pharmaceutical practices. Sterile solutions can be prepared by incorporating a composition described herein in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filter sterilization. Solutions can be formulated, e.g., using distilled water,13403741vl Page 83 of 173Attorney Docket: 2014191-0051physiological saline, or an isotonic solution containing glucose and other supplements such as D-sorbitol, D-mannose, D-mannitol, or sodium chloride as an aqueous solution for injection, optionally in combination with a suitable solubilizing agent, for example, an alcohol such as ethanol and / or a polyalcohol such as propylene glycol or polyethylene glycol, and / or a nonionic surfactant such as polysorbate 80™ or HCO-50, and the like. In the case of sterile powders for the preparation of sterile injectable solutions, methods for preparation include vacuum drying and freeze-drying that yield a powder of a composition described herein plus any additional desired ingredient (see below) from a previously sterile-fdtered solution thereof. The proper fluidity of a solution can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Prolonged absorption of injectable compositions can be brought about by including in the composition a reagent that delays absorption, for example, monostearate salts, and gelatin. In particular instances, a pharmaceutical composition can be formulated, for example, as a buffered solution at a suitable concentration and suitable for storage, e.g., at 2-8°C (c.., 4°C).

[0268] In various embodiments, a pharmaceutical composition of the present disclosure can be formulated as a solution, microemulsion, dispersion, liposome, or other ordered structure suitable for stable storage at high concentration. Generally, dispersions are prepared by incorporating a composition described herein into a sterile vehicle that contains a basic dispersion medium and the required other ingredients from those enumerated above.

[0269] In various instances, a pharmaceutical composition can be formulated to include a pharmaceutically acceptable carrier or excipient. Examples of pharmaceutically acceptable carriers include, without limitation, any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible.

[0270] In certain embodiments, compositions can be formulated with a carrier that will protect the therapeutic agent against rapid release, such as a controlled release formulation, including implants and microencapsulated delivery systems. Biodegradable, biocompatible polymers can be used, such as ethylene vinyl acetate, poly anhydrides, polyglycolic acid, collagen, polyorthoesters, and polylactic acid. Many methods for the preparation of such formulations are known in the art. See, e.g., J. R. Robinson (1978) “Sustained and Controlled Release Drug Delivery Systems,” Marcel Dekker, Inc., New York.13403741vl Page 84 of 173Attorney Docket: 2014191-0051

[0271] Route of administration can be parenteral, for example, administration by injection. Administration by injection can be by intravenous injection, intramuscular injection, intraperitoneal injection, subcutaneous injection. Administration can be systemic or local. In certain embodiments, a composition described herein can be therapeutically delivered to a subject by way of local administration. As used herein, “local administration” or “local delivery,” can refer to delivery that does not rely upon transport of the composition or therapeutic agent to its intended target tissue or site via the vascular system. For example, the composition may be delivered by injection or implantation of the composition or therapeutic agent or by injection or implantation of a device containing the composition or therapeutic agent. In certain embodiments, following local administration in the vicinity of a target tissue or site, the composition or therapeutic agent, or one or more components thereof, may diffuse to an intended target tissue or site that is not the site of administration.

[0272] A pharmaceutical composition can be administered parenterally in the form of an injectable formulation comprising a sterile solution or suspension in water or another pharmaceutically acceptable liquid. For example, a pharmaceutical composition can be formulated by suitably combining the therapeutic molecule with pharmaceutically acceptable vehicles or media, such as sterile water and physiological saline, vegetable oil, emulsifier, suspension agent, surfactant, stabilizer, flavoring excipient, diluent, vehicle, preservative, binder, followed by mixing in a unit dose form required for generally accepted pharmaceutical practices. Examples of oily liquid include sesame oil and soybean oil, and it may be combined with benzyl benzoate or benzyl alcohol as a solubilizing agent. Other items that may be included are a buffer such as a phosphate buffer, or sodium acetate buffer, a soothing agent such as procaine hydrochloride, a stabilizer such as benzyl alcohol or phenol, and an antioxidant. The formulated injection can be packaged in a suitable ampule.

[0273] In various embodiments, subcutaneous administration can be accomplished by means of a device, such as a syringe, a prefilled syringe, an auto-injector (e.g., disposable or reusable), a pen injector, a patch injector, a wearable injector, an ambulatory syringe infusion pump with subcutaneous infusion sets, or other device for combining with a therapeutic agent for subcutaneous injection.

[0274] An injection system of the present disclosure may employ a delivery pen as described in U. S. Pat. No. 5,308,341. Pen devices, most commonly used for self-delivery of13403741vl Page 85 of 173Attorney Docket: 2014191-0051insulin to patients with diabetes, are well known in the art. Such devices can include at least one injection needle, are typically pre-filled with one or more therapeutic unit doses of a solution that includes the therapeutic agent and are useful for rapidly delivering solution to a subject with as little pain as possible. One medication delivery pen includes a vial holder into which a vial of a therapeutic or other medication may be received. The pen may be an entirely mechanical device or it may be combined with electronic circuitry to accurately set and / or indicate the dosage of medication that is injected into the user. See, e.g., U. S. Pat. No. 6,192,891. In some embodiments, the needle of the pen device is disposable and the kits include one or more disposable replacement needles. Pen devices suitable for delivery of any one of the presently featured compositions are also described in, e.g., U. S. Pat. Nos. 6,277,099; 6,200,296; and 6,146,361, the disclosures of each of which are incorporated herein by reference in their entirety. A microneedle-based pen device is described in, e.g., U. S. Pat. No. 7,556,615, the disclosure of which is incorporated herein by reference in its entirety. See also the Precision Pen Injector (PPI) device, MOLLY™, manufactured by Scandinavian Health Ltd.

[0275] In certain embodiments, administration of a therapeutic agent as described herein is achieved by administering to a subject a nucleic acid encoding a therapeutic agent described herein. Nucleic acids encoding a therapeutic agent described herein can be incorporated into a gene construct to be used as a part of a gene therapy protocol to deliver nucleic acids that can be used to express and produce therapeutic agent within cells. Expression constructs of such components may be administered in any therapeutically effective carrier, e.g., any formulation or composition capable of effectively delivering the component gene to cells in vivo. Approaches include insertion of the subject gene in viral vectors including recombinant retroviruses, adenovirus, adeno-associated virus, lentivirus, and herpes simplex virus- 1 (HSV-1), or recombinant bacterial or eukaryotic plasmids. Viral vectors can transfect cells directly; plasmid DNAcan be delivered with the help of, for example, cationic liposomes (lipofectin) or derivatized, polylysine conjugates, gramicidin S, artificial viral envelopes or other such intracellular carriers, as well as direct injection of the gene construct or CaPCh precipitation. Examples of suitable retroviruses include adenovirus-derived vectors, adeno-associated virus (AAV), pLJ, pZIP, pWE, and pEM which are known to those skilled in the art.

[0276] In some embodiments, a composition can be formulated for storage at a temperature below 0°C (e.g., -20°C or -80°C). In some embodiments, the composition can be13403741vl Page 86 of 173Attorney Docket: 2014191-0051formulated for storage for up to 2 years (e.g., one month, two months, three months, four months, five months, six months, seven months, eight months, nine months, 10 months, 11 months, 1 year, or 2 years) at 2-8°C (e.g., 4°C). Thus, in some embodiments, the compositions described herein are stable in storage for at least 1 year at 2-8°C (e.g., 4°C).

[0277] A pharmaceutical composition can include a therapeutically effective amount of a therapeutic agent described herein. Such effective amounts can be readily determined by one of ordinary skill in the art. A therapeutically effective amount can be an amount at which any toxic or detrimental effects of the composition are outweighed by therapeutically beneficial effects. In some embodiments, a dose can also be chosen to reduce or avoid production of antibodies or other host immune responses against a therapeutic agent. Those of skill in the art will appreciate that data obtained from cell culture assays and animal studies can be used in formulating a range of dosage for use in humans. In various embodiments, the amount of active ingredient included in a pharmaceutical composition is such that a suitable dose within the designated range can be administered to subjects. The dose and method of administration can vary depending on weight, age, condition, and other characteristics of a patient, and can be suitably selected as needed by those skilled in the art.

[0278] Pharmaceutical compositions including certain therapeutic agents, e.g., therapeutic antibodies, can be administered as a fixed dose, or in a milligram per kilogram (mg / kg) dose. While in no way intended to be limiting, an exemplary single dose of certain pharmaceutical compositions described herein can include certain therapeutic agents as described herein in an amount equal to, e.g., 0.001 to 1000 mg / kg, 1-1000 mg / kg, 1-100 mg / kg, 0.5-50 mg / kg, 0.1-100 mg / kg, 0.5-25 mg / kg, 1-20 mg / kg, and 1-10 mg / kg body weight. Exemplary dosages of a composition described herein include, without limitation, 0.1 mg / kg, 0.5 mg / kg, 1 mg / kg, 2 mg / kg, 4 mg / kg, 8 mg / kg, or 20 mg / kg. The present disclosure is not limited to such ranges or dosages.

[0279] The present disclosure further includes methods of preparing pharmaceutical compositions of the present disclosure and kits including pharmaceutical compositions of the present disclosure.

[0280] In various embodiments, therapeutic agents of the present disclosure can be administered to a subject in a course of treatment that further includes administration of one or more additional therapeutic agents or therapies that are not therapeutic agents (e.g., surgery or13403741vl Page 87 of 173Attorney Docket: 2014191-0051radiation). Combination therapies of the present disclosure can include simultaneous exposure of a subject to therapeutic agents of two or more therapeutic regimens.

[0281] In certain embodiments, a therapeutic agent as described herein can be administered together with (e.g., at the same time and / or in the same composition as) an additional agent or therapy. In certain embodiments, a therapeutic agent of the present disclosure can be administered separately from an additional therapeutic agent or therapy (e.g., at a different time and / or in a different composition than the additional therapeutic agent or therapy). Dosing regimens of a therapeutic agent and one or more additional therapeutic agents with which it is administered in combination can be coordinated or independently determined. In various embodiments, an additional therapeutic agent or therapy administered in combination with a therapeutic agent as described herein can be administered at the same time as therapeutic agent, on the same day as therapeutic agent, or in the same week as therapeutic agent. In various embodiments, an additional therapeutic agent or therapy administered in combination with a therapeutic agent as described herein can be administered such that administration of the therapeutic agent and the additional therapeutic agent or therapy are separated by one or more hours before or after, one or more days before or after, one or more weeks before or after, or one or more months before or after administration of the therapeutic agent. In various embodiments, the administration frequency and / or dosage of one or more additional therapeutic agents can be the same as, similar to, or different from the administration frequency of a therapeutic agent. In some embodiments, the two or more regimens can be administered simultaneously; in some embodiments, such regimens can be administered sequentially (e.g., all “doses” of a first regimen are administered prior to administration of any doses of a second regimen); in some embodiments, such therapeutic agents are administered in overlapping dosing regimens.

[0282] In certain embodiments, administration of a therapeutic agent can be to a subject having previously received, scheduled to receive, or in the course of a treatment regimen including an additional cancer therapy. Administration of a therapeutic agent can, in some instances, improve delivery or efficacy of another therapeutic agent or therapy with which it is administered in combination.

[0283] It is contemplated that therapeutic agent combination therapies can demonstrate synergy and / or greater-than-additive effects between a therapeutic agent and one or more additional therapeutic agents with which it is administered in combination. A therapeutic agent13403741vl Page 88 of 173Attorney Docket: 2014191-0051can be administered in any effective amount as determined independently or as determined by the joint action of therapeutic agent and any of one or more additional therapeutic agents or therapies administered. Administration of the therapeutic agent may, in some embodiments, reduce the therapeutically effective dosage, required dosage, or administered dosage of the additional therapeutic agent or therapy relative to a reference regimen for administration of additional therapeutic agent or therapy or therapy absent the therapeutic agent. In certain embodiment, a composition described herein can replace or augment other previously or currently administered therapy. For example, upon treating with therapeutic agent, administration of one or more additional therapeutic agents or therapies can cease or diminish, e.g., be administered at lower levels.Kits

[0284] The present disclosure includes kits for detecting modification and / or accessibility of one or more genomic loci. In some embodiments, the present disclosure provides kits for quantifying one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci. Kits of the present disclosure can include, e.g., reagents such as buffers and / or antibodies useful in the detection and quantification of histone modifications. In certain embodiments, a kit of the present disclosure can include at least one antibody that selective binds a histone modification selected from H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, or pan acetylation. In certain embodiments, a kit of the present disclosure can include at least one antibody that selective binds H3K4me3 modifications. In certain embodiments, a kit of the present disclosure can include at least one antibody that selective binds H3K27ac modifications. A kit of the present disclosure can include instructional materials disclosing or describing the use of the kit in a method of determining ER status and / or treatment disclosed herein. In various embodiments, a kit of the present disclosure can include one or more therapeutic agents useful in the treatment of cancer, e.g., as disclosed herein, optionally in combination with instruction materials for treatment of cancer, e.g., breast cancer, ovarian cancer, or endometrial cancer based on ER status.

[0285] In some embodiments, a kit of the present disclosure comprises reagents for quantifying one or more histone modifications, chromatin accessibility, binding of one or more13403741vl Page 89 of 173Attorney Docket: 2014191-0051transcription factors, and / or DNA methylation at one or more genomic loci, wherein the one or more genomic loci are selected from those provided in Table 1.

[0286] In some embodiments, the kit comprises reagents for quantifying H3K4me3 for at least 5, 10, 20, 30, 40, or 50 genomic loci in Table 1. In some embodiments, the kit comprises reagents for quantifying H3K27ac for at least 5, 10, 20, 30, 40, or 50 genomic loci in Table 1. In some embodiments, the kit comprises one or more antibodies for use in ChlP-seq, optionally wherein the one or more antibodies specifically bind H3K4me3- or H3K27ac-modified histones.

[0287] In some embodiments, the kit comprises reagents for quantifying DNA methylation for at least 5, 10, 20, 30, 40, or 50 genomic loci in Table 1. In some embodiments, the kit comprises one or more methyl-binding domains for use in MBD-seq. In some embodiments, the kit comprises one or more antibodies that can bind methylated DNA (e.g., for use in MeDIP).

[0288] In some embodiments, the kit comprises reagents for isolation of cell-free DNA (cfDNA) from a liquid biopsy sample. In some embodiments, the kit comprises reagents for library preparation for sequencing. In some embodiments, the kit comprises reagents for sequencing. In some embodiments, the kit comprises instructions for determining if a subject has an ER-positive cancer.

[0289] In some embodiments, the kit comprises one or more reagents for enriching for cfDNA having a sequence that falls within or overlaps with one or more genomic loci for which H3K4me3 modifications, H3K27ac modifications, and / or DNA methylation are to be quantified. In some embodiments, one or more reagents for enriching comprise reagents for selectively amplifying. In some embodiments, reagents for selectively amplifying comprise oligonucleotide primers (e.g., oligonucleotide primers for PCR amplification). In some embodiments, one or more reagents for enriching comprise one or more reagents that preferentially bind to one or more sequences within one or more genomic loci. In some embodiments, one or more reagents that preferentially bind one or more sequences within one or more genomic loci comprise one or more oligonucleotides, each comprising a sequence that is complementary to the one or more genomic loci.Systems

[0290] In some embodiments, one or more non-transitory computer readable storage13403741vl Page 90 of 173Attorney Docket: 2014191-0051media is encoded with a computer program, wherein the program includes instructions that when executed by one or more processors cause the one or more processors to perform operations to perform a method of the present disclosure.

[0291] In some embodiments, a computer system includes a memory and one or more processors coupled to the memory, wherein the one or more processors are configured to perform a method of the present disclosure.

[0292] In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds a histone modification selected from H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, orH3K4me3, or pan acetylation. In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds H3K4me3 modifications. In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds H3K27ac modifications.

[0293] In some embodiments, a system includes reagents for isolation of cfDNA from a liquid biopsy sample. In some embodiments, the sequencer includes reagents for library preparation for sequencing. In some embodiments, the sequencer includes reagents for sequencing.

[0294] Illustrative embodiments of systems and methods disclosed herein are described with reference to determinations and / or models that may be performed or used by a computing device. That is, in some embodiments, methods disclosed herein are computer-implemented methods and, in some embodiments, a system as disclosed herein includes a processor and one or more non-transitory computer readable storage media (e.g., one or more memories) that have instructions stored thereon that, when executed by the processor, cause the processor to perform operations that include a method disclosed herein. For example, a cTF estimation model may be stored on a memory. Such a cTF estimation model may be utilized by a processor to perform a method (e.g., a diagnostic method). Methods of the present disclosure, or portions thereof, may be performed using a processor. The processor may be a part of a computing device and / or computing system.

[0295] Systems of the present disclosure may include a processor and / or a memory. The memory may store one or more programs that include instructions that when executed by a processor cause at least a portion of a method disclosed herein to be performed. The system may further include a machine-learned model. Additionally or alternatively, a remotely stored and / or13403741vl Page 91 of 173Attorney Docket: 2014191-0051operated machine-learned model may be accessed by a (e.g., the) processor. The processor and / or memory may be a part of a computing device and / or computing system.

[0296] One or non-transitory computer readable media may store one or more programs that include instructions that when executed by a (e.g., the) processor cause at least a portion of a method disclosed herein to be performed.

[0297] Methods and systems disclosed herein may utilize one or more models that are machine-learned models. A machine-learned model may be or include an artificial neural network. A machine-learned model may employ, for example, an attention-based model (e.g., a transformer model, such as, for example, a vision transformer), a transformer model (e.g., a vision transformer), a regression-based model (e.g., a logistic regression model), a regularization-based model (e.g., an elastic net model or a ridge regression model), an instancebased model (e.g., a support vector machine or a k-nearest neighbor model), a Bayesian -based model (e.g., a naive-based model or a Gaussian naive-based model), a clustering-based model (e.g., an expectation maximization model), an ensemble-based model (e.g., an adaptive boosting model, a random forest model, a bootstrap-aggregation model, or a gradient boosting machine model), or a neural -network-based model (e.g., a convolutional neural network, a recurrent neural network, autoencoder, a back propagation network, or a stochastic gradient descent network).

[0298] In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a k nearest neighbors methodology, a generalized regression forward selection methodology, a generalized regression pruned forward selection methodology, a fit stepwise methodology, a generalized regression lasso methodology, a generalized regression elastic net methodology, a generalized regression ridge methodology, a nominal logistic methodology, a support vector machines methodology, a discriminant methodology, a naive Bayes methodology, or a combination thereof. In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a generalized regression lasso methodology, a generalized regression elastic net methodology, a generalized regression ridge methodology, a nominal logistic methodology, a support vector machines methodology, a discriminant methodology, or a combination thereof. In some embodiments, a machine-learned model is or is derived from a13403741vl Page 92 of 173Attorney Docket: 2014191-0051decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a support vector machines methodology, or a combination thereof.

[0299] Certain embodiments described herein make use of computer algorithms in the form of software instructions executed by a computer processor. In certain embodiments, the software instructions include a machine learned module. A machine learned module refers to a computer implemented process (e.g., a software function) that implements one or more specific machine-learned models, such as or including an artificial neural network (ANN), a convolutional neural network (CNN), random forest, one or more decision trees, one or more support vector machines, or a combination thereof, in order to determine, for a given input (e.g., one or more inputs), one or more output values. In certain embodiments, the input includes an image. In certain embodiments, the input includes numerical data, tagged data, and / or functional relationships. In certain embodiments, the input includes alphanumeric data which can include numbers, words, phrases, or lengthier strings, for example. In certain embodiments, the one or more output values include values representing numeric values, words, phrases, or other alphanumeric strings.

[0300] In embodiments, a machine-learned model has been trained using supervised learning algorithm(s), unsupervised learning algorithm(s), semi-supervised learning algorithm(s) (e.g., partial supervision), weak supervision, transfer, multi-task learning, or any combination thereof. In embodiments, a machine-learned model employs a model that includes parameters (e.g., weights) that are tuned during training of the model. For example, the parameters may be adjusted to minimize a loss function, thereby improving the predictive capacity of the machine learning model. A machine-learned model may be further trained after an initial training period, for example, may be adapted to continuously train as it is used.

[0301] In certain embodiments, a machine learned model has been trained, for example, using datasets that include categories of data described herein. Such training may be used to determine various parameters of a machine learned model, for example implemented by a machine learning module, such as, for example, weights associated with layers in neural networks. In certain embodiments, once a machine learned model has been trained, e.g., to accomplish a specific task such as identifying certain output (e.g., extracting certain feature vector(s)), values of determined parameters are fixed and the (e.g., unchanging, static) machine learned model is used to process new data (e.g., different from the training data) and accomplish13403741vl Page 93 of 173Attorney Docket: 2014191-0051its trained task without further updates to its parameters (e.g., the machine learned model does not receive feedback and / or updates). In certain embodiments, a machine learned model may receive feedback, e.g., based on user review of accuracy, and such feedback may be used as additional training data, to dynamically update the machine learned model. In certain embodiments, two or more machine learning models may be combined and implemented as a single model, in a single module, and / or in a single software application. In certain embodiments, two or more machine learned models may also be implemented separately, e.g., as separate software applications. In certain embodiments, two or more machine learned modules may also be implemented separately, e.g., as separate software applications. A machine learned model may be or include software and / or hardware. For example, a machine learned model may be implemented entirely as software, or certain functions of an ANN module (e.g., CNN) may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC)). A machine learned module may be or include software and / or hardware. For example, a machine learned module may be implemented entirely as software, or certain functions of an ANN module (e.g., CNN) may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC)).

[0302] In certain embodiments, machine learning modules implementing machine learning techniques may be composed of individual nodes (e.g. units, neurons). A node may receive a set of inputs that may include at least a portion of a given input data for the machine learning module and / or at least one output of another node. A node may have at least one parameter to apply and / or a set of instructions to perform (e.g., mathematical functions to execute) over the set of inputs. In certain embodiments, node instructions may include a step to provide various relative importance to the set of inputs using various parameters, such as weights. The weights may be applied by performing scalar multiplication (e.g., or other mathematical function) between a set of inputs values and the parameters, resulting in a set of weighted inputs. In certain embodiments, a node may have a transfer function to combine the set of weighted inputs into one output value. A transfer function may be implemented by a summation of all the weighted inputs and the addition of an offset (e.g., bias) value. In certain embodiments, a node may have an activation function to introduce non-linearity into the output value. Non-limiting examples of the activation function include Rectified Linear Activation (ReLu), logistic (e.g., sigmoid), hyperbolic tangent (tanh), and softmax. In certain embodiments,13403741vl Page 94 of 173Attorney Docket: 2014191-0051a node may have a capability of remembering previous states (e.g., recurrent nodes). Previous states may be applied to the input and output values using a set of learning parameters.

[0303] In certain embodiments, the machine learning module includes a deep learning architecture composed of nodes organized into layers. For example, a layer is a set of nodes that receives data input (e.g., weighted or non-weighted input), transforms it (e.g., by carrying out instructions, e.g., applying a set of functions e.g., linear and / or non-linear functions), and passes transformed values as output (e.g., to the next layer). In certain embodiments, the set of nodes in a particular layer may share the same parameters and instructions without interacting with each other. A machine learning module may be composed of at least one layer (e.g., ordered).Examples of types of layers include convolutional layers (e.g., layers with a kernel, a matrix of parameters that is slid across an input to be multiplied with multiple input values to reduce them to a single output value); fully connected (FC) layers (e.g. all nodes are connected to all outputs of the previous layer); recurrent layers, long / short term memory (LSTM) layers, gated recurrent unit (GRU) layers (e.g., nodes with the various abilities to memorize and apply their previous inputs and / or outputs); batch normalization (BN) layers (e.g., layers that normalize a set of outputs from another layer, allowing for more independent learning of individual layers); activation layers (e.g., layers with nodes that only contain an activation function); and / or (un)pooling layers [e.g., layers that reduce (increase) dimensions of an input by summarizing (splitting) input values in defined patches).

[0304] In certain embodiments, the performance of a machine learning module may be characterized by its ability to produce an output data with specific accuracy. To achieve specific accuracy, a training process is performed to find optimal parameters, such as weights, for each node in each layer of the machine learning module. In certain embodiments, the training process of a machine learning module may involve using output data to calculate an objective function (e.g., cost function, loss function, error function) that needs to be optimized (e.g., minimized, maximized). For example, a machine learning objective function may be a combination of a loss function and regularization parameter. The loss function is related to how well the output is able to predict the input. The loss function may take various forms, like mean squared error, mean absolute error, binary cross-entropy, categorical cross-entropy, for example. The regularization term may be needed to prevent overfitting and improve generalization of the training process. Examples of regularization techniques include LI Regularization or Lasso Regression, L213403741vl Page 95 of 173Attorney Docket: 2014191-0051Regularization or Ridge Regression, and Dropout (e.g., dropping layer outputs at random during training process).

[0305] In certain embodiments, objective function optimization of a machine learning module may involve finding at least one (e.g., all) of the present global optima (e.g., as opposed to local optima). In certain embodiments, the algorithm for objective function optimization follows principles of mathematical optimization for a multi-variable function and relies on achieving specific accuracy of the process. Examples of objective function optimization algorithms include gradient descent, nonlinear conjugate gradient, random search, Levenberg-Marquardt algorithm, limited-memory Broyden-Fietcher-Goldfarb-Shanno algorithm, pattern search, basin hopping method, Krylov method, Adam method, genetic algorithm, particle swarm optimization, surrogate optimization, and simulated annealing.

[0306] Computations may be performed locally by a computing device. Computations performed over a network are also contemplated. FIG. 2 shows an illustrative network environment 200 for use in the methods and systems described herein. In brief overview, referring now to FIG. 2, a block diagram of an illustrative cloud computing environment 200 is shown and described. The cloud computing environment 200 may include one or more resource providers 202a, 202b, 202c (collectively, 202). Each resource provider 202 may include computing resources. In some implementations, computing resources may include any hardware and / or software used to process data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, illustrative computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 202 may be connected to any other resource provider 202 in the cloud computing environment 200. In some implementations, the resource providers 202 may be connected over a computer network 208. Each resource provider 202 may be connected to one or more computing device 204a, 204b, 204c (collectively, 204), over the computer network 208.

[0307] The cloud computing environment 200 may include a resource manager 206. The resource manager 206 may be connected to the resource providers 202 and the computing devices 204 over the computer network 208. In some implementations, the resource manager 206 may facilitate the provision of computing resources by one or more resource providers 202 to one or more computing devices 204. The resource manager 206 may receive a request for a13403741v 1 Page 96 of 173Attorney Docket: 2014191-0051computing resource from a particular computing device 204. The resource manager 206 may identify one or more resource providers 202 capable of providing the computing resource requested by the computing device 204. The resource manager 206 may select a resource provider 202 to provide the computing resource. The resource manager 206 may facilitate a connection between the resource provider 202 and a particular computing device 204. In some implementations, the resource manager 206 may establish a connection between a particular resource provider 202 and a particular computing device 204. In some implementations, the resource manager 206 may redirect a particular computing device 204 to a particular resource provider 202 with the requested computing resource.

[0308] FIG. 3 shows an example of a computing device 300 and a mobile computing device 350 that can be used in the methods and systems described in this disclosure. The computing device 300 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 350 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting.

[0309] The computing device 300 includes a processor 302, a memory 304, a storage device 306, a high-speed interface 308 connecting to the memory 304 and multiple high-speed expansion ports 310, and a low-speed interface 312 connecting to a low-speed expansion port 314 and the storage device 306. Each of the processor 302, the memory 304, the storage device 306, the high-speed interface 308, the high-speed expansion ports 310, and the low-speed interface 312, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 302 can process instructions for execution within the computing device 300, including instructions stored in the memory 304 or on the storage device 306 to display graphical information for a GUI on an external input / output device, such as a display 316 coupled to the high-speed interface 308. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade13403741vl Page 97 of 173Attorney Docket: 2014191-0051servers, or a multi -processor system). Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Thus, as the term is used herein, where a plurality of functions are described as being performed by “a processor”, this encompasses embodiments wherein the plurality of functions are performed by any number of processors (e.g., one or more processors) of any number of computing devices (e.g., one or more computing devices).Furthermore, where a function is described as being performed by “a processor”, this encompasses embodiments wherein the function is performed by any number of processors (e.g., one or more processors) of any number of computing devices (e.g., one or more computing devices) (e.g., in a distributed computing system).

[0310] The memory 304 stores information within the computing device 300. In some implementations, the memory 304 is a volatile memory unit or units. In some implementations, the memory 304 is a non-volatile memory unit or units. The memory 304 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0311] The storage device 306 is capable of providing mass storage for the computing device 300. In some implementations, the storage device 306 may be or contain a computer-readable medium, such as a hard disk device, an optical disk device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 302), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices such as computer- or machine-readable mediums (for example, the memory 304, the storage device 306, or memory on the processor 302).

[0312] The high-speed interface 308 manages bandwidth-intensive operations for the computing device 300, while the low-speed interface 312 manages lower bandwidth-intensive operations. Such allocation of functions is an example only. In some implementations, the highspeed interface 308 is coupled to the memory 304, the display 316 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 310, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 312 is coupled to the storage device 306 and the low-speed expansion port 314. The low-speed expansion port 314, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless13403741vl Page 98 of 173Attorney Docket: 2014191-0051Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0313] The computing device 300 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 320, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 322. It may also be implemented as part of a rack server system 324.Alternatively, components from the computing device 300 may be combined with other components in a mobile device (not shown), such as a mobile computing device 350. Each of such devices may contain one or more of the computing device 300 and the mobile computing device 350, and an entire system may be made up of multiple computing devices communicating with each other.

[0314] The mobile computing device 350 includes a processor 352, a memory 364, an input / output device such as a display 354, a communication interface 366, and a transceiver 368, among other components. The mobile computing device 350 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 352, the memory 364, the display 354, the communication interface 366, and the transceiver 368, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0315] The processor 352 can execute instructions within the mobile computing device 350, including instructions stored in the memory 364. The processor 352 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 352 may provide, for example, for coordination of the other components of the mobile computing device 350, such as control of user interfaces, applications run by the mobile computing device 350, and wireless communication by the mobile computing device 350.

[0316] The processor 352 may communicate with a user through a control interface 358 and a display interface 356 coupled to the display 354. The display 354 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 356 may include appropriate circuitry for driving the display 354 to present graphical and other information to a user. The control interface 358 may receive commands from a user and convert them for13403741vl Page 99 of 173Attorney Docket: 2014191-0051submission to the processor 352. In addition, an external interface 362 may provide communication with the processor 352, so as to enable near area communication of the mobile computing device 350 with other devices. The external interface 362 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0317] The memory 364 stores information within the mobile computing device 350. The memory 364 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 374 may also be provided and connected to the mobile computing device 350 through an expansion interface 372, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 374 may provide extra storage space for the mobile computing device 350, or may also store applications or other information for the mobile computing device 350. Specifically, the expansion memory 374 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 374 may be provided as a security module for the mobile computing device 350, and may be programmed with instructions that permit secure use of the mobile computing device 350. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.

[0318] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier and, when executed by one or more processing devices (for example, processor 352), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 364, the expansion memory 374, or memory on the processor 352). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 368 or the external interface 362.

[0319] The mobile computing device 350 may communicate wirelessly through the communication interface 366, which may include digital signal processing circuitry where necessary. The communication interface 366 may provide for communications under various13403741vl Page 100 of 173Attorney Docket: 2014191-0051modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA(time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 368 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth®, Wi-Fi™, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 370 may provide additional navigation- and location-related wireless data to the mobile computing device 350, which may be used as appropriate by applications running on the mobile computing device 350.

[0320] The mobile computing device 350 may also communicate audibly using an audio codec 360, which may receive spoken information from a user and convert it to usable digital information. The audio codec 360 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 350. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 350.

[0321] The mobile computing device 350 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 380. It may also be implemented as part of a smart-phone 382, personal digital assistant, or other similar mobile device.

[0322] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0323] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be13403741vl Page 101 of 173Attorney Docket: 2014191-0051implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0324] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0325] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0326] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0327] In some embodiments, a system of the present disclosure comprises reagents for quantifying one or more histone modifications, chromatin accessibility, binding of one or more13403741vl Page 102 of 173Attorney Docket: 2014191-0051transcription factors, and / or DNA methylation at one or more genomic loci, wherein the one or more genomic loci are selected from Table 1.

[0328] In some embodiments, the system comprises reagents for quantifying H3K4me3 for at least 5, 10, 20, 30, 40, or 50 genomic loci in Table 1. In some embodiments, the system comprises reagents for quantifying H3K27ac for at least 5, 10, 20, 30, 40, or 50 genomic loci in Table 1. In some embodiments, the system comprises one or more antibodies for use in ChlP-seq, optionally wherein the one or more antibodies specifically bind H3K4me3- or H3K27ac-modified histones.

[0329] In some embodiments, the system comprises reagents for quantifying DNA methylation for at least 5, 10, 20, 30, 40, or 50 genomic loci in Table 1. In some embodiments, the system comprises one or more methyl-binding domains for use in MBD-seq.

[0330] In some embodiments, the system comprises reagents for isolation of cell-free DNA (cfDNA) from a liquid biopsy sample. In some embodiments, the sequencer comprises reagents for library preparation for sequencing. In some embodiments, the sequencer comprises reagents for sequencing. In some embodiments, the system comprises instructions for determining if a subject has an ER-positive cancer.

[0331] In some embodiments, the system comprises one or more reagents for enriching for cfDNA having a sequence that falls within or overlaps with one or more genomic loci for which H3K4me3 modifications, H3K27ac modifications, and / or DNA methylation are to be quantified. In some embodiments, one or more reagents for enriching comprise reagents for selectively amplifying. In some embodiments, reagents for selectively amplifying comprise oligonucleotide primers (e.g., oligonucleotide primers for PCR amplification). In some embodiments, one or more reagents for enriching comprise one or more reagents that preferentially bind to one or more sequences within one or more genomic loci. In some embodiments, one or more reagents that preferentially bind one or more sequences within one or more genomic loci comprise one or more oligonucleotides, each comprising a sequence that is complementary to the one or more genomic loci.Certain Exemplary Embodiments

[0332] Without limitation to the foregoing description, the following is an enumerated list of non-limiting exemplary embodiments included in the present disclosure. Those of ordinary skill13403741vl Page 103 of 173Attorney Docket: 2014191-0051in the art will appreciate that one or more features discussed above may be included with or incorporated into any of the following numbered embodiments to form additional embodiments.1. A method of determining ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andwherein the one or more genomic loci comprise one or more expression-level correlated loci for ESRI.2. The method of embodiment 1, wherein the one or more genomic loci comprise one or more genomic loci that are within + / - 200 kB of ESRI, and optionally wherein the one or more genomic loci include one or more of the ESRI associated loci provided in Table 1.3. A method of determining the ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andwherein the one or more genomic loci comprise one or more expression-level correlated loci for ESRI and / or one or more ESRI related genes.4. The method of embodiment 3, wherein the one or more ESRI related genes include ENOL, YBX1, GALA 3, FOXA1, HAP LN 3, EN1, PIM1, CCDC170, or any combination thereof.13403741vl Page 104 of 173Attorney Docket: 2014191-00515. The method of embodiment 3 or 4, wherein the one or more genomic loci comprise:(i) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRJ, and / or(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for one or more of ENOL YBX1, GATA3, FOXAI, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.6. The method of embodiment 5, wherein the one or more expression-level correlated loci for ESRI, ENO J, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170 are proximal to (e.g., within + / - 200 kB of the transcription start site (TSS) of) ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170, respectively.7. The method embodiment of 5 or 6, wherein the one or more expression-level correlated loci for ESRI, ENO1, YBX1, GATA3, FOXAI, HAPLN3, EN1, P1M1, CCDC170, or any combination thereof include one or more promoter regions for one or more of ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.8. The method of any one of embodiments 5-7, wherein the one or more expression-level correlated loci for ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof include one or more enhancer regions for one or more of ESRI, ENO1, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.9. The method of any one of embodiments 5-8, wherein the one or more expression-level correlated loci for ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof, are genomic regions at which signal of one or more epigenetic biomarkers (i) has been shown to be correlated with ESRI expression, and / or (ii) has been shown to be correlated with ENO1, YBXI, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170 expression, respectively.13403741vl Page 105 of 173Attorney Docket: 2014191-005110. The method of any one of embodiments 5-9, wherein(i) the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRI include one or more ESRJ associated loci provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, or 75 or more expression-level correlated loci for ENO1 include one or more ENO 1 associated loci provided in Table 1;(iii) one or more, 5 or more, 10 or more, or 15 or more expression-level correlated loci for YBX1 include one or more YBX1 associated loci provided in Table 1;(iv) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, or 90 or more expression-level correlated loci for GATA3 include one or more GATA3 associated loci provided in Table 1;(v) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for EOXA1 include one or more OXA1 associated loci provided in Table 1;(vi) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 55 or more expression-level correlated loci for HAPLN3 include one or more HAPLN3 associated loci provided in Table 1;(vii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, or 125 or more expression-level correlated loci for EN1 include one or more ENl associated loci provided in Table 1;(viii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more expression-level correlated loci for P IM 1 include one or more PIM1 associated loci provided in Table 1;(ix) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 115 or more expression-level correlated loci for CCDC170 include one or more CCDC170 associated loci13403741vl Page 106 of 173Attorney Docket: 2014191-0051provided in Table 1; or(x) any combination of (i)-(ix).11. The method of any one of embodiments 5-10, wherein the one or more expression-level correlated loci for ESRI include:(i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with ESRI and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 60 or more “H3K27ac” analyte loci that are associated with ESRI and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ESRI and that are provided in Table 1; or (iv) any combination of (i)-(iii).12. The method of any one of embodiments 5-11, wherein the one or more expression-level correlated loci for ENO J include:(i) one or more, 5 or more, or 10 or more “H3K4me3” analyte loci that are associated with ENO1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “H3K27ac” analyte loci that are associated with ENO1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ENO1 and that are provided in Table 1; or (iv) any combination of (i)-(iii).13. The method of any one of embodiments 5-12, wherein the one or more expression-level correlated loci for YBX1 include:(i) the “H3K27ac” analyte loci that is associated with YBX1 and that is provided in Table 1;(ii) one or more, 5 or more, 10 or more, or 15 or more “MBD” analyte loci that are associated with YBX1 and that are provided in Table 1; or(iii) (i) and (ii).13403741vl Page 107 of 173Attorney Docket: 2014191-005114. The method of any one of embodiments 5-13, wherein the one or more expression-level correlated loci for GATA3 include:(i) one or more, 5 or more, 10 or more, or 15 or more “H3K4me3” analyte loci that are associated with GATA3 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 “H3K27ac” analyte loci that are associated with GATA3 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “MBD” analyte loci that are associated with GATA3 and that are provided in Table 1; or(iv) any combination of (i)-(iii).15. The method of any one of embodiments 5-14, wherein the one or more expression-level correlated loci for FOXA1 include:(i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with FO XA1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with FOXA1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with FOXA1 and that are provided in Table 1; or(iv) any combination of (i)-(iii).16. The method of any one of embodiments 5-15, wherein the one or more expression-level correlated loci for HAPLN3 include:(i) one or more, 2 or more, 3 or more, 4 or more, 5 or more, 5 or more, 7 or more, 8 or more, 9 or more, or 10 or more “H3K4me3” analyte loci that are associated with HAPLN3 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K27ac” analyte loci that are associated with HAPLN3 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “MBD” analyte loci13403741vl Page 108 of 173Attorney Docket: 2014191-0051that are associated with HAPLN3 and that are provided in Table 1; or(iv) any combination of (i)-(iii).17. The method of any one of embodiments 5-16, wherein the one or more expression-level correlated loci for EN1 include:(i) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “H3K4me3” analyte loci that are associated with EN1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with EN1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with EN1 and that are provided in Table 1; or (iv) any combination of (i)-(iii).18. The method of any one of embodiments 5-17, wherein the one or more expression-level correlated loci for PIM1 include:(i) one or more, 2 or more, or 3 or more “H3K4me3” analyte loci that are associated with PIM1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “H3K27ac” analyte loci that are associated with PIM1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with PIM1 and that are provided in Table 1; or(iv) any combination of (i)-(iii).19. The method of any one of embodiments 5-18, wherein the one or more expression-level correlated loci for CCDC170 include:(i) one or more, 5 or more, 10 or more, or 20 or more “H3K4me3” analyte loci that are associated with CCDC170 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more “H3K27ac” analyte loci that are associated with CCDC170 and that13403741vl Page 109 of 173Attorney Docket: 2014191-0051are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with CCDC170 and that are provided in Table 1; or(iv) any combination of (i)-(iii).20. The method of any one of embodiments 5-19, wherein:(i) the level of the one or more epigenetic biomarkers at the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci have been shown to be correlated withESKf expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8); or(ii) the level of the one or more epigenetic biomarkers at the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci have been shown to be correlated with ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, ENJ, PIM1, or CCDC170 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8).21. The method of any one of embodiments 1-20, wherein the one or more expression-level correlated loci have been determined to exhibit low or no signal of the one or more epigenetic biomarkers in one or more liquid biopsy samples obtained from one or more healthy subjects (e.g., low or no signal as determined using one or more assays described herein), optionally wherein the one or more expression-level correlated loci have been determined to exhibit low or no signal of the one or more epigenetic biomarkers in 10% or more of healthy subjects.22. The method of any one of embodiments 1-20, wherein the one or more histone modifications are quantified using a histone modification assay that measures one or more of H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, H3K4me3, and panacetylation.23. The method of embodiment 21, wherein the histone modification assay detects H3K4me313403741vl Page 110 of 173Attorney Docket: 2014191-0051modifications.24. The method of embodiment 22 or 23, wherein the histone modification assay detects H3K27ac modifications.25. The method of any one of embodiments 22-24, wherein the histone modification assay is selected from ChlP-seq (Chromatin ImmunoPrecipitation sequencing), CUT& RUN (Cleavage Under Targets and Release Using Nuclease) sequencing, and CUT& Tag (Cleavage Under Targets and Tagmentation) sequencing.26. The method of any one of embodiments 1-25, wherein chromatin accessibility is quantified using an ATAC-seq (Assay of Transpose Accessible Chromatin sequencing) assay, aNOMe-seq (Nucleosome Occupancy and Methylome sequencing) assay, a FAIRE-seq (Formaldehyde-Assisted Isolation of Regulatory Elements sequencing) assay, an MNase-seq (Micrococcal Nuclease digestion with sequencing) assay, a DNase hypersensitivity assay, or a fragmentomics assay.27. The method of any one of embodiments 1-26, wherein the binding of one or more transcription factors is quantified using a transcription factor binding assay that detects binding of one or more of p300, mediator complex, cohesin complex, RNA pol II, FOXA1, ESRI, PR, MYC, EN1, F0XM1, KLF4, AP-2, RARa, or RUNX1.28. The method of embodiment 27, wherein the transcription factor binding assay is ChlP-seq (Chromatin ImmunoPrecipitation sequencing), CUT& RUN (Cleavage Under Targets and Release Using Nuclease) sequencing, or CUT& Tag (Cleavage Under Targets and Tagmentation) sequencing.29. The method of any one of embodiments 1-28, wherein DNA methylation is quantified using Bisulfite sequencing (BS-Seq), Whole Genome Bisulfite Sequencing (WGBS), Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq), or Methyl-CpG-Binding Domain sequencing (MBD-seq).13403741vl Page 111 of 173Attorney Docket: 2014191-005130. The method of any one of embodiments 1-29, comprising quantifying two or more of the epigenetic biomarkers at the one or more genomic loci.31. The method of embodiment 30, comprising quantifying two or more histone modifications.32. The method of embodiment 31, comprising quantifying H3K4me3 and H3K27ac modifications.33. The method of embodiment 30, comprising quantifying one or more histone modifications and DNA methylation.34. The method of embodiment 30, comprising quantifying H3K4me3 and / or H3K27ac modifications and DNA methylation.35. The method of embodiment 34, comprising quantifying H3K4me3 modifications, H3K27ac modifications, and DNA methylation.36. The method of any one of embodiments 1-35, wherein the method comprises:(i) quantifying H3K4me3 modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “H3K4me3” analyte loci provided in Table 1;(ii) quantifying H3K27ac modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “H3K27ac” analyte loci provided in Table 1;(iii) quantifying DNA methylation at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “MBD” analyte loci provided in Table 1; or(iv) any combination of (i)-(iii).37. The method of any one of embodiments 1-36, comprising:13403741vl Page 112 of 173Attorney Docket: 2014191-0051(i) quantifying H3K4me3 modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the genomic loci using an assay that comprises enriching for cfDNA comprising one or more H3K4me3 modifications and sequencing the cfDNA enriched for H3K4me3 modifications (e.g., using a cfChlP-seq assay);(ii) quantifying H3K27ac modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the genomic loci using an assay that comprises enriching for cfDNA comprising one or more H3K27ac modifications and sequencing the cfDNA enriched for H3K27ac modifications (e.g., using a cfChlP-seq assay);(iii) quantifying DNA methylation at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) using an assay that comprises enriching for methylated cfDNA and sequencing the cfDNA enriched for methylated cfDNA (e.g., using a MBD-seq assay); or(iv) any combination of (i)-(iii).38. The method of embodiment 37, wherein:(i) cfDNA comprising H3K4me3 modifications is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody) that binds H3K4me3 modifications;(ii) cfDNA comprising H3K27ac modifications is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody) that binds H3K27ac modifications; and / or(iii) methylated cfDNA is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody or a methyl binding domain) that binds methylated DNA.39. The method of embodiment 38, wherein the agent that binds H3K4me3 modifications, the agent that binds H3K27ac modifications, and / or the agent that binds methylated DNA are attached (e.g., via a covalent or noncovalent bond) to a physical support (e.g., a bead, a magnetic bead, an agarose bead, or a magnetic epoxy bead) prior to incubating with the sample.40. The method of embodiment 38 or 39, wherein the method comprises incubating with two or more of (i) the agent that binds H3K4me3 modifications, (ii) the agent that binds H3K27ac13403741vl Page 113 of 173Attorney Docket: 2014191-0051modifications, and (iii) the agent that binds methylated DNA, andwherein the sample is incubated with the two or more agents in sequence or in parallel (e.g., wherein the sample is divided into fractions and each fraction is incubated with a different agent).41. The method of any...

Claims

Attorney Docket: 2014191-0051CLAIMSWhat is claimed is:

1. A method of determining ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(i v) DN A m ethyl ati on; an dwherein the one or more genomic loci comprise one or more expression-level correlated loci for ESR 1.

2. The method of claim 1, wherein the one or more genomic loci comprise one or more genomic loci that are within + / - 200 kB of ESRI, and optionally wherein the one or more genomic loci include one or more of the ESRJ associated loci provided in Table 1.

3. A method of determining the ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andwherein the one or more genomic loci comprise one or more expression-level correlated loci for ESRI and / or one or more ESRI related genes.

4. The method of claim 3, wherein the one or more ESRI related genes include ENO1, YBXI, GATA3, FOXA1, EIAPLN3, EN1, PIM1, CCDC170, or any combination thereof.13403741vl Page 154 of 173Attorney Docket: 2014191-00515. The method of claim 3 or 4, wherein the one or more genomic loci comprise:(i) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRI, and / or(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for one or more of ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.

6. The method of claim 5, wherein the one or more expression-level correlated loci for ESRI, ENO J, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170 are proximal to (e.g., within + / - 200 kB of the transcription start site (TSS) of) ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170, respectively.

7. The method of claim 5 or 6, wherein the one or more expression-level correlated loci for ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof include one or more promoter regions for one or more of ESRI, ENO1, YBX1, GA TA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.

8. The method of any one of claims 5-7, wherein the one or more expression-level correlated loci for ESRJ, ENO J, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof include one or more enhancer regions for one or more of ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof.

9. The method of any one of claims 5-8, wherein the one or more expression-level correlated loci for ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, CCDC170, or any combination thereof, are genomic regions at which signal of one or more epigenetic biomarkers (i) has been shown to be correlated with ESRI expression, and / or (ii) has been shown to be correlated with ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170 expression, respectively.

10. The method of any one of claims 5-9, wherein13403741vl Page 155 of 173Attorney Docket: 2014191-0051(i) the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci for ESRI include one or more ESRI associated loci provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, or 75 or more expression-level correlated loci for ENO1 include one or more ENO1 associated loci provided in Table 1;(iii) one or more, 5 or more, 10 or more, or 15 or more expression-level correlated loci for YBX1 include one or more YBX1 associated loci provided in Table 1;(iv) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, or 90 or more expression-level correlated loci for GATA3 include one or more GATA3 associated loci provided in Table 1;(v) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci for FOXA1 include one or more FOXA1 associated loci provided in Table 1;(vi) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 55 or more expression-level correlated loci for HAPLN3 include one or more HAPLN3 associated loci provided in Table 1;(vii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, or 125 or more expression-level correlated loci for EN1 include one or more EN1 associated loci provided in Table 1;(viii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more expression-level correlated loci for PIM1 include one or more PIM1 associated loci provided in Table 1;(ix) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 115 or more expression-level correlated loci for CCDC170 include one or more CCDC170 associated loci provided in Table 1; or(x) any combination of (i)-(ix).13403741vl Page 156 of 173Attorney Docket: 2014191-005111. The method of any one of claims 5-10, wherein the one or more expression-level correlated loci for ESRI include:(i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated withElSKY and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 60 or more “H3K27ac” analyte loci that are associated with ESRI and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ESRI and that are provided in Table 1; or (iv) any combination of (i)-(iii).

12. The method of any one of claims 5-11, wherein the one or more expression-level correlated loci for ENO1 include:(i) one or more, 5 or more, or 10 or more “H3K4me3” analyte loci that are associated with ENO1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “H3K27ac” analyte loci that are associated with ENO1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ENO1 and that are provided in Table 1; or (iv) any combination of (i)-(iii).

13. The method of any one of claims 5-12, wherein the one or more expression-level correlated loci for YBX1 include:(i) the “H3K27ac” analyte loci that is associated with YBX1 and that is provided in Table 1;(ii) one or more, 5 or more, 10 or more, or 15 or more “MBD” analyte loci that are associated with YBX1 and that are provided in Table 1; or(iii) (i) and (ii).

14. The method of any one of claims 5-13, wherein the one or more expression-level correlated13403741vl Page 157 of 173Attorney Docket: 2014191-0051loci for GATA3 include:(i) one or more, 5 or more, 10 or more, or 15 or more “H3K4me3” analyte loci that are associated with GATA3 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 “H3K27ac” analyte loci that are associated with GATA3 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “MBD” analyte loci that are associated with GATA3 and that are provided in Table 1; or(iv) any combination of (i)-(iii).

15. The method of any one of claims 5-14, wherein the one or more expression-level correlated loci for FOXA1 include:(i) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K4me3” analyte loci that are associated with FOXA1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with FOXA1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with FOXA1 and that are provided in Table 1; or(iv) any combination of (i)-(iii).

16. The method of any one of claims 5-15, wherein the one or more expression-level correlated loci for HAPLN3 include:(i) one or more, 2 or more, 3 or more, 4 or more, 5 or more, 5 or more, 7 or more, 8 or more, 9 or more, or 10 or more “H3K4me3” analyte loci that are associated with HAPLN3 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “H3K27ac” analyte loci that are associated with HAPLN3 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, or 20 or more “MBD” analyte loci that are associated with HAPLN3 and that are provided in Table 1; or(iv) any combination of (i)-(iii).13403741vl Page 158 of 173Attorney Docket: 2014191-005117. The method of any one of claims 5-16, wherein the one or more expression-level correlated loci for ENl include:(i) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “H3K4me3” analyte loci that are associated with ENl and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 70 or more “H3K27ac” analyte loci that are associated with ENl and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, or 30 or more “MBD” analyte loci that are associated with ENl and that are provided in Table 1; or (iv) any combination of (i)-(iii).

18. The method of any one of claims 5-17, wherein the one or more expression-level correlated loci for PIM1 include:(i) one or more, 2 or more, or 3 or more “H3K4me3” analyte loci that are associated with PIM1 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more “H3K27ac” analyte loci that are associated with PIM1 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD” analyte loci that are associated with ZMZ and that are provided in Table 1; or(iv) any combination of (i)-(iii).

19. The method of any one of claims 5-18, wherein the one or more expression-level correlated loci for CCDC170 include:(i) one or more, 5 or more, 10 or more, or 20 or more “H3K4me3” analyte loci that are associated with CCDC170 and that are provided in Table 1;(ii) one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, or 65 or more “H3K27ac” analyte loci that are associated with CCDC170 and that are provided in Table 1;(iii) one or more, 5 or more, 10 or more, 15 or more, 20 or more, or 25 or more “MBD”13403741vl Page 159 of 173Attorney Docket: 2014191-0051analyte loci that are associated with CCDC170 and that are provided in Table 1; or(iv) any combination of (i)-(iii).

20. The method of any one of claims 5-19, wherein:(i) the level of the one or more epigenetic biomarkers at the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, or 110 or more expression-level correlated loci have been shown to be correlated with ESRI expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8); or(ii) the level of the one or more epigenetic biomarkers at the one or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, or 120 or more expression-level correlated loci have been shown to be correlated with ESRI, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM1, or CCDC170 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8).

21. The method of any one of claims 1-20, wherein the one or more expression-level correlated loci have been determined to exhibit low or no signal of the one or more epigenetic biomarkers in one or more liquid biopsy samples obtained from one or more healthy subjects (e.g., low or no signal as determined using one or more assays described herein), optionally wherein the one or more expression-level correlated loci have been determined to exhibit low or no signal of the one or more epigenetic biomarkers in 10% or more of healthy subjects.

22. The method of any one of claims 1-20, wherein the one or more histone modifications are quantified using a histone modification assay that measures one or more of H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, H3K4me3, and pan-acetylation.

23. The method of claim 21, wherein the histone modification assay detects H3K4me3 modifications.

24. The method of claim 22 or 23, wherein the histone modification assay detects H3K27ac13403741vl Page 160 of 173Attorney Docket: 2014191-0051modifications.

25. The method of any one of claims 22-24, wherein the histone modification assay is selected from ChlP-seq (Chromatin ImmunoPrecipitation sequencing), CUT& RUN (Cleavage Under Targets and Release Using Nuclease) sequencing, and CUT& Tag (Cleavage Under Targets and Tagmentation) sequencing.

26. The method of any one of claims 1-25, wherein chromatin accessibility is quantified using an ATAC-seq (Assay of Transpose Accessible Chromatin sequencing) assay, a NOMe-seq (Nucleosome Occupancy and Methylome sequencing) assay, a FAIRE-seq (Formaldehyde-Assisted Isolation of Regulatory Elements sequencing) assay, an MNase-seq (Micrococcal Nuclease digestion with sequencing) assay, a DNase hypersensitivity assay, or a fragmentomics assay.

27. The method of any one of claims 1-26, wherein the binding of one or more transcription factors is quantified using a transcription factor binding assay that detects binding of one or more of p300, mediator complex, cohesin complex, RNApol II, FOXA1, ESRI, PR, MYC, EN1, F0XM1, KLF4, AP-2, RARa, or RUNX1.

28. The method of claim 27, wherein the transcription factor binding assay is ChlP-seq (Chromatin ImmunoPrecipitation sequencing), CUT& RUN (Cleavage Under Targets and Release Using Nuclease) sequencing, or CUT& Tag (Cleavage Under Targets and Tagmentation) sequencing.

29. The method of any one of claims 1-28, wherein DN A methylation is quantified using Bisulfite sequencing (BS-Seq), Whole Genome Bisulfite Sequencing (WGBS), Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq), or Methyl-CpG-Binding Domain sequencing (MBD-seq).

30. The method of any one of claims 1-29, comprising quantifying two or more of the epigenetic biomarkers at the one or more genomic loci.13403741vl Page 161 of 173Attorney Docket: 2014191-005131. The method of claim 30, comprising quantifying two or more histone modifications.

32. The method of claim 31, comprising quantifying H3K4me3 and H3K27ac modifications.

33. The method of claim 30, comprising quantifying one or more histone modifications and DNA methylation.

34. The method of claim 30, comprising quantifying H3K4me3 and / or H3K27ac modifications and DNA methylation.

35. The method of claim 34, comprising quantifying H3K4me3 modifications, H3K27ac modifications, and DNA methylation.

36. The method of any one of claims 1-35, wherein the method comprises:(i) quantifying H3K4me3 modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “H3K4me3” analyte loci provided in Table 1;(ii) quantifying H3K27ac modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “H3K27ac” analyte loci provided in Table 1;(iii) quantifying DNA methylation at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the “MBD” analyte loci provided in Table 1; or(iv) any combination of (i)-(iii).

37. The method of any one of claims 1-36, comprising:(i) quantifying H3K4me3 modifications at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the genomic loci using an assay that comprises enriching for cfDNA comprising one or more H3K4me3 modifications and sequencing the cfDNA enriched for H3K4me3 modifications (e.g., using a cfChlP-seq assay);13403741vl Page 162 of 173Attorney Docket: 2014191-0051(ii) quantifying H3K27ac modifications at one or more (e g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) of the genomic loci using an assay that comprises enriching for cfDNA comprising one or more H3K27ac modifications and sequencing the cfDNA enriched for H3K27ac modifications (e.g., using a cfChlP-seq assay);(iii) quantifying DNA methylation at one or more (e.g., 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more) using an assay that comprises enriching for methylated cfDNA and sequencing the cfDNA enriched for methylated cfDNA (e g., using aMBD-seq assay); or(iv) any combination of (i)-(iii).

38. The method of claim 37, wherein:(i) cfDNA comprising H3K4me3 modifications is enriched using a method that comprises incubating the sample with an agent (e g., an antibody) that binds H3K4me3 modifications;(ii) cfDNA comprising H3K27ac modifications is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody) that binds H3K27ac modifications; and / or(iii) methylated cfDNA is enriched using a method that comprises incubating the sample with an agent (e.g., an antibody or a methyl binding domain) that binds methylated DNA.

39. The method of claim 38, wherein the agent that binds H3K4me3 modifications, the agent that binds H3K27ac modifications, and / or the agent that binds methylated DNA are attached (e.g., via a covalent or noncovalent bond) to a physical support (e.g., a bead, a magnetic bead, an agarose bead, or a magnetic epoxy bead) prior to incubating with the sample.

40. The method of claim 38 or 39, wherein the method comprises incubating with two or more of (i) the agent that binds H3K4me3 modifications, (ii) the agent that binds H3K27ac modifications, and (iii) the agent that binds methylated DNA, andwherein the sample is incubated with the two or more agents in sequence or in parallel (e.g., wherein the sample is divided into fractions and each fraction is incubated with a different agent).13403741vl Page 163 of 173Attorney Docket: 2014191-005141. The method of any one of claims 37-40, wherein the sequencing is performed using a next generation sequencing method.

42. The method of any one of claims 37-41, wherein the method comprises attaching (e.g., covalently attaching) DNA sequencing adapters to cfDNA obtained from the subject.

43. The method of claim 42, wherein the DNA sequencing adapters are attached after cfDNA has been enriched for cfDNA comprising one or more H3K4me3 modifications, cfDNA comprising one or more H3K27ac modifications, methylated cfDNA, or any combination thereof.

44. The method of claim 42 or 43, comprising amplifying the cfDNA attached to the DNA sequencing adapters.

45. The method of any one of claims 37-44, comprising mapping sequence reads to a reference genome.

46. The method of any one of claims 37-45, wherein non-uniquely mapped and redundant sequence reads are discarded.

47. The method of any one of claims 37-46, wherein quantifying H3K4me3 modifications, H3K27ac modifications, methylated DNA, or any combination thereof at each of the one or more genomic loci comprises summing the number of sequence reads having at least one nucleotide overlap with each of the one or more genomic loci.

48. The method of claim 47, wherein the number of sequence reads at each of the one or more genomic loci is adjusted on the basis of sequencing depth (e.g., quantile normalizing sequence reads to a common reference distribution) and / or ChIP quality, and wherein the adjusting is done prior to or subsequent to summing the number of sequence reads having at least one nucleotide overlap with each of the one or more genomic loci.13403741vl Page 164 of 173Attorney Docket: 2014191-005149. The method of claim 47 or 48, wherein an estimate of local background signal is subtracted from the sequence reads at each genomic loci prior to summing.

50. The method of any one of claims 37-49, comprising calculating sequence read density at the one or more genomic loci.

51. The method of claim 50, wherein the sequence read density is calculated using a method that comprises:(i) summing background adjusted sequence counts at each of the one or more genomic loci, and(ii) dividing the sum of the background adjusted sequence counts by the combined sum of the length (e.g., the number of nucleotides) of the one or more genomic loci.

52. The method of claim 50, wherein the sequence read density is calculated using a method that comprises:(i) for each genomic loci, dividing the background adjusted fragment count by the length (e.g., number of nucleotides) of the genomic loci, and(ii) summing the resulting value of (i) for each of the one or more genomic loci.

53. The method of any one of claims 37-52, wherein sequence reads are normalized to aggregate counts in a given sample across a set of regions (e.g., 10,000 regions) previously determined to have DNAse hypersensitivity in most cell types.

54. The method of any one of claims 1-53, wherein the method comprises determining:(i) a point estimate of H3K27ac modifications,(ii) a point estimate of H3K4me3 modifications,(iii) a point estimate of methylated DNA; or(iv) any combination of (i)-(iii).

55. The method of claim 54, wherein the point estimate is a mean.13403741V1 Page 165 of 173Attorney Docket: 2014191-005156. The method of claim 55, wherein the point estimate is a geometric mean.

57. The method of any one of claims 1-56, comprising inputting the values obtained by quantifying the one or more epigenetic biomarkers into a model.

58. The method of claim 57, wherein the model was produced based on (i) a signal of one or more of the epigenetic markers at one or more of the genomic loci provided in Table 1 measured in one or more cancer cell lines, and / or (ii) a signal of one or more of the epigenetic markers at one or more of the genomic loci provided in Table 1 measured in one or more samples obtained from one of more subjects having the cancer.

59. The method of claim 58, wherein the model was produced based on the signal of one or more of the epigenetic markers at one or more of the genomic loci provided in Table 1 measured in (i) one or more ER-positive cancer cell lines and (i) one or more ER-negative cell lines (e.g., ERnegative cancer cell lines), optionally wherein the ER-positive cancer cell line and / or the ERnegative cancer cell line are each breast cancer cell lines.

60. The method of claim 58, wherein the model was produced based on the signal of one or more of the epigenetic markers at one or more of the genomic loci provided in Table 1 measured in one or more samples obtained from one of more subjects having a ER-positive cancer and one or more subjects having a ER-negative cancer or one or more healthy subjects.

61. The method of any one of claims 57-61, wherein the model inputs were point estimates of the one or more epigenetic biomarkers.

62. The method of claim 61, wherein the point estimates are means or geometric means of the one or more epigenetic biomarkers.

63. The method of any one of claims claim 57-62, wherein the model was produced by a method that comprised regressing the values obtained by quantifying the one or more epigenetic biomarkers against measured ER expression.13403741vl Page 166 of 173Attorney Docket: 2014191-005164. The method of claim 63, wherein the regressing was performed using an ordinary least squares (OLS) regression method.

65. The method of any one of claims 1-64, further comprising comparing the quantified amount of the one or more epigenetic biomarkers to a reference.

66. The method of claim 65, wherein the reference is a predetermined threshold, a measurement from a liquid biopsy sample, a measurement from liquid biopsy samples obtained from a cohort of subjects, and / or a normalized value.

67. The method of 66, wherein the predetermined threshold and / or the normalized value distinguish ER-positive and ER-negative cancers with an AUC of 0.5 or greater (e g., 0.8 or greater).

68. The method of any one of claims 65-67, wherein the reference is a measurement from a liquid biopsy sample obtained from a cohort of subjects who have previously been determined to have a ER-positive or a ER-negative cancer.

69. The method of any one of claims 65-67, wherein the reference is a measurement from a liquid biopsy sample obtained from a cohort of healthy subjects.

70. The method of any one of claims 1-69, wherein the method determines (i) whether the cancer is ER-positive (ER+) or ER-negative (ER-), (ii) the percentage of cells in the cancer that express ER (e.g., that would stain positive for ER expression using an IHC and / or ISH assay), and / or (iii) an Allred score between 0 and 8 for the cancer;optionally, wherein the ER-positive cancer correlates with an ER+ Allred score of 3, 4, 5, 6, 7 or 8 based on IHC testing and / or the ER-negative cancer correlates with an Allred score of 0, 1, or 2 based on IHC testing.

71. The method of any one of claims 1-70, wherein the liquid biopsy sample is a plasma sample,13403741vl Page 167 of 173Attorney Docket: 2014191-0051serum sample, or urine sample.

72. The method of any one of claims 1-71, wherein the method comprises purifying DNA(e.g., cfDNA) from about 1 mL, about 2 mL, about 3 mL, about 4 mL, or about 5 mL of the liquid biopsy sample (e.g., plasma sample).

73. A method of treating a subject having a cancer, the method comprising:administering a cancer therapy to the subject based on the ER status of the cancer, wherein the ER status of the cancer has been determined using the method of any one of claims 1-72.

74. The method of claim 73, wherein:(i) if the cancer is determined to be ER-positive, the method comprises administering a therapy for treating a ER-positive cancer, and(ii) if the cancer is determined to be ER-negative, the method comprises administering a therapy for treating a ER-negative cancer.

75. The method of claim 74, wherein the therapy for treating a ER-positive cancer comprises administering a ER-targeted agent.

76. A method of monitoring the ER status of a cancer in a subject, and optionally treating the cancer, the method comprising:determining the ER status of the cancer using the method of any one of claims 1-72 at a first and a second time point.

77. The method of claim 76, wherein the subject has been administered a cancer therapy before the first time point or at the first time point, or wherein the subject has been administered a cancer therapy after the first time point and before the second time point.

78. The method of claim 76 or 77, further comprising administering a cancer therapy to the subject based on the ER status of the cancer at the second time point and / or a change in ER status between the first time point and the second time point.13403741vl Page 168 of 173Attorney Docket: 2014191-005179. The method of claim 78, wherein the type, dose and / or frequency of administration of the cancer therapy is adjusted based on the ER status of the cancer at the second time point and / or the change in ER status between the first time point and the second time point.

80. The method of claim 78, wherein:(i) if the cancer is ER-positive at the second time point, the method comprises administering a ER-targeted therapeutic to the subject, and(ii) if the cancer is ER-negative at the second time point, the method does not comprise administering a ER-targeted therapeutic to the subject.

81. A kit comprising reagents for quantifying one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci, wherein the one or more genomic loci are selected from Table 1.

82. The kit of claim 81, wherein the kit comprises reagents for quantifying:(i) H3K4me3 modifications, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 or more genomic loci in Table 1;(ii) H3K27ac modifications, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1;(iii) DNA methylation, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1; or(iv) any combination of (i)-(iii);optionally, wherein the reagents for quantifying comprise one or more reagents for enriching for the one or more genomic loci for which H3K4me3 modifications, H3K27ac modifications, and / or DNA methylation is being quantified (e.g., reagents for selectively amplifying (e.g., PCR primers) or preferentially binding (e.g., using nucleic acid bait molecules) the one or more genomic loci).

83. The kit of claim 81 or 82, wherein the kit comprises one or more antibodies for use in ChlP-seq, optionally wherein the one or more antibodies specifically bind H3K4me3- or H3K27ac-13403741V1 Page 169 of 173Attorney Docket: 2014191-0051modified histones.

84. The kit of any one of claims 81-83, wherein the kit comprises one or more methyl-binding domains for use in MBD-seq or wherein the kit comprises one or more antibodies that bind methylated DNA for use in MeDIP-seq.

85. The kit of any one of claims 81-84, wherein the kit comprises reagents for isolation of cell-free DNA (cfDNA) from a liquid biopsy sample.

86. The kit of any one of claims 81-85, wherein the kit comprises reagents for library preparation for sequencing.

87. The kit of any one of claims 81-86, wherein the kit comprises reagents for sequencing.

88. The kit of any one of claims 81-87, wherein the kit comprises instructions for determining if a subject has a ER-positive cancer.

89. A non-transitory computer readable storage medium encoded with a computer program, wherein the program comprises instructions that when executed by one or more processors cause the one or more processors to perform operations to perform the method of any one of claims 1-72.

90. A computer system comprising a memory and one or more processors coupled to the memory, wherein the one or more processors are configured to perform operations to perform the method of any one of claims 1-72.

91. A system for determining the ER status of a cancer in a subject, the system comprising a sequencer configured to generate a sequencing dataset from a sample; and a non-transitory computer readable storage medium of claim 89 and / or a computer system of claim 90.

92. The system of claim 91, wherein the sequencer is configured to generate a Whole Genome13403741vl Page 170 of 173Attorney Docket: 2014191-0051Sequencing (WGS) dataset from the sample.

93. The system of claim 91 or 92, further comprising a sample preparation device configured to prepare the sample for sequencing from a biological sample, optionally a liquid biopsy sample.

94. The system of claim 93, wherein the sample preparation device comprises reagents for quantifying one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci in cell-free DNA (cfDNA) from the biological sample, optionally the liquid biopsy sample.

95. The system of claim 94, wherein the one or more genomic loci comprise one or more genomic loci provided in Table 1.

96. The system of claim 95, wherein the device comprises reagents for quantifying:(i) H3K4me3 modifications, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1;(ii) H3K27ac modifications, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1;(iii) DNA methylation, e.g., for at least 5, at least 10, at least 20, at least 30, at least 40, or at least 50 genomic loci in Table 1;(iv) any combination of (i)-(iii);optionally, wherein the reagents for quantifying comprise one or more reagents for enriching for the one or more genomic loci for which H3K4me3 modifications, H3K27ac modifications, and / or DNA methylation is being quantified (e.g., reagents for selectively amplifying (e.g., PCR primers) and / or preferentially binding (e.g., nucleic acid bait molecules) the one or more genomic loci).

97. The system of any one of claims 94-96, wherein the reagents comprise one or more antibodies for use in ChlP-seq, optionally wherein the one or more antibodies specifically bind H3K4me3- or H3K27ac-modified histones.13403741v1 Page 171 of 173Attorney Docket: 2014191-005198. The system of claim 96 or 97, wherein the reagents comprise one or more methyl -binding domains for use in MBD-seq.

99. The system of any one of claims 91-98, wherein the device comprises reagents for isolation of cell-free DNA (cfDNA) from the biological sample, optionally the liquid biopsy sample.

100. The system of any one of claims 91-99, wherein the device comprises reagents for library preparation for sequencing.

101. The system of any one of claims 91-100, wherein the sequencer comprises reagents for sequencing.13403741v1 Page 172 of 173