Systems and methods for predicting expression levels of target genes for indications
Patent Information
- Application Number
- PCT/US2026/020459
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-10-10
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure US2026020459_01102026_PF_FP_ABST
Abstract
Description
Attorney Docket: 2014191-0049SYSTEMS AND METHODS FOR PREDICTING EXPRESSION LEVELS OF TARGET GENES FOR INDICATIONSBACKGROUND
[0001] Cancer ranks as the primary or secondary cause of premature death in many countries. Antibody-drug conjugates (ADCs) were ushered into oncology clinical practice in 2000 with the FDA’s approval of Mylotarg™ for the treatment of acute myeloid leukemia (AML). ADC molecules marry the precision of antibody-mediated tumor antigen targeting with potent cytotoxic agents, thereby creating a targeted delivery vehicle for malignant tumors. In this manner, ADCs provide a means to reduce off-tumor toxicities by limiting payload exposure in normal tissues. While most ADC clinical candidates utilize cytotoxic chemotherapeutic payloads, recent ADC candidates have also incorporated targeted small molecules and immunomodulatory agents. Since Mylotarg™’ s first registration, hundreds of ADCs have been evaluated in the investigational setting and several have made it to regulatory approval including Adcetris™, Kadcyla™, Besponsa™, Enhertu™, Padcev™, Polivy™, Blenrep™, Todelvy™, Tivdak™, Zylonta™ and Elahere™
[0002] Target antigen expression is currently used as the primary biomarker for patient selection and ADC efficacy has been shown to correlate with the level of target antigen expression in some studies. However, IHC methods for quantifying target antigen expression are invasive, requiring tissue biopsies. In addition, for many targets, there is limited information about their expression across a variety of tumor types, as well as intratumoral heterogeneity and evolution of expression with tumor progression and metastasis. The effects of different intervening therapies on expression are also largely unknown. Data also suggests that there can be clinical efficacy in patients who also have low target antigen expression, shifting toward a threshold expression that is still unknown for many target antigens over which ADCs have therapeutic efficacy.
[0003] Target assessment has also been challenging as a result of known challenges with IHC assessment / interpretation. There therefore remains a need for new biomarkers and improved methods predicting expression levels for target genes in samples, for example for use in developing ADC therapies and / or identifying subjects for administration of particular ADC therapies.13404083v 1 Page 1 of 258Attorney Docket: 2014191-0049SUMMARY
[0004] Disclosed herein are, inter alia, systems and methods for producing models that predict an expression level of a target gene in a sample for an indication, such as cancer, based on one or more epigenetic biomarkers, such as histone modification and / or DNA methylation, for the sample, which may be used to determine whether and how expression level of the target is correlated with one or more epigenetic biomarkers for the indication. Digital samples that are generated in silica may be used to produce the models, thereby reduce the amount of empirical data that need to be collected and significantly shortening the amount of time needed to generate sufficient data. For example, sequencing data from cell samples (e.g., cell lines) and health volunteers corresponding to one or more epigenetic biomarkers may be sampled to form digital samples. Moreover, digital samples may be derived from liquid biopsy samples, for example liquid plasma samples, alleviating or minimizing biopsy burden faced in other methods such as IHC -based methods. Models may be produced by determining correlation between expression level and signal for one or more epigenetic biomarkers across a genomic region corresponding to a target gene, for example by breaking the genomic region up into tiles and looking for correlation within each of the tiles to determine expression-level correlated tiles that are subregions of a genomic region that have high signal to expression level correlation. A model may be produced (e.g, fit and / or trained) based on signal for one or more epigenetic biomarkers in expression-level correlated tiles and expression level for digital samples. Point estimates, such as a geometric mean, that characterize signal across a set of identified expression-level correlated tiles may be used when producing a model. Geometric mean in particular may be robust outliers and / or allow for some dependency between individual underlying measurements while still producing sufficiently performant models (e.g., that can be subsequently tuned and / or feature engineered to increase performance).
[0005] Models produced using methods and systems disclosed herein can be tested to determine if they are performant, for example within one or more predefined criteria. Because samples can be chosen that all correspond to a particular indication, determination that one or more epigenetic biomarkers are correlated with expression level for a target gene to a significant degree, such that a resulting model produced based on that correlation is performant, can be used to identify relevant target genes for indications, for example for use in antibody-drug conjugate (ADC) and / or ADC therapy development. Predictions of expression level for samples, for13404083 v 1 Page 2 of 258Attorney Docket: 2014191-0049example derived from patients that are or may be candidates for receiving an ADC therapy, may be made using models disclosed herein. Methods and systems disclosed herein may be used to determine that there is a predictable expression level of a target gene for an indication. An ADC therapy may target an antigen encoded by a target gene based on a predicted expression level for the target gene for an indication as determined by a system or method disclosed herein. Information derived from systems and methods disclosed herein may be used in developing ADCs for indications for which there is not an existing ADC on market by identifying one or more target genes or targeting a new gene for an indication not corresponding to any existing ADC for that indication on market. Systems and methods disclosed herein may be used to determine and / or predict an expression level of a target gene in a subject at one or more points in time. Systems and methods disclosed herein may be used to diagnose, prognose, and / or monitor subjects, for example subjects having a cancer.
[0006] Systems and methods disclosed herein may be used to rapidly (e.g., automatically) produce models that can be used, and / or be tuned and / or feature engineered to be used, to predict expression level of a target gene for a subject based on epigenetic biomarker signal for a sample (e g, a liquid biopsy sample, such as a plasma sample) derived from the subject. Models can be produced for a large number of target genes for a particular indication rapidly (e.g., automatically), for example to screen a large library of target genes quickly for each of one or more indications. Models for a library of target genes can be generated quickly based on new digital samples. Large datasets of digital samples for one or more epigenetic biomarkers can be generated quickly from relatively small amounts of source data obtained from cell sample sources, such as cell lines, and healthy volunteers for a particular indication of interest. Where source data is already available, for example because a cell sample source has already been sequenced for one or more epigenetic biomarkers of interest, this process can happen quite quickly. Even if such source data need be obtained experimentally, the time needed to obtain sufficient data can be reduced based on the ability to generate large numbers of digital samples from a small amount of source data, for example by combining it with healthy volunteer in many different combinations to generate unique digital samples. Models for target genes for new indications can be rapidly produced by generating new digital samples corresponding to the new indications (e.g., using different cell lines).
[0007] Systems and methods disclosed herein may produce models based on digital samples and / or use such models, for example to predict expression level of a target gene. Such13404083 v 1 Page 3 of 258Attorney Docket: 2014191-0049digital samples may include data, such as sequencing data, derived from samples such as liquid biopsy samples (e.g., liquid plasma samples) and / or that simulate data derived from such liquid biopsy samples [e.g., by combining cell sample (e.g., cell line derived) data and healthy volunteer data]. Using such digital samples can be advantageous in that a model produced therefrom (or produced and subsequently tuned as described herein) can predict an expression level for a target gene in a subject having an indication based data derived from a liquid biopsy sample, such as a liquid plasma sample, from the subject. Such liquid samples from a subject can be obtained using a minimally invasive procedure (e.g., blood draw) and relatively small sample volume (e.g., about 1-3 mL), so models that can be applied to such samples are desirable. The present disclosure includes the recognition that unexpectedly models produced using digital samples derived from a combination of data from cell samples, such as cell line samples, and healthy volunteer samples (e g., that together simulate data derived from liquid biopsy samples) can be performant when used to predict the tumor-specific transcriptional expression (e.g., RNA-seq) level of a target gene from liquid biopsy (e.g., plasma) samples from subjects with cancer. Such performance is unexpected given the differences between the sources of the data for the digital samples (cell data) and the input data. It is additionally unexpected given that previously all gene expression models had been bespoke with development against an orthogonal standard such as IHC - here we demonstrate the ability to automate the creation of the prediction of tumor-specific RNA expression. Systems and methods disclosed herein may be used to predict expression level of a target gene in a sample having an indication based on one or more epigenetic biomarkers detectable in a liquid biopsy sample (e.g., based on cfDNA in the sample). Systems and methods disclosed herein may be used to identify candidate target genes for an indication using one or more epigenetic biomarkers detectable in a liquid biopsy sample (e.g., based on cfDNA in the sample).
[0008] Where an indication is a cancer indication, certain samples, like liquid plasma samples, will have cell free DNA (cfDNA) in them that includes circulating tumor DNA (ctDNA). Each digital sample may therefore correspond to a particular ctDNA fraction (fraction of total cfDNA for the sample that is ctDNA). A model may correspond to a particular ctDNA fraction, for example based on having been produced using samples that all correspond to the ctDNA fraction. Models may be produced for samples having a high ctDNA fraction, for example a ctDNA fraction that is in a range of 5-10%. Clinically relevant ctDNA fractions are often lower, for example in a range of 1-3%. Nonetheless, the ability to produce a performant model at high1340408 v 1 Page 4 of 258Attorney Docket: 2014191-0049ctDNA fraction may be indicative that a target gene is relevant to an indication A method can be iterated, for example performing additional loops of identifying expression-level correlated tiles and producing a model therefrom, using samples at lower ctDNA fraction to see if sufficient model performance can be maintained at lower (e.g., clinically relevant) ctDNA fraction. In general, as ctDNA fraction is lowered, it is expected that fewer expression-level correlated tiles will be identified; generally lower ctDNA fraction samples will have less discernable signal against background so fewer tiles will show correlation. In some embodiments, a model -production loop is performed for a number of iterations using subsets of samples all corresponding to the same initial (e.g., relatively high) ctDNA fraction (e.g., 10% or 5%) and then, if the predictions from those iterations are within one or more predefined criteria, performed for a second number of iterations using subsets of samples all corresponding to a second, lower ctDNA fraction (e.g., 5% or 3%). Successful prediction within one or more predefined criteria may lead to tuning a model to perform at lower ctDNA fraction (e.g., clinically relevant fractions of 1-3%) and / or exhibit better performance (e.g., based on the one or more predefined criteria).
[0009] In some aspects, the present disclosure is directed to a method. The method may include receiving digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene. The method may include tiling the samples into tiles that together span a genomic region corresponding to the target gene. The method may include performing a loop for each of at least one subset of the samples. The loop may include determining a set of expression-level correlated tiles for the subset of the samples. Determining the set of expression-level correlated tiles may include testing each of the tiles for correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level across the subset of the samples, for example by testing for correlation for each sample in sequence (e.g., each tile for each sample) or by testing for correlation for each tile in sequence (e.g., each sample for each tile). The loop may include producing a model based on the signal for the one or more epigenetic biomarkers for the set of expression-level correlated tiles and the expression level for the target gene for each sample of the subset of the samples. The loop may include predicting, using the model, an expression level for the target gene for one or more of the samples not included in the subset of the samples. The method may be for making a preliminary prediction of expression of a genomic target for an indication (e.g., at a particular circulating tumor1340408 v 1 Page 5 of 258Attorney Docket: 2014191-0049DNA (ctDNA) fraction) based on one or more epigenetic biomarkers. The method may be for producing a model that predicts expression level of a target gene for samples (e.g., liquid biopsy samples) derived from one or more subjects (e g., a patient population or perspective patient population). The method may be for obtaining a proof of concept that a target gene is relevant to an indication, for example for developing an ADC and / or ADC therapy, and / or that expression level of a target gene can be predicted for an indication.
[0010] In some aspects, the present disclosure is directed to a system that includes a processor and one or more non-transitory computer readable media having instructions stored thereon that, when executed by the processor, cause the processor to perform operations that include at least a portion of a method disclosed herein (e.g., that include a method disclosed herein).
[0011] In some aspects, the present disclosure is directed to one or more non-transitory computer readable media having instructions stored thereon that, when executed by a processor, cause the processor to perform operations that include at least a portion of a method disclosed herein (e.g, that include a method disclosed herein).
[0012] In some aspects, the present disclosure is directed to a method of predicting expression level of a liquid biopsy sample for a subject. The method may include providing an expression level prediction model that has been produced from digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene, wherein the digital samples have been generated using data derived from cell samples specific to the indication and healthy volunteers. The method may include providing input data including signal for the one or more epigenetic biomarkers derived from a liquid biopsy sample for a subject. The method may include predicting expression level of the target gene for the subject from the input data using the model,
[0013] In some aspects, the present disclosure is directed to a method of producing a model that predicts expression level of a target gene. The method may include receiving digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene. The method may include producing a set of constituent models based on the digital samples, wherein13404083 v 1 Page 6 of 258Attorney Docket: 2014191-0049the set includes one constituent model for each of the target gene and the one or more related genes. The method may include producing an ensemble model based on a combination (e.g., using a ridge regression) of the constituent models in the set.
[0014] In some aspects, the present disclosure is directed to a method of producing a model that predicts expression level of a target gene. The method may include receiving digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene. The method may include determining genomic regions corresponding to the target gene and the one or more related genes where the signal for at least one of the one or more epigenetic biomarkers correlates with the expression level for the samples (e.g., determine expression-level correlated tiles corresponding to the target gene and each of the one or more related genes). The method may include producing a model based on the signal for the one or more epigenetic biomarkers for the genomic regions (e.g., for the expression-level correlated tiles) and the expression level.
[0015] In some aspects, the present disclosure is directed to a method of predicting expression level of a target gene in a subject. The method may include receiving sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene. The method may include determining, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles corresponding to the genomic region. The method may include aggregating, for each of the one or more epigenetic biomarkers, the tile signal for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal for the epigenetic biomarker. The method may include predicting (e.g., using a model) an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signal for each of the one or more epigenetic biomarkers. The method may include comprising tiling the sample sequencing data into tiles that together span a genomic region corresponding to a target gene to obtain tiled sample sequencing data, wherein determining the tile signal is performed using the tiled sample sequencing data.
[0016] In some aspects, the present disclosure is directed to a method of predicting expression level of a target gene in a subject. The method may include receiving sample13404083 v 1 Page 7 of 258Attorney Docket: 2014191-0049sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene. The method may include predicting (e.g., using a model) an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the signal for each of the one or more epigenetic biomarkers. The method may include determining, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles. The method may include aggregating, for each of the one or more epigenetic biomarkers, the tile signal for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal. In some embodiments, the prediction of the expression level based on the signal for each of the one or more epigenetic biomarkers includes predicting the expression level based on the aggregated tile signal for each of the one or more epigenetic biomarkers. The method may include tiling, the sample sequencing data into tiles that together span a genomic region corresponding to a target gene to obtain tiled sample sequencing data, wherein determining the tile signal is performed using the tiled sample sequencing data.
[0017] In some aspects, the present disclosure is directed to a method of predicting expression level of a target gene in a subject. The method may include receiving sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for two or more genes (e.g., for one or more genomic regions corresponding to the two or more genes), wherein the two or more genes comprise a target gene and one or more additional (e.g., related) genes. The method may include determining, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for each of the two or more genes for a set of expression-level correlated tiles (e.g., wherein the method comprises tiling, the sample sequencing data into tiles that together span one or more genomic regions corresponding to each of the two or more genes to obtain tiled sample sequencing data and the aggregating is performed using the tiled sample sequencing data). The method may include aggregating the tile signal for the one or more epigenetic biomarkers signal for the set of expression-level correlated tiles to obtain aggregated tile signals for each of the two or more genes. The method may include predicting an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signals for each of the two or more genes.13404083 v 1 Page 8 of 258Attorney Docket: 2014191-0049
[0018] In some aspects, the present disclosure is directed to a method of predicting expression level of a target gene in a subject. The method may include receiving sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for two or more genes (e.g., for one or more genomic regions corresponding to the two or more genes), wherein the two or more genes comprise a target gene and one or more additional (e.g., related) genes. The method may include predicting an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the signal for reach of the one or more epigenetic biomarkers for the two or more genes.
[0019] Any two or more of the features described in this specification, including in this summary section, may be combined to form implementations of the disclosure, whether specifically expressly described as a separate combination in this specification or not.
[0020] At least part of the methods, systems, and techniques described in this specification may be controlled by executing, on one or more processing devices, instructions that are stored on one or more non-transitory machine-readable storage media. Examples of non-transitory machine- readable storage media include read-only memory, an optical disk drive, memory disk drive, and random access memory. At least part of the methods, systems, and techniques described in this specification may be controlled using a computing system including one or more processing devices and memory storing instructions that are executable by the one or more processing devices to perform various control operations.1340408 v 1 Page 9 of 258Attorney Docket: 2014191-0049BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present teachings described herein will be more fully understood from the following description of various illustrative embodiments, when read together with the accompanying drawings. It should be understood that the drawings described below are for illustration purposes only and is not intended to limit the scope of the present teachings in any way. The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and may be better understood by referring to the following description taken in conjunction with the accompanying drawings.
[0022] FIG. 1 illustrates a method of identifying a candidate target gene and / or making preliminary prediction of expression level for a target gene for an indication, according to illustrative embodiments of the present disclosure.
[0023] FIG. 2 is a block diagram of an example network environment for use in the methods and systems described herein, according to illustrative embodiments of the present disclosure.
[0024] FIG. 3 is a block diagram of an example computing device and an example mobile computing device, for use in illustrative embodiments of the present disclosure.
[0025] FIG. 4 illustrates performance of predictive gene expression models trained across a panel of breast cancer samples. Each point represents a gene. The number of genes meeting different AUC thresholds across varying ctDNA fractions (100%, 30%, 10%, 5%, 3%, 1%, rightmost to left-most line, in order) is shown. Predictive accuracy generally declines at lower ctDNA fractions, but many genes remain robustly predicted at clinically relevant levels.
[0026] FIG. 5 illustrates performance of predictive gene expression models for clinically relevant drug targets in breast cancer. (A) Shows the performance of methods described herein for predicting the expression of certain representative drug targets at 10% ctDNA fraction, including targets of antibody-drug conjugates (ADCs), hormone therapy, and other oncology¬ drugs. Each point represents a gene, with the Spearman correlation coefficient between actual expression (as measured by RNA-seq) and predicted expression (as calculated using cfDNA epigenomic features) plotted on the x-axis and model AUC plotted on the y-axis. (B) Shows validation of B7-H4 (VTCN1) expression predictions across 35 breast cancer cell lines at 10% ctDNA. Each point represents a different cell line. Plotted on the x-axis is actual RNA expression (as measured by RNA-seq) and plotted on the y-axis is predicted RNA expression13404083v 1 Page 10 of 258Attorney Docket: 2014191-0049(predicted using cfDNA epigenomic features). (C) Shows stability of predictions across varying ctDNA fractions for certain genes. Plotted on the x-axis is tumor (ctDNA) fraction. Plotted on the y-axis is the Spearman correlation coefficient between RNA expression (as measured by RNA-seq) and predicted expression (as measured using epigenomic features). The red horizontal line indicates the Spearman correlation coefficient at a tumor fraction of 0.10.
[0027] FIG. 6 illustrates plasma-based gene expression predictions validate in matched tumor RNA-seq. (A) Shows observed vs. shuffled correlations between predicted and measured RNA-seq expression. Boxplots compare the distribution of correlations between model-predicted gene expression and RNA-seq from matched tumor biopsies (Observed) versus control (Shuffled, N=100). (B) Scatter plots show expression predicted using plasma samples (y-axis) vs. RNA-seq expression values measured in matched tumor biopsies (x-axis) for 12 genes, N=12, ctDNA >3%. Genes include ADC targets (e.g., HER.2, NECT1N4, B7-H4 in purple dotted box) and breast cancer-relevant markers. Blue regression lines indicate model fit, and shaded regions represent confidence intervals. (C) Compares expression predictions collected from plasma samples to expression measurements collected from non-matched tissue samples. “BRCA,” “PRAD,” and “SCLC” are abbreviations for breast cancer, prostate cancer adenocarcinoma, and small cell lung cancer, respectively. In both the left heatmap and the right heatmap, each row provides expression values for an individual gene (652 total) and each column provides expression values for a single sample. Each gene for which data is shown in the left and right heatmaps were selected for having an associated locus expression model prediction that was statistically significantly associated with a particular indication (e.g., prediction of expression for genes in the top block were higher in breast cancer than in prostate cancer or small cell lung cancer) when used to predict gene expression in plasma samples from patients selected from breast cancer, prostate cancer, or SCLC. The left heatmap shows predicted gene expression in plasma samples using the locus expression models described herein. The right heatmap shows RNA-seq expression from TCGA tumor biopsies across breast (BRCA), prostate (PRAD), and small cell lung cancer (SCLC), highlighting genes with cancer-specific expression patterns. Despite there being no direct matching between plasma and TCGA samples beyond selecting samples from these 3 indications, plasma predictions recovered tumor-type-specific gene expression, mirroring RNA-seq patterns observed in tumor tissue. Columns are sorted by adj pval of subtype specific differential test on LEMs (genes are shown in the same order in the left1340408 v 1 Page 11 of 258Attorney Docket: 2014191-0049and right heatmaps). In the heatmaps, higher expression values are represented by increasingly warm colors (i.e., most red indicates highest expression value).
[0028] FIG. 7 illustrates refining locus expression models improves accuracy and lowers ctDNA detection limits. (A), (C), and (E) show AUC curves for the classification of HER2, ER, and PR status, respectively, using plasma samples. (B), (D), and (F) show AUC values at different ctDNA fractions for the classification of HER2, ER, and PR status, respectively.“Refined” refers to a HER2 classifier that was improved through feature engineering.“Multigene” refers to ER and PR classifiers that were improved through incorporation of loci from additional genes (i.e., in addition to ESRI and PGR). “ISP” stands for in silico plasma, and refers to simulated plasma samples generated by diluting in silico plasma sequencing data from cancer patients with healthy plasma data.
[0029] FIG. 8A depicts comparison of performance of ER status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation) based on genomic loci near ESRI (a single gene model). Comparison was made by plotting area under the curve (AUC) for each model based on specificity and sensitivity of models.
[0030] FIG. 8B depicts comparison of performance of ER status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation) based on genomic loci near multiple genes associated with expression of ESRI (a multigene model). Comparison was made by plotting area under the curve ( AUC) for each model based on specificity and sensitivity of models.
[0031] FIG. 9 depicts comparison of performance of ER status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation) based on genomic loci near ESRI (a single gene model) or genomic loci near multiple genes associated with expression of ESRI (a multigene model). Model performance was evaluated at ctDNA concentrations ranging from 10% to 0.5% in diluted in silico plasma samples Comparison was made by determining an area under the1340408 v 1 Page 12 of 258Attorney Docket: 2014191-0049curve (AUC) for each model based on specificity and sensitivity of models
[0032] FIG. 10 depicts comparison of performance of HER2 status classifiers using estimated gene expression from target expression predictive models using signals from two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation). Comparison was made by plotting area under the curve (AUC) for each model based on specificity and sensitivity of models.
[0033] FIG. 11 illustrates performance of predictive gene expression models for clinically relevant drug targets in breast cancer. Performance of methods described herein for predicting the expression of certain representative drug targets at 10% ctDNA fraction is shown. Drug targets include targets of antibody-drug conjugates (ADCs), hormone therapy, and other oncology drugs. Gene expression models were evaluated for performance when estimating gene expression based on two analytes (H3K4me3 and H3K27Ac histone modifications) or three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation). Each point represents an AUC for a three-analyte model plotted on the x-axis, and a two-analyte model plotted on the y-axis.
[0034] FIG. 12 illustrates comparison of exemplary predictive models provided herein for Nectin4 transcript expression in silico. FIG. 12A is a graph showing validation of Nectin4 expression predictions across 35 breast cancer cell lines at 10% ctDNA. Each point represents a different cell line. Plotted on the x-axis is actual RNA expression (as measured by RNA-seq) and plotted on the y-axis is predicted RNA expression (predicted using ctDNA epigenomic features) with a model using data from three analytes (H3K4me3 and H3K27Ac histone modifications, and DNA methylation). 12B is a graph showing validation of Nectin4 expression predictions across 35 breast cancer cell lines at 10% ctDNA. Each point represents a different cell line. Plotted on the x-axis is actual RNA expression (as measured by RNA-seq) and plotted on the y-axis is predicted RNA expression (predicted using cfDNA epigenomic features) with a model using data from two analytes (H3K4me3 and H3K27Ac histone modifications). FIG. 12C is a graph showing stability of predictions across varying ctDNA fractions for Nectin4 when using data form three analytes. Plotted on the x-axis is tumor (ctDNA) fraction. Plotted on the y-axis is the Spearman correlation coefficient between RNA expression (as measured by RNA-seq) and predicted expression (as measured using epigenomic features). The red horizontal line indicates the Spearman correlation coefficient at a tumor fraction of 0.10 FIG. 12D is a graph showing1340408 v 1 Page 13 of 258Attorney Docket: 2014191-0049stability of predictions across varying ctDNA fractions for Nectin4 when using data form two analytes. Plotted on the x-axis is tumor (ctDNA) fraction. Plotted on the y-axis is the Spearman correlation coefficient between RNA expression (as measured by RNA-seq) and predicted expression (as measured using epigenomic features). The red horizontal line indicates the Spearman correlation coefficient at a tumor fraction of 0.10.DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0035] Disclosed herein are systems and methods that use digital samples for an indication to determine correlations for the samples between signal for one or more epigenetic biomarkers for regions of or near a target gene and expression level, for example mRNA or protein expression level, for the target gene to produce a model that predicts expression level of a sample corresponding to the indication based on sample signal for the one or more epigenetic biomarkers and systems and methods of using such models. Samples may be tiled in a genomic region that spans a target gene to determine epigenetic biomarker and expression level correlations thereby producing a set of expression-level correlation tiles. A model may be produced based on signal for the epigenetic biomarker(s) for the expression-level correlation tiles and the expression levels for the different samples. In some embodiments, signal for an epigenetic biomarker is aggregated across a set of expression-level tiles using one or more point estimates and a model is produced based on the one or more point estimates for each of the one or more epigenetic biomarkers. Byusing tiling-based approaches, subregions in or near a target gene having differential signal may be distinguished from subregions that do not, to improve predictive power of a model. Other digital samples (e.g., held out from the set used for model production) may be used to determine whether a produced model can sufficiently predict expression level for an unseen sample. The resulting models and predictions from systems and methods disclosed herein may be indication-specific (e.g., to a specific cancer type or subtype), for example based on which digital samples are used.
[0036] A model may be considered performant if predictions from the model fall within one or more predefined criteria, such as an R2or area under curve (AUC) criterion. Such a successful model may serve as a proof of concept predictor. An initial model produced by a method disclosed herein may be produced based on samples at a relatively high ctDNA fraction (e.g., 10% or higher). Such a model may be tuned, either through iterative application of a method13404083v 1 Page 14 of 258Attorney Docket: 2014191-0049disclosed herein, or using a different method (eg., manual method) (e.g., feature engineering), to be performant at lower ctDNA fraction. Lower ctDNA fractions are more clinically relevant and therefore model performance at low ctDNA fraction is generally desirable. For example, in target selection for ADC therapies, it may be preferable to choose a target gene that will show strong epigenetic biomarker to expression level correlation such that its expression is predictable in samples with low ctDNA fraction (and therefore lower epigenetic biomarker signal) thereby allowing patients with early stage cancer (who have smaller fractions of ctDNA) to be identified and treated with the resulting therapy.
[0037] Because all samples used can be for a particular indication, such a successful model may indicate that there is significant relationship between expression level of a target gene and an indication. Therefore, such methods may be used to identify target genes for indications. Identification of a target gene may be used to develop an antibody-drug conjugate (ADC). Predictions from expression level prediction models may be used in the development of an ADC or therapy using the ADC.
[0038] A method may include receiving (e.g., and generating) digital samples for an indication each including (i) signal (e.g., sequencing counts) for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene. One or more epigenetic biomarkers may be a plurality of epigenetic biomarkers. An epigenetic biomarker may be a histone modification or DNA methylation or other epigenetic biomarker (e.g., chromatin accessibility or transcription factor binding). A set of one or more epigenetic biomarkers may include H3K27ac modification, H3K4me3 modification, DNA methylation, or a combination thereof. An expression level may be an mRNA expression level or a protein expression level, for example. A target gene may be a gene that encodes an antibody¬ drug conjugate (ADC) target antigen. In some embodiments, an indication is a cancer indication. A cancer indication may be a cancer type. A cancer indication may be a cancer subtype
[0039] A method may further include tiling digital samples into tiles that together span a genomic region corresponding to a target gene. A loop may be performed to determine where in the genomic region epigenetic biomarker and expression level correlation exists based on a tile-by-tile testing, thereby allowing high signal subregions of a genomic region corresponding to a target gene to be identified and used to produce and test a model. A loop may be performed for each of at least one subset of samples, for example in a cross validated loop having folds where1340408 v 1 Page 15 of 258Attorney Docket: 2014191-0049one or more (e.g., exactly) one sample is held out from each subset and the held out sample is rotated through the folds. A plurality of iterations of a loop may be performed each using a different subset of the samples. All of the samples in different subsets used in multiple iterations of a loop may correspond to (e.g., have) a particular ctDNA fraction. A subset of samples used in a loop may all correspond to (e.g., have) a particular ctDNA fraction. A loop may be performed for each of at least one subset of samples for each of a set of ctDNA fractions. A set of ctDNA fractions may correspond to an expected range of ctDNA fractions for an indication. Where multiple iterations of a loop are performed, all of the samples in the subsets of sample may correspond to (e.g., have) a particular ctDNA fraction, for example a cross validated loop with a number of folds of sample subsets may be performed all corresponding to (e.g., having a particular ctDNA fraction).
[0040] A loop may include determining a set of expression-level correlated tiles, producing a model based on those tiles, and predicting an expression level for a sample not in a subset of samples used to produce the model using the model. A loop may include testing each tile for correlation between the signal corresponding to the tile for each of one or more epigenetic biomarkers (e.g., in a sample) and the expression level (e.g., for the sample) across a subset of samples to determine a set of expression-level correlated tiles for the subset of the samples. Correlation may be tested for each sample in sequence (e.g., each tile for each sample) or for each tile in sequence (e.g., each sample for each tile), for example. A loop may include producing a model based on signal for one or more epigenetic biomarkers for a set of expression-level correlated tiles determined for a subset of samples and expression level for each sample of the subset of samples. A model may be or include a regression (e.g., a multiple regression) based model [e.g., is a regression (e.g., a multiple regression)]. A regression used in a model may be an ordinary least squares (OLS) regression (e.g., an OLS multiple regression). A loop may include predicting, using a produced model, an expression level for a target gene for one or more (e.g., exactly one) samples not included in a subset of samples used to produce the model (e.g., a held out sample).
[0041] Methods disclosed herein can be automated, including steps of generating digital samples and / or predicting an expression-level correlation tiles and producing models with them. In this way, candidate target genes for different indications can be rapidly tested and / or identified by applying a uniform method, for example without needing to engage in significant feature13404083 v 1 Page 16 of 258Attorney Docket: 2014191-0049engineering. An automated method may use predefined steps that are not modified based on indication and / or target gene. A model produced by an automated method may then be tuned in a non-automated (e.g., manual) way, for example to improve performance.
[0042] FIG. 1 illustrates a method 100 according to the present disclosure that may be performed automatically. The method can be used to identify a candidate target gene, for example for an ADC for an indication and / or to make a preliminary prediction of expression level (e.g., mRNA or protein expression level) for a target gene for an indication, for example. In step 102, digital samples for an indication that include signal for one or more epigenetic biomarkers and an expression level for a target gene are received. In some embodiments, method 100 includes generating the digital samples. In step 104, the samples are tiled into tiles. In step 106, a loop is performed for at least one subset of the samples received in step 102 and tiled in step 104, for example a cross validated loop where each fold uses a different subset of the samples and rotates which exactly one of the samples is held out between the folds. The loop may be performed with a subset of samples all corresponding to (e.g., having) a particular ctDNA fraction. The loop may be performed a number of iterations for different subsets of samples all corresponding to (e.g., having) a particular ctDNA fraction. In some embodiments, the loop is performed a number of iterations using subsets of samples all corresponding to the same initial (e.g., relatively high) ctDNA fraction (e.g, 10% or 5%) and then, if the predictions from those iterations are within one or more predefined criteria, performed for a second number of iterations using subsets of samples all corresponding to a second, lower ctDNA fraction (e.g., 5% or 3%).
[0043] The loop includes performing steps 106a- 106c. In step 106a, one or more expression-level correlated tiles are determined from the tiles for each sample of the subset of samples used in the loop. The one or more expression-level correlated tiles may be determined by determining for each of the tiles whether there is correlation in the tile between expression level for a sample and signal for one or more epigenetic biomarkers for the sample (e.g., each of one or more epigenetic biomarkers) across the subset of samples. In some cases, adjacent overlapping tiles may both be correlated and collapsed together into one expression-level correlated tile. In step 106b, a model is produced using the expression-level correlated tiles (e.g., based on signal for the one or more epigenetic biomarkers corresponding to the expression-level correlated tiles and the expression level for the subset of the samples). Producing the model may include producing one or more point estimates (e.g., mean(s), e.g., geometric mean(s)) for signal across the1340408 v 1 Page 17 of 258Attorney Docket: 2014191-0049expression-level correlated tiles Producing the model may include producing a regression-based model, such as an OLS multiple regression using point estimates determined from signal for one or more epigenetic biomarkers for the subset of the samples in the expression-level correlated tiles determined in step 106a and the expression level for each of the subset of the samples. In step 106c, expression level (e.g., mRNA or protein expression level) is predicted for a sample received in step 102 and not included in the subset used to produce the model (e.g., a held out sample for that cross validation fold).
[0044] If all loops have been completed, then the method proceeds to step 108. If not, then another loop of step 106 (including steps 106a-106c) proceeds. In step 108, once all appropriate loops are completed, it is determined whether the predictions fall within one or more predefined criteria by comparing the predicted expression level to the actual expression level for each sample used for prediction. The actual expression level may be assumed from a measured expression level corresponding to the cell line used to generate the sample. Predefined criteria that may be used include an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%) and / or an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9). Successful prediction within one or more predefined criteria may lead to tuning a model to perform at lower ctDNA fraction (e.g., clinically relevant fractions of 1-3%) and / or exhibit better performance (e.g., based on the one or more predefined criteria) and / or may lead to repeating the method 100 with samples corresponding to (e.g., having) a lower ctDNA fraction. Method 100 may be used to test whether a target gene is correlated with expression for an indication. The outcome of method 100 may therefore be a determination that a target gene is a candidate target gene for an ADC for an indication. Such information may be used in developing ADCs for new indications for which there is not an existing ADC on market or targeting a new gene for an indication not targeted by any existing ADC on market.
[0045] When samples are tiled, the resulting tiles may be used to determine specific tiles where there is correlation between expression level and signal for one or more epigenetic biomarkers. Expression-level correlated tiles may be determined by testing tiles for samples (e.g., a subset of samples). Signal for one or more epigenetic biomarkers for samples in different tiles spanning a genomic region corresponding to a target gene may be compared to expression level for the samples to determine whether there is correlation. Such processes may identify the particular subregions in a genomic region where there is statistically meaningful signal that may13404083 v 1 Page 18 of 258Attorney Docket: 2014191-0049be predictive of expression level. In general, not every portion of a genomic region corresponding to a target gene will produce signal (e g., sequencing counts) that exhibits a statistically meaningful relationship with expression level (e.g., for at least one of a set of one or more epigenetic biomarkers being considered) and those regions can be ignored (e.g., at least with respect to the at least one of the set of one or more epigenetic biomarkers) when producing a model to predict expression level. Furthermore, by considering genomic regions or subregions that are too large, technical noise or other detrimental effects can reduce predictive power. By considering genomic regions (e.g., only genomic regions) that don’t exhibit signal in healthy volunteer samples, the signal associated with a target gene can be highly cancer-cell-specific. Tiling samples and then predicting an expression-level correlated tiles therefrom can mitigate these effects. Tiling approaches as disclosed herein may helpfully reduce dimensionality for further model tuning (e.g., feature engineering). A set of expression-level correlated tiles may be determined as tiles that each exhibit correlation between each of one or more epigenetic biomarkers and expression level or may be determined as tiles that each exhibit correlation between at least one of one or more epigenetic biomarkers and expression level.
[0046] In some embodiments, determining a set of expression-level correlated tiles includes determining, for one or more tiles, that a correlation between signal for at least one epigenetic biomarker with expression level exceeds a threshold for a correlation measure. Different correlation measures may be used in various embodiments. In some embodiments, a correlation measure is a Spearman correlation. In some embodiments, a threshold for a Spearman correlation used as a correlation measure is at least 0.1, at least 0.2, at least 0.3, at least 0.4, or at least 0.5. In some embodiments, a correlation measure is a Pearson correlation. In some embodiments, a fixed threshold may be used (e.g., of at least 0.1, at least 0.2, at least 0.3, at least 0.4, or at least 0.5). In some embodiments, a dynamic threshold is used, for example where the dynamic threshold is defined as a point estimate (e.g., median or mean) of correlations determined for a particular biomarker and gene. For example, a correlation measure (e.g., Pearson correlation) may be calculated for epigenetic biomarker signal and expression across a set of tiles and expression-level correlated tiles may be determined based on (e.g., as) those tile(s) for which the correlation measure is above the median. In some embodiments, a combination threshold is used that is a combination of a fixed threshold and a dynamic threshold, For example, a combination threshold may be a combination of a fixed threshold of 0.1 and a dynamic threshold of a median1340408 v 1 Page 19 of 258Attorney Docket: 2014191-0049correlation measure between biomarker signal and expression level for a set of tiles (i.e, both thresholds must be exceeded to determine a tile as an expression-level correlated tile).
[0047] In some embodiments, one or more tiles are determined to be in a set of expression-level correlated tiles further based on signal for at least one epigenetic biomarker in healthy volunteer samples (e.g., used to generate digital samples being used to determine the set of expression-level correlated tiles) being below a threshold. Such a criterion can be used to eliminate tiles where there is an apparent correlation between signal for an epigenetic biomarker and expression level but the correlation is not caused by the indication but rather is a latent correlation for the genome (e g., human genome). In some embodiments, healthy volunteer sample signal being below a threshold is determined by identifying peaks in epigenetic biomarker signal for healthy volunteer samples and excluding any tile with a peak in at least a threshold amount of healthy volunteer samples, for example with a peak in at least 10% of health volunteer samples. Such peaks may be identified using tools known in the art, such as, for example MACS2. In some embodiments, healthy volunteer sample signal being below a threshold is determined using pebbling with a p-value cutoff. Such pebbling may estimate likelihood that signal (e.g., sequencing counts) in a tile are above background.
[0048] In some embodiments, determining a set of expression-level correlated tiles includes determining signal for at least one epigenetic biomarker for one or more of a set of tiles is uncorrelated with expression level across a subset of samples and excluding the one or more of the tiles from the set of expression-level correlated tiles [e.g., for the at least one of the one or more epigenetic biomarkers (e.g., excluding the one or more of the tiles from the set of expression-level correlated tiles entirely)]. Such exclusion may exclude that tile for only each epigenetic biomarker for which the tile is determined to be uncorrelated. In some embodiments, a set of expression-level correlated tiles includes only tiles for which there is correlation between epigenetic biomarker signal and expression level across each of a set of epigenetic biomarkers being considered or includes tiles for which there is correlation between epigenetic biomarker signal and expression level for at least one of a set of epigenetic biomarkers being considered; in general, the latter approach is preferred.
[0049] Tiles where there is correlation between expression level and epigenetic biomarker signal may be collapsed together where ones of the tiles overlap. Such collapsing may be used to produce a set of expression-level correlated tiles that are mutually exclusive. There may be13404083 v 1 Page 20 of 258Attorney Docket: 2014191-0049correlation between signal for one or more epigenetic biomarkers and expression level for overlapping tiles where the correlation results from a common region shared by the overlapping tiles, where a correlative region spans over two or more tiles, or where there are a plurality of small nearly-spaced correlative regions, for example. In some embodiments, determining a set of expression-level correlated tiles includes (i) determining overlapping tiles have a correlation between signal corresponding to the tile for each of one or more epigenetic biomarkers and the expression level and (ii) collapsing the overlapping tiles such that the set of expression-level correlated tiles are mutually non-overlapping. Collapsing overlapping tiles may result in a set of expression-level correlated tiles includes tiles having non-uniform size.
[0050] A model for predicting expression level may be produced based on raw signal (e.g., normalized counts) and expression level or may be produced based on data derived from raw signal, for example one or more point estimates and expression level. A point estimate may be a mean, such as, for example a geometric mean. Only one point estimate may be used per epigenetic biomarker or multiple point estimates may be used per epigenetic biomarker. For example, DNA methylation may be positively associated or negatively associated and different point estimates may be used for expression-level correlation tiles corresponding to positively associated DNA methylation and expression-level correlation tiles corresponding to negatively associated DNA methylation. In some embodiments, producing a model includes, for each of one or more epigenetic biomarkers, determining a point estimate for signal for the epigenetic biomarker across a set of expression-level correlated tiles and producing the model based on the point estimate. In some embodiments, the point estimate is a mean. In some embodiments, the mean is a geometric mean. In some embodiments, a model is produced based on a respective point estimate of signal for each of one or more epigenetic biomarkers for a set of expression-level correlation tiles. In some embodiments, producing a model includes determining a plurality of point estimates for signal for at least epigenetic biomarker across a set of expression-level correlated tiles and the model is produced based on the point estimate. In some embodiments, the plurality of point estimates includes a first point estimate for positively associated tiles for an epigenetic biomarker and a second point estimate for negatively associated tiles for the epigenetic biomarker.
[0051] One or more predefined criteria may be used to characterize a relationship (e.g., correlation) between predicted expression level from a model and actual expression level (e.g., taken from a cell line used to make samples). A predefined criterion may be, for example, an R213404083 v 1 Page 21 of 258Attorney Docket: 2014191-0049being above a threshold (e.g., at least 50%, at least 60%, at least 70%, or at least 80%) or an area under curve (AUC) above a threshold (e.g., at least 0.6, at least 0.7, at least 0.8, or at least 0,9); a combination of these criteria may be used. In some embodiments, a method includes determining whether a relationship (e.g., correlation) between an expression level for a target gene for a sample predicted from a model and an expression level for the sample across a plurality predicted samples are within one or more predefined criteria, for example an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%) and / or an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9). Such a determination may be made at a particular ctDNA fraction or for each of a plurality of ctDNA fractions (e.g., of no more than 10%, no more than 8%, no more than 6%, or no more than 5%).
[0052] If a model is able to predict expression level within one or more predefined criteria, it may be used as an indication that a target gene is suitable to use to form an ADC that targets an antigen encoded by the target gene. If a model is able to predict expression level within one or more predefined criteria, the model may serve has a proof of concept predictor for expression level for a target gene for an indication. If a model is able to predict expression level within one or more predefined criteria, the model may be used to make preliminary predictions of expression level for a target gene for an indication.
[0053] A model may be an initial model that performs at a relatively high ctDNA fraction (e.g., 10% or 5%). A model may be tuned upon determining that the model is performant within one or more predefined criteria. A model may be tuned upon determining that predictions for expression level for a target gene are within one or more predefined criteria. A method may be iterated at a lower ctDNA fraction upon determining that a model is performant within one or more predefined criteria. A method may be iterated at a lower ctDNA fraction upon determining that predictions for expression level for a target gene are within one or more predefined criteria. In some embodiments, a method includes determining that a relationship between expression level predicted with a model and actual expression level is within the one or more predefined criteria and, responsive to that determination, (e.g., manually) tuning (e.g., feature engineering) the model. Tuning a model may include selecting a subset of one or more expression-level correlated tiles previously determined and tuning the model using the subset of the one or more expression-level correlated tiles. In some embodiments, a model is for an initial ctDNA fraction and tuning the model includes tuning the model to produce predictions at a ctDNA fraction lower than the initial1340408 v 1 Page 22 of 258Attorney Docket: 2014191-0049ctDNA fraction. In some embodiments, a method includes determining that a relationship between expression level predicted with a model and actual expression level is within one or more predefined criteria and, responsive to that determination, performing a loop with each of at least one subset of samples that correspond to a lower ctDNA fraction.
[0054] Samples may be tiled over a genomic region corresponding to a target gene. A genomic region may include a region corresponding to a transcript for a target gene. A genomic region may correspond to an exon for a target gene. A genomic region may include a buffer around the target gene (e.g., around an exon) (e.g., around a transcript). For example, a genomic region may include initial and ending buffer regions. For example, a buffer of ±100 kb, ±200 kb, ±300 kb, ±400 kb, or ±500 kb may be used, (e.g., around an exon) (e.g., around a transcript).
[0055] Samples may be tiled into tiles to determine regions of a target gene where one or more epigenetic biomarkers are correlated with an expression level. In some embodiments, tiles are overlapping. In some embodiments, no more than two tiles are mutually overlapping. In some embodiments, adjacent tiles overlap by at least 10% (e.g., at least 20%, at least 30% of a length of the tiles) and no more than 70% (e g., no more than 60% or no more than 50% of the length of the tiles. In some embodiments, adjacent tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp). In some embodiments, tiles have a length in a range of from 100 to 1000 bp (e.g., from 250 to 750 bp).
[0056] Signal for an epigenetic biomarker included in a sample may be normalized and / or pebbled, for example to allow different samples to be appropriately compared to each other and / or to account for background signal. Signal may be normalized and / or pebbled before performing one or more loops to determine correlation between expression level and epigenetic biomarker signal for sample tiles, produce a model, and use the model to predict expression level. Normalization may include a quantile normalization. In some embodiments, a method includes normalizing signal for each of one or more epigenetic biomarkers in samples prior to performing a loop. In some embodiments, a method includes pebbling signal for each of one or more epigenetic biomarkers in samples prior to performing the loop such that the loop is performed using the pebbled signal. In some embodiments, pebbling a signal for an epigenetic biomarker includes determining a background signal for the epigenetic biomarker for a genomic region based on signal (e.g., a number of fragments) in a sample in a pebbling region and subtracting the background signal from the signal. Normalization may be ctDNA fraction dependent (e.g.,13404083vl Page 23 of 258Attorney Docket: 2014191-0049different normalizations applied to samples corresponding to (e.g., having) different ctDNA fractions). A pebbling region may be larger than a genomic region corresponding to a target gene. A pebbling region may be at least 1 mega-base-pairs (Mbp), at least 2 Mbp, at least 3 Mbp, at least 4 Mbp, or at least 5 Mbp large. A pebbling region may include or overlap with a genomic region corresponding to a target gene, preferably include.
[0057] In some embodiments, one or more additional related (e.g., correlated) genes are considered in addition to a target gene when producing a model. For example, an ensemble model may be produced from constituent models each corresponding to a different gene or a single model may be produced that simultaneously considers signal across a larger genomic region corresponding to multiple genes. Such approaches may benefit from using predictive epigenomic signal to genomic regions outside a target gene that are related to the target gene
[0058] Related genes may be selected (e g., identified) in a number of ways, including using publicly accessible databases such as The Cancer Genome Atlas (TCGA). Related genes may be selected based on, for example, genes being in a transcriptional complex with a target gene, being master regulators for an indication, being related to a particular pathway relevant for an indication, or a combination thereof. As an example, if ESR I is a target gene, then FOXA1 and GATA3 may be selected as related genes for being in a same transcriptional complex. As another example or addition to the prior example, if ESRI is a target gene, then EN1 and PIM1 may be selected as related genes because they are master regulators of ER- breast cancer. A number of related genes may be selected, for example at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15. Related genes may be selected by selecting a certain number of genes having highest rank according to a measure of correlation (e.g., as ranked by a data source (e.g., TCGA)), for example, though not only the highest ranking set of related genes need be used.
[0059] A set of expression-level correlated tiles for one or more epigenetic biomarkers may be determined for each of the related genes in addition to the target gene from digital samples that include signal for the one or more epigenetic biomarkers for the related genes in addition to the target gene in a manner disclosed herein with respect to target genes, A model may be produced based on signal for one or more epigenetic biomarkers for a set of expression-level correlated tiles for each related gene and target gene together (e g., using a ridge regression) or constituent models for each related gene and target gene may be produced separately from individual sets of for each of one or more epigenetic biomarkers and then an ensemble model may be produced therefrom1340408 v 1 Page 24 of 258Attorney Docket: 2014191-0049(e.g., based on an ordinary least squares regression or a ridge regression of the predictions from each constituent model). The latter approach may simplify the model and may be less prone to overfitting. The former approach may benefit where certain epigenetic biomarkers and / or regions might be more correlated for genes that are under tight regulatory control. For example, in certain such cases, enhancers are more correlated than promoters or DNA methylation. Because models can be produced for a large numbers of target genes quickly using methods disclosed herein, a model may have already been produced and / or a set of expression-level correlated tiles may have already been determined for a related gene to a target gene and therefore multigene models can be rapidly produced as well.
[0060] In some embodiments, a method of producing a multigene expression level prediction model may include producing a first constituent model for a target gene using a method disclosed herein; selecting one or more related genes related to the target gene; and producing a respective second constituent model for each of the one or more related genes using a method disclosed herein (e.g, wherein the related gene is the target gene in the method); and producing an ensemble model based on the first constituent model and the respective second constituent model(s). In some embodiments, the ensemble model uses a ridge regression of the first constituent model and the second constituent model.
[0061] In some embodiments, a method includes producing a model according to a method disclosed herein and selecting one or more related genes related to the target gene, wherein the digital samples include signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes such that set of expression-level correlated tiles includes at least one tile corresponding to each of the one or more related genes.
[0062] In some embodiments, a method includes producing a model according to a method disclosed herein and selecting one or more related genes related to the target gene, wherein the digital samples include signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes, and determining the set of expression-level correlated tiles includes testing each of the tiles corresponding to the one or more related genes for correlation between the signal in the sample corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level for the sample across the subset of the samples.1340408 v 1 Page 25 of 258Attorney Docket: 2014191-0049
[0063] In some embodiments, a method of producing a model that predicts expression level of a target gene includes receiving digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene; producing a set of constituent models based on the digital samples, wherein the set includes one constituent model for each of the target gene and the one or more related genes; and producing an ensemble model based on a combination (e.g., using a ridge regression) of the constituent models in the set.
[0064] In some embodiments, a method of producing a model that predicts expression level of a target gene includes receiving digital samples for an indication each including (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g, mRNA expression level) for the target gene; determining genomic regions corresponding to the target gene and the one or more related genes where the signal for at least one of the one or more epigenetic biomarkers correlates with the expression level for the samples (e.g., determine expression-level correlated tiles corresponding to the target gene and each of the one or more related genes); producing a model based on the signal for the one or more epigenetic biomarkers for the genomic regions (e.g., for the expression-level correlated tiles) and the expression level.
[0065] In some embodiments, signal for digital samples and for a subject is sequencing counts. In some embodiments, selecting one or more related genes includes selecting one or more genes that are in a transcriptional complex with a target gene. In some embodiments, selecting one or more related genes includes selecting one or more genes that are master regulators for an indication. In some embodiments, selecting one or more related genes includes selecting one or more genes that are related to a pathway relevant for an indication. In some embodiments, one or more related genes includes at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes. In some embodiments, selecting one or more related genes includes selecting a number of genes having highest rank according to a measure of correlation [e.g,, or as ranked by a data source (e.g., TCGA)].
[0066] In some embodiments, digital samples have been generated using data derived from cell samples specific to an indication and healthy volunteers. In some embodiments, cell samples include tissue samples. In some embodiments, cell samples have been derived from one or more13404083 v 1 Page 26 of 258Attorney Docket: 2014191-0049cell lines, one or more patient-derived xenografts or a biopsy therefrom, one or more organoids, or a combination thereof. In some embodiments, a liquid biopsy sample is a plasma sample. In some embodiments, one or more epigenetic biomarkers include one or more histone modifications and / or DNA methylation. In some embodiments, one or more epigenetic biomarkers includes H3K27ac modification, H3K4me3 modification, and DNA methylation. In some embodiments, one or more epigenetic biomarkers includes H3K27ac modification and H3K4me3 modification.Subjects and Samples
[0067] Methods disclosed herein may use digital samples that are generated from existing data, for example existing sequencing data, for example from healthy volunteers and indicationspecific (e.g., cancer) sources (e.g., cell lines). The use of digital samples may reduce the amount of data that may need to be obtained or eliminate the need for new data to be obtained in order to perform a method, allowing the method to be performed faster than other methods. For example, as compared to other methods to determine (e.g., estimate) expression level for target genes and / or identify suitable target genes (e.g., for ADC development), the amount of invasive biopsy (whether liquid or solid biopsy) and / or sequencing that needs to be performed may be reduced or eliminated. Existing data, such as sequencing data, from healthy volunteers and cell lines may be combined to form an in silico dilution used as a digital sample. Large datasets of digital samples can be generated quickly and then analyzed using methods disclosed herein. Digital samples may be digital simulations of data derived from biopsy samples, for example liquid biopsy samples, for example liquid plasma samples. A digital sample may be referred to herein simply as a “sample” or may be referred to herein interchangeably as an “ / >z silica sample” or a “simulated sample.”
[0068] A plurality of digital samples used in a method disclosed herein may include in silico diluted (e.g., titrated) samples (e.g., in silico plasma (ISP) samples). Sequencing data for an in silico diluted digital sample may be obtained from a random sampling of fragments from a healthy (e.g., having a zero or below limit of detection ctDNA fraction) sample and a cancer sample (e.g., a sample of a cancer cell line) in a predetermined ratio of amounts. In some embodiments, a digital sample is a simulated dilution. In some embodiments, a digital sample corresponds to (e.g., has) a ctDNA fraction that is (e.g., has been determined by) a linear combination of ctDNA fractions of reference samples used to generate the sample. Expression level for a digital sample may be taken as the expression level for the cancer sample used to1340408 v 1 Page 27 of 258Attorney Docket: 2014191-0049generate the digital sample (e g., for the cell line sample used to generate the digital sample) Indi cation- specific sources used to derive indication-specific samples (e.g,, cancer samples) used to generate digital samples may be cell samples (e.g., tissue samples) such as, for example, cell lines, patient-derived xenografts or a biopsy therefrom, or organoids.
[0069] Digital samples can be generated from data from healthy volunteers and from indication-specific (e.g., cancer) cell lines. In some embodiments, a set of digital samples are or have been generated using at least 25 different cell lines. In some embodiments, a set of digital samples have been generated using at least 25 different healthy volunteers. In some embodiments, a subset of digital samples used in a loop are or have been generated using at least 25 different healthy volunteers. In some embodiments, a subset of digital samples used in a loop have been generated using at least 25 different cell lines. In some embodiments, digital samples are in silico diluted samples. In some embodiments, digital samples are in silico diluted plasma samples.
[0070] A method may include generating digital samples. Digital samples may be generated from healthy volunteer samples and from indication-specific (e.g., cancer) samples Indication-specific samples may be derived from cell lines. Generating a digital sample may include (i) randomly sampling data (e.g., sequencing counts for a region that includes a target gene) for at least one healthy sample and at least one cell line sample (e.g., from one or more indication-specific, e.g., cancer, cell lines) in a mixing ratio corresponding to a desired ctDNA fraction for the digital sample and (ii) predicting an expression level for the sample. Predicting an expression level for a digital sample may include using a previously measured expression level for at least one cel! line sample used to make the digital sample as the expression level for the digital sample. In some embodiments, signal for each of one or more epigenetic biomarkers for a sample is in silico mixed sequencing data (e.g., sequencing counts).
[0071] Training samples at a range of ctDNA fractions may be used in a method disclosed herein. In some embodiments, cell lines and solid tumor data can be considered high fraction (e.g., 100% ctDNA fraction). Samples from healthy subject (e.g., healthy plasma) can be considered low (e.g,, 0%) ctDNA fraction. ctDNA fraction may be determined from a method such as a VAF or CNV method (e.g. ichorCNA).
[0072] For training purposes, higher fraction samples can be mixed with healthy samples in silico to make digital samples (e.g., in silico plasma (ISP) samples) having a predetermined ctDNA fraction that is expected based on a mixing ratio used. Sequencing data for digital samplesI 340408 v 1 Page 28 of 258Attorney Docket: 2014191-0049(e.g., ISP samples) may be created by random sampling of sequencing data from the constituent samples used to make the digital sample. For example, a digital (e.g., ISP) sample may be made by randomly sampling a number of fragments (e.g., 5 million fragments) from sequencing data for a cell line and combining it with a number of fragments (e.g., 5 million fragments) from sequencing data from a healthy volunteer sample to form a new digital sample where signal for one or more epigenetic biomarkers for the digital sample is a combination of the signal for the cell line fragments and the healthy volunteer fragments and a ctDNA fraction for the sample is a linear combination of the ctDNA fractions for the cell line sample and the healthy volunteer sample based on the relative proportion of fragments used and an expression level that corresponds to the expression level for the cell line sample. Such samples are further described elsewhere herein.
[0073] Different numbers of empirical and / or in silico samples may be used in a method disclosed herein. Healthy and known cancer samples may be used. In some embodiments, a number of empirical samples used is in a range of from 25 to 10,000 samples (a higher or lower number outside this range could be used in some embodiments). In some embodiments, a number of healthy samples used is in a range of from 10 to 5,000. In some embodiments, a number of known cancer samples used is in a range of from 10 to 5,000. In some embodiments, a number of digital samples used is in a range of from 10 to 10,000 (optionally wherein no non-digital samples are used).
[0074] A sample used in methods and systems provided herein can be derived from any biological sample including any processed sample that includes circulating tumor DNA (ctDNA) derived from a biological sample. In various embodiments, a sample analyzed using methods and systems provided herein can be derived from a sample obtained from a mammalian subject. In various embodiments, a sample analyzed using methods and systems provided herein can be derived from a sample obtained from a human subject.
[0075] In various instances, a human subject is a subject diagnosed or seeking diagnosis as having, diagnosed as, or seeking diagnosis as at risk of having, and / or diagnosed as or seeking diagnosis as at immediate risk of having cancer, e.g., breast cancer, small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC), etc. In various instances, a human subject is a subject identified as needing ADC therapy In certain instances, a human subject is a subject identified as needing ADC therapy screening by a medical practitioner.1340408 v 1 Page 29 of 258Attorney Docket: 2014191-0049
[0076] The subject may not have undergone previous treatments for cancer, such as the treatments recited in this disclosure. In other embodiments, the subject has undergone previous treatments for cancer, such as the treatments recited in this disclosure.
[0077] In various embodiments a subject has one or more biomarkers and / or risk factors for cancer, e.g., breast cancer, small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC), etc. In certain embodiments, a human subject is identified as in need of ADC therapy screening based on an initial cancer diagnosis, e.g, a breast cancer, small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC), etc. diagnosis. In various instances, a human subject is a subject not yet diagnosed as having, not at risk of having, not at immediate risk of having, not diagnosed as having, and / or not seeking diagnosis for a cancer. Genetic factors may also contribute to cancer risk, as evidenced by individuals with a family history of cancer.
[0078] In various embodiments, a sample from a subject, e.g, a human can be obtained from a liquid biopsy. In certain embodiments, a sample and / or reference is obtained from serum, plasma, or urine. In certain embodiments, the sample is serum. In certain embodiments, a sample includes circulating tumor DNA (ctDNA). In certain embodiments, a sample is derived from about 1 mL of blood obtained from the subject In certain embodiments, a sample is derived from about 0.5-2 mL ofblood obtained from the subject, e.g., about 0.5 to 1.75 mL, about 0.5 to 1.5 mL, about 0,75 to 1.25 mL or about 0.9 to 1.1 mL ofblood.
[0079] In various embodiments, a sample is a sample of cell-free DNA (cfDNA). cfDNA is typically found in human biofluids (e.g, plasma, serum, or urine) in short, double-stranded fragments. The concentration of cfDNA is typically low, but can significantly increase under particular conditions, including without limitation pregnancy, autoimmune disorders, myocardial infarction, and cancer. Circulating tumor DNA (ctDNA) is the component of cell-free DNA specifically derived from cancer cells. ctDNA can be present in human biofluids bound to leukocytes and erythrocytes or not bound to leukocytes and erythrocytes. Various tests for detection of tumor-derived ctDNA are based on detection of genetic or epigenetic modifications that are characteristic of cancer (e.g, of a relevant cancer). Genetic or epigenetic factors characteristic of cancer can include, without limitation, oncogenic or cancer-associated mutations in tumor-suppressor genes, activated oncogenes, chromosomal disorders, histone modifications (e.g., histone methylation and / or histone acetylation), chromatin accessibility, binding of one or more transcription factors and / or DNA methylation.1340408 v 1 Page 30 of 258Attorney Docket: 2014191-0049
[0080] In various embodiments, ctDNA includes less than 30%, less than 20%, or less than 10% of the ctDNA in the liquid biopsy sample obtained from the subject, e.g., less than 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or less than 1% of the cfDNA in the sample. In some embodiments, the percentage of ctDNA in the liquid biopsy sample is assessed using ichorCNA which estimates the percentage of ctDNA in a sample probabilistically (see Adalsteinsson et al., Nat Commun (2017) 8(1): 1324 the entire contents of which are incorporated herein by reference).
[0081] cfDNA and ctDNA can provide a real-time or nearly real time metric of status of a source tissue. cfDNA and ctDNA demonstrate a half-life in blood of about 2 hours, such that a sample taken at a given time provides a relatively timely reflection of the status of a source tissue.
[0082] Various methods of isolating nucleic acids from a sample (e.g., of isolating cfDNA from blood or plasma) are known in the art. Nucleic acids can be isolated using, without limitation, standard DNA purification techniques, by direct gene capture (e.g., by clarification of a sample to remove assay-inhibiting agents and capturing a target nucleic acid, if present, from the clarified sample with a capture agent to produce a capture complex and isolating the capture complex to recover the target nucleic acid).
[0083] Reagents and protocols for obtaining and analyzing cfDNA and ctDNA, such as circulating in blood or other tissue, are commercially available as described in the Examples and well-known in the art (see, for example, Anker et al.. Cancer and Metastasis Rev (1999) 18:65-73; Wua et al., Clin Chim Acta (2002) 321:77-87; Fiegl et al., Cancer Res (2005) 15: 1141-1145; Pathak et al., Clin Chem (2006) 52:1833-1842; Schwarzenbach et al., Clin Cancer Res (2009) 15:1032-1038; Schwarzenbach et al., Nat Rev Cancer (2011) 11:426-437) the contents of each of which is separately incorporated herein by reference in their entirety).
[0084] In various embodiments, samples can be collected from individuals repeatedly over a period of time (e.g., once daily, weekly, monthly, annually, biannually, etc.). In various embodiments, such samples can be used to verify results from earlier detections and / or to identify an alteration in biological pattern because of, for example, disease progression, resistance to therapy, treatment, remission, and the like. For example, subject samples can be taken and monitored every month, every two months, or combinations of one, two, or three-month intervals according to the present disclosure. In various embodiments, samples can be collected for monitoring over time beginning at or at certain clinically determined stages, such as at resistance to a therapy, before radiographic progression, after radiographic progression, and / or at tissue1340408 v 1 Page 31 of 258Attorney Docket: 2014191-0049biopsy In addition, results from samples obtained at different points in time can be conveniently compared with each other, as well as with those of normal controls during the monitoring period, thereby providing the subject’s own values, as an internal, or personal, control for long-term monitoring.
[0085] Samples include materials prepared by processes including, without limitation, steps such as concentration, dilution, adjustment of pH, removal of high abundance polypeptides (e.g., albumin, gamma globulin, and transferrin, etc.), addition of preservatives, addition of calibrants, addition of protease inhibitors, addition of denaturants, desalting, concentration and / or extraction of sample nucleic acids, and / or amplification of sample nucleic acids (e.g., by PCR or other nucleic acid amplification techniques). Samples also include materials prepared by techniques that isolate, e.g., nucleosomes or transcription factors and / or nucleic acids associated with nucleosomes or transcription factors.
[0086] Removal from a sample of proteins that are not desirable for a relevant purpose or context e.g., high abundance, uninformative, or undetectable proteins) can be achieved using high affinity reagents, high molecular weight filters, ultracentrifugation and / or electrodialysis. High affinity reagents include antibodies or other reagents (e.g., aptamers) that selectively bind to high abundance proteins. Sample preparation can also include ion exchange chromatography, metal ion affinity chromatography, gel filtration, hydrophobic chromatography, chromatofocusing, adsorption chromatography, isoelectric focusing and related techniques. Molecular weight filters include membranes that separate molecules based on size and molecular weight. Such filters may further employ reverse osmosis, nanofiltration, ultrafiltration and microfiltration. Ultracentrifugation is the centrifugation of a sample at about 15,000-60,000 rpm while monitoring with an optical system the sedimentation (or lack thereof) of particles. Electrodialysis is a procedure which uses an electromembrane or semipermeable membrane in a process in which ions are transported through semi -permeable membranes from one solution to another under the influence of a potential gradient. Since the membranes used in electrodialysis may have the ability to selectively transport ions having positive or negative charge, reject ions of the opposite charge, or to allow- species to migrate through a semipermeable membrane based on size and charge, it renders electrodialysis useful for concentration, removal, or separation of electrolytes.
[0087] Separation and purification in the present disclosure may include any procedure known in the art, such as capillary'- electrophoresis (e.g., in capillary' or on-chip) or chromatography13404083 v 1 Page 32 of 258Attorney Docket: 2014191-0049(e.g., in capillary, column or on a chip). Electrophoresis is a method that can be used to separate ionic molecules under the influence of an electric field. Electrophoresis can be conducted in a gel, capillary, or in a microchannel on a chip. Examples of gels used for electrophoresis include starch, acrylamide, polyethylene oxides, agarose, or combinations thereof. A gel can be modified by its cross-linking, addition of detergents, or denaturants, immobilization of enzymes or antibodies (affinity electrophoresis) or substrates (zymography) and incorporation of a pH gradient. Examples of capillaries used for electrophoresis include capillaries that interface with an electrospray.
[0088] Capillary electrophoresis (CE) is preferred for separating complex hydrophilic molecules and highly charged solutes. CE technology can also be implemented on microfluidic chips. Depending on the types of capillary and buffers used, CE can be further segmented into separation techniques such as capillary' zone electrophoresis (CZE), capillary isoelectric focusing (CIEF), capillary' isotachophoresis (CITP) and capillary electrochromatography (CEC). An embodiment to couple CE techniques to electrospray ionization involves the use of volatile solutions, for example, aqueous mixtures containing a volatile acid and / or base and an organic such as an alcohol or acetonitrile.
[0089] Capillary' isotachophoresis (CITP) is a technique in which the analytes move through the capillary at a constant speed but are nevertheless separated by their respective mobilities. Capillary' zone electrophoresis (CZE), also known as free-solution CE (FSCE), is based on differences in the electrophoretic mobility of the analytes, determined by the charge on the analytes, and the frictional resistance the analytes encounter during migration, which is often directly proportional to the size of the analytes. Capillary isoelectric focusing (CIEF) allows weakly-ionizable amphoteric molecules, to be separated by electrophoresis in a pH gradient. CEC is a hybrid technique between traditional high performance liquid chromatography (HPLC) and CE.[00.90] Separation and purification techniques used in the present disclosure can include any chromatography procedures known in the art. Chromatography can be based on the differential adsorption and elution of certain analytes or partitioning of analytes between mobile and stationary' phases. Different examples of chromatography include, but not limited to, liquid chromatography (LC), gas chromatography (GC), high performance liquid chromatography (HPLC), etc.13404083 v 1 Page 33 of 258Attorney Docket: 2014191-0049
[0091] In some embodiments, whole blood is collected from a subject, and a plasma layer is separated by centrifugation. cfDNA may be then extracted from the plasma using methods known in the art.Histone Modifications, Chromatin Accessibility and Transcription Factor Binding
[0092] Histone methylation is understood to increase or decrease expression of associated coding sequences, depending on which histone residue is methylated. Histone methylation is an essential modification that can cause monomethylation (mel), dimethylation (me2), and trimethylation (me3) of several amino acids, thus directly affecting heterochromatin formation, gene imprinting, X chromosome inactivation, and gene transcriptional regulation. Histone methyltransferases promote monomethylation, dimethylation, or tri methylation of histones while histone demethylases promote demethylation of histones In general, lysine (Lys or K), arginine (Arg or R), and rarely histidine (His or H) are the most common histone methyl acceptors. Histone methylation only occurs at specific lysine and arginine sites of histone H3 and H4. In histone H3, lysine 4, 9, 26, 27, 36, 56, and 79 and arginine 2, 8, and 17 can be methylated. By comparison, histone H4 has fewer methylation sites, in which only lysine 5, 12, and 20 and arginine 3 can be methylated. Histone methylation is often associated with transcriptional activation or inhibition of downstream genes. The methylation of histone H3K4, R8, R17, K26, K36, K79, H4R3, and K12 can activate gene transcription. However, the methylation of histone H3K9, K27, K56, H4K5, and K20 can inhibit gene transcription. Thus, for example, H3K4 methylation generally activates gene expression, while H3K27 methylation generally represses gene expression.
[0093] Histone acetylation occurs predominantly at lysine residues and is generally- understood to increase expression of associated coding sequences. Without wishing to be bound by any theory', acetylation of lysine residues is thought to neutralize lysine’s positive charge and thereby cause histones to drift away from DNA, which has a negative charge. The released structure facilitates access to transcriptional machinery such as transcription factors and RNA polymerase II. Histone acetylation and deacetylation are generally catalyzed by histone acetyltransferases (HATs) and HDACs, respectively. Acetyl-CoA can be a source and co-factor of acetylation. In regulatory' regions, HATs can acetylate histones and recruit HAT-containing complexes to activate the transcriptional process. For instance, H3K9ac and H3K27ac levels can be associated with promoter and enhancer activities. Furthermore, H3K27ac enhances not only13404083 v 1 Page 34 of 258Attorney Docket: 2014191-0049the kinetics of transcriptional activation, but also accelerates the transition of RNA polymerase II from the initiation state to the elongation state.
[0094] Differential modification of a genomic locus (e.g., differential histone methylation and / or differential histone acetylation) can refer to, or be determined by or detected as, a comparative difference or change in modification status of one or more genomic loci between a first sample, condition, disease, or state and a second or reference sample, condition, disease, or state. Those of skill in the art will appreciate that a reference is typically produced by measurement using a methodology identical, similar, or comparable to that by which a compared non-reference measurement was taken.
[0095] Chromatin accessibility can refer to the degree to which nuclear macromolecules are able to physically contact DNA and is determined in part by the occupancy and modification status of nucleosomes. Modified histones can regulate chromatin accessibility through a variety of mechanisms, such as altering transcription factor (TF) binding through steric hindrance and modulating nucleosome affinity for active chromatin remodelers. The topological organization of nucleosomes across the genome is non-uniform: while histones can be densely arranged within facultative and constitutive heterochromatin, histones can be depleted at regulatory loci, including within enhancers, insulators and transcribed gene bodies. Active regulatory elements of the genome are generally accessible.
[0096] Differential accessibility of a genomic locus can refer to, or be determined by or detected as, a comparative difference or change in modification status of one or more genomic loci between a first sample, condition, disease, or state and a second or reference sample, condition, disease, or state. Those of skill in the art will appreciate that a reference is typically produced by measurement using a methodology identical, similar, or comparable to that by which a compared non-reference measurement was taken.
[0097] A reference can be a value or set of values that are predetermined or derived from a sample or set of samples. A reference can be a sample or set of samples. A reference value can be a predetermined threshold value, a value that varies in accordance with circumstances (e.g., according to patient subpopulation, age, weight, or other variables), or a ratio. Reference ratios can be ratios relating to the modification and / or accessibility of multiple loci within individual samples and / or references, or across or between samples and / or references. In various embodiments, a reference can have or represent a normal, non-diseased state. In some1340408 v 1 Page 35 of 258Attorney Docket: 2014191-0049embodiments, such as for staging of disease or for evaluating the efficacy of treatment, a reference can have or represent a diseased state, e.g., a cancer, stage of cancer, or subtype of cancer. In some embodiments, a reference can represent a particular level of target gene (e.g., ADC target) expression based on IHC testing, e.g., ADC target-positive or ADC target-negative cancer.
[0098] In certain instances, a reference is a non-contemporaneous sample from the same source, e.g., a prior sample from the same source, e.g, from the same subject. In certain instances, a reference for the modification status of one or more genomic loci (e.g., one or more differentially modified genomic loci) can be the modification status of the one or more genomic loci (e.g., one or more differenti lly modified genomic loci) in a sample (e.g., a sample from a subject), or a plurality of samples, known to represent a particular state (e.g., ADC target-positive or ADC target-negative cancer). In certain instances, a reference for the accessibility status of one or more genomic loci (e.g., one or more differentially accessible genomic loci) can be the accessibility status of the one or more genomic loci (e.g., one or more differentially accessible genomic loci) in a sample (e.g., a sample from a subject), or a plurality of samples, known to represent a particular state (e.g., ADC target-positive or ADC target-negative cancer).
[0099] In some illustrative but non-limiting embodiments of the present disclosure differential modification or differential accessibility can refer to a differential (e.g., between a sample and a reference) with an absolute log2(f old-change) that is greater than or equal to 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0 ormore, or any range inbetween, inclusive, e.g., as measured according to an assay provided herein.
[0100] Enhancers are genomic loci that can be differentially modified or differentially accessible in and / or between conditions, diseases, and other states. Enhancers are cis-acting DNA regulatory regions that are thought to bind trans-acting proteins that contribute to expression patterns of associated genes. Chromatin ImmunoPrecipitation sequencing (ChlP-seq) of histone modifications (e.g., acetylation) have identified millions of enhancers in mammalian genomes. The number of active enhancers in any given cell type is estimated to be in the tens of thousands. Certain transcription factors (TFs), sometimes referred to as ‘‘master” transcription factors, associate with active enhancers with important impacts on gene expression and cell function. Certain such transcription factors preferentially associate with enhancers that regulate genes required for establishing cell identity and function, including enhancer domains known as “super¬ enhancers”. Moreover, master TFs can participate in inter-connected auto-regulatory circuitries1340408 v 1 Page 36 of 258Attorney Docket: 2014191-0049or “cliques” that are self-reinforcing, show marked cell selectivity, and function to maintain cell state and / or cell survival.Techniques for Detecting and Quantifying Histone Modifications and Transcription Factor Binding
[0101] Various techniques of molecular biology are well known in the art and / or disclosed in the present application for detecting and quantifying histone modifications and / or transcription factor binding. In some embodiments, the methods, kits and systems of present disclosure involve the detection and quantification of histone modifications and / or transcription factor binding in samples, e.g., in liquid biopsy samples including cfDNA such as plasma samples including cfDNA. Chromatin ImmunoPrecipitation (ChIP) is one technique of molecular biology useful in detecting and quantifying histone modifications and transcription factor binding in samples. CUT& RUN or CUT& Tag are other more recent techniques that can also be used to detect and quantify histone modifications and transcription factor binding sites. ChlP-chip, ChlP-exo, ChIP Re-ChIP, and ChlPmentation are other alternative techniques that could be used.
[0102] ChIP can involve various steps including one or more of fixation, sonication, immunoprecipitation, and analysis of the immunoprecipitated DNA. ChIP has become a very widely used tissue-based technique for determining the in vivo location of binding sites of various transcription factors and histones. Because the proteins are captured at the sites of their binding with DNA, ChIP helps to detect DNA-protein interactions that take place in living cells. More importantly, ChIP can be coupled to many commonly used molecular biology techniques such as PCR and real-time PCR, PCR with single-stranded conformational polymorphism. Southern blot analysis, Western blot analysis, cloning, and microarray. The resulting versatility has increased the potential of this technique.
[0103] ChIP of tissue samples usually involves cross-linking of the chromatin-bound proteins by formaldehyde, followed by sonication or nuclease treatment to obtain small DNA fragments. Immunoprecipitation can be then carried out using specific antibodies to the DNA-binding protein of interest. The DNA can be then released from the proteins and analyzed using various methods. ChIP has also been used to study RNA-protein interactions. X-ChIP methods utilize fixed chromatin fragmented by sonication, while the N-ChIP methods utilize native chromatin, which can be unfixed and nuclease digested.13404083v 1 Page 37 of 258Attorney Docket: 2014191-0049
[0104] The first step of the technique can be the cross-linking of DNA and proteins. Formaldehyde i s one of the most used cross-linking agents. One advantage of using formaldehyde can be the ease of reversibility of the cross-links and its ability to form bonds that span approximately 2 angstroms. This means that formaldehyde can bind molecules in close association with each other. Generally, formaldehyde can be added to the medium in the cell culture flask or plate. It enters the cells through the cell membrane and cross-links the proteins to the chromatin. Formaldehyde fixation of tumor tissues has also been done. Other cross-linking agents that have been used include chemicals such as methylene blue and acridine orange, cisplatin, dimethylarsinic acid, potassium chromate, and ultraviolet (UV) light and lasers.
[0105] Harvested chromatin can be sonicated in one or more sonication cycles. DNA can be typically broken into to 100–500 bp fragments to pinpoint the location of the DNA sequence of interest. An alternative to sonication can be nuclease digestion of the chromatin, e.g., in N-ChIP methods. Purification of chromatin can be achieved using a cesium chloride (CsCl) gradient centrifugation.
[0106] Chromatin can be immunoprecipitated using one or more antibodies that bind a target epitope. For example, an antibody used in ChIP can selectively bind a particular transcription factor or one or more particular histone modifications, such as one or more particular histone acetylation modifications or histone methylation modifications. In some embodiments, an antibody used to bind a target epitope can be a “pan” antibody (e.g., a pan-acetylation antibody, a pan-methylation antibody, an antibody that binds a group of histone modifications associated with increased transcription activation, and / or an antibody that binds a group of histone modifications associated with increased transcription repression). The antibody against the protein of interest is allowed to bind to the protein-DNA complex, and the complex can be then precipitated. Immunosorbants commonly used to separate the antigen-antibody complex from the lysate include salmon sperm DNA-protein A-Sepharose®, protein G, magnetic beads, and other engineered immunoprecipitation systems known to those of skill in the art.
[0107] Immunoprecipitated DNA can be eluted. Once the DNA of interest is isolated, many detection and quantification methods can be used to study the isolated gene fragments. Commonly utilized methods include PCR, real-time PCR, slot blot hybridization, microarray techniques, and deep or next-generation sequencing. ChlP-seq combines chromatin immunoprecipitation (ChIP) with massively parallel DNA sequencing to identify the binding sites1340408 v 1 Page 38 of 258Attorney Docket: 2014191-0049of DNA-associated proteins. ChlP-seq can be used to map DNA-binding proteins, e.g., transcription factor binding sites and histone modifications in a genome-wide manner.
[0108] Cell-free Chromatin ImmunoPrecipitation sequencing (cfChlP-seq) involves applying ChlP-seq to samples that include cell-free DNA, e.g., liquid biopsy samples including cfDNA such as plasma samples including cfDNA (e.g., see Sadeh et al., Nat Biotechnol (2021) 39: 586-598 and Jang et al., Life Sci Alliance (2023) 6(12):e202302003 the entire contents of each of which are incorporated herein by reference). In some embodiments, cfChlP-seq uses antibodies or antibody fragments that bind specific histone modifications (e.g., H3K4me3 and / or H3K27ac) and / or transcription factors that are coupled (covalently or non-covalently) to beads, e.g., magnetic beads such as Dynabeads® magnetic beads and incubated with a volume, e.g., about 1 mL of thawed plasma obtained from a subject. Without limitation, exemplary antibodies that bind H3K4me3 include PA5-27029 (available from Thermo Fisher Scientific in Waltham, MA) and C15410003 (available from Diagenode in Denville, NJ) and exemplary antibodies that bind H3K27ac include ab21623 or ab4729 (both available from Abeam in Cambridge, UK) and C15210016 (available from Diagenode in Denville, NJ).
[0109] In some embodiments, the antibodies or antibody fragments can be covalently coupled to beads, e.g., epoxy beads. In some embodiments, the antibodies or antibody fragments can be non-covalently coupled to beads, e.g.. Protein A or Protein G beads such as Dynabeads® Protein A or Dynabeads® Protein G beads. After washing, a cfDNA library is then typically prepared from the captured cfDNA. Library' preparation can be done on-bead or after releasing the captured cfDNA by digestion of bound histones, e.g., using proteinase K. The cfDNA library is then sequenced to generate reads of captured cfDNA sequences, e.g., by next-generation sequencing (NGS) as is known in the art. The reads are then analyzed, e.g., aligned and counted using standard bioinformatic techniques as is known in the art. A cfChlP-seq bioinformatic pipeline can include, e.g., alignment of sequence reads to a reference genome with BWA or Bowtie2. Aligned reads can be used to call and quantify peaks as compared to a reference.
[0110] CUT& Tag involves antibody-based binding of a target protein, e.g., transcription factor or histone modification of interest, where antibody incubation is directly followed by the shearing of the chromatin and library preparation (see Kaya-Okur et al., Nat Comm (2019) 10:1930). CUT& Tag assays take advantage of a Tn5 transposase that is fused with Protein A to direct the enzyme to the antibody bound to its target on chromatin. Tn5 transposase is pre-loaded1340408 v 1 Page 39 of 258Attorney Docket: 2014191-0049with sequencing adapters (generating the assembled pA-Tn5 adapter transposome) to carry out antibody-targeted tagmentation. In a typical CUT& Tag assay samples are incubated with an antibody immobilized on Concanavalin A-coated magnetic beads to facilitate subsequent washing steps. Cells can be incubated with a primary antibody specific for the target protein of interest followed by incubation with a secondary antibody. Samples can then be incubated with assembled transposomes, which consist of Protein A fused to the Tn5 transposase enzyme that is conjugated to NGS adapters. After incubation, unbound transposome can be washed away using stringent conditions. Tn5 is a Mg2+-dependent enzyme so Mg2−can be added to activate the reaction, which results in the chromatin being cut close to the protein binding site and simultaneous addition of the NGS adapter DNA sequences. Chromatin cleavage and library' preparation can be achieved in one single step.
[0111] CUT& RUN is an epigenomic profiling strategy in which antibody-targeted controlled cleavage by micrococcal nuclease releases specific protein-DNA complexes into the supernatant for paired-end DNA sequencing (see Skene and Henikoff, Elife (2017) 6:1-35, Skene et al., Nat Protoc (2018) 13:1006-1019). As only targeted fragments enter into solution, and the vast majority of DNA is left behind, CUT& RUN has low background levels. In an example CUT& RUN assay, a sample is incubated with an antibody or antibody fragment that binds the target protein, e.g., transcription factor or histone modification of interest. The sample is then incubated with Protein- A-MNase after which CaCh can be added to initiate the calcium dependent nuclease activity of MNase to cleave the DNA around the target protein. The protein-A-MNase reaction can be quenched by adding chelating agents (EDTA and EGTA). Cleaved DNA fragments are then liberated, extracted, and used to construct a sequencing library.Techniques for Detecting and Quantifying DNA Methylation
[0112] DNA from a sample may be sequenced in order to obtain sequencing data for the sample. DNA from a sample may be sequenced in order to determine methylation status. Various methylation sequencing techniques may be used to produce methylation sequencing data. In some embodiments, DNA extracted from plasma is enriched for densely methylated fragments as part of a sequencing method, such as, for example, Methyl-CpG-Binding Domain Sequencing (MBD-seq). In some embodiments, after sequencing, FASTQ files are processed to produce a table of unique fragments that align to a genome for a subject from which a sample is derived (e.g.. the13404083v 1 Page 40 of 258Attorney Docket: 2014191-0049human genome for human subjects / samples). Sequencing data used to produce a cTF estimation model and / or used to estimate cTF for a sample may be preprocessed to get unique fragments from aligned sequencing reads. In some embodiments, sequencing data received and / or obtained may be initially deduplicated.
[0113] Various techniques of molecular biology are well known in the art and / or disclosed in the present application for detecting and quantifying DNA methylation, for example of cfDNA (e.g., ctDNA) in a sample. In some embodiments, the methods and systems of the present disclosure involve the detection and quantification of DNA methylation in samples, e.g., in liquid biopsy samples including cfDNA (e.g., ctDNA) such as plasma samples including ctDNA. Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq) and Methyl-CpG-Binding Domain sequencing (MBD-seq) are exemplary techniques of molecular biology useful in detecting and quantifying DNA methylation in samples, though others, such as Bisulfite sequencing (BS-Seq) and Whole Genome Bisulfite Sequencing (WGBS) exist. In some embodiments, an enrichment method, such as MBD-seq, is used to obtain sequencing data. In some embodiments, sequencing data from an enrichment method, such as MBD-seq, is received and processed.
[0114] DNA methylation typically refers to the methylation of the 5’ position of cytosine (mC) by DNA methyltransferases (DNMT). It is a major epigenetic modification in humans and many other species. In mammals, most DNA methylations occur within the context of CpG dinucleotides. DNA methylation is thought to be a repressive chromatin modification. Aberrant methylation can lead to many diseases including cancers (Robertson, Nat Rev Genet (2005) 6:597-610 and Bergman and Cedar, Nat Struct Mol Biol (2013) 20:274-281).
[0115] MeDIP-seq was first reported by Weber et al., Nat Genet (2005) 37:853-862. In a typical MeDIP-seq protocol, antibody or antibody-fragment that binds 5 -methyl cytidine (5mC) is used to enrich methylated DNA fragments, then these fragments are sequenced and analyzed. If using 5mC-specific antibodies or antibody fragments, methylated DNA is isolated from genomic DNA via immunoprecipitation. Anti-5mC antibodies are incubated with fragmented genomic DNA and precipitated, followed by DNA purification and sequencing.
[0116] Methyl -CpG-Binding Domain sequencing (MBD-seq) is similar to MeDIP-seq except that it uses methyl binding domain (MBD) proteins instead of antibodies or antibody fragments to bind methylated DNA In a typical MBD-seq protocol, genomic DNA is first sonicated and incubated with tagged MBD proteins that can bind methylated cytosines. The protein-DNA13404083v 1 Page 41 of 258Attorney Docket: 2014191-0049complex is then precipitated with antibody-conjugated beads that are specific to the MBD protein tag, followed by DNA purification and sequencing.Techniques for Detecting and Quantifying Chromatin Accessibility
[0117] Various techniques of molecular biology are well known in the art and / or disclosed in the present application for detecting and quantifying chromatin accessibility. In some embodiments, the methods, kits and systems of the present disclosure involve the detection and quantification of chromatin accessibility in samples, e.g., in liquid biopsy samples including cfDNA such as plasma samples including cfDNA ATAC-seq (Assay of Transpose Accessible Chromatin sequencing), NOMe-seq (Nucleosome Occupancy and Methylome sequencing), FAIRE-seq (Formaldehyde- Assisted Isolation of Regulatory Elements sequencing), MNase-seq (Micrococcal Nuclease digestion with sequencing), and DNase hypersensitivity assays are exemplary techniques of molecular biology useful in detecting and quantifying chromatin accessibility in samples. Sono-Seq is another alternative method that could be used (see Auerbach et al., Proc Natl Acad USA (2009) 106(35): 14926-14931).
[0118] DNase hypersensitivity assays can use the non-specific DNA endonuclease Deoxyribonuclease I (DNase I), which selectively digests accessible DNA regions. DNase I hypersensitivity sites (DUS) identified by DNase-seq include open chromatin regulatory regions. A typical DNase hypersensitivity assay can include a first step in which nuclei are isolated from cells using lysis buffer, and nuclei are digested using DNase I. DNA fragment sizes are measured to identify optimal digestion using gel electrophoresis. Biotinylated linkers can be ligated to the ends of digested DNA after polishing to make blunt ends, and the DNA can then be isolated. DNA with biotinylated linker can be digested by restriction endonuclease Mmel and captured by streptavidin coated Dynabeads® to generate short tags to which a second sequencing adaptor can be ligated. A second linker can be ligated and amplified to generate a library for sequencing. A DNase-seq bioinformatic pipeline can include, e.g., alignment of sequence reads to a reference genome with BWA or Bowtie2. Aligned reads can be used to call and quantify peaks as compared to a reference.
[0119] MNase-seq determines chromatin accessibility with micrococcal nuclease (MNase) that preferentially digests nucleosome-free, protein-unbound DNA. A typical MNase-seq assay can include a first step in which nuclei are isolated from either native or crosslinked chromatin and1340408 v 1 Page 42 of 258Attorney Docket: 2014191-0049digested using MNase with titration. In vivo formaldehyde crosslinking step that is designed to capture the interaction between proteins and DNA, This crosslinking allows bound proteins to shield their associated DNA from digestion by MNase. Following crosslinking, samples are digested with MNase, which can be specifically activated by addition of Ca2+ to the buffer. Digestion can be halted by chelating the reaction, at which point the samples are RNase treated, crosslinks are reversed, and proteins are digested away from the chromatin. DNA can then be isolated via a phenol-chloroform extraction. Uncut DNA is purified and mononucleosome bands are isolated and excised through gel electrophoresis. Isolated DNA can be amplified by adding adapters to generate a library, and sequenced. MNase-seq primarily sequences regions of DNA bound by histones or other proteins. Therefore, it indirectly determines which regions of DNA are accessible by directly determining which regions are bound to nucleosomes or proteins.
[0120] FAIRE-seq is a method in which nucleosome-depleted regions of DNA (NDRs) are isolated from chromatin. A typical FAIRE-seq assay can include a first step in which cells are fixed using formaldehyde so that histones are crosslinked to interacting DNA. Crosslinked chromatin can then be sheared by sonication that generates protein-free DNA and protein-crosslinked DNA fragments. Protein-free DNA can be isolated using a phenol-chloroform extraction: DNA crosslinked with protein stays in organic phase, while protein-free DNA stays in aqueous phase. Highly crosslinked DNA remains in the organic phase and the non-crosslinked DNA is pulled to the aqueous phase. Non-crosslinked DNA from the aqueous phase can then be amplified and sequenced. Reads enriched in the sequencing pool tend to have lower nucleosome and transcription factor binding and are therefore inferred to come from accessible regions.
[0121] NOMe-seq is a method to identify nucleosome-depleted regions of DNA (NDRs) with M. CviPI methyltransferase that methylates cytosine in GpC dinucleotides not protected by nucleosomes or other proteins. Unlike CmpG, GpCmin the human genome does not occur naturally in most cell types. GpC”1levels at open chromatin regions can be compared to background signals and used to detect and quantify NDRs. A typical NOMe-seq protocol can include a step in which samples are treated with M. CviPI and S-adenosylhomocysteine (SAM) to methylate accessible GpC sites. M. CviPI treated DNA can be sheared using a sonicator, so that DNA fragments can be sequenced. DNA is treated with bisulfite, which converts unmethylated cytosine to uracil using sodium bisulfite, while methylated cytosine is unaffected. A library is generated using adapters and sequenced. Accessible chromatin is expected to have high levels of GpCmbut low levels of13404083 v 1 Page 43 of 258Attorney Docket: 2014191-0049CmpG. Therefore, NOMe-seq identifies NDRs using the two separate methylation analyses that serve as independent (but opposite) measures, providing matched chromatin designations for each regulatory element.ATAC-seq uses hyperactive Tn5 transposase that preferentially cuts accessible chromatin regions and simultaneously inserts adapters to the fragmented region (Buenrostro et al., Nat Methods (2013) 10(12): 1213-1218 the entirety of which is incorporated herein by reference). A typical ATAC-seq assay can include a first step in which samples are incubated with Tn5 transposase. DNA can then be isolated and purified. DNA fragmented and tagged by Tn5 transposase can be purified and then amplified to generate a library and sequenced for analysis.Exemplary Genomic Loci
[0122] Among other things, the present disclosure identifies exemplary genomic loci that are differentially modified and / or differentially accessible depending on ER, PR, or HER2 expression status, and show that differential modifications and / or accessibility of said exemplary genomic loci can be used to accurately predict (e.g., determine) ER, PR, or HER2 expression status. See Tables 5-7 which show the chromosomal coordinates of each genomic locus.
[0123] The present disclosure is not limited to methods that use the exact same chromosomal coordinates that are recited in Tables 5-7. The present disclosure encompasses methods that use any of the genomic loci in Table 5-7 and also subregions thereof, i.e., references herein to methods that involve detecting and / or quantifying one or more histone modifications, chromatin accessibility, binding of one or more transcription factors, and / or DNA methylation at one or more genomic loci of Table 5-7 encompasses methods that detect these marks anywhere within these genomic loci including within any subregions. For example, where Table 5 references chrl:8720924-8721417 as a genomic locus for detecting and / or quantifying H3K27ac modification, this encompasses methods that detect and / or quantify H3K27ac modification at any position or sub-region of chrl:8720924-8721417, e.g., methods that detect and / or quantify H3K27ac modification within chrl:8721024-8721317, etc. In some embodiments, a subregion may span at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500 or at least 3000 contiguous base pairs that are located between the lower and upper coordinates of a genomic locus recited in Tables 5-7. In some embodiments, a subregion may span less than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500 or at least 3000 contiguous base pairs that13404083 v 1 Page 44 of 258Attorney Docket: 2014191-0049are located between the lower and upper coordinates of a genomic locus recited in Tables 5-7. In some embodiments, a subregion may have the same central coordinate as a genomic locus recited in Tables 5-7. In some embodiments, a subregion may have a different central coordinate as a genomic locus recited in Tables 5-7. It is also to be understood that the lower / upper coordinates of the genomic loci in Tables 5-7 are approximate and that the present disclosure encompasses methods where any one or more of the genomic loci are expanded by increasing the size of the genomic locus by 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40% or up to 50% in one or both directions.
[0124] In some embodiments a classifier is generated using a set of differentially modified and / or differentially accessible genomic loci that are correlated with increased ER, HER2, or PR expression. Sequence reads that fall into each selected genomic locus are analyzed and counted, e g., as described herein including the Examples. In some embodiments, counts from genomic loci that are correlated with increased ER, HER2, or PR expression are aggregated. Other ways of using the genomic loci and related sequencing data to generate and apply a classifier to determine ER, PR, or HER2 expression status are described herein and known in the art, e.g., without limitation, methods that use a learning statistical classifier system or a combination of learning statistical classifier systems.
[0125] In some embodiments, exemplary genomic loci from one or more of Tables 5-7 are used in a monomodal ER, PR, or HER2 expression status classifier, e.g., an expression status classifier that uses a single histone modification (e.g, H3K4me3 or H3K27ac) or DNA methylation at one or more genomic loci for purposes of determining ER, PR, or HER2 expression status. In some embodiments, exemplary genomic loci from any one of Table 5-7, or any combination thereof, are used in combination in a multimodal classifier, e.g., a ER, PR, or HER2 expression status classifier that uses more than one histone modification (e.g., H3K4me3 and H3K27ac) or one or more histone modifications (e g., H3K4me3 and / or H3K27ac) and DNA methylation at one or more genomic loci for purposes of determining ER, PR, or HER2 expression status.
[0126] In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor at one or more loci provided in one or more of Tables 5-7. In some embodiments, a method described herein comprises quantifying one or more of a histone1340408 v 1 Page 45 of 258Attorney Docket: 2014191-0049modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 or more loci listed in one or more of Tables 5-7. In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor at each of the loci provided in one or more of Tables 5-7. In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor for at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, 40%, 50%, 75%, or 100% of loci identified in one or more of Tables 5-7. In some embodiments, a method described herein comprises quantifying one or more of a histone modification, DNA methylation, chromatic accessibility and / or binding of a transcription factor for at least a percent of loci identified in one or more of Tables 5-7 having a lower bound selected from 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 1%, 2%, 3%, 4%, 5%, or 10%, and an upper bound selected from 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, 40%, 50%, 75%, or 100%.Differential H3K4me3 modification
[0127] Exemplary genomic loci whose H3K4 methylation state (in particular, H3K4 trimethylation or H3K4me3 state) is associated with ER, PR, or HER2 expression status (e.g., ER+, PR+, or HER2+- status) are provided in Tables 5-7 (see “H3K4me3” loci).
[0128] A person of skill in the art will recognize that the methods disclosed herein do not require that every H3K4me3 analyte genomic locus listed in Tables 5-7 be assessed for H3K4me3 modification. Instead, a subset of H3K4me3 analyte loci may be assessed for H3K4me3 modification. Subsets of the H3K4me3 analyte genomic loci of Tables 5-7 can be selected (e.g., for use in determining ER, PR, or HER2 expression status) based on various performance criteria, e.g., to select genomic loci that demonstrate differential modification with a particular level of statistical significance and / or a particular threshold of differential between relevant states (e.g., a measured log2(fold-change)). Subsets of the genomic loci may also be selected based on an algorithm, e.g., during the process of obtaining a classifier. Those of skill in the art will appreciate that such subsets of loci of Tables 5-7, and loci included in such subsets, are together, individually, and / or in randomly selected subsets, at least as informative (e.g., as statistically significant and / or reliable) for uses disclosed herein, e.g., for determining ER, PR, or HER2 expression status.1340408 v 1 Page 46 of 258Attorney Docket: 2014191-0049
[0129] In various embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if about I, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 or more H3K4me3 loci identified in Tables 5, 6, or 7 are differentially H3K4me3 modified (e.g., (a) a number of loci identified in Table 5 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, including, e.g., about 1 to about 120, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100; (b) a number of loci identified in Table 6 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, including, e.g., about 1 to about 120, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100; or (c) a number of loci identified in Table 7 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 50, about 60, about 70, or about 80, including, e.g., about 1 to about 50, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, or about 50) as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2 + cancer)).
[0130] In various embodiments, a sample or subject from which the sample is derived, is determined to have an ER+, PR+, or HER2+ cancer if one or more promoter regions of one or more genes (e.g., 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more) in any one of Tables 5-7 are differentially H3K4me3 modified as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).13404083vl Page 47 of 258Attorney Docket: 2014191-0049
[0131] In various embodiments, a promoter region refers to a region a certain number of nucleotides upstream of a gene (e.g,, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, or 1,000 nucleotides upstream of a gene). In some embodiments, a promoter region refers to an H3K4me3 locus identified in Tables 5-7.
[0132] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if H3K4me3 modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K4me3 loci that are identified in Tables 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression are increased relative to a reference (eg., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2 cancer)). In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if H3K4me3 modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K4me3 loci that are identified in Tables 5, 6, or 7 as having a negative association with ER, PR, or HER2 expression are decreased relative to a reference (e.g,, a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).
[0133] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or IIER2 expression status if:(a) H3K4me3 modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K4me3 loci that are identified in Tables 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression are increased relative to a reference (e g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)); and(b) H3K4me3 modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K4me3 loci that are identified in Tables 5, 6, or 7 as having a negative association with ER, PR, or HER2 expression are decreased relative to a reference (e g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of13404083 vl Page 48 of 258Attorney Docket: 2014191-0049subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, orHER2+ cancer)).
[0134] In various embodiments, differentially H3K4me3 modified refers to a methylation status characterized by an increase or decrease in a value measuring methylation (e.g, of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4- fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3 4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e- 5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring methylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase or decrease of 0.1 -fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 40-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold, 1.2-fold to 4.0-fold, 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.Differential H3K27ac modification
[0135] Exemplary genomic loci whose H3K27ac state is associated with ER, PR, and HER2 expression status (e.g., ER+, PR+, or HER2+ status) are provided in Tables 5-7 (see “H3K27ac” loci).
[0136] A person of skill in the art will recognize that the methods disclosed herein do not require that every H3K27ac genomic locus listed in Tables 5-7 be assessed for H3K27ac1340408 v 1 Page 49 of 258Attorney Docket: 2014191-0049modifications. Instead, a subset of H3K27ac loci may be assessed for H3K27ac modification Subsets of the H3K27ac genomic loci of Tables 5-7 can be selected (e.g., for use in determining ER, PR, or HER2 expression status) based on various performance criteria, e.g., to select genomic loci that demonstrate differential modification with a particular level of statistical significance and / or a particular threshold of differential between relevant states (e.g., a measured log2(fold-change)). Subsets of the genomic loci may also be selected based on an algorithm, e.g., during the process of obtaining a classifier. Those of skill in the art will appreciate that such subsets of loci of Tables 5-7, and loci included in such subsets, are together, individually, and / or in randomly selected subsets, at least as informative (e.g., as statistically significant and / or reliable) for uses disclosed herein, e.g., for determining ER, PR, or HER2 expression status.
[0137] In various embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if about 1, about 2, about3, about 4, about5, about 10, about 15, about 20, or about 25 or more H3K27ac loci identified in Tables 5, 6, or 7 (e g., (a) a number of loci identified in Table 5 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 120, about 150, about 200, about 250, about 300, about 350, about 400, including, e.g., about 1 to about 400, about 1 to about 300, about 1 to about 200, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100; (b) a number of loci identified in Table 6 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 150, about 200, about 250, about 300, about 400, about 500, about 600, including, e.g., about 1 to about 400, about 1 to about 300, about 1 to about 200, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100; or (c) a number of loci identified in Table 7 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 50, about 60, about 70, or about 80, including, e.g., about 1 to about 50, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, or about 50) are differentially H3K27ac modified as compared to a1340408 v 1 Page 50 of 258Attorney Docket: 2014191-0049reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or 1IER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).
[0138] In various embodiments, a sample or subject from which the sample is derived, is determined to have a particular ER, PR, or HER2 expression status if one or more enhancer regions of one or more (e.g., 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more) H3K27ac loci in Tables 5, 6, or 7 are differentially H3K27ac modified as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER, PR, or HER2 expression status or a cohort of subjects with aberrant an ER, PR, or HER2 expression status (e.g., a subject or cohort of subjects with an ER+, PR+, or HER2+ cancer)).
[0139] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if H3K27ac modifications for 1, 2, 3, 4, 5, 10, 15, 20, or 25, or more of the H3K27ac loci identified in Tables 5, 6, or 7 are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER, PR, or HER2 expression status or a cohort of subjects with an aberrant ER, PR, or HER2 expression status (e.g., a subj ect or cohort of subjects with an ER+, PR+, or HER2+ cancer)).
[0140] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if H3K27ac modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K27ac loci that are identified in Tables 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression are increased relative to a reference (e g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)). In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if H3K27ac modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K27ac loci that are identified in Tables 5, 6, or 7 as having a negative association with ER, PR, or HER2 expression are decreased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).1340408 v 1 Page 51 of 258Attorney Docket: 2014191-0049
[0141] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if:(a) H3K27ac modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K27ac loci that are identified in Tables 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)); and(b) H3K27ac modifications for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the H3K27ac loci that are identified in Tables 5, 6, or 7 as having a negative association with ER, PR, or HER2 expression are decreased relative to a reference (e g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).
[0142] In various embodiments, differentially H3K27ac modified refers to an acetylation status characterized by an increase or decrease in a value measuring acetylation (e.g,, of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16- fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16- fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or l e-6. In various embodiments, an increase or decrease in a value measuring acetylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase or decrease of 0.1-fold to 10-fold,13404083 v 1 Page 52 of 258Attorney Docket: 2014191-00490.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold, 1.2-fold to 4.0-fold, 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0- fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.
[0143] In some embodiments, one or more enhancer regions of a recited gene are provided in Tables 5-7. In some embodiments, one or more enhancer regions of a recited gene corresponds to: (i) one or more loci with increased or decreased H3K27ac modifications as compared to a reference (e.g., a sample from a healthy subject) within a certain number of nucleotides (e.g., 50,000 nucleotides) of the recited gene; and / or (ii) one or more loci with increased or decreased H3K27ac modifications as compared to a reference (e.g., a sample from a healthy subject) that are closest to the recited gene in the genome.Differential DNA methylation
[0144] Exemplary genomic loci whose DNA methylated state is associated with ER, PR, or HER2 expression status (e.g,, ER+, PR+, or HER2+ status) are provided in Table 5-7 (see " MBD” loci).
[0145] A person of skill in the art will recognize that the methods disclosed herein do not require that every MBD locus listed in Tables 5-7 be assessed for DNA methylation. Instead, a subset of MBD loci may be assessed for DNA methylation. Subsets of the MBD loci of Tables 5-7 can be selected (e.g., for use in determining ER, PR, or HER2 expression status) based on various performance criteria, e.g., to select genomic loci that demonstrate differential modification with a particular level of statistical significance and / or a particular threshold of differential between relevant states (e.g., a measured log2(fold-change)). Subsets of the genomic loci may also be selected based on an algorithm, e.g., during the process of obtaining a classifier. Those of skill in the art will appreciate that such subsets of loci of Tables 5-7, and loci included in such subsets, are together, individually, and / or in randomly selected subsets, at least as informative (e.g,, as statistically significant and / or reliable) for uses disclosed herein, e.g., for determining ER, PR, or HER2 expression status
[0146] In various embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if about 1, about 2,1340408 v 1 Page 53 of 258Attorney Docket: 2014191-0049about 3, about 4, about 5, about 10, about 15, about 20, or about 25 or more MBD loci identified in Tables 5, 6, or 7 (e.g, (a) a number of loci identified in Table 5 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 100, about 105, about 110, about 115, about 125, about 150, about 200, about 250, including, e.g., about 1 to about 200, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100; (b) a number of loci identified in Table 6 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about. 20, about 25, or about 50 and an upper bound of about 100, about 105, about 110, about 115, about 125, about 150, about 200, about 250, about 300, about 350, including, e.g., about 1 to about 300, about 1 to about 200, about 1 to about 100, about 5 to about 100, about 10 to about 100, about 15 to about 100, about 20 to about 100, about 25 to about 100, about 30 to about 100, about 35 to about 100, about 40 to about 100, or about 50 to about 100; or (c) a number of loci identified in Table 7 within a range having a lower bound of about 1, about 2, about 3, about 4, about 5, about 10, about 15, about 20, about 25, or about 50 and an upper bound of about 50, about 60, about 70, or about 80, including, e.g., about 1 to about 50, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, or about 50) are differentially DNA methylated as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER, PR, or HER2 status or a cohort of subjects with an aberrant ER, PR, or HER2 expression status (e.g., a subject or cohort of subjects with an ER+, PR+, or HER2+ cancer)).
[0147] In various embodiments, a sample or subject from which the sample is derived, is determined to have a particular ER, PR, or HER2 expression status if one or more MBD loci (e.g., 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more loci) in Table 5, 6, or 7 are differentially DNA methylated as compared to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER, PR, or HER2 expression status or a cohort of subjects with an aberrant ER, PR, or HER2 expression status (e.g., a subject or cohort of subjects with an ER+, PR+, or HER2+ cancer)).
[0148] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if DNA methylation13404083v 1 Page 54 of 258Attorney Docket: 2014191-0049for 1, 2, 3, 4, 5, 10, 15, 20, or 25 or more MBD loci that are identified in Tables 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression status are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with an aberrant ER, PR, or HER2 expression status or a cohort of subjects with an aberrant ER, PR, or HER2 expression status (e.g., a subject or cohort of subjects with an ER+, PR+, or HER2+ cancer)).
[0149] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Tables 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression are increased relative to a reference (e g, a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)). In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Tables 5, 6, or 7 as having a negative association with ER, PR, or HER2 expression are decreased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HERZ expression or a cohort of subjects with aberrant ER, PR or HERZ expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).
[0150] In some embodiments, a sample or subject from which the sample is obtained or derived, is determined to have a particular ER, PR, or HER2 expression status if:(a) DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Table 5, 6, or 7 as having a positive association with ER, PR, or HER2 expression are increased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)); and(b) DNA methylation for 1, 2, 3, 4, 5, 10, 15, or 25 or more of the MBD loci that are identified in Tables 5, 6, or 7 as having a negative association with ER, PR, or HER2 expression are decreased relative to a reference (e.g., a sample from (i) a healthy subject or cohort of healthy13404083v 1 Page 55 of 258Attorney Docket: 2014191-0049subjects or (ii) a subject with aberrant ER, PR, or HER2 expression or a cohort of subjects with aberrant ER, PR, or HER2 expression (e.g., a subject or cohort of subjects with a ER+, PR+, or HER2+ cancer)).
[0151] In various embodiments, differentially DNA methylated refers to a methylation status characterized by an increase or decrease in a value measuring methylation (e.g., of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring methylation can be, or is expressed as, a log2(fold-change), e.g., a log2(f old-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase of 0.1-fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold, 1.2-fold to 4.0-fold, 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6Differential chromatin accessibility or transcription factor binding
[0152] Genomic loci provided in Tables 5-7 can also demonstrate differential chromatin accessibility or transcription factor binding in different PR, ER, or HER2 expression states.
[0153] In various embodiments, without wishing to be bound by any particular scientific theory, histone methylation (e.g., H3K4me3) corresponds and / or is correlated with chromatin1340408 v 1 Page 56 of 258Attorney Docket: 2014191-0049accessibility. In various embodiments, without wishing to be bound by any particular scientific theory, histone acetylation (e.g., H3K27ac) corresponds and / or is correlated with chromatin accessibility. In various embodiments, without wishing to be bound by any particular scientific theory, DNA methylation corresponds and / or is correlated with chromatin accessibility.
[0154] In some embodiments, without wishing to be limited to any particular scientific theory, chromatin accessibility corresponds and / or is correlated with H3K4me3 modifications. As a result, in some embodiments, ER, PR, or HER2 expression status may be determined by detecting and quantifying chromatin accessibility at one or more genomic loci in Table 5, 6, or 7 in accordance with the section above discussing exemplary genomic loci with differential H3K4me3 modifications.
[0155] In some embodiments, without wishing to be limited to any particular scientific theory, chromatin accessibility corresponds and / or is correlated with H3K27ac modifications. As a result, in some embodiments, ER, PR, or HER2 expression status may be determined by detecting and quantifying chromatin accessibility at one or more genomic loci in Tables 5, 6, or 7 in accordance with the section above discussing exemplary' genomic loci with differential H3K27ac modifications.
[0156] In some embodiments, without wishing to be limited to any particular scientific theory, chromatin accessibility corresponds and / or is correlated with DNA methylation. As a result, in some embodiments, ER, PR, or HER2 expression status can be measured by detecting and quantifying chromatin accessibility at one or more genomic loci in Tables 5, 6, or 7 in accordance with the section above discussing exemplary genomic loci with differential DNA methylation.
[0157] In various embodiments, without wishing to be bound by any particular scientific theory, histone methylation (e.g., H3K4me3) corresponds and / or is correlated with transcription factor binding. In various embodiments, without wishing to be bound by any particular scientific theory', histone acetylation (e.g., H3K27ac) corresponds and / or is correlated with transcription factor binding. In various embodiments, without wishing to be bound by any particular scientific theory, DNA methylation corresponds and / or is correlated with transcription factor binding.
[0158] In some embodiments, without wishing to be limited to any particular scientific theory', binding of RNA pol II corresponds and / or is correlated with H3K4me3 modifications. As a result, in some embodiments, ER, PR, or HER2 expression status may be determined by detecting13404083 v 1 Page 57 of 258Attorney Docket: 2014191-0049and quantifying binding of RNA pol II at one or more genomic loci in Tables 5, 6, or 7 in accordance with the section above discussing exemplary genomic loci with differential H3K4me3 modifications.
[0159] In some embodiments, without wishing to be limited to any particular scientific theory, binding of p300, mediator complex, cohesin complex or RNA pol II corresponds and / or is correlated with H3K27ac modifications. As a result, in some embodiments, ER, PR, or HER2 expression status may be determined by detecting and quantifying binding of p300, mediator complex, cohesin complex or RNA pol II at one or more genomic loci in Tables 5-7 in accordance with the section above discussing exemplary genomic loci with differential H3K27ac modifications.Using Models
[0160] Model produced according to methods disclosed herein may be used to predict expression level of a target gene in a subject based on one or more samples from the subject, for example liquid biopsy samples. A method may include receiving sample sequencing data derived from a nucleic acid in a biological sample derived from a subject. In some embodiments, sample sequencing data includes a signal for each of one or more epigenetic biomarkers for a target gene. In some embodiments, sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene. A method may include determining, using sample sequencing data, a tile signal for each of one or more epigenetic biomarkers for a set of expression-level correlated tiles corresponding to a genomic region. A method may include aggregating, for each of one or more epigenetic biomarkers, a tile signal for the epigenetic biomarker for a set of expression-level correlated tiles to obtain an aggregated tile signal for the epigenetic biomarker.
[0161] A method may include predicting (e.g., using a model) an expression level for a target gene for a subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on aggregated tile signal for each of one or more epigenetic biomarkers, A method may include predicting (e.g., using a model) an expression level for a target gene for a subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on signal for each of the one or more epigenetic biomarkers.1340408 v 1 Page 58 of 258Attorney Docket: 2014191-0049
[0162] A method may include normalizing signal in sample sequencing data for each of one or more epigenetic biomarkers. Normalizing signal may comprises a quantile normalization. A model may be a regression (e.g., a multiple regression) based model [e.g., is a regression (e.g., a multiple regression)]. A regression may be an ordinary least squares (OLS) regression (e.g., an OLS multiple regression). In some embodiments, an expression level is an mRNA expression level. In some embodiments, an expression level is a protein expression level. In some embodiments, one or more epigenetic biomarkers is a plurality of epigenetic biomarkers. In some embodiments, the tiles in a set of expression-level correlated tiles are mutually non-overlapping. In some embodiments, a set of expression-level correlated tiles comprises tiles having non-uniform size.
[0163] In some embodiments, a method includes tiling sample sequencing data into tiles that together span a genomic region corresponding to a target gene to obtain tiled sample sequencing data. Determining tile signal may be performed using tiled sample sequencing data. In some embodiments, the tiles are overlapping (e.g., wherein no more than two of the tiles are mutually overlapping). In some embodiments, adjacent ones of the tiles overlap by at least 10% (eg, at least 20%, at least 30% of a length of the tiles) and no more than 70% (e.g., no more than 60% or no more than 50% of the length of the tiles). In some embodiments, tiles have a length in a range of from 100 to 1000 bp (e.g., from 250 to 750 bp) In some embodiments, adjacent ones of the tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp).
[0164] In some embodiments, predicting expression level comprises using a model that has been trained on digital samples, each comprising (a) a signal for one or more epigenetic biomarkers for a target gene, and (b) an expression level for the target gene. In some embodiments, a model has been produced based on signal for one or more epigenetic biomarkers for a set of expression-level correlated tiles and an expression level for the target gene for each sample of the digital samples.
[0165] A method may include determining a circulating tumor DNA (ctDNA) fraction for a sample, for example in order to predict expression level for a target gene using sample sequencing data for the sample. Predicting expression level may further be based on ctDNA fraction. Determining a ctDNA fraction may include determining ctDNA fraction using sample sequencing data. Determining ctDNA fraction may include determining ctDNA fraction based on a ctDNA signal for one or more epigenetic biomarkers (e.g., one or more of the biomarker(s) used1340408 v 1 Page 59 of 258Attorney Docket: 2014191-0049for predicting expression level) In some embodiments, ctDNA signal for one or more epigenetic biomarkers comprises a ctDNA signal from one or more histone modifications. In some embodiments, ctDNA signal for one or more epigenetic biomarkers comprises a ctDNA signal from DNA methylation.
[0166] A model may accept aggregated tile signal derived from sample sequencing data as input. A model may accept tiled sample sequencing data as input. A model may accept sample sequencing data as input. A model may (i) tile sample sequencing data, (ii) determine tile signal for each of one or more epigenetic biomarkers for a set of expression-level correlated tiles using the sample sequencing data, (iii) aggregate the tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles, and (iv) predict expression level for a target gene for a sample based on the aggregated tile signal.
[0167] In some embodiments, a method includes pebbling signal for each of one or more epigenetic biomarkers. Pebbling may comprise, for each of one or more epigenetic biomarkers, determining a background signal for the epigenetic biomarker for a genomic region based on a signal (e.g., a number of fragments) in the sample in a pebbling region, and subtracting the background signal from signal for the epigenetic biomarker.
[0168] In some embodiments, a method includes determining, for each of one or more epigenetic biomarkers, a point estimate for a tile signal for the epigenetic biomarker across a set of expression-level correlated tiles. A point estimate may be a mean, such as a geometric mean. In some embodiments, a method includes predicting an expression level for a target gene based on a respective point estimate of signal for each of one or more epigenetic biomarkers for a set of expression-level correlation tiles. In some embodiments, a method comprises determining, for each of one or more epigenetic biomarkers, a plurality of point estimates for signal for a set of expression-level correlated tiles. In some embodiments, a plurality of point estimates comprises a first point estimate for positively associated tiles of a set of expression-level correlated tiles for one of one or more epigenetic biomarkers and a second point estimate for negatively associated tiles of a set of expression-level correlated tiles for one of one or more epigenetic biomarkers.
[0169] A method of predicting expression level of a target gene in a subject may include predicting expression level of the target gene using sample sequencing data for multiple genes. A method of predicting expression level of a target gene in a subject may include predicting expression level of the target gene using sample sequencing data for the target gene and one or1340408 v 1 Page 60 of 258Attorney Docket: 2014191-0049more additional (e.g, related) genes. A method of predicting expression level of a target gene in a subject may include predicting expression level of the target gene using sample sequencing data for one or more genomic regions corresponding to a target gene and one or more additional (e.g., related) genes, for example a genomic region for the target gene and one or more genomic regions for the one or more additional genes (e g., one distinct genomic region for each of the one or more additional genes).
[0170] A method may include receiving sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, where the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for two or more genes (e.g., for one or more genomic regions corresponding to the two or more genes). In some embodiments, two or more genes comprise a target gene and one or more additional (e.g., related) genes. A method may include predicting an expression level for a target gene for a subject (e.g., having an indication, such as cancer) (e.g., at the time a sample was taken) based on signal for reach of one or more epigenetic biomarkers for two or more genes. A method may include determining, using sample sequencing data, a tile signal for each of one or more epigenetic biomarkers for each of two or more genes for a set of expression-level correlated tiles. A method may include aggregating tile signal for one or more epigenetic biomarkers signal for the set of expression-level correlated tiles to obtain aggregated tile signals for each of the two or more genes. A method may include tiling sample sequencing data into tiles that together span one or more genomic regions corresponding to each of the two or more genes to obtain tiled sample sequencing data, optionally where aggregating is performed using the tiled sample sequencing data. A method may include predicting an expression level for a target gene for a subject (e.g., having an indication, such as cancer) (e.g., at the time a sample was taken) based on aggregated tile signals for each of two or more genes.
[0171] In some embodiments, one or more or more additional genes comprise one or more genes related to a target gene. In some embodiments, one or more related genes comprises one or more genes that are in a transcriptional complex with a target gene. In some embodiments, one or more related genes comprises one or more genes that are master regulators for the indication. In some embodiments, one or more related genes comprises one or more genes that are related to a pathway relevant for an indication In some embodiments, one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes. In some embodiments,13404083 v 1 Page 61 of 258Attorney Docket: 2014191-0049one or more related genes comprises a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].
[0172] In some embodiments, a method includes predicting expression level for a target gene using an ensemble model. In some embodiments, an ensemble model comprises a first constituent model and a second constituent model. In some embodiments, an ensemble model comprises a respective constituent model for a target gene and for each of one or more additional genes. In some embodiments, a first constituent model is a model for a target gene and a second constituent model is a model for one or more of one or more additional genes. In some embodiments, an ensemble model uses a ridge regression of a plurality of constituent models, for example a first constituent model and a second constituent model. In some embodiments, a first constituent model comprises (e.g., is) a model for the target gene. In some embodiments, a second constituent model comprises (e.g., is) a single gene model for a gene related to a target gene.
[0173] In some embodiments, a target gene whose expression level is predicted by a method (e.g., using a model) is a gene that encodes a polypeptide associated with (e.g., targeted by) a therapy. In some embodiments, a target gene whose expression level is predicted by a method (e.g,, using a model) is a gene that encodes an antibody drug conjugate (ADC) target antigen. In some embodiments, the method includes predicting a expression level for a target gene for an indication. In some embodiments, the indication is a cancer indication (eg., a cancer type or subtype). The indication may be, for example, breast cancer, ovarian cancer, bladder cancer, renal cancer, prostate cancer, pancreatic cancer, or lung cancer (e.g., small cell lung cancer).
[0174] In some embodiments, a target gene whose expression level is predicted by a method (e.g., using a model) is listed in Table 1, and an indication corresponding to the expression level is prostate cancer. In some embodiments, a target gene whose expression level is predicted by a method (e.g., using a model) is listed in Table 2, and an indication corresponding to the expression level is breast cancer. In some embodiments, a target gene whose expression level is predicted by a method (e.g., using a model) is listed in Table 3, and an indication corresponding to the expression level is lung cancer (e.g,, small cell lung cancer). In some embodiments, a target gene whose expression level is predicted by a method (e.g., using a model) is listed in Table 4 and an indication corresponding to the expression level is cancer.
[0175] Predictions of gene expression using a model, method, and / or system disclosed herein can be made for different samples at different times. Such predictions can be used to1340408 v 1 Page 62 of 258Attorney Docket: 2014191-0049diagnose and / or monitor subjects. In some embodiments, a model is used to quantify expression level of one or more target genes (e.g., of diagnostic interest) for a subject that has or is suspected of having an indication, such as, for example, a cancer. Different samples, such as liquid biopsy (e.g., plasma) samples, may be taken from a subject and used to monitor the subject and / or diagnose the subject. Such target genes may be targets of one or more therapeutic agents (e.g., ADC therapies and / or radiotherapy) that have been, are being, or will be administered to a subject. Models disclosed herein may be used to track response of certain genes to treatment with a therapy. For example, a subject may be treated with an ER degrader and a model may be used to monitor reduction in expression of one or more ER target genes for the subject using data derived from liquid biopsy samples for the subject. Pathways could also be monitored. For example, an ensemble model that includes constituent models for multiple genes corresponding to a pathway (e.g., a target gene and one or more related genes) could be used. Alternatively or additionally, a model that has been produced based on expression-level correlated tiles for a plurality of genes corresponding to a pathway could be used. As an example, if there were five particular genes whose transcription correlated with response to DNA damage and a subject were treated with a DNA damaging agent as a therapy, the subject could be monitored to track change in expression level across the five genes to see if the subject was responding (e.g., if the tumor was responding appropriately). As another example, expression level of ER response genes could be used to monitor a subject that has breast cancer undergoing an endocrine therapy.
[0176] In some embodiments, a method is for predicting an expression level of a target gene in a subject. Such a method may include providing sample data for a subject, wherein the sample data include signal for one or more epigenetic biomarkers for a target gene (e.g., and optionally one or more related genes). Such a method may further include providing a model that has been produced using a method disclosed herein. Such a method may further include predicting an expression level of the target gene for the subject using the sample data with the model In some embodiments, the sample data have been derived from a liquid biopsy sample (e.g., from plasma) for the subject.
[0177] In some embodiments, a method is for monitoring a subject. Such a method may include providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time. Such a method may further include predicting a first13404083 v 1 Page 63 of 258Attorney Docket: 2014191-0049expression level for a target gene using the signal for the first sample and a second expression level for the target gene using the signal for the second sample using a model that has been produced using a method disclosed herein, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g., an ADC therapy) corresponding to the target gene. Such a method may further include determining a difference in expression level between the first sample and the second sample. In some embodiments, the change is a reduction in expression level and the therapy is a degrader for the target gene.
[0178] In some embodiments, a method is for monitoring a subject. Such a method may include providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time. Such a method may further include predicting a first expression level for a set of related genes using the signal for the first sample and a second expression level for the related gene using the signal for the second sample using a model that has been produced according to a method disclosed herein, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g., an ADC therapy) corresponding to at least one of the related genes. Such a method may further include determining a difference in expression level between the first sample and the second sample. In some embodiments, the related genes correspond to a pathway.
[0179] In some embodiments, a method is for characterizing cancer recurrence and / or progression. Such a method may include predicting an expression level of a target gene for a subject having a cancer based on a first sample for the subject taken at a first time point using a model that has been produced using a method disclosed herein. Such a method may further include predicting an expression level of a target gene for the subject based on a second sample for the subject taken at a second time point after the first time point using a method disclosed herein. Such a method may further include determining a difference in the expression level at the second time point and at the first time point.
[0180] In some embodiments, a method is for monitoring cancer in a subject. Such a method may include predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein. Such a method may further include determining whether there is a difference in expression level of the target gene for the subject over time1340408 v 1 Page 64 of 258Attorney Docket: 2014191-0049
[0181] In some embodiments, a method is for determining effectiveness of a therapeutic agent in a subject having cancer. Such a method may include predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein. Such a method may further include determining whether there is a difference in expression level of the target gene for the subject over time.
[0182] In some embodiments, a method is for monitoring response of a subject having an indication to a therapeutic agent for the indication. Such a method may include predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein. Such a method may further include determining whether there is a difference in expression level of the target gene for the subject over time.
[0183] In some embodiments, a method includes administering a therapy including a therapeutic agent to a subject between when two or more samples were obtained. In some embodiments, a method includes administering a therapy to the subject when a difference in expression level of a target gene is determined to be at least as large as a threshold difference. In some embodiments, a method includes altering administration of a therapy to a subject when a difference in expression level over time is determined to be at least as large as a threshold difference. In some embodiments, altering administration includes increasing a dosage and / or frequency of administration.
[0184] In some embodiments, a method is for prognosing cancer. Such a method may include predicting an expression level of a target gene for a subject having a cancer based on a sample for the subject using a model that has been produced using a method disclosed herein. Such a method may further include prognosing cancer in the subject based on the determined expression level. In some embodiments, a method includes administering a therapy based on a prognosis.
[0185] In some embodiments, a method is for diagnosing cancer in a subject. Such a method may include predicting an expression level of a target gene for a subject having a cancer based on a sample for the subject using a model that has been produced using a method disclosed herein. Such a method may further include determining that the expression level exceeds a threshold. In some embodiments, a method includes initiating administration of a therapy based1340408 v 1 Page 65 of 258Attorney Docket: 2014191-0049on the expression level. In some embodiments, a method includes selecting a dosing regimen for a therapy based on the expression level,
[0186] In some embodiments, a method is for determining whether a cancer has been removed from a subject, the method including, after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, predicting expression level of a target gene for the subject based on a sample for the subject using a model that has been produced using a method disclosed herein. In some embodiments, a method includes continuing administration of a therapy based on the expression level. In some embodiments, a method includes ceasing administration of a therapy based on the expression levelFormulation and Administration of ADC Therapy
[0187] The present disclosure includes methods where an antibody-drug conjugate (ADC) therapy is administered to a subject having an indication. The ADC therapy may target an antigen encoded by a target gene identified using a method disclosed herein. The ADC therapy may target an antigen encoded by a target gene based on a predicted expression level for the target gene for an indication as determined by a method disclosed herein. The subject may be selected (e.g., identified) using a model produced as part of a method disclosed herein or a model derived (e.g., tuned) from a model produced as part of a method disclosed herein, A sample from the subject may be taken and used to quantify one or more epigenetic biomarkers for the sample for a target gene that encodes an antigen targeted by the ADC. The quantified one or more epigenetic biomarkers may be used to predict an expression level for the subject and the subject may be selected to be administered the ADC therapy based on the predicted expression level. The indication may be cancer. Digital samples used in a method disclosed herein may be chosen (e.g., generated) based on a desired indication for a new ADC therapy (e.g., using a new ADC). A subject may be selected to be administered an ADC therapy based on the subject having an indication that was used to identify a target gene used for developing the ADC therapy. Models used to identify target genes as candidates for one or more new ADC therapies may be used (e.g., as-is or after tuning) to predict expression levels for subjects as part of subject selection for ADC therapy administration.
[0188] A target for an ADC may identified using a method disclosed herein. Methods disclosed herein may be used to identify one or more targets for an ADC based on or using1340408 v 1 Page 66 of 258Attorney Docket: 2014191-0049predicted expression level(s) for the target(s) determined using the methods (e.g, based on a determination that one or more epigenetic biomarkers are predictive of expression level for the target gene). In general, ADC therapy provided herein will be available, appropriate, and / or preferred for a determined ADC target status. Those of skill in the art will be aware of recommended and / or governmentally approved formulations and / or dosages for various ADC therapies provided herein.
[0189] In some embodiments, a method includes forming (e.g., manufacturing and / or formulating) an ADC for a target gene (e.g., for use in an ADC therapy). The method may include forming an ADC that binds to a target antigen encoded by a target gene that has been determined (e.g., identified) using a method disclosed herein. In some embodiments, the method includes forming an ADC that binds to a target antigen encoded by a target gene that has been determined based on a result of a method disclosed herein. An ADC therapy may be formed such that administration of the ADC therapy may treat an indication for digital samples used in a method disclosed herein. In some embodiments, determining the gene that encodes an ADC target antigen using the method disclosed herein.
[0190] In some embodiments, a method of selecting a target antigen for an ADC includes performing a method disclosed herein. In some embodiments, the method includes forming (e.g., manufacturing and / or formulating) the ADC,
[0191] In some embodiments, a method includes treating a subject having an indication (e g, cancer) with an ADC therapy. The method may include administering an ADC therapy to a subject having an indication. The ADC therapy may target an antigen encoded by a target gene that has been identified for the indication using a method disclosed herein. In some embodiments, the subject has been selected at least in part based on a quantified level of one or more epigenetic biomarkers for the target gene. In some embodiments, the quantified level has been determined using cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject (e.g., that has a ctDNA fraction of no more than 5%, no more than 3%, no more than 2%, or no more than 1%). In some embodiments, the method includes obtaining the quantified level of the one or more epigenetic biomarkers. In some embodiments, the subject has been selected to receive the ADC therapy at least in part based on a predicted expression level that has been predicted using a model disclosed herein or a model derived therefrom. The model may use the quantified level of the one or more epigenetic biomarkers to produce the predicted expression level13404083 v 1 Page 67 of 258Attorney Docket: 2014191-0049
[0192] The present disclosure includes pharmaceutical compositions for delivery of one or more ADCs to a subject. As disclosed herein, a pharmaceutical composition may be in any form known in the art, including formulations for administration according to any route known in the art. A suitable means of administration can be selected based on the age and condition of a subject
[0193] Pharmaceutical composition forms of the present disclosure can include, e.g., liquid, semi-solid and solid dosage forms. Pharmaceutical composition forms of the present disclosure can include, e.g., liquid solutions (e.g., injectable and infusible solutions), dispersions or suspensions, tablets, pills, powders, and liposomes. Selection or use of any particular form may depend, in part, on the intended mode of administration and therapeutic application. Accordingly, the compositions can be formulated for administration by a parenteral mode (e.g., intravenous, subcutaneous, intraperitoneal, or intramuscular injection) or a non-parenteral mode. As used herein, parenteral administration refers to modes of administration other than enteral and topical administration, usually by injection or infusion.
[0194] In some embodiments, the compositions provided herein are present in unit dosage form, which unit dosage form can be suitable for self-administration. Such a unit dosage form may be provided within a container, e.g., a pill, vial, cartridge, prefilled syringe, or disposable pen.
[0195] A pharmaceutical composition of the present disclosure can be in an injectable or infusible form. For example, the present disclosure includes sterile formulations for injection or infusion, which can be formulated in accordance with conventional pharmaceutical practices. Sterile solutions can be prepared by incorporating a composition described herein in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filter sterilization. Solutions can be formulated, e.g., using distilled water, physiological saline, or an isotonic solution containing glucose and other supplements such as D-sorbitol, D-mannose, D-mannitol, or sodium chloride as an aqueous solution for injection, optionally in combination with a suitable solubilizing agent, for example, an alcohol such as ethanol and / or a polyalcohol such as propylene glycol or polyethylene glycol, and / or a nonionic surfactant such as polysorbate 80™ or HCO-50, and the like. In the case of sterile powders for the preparation of sterile injectable solutions, methods for preparation include vacuum drying and freeze-drying that yield a powder of a composition described herein plus any additional desired ingredient (see below) from a previously sterile-filtered solution thereof. The proper fluidity of a solution can be maintained, for example, by the use of a coating such as lecithin, by the13404083 v 1 Page 68 of 258Attorney Docket: 2014191-0049maintenance of the required particle size in the case of dispersion and by the use of surfactants Prolonged absorption of injectable compositions can be brought about by including in the composition a reagent that delays absorption, for example, monostearate salts, and gelatin. In particular instances, a pharmaceutical composition can be formulated, for example, as a buffered solution at a suitable concentration and suitable for storage, e.g., at 2-8°C (e.g., 4°C).
[0196] In various embodiments, a pharmaceutical composition of the present disclosure can be formulated as a solution, microemulsion, dispersion, liposome, or other ordered structure suitable for stable storage at high concentration. Generally, dispersions are prepared by incorporating a composition described herein into a sterile vehicle that contains a basic dispersion medium and the required other ingredients from those enumerated above.
[0197] In various instances, a pharmaceutical composition can be formulated to include a pharmaceutically acceptable carrier or excipient. Examples of pharmaceutically acceptable carriers include, without limitation, any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible
[0198] In certain embodiments, compositions can be formulated with a carrier that will protect the ADC against rapid release, such as a controlled release formulation, including implants and microencapsulated delivery systems. Biodegradable, biocompatible polymers can be used, such as ethylene vinyl acetate, polyanhydrides, polyglycolic acid, collagen, polyorthoesters, and polylactic acid. Many methods for the preparation of such formulations are known in the art See, e.g., J. R. Robinson (1978) “Sustained and Controlled Release Drug Delivery Systems,” Marcel Dekker, Inc., New York.
[0199] Route of administration can be parenteral, for example, administration by injection. Administration by injection can be by intravenous injection, intramuscular injection, intraperitoneal injection, subcutaneous injection Administration can be systemic or local. In certain embodiments, a composition described herein can be therapeutically delivered to a subject by way of local administration. As used herein, “local administration” or “local delivery,” can refer to delivery that does not rely upon transport of the composition or ADC to its intended target tissue or site via the vascular system. For example, the composition may be delivered by injection or implantation of the composition or ADC or by injection or implantation of a device containing the composition or ADC. In certain embodiments, following local administration in the vicinity1340408 v 1 Page 69 of 258Attorney Docket: 2014191-0049of a target tissue or site, the composition or ADC, or one or more components thereof, may diffuse to an intended target tissue or site that is not the site of administration.
[0200] A pharmaceutical composition can be administered parenterally in the form of an injectable formulation comprising a sterile solution or suspension in water or another pharmaceutically acceptable liquid. For example, a pharmaceutical composition can be formulated by suitably combining the therapeutic molecule with pharmaceutically acceptable vehicles or media, such as sterile water and physiological saline, vegetable oil, emulsifier, suspension agent, surfactant, stabilizer, flavoring excipient, diluent, vehicle, preservative, binder, followed by mixing in a unit dose form required for generally accepted pharmaceutical practices. Examples of oily liquid include sesame oil and soybean oil, and it may be combined with benzyl benzoate or benzyl alcohol as a solubilizing agent. Other items that may be included are a buffer such as a phosphate buffer, or sodium acetate buffer, a soothing agent such as procaine hydrochloride, a stabilizer such as benzyl alcohol or phenol, and an antioxidant. The formulated injection can be packaged in a suitable ampule.
[0201] In various embodiments, subcutaneous administration can be accomplished by means of a device, such as a syringe, a prefilled syringe, an auto-injector (e.g., disposable or reusable), a pen injector, a patch injector, a wearable injector, an ambulatory syringe infusion pump with subcutaneous infusion sets, or other device for combining with a ADC for subcutaneous injection.
[0202] An injection system of the present disclosure may employ a delivery pen as described in U. S. Pat. No. 5,308,341. Pen devices, most commonly used for self-delivery of insulin to patients with diabetes, are well known in the art. Such devices can include at least one injection needle, are typically pre-filled with one or more therapeutic unit doses of a solution that includes the ADC and are useful for rapidly delivering solution to a subject with as little pain as possible. One medication delivery pen includes a vial holder into which a vial of a therapeutic or other medication may be received. The pen may be an entirely mechanical device or it may be combined with electronic circuitry to accurately set and / or indicate the dosage of medication that is injected into the user. See, e.g., U. S. Pat. No. 6,192,891. In some embodiments, the needle of the pen device is disposable and the kits include one or more disposable replacement needles. Pen devices suitable for delivery of any one of the presently featured compositions are also described in, e.g., U. S. Pat. Nos. 6,277,099; 6,200,296; and 6,146,361, the disclosures of each of which are1340408 v 1 Page 70 of 258Attorney Docket: 2014191-0049incorporated herein by reference in their entirety. A microneedle-based pen device is described in, e.g., U. S. Pat. No, 7,556,615, the disclosure of which is incorporated herein by reference in its entirety. See also the Precision Pen Injector (PPI) device, MOLLY™, manufactured by¬ Scandinavian Health Ltd.
[0203] In some embodiments, a composition can be formulated for storage at a temperature below 0°C (e.g., -20°C or -80°C). In some embodiments, the composition can be formulated for storage for up to 2 years (c., one month, two months, three months, four months, five months, six months, seven months, eight months, nine months, 10 months, 11 months, 1 year, or 2 years) at 2-8°C (e.g., 4°C). Thus, in some embodiments, the compositions described herein are stable in storage for at least 1 year at 2-8°C (e.g., 4°C).
[0204] A pharmaceutical composition can include a therapeutically effective amount of a ADC described herein. Such effective amounts can be readily determined by one of ordinary skill in the art. A therapeutically effective amount can be an amount at which any toxic or detrimental effects of the composition are outweighed by therapeutically beneficial effects. In some embodiments, a dose can also be chosen to reduce or avoid production of antibodies or other host immune responses against an ADC. Those of skill in the art will appreciate that data obtained from cell culture assays and animal studies can be used in formulating a range of dosage for use in humans. In various embodiments, the amount of active ingredient included in a pharmaceutical composition is such that a suitable dose within the designated range can be administered to subjects. The dose and method of administration can vary depending on weight, age, condition, and other characteristics of a patient, and can be suitably selected as needed by those skilled in the art.
[0205] Pharmaceutical compositions including certain ADCs can be administered as a fixed dose, or in a milligram per kilogram (mg / kg) dose. While in no way intended to be limiting, an exemplary single dose of certain pharmaceutical compositions described herein can include certain ADCs as described herein in an amount equal to, e.g., 0.001 to 1000 mg / kg, 1-1000 mg / kg, 1-100 mg / kg, 0.5-50 mg / kg, 0 1-100 mg / kg, 0,5-25 mg / kg, 1-20 mg / kg, and 1-10 mg / kg bodyweight. Exemplary dosages of a composition described herein include, without limitation, 0.1 mg / kg, 0.5 mg / kg, 1 mg / kg, 2 mg / kg, 4 mg / kg, 8 mg / kg, or 20 mg / kg. The present disclosure is not limited to such ranges or dosages.1340408 v 1 Page 71 of 258Attorney Docket: 2014191-0049
[0206] The present disclosure further includes methods of preparing pharmaceutical compositions of the present disclosure and kits including pharmaceutical compositions of the present disclosure.
[0207] In various embodiments, ADCs of the present disclosure can be administered to a subject in a course of treatment that further includes administration of one or more additional therapeutic agents or therapies that are not ADCs (e.g., surgery or radiation). Combination therapies of the present disclosure can include simultaneous exposure of a subject to therapeutic agents of two or more therapeutic regimens.
[0208] In certain embodiments, an ADC as described herein can be administered together with (e.g., at the same time and / or in the same composition as) an additional agent or therapy. In certain embodiments, an ADC of the present disclosure can be administered separately from an additional therapeutic agent or therapy (e.g, at a different time and / or in a different composition than the additional therapeutic agent or therapy). Dosing regimens of an ADC and one or more additional therapeutic agents with which it is administered in combination can be coordinated or independently determined. In various embodiments, an additional therapeutic agent or therapy administered in combination with an ADC as described herein can be administered at the same time as the ADC, on the same day as the ADC, or in the same week as the ADC. In various embodiments, an additional therapeutic agent or therapy administered in combination with an ADC as described herein can be administered such that administration of the ADC and the additional therapeutic agent or therapy are separated by one or more hours before or after, one or more days before or after, one or more weeks before or after, or one or more months before or after administration of the ADC. In various embodiments, the administration frequency and / or dosage of one or more additional therapeutic agents can be the same as, similar to, or different from the administration frequency of the ADC. In some embodiments, the two or more regimens can be administered simultaneously; in some embodiments, such regimens can be administered sequentially (e.g, all “doses” of a first regimen are administered prior to administration of any doses of a second regimen); in some embodiments, such therapeutic agents are administered in overlapping dosing regimens.
[0209] In certain embodiments, administration of an ADC can be to a subject having previously received, scheduled to receive, or in the course of a treatment regimen including an13404083 v 1 Page 72 of 258Attorney Docket: 2014191-0049additional cancer therapy. Administration of an ADC can, in some instances, improve delivery or efficacy of another therapeutic agent or therapy with which it is administered in combination.
[0210] It is contemplated that therapeutic agent combination therapies can demonstrate synergy and / or greater-than-additive effects between an ADC and one or more additional therapeutic agents with which it is administered in combination. An ADC can be administered in any effective amount as determined independently or as determined by the joint action of the ADC and any of one or more additional therapeutic agents or therapies administered. Administration of the ADC may, in some embodiments, reduce the therapeutically effective dosage, required dosage, or administered dosage of the additional therapeutic agent or therapy relative to a reference regimen for administration of additional therapeutic agent or therapy or therapy absent the ADC. In certain embodiment, a composition described herein can replace or augment other previously or currently administered therapy For example, upon treating with an ADC, administration of one or more additional therapeutic agents or therapies can cease or diminish, e.g., be administered at lower levels.Systems
[0211] In some embodiments, one or more non-transitory computer readable storage media is encoded with a computer program, wherein the program includes instructions that when executed by one or more processors cause the one or more processors to perform operations to perform a method of the present disclosure.
[0212] In some embodiments, a computer system includes a memory and one or more processors coupled to the memory, wherein the one or more processors are configured to perform a method of the present disclosure.
[0213] In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds a histone modification selected from H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, or pan acetylation. In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds H3K4me3 modifications. In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds H3K27ac modifications.
[0214] In some embodiments, a system includes reagents for isolation of cfDNA from a liquid biopsy sample. In some embodiments, the sequencer includes reagents for library1340408 v 1 Page 73 of 258Attorney Docket: 2014191-0049preparation for sequencing In some embodiments, the sequencer includes reagents for sequencing.
[0215] Illustrative embodiments of systems and methods disclosed herein are described with reference to determinations and / or models that may be performed or used by a computing device. That is, in some embodiments, methods disclosed herein are computer-implemented methods and, in some embodiments, a system as disclosed herein includes a processor and one or more non-transitory computer readable storage media (e.g., one or more memories) that have instructions stored thereon that, when executed by the processor, cause the processor to perform operations that include a method disclosed herein For example, a cTF estimation model may be stored on a memory. Such a cTF estimation model may be utilized by a processor to perform a method (e.g., a diagnostic method). Methods of the present disclosure, or portions thereof, may be performed using a processor. The processor may be a part of a computing device and / or computing system.
[0216] Systems of the present disclosure may include a processor and / or a memory. The memory may store one or more programs that include instructions that when executed by a processor cause at least a portion of a method disclosed herein to be performed. The system may further include a machine-learned model. Additionally or alternatively, a remotely stored and / or operated machine-learned model may be accessed by a (e.g,, the) processor. The processor and / or memory may be a part of a computing device and / or computing system.
[0217] One or non-transitory computer readable media may store one or more programs that include instructions that when executed by a (e.g., the) processor cause at least a portion of a method disclosed herein to be performed.
[0218] Methods and systems disclosed herein may utilize one or more models that are machine-learned models. A machine-learned model may be or include an artificial neural network. A machine-learned model may employ, for example, an attention-based model (e.g., a transformer model, such as, for example, a vision transformer), a transformer model (e.g., a vision transformer), a regression-based model (e.g., a logistic regression model), a regularization-based model (e.g., an elastic net model or a ridge regression model), an instance-based model (e.g., a support vector machine or a k-nearest neighbor model), a Bayesian-based model (e.g., a naive¬ based model or a Gaussian naive-based model), a clustering-based model (e g., an expectation maximization model), an ensemble-based model (e.g., an adaptive boosting model, a random forest1340408 v 1 Page 74 of 258Attorney Docket: 2014191-0049model, a bootstrap-aggregation model, or a gradient boosting machine model), or a neural-network-based model (e.g., a convolutional neural network, a recurrent neural network, autoencoder, a back propagation network, or a stochastic gradient descent network).
[0219] In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a k nearest neighbors methodology, a generalized regression forward selection methodology, a generalized regression pruned forward selection methodology, a fit stepwise methodology, a generalized regression lasso methodology, a generalized regression elastic net methodology, a generalized regression ridge methodology, a nominal logistic methodology, a support vector machines methodology, a discriminant methodology, a naive Bayes methodology, or a combination thereof. In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a generalized regression lasso methodology, a generalized regression elastic net methodology, a generalized regression ridge methodology, a nominal logistic methodology, a support vector machines methodology, a discriminant methodology, or a combination thereof. In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a support vector machines methodology, or a combination thereof.
[0220] Certain embodiments described herein make use of computer algorithms in the form of software instructions executed by a computer processor. In certain embodiments, the software instructions include a machine learned module. A machine learned module refers to a computer implemented process (e.g., a software function) that implements one or more specific machine-learned models, such as or including an artificial neural network (ANN), a convolutional neural network (CNN), random forest, one or more decision trees, one or more support vector machines, or a combination thereof, in order to determine, for a given input (e.g., one or more inputs), one or more output values. In certain embodiments, the input includes an image. In certain embodiments, the input includes numerical data, tagged data, and / or functional relationships. In certain embodiments, the input includes alphanumeric data which can include numbers, words, phrases, or lengthier strings, for example. In certain embodiments, the one or more output values include values representing numeric values, words, phrases, or other alphanumeric strings.1340408 v 1 Page 75 of 258Attorney Docket: 2014191-0049
[0221] In embodiments, a machine-learned model has been trained using supervised learning algorithm(s), unsupervised learning algorithm(s), semi-supervised learning algorithm(s) (e.g., partial supervision), weak supervision, transfer, multi-task learning, or any combination thereof. In embodiments, a machine-learned model employs a model that includes parameters (e.g., weights) that are tuned during training of the model. For example, the parameters may be adjusted to minimize a loss function, thereby improving the predictive capacity of the machine learning model. A machine-learned model may be further trained after an initial training period, for example, may be adapted to continuously train as it is used.
[0222] In certain embodiments, a machine learned model has been trained, for example, using datasets that include categories of data described herein. Such training may be used to determine various parameters of a machine learned model, for example implemented by a machine learning module, such as, for example, weights associated with layers in neural networks. In certain embodiments, once a machine learned model has been trained, e.g., to accomplish a specific task such as identifying certain output (e.g., extracting certain feature vector(s)), values of determined parameters are fixed and the (e.g., unchanging, static) machine learned model is used to process new data (e.g., different from the training data) and accomplish its trained task without further updates to its parameters (e.g., the machine learned model does not receive feedback and / or updates). In certain embodiments, a machine learned model may receive feedback, e.g., based on user review of accuracy, and such feedback may be used as additional training data, to dynamically update the machine learned model. In certain embodiments, two or more machine learning models may be combined and implemented as a single model, in a single module, and / or in a single software application. In certain embodiments, two or more machine learned models may also be implemented separately, e.g., as separate software applications. In certain embodiments, two or more machine learned modules may also be implemented separately, e.g., as separate software applications. A machine learned model may be or include software and / or hardware. For example, a machine learned model may be implemented entirely as software, or certain functions of an ANN module (e.g., CNN) may be carried out via specialized hardware (e.g, via an application specific integrated circuit (ASIC)). A machine learned module may be or include software and / or hardware. For example, a machine learned module may be implemented entirely as software, or certain functions of an ANN module (e.g., CNN) may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC)).1340408 v 1 Page 76 of 258Attorney Docket: 2014191-0049
[0223] In certain embodiments, machine learning modules implementing machine learning techniques may be composed of individual nodes (e.g. units, neurons). A node may receive a set of inputs that may include at least a portion of a given input data for the machine learning module and / or at least one output of another node. A node may have at least one parameter to apply and / or a set of instructions to perform (e.g., mathematical functions to execute) over the set of inputs. In certain embodiments, node instructions may include a step to provide various relative importance to the set of inputs using various parameters, such as weights. The weights may be applied by performing scalar multiplication (e.g., or other mathematical function) between a set of inputs values and the parameters, resulting in a set of weighted inputs In certain embodiments, a node may have a transfer function to combine the set of weighted inputs into one output value. A transfer function may be implemented by a summation of all the weighted inputs and the addition of an offset (e.g., bias) value. In certain embodiments, a node may have an activation function to introduce non-linearity into the output value. Non-limiting examples of the activation function include Rectified Linear Activation (ReLu), logistic (e.g., sigmoid), hyperbolic tangent (tanh), and softmax. In certain embodiments, a node may have a capability of remembering previous states (e g, recurrent nodes). Previous states may be applied to the input and output values using a set of learning parameters.
[0224] In certain embodiments, the machine learning module includes a deep learning architecture composed of nodes organized into layers. For example, a layer is a set of nodes that receives data input (e.g., weighted or non-weighted input), transforms it (e.g., by carrying out instructions, e.g., applying a set of functions e.g., linear and / or non-linear functions), and passes transformed values as output (e.g., to the next layer). In certain embodiments, the set of nodes in a particular layer may share the same parameters and instructions without interacting with each other. A machine learning module may be composed of at least one layer (e.g., ordered). Examples of types of layers include convolutional layers (e g., layers with a kernel, a matrix of parameters that is slid across an input to be multiplied with multiple input values to reduce them to a single output value); fully connected (FC) layers (e.g. all nodes are connected to all outputs of the previous layer); recurrent layers, long / short term memory (LSTM) layers, gated recurrent unit (GRU) layers (e g., nodes with the various abilities to memorize and apply their previous inputs and / or outputs); batch normalization (BN) layers (e.g., layers that normalize a set of outputs from another layer, allowing for more independent learning of individual layers); activation layers (e.g.,13404083 v 1 Page 77 of 258Attorney Docket: 2014191-0049layers with nodes that only contain an activation function); and / or (un)pooling layers [e.g., layers that reduce (increase) dimensions of an input by summarizing (splitting) input values in defined patches).
[0225] In certain embodiments, the performance of a machine learning module may be characterized by its ability to produce an output data with specific accuracy. To achieve specific accuracy, a training process is performed to find optimal parameters, such as weights, for each node in each layer of the machine learning module. In certain embodiments, the training process of a machine learning module may involve using output data to calculate an objective function (e.g, cost function, loss function, error function) that needs to be optimized (e g, minimized, maximized). For example, a machine learning objective function may be a combination of a loss function and regularization parameter. The loss function is related to how well the output is able to predict the input. The loss function may take various forms, like mean squared error, mean absolute error, binary cross-entropy, categorical cross-entropy, for example. The regularization term may be needed to prevent overfitting and improve generalization of the training process Examples of regularization techniques include LI Regularization or Lasso Regression, L2 Regularization or Ridge Regression, and Dropout (e g., dropping layer outputs at random during training process).
[0226] In certain embodiments, objective function optimization of a machine learning module may involve finding at least one (e.g., all) of the present global optima (e.g., as opposed to local optima). In certain embodiments, the algorithm for objective function optimization follows principles of mathematical optimization for a multi-variable function and relies on achieving specific accuracy of the process. Examples of objective function optimization algorithms include gradient descent, nonlinear conjugate gradient, random search, Levenberg-Marquardt algorithm, limited-memory Broyden-Fletcher-Goldfarb-Shanno algorithm, pattern search, basin hopping method, Krylov method, Adam method, genetic algorithm, particle swarm optimization, surrogate optimization, and simulated annealing.
[0227] Computations may be performed locally by a computing device. Computations performed over a network are also contemplated. FIG. 2 shows an illustrative network environment 200 for use in the methods and systems described herein. In brief overview, referring now to FIG. 2, a block diagram of an illustrative cloud computing environment 200 is shown and described. The cloud computing environment 200 may include one or more resource providers1340408 v 1 Page 78 of 258Attorney Docket: 2014191-0049202a, 202b, 202c (collectively, 202). Each resource provider 202 may include computing resources. In some implementations, computing resources may include any hardware and / or software used to process data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, illustrative computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 202 may be connected to any other resource provider 202 in the cloud computing environment 200. In some implementations, the resource providers 202 may be connected over a computer network 208. Each resource provider 202 may be connected to one or more computing device 204a, 204b, 204c (collectively, 204), over the computer network 208.
[0228] The cloud computing environment 200 may include a resource manager 206. The resource manager 206 may be connected to the resource providers 202 and the computing devices 204 over the computer network 208. In some implementations, the resource manager 206 may facilitate the provision of computing resources by one or more resource providers 202 to one or more computing devices 204. The resource manager 206 may receive a request for a computing resource from a particular computing device 204, The resource manager 206 may identify one or more resource providers 202 capable of providing the computing resource requested by the computing device 204. The resource manager 206 may select a resource provider 202 to provide the computing resource. The resource manager 206 may facilitate a connection between the resource provider 202 and a particular computing device 204. In some implementations, the resource manager 206 may establish a connection between a particular resource provider 202 and a particular computing device 204. In some implementations, the resource manager 206 may redirect a particular computing device 204 to a particular resource provider 202 with the requested computing resource.
[0229] FIG, 3 shows an example of a computing device 300 and a mobile computing device 350 that can be used in the methods and systems described in this disclosure. The computing device 300 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 350 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other1340408 v 1 Page 79 of 258Attorney Docket: 2014191-0049similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting,
[0230] The computing device 300 includes a processor 302, a memory 304, a storage device 306, a high-speed interface 308 connecting to the memory 304 and multiple high-speed expansion ports 310, and a low-speed interface 312 connecting to a low-speed expansion port 314 and the storage device 306. Each of the processor 302, the memory 304, the storage device 306, the high-speed interface 308, the high-speed expansion ports 310, and the low-speed interface 312, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 302 can process instructions for execution within the computing device 300, including instructions stored in the memory 304 or on the storage device 306 to display graphical information for a GUI on an external input / output device, such as a display 316 coupled to the high-speed interface 308. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Thus, as the term is used herein, where a plurality of functions are described as being performed by “a processor”, this encompasses embodiments wherein the plurality of functions are performed by any number of processors (e.g., one or more processors) of any number of computing devices (e g., one or more computing devices). Furthermore, where a function is described as being performed by "‘a processor”, this encompasses embodiments wherein the function is performed by any number of processors (e.g., one or more processors) of any number of computing devices (e.g., one or more computing devices) (e.g., in a distributed computing system).
[0231] The memory 304 stores information within the computing device 300. In some implementations, the memory 304 is a volatile memory unit or units. In some implementations, the memory 304 is a non-volatile memory unit or units. The memory 304 may also be another form of computer-readable medium, such as a magnetic or optical disk.
[0232] The storage device 306 is capable of providing mass storage for the computing device 300. In some implementations, the storage device 306 may be or contain a computer-readable medium, such as a hard disk device, an optical disk device, a flash memory or other1340408 v 1 Page 80 of 258Attorney Docket: 2014191-0049similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 302), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices such as computer- or machine-readable mediums (for example, the memory 304, the storage device 306, or memory on the processor 302).
[0233] The high-speed interface 308 manages bandwidth-intensive operations for the computing device 300, while the low-speed interface 312 manages lower bandwidth-intensive operations. Such allocation of functions is an example only In some implementations, the highspeed interface 308 is coupled to the memory 304, the display 316 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 310, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 312 is coupled to the storage device 306 and the low-speed expansion port 314. The low-speed expansion port 314, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
[0234] The computing device 300 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 320, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 322. It may also be implemented as part of a rack server system 324. Alternatively, components from the computing device 300 may be combined with other components in a mobile device (not shown), such as a mobile computing device 350. Each of such devices may contain one or more of the computing device 300 and the mobile computing device 350, and an entire system may be made up of multiple computing devices communicating with each other.
[0235] The mobile computing device 350 includes a processor 352, a memory 364, an input / output device such as a display 354, a communication interface 366, and a transceiver 368, among other components. The mobile computing device 350 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 352, the memory 364, the display 354, the communication interface 366, and the transceiver 368,13404083 v 1 Page 81 of 258Attorney Docket: 2014191-0049are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.
[0236] The processor 352 can execute instructions within the mobile computing device 350, including instructions stored in the memory 364. The processor 352 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 352 may provide, for example, for coordination of the other components of the mobile computing device 350, such as control of user interfaces, applications run by the mobile computing device 350, and wireless communication by the mobile computing device 350.
[0237] The processor 352 may communicate with a user through a control interface 358 and a display interface 356 coupled to the display 354. The display 354 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology The display interface 356 may include appropriate circuitry for driving the display 354 to present graphical and other information to a user. The control interface 358 may receive commands from a user and convert them for submission to the processor 352. In addition, an external interface 362 may provide communication with the processor 352, so as to enable near area communication of the mobile computing device 350 with other devices. The external interface 362 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.
[0238] The memory 364 stores information within the mobile computing device 350. The memory 364 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 374 may also be provided and connected to the mobile computing device 350 through an expansion interface 372, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 374 may provide extra storage space for the mobile computing device 350, or may also store applications or other information for the mobile computing device 350 Specifically, the expansion memory 374 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 374 may be provided as a security module for the mobile computing device 350, and may be programmed with instructions that permit secure use of the mobile computing device 350. In addition, secure applications may be provided via the SIMM cards, along with1340408 v 1 Page 82 of 258Attorney Docket: 2014191-0049additional information, such as placing identifying information on the SIMM card in a non-hackable manner.
[0239] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier and, when executed by one or more processing devices (for example, processor 352), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 364, the expansion memory 374, or memory on the processor 352). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 368 or the external interface 362.
[0240] The mobile computing device 350 may communicate wirelessly through the communication interface 366, which may include digital signal processing circuitry where necessary. The communication interface 366 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 368 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth®, Wi-Fi™, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 370 may provide additional navigation- and location-related wireless data to the mobile computing device 350, which may be used as appropriate by applications running on the mobile computing device 350.
[0241] The mobile computing device 350 may also communicate audibly using an audio codec 360, which may receive spoken information from a user and convert it to usable digital information. The audio codec 360 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 350. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 350.1340408 v 1 Page 83 of 258Attorney Docket: 2014191-0049
[0242] The mobile computing device 350 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 380. It may also be implemented as part of a smart-phone 382, personal digital assistant, or other similar mobile device.
[0243] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0244] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0245] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g, a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0246] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware1340408 v 1 Page 84 of 258Attorney Docket: 2014191-0049component (e.g., an application server), or that includes a front end component (e g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0247] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.Certain Exemplary Embodiments
[0248] Without limitation to the foregoing description, the following is an enumerated list of non-limiting exemplary embodiments included in the present disclosure, written in two modules. Those of ordinary skill in the art will appreciate that one or more features discussed above may be included with or incorporated into any of the following numbered embodiments in the First Module and / or Second Module to form additional embodiments. Furthermore, one or more features of any embodiment of the First Module may be combined with one or more features of any embodiment of the Second Module to form a further embodiment, and vice versa.First Module1. A method of making a preliminary prediction of expression of a genomic target for an indication (e.g., at a particular circulating tumor DNA (ctDNA) fraction) and / or producing a model that predicts expression level of a target gene, the method comprising:receiving digital samples for an indication each comprising (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene;tiling the samples into tiles that together span a genomic region corresponding to the target gene;performing a loop for each of at least one subset of the samples, the loop comprising:13404083 v 1 Page 85 of 258Attorney Docket: 2014191-0049determining a set of expression-1 eve! correlated tiles for the subset of the samples, wherein the determining of the set comprises testing each of the tiles for correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level across the subset of the samples,producing a model based on the signal for the one or more epigenetic biomarkers for the set of expression-level correlated tiles and the expression level for the target gene for each sample of the subset of the samples, and predicting, using the model, an expression level for the target gene for one or more of the samples not included in the subset of the samples.2. The method of embodiment 1, wherein producing the model comprises, for each of the one or more epigenetic biomarkers, determining a point estimate for the signal for the epigenetic biomarker across the set of expression-level correlated tiles and producing the model based on the point estimate.3. The method of embodiment 2, wherein the point estimate is a mean.4. The method of embodiment 3, wherein the mean is a geometric mean.5. The method of any one of embodiments 1-4, wherein the model is produced based on a respective point estimate of the signal for each of the one or more epigenetic biomarkers for the set of expression-level correlation tiles.6. The method of any one of embodiments 1 -5, wherein producing the model comprises determining a plurality of point estimates for the signal for at least one of the one or more epigenetic biomarkers across the set of expression-level correlated tiles and the model is produced based on the point estimate.7. The method of embodiment 6, wherein the plurality of point estimates comprises a first point estimate for positively associated tiles for one of the one or more epigenetic biomarkers13404083vl Page 86 of 258Attorney Docket: 2014191-0049and a second point estimate for negatively associated tiles for the one of the one or more epigenetic biomarkers.8. The method of any one of embodiments 1-7, wherein determining the set of expression-level correlated tiles comprises determining signal for at least one of the one or more epigenetic biomarkers for one or more of the tiles is uncorrelated with the expression level across the subset of the samples and excluding the one or more of the tiles from the set of expression-level correlated tiles for the at least one of the one or more epigenetic biomarkers (e.g., excluding the one or more of the tiles from the set of expression-level correlated tiles entirely)9. The method of any one of embodiments 1-8, wherein determining the set of expressionlevel correlated tiles comprises determining, for one or more of the tiles, that a correlation between signal for at least one of the one or more epigenetic biomarkers with the expression level exceeds a threshold for a correlation measure (e g., has a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).10. The method of embodiment 9, wherein the one or more of the tiles are determined to be in the set of expression-level correlated tiles further based on the signal for at least one of the one or more epigenetic biomarkers in healthy volunteer samples (e.g., used to generate the subset of the samples) being below a threshold11. The method of any one of embodiments 1-10, wherein determining the set of express! on- level correlated tiles comprises: (i) determining overlapping ones of the tiles have a correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level, and (ii) collapsing the overlapping ones of the tiles such that the set of expression-level correlated tiles are mutually non-overlapping (e.g., and the set of expression-level correlated tiles comprises tiles having non-uniform size).12. The method of any one of embodiments 1-11, wherein the model is a regression (e.g., a multiple regression) based model [e.g., is a regression (e.g., a multiple regression)].1340408 v 1 Page 87 of 258Attorney Docket: 2014191-004913 The method of embodiment 12, wherein the regression is an ordinary least squares (OLS) regression (e.g., an OLS multiple regression).14 The method of any one of embodiments 1-13, wherein all of the samples in the subset of the samples correspond to a particular ctDNA fraction.15. The method of any one of embodiments 1-14, comprising performing a plurality of iterations of the loop each using a different subset of the samples where all of the samples in the different subset correspond to a particular ctDNA fraction (e g., and the one or more of the samples not included in the subset of the samples is a sample from a subset of the samples used in a different one of the plurality of iterations).16. The method of embodiment 15, comprising performing the loop for each of at least one subset of the samples for each of a set of ctDNA fractions (e.g., performing a cross validated loop using the samples for each of a set of ctDNA fractions).17. The method of embodiment 16, wherein the set of ctDNA fractions corresponds to an expected range of ctDNA fractions for the indication18. The method of any one of embodiments 1-17, wherein all of the samples in the at least one subset of the samples correspond to a particular ctDNA fraction.19. The method of any one of embodiments 1-18, wherein the at least one subset of the samples is a plurality of subsets of the samples [e.g, corresponding to folds in a cross-validation loop (e.g., for at least one ctDNA fraction)].20. The method of embodiment 18, comprising determining whether a relationship (e.g., correlation) between the predicted expression level for the target gene and the expression level for the one or more of the samples not included in the subset of the samples across the plurality of subsets of the samples are within one or more predefined criteria (e.g., at a ctDNA fraction of no more than 10%, no more than 8%, no more than 6%, or no more than 5%).13404083 v 1 Page 88 of 258Attorney Docket: 2014191-004921. The method of embodiment 20, wherein the one or more predefined criteria comprises an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%) and / or an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9)22. The method of embodiment 20 or embodiment 21, comprising determining that the relationship is within the one or more predefined criteria and, responsive to that determination, (e.g., manually) tuning (e.g., feature engineering) the model.23. The method of embodiment 22, wherein tuning the model comprises selecting a subset of the one or more of the expression-level correlated tiles and the method comprises tuning the model using the subset of the one or more of the expression-level correlated tiles.24. The method of embodiment 22 or embodiment 23, wherein the model is for an initial ctDNA fraction and tuning the model comprises tuning the model to produce predictions at a ctDNA fraction lower than the initial ctDNA fraction.25. The method of any one of embodiments 20-24, comprising determining that the relationship is within the one or more predefined criteria and, responsive to that determination, performing the loop with each of at least one subset of the samples that correspond to a. lower ctDNA fraction.26. The method of any one of embodiments 1-25, wherein the one or more of the samples not included in the subset of the samples is exactly one of the digital samples.27. The method of any one of embodiments 1-26, wherein the expression level is an mRNA expression level.28. The method of any one of embodiments 1-26, wherein the expression level is a protein expression level.13404083 v 1 Page 89 of 258Attorney Docket: 2014191-004929 The method of any one of embodiments 1 -28, wherein the one or more epigenetic biomarkers is a plurality of epigenetic biomarkers,30 The method of any one of embodiments 1-29, wherein the one or more epigenetic biomarkers comprises one or more histone modifications.31. The method of any one of embodiments I -30, wherein the one or more epigenetic biomarkers comprises DNA methylation.32. The method of any one of embodiments 1-31, wherein the one or more epigenetic biomarkers comprises H3K27ac modification, H3K4me3 modification, and DNA methylation.33. The method of any one of embodiments 1-32, wherein the genomic region comprises a region corresponding to a transcript for the target gene.34. The method of any one of embodiments 1-33, wherein the genomic region corresponds to an exon for the target gene.35. The method of embodiment 33 or embodiment 34, wherein the genomic region comprises initial and ending buffer regions (e.g., of ±200 kb) (e.g., around the exon) (e.g., around the transcript).36. The method of any one of embodiments 1-35, wherein the tiles are overlapping (e g., wherein no more than two of the tiles are mutually overlapping).37. The method of any one of embodiments 1-36, wherein adjacent ones of the tiles overlap by at least 10% (e.g., at least 20%, at least 30% of a length of the tiles) and no more than 70% (e.g., no more than 60% or no more than 50% of the length of the tiles).38. The method of any one of embodiments 1-37, wherein the tiles have a length in a range of from 100 to 1000 bp (e.g, from 250 to 750 bp).13404083vl Page 90 of 258Attorney Docket: 2014191-004939. The method of any one of embodiments 1-38, wherein adjacent ones of the tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp).40. The method of any one of embodiments 1-39, comprising generating the digital samples.41. The method of embodiment 40, wherein generating the digital samples comprises, for each of the digital samples, (i) randomly sampling data for at least one healthy sample and at least one cell line sample (e.g., from one or more indication-specific, e.g, cancer, cell lines) in a mixing ratio corresponding to a desired ctDNA fraction for the digital sample and (ii) determining the expression level for the sample.42. The method of embodiment 41, wherein determining the expression level comprises using a previously measured expression level for the at least one cell line sample as the expression level for the sample43. The method of any one of embodiments 1-42, wherein the signal is sequencing counts.44. The method of any one of embodiments 1-43, wherein the signal for each of the one or more epigenetic biomarkers for the samples is in silico mixed sequencing data.45. The method of any one of embodiments 1-44, wherein the digital samples have been generated using at least 25 different cell lines.46. The method of any one of embodiments 1-45, wherein the subset of the samples have been generated using at least 25 different healthy volunteers.47. The method of any one of embodiments 1-46, wherein the subset of the samples have been generated using at least 25 different cell lines.1340408 v 1 Page 91 of 258Attorney Docket: 2014191-004948 The method of any one of embodiments 1 -47, wherein the digital samples have been generated using at least 25 different healthy volunteers49 The method of any one of embodiments 1-48, wherein the digital samples comprise in silico diluted samples.50. The method of any one of embodiments 1 -49, wherein the digital samples comprise in silico diluted plasma samples.51. The method of any one of embodiments 1-50, comprising normalizing the signal for each of the one or more epigenetic biomarkers prior to performing the loop such that the loop is performed using the normalized signal.52. The method of embodiment 51, wherein normalizing the signal comprises a quantile normalization.53. The method of any one of embodiments 1-52, comprising pebbling the signal for each of the one or more epigenetic biomarkers prior to performing the loop such that the loop is performed using the pebbled signal.54. The method of embodiment 53, wherein the pebbling comprises, for each of the one or more epigenetic biomarkers, determining a background signal for the epigenetic biomarker for the genomic region based on signal (e.g., a number of fragments) in the sample in a pebbling region and subtracting the background signal from the signal.55. The method of any one of embodiments 1-54, wherein the target gene is a gene that encodes an antibody drug conjugate (ADC) target antigen.56. The method of any one of embodiments 1-55, wherein the indication is a cancer indication (e.g., a cancer type or subtype).1340408 v 1 Page 92 of 258Attorney Docket: 2014191-004957 The method of embodiment 56, wherein the indication is breast cancer, prostate cancer, or small cell lung cancer (SCLC).58 A method of producing a multigene expression level prediction model, the method comprising:producing a first constituent model for a target gene using a method according to any one of embodiments 1-57;selecting one or more related genes related to the target gene; andproducing a second constituent model for each of the one or more related genes using a method according to any one of embodiments 1-57 (e.g., wherein the related gene is the target gene in the method); andproducing an ensemble model based on the first constituent model and the second constituent model.59. The method of embodiment 58, wherein the ensemble model uses a ridge regression of the first constituent model and the second constituent model.60. The method of any one of embodiments 1-57, comprising selecting one or more related genes related to the target gene, wherein the digital samples comprise signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes such that set of expression-level correlated tiles comprises at least one tile corresponding to each of the one or more related genes.61. The method of any one of embodiments 1-57, comprising selecting one or more related genes related to the target gene, wherein the digital samples comprise signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes, and determining the set of expression-level correlated tiles comprises testing each of the tiles corresponding to the one or more related genes for correlation between the signal in the sample corresponding to the tile for1340408 v 1 Page 93 of 258Attorney Docket: 2014191-0049each of the one or more epigenetic biomarkers and the expression level for the sample across the subset of the samples.62 The method of anv one of embodiments 58-61, wherein selecting the one or more related genes comprises selecting one or more genes that are in a transcriptional complex with the target gene.63. The method of any one of embodiments 58-61, wherein selecting the one or more related genes comprises selecting one or more genes that are master regulators for the indication.64. The method of any one of embodiments 58-61, wherein selecting the one or more related genes comprises selecting one or more genes that are related to a pathway relevant for the indication.65. The method of any one of embodiments 58-64, wherein the one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes.66. The method of any one of embodiments 58-65, wherein selecting the one or more related genes comprises selecting a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].67. The method of any one of embodiments 1-66, wherein:(a) the target gene or genomic target is listed in Table 1, and the indication is prostate cancer;(b) the target gene or genomic target is listed in Table 2, and the indication is breast cancer;(c) the target gene or genomic target is listed in Table 3, and the indication is SCLC; or(d) the target gene or genomic target is listed in Table 4 and the indication is cancer.1340408 v 1 Page 94 of 258Attorney Docket: 2014191-004968 A method of predicting expression level of a liquid biopsy sample for a subject, the method comprising:providing an expression level prediction model that has been produced from digital samples for an indication each comprising (i) signal (e.g., sequencing counts) for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene, wherein the digital samples have been generated using data derived from cell samples specific to the indication and healthy volunteers;providing input data comprising signal (e.g., sequencing counts) for the one or more epigenetic biomarkers derived from a liquid biopsy sample for a subject; andpredicting expression level of the target gene for the subject from the input data using the model,69. The method of embodiment 68, wherein the cell samples comprise tissue samples.70. The method of embodiment 68 or embodiment 69, wherein the cell samples have been derived from one or more cell lines, one or more patient-derived xenografts or a biopsy therefrom, one or more organoids, or a combination thereof.71. The method of any one of embodiments 68-70, wherein the liquid biopsy sample is a plasma sample72. The method of any one of embodiments 68-71, wherein the one or more epigenetic biomarkers comprise one or more histone modifications and / or DNA methylation.73. The method of any one of embodiments 68-72, wherein the one or more epigenetic biomarkers comprises H3K27ac modification, H3K4me3 modification, and DNA methylation.74. The method of any one of embodiments 68-73, wherein the model has been produced using a method according to any one of embodiments 1-57.1340408 v 1 Page 95 of 258Attorney Docket: 2014191-004975 A method of predicting (e.g., determining) the expression of a genomic target for an indication (e.g., at a particular circulating tumor DNA (ctDNA) fraction), the method comprising quantifying one or more epigenetic biomarkers at one or more expression-level correlated loci for the genomic target.76. The method of embodiment 75, where the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci has been shown to be correlated with expression of the genomic target (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).77. The method of embodiment 75 or 76, wherein the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci in healthy volunteer samples is below a threshold.78. The method of any one of embodiments 75-77, wherein the one or more expression-level correlated loci can be or are determined using a method recited in any one of embodi ments 1-109.79. The method of any one of embodiments 75-78, wherein the expression level is an mRNA expression level.80. The method of any one of embodiments 75-79, wherein:(a) the genomic target is listed in Table 1, and the indication is prostate cancer;(b) the genomic target is listed in Table 2, and the indication is breast cancer;(c) the genomic target is listed in Table 3, and the indication is SCLC; or(d) the genomic target is listed in Table 4 and the indication is cancer.81. A method of determining the ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:1340408 v 1 Page 96 of 258Attorney Docket: 2014191-0049(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andthe one or more genomic loci comprise one or more expression-level correlated loci for ESRI that are provided in Table 5.82. A method of determining the ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andwherein the one or more expression-level correlated loci for the genomic target include one or more expression-level correlated loci for ESRI, EN01, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM, CCDC170, or any combination thereof.83. The method of embodiment 82, wherein:(i) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with ESR1 expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5); or(ii) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with ESR1, ENO1, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM, or CCDC170 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).84. The method of any one of embodiment 82 or embodiment 83, wherein the one or more genomic loci comprise one or more of the genomic loci listed in Table 5.1340408 v 1 Page 97 of 258Attorney Docket: 2014191-004985 A method of determining the PR status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii ) binding of one or more transcription factors, and / or(iv) DNA methylation;wherein the one or more expression-level correlated loci for the genomic target include one or more expression-level correlated loci for PGR1, SCUBE2, SERPINA11, CA12, ABAT, MAPT-IT1, MAPT, GREB1, PTPRT, PREX1, NEK10, LRIG1, TPRG1, SPEF2, RGS22, NXNL2, FGD3, SUSD3, or GRPR, or any combination thereof86. The method of embodiment 85, wherein:(i) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with PGR1 expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5); or(ii) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with SCUBE2, SERPINA11, CA12, ABAT, MAPT-IT1, MAPT, GREB1, PTPRT, PREX1, NEK10, LRIG1, TPRG1, SPEF2, RGS22, NXNL2, FGD3, SUSD3, GRPR, or PGR1 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).87. The method of embodiment 85 or embodiment 86, wherein the one or more expression¬ level correlated loci comprise one or more loci listed in Table 6.88. A method of determining the HER2 status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:13404083 vl Page 98 of 258Attorney Docket: 2014191-0049(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andthe one or more genomic loci comprise one or more of the genomic loci provided in Table 7.89. The method of any one of embodiments 75-88, wherein the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci in healthy volunteer samples is below a threshold.90. The method of any one of embodiments 75-89, wherein the one or more expression-level correlated loci can be or are determined using a method recited in any one of embodiments 1- 109.91. The method of any one of embodiments 75-90, wherein the expression level is an mRNA expression level.92. A method of predicting (e.g., determining) expression of a genomic target for an indication (e.g., at a particular circulating tumor DNA (ctDNA) fraction), the method comprising quantifying one or more epigenetic biomarkers at one or more expression-level correlated loci for the genomic target.93. The method of embodiment 92, where the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci have been shown to be correlated with expression of the genomic target (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).94. The method of embodiment 92 or embodiment 93, wherein the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci in healthy volunteer samples is below a threshold.1340408 v 1 Page 99 of 258Attorney Docket: 2014191-004995. The method of any one of embodiments 75-94, wherein the one or more expression-level correlated loci can be or are determined using a method recited in any one of embodiments 1-109.96. The method of any one of embodiments 75-95, wherein the expression level is an mRNA expression level.97. A method of producing a model that predicts expression level of a target gene, the method comprising:receiving digital samples for an indication each comprising (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene;producing a set of constituent models based on the digital samples, wherein the set comprises one constituent model for each of the target gene and the one or more related genes; andproducing an ensemble model based on a combination (e.g., using a ridge regression) of the constituent models in the set.98. A method of producing a model that predicts expression level of a target gene, the method comprising:receiving digital samples for an indication each comprising (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene;determining genomic regions corresponding to the target gene and the one or more related genes where the signal for at least one of the one or more epigenetic biomarkers correlates with the expression level for the samples (e.g., determine expression-level correlated tiles corresponding to the target gene and each of the one or more related genes); and1340408 v 1 Page 100 of 258Attorney Docket: 2014191-0049producing a model based on the signal for the one or more epigenetic biomarkers for the genomic regions (e.g., for the expression-level correlated tiles) and the expression level.99 The method of embodiment 97 or embodiment 98, wherein the digital samples have been generated using data derived from cell samples specific to the indication and healthy volunteers.100. The method of any one of embodiments 97-99, wherein the cell samples comprise tissue samples.101. The method of any one of embodiments 97-100, wherein the cell samples have been derived from one or more cell lines, one or more patient-derived xenografts or a biopsy therefrom, one or more organoids, or a combination thereof.102. The method of any one of embodiments 97-101, wherein the liquid biopsy sample is a plasma sample.103. The method of any one of embodiments 97-102, wherein the one or more epigenetic biomarkers comprise one or more histone modifications and / or DNA methylation.104. The method of any one of embodiments 97-103, wherein the one or more epigenetic biomarkers comprises H3K27ac modification, H3K4me3 modification, and DNA methylation.105. The method of any one of embodiments 97-104, wherein the signal for the digital samples and for the subject is sequencing counts.106. The method of any one of embodiments 97-105, wherein the model has been produced using a method according to any one of embodiments 1-104,107. The method of any one of embodiments 1-106, wherein the preliminary prediction of expression of a genomic target comprises predicting expression status of a target gene (e.g..1340408 v 1 Page 101 of 258Attorney Docket: 2014191-0049presence or absence of expression of a target gene, or expression of a target gene above a threshold level).108. The method of embodiment 107, wherein the genomic target is ERBB2, and the preliminary prediction of expression is a preliminary prediction of human epidermal growth factor receptor 2 (HER2) expression status (e.g., HER2 IHC 3+ / 2+ISH+ vs. HER2 2+ / 1+ / 0).109. The method of embodiment 107, wherein the genomic target is ESRI, and the preliminary prediction of expression is a preliminary prediction of estrogen receptor (ER) expression status (e.g., as determined using IHC).110. The method of embodiment 107, wherein the genomic target is PGR1, and the preliminary' prediction of expression is a preliminary prediction of progesterone receptor (PR) expression status (e.g., as determined using IHC).111. A method of predicting an expression level of a target gene in a subject, the method comprising:providing sample data for a subject, wherein the sample data comprise signal for one or more epigenetic biomarkers for a target gene (e.g., and optionally one or more related genes);providing a model that has been produced using a method according to any one of embodiments 1-110; andpredicting an expression level of the target gene for the subject using the sample data with the model.112. The method of embodiment 111, wherein the sample data have been derived from a liquid biopsy sample (e.g., from plasma) for the subject.113. A method of monitoring a subject, the method comprising:providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time;13404083vl Page 102 of 258Attorney Docket: 2014191-0049predicting a first expression level for a target gene using the signal for the first sample and a second expression level for the target gene using the signal for the second sample using a model that has been produced using a method according to any one of embodiments 1-109, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g., an ADC therapy) corresponding to the target gene, anddetermining a difference in expression level between the first sample and the second sample.114 The method of embodiment 113, wherein the change is a reduction in expression level and the therapy is a degrader for the target gene.115. A method of monitoring a subject, the method comprising:providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time;predicting a first expression level for a set of related genes using the signal for the first sample and a second expression level for the related gene using the signal for the second sample using a model that has been produced according to a method according to any one of embodiments 1-109, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e g., an ADC therapy) corresponding to at least one of the related genes; anddetermining a difference in expression level between the first sample and the second sample.116. The method of embodiment 115, wherein the related genes correspond to a pathway.117 A method of characterizing cancer recurrence and / or progression, the method comprising:predicting an expression level of a target gene for a subject having a cancer based on a first sample for the subject taken at a first time point using a model that has been produced using a method according to any one of embodiments 1-110;13404083vl Page 103 of 258Attorney Docket: 2014191-0049predicting an expression level of a target gene for the subject based on a second sample for the subject taken at a second time point after the first time point using a method according to any one of embodiments 1-110; anddetermining a difference in the expression level at the second time point and at the first time point.118. A method of monitoring cancer in a subject, the method comprising:predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of embodiments 1-110; anddetermining whether there is a difference in expression level of the target gene for the subject over time.119. A method of determining effectiveness of a therapeutic agent in a subject having cancer, the method comprising:predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of embodiments 1-110, anddetermining whether there is a difference in expression level of the target gene for the subject over time.120. A method of monitoring response of a subject having an indication to a therapeutic agent for the indication, the method comprising:predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of embodiments 1-110; anddetermining whether there is a difference in expression level of the target gene for the subject over time.1340408 v 1 Page 104 of 258Attorney Docket: 2014191-0049121. The method of any one of embodiments 117-120, comprising administering a therapy comprising a therapeutic agent to the subject between when two or more of the samples were obtained.122. The method of any one of embodiments 111-121, comprising administering a therapy to the subject when the difference is determined to be at least as large as a threshold difference.123. The method of any one of embodiments 111-122, comprising altering administration of a therapy to the subject when the difference is determined to be at least as large as a threshold difference.124. The method of embodiment 123, wherein altering administration comprises increasing a dosage and / or frequency of administration.125. A method of prognosing cancer, the method comprising:predicting an expression level of a target gene for a subject having a cancer based on a sample for the subject using a model that has been produced using a method according to any one of embodiments 1-110; andprognosing cancer in the subject based on the determined expression level.126. The method of embodiment 125, comprising administering a therapy based on the prognosis.127. A method of diagnosing cancer in a subject, the method comprising:predicting an expression level of a target gene for a subject having a cancer based on a sample for the subject using a model that has been produced using a method according to any one of embodiments 1-110; anddetermining that the expression level exceeds a threshold.128. The method of embodiment 127, comprising initiating administration of a therapy based on the expression level.1340408 v 1 Page 105 of 258Attorney Docket: 2014191-0049129. The method of embodiment 128, comprising selecting a dosing regimen for the therapy based on the expression level.130. A method of determining whether a cancer has been removed from a subject, the method comprising, after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, predicting expression level of a target gene for the subject based on a sample for the subject using a model that has been produced using a method according to any one of embodiments 1-110.131. The method of embodiment 130, comprising continuing administration of a therapy based on the expression level.132. The method of embodiment 130, comprising ceasing administration of a therapy based on the expression level.133. A system comprising a processor and one or more non-transitory computer readable media having instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising the method according to any one of embodiments 1-132.134. One or more non-transitory' computer readable media having instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising the method according to any one of embodiments 1-132.Second Module1. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data1340408 v 1 Page 106 of 258Attorney Docket: 2014191-0049comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene;determining, by the processor, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles corresponding to the genomic region;aggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal for the epigenetic biomarker; andpredicting (e g., using a model), by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signal for each of the one or more epigenetic biomarkers.2. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene; andpredicting (e.g., using a model), by the processor, an expression level for the target gene for the subject (e.g, having an indication, such as cancer) (e.g., at the time the sample was taken) based on the signal for each of the one or more epigenetic biomarkers.3. The method of embodiment 2, comprising:determining, by the processor, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles corresponding to the genomic region; andaggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal,1340408 v 1 Page 107 of 258Attorney Docket: 2014191-0049wherein predicting the expression level of the target gene based on the signal for each of the one or more epigenetic biomarkers is based on the aggregated tile signal for each of the one or more epigenetic biomarkers.4. The method of embodiment 1 or embodiment 3, comprising tiling, by the processor, the sample sequencing data into tiles that together span the genomic region to obtain tiled sample sequencing data, wherein determining the tile signal is performed using the tiled sample sequencing data.5. The method of any one of embodiments 1-4, wherein the method comprises determining a circulating tumor DNA (ctDNA) fraction for the sample.6. The method of embodiment 5, wherein predicting the expression level is further based on the ctDNA fraction.7. The method of embodiment 5 or 6, wherein determining the ctDNA fraction comprises determining the ctDNA fraction using the sample sequencing data.8. The method of embodiment 7, wherein determining the ctDNA fraction comprises determining the ctDNA fraction based on a ctDNA signal for the one or more epigenetic biomarkers.9. The method of embodiment 8, wherein the ctDNA signal for the one or more epigenetic biomarkers comprises a ctDNA signal from one or more histone modifications.10. The method of embodiment 8, wherein the ctDNA signal for the one or more epigenetic biomarkers comprises a ctDNA signal from DNA methylation.11. The method of any one of embodiments 1-10, wherein predicting the expression level comprises using a model that has been trained on digital samples, each comprising (a) a signal13404083 v 1 Page 108 of 258Attorney Docket: 2014191-0049for the one or more epigenetic biomarkers for the target gene, and (b) an expression level for the target gene12 The method of embodiment 11, wherein the model has been produced by:receiving the digital samples;tiling the digital samples into tiles that together span a genomic region corresponding to the target gene; anddetermining a set of expression-level correlated tiles for the digital samples, wherein the determining of the set comprises testing each of the tiles for correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level across the digital samples.13. The method of embodiment 11 or 12, wherein the model has been produced based on the signal for the one or more epigenetic biomarkers for the set of expression-level correlated tiles and the expression level for the target gene for each sample of the digital samples.14. The method of any one of embodiments 1-13, comprising determining, for each of the one or more epigenetic biomarkers, a point estimate for the tile signal for the epigenetic biomarker across the set of expression-level correlated tiles.15. The method of embodiment 14, w’herein the point estimate is a mean.16. The method of embodiment 15, wherein the mean is a geometric mean.17. The method of any one of embodiments 1 -16, comprising predicting the expression level for the target gene based on a respective point estimate of the signal for each of the one or more epigenetic biomarkers for the set of expression-level correlation tiles.18. The method of any one of embodiments 1-17, comprising inputting the sample sequencing data into a model, wherein the model (i) tiles the sequencing data, (ii) determines the tile signal for each of the one or more epigenetic biomarkers for a set of expression-level1340408 v 1 Page 109 of 258Attorney Docket: 2014191-0049correlated tiles, (iii) aggregates the tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles, and (iv) predicts the expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signal.19. The method of any one of embodiments 1-18, comprising determining, for each of the one or more epigenetic biomarkers, a plurality of point estimates for the signal for the set of expression-level correlated tiles.20. The method of embodiment 19, wherein the plurality of point estimates comprises a first point estimate for positively associated tiles of the set of expression-level correlated tiles for one of the one or more epigenetic biomarkers and a second point estimate for negatively associated tiles of the set of expression-level correlated tiles for the one of the one or more epigenetic biomarkers.21. The method of any one of embodiments 1-20, wherein the model has been trained by a method comprising determining that a signal for at least one of the one or more epigenetic biomarkers for one or more of the tiles is uncorrelated with the expression level, and excluding the one or more of the tiles from the set of expression-level correlated tiles for the at least one of the one or more epigenetic biomarkers (e.g., excluding the one or more of the tiles from the set of expression-level correlated tiles entirely ).22. The method of any one of embodiments 11-21, wherein the set of expression-level correlated tiles comprises tiles where a correlation between signal for at least one of the one or more epigenetic biomarkers with the expression level exceeds a threshold for a correlation measure (e.g., has a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).23. The method of embodiment 22, wherein the tiles are tiles where the signal for at least one of the one or more epigenetic biomarkers in healthy volunteer samples (e g., used to generate the subset of the samples) is below a threshold.13404083vl Page 110 of 258Attorney Docket: 2014191-004924 The method of any one of embodiments 1-10, wherein the tiles in the set of expression-level correlated tiles are mutually non-overlapping (e.g., and the set of expression-level correlated tiles comprises tiles having non-uniform size).25. The method of any one of embodiments 1-24, wherein the model is a regression (e.g., a multiple regression) based model [e.g., is a regression (e.g., a multiple regression)].26. The method of embodiment 25, wherein the regression is an ordinary least squares (OLS) regression (eg., an OLS multiple regression).27. The method of any one of embodiments 11-26, wherein the digital samples comprise a subset of digital samples (e.g, wherein all of the digital samples in the subset of the digital samples) corresponding to a particular ctDNA fraction.28. The method of embodiment 27, wherein the model has been trained, in part, by determining whether a relationship (e.g., correlation) between the predicted expression level for the target gene and the expression level for the one or more of the samples not included in the subset of the samples across the plurality of subsets of the samples are within one or more predefined criteria (e.g., at a ctDNA fraction of no more than 10%, no more than 8%, no more than 6%, or no more than 5%).29. The method of embodiment 28, wherein the one or more predefined criteria comprises one or more predefined criteria for the model, comprising an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%).30. The method of embodiment 28 or 29, wherein the one or more predefined criteria comprises one or more predefined criteria for the model, comprising an area under curve ( AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9).13404083 v 1 Page 111 of 258Attorney Docket: 2014191-004931 The method of embodiment 29 or embodiment 30, wherein the model has been tuned (e.g., manually) responsive to a determination that the relationship is within the one or more predefined criteria.32. The method of embodiment 31, wherein model has been tuned by (i) selecting a subset of the one or more of the expression-level correlated tiles and (ii) using the subset of the one or more of the expression-level correlated tiles.33. The method of embodiment 31 or embodiment 32, wherein the model is for an initial ctDNA fraction, and tuning the model has been tuned to produce predictions at a ctDNA fraction lower than the initial ctDNA fraction.34. The method of any one of embodiments 29-33, wherein the model has been trained, in part, by determining that the relationship is within the one or more predefined criteria and, responsive to that determination, performing a training loop with each of at least one subset of the samples that correspond to a lower ctDNA fraction.35. The method of any one of embodiments 21-34, wherein the one or more of the digital samples excluded from the subset of the samples is exactly one of the digital samples.36. The method of any one of embodiments 1-35, wherein the expression level is an mRNA expression level.37. The method of any one of embodiments 1-36, wherein the expression level is a protein expression level.38. The method of any one of embodiments 1-37, wherein the one or more epigenetic biomarkers is a plurality of epigenetic biomarkers.1340408 v 1 Page 112 of 258Attorney Docket: 2014191-004939 The method of any one of embodiments 1 -38, wherein the one or more epigenetic biomarkers comprises one or more histone modifications (e.g., H3K27ac modification and / or H3K4me3 modification).40. The method of any one of embodiments 1-39, wherein the one or more epigenetic biomarkers comprises DNA methylation.41. The method of any one of embodiments 1-40, wherein the one or more epigenetic biomarkers comprises (i) H3K27ac modification and H3K4me3 modification or (ii) H3K27ac modification, H3K4me3 modification, and DNA methylation.42. The method of any one of embodiments 1-41, wherein the genomic region comprises a region corresponding to a transcript for the target gene.43. The method of any one of embodiments 1-42, wherein the genomic region corresponds to an exon for the target gene,44. The method of embodiment 42 or embodiment 43, wherein the genomic region comprises initial and ending buffer regions (e.g., of ±200 kb) (e.g., around the exon) (e.g., around the transcript).45. The method of any one of embodiments 4-44, wherein the tiles are overlapping (e.g., wherein no more than two of the tiles are mutually overlapping).46. The method of any one of embodiments 4-45, wherein adjacent ones of the tiles overlap by at least 10% (e.g., at least 20%, at least 30% of a length of the tiles) and no more than 70% (e.g, no more than 60% or no more than 50% of the length of the tiles).47. The method of any one of embodiments 4-46, wherein the tiles have a length in a range of from 100 to 1000 bp (e.g., from 250 to 750 bp).1340408 v 1 Page 113 of 258Attorney Docket: 2014191-004948 The method of any one of embodiments 4-47, wherein adjacent ones of the tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp),49 The method of any one of embodiments 11-48, wherein the digital samples have been generated using at least 25 different cell lines.50. The method of any one of embodiments 11-49, wherein the digital samples have been generated using at least 25 different healthy volunteers.51. The method of any one of embodiments 11-50, wherein the digital samples have been generated using at least 25 different cell lines.52. The method of any one of embodiments 11-51, wherein the digital samples have been generated using at least 25 different healthy volunteers.53. The method of any one of embodiments 11-52, wherein the digital samples comprise in silico diluted samples.54. The method of any one of embodiments 11-53, wherein the digital samples comprise in silico diluted plasma samples55. The method of any one of embodiments 1-54, comprising normalizing the signal for each of the one or more epigenetic biomarkers.56. The method of embodiment 55, wherein normalizing the signal comprises a quantile normalization.57. The method of any one of embodiments 1-56, comprising pebbling the signal for each of the one or more epigenetic biomarkers.1340408 v 1 Page 114 of 258Attorney Docket: 2014191-004958 The method of embodiment 57, wherein the pebbling comprises, for each of the one or more epigenetic biomarkers, determining a background signal for the epigenetic biomarker for the genomic region based on a signal (e.g., a number of fragments) in the sample in a pebbling region, and subtracting the background signal from the signal.59. The method of any one of embodiments 1-58, wherein the target gene is a gene that encodes a polypeptide associated with (e g., targeted by) a therapy.60. The method of any one of embodiments 1-59, wherein the target gene is a gene that encodes an antibody drug conjugate (ADC) target antigen.61. The method of any one of embodiments 1 -60, comprising predicting the expression level for an indication.62. The method of embodiment 61, wherein the indication is a cancer indication (e g., a cancer type or subtype).63. The method of embodiment 62, wherein the indication is breast cancer, ovarian cancer, bladder cancer, renal cancer, prostate cancer, pancreatic cancer, or lung cancer (e.g., small cell lung cancer).64. The method of any one of embodiments 1-63, wherein the sample sequencing data comprises a signal for each of the one or more epigenetic biomarkers for one or more genomic regions corresponding to one or more additional genes and the method comprises:determining, by the processor, a tile signal for each of the one or more epigenetic biomarkers for a second set of expression-level correlated tiles corresponding to one or more genomic regions corresponding to one or more additional genes; andaggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal for the epigenetic biomarker,1340408 v 1 Page 115 of 258Attorney Docket: 2014191-0049wherein predicting the expression level for the target gene is further based on the tile signal for each of the one or more epigenetic biomarkers for the second set of expression-level correlated tiles.65. The method of embodiment 64, wherein the one or more additional genes comprise one or more genes related to the target gene.66. The method of embodiment 65, wherein the one or more related genes comprises one or more genes that are in a transcriptional complex with the target gene.67. The method of embodiment 65 or embodiment 66, wherein the one or more related genes comprises one or more genes that are master regulators for the indication.68. The method of any one of embodiments 65-67, wherein the one or more related genes comprises one or more genes that are related to a pathway relevant for the indication.69. The method of any one of embodiments 65-68, wherein the one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes.70. The method of any one of embodiments 69-74, wherein the one or more related genes comprises a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].71. The method of any one of embodiments 64-70, wherein predicting the expression level for the target gene comprises using an ensemble model.72. The method of embodiment 71, wherein the ensemble model comprises a first constituent model and a second constituent model (e.g., comprises a constituent model for the target gene and for each of the one or more additional genes).1340408 v 1 Page 116 of 258Attorney Docket: 2014191-004973 The method of embodiment 72, wherein the first constituent model is a model for the target gene and the second constituent model is a model for one or more of the one or more additional genes.74. The method of embodiment 72 or embodiment 73, wherein the ensemble model uses a ridge regression of the first constituent model and the second constituent model.74. The method of any one of embodiments 72-74, wherein the first constituent model comprises a model for the target gene.75. The method of any one of embodiments 72-74, wherein the second constituent model comprises a single gene model for a gene related to the target gene.76. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for two or more genes (e.g,, for one or more genomic regions corresponding to the two or more genes), wherein the two or more genes comprise a target gene and one or more additional (e g., related) genes;determining, by the processor, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for each of the two or more genes for a set of expression-level correlated tiles (e.g., wherein the method comprises tiling, by the processor, the sample sequencing data into tiles that together span one or more genomic regions corresponding to each of the two or more genes to obtain tiled sample sequencing data and the aggregating is performed using the tiled sample sequencing data):aggregating, by the processor, the tile signal for the one or more epigenetic biomarkers signal for the set of expression-level correlated tiles to obtain aggregated tile signals for each of the two or more genes; and13404083 v 1 Page 117 of 258Attorney Docket: 2014191-0049predicting, by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signals for each of the two or more genes.77. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for two or more genes (e.g., for one or more genomic regions corresponding to the tw o or more genes), wherein the two or more genes comprise a target gene and one or more additional (e.g., related) genes; and predicting, by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the signal for reach of the one or more epigenetic biomarkers for the two or more genes.78. The method of embodiment 76 or embodiment 77, wherein the one or more additional genes comprise one or more genes related to the target gene.79. The method of embodiment 78, wherein the one or more related genes comprises one or more genes that are in a transcriptional complex with the target gene.80. The method of embodiment 78 or embodiment 79, wherein the one or more related genes comprises one or more genes that are master regulators for the indication.81. The method of any one of embodiments 78-80, wherein the one or more related genes comprises one or more genes that are related to a pathway relevant for the indication.82. The method of any one of embodiments 78-81, wherein the one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes.13404083 v 1 Page 118 of 258Attorney Docket: 2014191-004983 The method of any one of embodiments 78-82, wherein the one or more related genes comprises a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].84. The method of any one of embodiments 76-83, wherein predicting the expression level for the target gene comprises using an ensemble model.85. The method of embodiment 84, wherein the ensemble model comprises a first constituent model and a second constituent model (e.g., comprises a constituent model for the target gene and for each of the one or more additional genes).86. The method of embodiment 85, wherein the first constituent model is a model for the target gene and the second constituent model is a model for one or more of the one or more additional genes.87. The method of embodiment 85 or embodiment 86, wherein the ensemble model uses a ridge regression of the first constituent model and the second constituent model.88. The method of any one of embodiments 85-87, wherein the first constituent model comprises a model for the target gene.89. The method of any one of embodiments 85-88, wherein the second constituent model comprises a single gene model for a gene related to the target gene.90. The method of any one of embodiments 1-89, wherein:(a) the target gene or genomic target is listed in Table 1, and the indication is prostate cancer;(b) the target gene or genomic target is listed in Table 2, and the indication is breast cancer;(c) the target gene or genomic target is listed in Table 3, and the indication is lung cancer (e.g., small cell lung cancer); or1340408 v 1 Page 119 of 258Attorney Docket: 2014191-0049(d) the target gene or genomic target is listed in Table 4 and the indication is cancer.91. The method of any one of embodiments 1-90, wherein the expression level of the target gene comprises presence or absence of expression of the target gene.92. The method of any one of embodiments 1-90, wherein the expression level of the target gene comprises expression of the target gene above a threshold level.93. The method of any one of embodiments 1-92, wherein the target gene is ERBB2, and the expression level is human epidermal growth factor receptor 2 (HER2) expression status (e.g., HER2 IHC 3+ / 2+ISH+ vs. HER2 2+ / 1+ / 0).94. The method of any one of embodiments 1-92, wherein the target gene is ESRI, and the expression level is estrogen receptor (ER) expression status.95. The method of any one of embodiments 1 -92, wherein the target gene is PGR1, and the expression level is progesterone receptor (PR) expression status.96. A system comprising a processor and one or more non-transitory computer readable media having instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising the method according to any one of embodiments 1- 95.97. One or more non-transitory computer readable media having instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising the method according to any one of embodiments 1-95.98. A method of monitoring a subject, the method comprising:providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time;1340408 v 1 Page 120 of 258Attorney Docket: 2014191-0049predicting a first expression level for a target gene using the signal for the first sample and a second expression level for the target gene using the signal for the second sample using a method according to any one of embodiments 1-95, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g., an ADC therapy) corresponding to the target gene; anddetermining a difference in expression level between the first sample and the second sample.99. The method of embodiment 98, wherein the change is a reduction in expression level and the therapy is a degrader for the target gene.100. A method of monitoring a subject, the method comprising:providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time;predicting a first expression level for a set of related genes using the signal for the first sample and a second expression level for the related gene using the signal for the second sample using a method according to any one of embodiments 1-95, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g....
Claims
Attorney Docket: 2014191-0049CLAIMSWhat is claimed is:
1. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene;determining, by the processor, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles corresponding to the genomic region;aggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal for the epigenetic biomarker; andpredicting (e g., using a model), by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signal for each of the one or more epigenetic biomarkers.
2. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for a genomic region corresponding to a target gene; andpredicting (e.g., using a model), by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the signal for each of the one or more epigenetic biomarkers.
3. The method of claim 2, comprising:13404083 v 1 Page 224 of 258Attorney Docket: 2014191-0049determining, by the processor, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles corresponding to the genomic region; andaggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal,wherein predicting the expression level of the target gene based on the signal of each of the one or more epigenetic biomarkers comprises predicting the expression level based on the aggregated tile signal for each of the one or more epigenetic biomarkers.
4. The method of claim 1 or claim 3, comprising tiling, by the processor, the sample sequencing data into tiles that together span the genomic region to obtain tiled sample sequencing data, wherein determining the tile signal is performed using the tiled sample sequencing data.
5. The method of any one of claims 1-4, wherein the method comprises determining a circulating tumor DNA (ctDNA) fraction for the sample.
6. The method of claim 5, wherein predicting the expression level is further based on the ctDNA fraction.
7. The method of claim 5 or 6, wherein determining the ctDNA fraction comprises determining the ctDNA fraction using the sample sequencing data.
8. The method of claim 7, wherein determining the ctDNA fraction comprises determining the ctDNA fraction based on a ctDNA signal for the one or more epigenetic biomarkers.
9. The method of claim 8, wherein the ctDNA signal for the one or more epigenetic biomarkers comprises a ctDNA signal from one or more histone modifications.13404083 v 1 Page 225 of 258Attorney Docket: 2014191-004910 The method of claim 8, wherein the ctDNA signal for the one or more epigenetic biomarkers comprises a ctDNA signal from DNA methylation.11 The method of any one of claims 1-10, wherein predicting the expression level comprises using a model that has been trained on digital samples, each comprising (a) a signal for the one or more epigenetic biomarkers for the target gene, and (b) an expression level for the target gene.
12. The method of claim 11, wherein the model has been produced by:receiving the digital samples;tiling the digital samples into tiles that together span a genomic region corresponding to the target gene; anddetermining a set of expression-level correlated tiles for the digital samples, wherein the determining of the set comprises testing each of the tiles for correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level across the digital samples13. The method of claim 11 or 12, wherein the model has been produced based on the signal for the one or more epigenetic biomarkers for the set of expression-level correlated tiles and the expression level for the target gene for each sample of the digital samples.
14. The method of any one of claims 1-13, comprising determining, for each of the one or more epigenetic biomarkers, a point estimate for the tile signal for the epigenetic biomarker across the set of expression-level correlated tiles.
15. The method of claim 14, wherein the point estimate is a mean.
16. The method of claim 15, wherein the mean is a geometric mean17. The method of any one of claims 1-16, comprising predicting the expression level for the target gene based on a respective point estimate of the signal for each of the one or more epigenetic biomarkers for the set of expression-level correlation tiles.13404083v1 Page 226 of 258Attorney Docket: 2014191-004918. The method of any one of claims 1-17, comprising inputting the sample sequencing data into a model, wherein the model (i) tiles the sequencing data, (ii) determines the tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles, (iii) aggregates the tile signal for each of the one or more epigenetic biomarkers for a set of expression-level correlated tiles, and (iv) predicts the expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signal.
19. The method of any one of claims 1-18, comprising determining, for each of the one or more epigenetic biomarkers, a plurality of point estimates for the signal for the set of expressionlevel correlated tiles20. The method of claim 19, wherein the plurality of point estimates comprises a first point estimate for positively associated tiles of the set of expression-level correlated tiles for one of the one or more epigenetic biomarkers and a second point estimate for negatively associated tiles of the set of expression-level correlated tiles for the one of the one or more epigenetic biomarkers.
21. The method of any one of claims 1-20, wherein the model has been trained by a method comprising determining that a signal for at least one of the one or more epigenetic biomarkers for one or more of the tiles is uncorrelated with the expression level, and excluding the one or more of the tiles from the set of expression-level correlated tiles for the at least one of the one or more epigenetic biomarkers (e.g., excluding the one or more of the tiles from the set of expressionlevel correlated tiles entirely).
22. The method of any one of claims 11-21, wherein the set of expression-level correlated tiles comprises tiles where a correlation between signal for at least one of the one or more epigenetic biomarkers with the expression level exceeds a threshold for a correlation measure (e g., has a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).13404083 v 1 Page 227 of 258Attorney Docket: 2014191-004923 The method of claim 22, wherein the tiles are tiles where the signal for at least one of the one or more epigenetic biomarkers in healthy volunteer samples (e.g., used to generate the subset of the samples) is below a threshold.
24. The method of any one of claims 1-10, wherein the tiles in the set of expression-level correlated tiles are mutually non-overlapping (e.g., and the set of expression-level correlated tiles comprises tiles having non-uniform size).
25. The method of any one of claims 1 -24, wherein the model is a regression (e.g., a multiple regression) based model [e.g., is a regression (e.g., a multiple regression)].
26. The method of claim 25, wherein the regression is an ordinary least squares (OLS) regression (e g., an OLS multiple regression).
27. The method of any one of claims 11-26, wherein the digital samples comprise a subset of digital samples (e.g., wherein all of the digital samples in the subset of the digital samples) corresponding to a particular ctDNA fraction.
28. The method of claim 27, wherein the model has been trained, in part, by determining whether a relationship (e.g., correlation) between the predicted expression level for the target gene and the expression level for the one or more of the samples not included in the subset of the samples across the plurality of subsets of the samples are within one or more predefined criteria (e.g., at a ctDNA fraction of no more than 10%, no more than 8%, no more than 6%, or no more than 5%).
29. The method of claim 28, wherein the one or more predefined criteria comprises one or more predefined criteria for the model, comprising an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%).13404083v1 Page 228 of 258Attorney Docket: 2014191-004930 The method of claim 28 or 29, wherein the one or more predefined criteria comprises one or more predefined criteria for the model, comprising an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9).
31. The method of claim 29 or claim 30, wherein the model has been tuned (e.g., manually) responsive to a determination that the relationship is within the one or more predefined criteria,32. The method of claim 31, wherein model has been tuned by (i) selecting a subset of the one or more of the expression-level correlated tiles and (ii) using the subset of the one or more of the expression-level correlated tiles.
33. The method of claim 31 or claim 32, wherein the model is for an initial ctDNA fraction, and tuning the model has been tuned to produce predictions at a ctDNA fraction lower than the initial ctDNA fraction34. The method of any one of claims 29-33, wherein the model has been trained, in part, by determining that the relationship is within the one or more predefined criteria and, responsive to that determination, performing a training loop with each of at least one subset of the samples that correspond to a lower ctDNA fraction.
35. The method of any one of claims 21-34, wherein the one or more of the digital samples excluded from the subset of the samples is exactly one of the digital samples.
36. The method of any one of claims 1-35, wherein the expression level is an mRNA expression level.
37. The method of any one of claims 1-36, wherein the expression level is a protein expression level.
38. The method of any one of claims 1-37, wherein the one or more epigenetic biomarkers is a plurality of epigenetic biomarkers.13404083vl Page 229 of 258Attorney Docket: 2014191-004939. The method of any one of claims 1-38, wherein the one or more epigenetic biomarkers comprises one or more histone modifications (e.g., H3K27ac modification and / or H3K4me3 modification).
40. The method of any one of claims 1-39, wherein the one or more epigenetic biomarkers comprises DNA methylation.
41. The method of any one of ci aims 1 -40, wherein the one or more epigenetic biomarkers comprises (i) H3K27ac modification and H3K4me3 modification or (ii) H3K27ac modification, H3K4me3 modification, and DNA methylation.
42. The method of any one of claims 1-41, wherein the genomic region comprises a region corresponding to a transcript for the target gene.
43. The method of any one of claims 1-42, wherein the genomic region corresponds to an exon for the target gene.
44. The method of claim 42 or claim 43, wherein the genomic region comprises initial and ending buffer regions (e.g, of ±200 kb) (e.g., around the exon) (eg., around the transcript).
45. The method of any one of claims 4-44, wherein the tiles are overlapping (e.g., wherein no more than two of the tiles are mutually overlapping).
46. The method of any one of claims 4-45, wherein adjacent ones of the tiles overlap by at least 10% (e.g., at least 20%, at least 30% of a length of the tiles) and no more than 70% (e.g., no more than 60% or no more than 50% of the length of the tiles).
47. The method of any one of claims 4-46, wherein the tiles have a length in a range of from 100 to 1000 bp (e.g., from 250 to 750 bp).13404083v 1 Page 23 O of 258Attorney Docket: 2014191-004948 The method of any one of ciaims 4-47, wherein adjacent ones of the tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp)49 The method of any one of claims 11-48, wherein the digital samples have been generated using at least 25 different cell lines.
50. The method of any one of claims 11-49, wherein the digital samples have been generated using at least 25 different healthy volunteers.
51. The method of any one of claims 11-50, wherein the digital samples have been generated using at least 25 different cell lines.
52. The method of any one of claims 11-51, wherein the digital samples have been generated using at least 25 different healthy volunteers.
53. The method of any one of clai s 11-52, wherein the digital samples comprise in silico diluted samples.
54. The method of any one of claims 11-53, wherein the digital samples comprise in silico diluted plasma samples.
55. The method of any one of claims 1-54, comprising normalizing the signal for each of the one or more epigenetic biomarkers.
56. The method of claim 55, wherein normalizing the signal comprises a quantile normalization.
57. The method of any one of claims 1-56, comprising pebbling the signal for each of the one or more epigenetic biomarkers.1340408 v 1 Page 231 of 258Attorney Docket: 2014191-004958 The method of claim 57, wherein the pebbling comprises, for each of the one or more epigenetic biomarkers, determining a background signal for the epigenetic biomarker for the genomic region based on a signal (e.g., a number of fragments) in the sample in a pebbling region, and subtracting the background signal from the signal.
59. The method of any one of claims 1-58, wherein the target gene is a gene that encodes a polypeptide associated with (e.g., targeted by) a therapy.
60. The method of any one of claims 1-59, wherein the target gene is a gene that encodes an antibody drag conjugate (ADC) target antigen.
61. The method of any one of claims 1-60, comprising predicting the expression level for an indication.
62. The method of claim 61, wherein the indication is a cancer indication (e.g., a cancer type or subtype).
63. The method of claim 62, wherein the indication is breast cancer, ovarian cancer, bladder cancer, renal cancer, prostate cancer, pancreatic cancer, or lung cancer (e.g., small cell lung cancer).
64. The method of any one of claims 1-63, wherein the sample sequencing data comprises a signal for each of the one or more epigenetic biomarkers for one or more genomic regions corresponding to one or more additional genes and the method comprises:determining, by the processor, a tile signal for each of the one or more epigenetic biomarkers for a second set of expression-level correlated tiles corresponding to one or more genomic regions corresponding to one or more additional genes; andaggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the for the epigenetic biomarker for the set of expression-level correlated tiles to obtain an aggregated tile signal for the epigenetic biomarker,1340408 v 1 Page 232 of 258Attorney Docket: 2014191-0049wherein predicting the expression level for the target gene is further based on the tile signal for each of the one or more epigenetic biomarkers for the second set of expression-level correlated tiles.
65. The method of claim 64, wherein the one or more additional genes comprise one or more genes related to the target gene.
66. The method of claim 65, wherein the one or more related genes comprises one or more genes that are in a transcriptional complex with the target gene.
67. The method of claim 65 or claim 66, wherein the one or more related genes comprises one or more genes that are master regulators for the indication.
68. The method of any one of claims 65-67, wherein the one or more related genes comprises one or more genes that are related to a pathway relevant for the indication.
69. The method of any one of claims 65-68, wherein the one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes.
70. The method of any one of claims 69-74, wherein the one or more related genes comprises a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].
71. The method of any one of claims 64-70, wherein predicting the expression level for the target gene comprises using an ensemble model.
72. The method of claim 71, wherein the ensemble model comprises a first constituent model and a second constituent model (e.g., comprises a constituent model for the target gene and for each of the one or more additional genes).1340408 v 1 Page 233 of 258Attorney Docket: 2014191-004973 The method of claim 72, wherein the first constituent model is a model for the target gene and the second constituent model is a model for one or more of the one or more additional genes.74 The method of claim 72 or claim 73, wherein the ensemble model uses a ridge regression of the first constituent model and the second constituent model.
74. The method of any one of claims 72-74, wherein the first constituent model comprises a model for the target gene.
75. The method of any one of claims 72-74, wherein the second constituent model comprises a single gene model for a gene related to the target gene.
76. A method of predicting expression level of a target gene in a subject, the method comprising:receiving, by a processor of a computing device, sample sequencing data derived from a nucleic acid in a biological sample derived from a subject, wherein the sample sequencing data comprises a signal for each of one or more epigenetic biomarkers for two or more genes (e.g., for one or more genomic regions corresponding to the two or more genes), wherein the two or more genes comprise a target gene and one or more additional (e.g., related) genes; and predicting, by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the signal for reach of the one or more epigenetic biomarkers for the two or more genes[e.g., wherein the method comprises:determining, by the processor, using the sample sequencing data, a tile signal for each of the one or more epigenetic biomarkers for each of the two or more genes for a set of expression-level correlated tiles (e.g., wherein the method comprises tiling, by the processor, the sample sequencing data into tiles that together span one or more genomic regions corresponding to each of the two or more genes to obtain tiled sample sequencing data and the aggregating is performed using the tiled sample sequencing data);1340408 v 1 Page 234 of 258Attorney Docket: 2014191-0049aggregating, by the processor, for each of the one or more epigenetic biomarkers, the tile signal for the one or more epigenetic biomarkers for the set of expression-level correlated tiles to obtain aggregated tile signals for each of the two or more genes; andpredicting, by the processor, an expression level for the target gene for the subject (e.g., having an indication, such as cancer) (e.g., at the time the sample was taken) based on the aggregated tile signals for each of the two or more genes],77. The method of claim 76, wherein the one or more additional genes comprise one or more genes related to the target gene.
78. The method of claim 77, wherein the one or more related genes comprises one or more genes that are in a transcriptional complex with the target gene.
79. The method of claim 77 or claim 78, wherein the one or more related genes comprises one or more genes that are master regulators for the indication.
80. The method of any one of claims 77-79, wherein the one or more related genes comprises one or more genes that are related to a pathway relevant for the indication.
81. The method of any one of claims 77-80, wherein the one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes.
82. The method of any one of claims 77-81, wherein the one or more related genes comprises a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].
83. The method of any one of claims 76-82, wherein predicting the expression level for the target gene comprises using an ensemble model.1340408 v 1 Page 235 of 258Attorney Docket: 2014191-004984 The method of claim 83, wherein the ensemble model comprises a first constituent model and a second constituent model (e.g., comprises a constituent model for the target gene and for each of the one or more additional genes).
85. The method of claim 84, wherein the first constituent model is a model for the target gene and the second constituent model is a model for one or more of the one or more additional genes.
86. The method of claim 84 or claim 85, wherein the ensemble model uses a ridge regression of the first constituent model and the second constituent model.
87. The method of any one of claims 84-86, wherein the first constituent model comprises a model for the target gene.
88. The method of any one of claims 84-87, wherein the second constituent model comprises a single gene model for a gene related to the target gene.
89. The method of any one of claims 1-88, wherein:(a) the target gene or genomic target is listed in Table 1, and the indication is prostate cancer;(b) the target gene or genomic target is listed in Table 2, and the indication is breast cancer;(c) the target gene or genomic target is listed in Table 3, and the indication is lung cancer (e.g., small cell lung cancer); or(d) the target gene or genomic target is listed in Table 4 and the indication is cancer.
90. A method of making a preliminary prediction of expression of a genomic target for an indication (e.g., at a particular circulating tumor DNA (ctDNA) fraction) and / or producing a model that predicts expression level of a target gene, the method comprising:receiving digital samples for an indication each comprising (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene;13404083vl Page 236 of 258Attorney Docket: 2014191-0049tiling the samples into tiles that together span a genomic region corresponding to the target gene;performing a loop for each of at least one subset of the samples, the loop comprising:determining a set of expression-level correlated tiles for the subset of the samples, wherein the determining of the set comprises testing each of the tiles for correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level across the subset of the samples,producing a model based on the signal for the one or more epigenetic biomarkers for the set of expression-level correlated tiles and the expression level for the target gene for each sample of the subset of the samples, and predicting, using the model, an expression level for the target gene for one or more of the samples not included in the subset of the samples.
91. The method of claim 90, wherein producing the model comprises, for each of the one or more epigenetic biomarkers, determining a point estimate for the signal for the epigenetic biomarker across the set of expression-level correlated tiles and producing the model based on the point estimate.
92. The method of claim 91, wherein the point estimate is a mean.
93. The method of claim 92, wherein the mean is a geometric mean.
94. The method of any one of claims 90-93, wherein the model is produced based on a respective point estimate of the signal for each of the one or more epigenetic biomarkers for the set of expression-level correlation tiles.
95. The method of any one of claims 90-94, wherein producing the model comprises determining a plurality of point estimates for the signal for at least one of the one or more epigenetic biomarkers across the set of expression-level correlated tiles and the model is produced based on the point estimate.13404083vl Page 237 of 258Attorney Docket: 2014191-004996. The method of claim 95, wherein the plurality of point estimates compri ses a first point estimate for positively associated tiles for one of the one or more epigenetic biomarkers and a second point estimate for negatively associated tiles for the one of the one or more epigenetic biomarkers.
97. The method of any one of claims 90-96, wherein determining the set of expression-level correlated tiles comprises determining signal for at least one of the one or more epigenetic biomarkers for one or more of the tiles is uncorrelated with the expression level across the subset of the samples and excluding the one or more of the tiles from the set of expression-level correlated tiles for the at least one of the one or more epigenetic biomarkers (e.g., excluding the one or more of the tiles from the set of expression-level correlated tiles entirely).
98. The method of any one of claims 90-97, wherein determining the set of expression-level correlated tiles comprises determining, for one or more of the tiles, that a correlation between signal for at least one of the one or more epigenetic biomarkers with the expression level exceeds a threshold for a correlation measure (e.g., has a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).
99. The method of claim 98, wherein the one or more of the tiles are determined to be in the set of expression-level correlated tiles further based on the signal for at least one of the one or more epigenetic biomarkers in healthy volunteer samples (e.g., used to generate the subset of the samples) being below a threshold.
100. The method of any one of claims 90-99, wherein determining the set of expression-level correlated tiles comprises: (i) determining overlapping ones of the tiles have a correlation between the signal corresponding to the tile for each of the one or more epigenetic biomarkers and the expression level; and (ii) collapsing the overlapping ones of the tiles such that the set of expression-level correlated tiles are mutually non-overlapping (e.g., and the set of expression-level correlated tiles comprises tiles having non-uniform size).1340408 v 1 Page 238 of 258Attorney Docket: 2014191-0049101. The method of any one of claims 90-100, wherein the model is a regression (e.g, a multiple regression) based model [e.g., is a regression (e.g., a multiple regression)].
102. The method of claim 101, wherein the regression is an ordinary least squares (OLS) regression (e.g., an OLS multiple regression).
103. The method of any one of claims 90-102, wherein all of the samples in the subset of the samples correspond to a particular ctDNA fraction.
104. The method of any one of claims 90-103, comprising performing a plurality of iterations of the loop each using a different subset of the samples where all of the samples in the different subset correspond to a particular ctDNA fraction (e.g., and the one or more of the samples not included in the subset of the samples is a sample from a subset of the samples used in a different one of the plurality of iterations).
105. The method of claim 104, comprising performing the loop for each of at least one subset of the samples for each of a set of ctDNA fractions (e.g., performing a cross validated loop using the samples for each of a set of ctDNA fractions).
106. The method of claim 105, wherein the set of ctDNA fractions corresponds to an expected range of ctDNA fractions for the indication.
107. The method of any one of claims 90-106, wherein all of the samples in the at least one subset of the samples correspond to a particular ctDNA fraction.
108. The method of any one of claims 90-107, wherein the at least one subset of the samples is a plurality of subsets of the samples [e.g., corresponding to folds in a cross-validation loop (eg., for at least one ctDNA fraction)].
109. The method of claim 108, comprising determining whether a relationship (e.g., correlation) between the predicted expression level for the target gene and the expression level1340408 v 1 Page 239 of 258Attorney Docket: 2014191-0049for the one or more of the samples not included in the subset of the samples across the plurality of subsets of the sampl es are within one or more predefined criteria (e.g., at a ctDNA fraction of no more than 10%, no more than 8%, no more than 6%, or no more than 5%).
110. The method of claim 109, wherein the one or more predefined criteria comprises an R2of at least 50% (e.g., at least 60%, at least 70%, or at least 80%) and / or an area under curve (AUC) of at least 0.6 (e.g., at least 0.7, at least 0.8, or at least 0.9).111 The method of claim 109 or claim 110, comprising determining that the relationship is within the one or more predefined criteria and, responsive to that determination, (e.g., manually) tuning (e.g., feature engineering) the model.
112. The method of claim 111, wherein tuning the model comprises selecting a subset of the one or more of the expression-level correlated tiles and the method comprises tuning the model using the subset of the one or more of the expression-level correlated tiles.
113. The method of claim 111 or claim 112, wherein the model is for an initial ctDNA fraction and tuning the model comprises tuning the model to produce predictions at a ctDNA fraction lower than the initial ctDNA fraction.
114. The method of any one of claims 109-113, comprising determining that the relationship is within the one or more predefined criteria and, responsive to that determination, performing the loop with each of at least one subset of the samples that correspond to a lower ctDNA fraction.
115. The method of any one of claims 90-114, wherein the one or more of the samples not included in the subset of the samples is exactly one of the digital samples.
116. The method of any one of claims 90-115, wherein the expression level is an mRNA expression level.1340408 v 1 Page 240 of 258Attorney Docket: 2014191-0049117. The method of any one of claims 90-116, wherein the expression level is a protein expression level.
118. The method of any one of claims 90-117, wherein the one or more epigenetic biomarkers is a plurality of epigenetic biomarkers.
119. The method of any one of claims 90-118, wherein the one or more epigenetic biomarkers comprises one or more histone modifications.
120. The method of any one of claims 90-119, wherein the one or more epigenetic biomarkers comprises DNA methylation,121. The method of any one of claims 90-120, wherein the one or more epigenetic biomarkers comprises H3K27ac modification, H3K4me3 modification, and DNA methylation.
122. The method of any one of claims 90-121, wherein the genomic region comprises a region corresponding to a transcript for the target gene.
123. The method of any one of claims 90-122, wherein the genomic region corresponds to an exon for the target gene.
124. The method of claim 122or claim 123, wherein the genomic region comprises initial and ending buffer regions (e.g., of ±200 kb) (e.g., around the exon) (e.g., around the transcript).
125. The method of any one of claims 90-124, wherein the tiles are overlapping (e.g., wherein no more than two of the tiles are mutually overlapping).
126. The method of any one of claims 90-125, wherein adjacent ones of the tiles overlap by at least 10% (e.g., at least 20%, at least 30% of a length of the tiles) and no more than 70% (e.g., no more than 60% or no more than 50% of the length of the tiles).1340408 v 1 Page 241 of 258Attorney Docket: 2014191-0049127. The method of any one of claims 90-126, wherein the tiles have a length in a range of from 100 to 1000 bp (e.g., from 250 to 750 bp).
128. The method of any one of claims 90-127, wherein adjacent ones of the tiles overlap by an amount in a range of from 10 to 500 bp (e.g., from 100 to 300 bp).
129. The method of any one of claims 90-128, comprising generating the digital samples.130 The method of claim 129, wherein generating the digital samples comprises, for each of the digital samples, (i) randomly sampling data for at least one healthy sample and at least one cell line sample (e g., from one or more indication- specific, e.g., cancer, cell lines) in a mixing ratio corresponding to a desired ctDNA fraction for the digital sample and (ii) determining the expression level for the sample.
131. The method of claim 130, wherein determining the expression level comprises using a previously measured expression level for the at least one cell line sample as the expression level for the sample.
132. The method of any one of claims 90-131, wherein the signal is sequencing counts.
133. The method of any one of claims 90-132, wherein the signal for each of the one or more epigenetic biomarkers for the samples is in silico mixed sequencing data.
134. The method of any one of claims 90-133, wherein the digital samples have been generated using at least 25 different cell lines.135 The method of any one of claims 90-134, wherein the subset of the samples have been generated using at least 25 different healthy volunteers.
136. The method of any one of claims 90-135, wherein the subset of the samples have been generated using at least 25 different cell lines.13404083vl Page 242 of 258Attorney Docket: 2014191-0049137. The method of any one of claims 90-136, wherein the digital samples have been generated using at least 25 different healthy volunteers.
138. The method of any one of claims 90-137, wherein the digital samples comprise in silico diluted samples.
139. The method of any one of claims 90-138, wherein the digital samples comprise in silico diluted plasma samples.
140. The method of any one of claims 90-139, comprising normalizing the signal for each of the one or more epigenetic biomarkers prior to performing the loop such that the loop is performed using the normalized signal.
141. The method of claim 140, wherein normalizing the signal comprises a quantile normalization.142 The method of any one of claims 90-141, comprising pebbling the signal for each of the one or more epigenetic biomarkers prior to performing the loop such that the loop is performed using the pebbled signal.
143. The method of claim 142, wherein the pebbling comprises, for each of the one or more epigenetic biomarkers, determining a background signal for the epigenetic biomarker for the genomic region based on signal (e.g., a number of fragments) in the sample in a pebbling region and subtracting the background signal from the signal.144 The method of any one of claims 90-143, wherein the target gene is a gene that encodes an antibody drug conjugate (ADC) target antigen.
145. The method of any one of claims 90-144, wherein the indication is a cancer indication (e.g., a cancer type or subtype).1340408 v 1 Page 243 of 258Attorney Docket: 2014191-0049146. The method of claim 145, wherein the indication is breast cancer, prostate cancer, or lung cancer (e.g., small cell lung cancer).
147. A method of producing a multigene expression level prediction model, the method comprising:producing a first constituent model for a target gene using a method according to any one of claims 90-146;selecting one or more related genes related to the target gene; andproducing a second constituent model for each of the one or more related genes using a method according to any one of claims 90-146(e.g., wherein the related gene is the target gene in the method); andproducing an ensemble model based on the first constituent model and the second constituent model.
148. The method of claim 148, wherein the ensemble model uses a ridge regression of the first constituent model and the second constituent model.
149. The method of any one of claims 90-148, comprising selecting one or more related genes related to the target gene, wherein the digital samples comprise signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes such that set of expression-level correlated tiles comprises at least one tile corresponding to each of the one or more related genes.
150. The method of any one of claims 90-149, comprising selecting one or more related genes related to the target gene, wherein the digital samples comprise signal for each of the one or more epigenetic biomarkers for the one or more related genes and the tiles further span a respective genomic region for each of the one or more related genes, and determining the set of expression-level correlated tiles comprises testing each of the tiles corresponding to the one or more related genes for correlation between the signal in the sample corresponding to the tile for13404083 v 1 Page 244 of 258Attorney Docket: 2014191-0049each of the one or more epigenetic biomarkers and the expression level for the sample across the subset of the samples.
151. The method of any one of claims 147-150, wherein selecting the one or more related genes comprises selecting one or more genes that are in a transcriptional complex with the target gene.
152. The method of any one of claims 147-151, wherein selecting the one or more related genes comprises selecting one or more genes that are master regulators for the indication.
153. The method of any one of claims 147-152, wherein selecting the one or more related genes comprises selecting one or more genes that are related to a pathway relevant for the indication.
154. The method of any one of claims 147-153, wherein the one or more related genes comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 15 genes.155 The method of any one of claims 147-154, wherein selecting the one or more related genes comprises selecting a number of genes having highest rank according to a measure of correlation [e.g., or as ranked by a data source (e.g., TCGA)].
156. The method of any one of claims 90- 155, wherein:(a) the target gene or genomic target is listed in Table 1, and the indication is prostate cancer;(b) the target gene or genomic target is listed in Table 2, and the indication is breast cancer;(c) the target gene or genomic target is listed in Table 3, and the indication is SCLC; or(d) the target gene or genomic target is listed in Table 4 and the indication is cancer.13404083 v 1 Page 245 of 258Attorney Docket: 2014191-0049157. A method of predicting expression level of a liquid biopsy sample for a subject, the method comprising:providing an expression level prediction model that has been produced from digital samples for an indication each comprising (i) signal (e.g., sequencing counts) for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and (ii) an expression level for the target gene, wherein the digital samples have been generated using data derived from cell samples specific to the indication and healthy volunteers;providing input data comprising signal (e.g., sequencing counts) for the one or more epigenetic biomarkers derived from a liquid biopsy sample for a subject; andpredicting expression level of the target gene for the subject from the input data using the model (e.g., using a method according to any one of claims 1-89).
158. The method of claim 157, wherein the cell samples comprise tissue samples.
159. The method of claim 157 or claim 158, wherein the cell samples have been derived from one or more cell lines, one or more patient-derived xenografts or a biopsy therefrom, one or more organoids, or a combination thereof.
160. The method of any one of claims 157-159, wherein the liquid biopsy sample is a plasma sample.
161. The method of any one of claims 157-160, wherein the one or more epigenetic biomarkers comprise one or more histone modifications and / or DNA methylation.
162. The method of any one of claims 157-161, wherein the one or more epigenetic biomarkers comprises H3K27ac modification, H3K4me3 modification, and DNA methylation.
163. The method of any one of claims 157-162, wherein the model has been produced using a method according to any one of claims 1-162.1340408 v 1 Page 246 of 258Attorney Docket: 2014191-0049164. A method of predicting (e.g., determining) the expression of a genomic target for an indication (e.g., at a particular circulating tumor DNA (ctDNA) fraction), the method comprising quantifying one or more epigenetic biomarkers at one or more expression-level correlated loci for the genomic target.
165. The method of claim 164, where the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci has been shown to be correlated with expression of the genomic target (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0,5).
166. The method of claim 164or 165, wherein the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci in healthy volunteer samples is below a threshold.
167. The method of any one of claims 164-166, wherein the one or more expression-level correlated loci can be or are determined using a method recited in any one of claims 1-109.168 The method of any one of claims 164-167, wherein the expression level is an mRNA expression level.
169. The method of any one of claims 164-168, wherein:(a) the genomic target is listed in Table 1, and the indication is prostate cancer;(b) the genomic target is listed in Table 2, and the indication is breast cancer;(c) the genomic target is listed in Table 3, and the indication is SCLC; or(d) the genomic target is listed in Table 4 and the indication is cancer.170 A method of determining the ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications.1340408 v 1 Page 247 of 258Attorney Docket: 2014191-0049(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andthe one or more genomic loci comprise one or more expression-level correlated loci for ESRI that are provided in Table 5.
171. A method of determining the ER status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andwherein the one or more expression-level correlated loci for the genomic target include one or more expression-level correlated loci for ESRI, ENOl, YBX1, GAT A3, FOXA1, HAPLN3, EN1, PIM, CCDC170, or any combination thereof.
172. The method of claim 171, wherein:(i) the level of the one or more epigenetic biomarkers at the one or more expressionlevel correlated loci have been shown to be correlated with ESR1 expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5), or(ii) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with ESRI, ENOl, YBX1, GATA3, FOXA1, HAPLN3, EN1, PIM, or CCDC170 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).
173. The method of any one of claim 171 or claim 172, wherein the one or more genomic loci comprise one or more of the genomic loci listed in Table 5.
174. A method of determining the PR status of a cancer in a subject, the method comprising:1340408 v 1 Page 248 of 258Attorney Docket: 2014191-0049quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications,(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation;wherein the one or more expression-level correlated loci for the genomic target include one or more expression-level correlated loci for PGR1, SCUBE2, SERPINA11, CA12, ABAT, MAPT-IT1, MAPT, GREB1, PTPRT, PREXI, NEK10, LRIG1, TPRG1, SPEF2, RGS22, NXNL2, FGD3, SUSD3, or GRPR, or any combination thereof.
175. The method of claim 174, wherein:(i) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with PGR1 expression (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5); or(ii) the level of the one or more epigenetic biomarkers at the one or more expression¬ level correlated loci have been shown to be correlated with SCUBE2, SERPINA11, CA12, ABAT, MAPT-IT1, MAPT, GREB1, PTPRT, PREX1, NEK10, LRIG1, TPRG1, SPEF2, RGS22, NXNL2, FGD3, SUSD3, GRPR, or PGR1 expression, or any combination thereof (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).
176. The method of claim 174 or claim 175, wherein the one or more expression-level correlated loci comprise one or more loci listed in Table 6.
177. A method of determining the HER2 status of a cancer in a subject, the method comprising:quantifying, at one or more genomic loci in cell-free DNA (cfDNA) from a liquid biopsy sample obtained or derived from the subject, one or more epigenetic biomarkers, wherein the one or more epigenetic biomarkers comprise:(i) one or more histone modifications.1340408 v 1 Page 249 of 258Attorney Docket: 2014191-0049(ii) chromatin accessibility,(iii) binding of one or more transcription factors, and / or(iv) DNA methylation; andthe one or more genomic loci comprise one or more of the genomic loci provided in Table 7.
178. The method of any one of claims 164-177, wherein the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci in healthy volunteer samples is below a threshold.
179. The method of any one of claims 164-178, wherein the one or more expression-level correlated loci can be or are determined using a method recited in any one of claims 1-109.
180. The method of any one of claims 164-179, wherein the expression level is an mRNA expression level.
181. A method of predicting (e.g., determining) expression of a genomic target for an indication (e.g., at a particular circulating tumor DNA (ctDNA) fraction), the method comprising quantifying one or more epigenetic biomarkers at one or more expression-level correlated loci for the genomic target.
182. The method of claim 181, where the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci have been shown to be correlated with expression of the genomic target (e.g., have a Spearman correlation of at least 0.2, at least 0.3, at least 0.4, or at least 0.5).183 The method of claim 181 or claim 182, wherein the level of the one or more epigenetic biomarkers at the one or more expression-level correlated loci in healthy volunteer samples is below a threshold.13404083 v 1 Page 250 of 258Attorney Docket: 2014191-0049184. The method of any one of ciaims 164-183, wherein the one or more expression-level correlated loci can be or are determined using a method recited in any one of claims 90183.
185. The method of any one of claims 164-184, wherein the expression level is an mRNA expression level.
186. A method of producing a model that predicts expression level of a target gene, the method comprising:receiving digital samples for an indication each comprising (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene;producing a set of constituent models based on the digital samples, wherein the set comprises one constituent model for each of the target gene and the one or more related genes; andproducing an ensemble model based on a combination (e.g,, using a ridge regression) of the constituent models in the set.
187. A method of producing a model that predicts expression level of a target gene, the method comprising:receiving digital samples for an indication each comprising (i) signal for each of one or more epigenetic biomarkers for a target gene corresponding to the indication and one or more selected related genes related to the target gene and (ii) an expression level (e.g., mRNA expression level) for the target gene;determining genomic regions corresponding to the target gene and the one or more related genes where the signal for at least one of the one or more epigenetic biomarkers correlates with the expression level for the samples (e.g., determine expression-level correlated tiles corresponding to the target gene and each of the one or more related genes); and producing a model based on the signal for the one or more epigenetic biomarkers for the genomic regions (e.g., for the expression-level correlated tiles) and the expression level.1340408 v 1 Page 251 of 258Attorney Docket: 2014191-0049188. The method of claim 186 or claim 187, wherein the digital samples have been generated using data derived from cell samples specific to the indication and healthy volunteers.
189. The method of any one of claims 186-188, wherein the cell samples comprise tissue samples.
190. The method of any one of claims 186-189, wherein the cell samples have been derived from one or more cell lines, one or more patient-derived xenografts or a biopsy therefrom, one or more organoids, or a combination thereof.
191. The method of any one of claims 186-190, wherein the liquid biopsy sample is a plasma sample192. The method of any one of claims 186-191, wherein the one or more epigenetic biomarkers comprise one or more histone modifications and / or DNA methylation.
193. The method of any one of claims 186-192, wherein the one or more epigenetic biomarkers comprises H3K27ac modification, H3K4me3 modification, and DNA methylation,194. The method of any one of claims 186-193, wherein the signal for the digital samples and for the subject is sequencing counts.
195. The method of any one of claims 186-194, wherein the model has been produced using a method according to any one of claims 77-191.
196. The method of any one of claims 186195, wherein the preliminary prediction of expression of a genomic target comprises predicting expression status of a target gene (e.g., presence or absence of expression of a target gene, or expression of a target gene above a threshold level).1340408 v 1 Page 252 of 258Attorney Docket: 2014191-0049197. The method of claim 196, wherein the genomic target is ERBB2, and the preliminary prediction of expression is a preliminary prediction of human epidermal growth factor receptor 2 (HER2) expression status (e.g., HER2 IHC 3+ / 2+ISH+ vs. HER2 2+ / 1+ / 0).
198. The method of claim 196, wherein the genomic target is ESRI, and the preliminary prediction of expression is a preliminary prediction of estrogen receptor (ER) expression status (e g., as determined using IHC)199 The method of claim 196, wherein the genomic target is PGR1, and the preliminary prediction of expression is a preliminary prediction of progesterone receptor (PR) expression status (e.g., as determined using IHC).
200. A method of predicting an expression level of a target gene in a subject, the method comprising:providing sample data for a subject, wherein the sample data comprise signal for one or more epigenetic biomarkers for a target gene (e.g., and optionally one or more related genes);providing a model that has been produced using a method according to any one of claims 90-199; andpredicting an expression level of the target gene for the subject using the sample data with the model.
201. The method of claim 200, wherein the sample data have been derived from a liquid biopsy sample (e.g., from plasma) for the subject.
202. A method of monitoring a subject, the method comprising:providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time;predicting a first expression level for a target gene using the signal for the first sample and a second expression level for the target gene using the signal for the second sample using a model that has been produced using a method according to any one of claims 1-109, wherein13404083 v 1 Page 253 of 258Attorney Docket: 2014191-0049between the first time and the second time the subject has been treated with a therapeutic agent (eg, an ADC therapy) corresponding to the target gene; anddetermining a difference in expression level between the first sample and the second sample.
203. The method of claim 202. wherein the change is a reduction in expression level and the therapy is a degrader for the target gene.204 A method of monitoring a subject, the method comprising:providing (e.g., by obtaining) signal for one or more epigenetic biomarkers derived from a first sample from a subject taken at a first time and from a second sample from the subject taken at a second time after the first time;predicting a first expression level for a set of related genes using the signal for the first sample and a second expression level for the related gene using the signal for the second sample using a model that has been produced according to a method according to any one of claims 1-201, wherein between the first time and the second time the subject has been treated with a therapeutic agent (e.g., an ADC therapy) corresponding to at least one of the related genes; and determining a difference in expression level between the first sample and the second sample.
205. The method of claim 204, wherein the related genes correspond to a pathway.
206. A method of characterizing cancer recurrence and / or progression, the method comprising:predicting an expression level of a target gene for a subject having a cancer based on a first sample for the subject taken at a first time point using a model that has been produced using a method according to any one of claims 90-201,predicting an expression level of a target gene for the subject based on a second sample for the subject taken at a second time point after the first time point using a method according to any one of claims 90-201; and13404083 v 1 Page 254 of 258Attorney Docket: 2014191-0049determining a difference in the expression level at the second time point and at the first time point.
207. A method of monitoring cancer in a subject, the method comprising:predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of claims 1-201; anddetermining whether there is a difference in expression level of the target gene for the subject over time.
208. A method of determining effectiveness of a therapeutic agent in a subject having cancer, the method comprising:predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of claims 1-201; anddetermining whether there is a difference in expression level of the target gene for the subject over time.
209. A method of monitoring response of a subject having an indication to a therapeutic agent for the indication, the method comprising:predicting an expression level for a target gene based on a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of claims 1-201; anddetermining whether there is a difference in expression level of the target gene for the subject over time.210 The method of any one of claims 202-209, comprising administering a therapy comprising a therapeutic agent to the subject between when two or more of the samples were obtained.13404083 v 1 Page 255 of 258Attorney Docket: 2014191-0049211. The method of any one of claims 202-210, comprising administering a therapy to the subject when the difference is determined to be at least as large as a threshold difference.
212. The method of any one of claims 202-211, comprising altering administration of a therapy to the subject when the difference is determined to be at least as large as a threshold difference.
213. The method of claim 212, wherein altering administration comprises increasing a dosage and / or frequency of administration214. A method of prognosing cancer, the method compri sing:predicting an expression level of a target gene for a subject having a cancer based on a sample for the subject using a model that has been produced using a method according to any one of claims 1-201; andprognosing cancer in the subject based on the determined expression level.
215. The method of claim 214, comprising administering a therapy based on the prognosis.
216. A method of diagnosing cancer in a subject, the method comprising:predicting an expression level of a target gene for a subject having a cancer based on a sample for the subject using a model that has been produced using a method according to any one of claims 1-202; anddetermining that the expression level exceeds a threshold.
217. The method of claim 216, comprising initiating administration of a therapy based on the expression level.
218. The method of claim 217, comprising selecting a dosing regimen for the therapy based on the expression level.13404083 v 1 Page 256 of 258Attorney Docket: 2014191-0049219. A method of determining whether a cancer has been removed from a subject, the method comprising, after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, predicting expression level of a target gene for the subject based on a sample for the subject using a model that has been produced using a method according to any one of claims 1-202.
220. The method of claim 219, comprising continuing administration of a therapy based on the expression level.
221. The method of claim 220, comprising ceasing administration of a therapy based on the expression level.
222. A system comprising a processor and one or more non-transitory computer readable media having instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising the method according to any one of claims 1-221.
223. One or more non-transitory computer readable media having instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising the method according to any one of claims 1-221.13404083 v 1 Page 257 of 258