Circulating tumor fraction estimation models, methods of producing them, and methods of using them
Patent Information
- Application Number
- PCT/US2026/021381
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure US2026021381_01102026_PF_FP_ABST
Abstract
Description
[0001] Attorney Docket No. : 2014191-0048
[0002] CIRCULATING TUMOR FRACTION ESTIMATION MODELS, METHODS OF PRODUCING THEM, AND METHODS OF USING THEM
[0003] BACKGROUND
[0004] [1] Many liquid biological samples have cell free DNA (cfDNA) in them, such as blood or plasma samples. Such cfDNA may come from a variety of sources. If a subject has a tumor, cfDNA that is derived from a solid tumor may be observed. Cell-free circulating tumor DNA (ctDNA) may be present in a sample from a subject with cancer. The amount of such ctDNA relative to the total amount of cfDNA in a sample is referred to as circulating tumor fraction (cTF). There is increasing interest in cTF as a biomarker and as a control variable for cancer classifiers. Therefore, it is desirable to estimate the cTF in biological samples. Accordingly, methods have been developed to estimate cTF using sequencing data including variant allele fraction (VAF) methods and copy number variation (CNV) methods, such as ichorCNA. Prior methods, however, are unable to accurately predict cTF at clinically relevant fractions (e.g., below 3%) and are limited to use with specific sequencing methods. There is a continued need, therefore, for estimation models to estimate cTF for biological samples.
[0005] SUMMARY
[0006] [2] Disclosed herein are, inter alia, cTF estimation models and methods of producing and using them. The cTF estimation models can estimate cTF in samples, such as liquid biopsy samples, such as plasma samples, from subjects using sequencing data for the samples. Sequencing data can be obtained using various techniques. Sequencing data may be epigenetic modification sequencing data that comprises information about epigenetic modification of nucleic acids or their associated histones in a sample. Epigenetic modification sequencing data may comprise histone modification sequencing data or methylation sequencing data. Histone modification sequencing data may reflect sequencing data of genetic locations in a nucleic acid, where the nucleic acid was bound to a histone comprising a particular chemical modification. Sequencing data may be methylation sequencing data that has been derived from a sequencing method that produces counts where one or more genetic locations are methylated. In some embodiments, sequencing data is from a region-based sequencing method where counts in the sequencing data indicate whether a particular region (e.g., locus) is methylated or not. In some embodiments, sequencing data has been produced via an enrichment method that enriches for Page 1 of 124
[0007] 13414664vlAttorney Docket No. : 2014191-0048
[0008] particular epigenetic modification (e.g., histone modification and / or methylation). Various enrichment methods that produce such region-based counts exist, such as methyl-binding domain sequencing (MBD-seq) (sometimes also called Methyl-CpG-Binding Domain Sequencing). A set of candidate regions where methylation may be indicative of cancer may be selected (e.g., identified) and then a determination of whether methylation is indicative of cancer may be made using the sequencing data to identify signal regions. Alternatively or additionally, an enrichment method may comprise enrichment based on a histone modification, for example by using an antibody that binds to a particular modified histone, precipitating an antibody bound to said histone in a complex with a nucleic acid, disassociating the histone from the nucleic acid, and sequencing the nucleic acid (e.g., chromatin immunoprecipitation sequencing (e.g.ChlP-seq)). Similarly, candidate regions may be selected based on potentially differential histone modification regions in cancer, and whether said differential histone modification is indicative of cancer may be determined using the sequencing data. Sequencing data may be normalized in order to make accurate determinations of signal regions, for example per-region-based sequencing data may be normalized by both region length and library size (number of fragments). Noise regions in the genome that are indicative of technical noise (e.g., batch effects) may also be determined. The normalized sequencing data and / or one or more metrics calculated using the normalized sequencing data may be used to produce a cTF estimation model, for example based on normalized counts of the sequencing data for the signal regions and / or noise regions.
[0009] [3] Methods for estimating cTF in a sample from a subject based on epigenetic modifications, such as histone modifications, have not existed previously. To date, most cTF estimation methods have focused on direct chromosomal abnormalities indicative of cancer, such as insertions, deletions and / or inversions. Without wishing to be bound by any particular theory, the present disclosure utilizes an insight that localization of certain modified histones may be correlated with (e.g., inversely correlated with) epigenetic modifications (e.g., methylation or modification of a different histone position) known to be indicative of cancer at particular regions. A region may be determined to have been bound by a particular modified histone, which may be indicative of non-alteration of said region by a different epigenetic modification that may correlate with cancer at the region. Such epigenetic modification may be H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, pan acetylation, or a combination thereof. In certain embodiments, such epigenetic modification comprises or is an H3K4me3
[0010] Page 2 of 124
[0011] 13414664vlAttorney Docket No. : 2014191-0048
[0012] modification. A cTF estimation model may be then produced based on region-based histone modification sequencing data.
[0013] [4] Moreover, methods for estimating cTF based on methylation disclosed herein differ from other methods that may use base specific sequencing data such as bisulfite (BS) or enzymatically (EM) converted methylation sequencing. BS- or EM-seq produce reads that have specific bases that are either methylated or unmethylated. With these data each CpG site can be assigned a methylation fraction, or an individual read can be assigned a probability of originating from tumor. Therefore, sequencing data are considered either on a per-base or a per-sample basis, not a per-region basis. Accordingly, methods for estimating cTF using base specific methylation sequencing data cannot be applied to region-based methylation sequencing data. Region-based methylation sequencing methods instead indicate that a region is methylated or not (which may be the result of a single site or multiple sites in the region being methylated) and cTF estimation models used with sequencing data from region-based methylation sequencing methods should account for this difference.
[0014] [5] Enrichment methodologies, such as MBD-seq, differ from certain other sequencing methods for detecting and quantifying DNA methylation, such as bisulfite sequencing. Such enrichment methods may produce counts that correspond to whether a region is methylated or not methylated whereas other sequencing methods may produce counts that correspond to whether a particular site is methylated or not methylated. Histone modification-based enrichment methods similarly can produce region-based counts for a particular histone modification. Such a difference means that models derived for per-site sequencing methods are inapplicable to use with sequencing data from per-region based sequencing methods such as, for example, MBD-seq or ChlP-seq. Moreover, in some embodiments, models disclosed herein are produced based on data that is specific to per-region sequencing methods (e.g., enrichment methods), for example data that has been normalized to account for the use of a per-region based sequencing method. Per-site based models may suffer from increased noise leading to higher limits of detection and / or require more expensive and / or time consuming sequencing in order to produce sufficient sequencing data for model production and use.
[0015] [6] It is an insight of the present disclosure that region-based epigenetic modifications sequencing data can be used for cTF estimation. In particular, use of region-based methylation sequencing data in cTF estimation models can provide one or more advantages. Alternatively or Page 3 of 124
[0016] 13414664vlAttorney Docket No. : 2014191-0048
[0017] additionally, histone modification region-based sequencing data may be used to provide one or more advantages in cTF estimation models. Currently, whole genome sequencing may go down to sufficient depth (e.g., lOOx or lOOOx) and therefore may not be as sensitive as region-based methods, which may allow for improved limit of detection (LoD) in cTF estimation models disclosed herein. Region-based sequencing data can also provide a smoothing effect that can improve model performance, especially at low cTF, because such data may be more genome-wide than site-based methods and / or be less sensitive to confounding factors that may cause single base site methylation or non-methylation. For example, methylation at a particular site may be not as well correlated with cancer whereas methylation in a region around that site may be well correlated (or at least as well correlated). Moreover, a particular histone modification may be correlated with cancer, and / or a histone modification may be correlated (e.g., inversely correlated) with methylation at a particular region that is correlated with cancer. Similarly, CNV- or VAF- based estimation methods may be inherently noisier, for example because subjects may have only a limited number of CNVs or single nucleotide polymorphisms (SNPs).
[0018] [7] cTF estimation models disclosed herein may be preferrable, alternatively or additionally, because they can be readily integrated into existing workflows. Region-based methylation sequencing data, such as MBD-seq data, may already be obtained as part of a workflow for classifying (e.g., scoring) cancer. Alternatively or additionally, other region-based sequencing approaches, such as ChlP-seq, may be used. Such data can be reused to also estimate cTF, or can be used to estimate cTF as a gate to determine whether to proceed with further workflow, for example because no further analysis or work may be warranted if a sample has no cTF. Region-based sequencing data for one or more epigenetic modifications may be used to estimate cTF. In some embodiments, ChlP-seq data or MBD-seq data may be used. In some embodiments, ChlP-seq data for one or more histone modifications is used, and further integrated in a variety of workflows that may use histone modification data, for example, for predicting an outcome for a subject with cancer. In some embodiments, cTF estimation models that use regionbased sequencing data, such as MBD-seq data, allows cTF to be estimated without having to perform an additional reaction. Accordingly, sample volumes can be reduced by using cTF estimation models disclosed herein as compared to cTF estimation models that use data from different sequencing methods. Reduced sample volumes are desirable because there are often practical limits to the amount of available sample or at least more sample volume would require Page 4 of 124
[0019] 13414664vlAttorney Docket No. : 2014191-0048
[0020] more extensive biopsy (e.g., blood draw) from a subject. Tn some embodiments, sufficient data may be obtained to estimate cTF from no more than 1 mb of liquid sample (e.g., blood) from a subject. Similarly, cTF estimation models disclosed herein can allow for serial performance of wet lab sequencing steps, avoiding the need to split a sample into separate aliquots.
[0021] [8] In some embodiments, the present disclosure is directed to a method of producing (e.g., training and / or fitting) a circulating tumor fraction (cTF) estimation model. The cTF estimation model may be a cancer-specific or pan-cancer model. The cTF estimation model may be for estimating cTF in human samples or non-human samples. The method may include selecting (e.g., identifying) candidate regions in a genome, for example where methylation may be indicative of cancer (e.g., correlates with cTF). Alternatively or additionally, selecting (e.g., identifying) candidate regions in a genome may comprise selecting regions with differential histone modification between cancer and healthy subjects. The method may include receiving, by obtaining or receiving previously obtained, sequencing data, for example methylation sequencing data, for a plurality of samples, for example of one cancer type or subtypes or multiple cancer types or subtypes. The method may comprise receiving, by obtaining or receiving previously obtained, histone modification sequencing data for a plurality of cancer samples. The plurality of samples each has a known cTF that is known empirically, known by expectation, or known by assumption. The plurality of samples may include empirical samples and / or in silico sample (e.g., in silico diluted samples). The method may include normalizing counts of the sequencing data for each of the candidate regions for each of the samples. The method may include determining signal regions for the genome, for example where methylation is indicative of cancer (e.g., correlates with cTF), based on the normalized counts of the sequencing data for the candidate regions. Signal regions may be determined based on regions where histone modification may be indicative of cancer. The method may include determining noise regions for the genome, for example where correlation is observed in healthy samples but not in cancer samples, using the normalized counts for the candidate regions. The method may include producing (e.g., fitting and / or training) a cTF estimation model using at least using at least the normalized counts for the signal regions and the known cTF for the samples (e.g., and the normalized counts for the noise regions, if determined). Such a cTF estimation model may determine an estimated cTF in a sample based on input sequencing data or data (e.g., one or more metrics) (e.g., one or more point estimates) derived therefrom.
[0022] Page 5 of 124
[0023] 13414664vlAttorney Docket No. : 2014191-0048
[0024] [9] In some embodiments, the present disclosure is directed to a method of estimating cTF in a sample, for example using a cTF estimation model that has been produced using a method disclosed herein. The method may include receiving, for example by obtaining, sequencing data for a sample (e.g., a liquid sample) from a subject. The method may include producing normalized counts of the sequencing data for signal regions in the genome of the subject. The method may include estimating cTF with a cTF estimation model based at least on the normalized counts. The method may include inputting the normalized counts for the signal regions, or data (e.g., one or more metrics) (e g., one or more point estimates) derived therefrom, into a cTF estimation model and obtaining an estimated cTF as output. Determining an estimated cTF using a cTF estimation model may include determining one or more preliminary estimated cTFs, for example with constituent models of an ensemble model, and using the one or more preliminary estimated cTFs to determine the estimated cTF (e.g., by combining preliminary estimated cTFs together, for example using a mean or other point estimator). Determining an estimated cTF using a cTF estimation model may include determining whether ctDNA is present in a sample using a first model of a cTF estimation model and then, if ctDNA is determined to be present, estimating cTF using a second model of a cTF estimation model. Determining an estimated cTF may include determining a preliminary estimated cTF and then using a first model of a cTF estimation model to estimate cTF if the preliminary estimated cTF is exceeds a threshold and using a second model of the cTF estimation model to estimate the cTF if the preliminary estimated cTF does not exceed the threshold.
[0025]
[0010] Estimated cTF as estimated using a cTF estimation model disclosed herein may be used to monitor a subject, diagnose cancer in a subject, prognose cancer in a subject, or determine origin of cancer in a subject, for example. Monitoring may include determining cTF from different samples from a subject taken at different times. In some embodiments, the present disclosure is directed to a method of characterizing cancer recurrence and / or progression using estimated cTF from samples for a subject taken at different times [e.g., by comparing the estimated cTFs (e.g., by determining a difference (e.g., change or lack thereof) in estimated cTF)]. In some embodiments, the present disclosure is directed to a method of monitoring cancer in a subject using estimated cTF determined from a series of two or more samples for a subject taken over a period of time (e.g., by determining whether there is a difference in cTF for the samples over time). In some embodiments, the present disclosure is directed to a method of prognosing cancer (e.g., a Page 6 of 124
[0026] 13414664vlAttorney Docket No. : 2014191-0048
[0027] stage of cancer, e g., as early or late stage) in a subject using estimated cTF, for example whether the estimated cTF exceeds a threshold. In some embodiments, the present disclosure is directed to a method of diagnosing cancer in a subject using estimated cTF, for example whether the estimated cTF exceeds a threshold. In some embodiments, the present disclosure is directed to a method of determining whether a cancer has been removed from a subject, for example after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, by estimating cTF in a sample from the subject, for example after administration of the therapy and / or surgery, respectively. In some embodiments, the present disclosure is directed to a method of predicting a tumor of origin for ctDNA in a sample from a subject by estimating cTF using different cancer-specific models and comparing the estimated cTFs. In some embodiments, the present disclosure is directed to a method of making a preliminary diagnosis of cancer in a subject by estimating cTF using different cancer-specific models and comparing the estimated cTFs.
[0028]
[0011] Any two or more of the features described in this specification, including in this summary section, may be combined to form implementations of the disclosure, whether specifically expressly described as a separate combination in this specification or not.
[0029]
[0012] At least part of the methods, systems, and techniques described in this specification may be controlled by executing, on one or more processing devices, instructions that are stored on one or more non-transitory machine-readable storage media. Examples of non-transitory machine-readable storage media include read-only memory, an optical disk drive, memory disk drive, and random access memory. At least part of the methods, systems, and techniques described in this specification may be controlled using a computing system including one or more processing devices and memory storing instructions that are executable by the one or more processing devices to perform various control operations.
[0030] BRIEF DESCRIPTION OF THE DRAWINGS
[0031]
[0013] The present teachings described herein will be more fully understood from the following description of various illustrative embodiments, when read together with the accompanying drawings. It should be understood that the drawing described below is for illustration purposes only and is not intended to limit the scope of the present teachings in any way. The foregoing and other objects, aspects, features, and advantages of the disclosure will
[0032] Page 7 of 124
[0033] 13414664vlAttorney Docket No. : 2014191-0048
[0034] become more apparent and may be better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0035]
[0014] FIG. 1 illustrates a method of forming a cTF estimation model, according to illustrative embodiments of the present disclosure; and
[0036]
[0015] FIG. 2 is a block diagram of an example network environment for use in the methods and systems described herein, according to illustrative embodiments of the present disclosure;
[0037]
[0016] FIG. 3 is a block diagram of an example computing device and an example mobile computing device, for use in illustrative embodiments of the present disclosure.
[0038]
[0017] FIG. 4 is a plot showing an example fit for a particular signal region, according to illustrative embodiments of the present disclosure;
[0039]
[0018] FIG. 5 is a plot showing a noise region, according to illustrative embodiments of the present disclosure;
[0040]
[0019] FIG. 6 is a plot showing a regression for an SNR-based cTF estimation model, according to illustrative embodiments of the present disclosure;
[0041]
[0020] FIG. 7 is a plot comparing two estimation models for breast cancer, according to illustrative embodiments of the present disclosure;
[0042]
[0021] FIG. 8 is four plots showing performance of an expectation maximization model for breast cancer, according to illustrative embodiments of the present disclosure;
[0043]
[0022] FIG. 9 is a series of plots illustrating that a LoD of about 0.5% is stable assuming a limit of blank (LoB) is about 0.3% for a breast cancer model, according to illustrative embodiments of the present disclosure;
[0044]
[0023] FIG. 10 is a plot illustrating that a LoB of 0.3% is consistent with a 95% specificity in unseen healthy samples for a breast cancer model, according to illustrative embodiments of the present disclosure;
[0045]
[0024] FIG. 11 is a plot illustrating comparison of limits of detection across different numbers of fragments for breast cancer models, according to illustrative embodiments of the present disclosure; and
[0046]
[0025] FIG. 12 is a plot illustrating limits of detection across a range of cTF fractions between a cTF estimation model according to illustrative embodiments of the present disclosure and known model based on copy number alterations (ichorCNA).
[0047] Page 8 of 124
[0048] 13414664vlAttorney Docket No. : 2014191-0048
[0049] DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0050]
[0026] Disclosed herein are, inter alia, systems and methods for producing and using cTF estimation models that can estimate cTF of a sample, such as a liquid biopsy sample, using sequencing data, or data derived from sequencing data (e.g., one or more metrics), for the sample. Such models may be trained using sequencing data from a plurality of samples having known cTF. Such a plurality of samples may be derived from samples obtained from subjects known to have cancer, healthy subjects, cell lines, and / or in silica diluted samples (e.g., samples that are synthetically derived in silica from two or more other samples). Known cTF for a sample may be empirically known, assumed, or expected. In some embodiments, a cTF estimation model is an epigenetic modification-based cTF estimation model. In some embodiments, a cTF estimation model is a histone modification-based cTF estimation model. In some embodiments, a cTF estimation model is a methylation-based cTF estimation model.
[0051]
[0027] Sequencing data used to produce a cTF estimation method may be epigenetic modification sequencing data derived from a sequencing method that produces a signal corresponding to one or more epigenetic modifications. In some embodiments, epigenetic modification sequencing data may be derived from a methylation sequencing method or a histone modification sequencing method. In some embodiments, sequencing data may be histone modification sequencing data. Such histone modification sequencing data may include sequencing reads for a nucleic acid that was bound to a histone bearing a particular chemical modification in a subject and in a sample from a subject. Such epigenetic modification may be H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, pan acetylation, or a combination thereof. In certain embodiments, such epigenetic modification comprises or is an H3K4me3 modification. In some embodiments, sequencing data has been derived from an enrichment method (e.g., a modified histone enrichment method). In some embodiments, sequencing data has been derived from chromatin immunoprecipitation sequencing (ChlP-seq). Sequencing data may be normalized. Sequencing data may be useful to identify signal regions, at which binding of a particular modified histone may be indicative (e g., predictive) of cancer. Without wishing to be bound by any particular theory, histone modification sequencing data may provide an insight into signal regions absent of a different epigenetic modification (e.g., methylation or a different histone modification) that is indicative (e.g., predictive) of cancer at those particular regions.
[0052] Page 9 of 124
[0053] 13414664vlAttorney Docket No. : 2014191-0048
[0054]
[0028] Sequencing data used to produce a cTF estimation method may be methylation sequencing data derived from a methylation sequencing method. Such methylation sequencing data may include counts that indicate whether a particular region is methylated or not or may indicate whether a particular site is methylated or not. In some embodiments, sequencing data has been derived from an enrichment method (e.g., a methylation enrichment method). In some embodiments, sequencing data is methyl-binding domain sequencing (MBD-seq) data. Sequencing data may be normalized and used to determine signal regions where methylation in the region is indicative (e.g., predictive) of cancer, for example where a linear or non-linear correlation between normalized counts in methylation sequencing data and cTF exists, from a set of selected candidate regions in a genome.
[0055]
[0029] In some embodiments, candidate regions may be regions with differential epigenetic modifications that may be associated with cancer. In some embodiments, differential regions with epigenetic modifications may comprise regions with differential histone modifications and / or differential methylation between two or more cell types or sources (e.g. cancer vs non-cancer). In some embodiments, candidate regions may be regions with differential histone modifications (DHMRs), for example, cancer-specific DHMRs (cDHMRs). In some embodiments, signal regions are regions determined to be DHMRs, for example cDHMRs. In some embodiments, candidate regions are regions that are or may be differentially methylated regions (DMRs), for example cancer-specific differentially methylated regions (cDMRs). In some embodiments, signal regions are regions determined to be DMRs, for example cDMRs.
[0056]
[0030] Technical noise may be accounted for in a cTF estimation model using sequencing data to determine noise regions and then using normalized counts for the noise regions from the sequencing data when producing the cTF estimation model. Using data described herein. Normalized sequencing data (e.g., normalized counts of sequencing data for samples for signal regions and / or noise regions) or data derived therefrom (e.g., one or more metrics) (e.g., one or more point estimates) may be used to produce a cTF estimation model, for example by fitting the data, that uses sequencing data or data derived therefrom as input to produce an estimated cTF.
[0057]
[0031] Once produced, a cTF estimation model may be used to estimate cTF of new samples using sequencing data received (e.g., obtained) for the samples. Estimated cTF may be used to monitor a subject, diagnose cancer in a subject, prognose cancer in a subject, or determine origin of cancer in a subject, for example. Monitoring may include determining cTF from different Page 10 of 124
[0058] 13414664vlAttorney Docket No. : 2014191-0048
[0059] samples from a subject taken at different times. Changes in cTF, presence of a non-zero cTF, or differences in cTF estimated by cancer-specific cTF estimation models may be used to determine one or more characteristics of a cancer that a subject may or does have. A subject may be a patient.
[0060]
[0032] A cTF estimation model may estimate cTF in one cancer type (e.g., breast cancer), or even one cancer subtype (e g., HER2 negative breast cancer), for example based on having been produced (e.g., trained and / or fit) using samples only from that cancer type. A cTF estimation model may estimate cTF across multiple cancer types, that is may be a pan-cancer model, for example based on having been produced (e.g., trained and / or fit) using samples from a variety of different cancer types. Methods disclosed herein are tumor naive in that no information about tumor of origin for ctDNA need be known in order to estimate cTF, down to the limit of detection for the model, which may be no more than 3%, no more than 2%, or no more than 1% cTF.
[0061]
[0033] cTF estimation models disclosed herein may have low limit of detection. For example, in some embodiments, a cTF estimation model may be able to estimate cTF down to a limit of no more than 3%, no more than 2%, no more than 1%, no more than 0.75%, or about 0.5%, and optionally at least 0.5%. Moreover, cTF estimation models disclosed herein have such low limit of detection while also being tumor naive. Achieving low limit of detection while being tumor naive eliminates the need to perform invasive and / or risky procedures (e.g., radiological imaging and / or solid biopsy) in order to estimate cTF. Such low limit of detection is useful for a number of applications, including early cancer screening and more sensitive cancer monitoring among others. The tumor naive nature of cTF estimation models disclosed herein also allows for cancer screening and / or monitoring to be easily integrated into workflows for samples (e.g., liquid biopsy samples, e.g., plasma samples) because no prior knowledge of a tumor or cancer of a subject is needed in order to estimate cTF for a sample from the subject and such models may use already use same sequencing data as used for other analysis (e.g., estimators) used in a workflow. Tumor informed models require independently obtaining some knowledge that a subject has cancer and what the nature of the subject’s tumor is in order to be used. cTF estimation models having such low limit of detections may also be useful in validating other related cancer estimators and / or understanding performance of such estimators.
[0062] Producing cTF Estimation Models
[0063]
[0034] Circulating tumor fraction (cTF) estimation models as disclosed herein can estimate Page 11 of 124
[0064] 13414664vlAttorney Docket No. : 2014191-0048
[0065] the amount of cell free DNA (cfDNA) that is believed to be derived from cancer, for example from a solid tumor. cTF estimation models as disclosed herein can estimate circulating tumor DNA (ctDNA) as a fraction of the total cfDNA sequenced from a plasma sample. cTF ranges from 0 to 1 or can be expressed as a percentage from 0-100%.
[0066]
[0035] Models for cTF estimation disclosed herein estimate cTF based on one or more epigenetic modifications. At a high level, a methylation-based estimated cTF estimation model uses differences in DNA methylation patterns between healthy and cancerous tissues. As an example, plasma cfDNA represents an admixture of DNA that originates from different potential tissue types in a hematological background. Magnitude of methylation across many regions of the genome may be proportional, linearly or non-linearly, to the amount of cfDNA that originates from a particular tissue. As an example, a healthy patient should have very little, if any, detectable methylated DNA that is specific to breast tissue in their plasma, whereas a breast cancer patient with actively shedding tumor DNA into the bloodstream should.
[0067]
[0036] A histone modification-based model for cTF estimation may use differential histone modification signals in healthy and cancer subject. Histone modifications may differ between cancerous and non-cancerous tissues based on known properties of particular modified histones pertinent to driving oncogene activity or inhibiting tumor-suppressor genes. Histone modifications may also be indicative of absence of cancer-specific methylation (e.g., wherein cancer-specific methylation may be presence of either increased or decreased methylation in a particular region).
[0068]
[0037] cTF estimation models disclosed herein, and methods of producing them, have been developed based on these insights.
[0069]
[0038] In some embodiments, a method of producing a cTF estimation model includes the following steps. In a first step, candidate regions that may be differentially methylated are selected (e.g., identified). In a second step, sequencing data for a range of samples with known cTF are normalized for the candidate regions. In a third step, differentially methylated regions (DMRs) of the genome for the samples that correlate with cTF are determined to be signal regions. In a fourth step, which is performed in some methods but is not a necessary step, regions of the genome that show correlation only in the healthy samples are determined. In a fifth step, a cTF estimation model is produced (e.g., fitted). Such a model may be an ensemble model that includes a plurality of constituent models.
[0070] Page 12 of 124
[0071] 13414664vlAttorney Docket No. : 2014191-0048
[0072]
[0039] In some embodiments, in a first step, a method of producing a cTF estimation model comprises selecting (e.g., identifying) candidate regions that may have been bound by a modified histone. In a second step, sequencing data for a range of samples with known cTF are normalized. In a third step, DHMRs of the genome for the samples that correlate with cTF are determined to be signal regions. In a fourth step, which is performed in some methods but is not a necessary step, regions of the genome that show correlation only with the median of the signal in the healthy samples are determined. In a fifth step, a cTF estimation model is produced (e.g., fitted).
[0073]
[0040] In some embodiments, the present disclosure provides a method of producing a circulating tumor fraction (cTF) estimation model for estimating an amount of circulating tumor DNA in a sample (e.g., plasma sample) of a subject. The method may include selecting candidate regions (e.g., loci) in a genome where methylation may be indicative of cancer (e.g., may correlate with cTF). Alternatively or additionally, the method may include selecting candidate regions (e.g., loci) in a genome where histone modification quantity may be indicative of cancer (e.g., may correlate with cTF). Selecting the candidate regions may include identifying the candidate regions. The genome is a genome of interest and samples used to produce a model will correspond to the genome, for example a cTF estimation model being produced or that has been produced may be for estimating cTF in one or more types of human cancer and therefore candidate regions are regions of the human genome. The method may further include receiving sequencing data for a plurality of samples corresponding to the genome. The method may include obtaining the sequencing data (e.g., by running a sequencing assay). In some embodiments, the sequencing data may be methylation sequencing data derived from a methylation sequencing technique, such as, for example MBD-seq. In some embodiments, the sequencing data may be histone modification sequencing data derived from a histone modification sequencing technique, such as, for example, ChlP-seq. Each of the plurality of samples has a known cTF that may be empirically known, assumed, or expected, as described further subsequently. The method may further include normalizing counts of the sequencing data for each of the candidate regions for each of the samples. The method may further include determining signal regions (e.g., loci) in the genome where methylation of the signal region is indicative (e.g., predictive) of cancer (e.g., that correlates with cTF) using the normalized counts for the candidate regions and the known cTF for the samples. Alternatively or additionally, a signal region with a histone modification indicative (e.g., predictive) of cancer (e.g., that correlates with cTF) may be determined using the normalized Page 13 of 124
[0074] 13414664vlAttorney Docket No. : 2014191-0048
[0075] counts for the candidate regions at known cTF. The method may further include determining noise regions for the genome using the normalized counts of the sequencing data, for example from candidate regions that are not signal regions. The method may further include producing (e.g., fitting) a cTF estimation model that determines an estimated cTF in a sample based on input sequencing data or data (e.g., one or more metrics) (e.g., one or more point estimates) derived therefrom. The producing may be performed using at least the normalized counts for the signal regions and the known cTF for the samples, for example and also the normalized counts for the noise regions for the samples.
[0076]
[0041] FIG. 1 illustrates an example method disclosed herein. In step 102, candidate regions are selected. In step 104, sequencing data, for example MBD-seq data, for samples with known cTF is received. In step 106, sequencing data is normalized, for example by region length and number of fragments, to produce normalized counts. In step 108, signal regions are determined using normalized counts for the candidate regions from the sequencing data. In optional step 110, noise regions are determined using normalized counts for the candidate regions from the sequencing data. In step 112, a cTF estimation model is produced using normalized counts for the signal regions and, if optional step 110 is performed, the noise regions and the known cTFs for the samples. In step 114, the cTF estimation model may be used to estimate cTF of a new sample using sequencing data for that sample.
[0077]
[0042] The following description provides additional details regarding embodiments of candidate regions and how they may be selected (e.g., identified), training samples and sequencing data normalization processes, signal regions and how they may be determined, noise regions and how they may be determined, and subsequent cTF estimation model structure and production.
[0078] Candidate Regions
[0079]
[0043] In some embodiments, candidate regions in a genome may be selected to determine whether an epigenetic modification (e.g., a histone modification or methylation modification) at the regions is indicative (e.g., predictive) of cancer (e.g., correlates with cTF), e.g., that any of the candidate regions may be signal regions. In some embodiments, candidate regions in a genome are first selected to then determine whether any of them is a signal region where histone modification of the signal region is indicative (e.g., predictive) of cancer (e.g., correlates with cTF). Candidate regions of a genome may be regions where histone modification may be indicative of cancer (e.g.,
[0080] Page 14 of 124
[0081] 13414664vlAttorney Docket No. : 2014191-0048
[0082] correlate with cTF). In some embodiments, candidate regions in a genome are first selected to then determine whether any of them is a signal region where methylation of the signal region is indicative (e.g., predictive) of cancer (e.g., correlates with cTF). Candidate regions of a genome may be regions where methylation may be indicative of cancer (e.g., correlate with cTF). Selecting candidate regions may be or include identifying the candidate regions. Candidate regions can be selected in different manners. For example, candidate regions may be selected as uniformly sized bins across a genome of interest (e.g., the human genome) or in specific target areas within a genome of interest (e.g., the human genome). Candidate regions may be selected as regions of dense CpG sites (cystine-guanine dinucleotides), referred to as CpG islands. Such CpG islands may be identified by applying a definition to genomic data. A commonly used definition for a CpG island is a region of at least 200 bp length, a CG percentage greater than 50%, and an observed-to-expected CpG ratio (e.g., # of CpGs / (# of C * # of G / sequence length)) greater than 60%. A reference source may be used to select CpG islands as candidate regions, such as, for example, University of California Santa Cruz (UCSC) publishes a library of CpG islands. Candidate regions may be selected by identification as a set of consensus regions found by combining enrichment peaks and identifying those that are differentially methylated between conditions. Candidate regions may be genomic loci, for example each independently having a length in a range of 10-10,000 bp. Information about differential methylation of CpG islands may be used to select candidate regions for histone modification-based cTF estimation, for example, in instances where histone modification is correlated (e.g., inversely correlated) with differential methylation.
[0083] [441 Site-based methods, for example that consider single CpG sites that may be hypo-or hypermethylated in cancer relative to healthy and correlate with cTF on a site-by-site basis, may perform poorer than region-based methods. For example, CpG site methylation patterns are often locally similar (i.e. adjacent CpG sites usually are hypo or hypermethylated together). The present disclosure recognizes that aggregating sites together into regions to consider a region as a whole can give better sensitivity.
[0084]
[0045] In some embodiments, selecting candidate regions includes (e.g., consists of) selecting a set of CpG islands (e.g., reference CpG islands), for example CpG islands that have been identified (e.g., named) by a reference source. In some embodiments, selecting candidate regions includes (e.g., consists of) selecting bins of uniform size within a genome. Such bins may Page 15 of 124
[0085] 13414664vlAttorney Docket No. : 2014191-0048
[0086] be overlapping or non-overlapping. Tn some embodiments, selecting candidate regions includes (e.g., consists of) determining regions having differential epigenetic modification (e.g., differential histone modification and / or methylation) between conditions (e.g., healthy state and cancer state). In some embodiments, selecting candidate regions includes (e.g., consists of) determining regions having differential histone modification between conditions (e.g., healthy state and cancer state). In some embodiments, selecting candidate regions includes (e.g., consists of) determining regions having differential methylation between conditions (e.g., healthy state and cancer state). Regions with differential epigenetic modifications may be determined using sequencing data for one or more healthy samples and sequencing data for one or more samples with detectable cTF (e.g., high cTF). The regions having differential methylation between conditions may be determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples [e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% cTF (e.g., as determined by an ichor method)] (e.g., derived from one or more cell lines). In some embodiments, samples with low cTF (e.g., <0.1% cTF) may be additionally used.
[0087] Training samples and sequencing data normalization
[0088]
[0046] Training samples at a range of tumor fractions may be used to produce a cTF estimation model and / or identify signal regions that are used to produce a cTF estimation model. In some embodiments, cell lines and solid tumor data can be considered high fraction (e.g., 100% cTF) (e.g., heterogeneity and / or sub-pure tissue effects are ignored). Samples from healthy subject (e.g., healthy plasma) can be considered low (e.g., 0% cTF). Other samples may have an intermediate cTF. Such a cTF may be known because it is assumed from an orthogonal method such as a VAF or CNV method (e.g. ichorCNA).
[0089]
[0047] For training purposes, higher fraction samples can be mixed with healthy samples in silico to make digital samples (e.g., in silica plasma (ISP) samples) having a predetermined (e.g., artificially lower fraction) known cTF that is expected based on a mixing ratio used. Sequencing data for digital samples (e.g., ISP samples) may be created by random sampling of sequencing data from the constituent samples used to make the digital sample. For example, an ISP sample may be made by randomly sampling 5 million fragments from a 50% ctDNA plasma and combining it with 5 million fragments from a healthy sample to form a new in silico sample Page 16 of 124
[0090] 13414664vlAttorney Docket No. : 2014191-0048
[0091] with a 25% known (expected) cTF. Such samples are further described elsewhere herein. TSP samples used for training may include samples over a range of simulated cTF, for example in a range of from 0.05% to 50%, from 0.05% to 40%, from 0.05% to 30%, from 0.05% to 20%, from 0.05% to 15%, from 0.05% to 10%, from 0.05% to 8%, from 0.05% to 6%, from 0.05% to 5%, from 0.05% to 4%, from 0.05% to 3%, from 0.05% to 2%, from 0.05% to 1.5%, from 0.05% to 1%, from 0.05% to 0.75%, from 0.05% to 0.5%, 0.1% to 50%, from 0.1% to 40%, from 0.1% to 30%, from 0.1% to 20%, from 0.1% to 15%, from 0.1% to 10%, from 0.1% to 8%, from 0.1% to 6%, from 0.1% to 5%, from 0.1% to 4%, from 0.1% to 3%, from 0.1% to 2%, from 0.1% to 1.5%, from 0.1% to 1%, from 0.1% to 0.75%, or from 0.1% to 0.5%.
[0092]
[0048] Different numbers of empirical and / or in silico samples may be used to produce a cTF estimation model. Healthy and known cancer samples may be used. In some embodiments, a number of empirical samples used to produce a cTF estimation model is in a range of from 25 to 10,000 samples (a higher or lower number outside this range could be used in some embodiments). In some embodiments, a number of healthy samples used to produce a cTF estimation model is in a range of from 10 to 5,000. In some embodiments, a number of known cancer samples used to produce a cTF estimation model is in a range of from 10 to 5,000. In some embodiments, a number of in silico samples used to produce a cTF estimation model is in a range of from 10 to 10,000. Larger number of samples may mitigate effects of sample quality, for example where individual estimated known cTFs for samples determined by an orthogonal method may be imprecise and / or where sequencing data for the samples may have anomalous counts.
[0093]
[0049] Counts of sequencing data in candidate regions for training samples may be used in order to identify signal regions. Counts may be normalized to account for the sequencing method used, for example enrichment sequencing methods such as MBD-seq will produce counts if a region is methylated instead of on a site by site basis like other sequencing methods such as bisulfite sequencing methods. Other region-based enrichment sequencing methods, such as ChlP-seq, may similarly produce differentially modified counts across a region of nucleotides that has been bound to a histone. Where a per-region-based sequencing method that produces counts based on whether a region is modified (e.g., bound by a modified histone and / or methylated) or not is used, regions of larger size will tend to produce higher counts as compared to shorter regions as will more fragments being counted. In per-site sequencing methods, such as bisulfite methods, methylation proportion is typically reported instead and does not use the same or similar Page 17 of 124
[0094] 13414664vlAttorney Docket No. : 2014191-0048
[0095] normalizations (e.g., does not normalize by region length). Therefore, in some embodiments, normalizing counts for sequencing data includes, for each candidate region for each sample, normalizing the counts based on number of fragments in the sequencing data for the sample and based on length of region (a double normalization). In some embodiments, counts for a candidate region correspond to (e.g., are) a number of fragments in sequencing data that overlap (e.g., wholly or partially) with the candidate region. In some embodiments, counts are per-region-based counts (e.g., ChlP-seq counts) that indicate that a region of a sample has been bound to a modified histone. In some embodiments, counts are per-region-based counts (e.g., MBD-seq counts) (e.g., not persite counts) that indicate that a region of a sample is methylated. In some embodiments, normalizing further includes accounting for copy number alterations reflected in the sequencing data (e.g., using shallow whole genome sequencing data).
[0096]
[0050] In some embodiments, per sample and / or per candidate region counts are summed. These may then be normalized to the size of the region and the library size of the sample as defined by:
[0097] / 1000 * (cir+ 1)\
[0098] logCPKMi r= log - - Z
[0099]
[0100] \ ly * ZjkD /
[0101] where ci ris unique counts for sample i and region r, LS is the library size of the sample as defined by the number of unique fragments sequenced, and lris the length of the candidate region. The addition of 1 to the count number a,ravoids errors that may otherwise intrinsically results subsequently, for example due to infinite or null values when performing a logarithmic operation. Thus, counts used (e.g., for normalization) may be pseudocounts [e.g., offset true counts that are offset (e.g., by uniform addition) in order to avoid errors subsequent in model fitting and / or normalization]. Thus, logCPKM corresponds to the rate of captured fragments per base pair and fragment sequenced. In some embodiments, normalizing includes performing a normalization according to:
[0102] (f * (cir + d}\
[0103] logCPKMir= log ^l’r J
[0104]
[0105] where ci ris unique counts for sample i and region r, d is a constant (e.g., a non-zero constant, e.g., 1), / is a non-zero constant (e.g., 1,000), LS is sample library size as defined by the number of unique fragments sequenced, and lris region length for region r.
[0106] Page 18 of 124
[0107] 13414664vlAttorney Docket No. : 2014191-0048
[0108] Signal Regions
[0109]
[0051] Signal regions that are predictive of cTF may be determined from a set of candidate regions. In some embodiments, signal regions are determined as those where normalized counts correlate with known cTF for a set of samples. In some embodiments, in training samples, logCPKM values per sample are correlated with the known cTF of the samples. Differentially methylated regions (DMRs) that have a strong correlation with cTF and where hypo or hypermethylation is still observable down to a target fraction (e.g., down to 1% cTF) may be selected as signal regions. In some embodiments, regions with differential histone modification (DHMRs), low cTF regions (e.g., down to 1% cTF, e.g., down to 0.1% cTF) may be selected as candidate regions if they have a strong correlation with cTF (e.g., as determined by a correlation coefficient). In some embodiments, non-zero cTF samples should still have higher or lower counts than a pre-specified quantile of healthy samples. For example, a hypermethylated site may have more normalized counts in a region than the 95thpercentile of healthy sample counts for the same region. In another example, a cancer-correlated histone modification may have more normalized counts in a region than the 95thpercentile of healthy sample counts for the same region. Differential regions with epigenetic modifications, e.g., DHMRs and or / DMRs are determined based on cutoffs for significance, such as correlation p-values or limit of detection above healthy signal. For each region, there may also be linear or nonlinear function parameters (e.g., slope, intercept) that are kept for downstream modeling. In some embodiments, a range of 25-100 signal regions may be determined. Signal regions may be specific to a particular cancer type or subtype or may be pan-cancer regions that are signal regions for a variety of cancer types. Signal regions may correspond to (e.g., are, contain, or overlap with) CpG islands in a genome, for example where candidate regions are selected as CpG islands or where bins of uniform size are used and observed differential methylation and / or histone modification is more likely to occur in bins that correspond with CpG islands. Signal regions may be genomic loci, for example each independently having a length in a range of 10-10,000 bp.
[0110]
[0052] In some embodiments, signal regions are determined based on normalized counts for a candidate region having a linear relationship with cTF and the normalized counts for the candidate region for samples at a threshold cTF (e.g., a low threshold) being higher than an upper quantile or lower than a lower quantile for healthy samples. For example, normalized counts for the candidate region for the samples at the threshold cTF being higher than an upper quantile for Page 19 of 124
[0111] 13414664vlAttorney Docket No. : 2014191-0048
[0112] healthy samples with the upper quantile in a range of 85%-99% (e.g., is 90%, 95%, 97%, or 99%) may be used, in part, to determine signal regions. The normalized counts for the candidate region for the samples at the threshold cTF being lower than a lower quantile for healthy samples with the lower quantile being in a range of 1%- 15% (e.g., is 10%, 5%, 3%, or 1%) may be used, in part, to determine signal regions. Alternatively or additionally, in some embodiments, a candidate region is determined to be a signal region based on a minimum number (e.g., at least 80%, at least 85%, at least 90%, or at least 95%) of cancer samples having a normalized counts that are greater than an upper quantile in a range of 85%-99% (e.g., is 90%, 95%, 97%, or 99%) for healthy samples.
[0113]
[0053] In some embodiments, determining signal regions includes determining that ones of candidate regions comprise differential histone modification (e.g., have been bound to a modified histone to a greater or lesser extent than the same region in a healthy sample) over a range of cTF (e.g., at range of at least from 3% to 30% cTF, at least from 2% to 30% cTF, at least from 1% to 30% cTF, at least from 0.5% to 30% cTF). In some embodiments, determining signal regions includes determining that one or more candidate regions comprise differential histone modification (e.g., have been bound to a modified histone to a greater or lesser extent than the same region in a healthy sample) observable down to a target cTF in a range of from 0.5% to 3% cTF (e.g., no more than 2% cTF, no more than 1% cTF, no more than 0.5% cTF, or in a range of from 0.5% to 1% cT).
[0114]
[0054] In some embodiments, determining signal regions includes determining that ones of candidate regions are differentially (e.g., hypo- or hyper-) methylated over a range of cTF (e.g., at range of at least from 3% to 30% cTF, at least from 2% to 30% cTF, at least from 1% to 30% cTF, at least from 0.5% to 30% cTF). In some embodiments, determining signal regions includes determining that one or more candidate regions have differential (e.g., hypo- or hyper-) methylation observable down to a target cTF in a range of from 0.5% to 3% cTF (e.g., no more than 2% cTF, no more than 1% cTF, no more than 0.5% cTF, or in a range of from 0.5% to 1% cT).
[0115]
[0055] In some embodiments, determining signal regions includes determining (e.g., for each of the candidate regions) that there exists a linear relationship between normalized counts for a candidate region and cTF. In some embodiments, determining signal regions includes, for each
[0116] Page 20 of 124
[0117] 13414664vlAttorney Docket No. : 2014191-0048
[0118] of a set of candidate regions, determining a relationship between normalized counts in sequencing for the candidate region and the known cTF for a set of samples.
[0119] [561 In some embodiments, signal regions are cancer-specific signal regions that correspond either to a particular cancer type or a particular cancer sub-type. For a pan-cancer cTF estimation model, determining signal regions may include determining cancer-specific signal regions using sequencing data for samples of different cancer types separately (on a “per-cancer” basis) and then determining one or more signal regions based on regions that are determined to be cancer-specific signal regions for a minimum number of cancer types or subtypes, for example at least two, at least three, at least four, or at least five cancer types or subtypes and / or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or all cancer types or subtypes represented by the range of samples used to determine the cancer-specific signal regions. In this way, genomic regions that are differentially modified (e.g., differentially methylated, e.g., comprise differential histone modifications) regions for a range of cancers may be determined and used as signal regions for producing a cTF estimation model. A pan-cancer cTF estimation model that is produced based on using all such cancer-specific signal regions identified as signal regions may perform poorer, for example due to effects resulting from only a small subset of signal regions having counts above background in sequencing data used from a sample for estimating cTF. For example, in a pancancer cTF estimation model where only 10% of signal regions are signal regions specific to cancer type A, a cTF estimate of a sample from a subject having cancer type A may be inaccurate because 90% of signal regions may have normalized counts indistinguishable from background and / or indistinguishable from healthy (e.g., cTF of 0) samples. In some embodiments, regions that show differential modification (e.g., methylation and / or histone modification) across only one, or a subset of desired, cancer type or subtype may be excluded from a determined set of signal regions and not used in producing a cTF estimation model. In some embodiments, a pan-cancer model includes a plurality of cancer-specific constituent models and cTF estimated by the model corresponds to a preliminary estimated cTF identified by one of the constituent models (e.g., a maximum of the preliminary estimated cTFs of the constituent models).
[0120] Noise Regions
[0121] [57J In some embodiments, there may be noise regions of a genome that capture technical noise in sample processing and / or batch effects. Such noise and / or effects may affect
[0122] Page 21 of 124
[0123] 13414664vlAttorney Docket No. : 2014191-0048
[0124] (e g., skew) estimates of cTF from a model. Noise regions may be determined from regions of a genome that are not signal regions. In some embodiments, candidate regions that are not determined to be signal regions are tested to determine whether they are noise regions. In some embodiments, noise regions are determined from candidate regions that are determined to not be signal regions. In a noise region, a metric used to determine a signal region, for example normalized counts or a point estimate such as a median normalized count per sample, may show a correlation in healthy samples but not in cancer samples (having non-zero cTF). One would expect there not to be a correlation in the healthy samples in such cases, regardless of whether there was a correlation in cancer samples. A correlation in healthy samples is evidence of technical noise in sample processing and / or batch effects. A method may include determining noise regions and using the noise regions in producing a cTF estimation model. In some embodiments, a method includes determining noise regions in a genome. In some embodiments, a method includes determining noise regions in a genome that indicate correlation only in healthy samples of a set of samples. Correlation may be determined where, for example, correlation for a metric (e.g., normalized counts or a metric derived therefrom) in a region in healthy samples has a significance with a p value of <0.001 and insignificance with a p value of >0.005 for cancer samples and, optionally, does not have a batch effect in healthy samples with a p value of >0.1. In some embodiments, a method includes determining noise regions in a genome where there is correlation between normalized counts in healthy samples and no correlation between normalized counts in cancerous samples. In some embodiments, a method includes determining noise regions in a genome that are indicative of batch effects in sequencing data, before or after normalization. Determining noise regions may include, for each noise region, determining that a metric (e.g., normalized counts) for samples having a known cTF at or below a low threshold (e.g., of 0) has a positive correlation and that the metric for samples having a known cTF above a threshold is uncorrelated.
[0125]
[0058] Sequencing data for noise regions, once the noise regions have been determined, may be used to produce a cTF estimation model. For example, a metric may be calculated using sequencing data for noise regions and that metric may be used in a cTF estimation model. In some embodiments, a signal-to-noise ratio (SNR) for each of a set of samples is determined. The signal-to-noise ratio may be based on a point estimate (e g., median) count value across each of the noise regions for a sample and a point estimate (e.g., median) count value across each of the signal Page 22 of 124
[0126] 13414664vlAttorney Docket No. : 2014191-0048
[0127] regions for the sample. The count values may be normalized counts normalized using normalization described elsewhere with respect to finding signal regions. A cTF estimation model (e.g., a constituent model thereof) may be produced based on determining a relationship between the signal-to-noise ratio and the known cTF for each of the samples. An SNR may be determined for a sample based on normalized counts for noise regions and normalized counts for signal regions. Producing normalized counts for noise regions for samples used to produce a cTF estimation model may include, for each noise region, normalizing counts for the noise region based on number of fragments in the sequencing data for the sample and based on length of the noise region. Producing the normalized counts for noise regions may include, for each noise region, performing a normalization according to
[0128] f * \Ci>r+ d)
[0129] logCPKMr= log
[0130]
[0131] lr* LS
[0132] where ciris unique counts in the sequencing data for sample z in region r, d is a constant (e.g., a non-zero constant, e.g., l), / is a non-zero constant (e.g., 1000), LS is sample library size as defined by the number of unique fragments sequenced for the sample, and lris region length for region r. In some embodiments, SNR for samples used to produce a cTF estimation model is determined as a ratio of a point estimate of normalized counts for signal regions and a point estimate of normalized counts for noise regions, preferably where the point estimates are the median.
[0133] cTF Estimation Models Structure and Production
[0134]
[0059] A cTF estimation model can determine an estimated cTF in a sample based on input sequencing data or data derived therefrom. Data derived from sequencing data may include one or more metrics. Data derived from sequencing data may include one or more point estimates. Producing a cTF estimation model may be performed using at least normalized counts for signal regions and known cTF for samples. Known cTF may be empirically known, assumed, or expected, for example depending on whether a sample is a healthy sample (from a subject known to not have cancer), a sample with an independently empirically known or determined cTF, or an in silico diluted sample. Producing a cTF estimation model may including training the model using samples, from subjects (e.g., via biopsy) (e.g., including one or more cell lines) and / or in silico dilutions.
[0135] Page 23 of 124
[0136] 13414664vlAttorney Docket No. : 2014191-0048
[0137]
[0060] A cTF estimation model may be produced to accept input that includes one or more metrics. One or more metrics may include one or more metrics calculated from sequencing data, for example from normalized counts of sequencing data for signal regions and / or noise regions. For example, a cTF estimation model may be produced that uses input that includes an SNR metric calculated using normalized counts for signal regions and noise regions. A cTF estimation model may be produced to accept input that includes one or more features. One or more features may include one or more features determined from sequencing data, for example normalized counts of sequencing data for signal regions and / or noise regions. For example, a cTF estimation model may be produced that uses input that includes an SNR metric determined using normalized counts for signal regions and noise regions.
[0138]
[0061] Different model types may be produced. Producing a cTF estimation model may include producing a plurality of preliminary models, assessing the preliminary models for performance, and producing the cTF estimation model from one or more of the preliminary models that are assessed to have appropriate (e.g., sufficient and / or good) performance. One or more model types may behave better in one or more certain cancer types or subtypes than one or more others. A regression model may be produced multivariate, regularized, logistic, and or tree-based regression can be used where a feature matrix has n samples and m features that are used to fit a model to predict they response variable for cTF. Feature aggregation and regression can be used to produce a model. Features can be combined by taking sum, mean (or other average), median, or other point estimate per sample. These values can then be used to fit a model against cTF. As discussed above, noise regions may be used to estimate noise in data and a noise-based metric, such as SNR, may be used to fit a model. An expectation maximization (EM) model may be used where a relationship is defined between normalized counts and cTF for each signal region independently and then the model can be used to perform an expectation maximization that minimizes a global loss function for expected normalized counts across all signal regions considered independently simultaneously in order to estimate cTF. That is, in some embodiments, an expectation maximization model is produced that uses a difference between expected number of normalized counts in each signal region and observed counts in a sample can be determined as a loss for the regions and sums the loss across all signal regions to estimate cTF as the cTF with a minimum loss. Deep learning models may also be produced where features can be used as tensors to train a neural net model to predict cTF. Features used in a model can be normalized counts or Page 24 of 124
[0139] 13414664vlAttorney Docket No. : 2014191-0048
[0140] another metric as discussed elsewhere herein. Tn some embodiments, producing a cTF estimation model includes selecting a type of model based on cancer type for samples used to produce (e.g., fit and / or train) the model.
[0141]
[0062] Producing a cTF estimation model may include fitting a model based on normalized counts for sequencing data and known cTF. Fitting may include a linear optimization and / or nonlinear optimization. Fitting may include making a linear regression and / or non-linear regression. In some embodiments, producing a model is based on, for each of a plurality of samples, a point estimate of normalized counts for signal regions for the sample, preferably wherein the point estimate is the median.
[0142]
[0063] In some embodiments, producing a model includes producing an expectation maximization model. In some embodiments, an expectation maximalization model includes a set of relationships (e g., regressions) between signal regions and cTF across which an expectation maximization can be performed (e.g., on a per-region basis). In some embodiments, producing a model includes producing (e.g., defining and / or fitting) a set of relationships (e.g., regressions) between signal regions and cTF across which an expectation maximization can be performed (e.g., on a per-region basis). Such a model may consider signal regions independently (e g., through independent relationships for each signal region). An expectation maximalization model may be produced to maximize expectation across (e.g., all) signal regions simultaneously. In some embodiments, producing a model includes fitting a cTF regression for each signal region. In some embodiments, producing a model includes defining a global loss function for a set of relationships (e.g., regressions) between cTF and normalized counts for a signal region, for each of a set of signal regions (e.g., each signal region independently having a corresponding relationship). In some embodiments, a model considers relationships between cTF and normalized counts for a signal region for each signal region independently. In some embodiments, a model considers relationships between cTF and normalized counts for a signal region for each signal region simultaneously. In some embodiments, producing a cTF estimation model includes independently fitting signal regions (e.g., in an expectation maximization model). In some embodiments, a cTF estimation model determines estimated cTF based on independent consideration of normalized counts for the signal regions in a sample. In some embodiments, a cTF estimation model estimates a per-region-based estimated cTF for each signal region individually and then determines an estimated cTF based on the per-region-based estimated cTF.
[0143] Page 25 of 124
[0144] 13414664vlAttorney Docket No. : 2014191-0048
[0145]
[0064] In some embodiments, producing a cTF estimation model includes a multivariate, regularized, logistic, and / or tree-based regression (e.g., that uses a feature matrix). In some embodiments, producing a cTF estimation model includes producing a feature aggregation and regression model that uses a per-sample-based point estimate (e.g., sum, mean, or median) (e.g., for signal regions, optionally and noise regions). In some embodiments, producing a cTF estimation model includes producing a deep learning model (e.g., an artificial neural network) (e.g., trained on normalized sequencing data for signal regions, optionally and noise regions). Producing a cTF estimation model may include training the deep learning model using tensors derived from normalized sequencing data for signal regions.
[0146]
[0065] In some embodiments, a cTF estimation model is an ensemble model. Ensemble models may have better performance at least in some cancer types or subtypes. An ensemble method may include a plurality of constituent models. Such constituent models may be structured to produce preliminary estimated cTF estimates that are then combined to produce a cTF estimate based on given input. A combination may be a point estimate of the preliminary estimated cTFs, such as, for example, the median or a mean (e.g., arithmetic or geometric mean). Such constituent models may include models of different types, for example including a (e.g., linear or non-linear) regression model and an expectation maximalization model or a deep learning model and an expectation maximalization model. Producing a cTF estimation model may include producing constituent models and determining a combination for preliminary estimated cTF from the constituent models. In some embodiments, a cTF estimation model includes an ensemble model (e.g., including two, three, or more constituent models). In some embodiments, producing a cTF estimation model includes producing a plurality of constituent models and determining a set (e.g., subset) of the plurality of constituent models to include in an ensemble model (e.g., including two, three, or more constituent models). In some embodiments, an ensemble model uses a geometric mean of constituent models to estimate cTF.
[0147]
[0066] An ensemble model may include constituent models for different epigenetic modifications, such as a first constituent model for DNA methylation and a second constituent model for histone modification. An ensemble model may include both multiple constituent models for an epigenetic modification (for each of one or more epigenetic modifications) and constituent models for different epigenetic modifications, for example may include one or more DNA
[0148] Page 26 of 124
[0149] 13414664vlAttorney Docket No. : 2014191-0048
[0150] methylation based constituent models and one or more histone modification based constituent models for each of one or more particular histone modifications.
[0151] [671 In some embodiments, a cTF estimation model combines a plurality of preliminary estimated cTFs, each determined by a separate constituent model of the cTF estimation model, to estimate cTF. In some embodiments, combining includes determining a geometric mean of the preliminary estimated cTFs. In some embodiments, the preliminary estimated cTFs include a per-region-based estimated cTF and a per-sample-based estimated cTF. In some embodiments, a constituent model is a per-sample-based model, for example a point estimate based model. In some embodiments, a constituent model is a per-region-based model, for example an expectation maximization model. In some embodiments, a constituent model is an SNR-based model. In some embodiments, a cTF estimation model estimates a plurality of preliminary estimated cTFs comprising a per-region-based estimated cTF estimated using a first model of the cTF estimation model and a per-sample-based estimated cTF estimated using a second model of the cTF estimation model, wherein the cTF is estimated by a combination (e.g., geometric mean) that includes (e.g., is of) the per-region-based estimated cTF and the per-sample-based estimated cTF.
[0152]
[0068] In some embodiments, an ensemble model includes (e.g., a plurality of constituent models includes) a first model that estimates cTF based on each signal region together (e.g., a point estimate regression model) and a second model that estimates cTF based on each of the signal regions individually (e.g., an expectation maximization model). In some embodiments, an ensemble model includes (e.g., a plurality of constituent models includes) a first model that estimates cTF based on a metric determined from each signal region together [e.g., a point estimate (e.g., median) of normalized counts for all of the signal regions] and a second model that estimates cTF based on individual metrics for each of the signal regions (e.g., an expectation maximization model). In some embodiments, an ensemble model includes an expectation maximization model and (i) a signal-to-noise ratio (SNR) based model (e.g., an SNR regression model) and / or (ii) a point estimate regression based model. In some embodiments, producing an ensemble model includes producing an expectation maximization model and (i) a signal-to-noise ratio (SNR) based model (e.g., an SNR regression model) and / or (ii) a point estimate regression based model. In some embodiments, producing a model includes producing a pan-cancer ensemble model comprising at least one constituent model for a plurality of cancer types (e.g., wherein each of the at least one constituent models includes an ensemble model).
[0153] Page 27 of 124
[0154] 13414664vlAttorney Docket No. : 2014191-0048
[0155]
[0069] An ensemble model may include a first constituent model that is a logistic regression based model (e.g., an L2 based logistic regression based model) and a second constituent model that is a histogram-based gradient boosting regression tree (HGBR) based model. An ensemble model may include a first constituent model, for example logistic regression model, that has been trained on healthy samples, low cTF samples (e g., <1%), for example as determined by physical measurement or a cTF estimation model, and digital samples (e.g., ISP samples) (e.g., having a cTF in a range of from 0.1% to 0.5%) and a second model, for example an HGBR based model that has been trained on digital samples (e.g., ISP samples). Such an ensemble model may produce a cTF estimation of 0 if ctDNA is not detected, for example using the first constituent model, and otherwise use a cTF estimate of an HGBR based model as the cTF estimate.
[0156]
[0070] In some embodiments, a cTF estimation model is an ensemble model that includes a first constituent model for detecting whether ctDNA is present in a sample or not and a second constituent model for estimating cTF. In this way, models that are more performant at ctDNA detection can be used for detection and models that are more performant at estimation can be used for estimation. Such a first constituent model may be a logistic regression based model. Such a second constituent model may be a HGBR based model.
[0157]
[0071] In some embodiments, a cTF estimation model determines a point estimate of normalized counts for signal regions and uses the point estimate to estimate the cTF for a sample (e.g., in one or more constituent models of the cTF estimation model), preferably where the point estimate is the median.
[0158]
[0072] In some embodiments, producing a cTF estimation model includes producing a first model that detects if ctDNA is present in a sample and a second model that estimates cTF if ctDNA is present as determined using the first model. In some embodiments, producing a cTF estimation model includes producing a first model to detect that ctDNA is present in a sample and a second model that estimates cTF. In some embodiments, producing a cTF estimation model includes producing a first model includes an ensemble model that uses a geometric mean of preliminary estimated cTFs from a set of constituent models. In some embodiments, a first model is a signal-to-noise ratio (SNR) based model and a second model is an expectation maximalization model. In some embodiments, producing a cTF estimation model includes producing a first model that estimates a preliminary estimated cTF and a second model that estimates cTF depending on the Page 28 of 124
[0159] 13414664vlAttorney Docket No. : 2014191-0048
[0160] preliminary estimated cTF. For example, in some embodiments, producing a cTF estimation includes producing one model to estimate cTF if a preliminary estimated cTF is above a threshold and another model to estimate cTF if preliminary estimated cTF is below a threshold.
[0161]
[0073] A cTF estimation model may use input data corresponding to a first epigenetic modification while having been trained on data generated based on a second epigenetic modification. For example, estimates of cTF for training samples may have been generated using a DNA methylation based cTF estimation model and then corresponding histone modification data for the training samples and the cTF estimates from the DNA methylation based cTF estimation model may be used to train a histone modification based cTF estimation model. As another example, estimates of cTF for training samples may have been generated using a histone modification based cTF estimation model and then corresponding DNA methylation data for the training samples and the cTF estimates from the histone modification based cTF estimation model may be used to train a DNA methylation based cTF estimation model. Such training data may be used in combination with healthy training samples and / or digital samples (e.g., ISP samples). A digital sample (e.g., ISP sample) may be a sample for which the cTF estimate has been produced using a cTF estimation model (e.g., a DNA methylation based cTF estimation model or histone modification based cTF estimation model).
[0162] Estimating cTF
[0163]
[0074] Once a model has been produced, it may be used to estimate cTF of new samples from subjects, for example liquid plasma samples or other liquid biopsy samples or samples derived therefrom. Cancer-specific or pan-cancer models may be used to estimate cTF. In some embodiments, a model is a pan-cancer model, for example having been produced using data from samples from different subjects having different cancers, and an estimate of cTF may be indicative of whether a subject has cancer. In some embodiments, a model is a cancer-specific model and an estimate of cTF using the model may be indicative of whether a subject has a particular cancer. Thus, a particular model may be chosen for use to estimate cTF based on a cancer that a subject is suspected to have, has, or has had in the past.
[0164]
[0075] To estimate cTF using a cTF estimation model, appropriate sequencing data is provided and / or received, for example at a processor of a computing device. In some embodiments, a method includes obtaining the sequencing data. The sequencing data can be Page 29 of 124
[0165] 13414664vlAttorney Docket No. : 2014191-0048
[0166] processed to obtain appropriate input for a cTF estimation model and estimated cTF can be output from the model. Processing sequencing data may include, for example, producing normalized counts for signal regions in a subject’s genome that are used in a cTF estimation model and / or producing normalized counts for noise regions in a subject’s genome that are used in the model. Further processing may occur after normalization as part of estimating cTF, for example normalized counts for signal regions and / or normalized counts for noise regions may be converted to one or more metrics that are used in the model. Examples of such metrics include point estimate metrics, such as median normalized counts for all signal regions or an SNR. Such metrics may be per-sample metrics, like the aforementioned exemplary point estimate metrics, or per-region metrics, such as those used in an expectation maximization model.
[0167]
[0076] In some embodiments, a (e.g., computer-implemented) method of estimating cTF in a sample includes receiving (e.g., obtaining) sequencing data for a liquid sample from a subject. The method may further include producing normalized counts of the sequencing data for signal regions in the genome of the subject. The method may further include estimating cTF for the sample with a cTF estimation model based at least on the normalized counts (e.g., by first processing the normalized counts into one or more metrics and then obtaining an estimated cTF using the one or more metrics). The sequencing data may have been derived from an enrichment method (e.g., a histone modification enrichment method, e.g., a methylation enrichment method). The sequencing data may be chromatin immunoprecipitation sequencing (ChlP-seq) data. The sequencing data may be methyl-binding domain sequencing (MBD-seq) data. The signal regions may be regions determined to be cancer-specific regions with differential histone modifications ns (cDHMRs). The signal regions may be regions determined to be cancer-specific differentially methylated regions (cDMRs).
[0168]
[0077] Producing normalized counts for a sample may include, for each signal region, normalizing the counts based on number of fragments in the sequencing data for the sample and based on length of the signal region. Producing normalized counts for a sample may include, for each signal region, may include performing a normalization according to
[0169] (f * (cr+ d)\
[0170] logCPKMr= log
[0171]
[0172] J
[0173] where cris unique counts in the sequencing data for the sample in region r, d is a constant (e.g., a non-zero constant, e.g., 1), / is anon-zero constant (e.g., 1000), LS is sample library size as defined Page 30 of 124
[0174] 13414664vlAttorney Docket No. : 2014191-0048
[0175] by the number of unique fragments sequenced for the sample, and lris region length for region r. In some embodiments, for each of signal region, the counts for the signal region in the sequencing data are a number of fragments in the sequencing data that overlap with signal region.
[0176]
[0078] In some embodiments, estimating cTF for a sample includes includes estimating a per-region-based estimated cTF for each signal region individually (e.g., and simultaneously) and then determining the estimated cTF based on the per-region-based estimated cTF using the cTF estimation model. Such a method may be achieved using an expectation maximization model as or in a cTF estimation model. In some embodiments, estimating cTF includes performing an expectation maximization of cTF for a sample with a cTF estimation model based on normalized counts. In some embodiments, cTF is estimated with an expectation maximization model that has been trained to maximize expectation on a per-region basis.
[0177]
[0079] In some embodiments, estimating cTF includes determining a point estimate of normalized counts for signal regions for a sample and a cTF estimation model uses the point estimate to estimate the cTF for the sample. The point estimate may be the median or a different metric, for example an SNR for a sample. A point estimate may be used in one or more constituent models of a cTF estimation model.
[0178]
[0080] Estimating cTF for a sample may include estimating a plurality of preliminary estimated cTFs and then combining the preliminary estimated cTFs together. A plurality of preliminary estimated cTFs may include a per-region-based estimated cTF estimated using a first model of a cTF estimation model and a per-sample-based estimated cTF estimated using a second model of the cTF estimation model. cTF may be estimated by a combination that includes (e.g., is of) a per-region-based estimated cTF and a per-sample-based estimated cTF. For example, an expectation maximization constituent model may be used to generate a per-region-based estimated cTF and a point estimate based model may be used to generate a per-sample-based estimated cTF and these may be combined together, optionally with one or more additional preliminary estimated cTFs, to estimate cTF for a sample. Different combinations of preliminary estimated cTFs produced by a model may be used, for example a mean (e.g., arithmetic mean or geometric mean) or the median may be used. It has been found that, in certain embodiments, geometric mean is preferred because it produces more accurate estimates.
[0179]
[0081] Sample sequencing data for noise regions in a subject’s genome may be used with a cTF estimation model to estimate cTF for the sample. For example, if a cTF estimation model Page 31 of 124
[0180] 13414664vlAttorney Docket No. : 2014191-0048
[0181] uses (e.g., in a constituent model) SNR to estimate cTF, then sequencing data for noise regions for a sample will be used in order to calculate the SNR for the sample and estimate (e.g., preliminary) cTF. In some embodiments, a method includes producing normalized counts of sequencing data for noise regions in the genome of the subject. Such normalization may occur in a similar manner to normalization for sequencing data in signal regions. For example, in some embodiments, producing normalized counts for noise regions includes, for each of the noise regions, normalizing counts for the noise region based on number of fragments in the sequencing data for the sample and based on length of the noise region. Producing normalized counts for the noise regions may include, for each of the noise regions, performing a normalization according to
[0182] ( f * (cr+ d)\
[0183] logCPKMr= log , "
[0184]
[0185] y tr* LS y
[0186] where cris unique counts in the sequencing data for the sample in region r, d is a constant (e.g., a non-zero constant, e.g., 1), / is anon-zero constant (e.g., 1000), LS is sample library size as defined by the number of unique fragments sequenced for the sample, and lris region length for region r. In some embodiments, a method includes determining a signal-to-noise ratio (SNR) for a sample based on normalized counts for the noise regions and normalized counts for signal regions. A cTF estimation model may use SNR to estimate the cTF, for example may produce an SNR-based preliminary estimated cTF that is combined with one or more other preliminary estimated cTFs (e.g., a per-region-based estimated cTF and / or per-sample-based estimated cTF) to estimate the cTF for the sample. Producing normalized counts for noise regions may include, for each of the noise regions, normalizing counts for the noise region based on number of fragments in the sequencing data for the sample and based on length of the noise region. SNR may be determined as a ratio of a point estimate of normalized counts for signal regions and a point estimate of normalized counts for noise regions, for example where the point estimates are the respective medians.
[0187]
[0082] A cTF estimation model used to estimate cTF for samples may include or be an ensemble model. In some embodiments, a cTF estimation model produces and combines a plurality of preliminary estimated cTFs, each determined by a separate constituent model of the cTF estimation model, to estimate the cTF. The combining may include determining a geometric mean of the preliminary estimated cTFs. The preliminary estimated cTFs may include a per-region-based estimated cTF and a per-sample-based estimated cTF. The per-sample-based Page 32 of 124
[0188] 13414664vlAttorney Docket No. : 2014191-0048
[0189] estimated cTF may be determined using a first constituent model that is a point estimate based model. The per-region-based estimated cTF may be determined using an expectation maximization model. The preliminary estimated cTFs may include an SNR-based preliminary estimated cTF.
[0190]
[0083] Different models (e.g., constituent models) may perform differently either for different cancer types. Some cTF estimation models are better at detecting ctDNA in a sample while others have higher numerical accuracy for estimating cTF. Which model may be better at detecting and which may have higher numerical accuracy may depend on for which cancer type(s), or subtype, a model is trained and / or used. For example, it has been found that for certain lung and breast cancer models that an expectation maximization model has higher numerical accuracy for estimating cTF whereas an ensemble cTF estimation model, for example based on a geometric mean of preliminary estimated cTFs from an expectation maximization model, an SNR-based model, and a model based on a regression of median normalized counts across all signal regions (as a per-sample-based point estimate), is better at detecting ctDNA in a sample, especially near the limit of detection. Thus, a cTF estimation model may include a first model that is used to detect whether a sample has a non-zero cTF (has ctDNA in it as reflected in sequencing data) and a second model that, if so detected, is used to produce an estimated cTF for a sample. As another example, it has been found for certain prostate cancer models that an SNR model is better at detecting a non-zero cTF while a model based on a regression of median normalized counts across all signal regions (as a per-sample-based point estimate) has higher numerical accuracy for estimating cTF at least for certain ranges of cTF. Numerical accuracy of a model may depend on what range of cTF is being estimated. For example, referring again to certain prostate cancer models, it has been found that an SNR-based model has better numerical accuracy at low cTF (e.g., at or below 3%) and a model based on a regression of median normalized counts across all signal regions (as a per-sample-based point estimate) at high cTF (e.g., above 3%). Accordingly, estimating cTF of a sample using a cTF estimation model may include detecting whether ctDNA is present in the sample with a first model of the estimation model and then, depending on a preliminary estimated cTF estimate, using one of a plurality of second models of the estimation model to estimate cTF.
[0191]
[0084] In some embodiments, estimating cTF includes using a first model of a cTF estimation model to detect if ctDNA is present in a sample and subsequently using a second model Page 33 of 124
[0192] 13414664vlAttorney Docket No. : 2014191-0048
[0193] of the cTF estimation model to estimate the cTF if ctDNA is present or estimating the cTF at 0 if ctDNA is determined to be not present using the first model. In some embodiments, estimating cTF includes using a first model of a cTF estimation model to detect that ctDNA is present in the sample and, upon determining that ctDNA is present, using a second model of the cTF estimation model to estimate the cTF. In some embodiments, a first model includes an ensemble model that uses a geometric mean of preliminary estimated cTFs and a second model is an expectation maximalization model, for example in embodiments where the sample is of a subject suspected to have, that has, or that has had breast cancer or lung cancer, for example where the cTF estimation model is a breast cancer cTF estimation model or lung cancer cTF estimation model. In some embodiments, a first model is a signal-to-noise ratio (SNR) based model and the second model is an expectation maximalization model, for example in embodiments where the sample is of a subject suspected to have, that has, or that has had breast cancer or lung cancer, for example where the cTF estimation model is a prostate cancer cTF estimation model. In some embodiments, using a first model includes estimating a preliminary estimated cTF and the second model used to estimate the cTF depends on the preliminary estimated cTF. In some embodiments, estimating cTF with a cTF estimation model includes using one model to estimate cTF if a preliminary estimated cTF is above a threshold and using another model to estimate the cTF if the preliminary estimated cTF is below a threshold.
[0194]
[0085] In some embodiments, estimating cTF of a sample includes using a cTF estimation model that uses a multivariate, regularized, logistic, and / or tree-based regression (e.g., that uses a feature matrix, e.g., a histogram-based gradient boosting regression tree). In some embodiments, estimating cTF of a sample includes using a cTF estimation model that uses a histogram-based gradient boosting regression tree. In some embodiments, estimating cTF of a sample includes using a cTF estimation model that includes a feature aggregation and regression model that uses a per-sample-based point estimate (e.g., sum, mean, or median) (e.g., for signal regions). In some embodiments, estimating cTF of a sample includes using a cTF estimation model that includes a deep learning model (e.g., an artificial neural network) that has been trained using tensors derived from normalized sequencing data for the signal regions.
[0195]
[0086] In some embodiments, a cTF estimation model uses signal regions (e.g., normalized counts for the signal regions) independently to obtain estimated cTF. In some embodiments, a
[0196] Page 34 of 124
[0197] 13414664vlAttorney Docket No. : 2014191-0048
[0198] cTF estimation model determines estimated cTF based on independent consideration of normalized counts for signal regions in a sample.
[0199] [871 In some embodiments, a cTF estimation model has been trained using data from cancer-specific samples and the subject is suspected of having and / or known to have or have had the cancer to which the model is specific. In some embodiments, a cTF estimation model has been trained using data from samples having different cancers (e.g., wherein the model is a pan-cancer model).
[0200] Use of Estimated cTF
[0201]
[0088] An estimate of cTF in a sample provides useful information that may be used as or in a biomarker, as a control variable in a classifier, to characterize and / or understand performance of classifiers, or to aid in diagnosing, prognosing, and / or monitoring a subject (e g., over time), for example. A subject may be monitored to monitor treatment efficacy and / or cancer progression. Diagnostic, prognostic, and / or monitoring information may be used to initiate, alter, and / or cease treatment of a subject with one or more therapies.
[0202]
[0089] Estimated cTF may be used to diagnose, prognose, and / or monitor a subject. As one example, estimated cTF may co-vary with another parameter that is indicative of treatment efficacy such that monitoring estimated cTF over time can be used as a proxy for monitoring such other parameter (or parameters). It may be easier to estimate cTF from samples than to characterize other parameters, such that accurate estimation of cTF allows for easier monitoring than alternative methods. As another example, it is generally expected that cTF increases with cancer progression (as more cancer is present, more ctDNA is likely to be present in a liquid biopsy) and therefore cTF may be useful in prognosing cancer and / or monitoring cancer progression. An increase in estimated cTF or a characteristic (e.g., rate) of increase of estimated cTF may be indicative of changes in a cancer of a subject such as, for example, tumor volume and / or metastasis. Similarly, if a subject’s sample has a non-zero cTF, that likely suggests the subject has cancer and therefore estimated cTF may be used to diagnose a subject with cancer. Because cTF estimation models disclosed herein are usable at very low cTF (e.g., less than 3%, less than 2%, or less than 1%), earlier diagnosis and / or earlier stage prognosis may be achievable than with other methods. Moreover, such early diagnosis or prognosis may not be possible with
[0203] Page 35 of 124
[0204] 13414664vlAttorney Docket No. : 2014191-0048
[0205] tumor-informed methods which require a priori knowledge that a subject has a tumor (and, in some cases, what the nature of the tumor).
[0206]
[0090] Epigenetic modifications may not be uniform in all indications of a similar type. For example, extent and / or location of DNA methylation can vary based on cancer type and / or subtype. Similarly, type and number of histone modifications across the same genomic region (e.g., relative to a reference genome) may vary between cancers. Accordingly, comparing estimated cTF from different cTF estimation models for different cancer types (e.g., trained on different single-cancer-type sample sets) may yield useful information regarding which type of cancer a subject may have. For example, cTF may be estimated using a breast cancer cTF estimation model, a lung cancer cTF estimation model, and a prostate cancer cTF estimation model and then the estimates may be compared. A difference in estimated cTF may indicate that a subject is more likely to have one type of cancer rather than another. A subsequent diagnostic assessment may be made based on the difference. Accordingly, in some embodiments, a preliminary diagnostic assessment may be made using estimated cTF estimated using a method disclosed herein and a cancer-specific subsequent diagnostic assessment may be performed that depends on the estimated cTF may then be performed. Such diagnostic assessments can include invasive solid biopsy procedures whereas cTF may be estimated using a method disclosed herein from a liquid biopsy sample of small volume. Therefore, cTF can be estimated easily, even alongside other testing using the same liquid biopsy, before a more invasive test is ordered. By making an early preliminary diagnosis of cancer type or subtype using estimated cTF, invasive, expensive, and / or time consuming diagnostic assessments that may be for an incorrect type of cancer may be avoided or postponed. For example, estimated cTF from different cTF estimation models may provide an indication that a subject more likely has breast cancer than lung cancer and therefore an initial subsequent diagnostic assessment should be for breast cancer rather than lung cancer.
[0207]
[0091] A subject may be monitored by estimating cTF from different samples taken over time. Such samples may be taken before and / or after a cancer removal procedure (e.g., chemotherapy, immunotherapy, and / or removal surgery) to determine whether cancer has been successfully removed. Samples may be taken after sufficient passing of time to allow for residual ctDNA to be cleared from a subject in order to ensure accurate cTF estimation.
[0208]
[0092] Estimated cTF may be used as a control variable for a classifier, such as a cancer type or cancer subtype classifier. For example, a cancer type or subtype classifier may yield false Page 36 of 124
[0209] 13414664vlAttorney Docket No. : 2014191-0048
[0210] positives due to effects caused by different levels of ctDNA in samples used in the classifier. Estimated cTF may be used as a control variable to mitigate such effects. As another example, a cancer type or subtype classifier may yield a false positive where no ctDNA is actually present in a sample (or is present below a limit of detection). Accordingly, estimating cTF using a method disclosed herein may be used to catch a false positive for a classifier (e.g., a cancer type or cancer subtype classifier) or be used as a screen before classification. Estimated cTF may be used to quantify to what limit of cTF a classifier (e.g., a cancer type or cancer subtype classifier) is performant. As such classifiers may be important to defining patient population (e.g., for a therapy), understanding limits of performance, such as minimum cTF, may be important to accurately defining a patient population. (Likewise, a cTF estimation model as disclosed herein may be used to define an actual or prospective patient population.)
[0211]
[0093] In some embodiments, a method is for characterizing cancer recurrence and / or progression. Such a method may include estimating cTF in a first sample for a subject taken at a first time point using a method disclosed herein. Such a method may further include estimating cTF in a second sample for the subject taken at a second time point after the first time point using a method disclosed herein. A difference in the estimated cTF at the second time point and at the first time point may then be determined. The difference may be used to characterize cancer recurrence and / or progression. In some embodiments, a method is for monitoring cancer in a subject, which may include estimating cTF in a series of two or more samples for a subject, each taken at a different time point, using a method disclosed herein, and determining whether there is a difference in cTF for the samples over time. In some embodiments, such a method for characterizing and / or monitoring includes administering a therapy (e.g., initiating administration of the therapy) to the subject when the difference is determined to be at least as large as a threshold difference. In some embodiments, such a method for characterizing and / or monitoring includes altering administration of a therapy to the subject when the difference is determined to be at least as large as a threshold difference, for example increasing a dosage and / or frequency of administration.
[0212]
[0094] In some embodiments, a method is for prognosing cancer. The method may include estimating cTF in a sample for a subject using a method disclosed herein and prognosing cancer in the subject based on the estimated cTF in the sample, for example based on whether the estimated cTF falls into and / or outside one or more particular ranges of cTF. Prognosing may Page 37 of 124
[0213] 13414664vlAttorney Docket No. : 2014191-0048
[0214] include estimating or determining a stage of cancer. Prognosing may include estimating or determining a cancer as early or late stage. In some embodiments, the method includes administering a therapy based on the prognosis.
[0215]
[0095] In some embodiments, a method is for diagnosing cancer in a subject. The method may include estimating cTF in a sample for a subject using a method disclosed herein; and determining that the estimated cTF exceeds a threshold. In some embodiments, the method includes initiating administration of a therapy based on the estimated cTF. In some embodiments, the method includes selecting a dosing regimen for the therapy based on the estimated cTF.
[0216]
[0096] In some embodiments, a method for determining whether a cancer has been removed from a subject includes estimating cTF in a sample for the subject using a method disclosed herein. The method can be performed after a subject has been administered a therapy to remove cancer. The method may be performed after a subject had a surgical removal of cancer. In some embodiments, the method includes continuing administration of a therapy based on the estimated cTF (e.g., based on a determination that the cancer has not been sufficiently removed based on the estimated cTF). In some embodiments, the method includes ceasing administration of a therapy based on the estimated cTF (e.g., based on a determination that the cancer has been sufficiently removed based on the estimated cTF).
[0217]
[0097] In some embodiments, a method is for predicting a tumor of origin for ctDNA in a sample from a subject and / or making a preliminary diagnosis of cancer in a subject. The method may include estimating cTF in a sample for a subject using a first method disclosed herein for a first type of cancer, for example where the cTF estimation model for the first method has been trained for a first type of cancer [e.g., using only healthy samples and samples corresponding to the first type of cancer (e.g., pure or in silico diluted samples)]. The method may include estimating cTF in the using a second method disclosed herein for a second type of cancer, for example where the cTF estimation model for the second method has been trained for a second type of cancer [e.g., using only healthy samples and samples corresponding to the second type of cancer (eg., pure or in silico diluted samples)]. The method may include making a preliminary determination of a tumor of origin based on a difference in the cTF estimated using the first method and the cTF estimated using the second method. In some embodiments, a tumor of origin may be preliminarily determined based on the cTF estimated using the first method exceeding a threshold difference with the cTF estimated using the second method, and optionally also exceeds a threshold Page 38 of 124
[0218] 13414664vlAttorney Docket No. : 2014191-0048
[0219] value. In some embodiments, a tumor of origin may be estimated based on which of the cTF estimated using the first method and the cTF estimated using the second method is higher and which is lower. The method may include making a preliminary diagnosis of a cancer in the subject based on a difference in the cTF estimated using the first method and the cTF estimated using the second method. In some embodiments, a preliminary diagnosis of a cancer may be based on the cTF estimated using the first method exceeding a threshold difference with the cTF estimated using the second method, and optionally also exceeds a threshold value. In some embodiments, a preliminary diagnosis of a cancer may be based on which of the cTF estimated using the first method and the cTF estimated using the second method is higher and which is lower). The method may include selecting and performing a subsequent diagnostic assessment of the subject for a particular cancer, wherein the particular cancer is selected based on the preliminary diagnosis and / or the preliminary determination of the tumor of origin.
[0220]
[0098] The present disclosure includes methods where a therapeutic agent or regimen is administered to a subject based on an estimate of cTF for the subject and / or a change in estimate of cTF for the subject over time. In some embodiments, a cTF estimation model is cancer-specific in that it has been trained on samples for a specific cancer indication (e.g., type or subtype). In general, a therapeutic agent or regimen provided will be available, appropriate, and / or preferred based on the cTF estimation model used and / or the cTF estimated for a sample from a subject. Those of skill in the art will be aware of recommended and / or governmentally approved formulations and / or dosages for various therapeutic agents.
[0221]
[0099] The present disclosure includes pharmaceutical compositions for delivery of one or more therapeutic agents to a subject. As disclosed herein, a pharmaceutical composition may be in any form known in the art, including formulations for administration according to any route known in the art. A suitable means of administration can be selected based on the age and condition of a subject, for example as determined or indicated by a cTF estimation.
[0222]
[0100] Pharmaceutical composition forms of the present disclosure can include, e.g., liquid, semi-solid and solid dosage forms. Pharmaceutical composition forms of the present disclosure can include, e.g., liquid solutions (e.g., injectable and infusible solutions), dispersions or suspensions, tablets, pills, powders, and liposomes. Selection or use of any particular form may depend, in part, on the intended mode of administration and therapeutic application. Accordingly, the compositions can be formulated for administration by a parenteral mode (e.g., intravenous,
[0223] Page 39 of 124
[0224] 13414664vlAttorney Docket No. : 2014191-0048
[0225] subcutaneous, intraperitoneal, or intramuscular injection) or a non-parenteral mode. As used herein, parenteral administration refers to modes of administration other than enteral and topical administration, usually by injection or infusion.
[0226]
[0101] In some embodiments, the compositions provided herein are present in unit dosage form, which unit dosage form can be suitable for self-administration. Such a unit dosage form may be provided within a container, e.g., a pill, vial, cartridge, prefilled syringe, or disposable pen.
[0227]
[0102] A pharmaceutical composition of the present disclosure can be in an injectable or infusible form. For example, the present disclosure includes sterile formulations for injection or infusion, which can be formulated in accordance with conventional pharmaceutical practices. Sterile solutions can be prepared by incorporating a composition described herein in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filter sterilization. Solutions can be formulated, e.g., using distilled water, physiological saline, or an isotonic solution containing glucose and other supplements such as D-sorbitol, D-mannose, D-mannitol, or sodium chloride as an aqueous solution for injection, optionally in combination with a suitable solubilizing agent, for example, an alcohol such as ethanol and / or a polyalcohol such as propylene glycol or polyethylene glycol, and / or a nonionic surfactant such as polysorbate 80™ or HCO-50, and the like. In the case of sterile powders for the preparation of sterile injectable solutions, methods for preparation include vacuum drying and freeze-drying that yield a powder of a composition described herein plus any additional desired ingredient (see below) from a previously sterile-filtered solution thereof. The proper fluidity of a solution can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Prolonged absorption of injectable compositions can be brought about by including in the composition a reagent that delays absorption, for example, monostearate salts, and gelatin. In particular instances, a pharmaceutical composition can be formulated, for example, as a buffered solution at a suitable concentration and suitable for storage, e.g., at 2-8°C (e.g., 4°C).
[0228]
[0103] In various embodiments, a pharmaceutical composition of the present disclosure can be formulated as a solution, microemulsion, dispersion, liposome, or other ordered structure suitable for stable storage at high concentration. Generally, dispersions are prepared by incorporating a composition described herein into a sterile vehicle that contains a basic dispersion medium and the required other ingredients from those enumerated above.
[0229] Page 40 of 124
[0230] 13414664vlAttorney Docket No. : 2014191-0048
[0231]
[0104] In various instances, a pharmaceutical composition can be formulated to include a pharmaceutically acceptable carrier or excipient. Examples of pharmaceutically acceptable carriers include, without limitation, any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible.
[0232]
[0105] In certain embodiments, compositions can be formulated with a carrier that will protect the therapeutic agent against rapid release, such as a controlled release formulation, including implants and microencapsulated delivery systems. Biodegradable, biocompatible polymers can be used, such as ethylene vinyl acetate, polyanhydrides, polyglycolic acid, collagen, polyorthoesters, and polylactic acid. Many methods for the preparation of such formulations are known in the art. See, e.g., J. R. Robinson (1978) “Sustained and Controlled Release Drug Delivery Systems,” Marcel Dekker, Inc., New York.
[0233]
[0106] Route of administration can be parenteral, for example, administration by injection. Administration by injection can be by intravenous injection, intramuscular injection, intraperitoneal injection, subcutaneous injection. Administration can be systemic or local. In certain embodiments, a composition described herein can be therapeutically delivered to a subject by way of local administration. As used herein, “local administration” or “local delivery,” can refer to delivery that does not rely upon transport of the composition or therapeutic agent to its intended target tissue or site via the vascular system. For example, the composition may be delivered by injection or implantation of the composition or therapeutic agent or by injection or implantation of a device containing the composition or therapeutic agent. In certain embodiments, following local administration in the vicinity of a target tissue or site, the composition or therapeutic agent, or one or more components thereof, may diffuse to an intended target tissue or site that is not the site of administration.
[0234]
[0107] A pharmaceutical composition can be administered parenterally in the form of an injectable formulation including a sterile solution or suspension in water or another pharmaceutically acceptable liquid. For example, a pharmaceutical composition can be formulated by suitably combining the therapeutic molecule with pharmaceutically acceptable vehicles or media, such as sterile water and physiological saline, vegetable oil, emulsifier, suspension agent, surfactant, stabilizer, flavoring excipient, diluent, vehicle, preservative, binder, followed by mixing in a unit dose form required for generally accepted pharmaceutical practices. Examples of Page 41 of 124
[0235] 13414664vlAttorney Docket No. : 2014191-0048
[0236] oily liquid include sesame oil and soybean oil, and it may be combined with benzyl benzoate or benzyl alcohol as a solubilizing agent. Other items that may be included are a buffer such as a phosphate buffer, or sodium acetate buffer, a soothing agent such as procaine hydrochloride, a stabilizer such as benzyl alcohol or phenol, and an antioxidant. The formulated injection can be packaged in a suitable ampule.
[0237]
[0108] In various embodiments, subcutaneous administration can be accomplished by means of a device, such as a syringe, a prefilled syringe, an auto-injector (e.g., disposable or reusable), a pen injector, a patch injector, a wearable injector, an ambulatory syringe infusion pump with subcutaneous infusion sets, or other device for combining with a therapeutic agent for subcutaneous injection.
[0238]
[0109] An injection system of the present disclosure may employ a delivery pen as described in U.S. Pat. No. 5,308,341. Pen devices, most commonly used for self-delivery of insulin to patients with diabetes, are well known in the art. Such devices can include at least one injection needle, are typically pre-filled with one or more therapeutic unit doses of a solution that includes the therapeutic agent and are useful for rapidly delivering solution to a subject with as little pain as possible. One medication delivery pen includes a vial holder into which a vial of a therapeutic or other medication may be received. The pen may be an entirely mechanical device or it may be combined with electronic circuitry to accurately set and / or indicate the dosage of medication that is injected into the user. See, e.g., U.S. Pat. No. 6,192,891. In some embodiments, the needle of the pen device is disposable. Pen devices suitable for delivery of any one of the presently featured compositions are also described in, e.g., U.S. Pat. Nos. 6,277,099; 6,200,296; and 6,146,361, the disclosures of each of which are incorporated herein by reference in their entirety. A microneedle-based pen device is described in, e.g., U.S. Pat. No. 7,556,615, the disclosure of which is incorporated herein by reference in its entirety. See also the Precision Pen Injector (PPI) device, MOLLY™, manufactured by Scandinavian Health Ltd.
[0239] [HO] In certain embodiments, administration of a therapeutic agent as described herein is achieved by administering to a subject a nucleic acid encoding a therapeutic agent described herein. Nucleic acids encoding a therapeutic agent described herein can be incorporated into a gene construct to be used as a part of a gene therapy protocol to deliver nucleic acids that can be used to express and produce therapeutic agent within cells. Expression constructs of such components may be administered in any therapeutically effective carrier, e.g., any formulation or Page 42 of 124
[0240] 13414664vlAttorney Docket No. : 2014191-0048
[0241] composition capable of effectively delivering the component gene to cells in vivo. Approaches include insertion of the subject gene in viral vectors including recombinant retroviruses, adenovirus, adeno-associated virus, lentivirus, and herpes simplex virus- 1 (HSV-1), or recombinant bacterial or eukaryotic plasmids. Viral vectors can transfect cells directly; plasmid DNA can be delivered with the help of, for example, cationic liposomes (lipofectin) or derivatized, polylysine conjugates, gramicidin S, artificial viral envelopes or other such intracellular carriers, as well as direct injection of the gene construct or CaPC>4 precipitation. Examples of suitable retroviruses include adenovirus-derived vectors, adeno-associated virus (AAV), pLJ, pZIP, pWE, and pEM which are known to those skilled in the art.
[0242] [Hl] In some embodiments, a composition can be formulated for storage at a temperature below 0°C (e.g., -20°C or -80°C). In some embodiments, the composition can be formulated for storage for up to 2 years (e g., one month, two months, three months, four months, five months, six months, seven months, eight months, nine months, 10 months, 11 months, 1 year, or 2 years) at 2-8°C (e.g., 4°C). Thus, in some embodiments, the compositions described herein are stable in storage for at least 1 year at 2-8°C (e.g., 4°C).
[0243]
[0112] A pharmaceutical composition can include a therapeutically effective amount of a therapeutic agent described herein. Such effective amounts can be readily determined by one of ordinary skill in the art. A therapeutically effective amount can be an amount at which any toxic or detrimental effects of the composition are outweighed by therapeutically beneficial effects. In some embodiments, a dose can also be chosen to reduce or avoid production of antibodies or other host immune responses against a therapeutic agent. Those of skill in the art will appreciate that data obtained from cell culture assays and animal studies can be used in formulating a range of dosage for use in humans. In various embodiments, the amount of active ingredient included in a pharmaceutical composition is such that a suitable dose within the designated range can be administered to subjects. The dose and method of administration can vary depending on weight, age, condition, and other characteristics of a patient, and can be suitably selected as needed by those skilled in the art.
[0244]
[0113] Pharmaceutical compositions including certain therapeutic agents, e.g., therapeutic antibodies, can be administered as a fixed dose, or in a milligram per kilogram (mg / kg) dose. While in no way intended to be limiting, an exemplary single dose of certain pharmaceutical compositions described herein can include certain therapeutic agents as described herein in an Page 43 of 124
[0245] 13414664vlAttorney Docket No. : 2014191-0048
[0246] amount equal to, e.g., 0.001 to 1000 mg / kg, 1-1000 mg / kg, 1-100 mg / kg, 0.5-50 mg / kg, 0.1-100 mg / kg, 0.5-25 mg / kg, 1-20 mg / kg, and 1-10 mg / kg body weight. Exemplary dosages of a composition described herein include, without limitation, 0.1 mg / kg, 0.5 mg / kg, 1 mg / kg, 2 mg / kg, 4 mg / kg, 8 mg / kg, or 20 mg / kg. The present disclosure is not limited to such ranges or dosages.
[0247]
[0114] In various embodiments, therapeutic agents of the present disclosure can be administered to a subject in a course of treatment that further includes administration of one or more additional therapeutic agents or therapies that are not therapeutic agents (e.g., surgery or radiation). Combination therapies of the present disclosure can include simultaneous exposure of a subject to therapeutic agents of two or more therapeutic regimens.
[0248]
[0115] In certain embodiments, a therapeutic agent as described herein can be administered together with (e.g., at the same time and / or in the same composition as) an additional agent or therapy. In certain embodiments, a therapeutic agent of the present disclosure can be administered separately from an additional therapeutic agent or therapy (e.g., at a different time and / or in a different composition than the additional therapeutic agent or therapy). Dosing regimens of a therapeutic agent and one or more additional therapeutic agents with which it is administered in combination can be coordinated or independently determined. In various embodiments, an additional therapeutic agent or therapy administered in combination with a therapeutic agent as described herein can be administered at the same time as therapeutic agent, on the same day as therapeutic agent, or in the same week as therapeutic agent. In various embodiments, an additional therapeutic agent or therapy administered in combination with a therapeutic agent as described herein can be administered such that administration of the therapeutic agent and the additional therapeutic agent or therapy are separated by one or more hours before or after, one or more days before or after, one or more weeks before or after, or one or more months before or after administration of the therapeutic agent. In various embodiments, the administration frequency and / or dosage of one or more additional therapeutic agents can be the same as, similar to, or different from the administration frequency of a therapeutic agent. In some embodiments, the two or more regimens can be administered simultaneously; in some embodiments, such regimens can be administered sequentially (e.g., all “doses” of a first regimen are administered prior to administration of any doses of a second regimen); in some embodiments, such therapeutic agents are administered in overlapping dosing regimens.
[0249] Page 44 of 124
[0250] 13414664vlAttorney Docket No. : 2014191-0048
[0251]
[0116] In certain embodiments, administration of a therapeutic agent can be to a subject having previously received, scheduled to receive, or in the course of a treatment regimen including an additional cancer therapy. Administration of a therapeutic agent can, in some instances, improve delivery or efficacy of another therapeutic agent or therapy with which it is administered in combination.
[0252]
[0117] It is contemplated that therapeutic agent combination therapies can demonstrate synergy and / or greater-than-additive effects between a therapeutic agent and one or more additional therapeutic agents with which it is administered in combination. A therapeutic agent can be administered in any effective amount as determined independently or as determined by the joint action of therapeutic agent and any of one or more additional therapeutic agents or therapies administered. Administration of the therapeutic agent may, in some embodiments, reduce the therapeutically effective dosage, required dosage, or administered dosage of the additional therapeutic agent or therapy relative to a reference regimen for administration of additional therapeutic agent or therapy or therapy absent the therapeutic agent. In certain embodiment, a composition described herein can replace or augment other previously or currently administered therapy. For example, upon treating with therapeutic agent, administration of one or more additional therapeutic agents or therapies can cease or diminish, e.g., be administered at lower levels.
[0253] Subjects and Samples
[0254]
[0118] Estimation models may be produced using a plurality of biological samples. Such samples may be liquid biopsy samples, for example from a blood draw (e.g., derived from whole blood). Such samples may be plasma samples, for example including empirical (e.g., obtained and tested) and / or simulated plasma samples (e.g., obtained by a linear combination of two empirical samples). A sample may be derived from a cell line. A sample may be derived from a subject. In some embodiments, samples span a range of known cTF (e.g., from 0% cTF to 100% cTF, from 1% cTF to 100% cTF, from 2% cTF to 100%, or from 3% cTF to 100% cTF). In some embodiments, a cTF of a sample is known from a reference method (e.g., an orthogonal reference method), such as, for example ichorCNA. Other orthogonal reference methods such as VAF and / or CNV methods may be used.
[0255] Page 45 of 124
[0256] 13414664vlAttorney Docket No. : 2014191-0048
[0257]
[0119] A plurality of samples used to produce a cTF estimation model may include in silico diluted (e.g., titrated) samples (e.g., in silico plasma (ISP) samples). Sequencing data for an in silico diluted sample may be obtained from a random sampling of fragments from a healthy (e.g., zero or below limit of detection cTF) sample and a sample with a known non-zero cTF in a predetermined ratio of amounts. In some embodiments, a sample is a simulated dilution and a known cTF for the sample has been determined by linear combination of known cTFs of reference samples used to generate the sample.
[0258]
[0120] A sample analyzed using methods and systems provided herein can be any biological sample including any processed sample that includes circulating tumor DNA (ctDNA) derived from a biological sample. In various embodiments, a sample used in (e.g., analyzed using) methods and systems provided herein can be a sample obtained from a mammalian subject. In various embodiments, a sample used in (e.g., analyzed using) methods and systems provided herein can be a sample obtained from a human subject. A sample may come from a donor subject, a patient subject, or a cell line, for example.
[0259]
[0121] In various instances, a human subject is a subject diagnosed or seeking diagnosis as having, diagnosed as, or seeking diagnosis as at risk of having, and / or diagnosed as or seeking diagnosis as at immediate risk of having, cancer. In various instances, a human subject is a subject identified as needing cancer screening. In certain instances, a human subject is a subject identified as needing cancer screening by a medical practitioner.
[0260]
[0122] A subject may not have undergone previous treatments for cancer. In other embodiments, a subject has undergone previous treatments for cancer, such as treatments recited in this disclosure.
[0261]
[0123] In various embodiments, a subject has one or more biomarkers and / or risk factors for cancer. In various instances, a human subject is a subject not yet diagnosed as having, not at risk of having, not at immediate risk of having, not diagnosed as having, and / or not seeking diagnosis for a cancer. Genetic factors may also contribute to cancer risk, as evidenced by individuals with a family history of cancer.
[0262]
[0124] In various embodiments, a sample from a subject, e.g., a human, can be obtained from a liquid biopsy. In certain embodiments, a sample and / or reference is obtained from serum, plasma, or urine. In certain embodiments, the sample is serum. In certain embodiments, a sample includes circulating tumor DNA (ctDNA). In certain embodiments, a sample is derived from about Page 46 of 124
[0263] 13414664vlAttorney Docket No. : 2014191-0048
[0264] 1 mL of blood obtained from the subject. In certain embodiments, a sample is derived from about 0.5-2 mL of blood obtained from the subject, e.g., about 0.5 to 1.75 mL, about 0.5 to 1.5 mL, about 0.75 to 1.25 mL or about 0.9 to 1.1 mL of blood. cTF estimation models disclosed herein may be used with samples having low sample volumes, for example no more than 3 mL, no more than 2 mL, or no more than 1 mL.
[0265]
[0125] In various embodiments, a sample is a sample of or containing cell-free DNA (cfDNA). cfDNA is typically found in human biofluids (e.g., plasma, serum, or urine) in short, double-stranded fragments. The concentration of cfDNA is typically low, but can significantly increase under particular conditions, including without limitation pregnancy, autoimmune disorders, myocardial infarction, and cancer. Circulating tumor DNA (ctDNA) is the component of cell-free DNA specifically derived from cancer cells. ctDNA can be present in human biofluids bound to leukocytes and erythrocytes or not bound to leukocytes and erythrocytes. Various tests for detection of tumor-derived ctDNA are based on detection of genetic or epigenetic modifications that are characteristic of cancer (e.g., of a relevant cancer). Genetic or epigenetic factors characteristic of cancer can include, without limitation, oncogenic or cancer-associated mutations in tumor-suppressor genes, activated oncogenes, chromosomal disorders, histone modifications (e.g., histone methylation and / or histone acetylation), chromatin accessibility, binding of one or more transcription factors and / or DNA methylation.
[0266]
[0126] In various embodiments, ctDNA includes less than 30%, less than 20%, or less than 10% of the cfDNA in the liquid biopsy sample obtained from the subject, e.g., less than 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or less than 1% of the cfDNA in the sample. In various embodiments, ctDNA includes at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the cfDNA in the liquid biopsy sample obtained from the subject.
[0267]
[0127] cfDNA and ctDNA can provide a real-time or nearly real time metric of status of a source tissue. cfDNA and ctDNA demonstrate a half-life in blood of about 2 hours, such that a sample taken at a given time provides a relatively timely reflection of the status of a source tissue.
[0268]
[0128] Various methods of isolating nucleic acids from a sample (e.g., of isolating cfDNA from blood or plasma) are known in the art. Nucleic acids can be isolated using, without limitation, standard DNA purification techniques, by direct gene capture (e.g., by clarification of a sample to remove assay-inhibiting agents and capturing a target nucleic acid, if present, from the clarified sample with a capture agent to produce a capture complex and isolating the capture complex to Page 47 of 124
[0269] 13414664vlAttorney Docket No. : 2014191-0048
[0270] recover the target nucleic acid). Alternatively or additionally, nucleic acids can be isolated by performing an initial isolation of a protein-nucleic acid complex (e.g., a nucleosome) by using a capture agent that binds to a complex, a nucleic acid in a complex, or a non-nucleic acid component of a complex (e.g., a polypeptide, e.g., a histone), with a subsequent isolation of a nucleic acid from an isolated complex, for example, by denaturing a non-nucleic acid component of the complex (e.g., a polypeptide, e.g., a histone) and releasing a nucleic acid from the complex.
[0271]
[0129] Reagents and protocols for obtaining and analyzing cfDNA and ctDNA, such as circulating in blood or other tissue, are commercially available as described in the Examples and well-known in the art (see, for example, Anker et al., Cancer and Metastasis Rev (1999) 18:65-73; Wua et al., Clin Chim Acta (2002) 321:77-87; Fiegl et al., Cancer Res (2005) 15:1141-1145; Pathak et al., Clin Chem (2006) 52:1833-1842; Schwarzenbach et al., Clin Cancer Res (2009) 15:1032-1038; Schwarzenbach et al., Nat Rev Cancer (2011) 11:426-437) the contents of each of which is separately incorporated herein by reference in their entirety).
[0272]
[0130] In various embodiments, samples can be collected from individuals repeatedly over a period of time (e.g., once daily, weekly, monthly, annually, biannually, etc.). In various embodiments, such samples can be used to verify results from earlier detections and / or to identify an alteration in biological pattern because of, for example, disease progression, resistance to therapy, treatment, or remission. For example, subject samples can be taken and monitored every month, every two months, or combinations of one, two, or three-month intervals according to the present disclosure. In various embodiments, samples can be collected for monitoring over time beginning at or at certain clinically determined stages, such as at resistance to a therapy, before radiographic progression, after radiographic progression, and / or at tissue biopsy. In addition, AR activity obtained at different points in time can be conveniently compared with each other, as well as with those of normal controls during the monitoring period, thereby providing the subject’s own values, as an internal, or personal, control for long-term monitoring.
[0273]
[0131] Samples may include materials prepared by processes including, without limitation, steps such as concentration, dilution, adjustment of pH, removal of high abundance polypeptides (e.g., albumin, gamma globulin, and transferrin, etc.), addition of preservatives, addition of calibrants, addition of protease inhibitors, addition of denaturants, desalting, concentration and / or extraction of sample nucleic acids, and / or amplification of sample nucleic acids (e.g., by PCR or other nucleic acid amplification techniques). Samples may also include materials prepared by Page 48 of 124
[0274] 13414664vlAttorney Docket No. : 2014191-0048
[0275] techniques that isolate, e.g., nucleosomes or transcription factors and / or nucleic acids associated with nucleosomes or transcription factors.
[0276] [1321 Removal from a sample of proteins that are not desirable for a relevant purpose or context (e.g., high abundance, uninformative, or undetectable proteins) can be achieved using high affinity reagents, high molecular weight filters, ultracentrifugation and / or electrodialysis. High affinity reagents include antibodies or other reagents (e.g., aptamers) that selectively bind to high abundance proteins. Sample preparation can also include ion exchange chromatography, metal ion affinity chromatography, gel filtration, hydrophobic chromatography, chromatofocusing, adsorption chromatography, isoelectric focusing and related techniques. Molecular weight filters include membranes that separate molecules based on size and molecular weight. Such filters may further employ reverse osmosis, nanofiltration, ultrafiltration and microfiltration. Ultracentrifugation is the centrifugation of a sample at about 15,000-60,000 rpm while monitoring with an optical system the sedimentation (or lack thereof) of particles. Electrodialysis is a procedure which uses an electromembrane or semipermeable membrane in a process in which ions are transported through semi-permeable membranes from one solution to another under the influence of a potential gradient. Since the membranes used in electrodialysis may have the ability to selectively transport ions having positive or negative charge, reject ions of the opposite charge, or to allow species to migrate through a semipermeable membrane based on size and charge, it renders electrodialysis useful for concentration, removal, or separation of electrolytes.
[0277]
[0133] Separation and purification in the present disclosure may include any procedure known in the art, such as capillary electrophoresis (e.g., in capillary or on-chip) or chromatography (e.g., in capillary, column or on a chip). Electrophoresis is a method that can be used to separate ionic molecules under the influence of an electric field. Electrophoresis can be conducted in a gel, capillary, or in a microchannel on a chip. Examples of gels used for electrophoresis include starch, acrylamide, polyethylene oxides, agarose, or combinations thereof. A gel can be modified by its cross-linking, addition of detergents, or denaturants, immobilization of enzymes or antibodies (affinity electrophoresis) or substrates (zymography) and incorporation of a pH gradient. Examples of capillaries used for electrophoresis include capillaries that interface with an electrospray.
[0278]
[0134] Capillary electrophoresis (CE) is preferred for separating complex hydrophilic molecules and highly charged solutes. CE technology can also be implemented on microfluidic Page 49 of 124
[0279] 13414664vlAttorney Docket No. : 2014191-0048
[0280] chips. Depending on the types of capillary and buffers used, CE can be further segmented into separation techniques such as capillary zone electrophoresis (CZE), capillary isoelectric focusing (CIEF), capillary isotachophoresis (CITP) and capillary electrochromatography (CEC). An embodiment to couple CE techniques to electrospray ionization involves the use of volatile solutions, for example, aqueous mixtures containing a volatile acid and / or base and an organic such as an alcohol or acetonitrile.
[0281]
[0135] Capillary isotachophoresis (CITP) is a technique in which the analytes move through the capillary at a constant speed but are nevertheless separated by their respective mobilities. Capillary zone electrophoresis (CZE), also known as free-solution CE (FSCE), is based on differences in the electrophoretic mobility of the analytes, determined by the charge on the analytes, and the frictional resistance the analytes encounter during migration, which is often directly proportional to the size of the analytes. Capillary isoelectric focusing (CIEF) allows weakly-ionizable amphoteric molecules, to be separated by electrophoresis in a pH gradient. CEC is a hybrid technique between traditional high performance liquid chromatography (HPLC) and CE.
[0282]
[0136] Separation and purification techniques used in the present disclosure can include any chromatography procedures known in the art. Chromatography can be based on the differential adsorption and elution of certain analytes or partitioning of analytes between mobile and stationary phases. Different examples of chromatography include, but not limited to, liquid chromatography (LC), gas chromatography (GC), high performance liquid chromatography (HPLC), etc.
[0283]
[0137] In some embodiments, whole blood is collected from a subject, and a plasma layer is separated by centrifugation. cfDNA may be then extracted from the plasma using methods known in the art.
[0284]
[0138] Certain methods and systems disclosed herein use samples (e.g., liquid biopsy samples) having known cTF. Known cTF may be known because it is empirically known, assumed, or expected. In some embodiments, cTF of a liquid biopsy sample is assessed using ichorCNA which estimates the percentage of ctDNA in a sample probabilistically to obtain an assumed known cTF (see Adalsteinsson et al., Nat Commun. (2017) 8(1): 1324 the entire contents of which are incorporated herein by reference). Other orthogonal methods for estimating cTF, such as VAF or CNV methods, may be used to obtain assumed known cTFs for samples. A cTF Page 50 of 124
[0285] 13414664vlAttorney Docket No. : 2014191-0048
[0286] for a sample may be known empirically, for example if the sample is derived from a healthy (e.g., cancer free) subject and therefore has a cTF of 0% (0) or if the sample is derived from a cancer cell line and therefore has a cTF of 100% (1). A cTF may be an expected known cTF for example where the sample is an in silico dilution (e.g., ISP) sample as described elsewhere herein, in which case the cTF is expected based on the predetermined dilution ratio.
[0287] Techniques for Detecting and Quantifying Histone Modifications
[0288]
[0139] DNA from a sample may be sequenced in order to obtain sequencing data for a sample. DNA from a sample may be sequenced in order to determine histone modification status.
[0289]
[0140] Various methylation sequencing techniques may be used for detecting and quantifying histone modifications to produce histone modification sequencing data.
[0290]
[0141] In some embodiments, the methods, kits and systems of present disclosure involve the detection and quantification of histone modifications, e.g., in liquid biopsy samples including cfDNA such as plasma samples including cfDNA. Chromatin ImmunoPrecipitation (ChIP) is one technique of molecular biology useful in detecting and quantifying histone modifications in samples. CUT&RUN or CUT&Tag are other more recent techniques that can also be used to detect and quantify histone modifications. ChlP-chip, ChlP-exo, ChIP Re-ChIP, and ChlPmentation are other alternative techniques that could be used.
[0291]
[0142] ChIP can involve various steps including one or more of fixation, sonication, immunoprecipitation, and analysis of the immunoprecipitated DNA. ChIP has become a very widely used tissue-based technique for determining the in vivo location of binding sites of various histones. Because the proteins are captured at the sites of their binding with DNA, ChIP helps to detect DNA-protein interactions that take place in living cells. More importantly, ChIP can be coupled to many commonly used molecular biology techniques such as PCR and real-time PCR, PCR with single-stranded conformational polymorphism, Southern blot analysis, Western blot analysis, cloning, and microarray. The resulting versatility has increased the potential of this technique.
[0292]
[0143] ChIP of tissue samples usually involves cross-linking of the chromatin-bound proteins by formaldehyde, followed by sonication or nuclease treatment to obtain small DNA fragments. Immunoprecipitation can be then carried out using specific antibodies to the DNA-binding protein of interest. The DNA can be then released from the proteins and analyzed using Page 51 of 124
[0293] 13414664vlAttorney Docket No. : 2014191-0048
[0294] various methods. ChIP has also been used to study RNA-protein interactions. X-ChIP methods utilize fixed chromatin fragmented by sonication, while the N-ChIP methods utilize native chromatin, which can be unfixed and nuclease digested.
[0295]
[0144] The first step of the technique can be the cross-linking of DNA and proteins. Formaldehyde is one of the most used cross-linking agents. One advantage of using formaldehyde can be the ease of reversibility of the cross-links and its ability to form bonds that span approximately 2 angstroms. This means that formaldehyde can bind molecules in close association with each other. Generally, formaldehyde can be added to the medium in the cell culture flask or plate. It enters the cells through the cell membrane and cross-links the proteins to the chromatin. Formaldehyde fixation of tumor tissues has also been done. Other cross-linking agents that have been used include chemicals such as methylene blue and acridine orange, cisplatin, dimethylarsinic acid, potassium chromate, and ultraviolet (UV) light and lasers.
[0296]
[0145] Harvested chromatin can be sonicated in one or more sonication cycles. DNA can be typically broken into to 100-500 bp fragments to pinpoint the location of the DNA sequence of interest. An alternative to sonication can be nuclease digestion of the chromatin, e.g., in N-ChIP methods. Purification of chromatin can be achieved using a cesium chloride (CsCl) gradient centrifugation.
[0297]
[0146] Chromatin can be immunoprecipitated using one or more antibodies that bind a target epitope. For example, an antibody used in ChIP can selectively bind one or more particular histone modifications, such as one or more particular histone acetylation modifications or histone methylation modifications. In some embodiments, an antibody used to bind a target epitope can be a “pan” antibody (e.g., a pan-acetylation antibody, a pan-methylation antibody, an antibody that binds a group of histone modifications associated with increased transcription activation, and / or an antibody that binds a group of histone modifications associated with increased transcription repression). The antibody against the protein of interest is allowed to bind to the protein-DNA complex, and the complex can be then precipitated. Immunosorbants commonly used to separate the antigen-antibody complex from the lysate include salmon sperm DNA-protein A-Sepharose®, protein G, magnetic beads, and other engineered immunoprecipitation systems known to those of skill in the art.
[0298]
[0147] Immunoprecipitated DNA can be eluted. Once the DNA of interest is isolated, many detection and quantification methods can be used to study the isolated gene fragments.
[0299] Page 52 of 124
[0300] 13414664vlAttorney Docket No. : 2014191-0048
[0301] Commonly utilized methods include PCR, real-time PCR, slot blot hybridization, microarray techniques, and deep or next-generation sequencing. ChlP-seq combines chromatin immunoprecipitation (ChIP) with massively parallel DNA sequencing to identify the binding sites of DNA-associated proteins. ChlP-seq can be used to map DNA-binding proteins, e.g., histone modifications, in a genome-wide manner.
[0302]
[0148] Cell-free Chromatin ImmunoPrecipitation sequencing (cfChlP-seq) involves applying ChlP-seq to samples that include cell-free DNA, e.g., liquid biopsy samples including cfDNA such as plasma samples including cfDNA (e.g., see Sadeh et al., Nat Biotechnol (2021) 39: 586-598 and Jang et al., Life Sci Alliance (2023) 6(12):e202302003 the entire contents of each of which are incorporated herein by reference). In some embodiments, cfChlP-seq uses antibodies or antibody fragments that bind specific histone modifications (e.g., H3K4me3 and / or H3K27ac) that are coupled (covalently or non-covalently) to beads, e.g., magnetic beads such as Dynabeads® magnetic beads and incubated with a volume, e.g., about 1 mL of thawed plasma obtained from a subject. Without limitation, exemplary antibodies that bind H3K4me3 include PA5-27029 (available from Thermo Fisher Scientific in Waltham, MA) and C15410003 (available from Diagenode in Denville, NJ) and exemplary antibodies that bind H3K27ac include ab21623 or ab4729 (both available from Abeam in Cambridge, UK) and Cl 5210016 (available from Diagenode in Denville, NJ).
[0303]
[0149] In some embodiments, the antibodies or antibody fragments can be covalently coupled to beads, e.g., epoxy beads. In some embodiments, the antibodies or antibody fragments can be non-covalently coupled to beads, e.g., Protein A or Protein G beads such as Dynabeads® Protein A or Dynabeads® Protein G beads. After washing, a cfDNA library is then typically prepared from the captured cfDNA. Library preparation can be done on-bead or after releasing the captured cfDNA by digestion of bound histones, e.g., using proteinase K. The cfDNA library is then sequenced to generate reads of captured cfDNA sequences, e.g., by next-generation sequencing (NGS) as is known in the art. The reads are then analyzed, e.g., aligned and counted using standard bioinformatic techniques as is known in the art. A cfChlP-seq bioinformatic pipeline can include, e.g., alignment of sequence reads to a reference genome with BWA or Bowtie2. Aligned reads can be used to call and quantify peaks as compared to a reference.
[0304] Page 53 of 124
[0305] 13414664vlAttorney Docket No. : 2014191-0048
[0306] Differential Histone Modification
[0307]
[0150] In various embodiments, differential histone modification refers to a status characterized by an increase or decrease in a value measuring binding of a histone to a region (e.g., of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3 -fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring methylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase of 0.1-fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold. 1.2-fold to 4.0-fold. 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.
[0308] Techniques for Detecting and Quantifying DNA Methylation
[0309]
[0151] DNA from a sample may be sequenced in order to obtain sequencing data for the sample. DNA from a sample may be sequenced in order to determine methylation status. Various methylation sequencing techniques may be used to produce methylation sequencing data. In some embodiments, DNA extracted from plasma is enriched for densely methylated fragments as part of a sequencing method, such as, for example, Methyl-CpG-Binding Domain Sequencing (MBD-seq). In some embodiments, after sequencing, FASTQ files are processed to produce a table of Page 54 of 124
[0310] 13414664vlAttorney Docket No. : 2014191-0048
[0311] unique fragments that align to a genome for a subject from which a sample is derived (e g., the human genome for human subjects / samples). Sequencing data used to produce a cTF estimation model and / or used to estimate cTF for a sample may be preprocessed to get unique fragments from aligned sequencing reads. In some embodiments, sequencing data received and / or obtained may be initially deduplicated.
[0312]
[0152] Various techniques of molecular biology are well known in the art and / or disclosed in the present application for detecting and quantifying DNA methylation, for example of cfDNA (e.g., ctDNA) in a sample. In some embodiments, the methods and systems of the present disclosure involve the detection and quantification of DNA methylation in samples, e.g., in liquid biopsy samples including cfDNA (e.g., ctDNA) such as plasma samples including cfDNA. Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq) and Methyl-CpG-Binding Domain sequencing (MBD-seq) are exemplary techniques of molecular biology useful in detecting and quantifying DNA methylation in samples, though others, such as Bisulfite sequencing (BS-Seq) and Whole Genome Bisulfite Sequencing (WGBS) exist. In some embodiments, an enrichment method, such as MBD-seq, is used to obtain sequencing data. In some embodiments, sequencing data from an enrichment method, such as MBD-seq, is received and processed.
[0313]
[0153] DNA methylation typically refers to the methylation of the 5’ position of cytosine (mC) by DNA methyltransferases (DNMT). It is a major epigenetic modification in humans and many other species. In mammals, most DNA methylations occur within the context of CpG dinucleotides. DNA methylation is thought to be a repressive chromatin modification. Aberrant methylation can lead to many diseases including cancers (Robertson, Nat Rev Genet (2005) 6:597-610 and Bergman and Cedar, Nat Struct Mol Biol (2013) 20:274-281).
[0314]
[0154] MeDIP-seq was first reported by Weber et al., Nat Genet (2005) 37:853-862. In a typical MeDIP-seq protocol, antibody or antibody-fragment that binds 5-methylcytidine (5mC) is used to enrich methylated DNA fragments, then these fragments are sequenced and analyzed. If using 5mC-specific antibodies or antibody fragments, methylated DNA is isolated from genomic DNA via immunoprecipitation. Anti-5mC antibodies are incubated with fragmented genomic DNA and precipitated, followed by DNA purification and sequencing.
[0315]
[0155] Methyl-CpG-Binding Domain sequencing (MBD-seq) is similar to MeDIP-seq except that it uses methyl binding domain (MBD) proteins instead of antibodies or antibody fragments to bind methylated DNA. In a typical MBD-seq protocol, genomic DNA is first sonicated Page 55 of 124
[0316] 13414664vlAttorney Docket No. : 2014191-0048
[0317] and incubated with tagged MBD proteins that can bind methylated cytosines. The protein-DNA complex is then precipitated with antibody-conjugated beads that are specific to the MBD protein tag, followed by DNA purification and sequencing.
[0318] Differential DNA Methylation
[0319]
[0156] In various embodiments, differentially DNA methylated refers to a methylation status characterized by an increase or decrease in a value measuring methylation (e.g., of read counts and / or normalized read counts for a given genomic locus), and / or a mean, median and / or mode thereof, and / or a log thereof (e.g., log base 2 (log2)), of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, 50-fold, or greater, or any range in between, inclusive, such as 1% to 50%, 50% to 2-fold, 25% to 50-fold, 25% to 30-fold, 25% to 20-fold, 25% to 16-fold, 30% to 16-fold, 50% to 16-fold, 70% to 16-fold, 2-fold to 16-fold, 2.2-fold to 16-fold, 2.6-fold to 16-fold, 3-fold to 16-fold, 3.4-fold to 16-fold, 4-fold to 16-fold, 4.5-fold to 16-fold, 5.2-fold to 16-fold, 6-fold to 16-fold, 7-fold to 16-fold, or 8-fold to 16-fold, as compared to a reference, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6. In various embodiments, an increase or decrease in a value measuring methylation can be, or is expressed as, a log2(fold-change), e.g., a log2(fold-change) of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, or greater, or any range in between, inclusive, such as an increase of 0.1-fold to 10-fold, 0.2-fold to 5-fold, 0.2-fold to 4.0-fold, 0.4-4.0-fold, 0.4-fold to 4.0-fold, 0.6-fold to 4.0-fold, 0.8-fold to 4.0-fold, 1.0-fold to 4.0-fold. 1.2-fold to 4.0-fold. 1.4-fold to 4.0-fold, 1.6-fold to 4.0-fold, 1.8-fold to 4.0-fold, 2.0-fold to 4.0-fold, 2.2-fold to 4.0-fold, 2.4-fold to 4.0-fold, 2.6-fold to 4.0-fold, 2.8-fold to 4.0-fold, or 3.0-fold to 4.0-fold, optionally where the statistical significance of the increase or decrease is at least 5e-2, le-2, 5e-3, le-3, 5e-4, le-4, 5e-5, le-5, 5e-6, or le-6.
[0320] Systems
[0321]
[0157] Among other things, provided herein are systems for detecting various histone modifications of one or more genomic loci. In some embodiments, systems for detecting various Page 56 of 124
[0322] 13414664vlAttorney Docket No. : 2014191-0048
[0323] histone modifications may use ChTP-seq. A histone modification may be quantified using provided systems, Alternatively or additionally, the present disclosure includes systems for detecting methylation of one or more genomic loci, for example using MBD-seq. In some embodiments, the present disclosure provides systems for quantifying one or more DNA methylation at one or more genomic loci. Systems of the present disclosure can include a sequencer configured to generate a sequencing dataset from a sample; and a non-transitory computer readable storage medium and / or a computer system.
[0324]
[0158] In some embodiments, the non-transitory computer readable storage medium is encoded with a computer program, wherein the program includes instructions that when executed by one or more processors cause the one or more processors to perform operations to perform a method of the present disclosure.
[0325]
[0159] In some embodiments, the computer system includes a memory and one or more processors coupled to the memory, wherein the one or more processors are configured to perform a method of the present disclosure.
[0326]
[0160] In some embodiments, the sequencer is configured to generate a whole genome sequencing (WGS) dataset from the sample. In some embodiments, the sequencer is configured to generate a chromatin immunoprecipitation sequencing (ChlP-seq) dataset from the sample. In some embodiments, the sequencer is configured to generate a methyl-CpG-binding domain sequencing (MBD-seq) dataset from the sample. In some embodiments, the system also includes a sample preparation device configured to prepare the sample for sequencing from a biological sample, optionally a liquid biopsy sample. The sample preparation device may include reagents for quantifying epigenetic modification (e.g., histone modification and / or DNA methylation) at one or more genomic loci in cell-free DNA (cfDNA) from the biological sample, optionally the liquid biopsy sample.
[0327]
[0161] Systems of the present disclosure can include, e.g., reagents such as buffers and / or antibodies useful in the detection and quantification of histone modifications. In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds a histone modification selected from H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, or pan acetylation. In certain embodiments, a system of the present disclosure can include at least one antibody that selective binds H3K4me3 modifications. In certain embodiments, a system of the present disclosure can include at least one antibody that Page 57 of 124
[0328] 13414664vlAttorney Docket No. : 2014191-0048
[0329] selective binds H3K27ac modifications. A system of the present disclosure can include instructional materials disclosing or describing the use of the system in a method of estimating cTF and / or cancer treatment disclosed herein.
[0330]
[0162] In some embodiments, the system includes reagents for isolation of cfDNA from a liquid biopsy sample. In some embodiments, the sequencer includes reagents for library preparation for sequencing. In some embodiments, the sequencer includes reagents for sequencing.
[0331]
[0163] Illustrative embodiments of systems and methods disclosed herein are described with reference to determinations and / or models that may be performed or used by a computing device. That is, in some embodiments, methods disclosed herein are computer-implemented methods and, in some embodiments, a system as disclosed herein includes a processor and one or more non-transitory computer readable storage media (e.g, one or more memories) that have instructions stored thereon that, when executed by the processor, cause the processor to perform operations that include a method disclosed herein. For example, a cTF estimation model may be stored on a memory. Such a cTF estimation model may be utilized by a processor to perform a method (e.g., a diagnostic method). Methods of the present disclosure, or portions thereof, may be performed using a processor. The processor may be a part of a computing device and / or computing system.
[0332]
[0164] Systems of the present disclosure may include a processor and / or a memory. The memory may store one or more programs that include instructions that when executed by a processor cause at least a portion of a method disclosed herein to be performed. The system may further include a machine-learned model. Additionally or alternatively, a remotely stored and / or operated machine-learned model may be accessed by a (e.g., the) processor. The processor and / or memory may be a part of a computing device and / or computing system.
[0333]
[0165] One or non-transitory computer readable media may store one or more programs that include instructions that when executed by a (e.g., the) processor cause at least a portion of a method disclosed herein to be performed.
[0334]
[0166] Methods and systems disclosed herein may utilize one or more machine-learned models. A machine-learned model may be or include an artificial neural network. A machine-learned model may employ, for example, an attention-based model (e.g., a transformer model, such as, for example, a vision transformer), a transformer model (e.g., a vision transformer), a Page 58 of 124
[0335] 13414664vlAttorney Docket No. : 2014191-0048
[0336] regression -based model (e.g., a logistic regression model), a regularization-based model (e.g., an elastic net model or a ridge regression model), an instance-based model (e.g., a support vector machine or a k-nearest neighbor model), a Bayesian-based model (e.g., a naive-based model or a Gaussian naive-based model), a clustering-based model (e.g., an expectation maximization model), an ensemble-based model (e.g., an adaptive boosting model, a random forest model, a bootstrap-aggregation model, or a gradient boosting machine model), a histogram-based gradient boosting regression tree model, or a neural-network-based model (e.g., a convolutional neural network, a recurrent neural network, autoencoder, a back propagation network, or a stochastic gradient descent network).
[0337]
[0167] In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a k nearest neighbors methodology, a generalized regression forward selection methodology, a generalized regression pruned forward selection methodology, a fit stepwise methodology, a generalized regression lasso methodology, a generalized regression elastic net methodology, a generalized regression ridge methodology, a nominal logistic methodology, a support vector machines methodology, a discriminant methodology, a naive Bayes methodology, or a combination thereof. In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a generalized regression lasso methodology, a generalized regression elastic net methodology, a generalized regression ridge methodology, a histogram-based gradient boosting regression tree methodology, a nominal logistic methodology, a support vector machines methodology, a discriminant methodology, or a combination thereof. In some embodiments, a machine-learned model is or is derived from a decision tree methodology, a neural boosted methodology, a bootstrap forest methodology, a boosted tree methodology, a support vector machines methodology, or a combination thereof.
[0338]
[0168] Certain embodiments described herein make use of computer algorithms in the form of software instructions executed by a computer processor. In certain embodiments, the software instructions include a machine learned module. A machine learned module refers to a computer implemented process (e.g., a software function) that implements one or more specific machine-learned models, such as or including an artificial neural network (ANN), a convolutional neural network (CNN), random forest, one or more decision trees, one or more support vector Page 59 of 124
[0339] 13414664vlAttorney Docket No. : 2014191-0048
[0340] machines, or a combination thereof, in order to determine, for a given input (e.g., one or more inputs), one or more output values. In certain embodiments, the input includes an image. In certain embodiments, the input includes numerical data, tagged data, and / or functional relationships. In certain embodiments, the input includes alphanumeric data which can include numbers, words, phrases, or lengthier strings, for example. In certain embodiments, the one or more output values include values representing numeric values, words, phrases, or other alphanumeric strings.
[0341]
[0169] In embodiments, a machine-learned model has been trained using supervised learning algorithm(s), unsupervised learning algorithm(s), semi-supervised learning algorithm(s) (e.g., partial supervision), weak supervision, transfer, multi-task learning, or any combination thereof. In embodiments, a machine-learned model employs a model that includes parameters (e.g., weights) that are tuned during training of the model. For example, the parameters may be adjusted to minimize a loss function, thereby improving the predictive capacity of the machine learning model. A machine-learned model may be further trained after an initial training period, for example, may be adapted to continuously train as it is used.
[0342]
[0170] In certain embodiments, a machine learned model has been trained, for example, using datasets that include categories of data described herein. Such training may be used to determine various parameters of a machine learned model, for example implemented by a machine learning module, such as, for example, weights associated with layers in neural networks. In certain embodiments, once a machine learned model has been trained, e.g., to accomplish a specific task such as identifying certain output (e.g., extracting certain feature vector(s)), values of determined parameters are fixed and the (e.g., unchanging, static) machine learned model is used to process new data (e.g., different from the training data) and accomplish its trained task without further updates to its parameters (e.g., the machine learned model does not receive feedback and / or updates). In certain embodiments, a machine learned model may receive feedback, e.g., based on user review of accuracy, and such feedback may be used as additional training data, to dynamically update the machine learned model. In certain embodiments, two or more machine learning models may be combined and implemented as a single model, in a single module, and / or in a single software application. In certain embodiments, two or more machine learned models may also be implemented separately, e.g., as separate software applications. In certain embodiments, two or more machine learned modules may also be implemented separately, e.g., as separate software applications. A machine learned model may be or include software and / or hardware. For example,
[0343] Page 60 of 124
[0344] 13414664vlAttorney Docket No. : 2014191-0048
[0345] a machine learned model may be implemented entirely as software, or certain functions of an ANN module (e.g., CNN) may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC)). A machine learned module may be or include software and / or hardware. For example, a machine learned module may be implemented entirely as software, or certain functions of an ANN module (e.g., CNN) may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC)).
[0346]
[0171] In certain embodiments, machine learning modules implementing machine learning techniques may be composed of individual nodes (e.g. units, neurons). A node may receive a set of inputs that may include at least a portion of a given input data for the machine learning module and / or at least one output of another node. A node may have at least one parameter to apply and / or a set of instructions to perform (e.g., mathematical functions to execute) over the set of inputs. In certain embodiments, node instructions may include a step to provide various relative importance to the set of inputs using various parameters, such as weights. The weights may be applied by performing scalar multiplication (e.g., or other mathematical function) between a set of inputs values and the parameters, resulting in a set of weighted inputs. In certain embodiments, a node may have a transfer function to combine the set of weighted inputs into one output value. A transfer function may be implemented by a summation of all the weighted inputs and the addition of an offset (e.g., bias) value. In certain embodiments, a node may have an activation function to introduce non-linearity into the output value. Non-limiting examples of the activation function include Rectified Linear Activation (ReLu), logistic (e.g., sigmoid), hyperbolic tangent (tanh), and softmax. In certain embodiments, a node may have a capability of remembering previous states (e.g., recurrent nodes). Previous states may be applied to the input and output values using a set of learning parameters.
[0347]
[0172] In certain embodiments, the machine learning module includes a deep learning architecture composed of nodes organized into layers. For example, a layer is a set of nodes that receives data input (e.g., weighted or non-weighted input), transforms it (e.g., by carrying out instructions, e.g., applying a set of functions e.g., linear and / or non-linear functions), and passes transformed values as output (e.g., to the next layer). In certain embodiments, the set of nodes in a particular layer may share the same parameters and instructions without interacting with each other. A machine learning module may be composed of at least one layer (e.g., ordered). Examples of types of layers include convolutional layers (e.g., layers with a kernel, a matrix of parameters Page 61 of 124
[0348] 13414664vlAttorney Docket No. : 2014191-0048
[0349] that is slid across an input to be multiplied with multiple input values to reduce them to a single output value); fully connected (FC) layers (e.g. all nodes are connected to all outputs of the previous layer); recurrent layers, long / short term memory (LSTM) layers, gated recurrent unit (GRU) layers (e.g., nodes with the various abilities to memorize and apply their previous inputs and / or outputs); batch normalization (BN) layers (e.g., layers that normalize a set of outputs from another layer, allowing for more independent learning of individual layers); activation layers (e.g., layers with nodes that only contain an activation function); and / or (un)pooling layers [e.g., layers that reduce (increase) dimensions of an input by summarizing (splitting) input values in defined patches).
[0350]
[0173] In certain embodiments, the performance of a machine learning module may be characterized by its ability to produce an output data with specific accuracy. To achieve specific accuracy, a training process is performed to find optimal parameters, such as weights, for each node in each layer of the machine learning module. In certain embodiments, the training process of a machine learning module may involve using output data to calculate an objective function (e.g., cost function, loss function, error function) that needs to be optimized (e.g., minimized, maximized). For example, a machine learning objective function may be a combination of a loss function and regularization parameter. The loss function is related to how well the output is able to predict the input. The loss function may take various forms, like mean squared error, mean absolute error, binary cross-entropy, categorical cross-entropy, for example. The regularization term may be needed to prevent overfitting and improve generalization of the training process. Examples of regularization techniques include LI Regularization or Lasso Regression, L2 Regularization or Ridge Regression, and Dropout (e.g., dropping layer outputs at random during training process).
[0351]
[0174] In certain embodiments, objective function optimization of a machine learning module may involve finding at least one (e.g., all) of the present global optima (e.g., as opposed to local optima). In certain embodiments, the algorithm for objective function optimization follows principles of mathematical optimization for a multi-variable function and relies on achieving specific accuracy of the process. Examples of objective function optimization algorithms include gradient descent, nonlinear conjugate gradient, random search, Levenberg-Marquardt algorithm, limited-memory Broyden-Fietcher-Goldfarb-Shanno algorithm, pattern search, basin hopping
[0352] Page 62 of 124
[0353] 13414664vlAttorney Docket No. : 2014191-0048
[0354] method, Krylov method, Adam method, genetic algorithm, particle swarm optimization, surrogate optimization, and simulated annealing.
[0355] [1751 Computations may be performed locally by a computing device. Computations performed over a network are also contemplated. FIG. 2 shows an illustrative network environment 200 for use in the methods and systems described herein. In brief overview, referring now to FIG. 2, a block diagram of an illustrative cloud computing environment 200 is shown and described. The cloud computing environment 200 may include one or more resource providers 202a, 202b, 202c (collectively, 202). Each resource provider 202 may include computing resources. In some implementations, computing resources may include any hardware and / or software used to process data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, illustrative computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 202 may be connected to any other resource provider 202 in the cloud computing environment 200. In some implementations, the resource providers 202 may be connected over a computer network 208. Each resource provider 202 may be connected to one or more computing device 204a, 204b, 204c (collectively, 204), over the computer network 208.
[0356]
[0176] The cloud computing environment 200 may include a resource manager 206. The resource manager 206 may be connected to the resource providers 202 and the computing devices 204 over the computer network 208. In some implementations, the resource manager 206 may facilitate the provision of computing resources by one or more resource providers 202 to one or more computing devices 204. The resource manager 206 may receive a request for a computing resource from a particular computing device 204. The resource manager 206 may identify one or more resource providers 202 capable of providing the computing resource requested by the computing device 204. The resource manager 206 may select a resource provider 202 to provide the computing resource. The resource manager 206 may facilitate a connection between the resource provider 202 and a particular computing device 204. In some implementations, the resource manager 206 may establish a connection between a particular resource provider 202 and a particular computing device 204. In some implementations, the resource manager 206 may redirect a particular computing device 204 to a particular resource provider 202 with the requested computing resource.
[0357] Page 63 of 124
[0358] 13414664vlAttorney Docket No. : 2014191-0048
[0359]
[0177] FIG. 3 shows an example of a computing device 300 and a mobile computing device 350 that can be used in the methods and systems described in this disclosure. The computing device 300 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 350 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting.
[0360]
[0178] The computing device 300 includes a processor 302, a memory 304, a storage device 306, a high-speed interface 308 connecting to the memory 304 and multiple high-speed expansion ports 310, and a low-speed interface 312 connecting to a low-speed expansion port 314 and the storage device 306. Each of the processor 302, the memory 304, the storage device 306, the high-speed interface 308, the high-speed expansion ports 310, and the low-speed interface 312, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 302 can process instructions for execution within the computing device 300, including instructions stored in the memory 304 or on the storage device 306 to display graphical information for a GUI on an external input / output device, such as a display 316 coupled to the high-speed interface 308. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Thus, as the term is used herein, where a plurality of functions are described as being performed by “a processor”, this encompasses embodiments wherein the plurality of functions are performed by any number of processors (e.g., one or more processors) of any number of computing devices (e g., one or more computing devices). Furthermore, where a function is described as being performed by “a processor”, this encompasses embodiments wherein the function is performed by any number of processors (e.g., one or more processors) of any number of computing devices (e.g., one or more computing devices) (e g., in a distributed computing system).
[0361] Page 64 of 124
[0362] 13414664vlAttorney Docket No. : 2014191-0048
[0363]
[0179] The memory 304 stores information within the computing device 300. In some implementations, the memory 304 is a volatile memory unit or units. In some implementations, the memory 304 is a non-volatile memory unit or units. The memory 304 may also be another form of computer-readable medium, such as a magnetic or optical disk.
[0364]
[0180] The storage device 306 is capable of providing mass storage for the computing device 300. In some implementations, the storage device 306 may be or contain a computer-readable medium, such as a hard disk device, an optical disk device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 302), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices such as computer- or machine-readable mediums (for example, the memory 304, the storage device 306, or memory on the processor 302).
[0365]
[0181] The high-speed interface 308 manages bandwidth-intensive operations for the computing device 300, while the low-speed interface 312 manages lower bandwidth-intensive operations. Such allocation of functions is an example only. In some implementations, the highspeed interface 308 is coupled to the memory 304, the display 316 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 310, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 312 is coupled to the storage device 306 and the low-speed expansion port 314. The low-speed expansion port 314, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
[0366]
[0182] The computing device 300 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 320, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 322. It may also be implemented as part of a rack server system 324. Alternatively, components from the computing device 300 may be combined with other components in a mobile device (not shown), such as a mobile computing device 350. Each of such devices may contain one or more of the computing device 300 and the mobile computing device Page 65 of 124
[0367] 13414664vlAttorney Docket No. : 2014191-0048
[0368] 350, and an entire system may be made up of multiple computing devices communicating with each other.
[0369] [1831 The mobile computing device 350 includes a processor 352, a memory 364, an input / output device such as a display 354, a communication interface 366, and a transceiver 368, among other components. The mobile computing device 350 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 352, the memory 364, the display 354, the communication interface 366, and the transceiver 368, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.
[0370]
[0184] The processor 352 can execute instructions within the mobile computing device 350, including instructions stored in the memory 364. The processor 352 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 352 may provide, for example, for coordination of the other components of the mobile computing device 350, such as control of user interfaces, applications run by the mobile computing device 350, and wireless communication by the mobile computing device 350.
[0371]
[0185] The processor 352 may communicate with a user through a control interface 358 and a display interface 356 coupled to the display 354. The display 354 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 356 may include appropriate circuitry for driving the display 354 to present graphical and other information to a user. The control interface 358 may receive commands from a user and convert them for submission to the processor 352. In addition, an external interface 362 may provide communication with the processor 352, so as to enable near area communication of the mobile computing device 350 with other devices. The external interface 362 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.
[0372]
[0186] The memory 364 stores information within the mobile computing device 350. The memory 364 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 374 may also be provided and connected to the mobile computing device 350 through an expansion interface 372, which may include, for example, a SIMM (Single In Line Memory Module) card Page 66 of 124
[0373] 13414664vlAttorney Docket No. : 2014191-0048
[0374] interface. The expansion memory 374 may provide extra storage space for the mobile computing device 350, or may also store applications or other information for the mobile computing device 350. Specifically, the expansion memory 374 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 374 may be provided as a security module for the mobile computing device 350, and may be programmed with instructions that permit secure use of the mobile computing device 350. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.
[0375]
[0187] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier and, when executed by one or more processing devices (for example, processor 352), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 364, the expansion memory 374, or memory on the processor 352). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 368 or the external interface 362.
[0376]
[0188] The mobile computing device 350 may communicate wirelessly through the communication interface 366, which may include digital signal processing circuitry where necessary. The communication interface 366 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 368 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth®, Wi-Fi™, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 370 may provide additional navigation- and location-related wireless data to the mobile computing device 350, which may be used as appropriate by applications running on the mobile computing device 350.
[0377] Page 67 of 124
[0378] 13414664vlAttorney Docket No. : 2014191-0048
[0379]
[0189] The mobile computing device 350 may also communicate audibly using an audio codec 360, which may receive spoken information from a user and convert it to usable digital information. The audio codec 360 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 350. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 350.
[0380]
[0190] The mobile computing device 350 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 380. It may also be implemented as part of a smart-phone 382, personal digital assistant, or other similar mobile device.
[0381]
[0191] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0382]
[0192] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0383]
[0193] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a Page 68 of 124
[0384] 13414664vlAttorney Docket No. : 2014191-0048
[0385] pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0386]
[0194] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0387]
[0195] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0388] Certain Exemplary Embodiments
[0389]
[0196] Without limitation to the foregoing description, the following is an enumerated list of non-limiting exemplary embodiments included in the present disclosure. Those of ordinary skill in the art will appreciate that one or more features discussed above may be included with or incorporated into any of the following numbered embodiments to form additional embodiments.
[0390] 1. A (e.g., computer-implemented) method of producing (e.g., forming) a circulating tumor fraction (cTF) estimation model [e.g., for estimating an amount of circulating tumor DNA in a sample (e.g., plasma sample) of a subject], the method comprising:
[0391] selecting (e.g., identifying) candidate regions (e.g., loci) in a genome (e.g., the human genome) [e.g., where methylation may be indicative of cancer (e.g., may correlate with cTF)];
[0392] Page 69 of 124
[0393] 13414664vlAttorney Docket No. : 2014191-0048
[0394] receiving (e.g., obtaining) sequencing data (e.g., methylation sequencing data) for a plurality of samples corresponding to the genome (e.g., for a plurality of human samples), wherein each of the samples has a known (e.g., empirically known, assumed, or expected) cTF;
[0395] normalizing counts of the sequencing data for each of the candidate regions for each of the samples;
[0396] determining signal regions (e.g., loci) in the genome [e.g., where methylation is indicative (e.g., predictive) of cancer (e.g., of ctDNA being present) (e.g., that correlates with cTF)] using the normalized counts for the candidate regions and the known cTF for the samples; and producing (e.g., fitting and / or training) a cTF estimation model [e.g., that determines an estimated cTF in a sample based on input sequencing data or data (e.g., one or more metrics) (e.g., one or more point estimates) derived therefrom] using at least the normalized counts for the signal regions and the known cTF for the samples.
[0397] 2. The method of embodiment 1, wherein the sequencing data has been derived from an enrichment method.
[0398] 3. The method of embodiment 1 or embodiment 2, wherein the sequencing data is methyl-binding domain sequencing (MBD-seq) data.
[0399] 4. The method of any one of embodiments 1-3, wherein the signal regions are regions determined to be cancer-specific differentially methylated regions (cDMRs).
[0400] 5. The method of any one of embodiments 1-4, wherein the candidate regions are regions that are or may be cancer-specific differentially methylated regions (cDMRs).
[0401] 6. The method of any one of embodiments 1-5, wherein the normalizing comprises, for each of the candidate regions for each of the samples, normalizing the counts based on number of fragments in the sequencing data for the sample and based on length of the candidate region.
[0402] 7. The method of any one of embodiments 1-6, wherein the normalizing comprises performing a normalization according to
[0403] Page 70 of 124
[0404] 13414664vlAttorney Docket No. : 2014191-0048
[0405] (f * (ci r+ dX\
[0406] logCPKM^ = log / l’r 7
[0407]
[0408] y Lr* LD j
[0409] where ci ris unique counts for sample i and region r, d is a constant (e.g., a non-zero constant, e.g., 1 ), / is a non-zero constant (e.g., 1000), LS is sample library size as defined by the number of unique fragments sequenced, and lris region length for region r.
[0410] 8. The method of any one of embodiments 1-7, wherein, for each of the candidate regions, the counts for the candidate region correspond to (e.g., are) a number of fragments in the sequencing data that overlap with candidate region.
[0411] 9. The method of any one of embodiments 1-8, wherein the normalizing comprises accounting for copy number alterations reflected in the sequencing data (e.g., using shallow whole genome sequencing data).
[0412] 10. The method of any one of embodiments 1-9, wherein the counts are pseudocounts [e.g., offset true counts that are offset (e.g., by uniform addition) in order to avoid errors subsequent in model fitting and / or normalization],
[0413] 11. The method of any one of embodiments 1-10, wherein the counts are per-region-based counts (e.g., MBD-seq counts) (e.g., not per-site counts) that indicate that a region is methylated.
[0414] 12. The method of any one of embodiments 1-11, wherein the signal regions are determined based on the normalized counts for a candidate region having a linear relationship with cTF and the normalized counts for the candidate region for samples at a threshold cTF (e.g., a low threshold) being higher than an upper quantile or lower than a lower quantile for healthy samples.
[0415] 13. The method of embodiment 12, wherein the threshold cTF is in a range of from 0.5 to 2% cTF (e.g., of from 1-1.5% cTF or of 1% cTF).
[0416] 14. The method of embodiment 12 or embodiment 13, wherein the normalized counts for the candidate region for the samples at the threshold cTF being higher than an upper quantile for Page 71 of 124
[0417] 13414664vlAttorney Docket No. : 2014191-0048
[0418] healthy samples with the upper quantile being in a range of 85%-99% (e.g., is 90%, 95%, 97%, or 99%) is used as a condition for determining the signal regions or wherein the normalized counts for the candidate region for the samples at the threshold cTF being lower than a lower quantile for healthy samples with the lower quantile being in a range of 1%- 15% (e.g., is 10%, 5%, 3%, or 1%) is used as a condition for determining the signal regions.
[0419] 15. The method of any one of embodiments 1-14, wherein determining the signal regions comprises determining that ones of the candidate regions are differentially (e.g., hypo- or hyper-) methylated over a range of cTF (e.g., at range of at least from 3% to 30% cTF, at least from 2% to 30% cTF, at least from 1% to 30% cTF, at least from 0.5% to 30% cTF).
[0420] 16. The method of any one of embodiments 1-15, wherein determining the signal regions comprises determining that one or more of the candidate regions have differential (e.g., hypo- or hyper-) methylation observable down to a target cTF in a range of from 0.5% to 3% cTF (e.g., no more than 2% cTF, no more than 1% cTF, no more than 0.5% cTF, or in a range of from 0.5% to 1% cT).
[0421] 17. The method of any one of embodiments 1-16, wherein determining the signal regions comprises determining (e.g., for each of the candidate regions) that there exists a linear relationship between the normalized counts for a candidate region and cTF.
[0422] 18. The method of any one of embodiments 1-17, wherein the signal regions correspond to (e.g., are, contain, or overlap with) CpG islands in the genome.
[0423] 19. The method of any one of embodiments 1-18, wherein determining the signal regions comprises, for each of the candidate regions, determining a relationship between the normalized counts for the candidate region and the known cTF for the samples.
[0424] 20. The method of any one of embodiments 1-19, wherein producing the model is based on, for each of the plurality of samples, a point estimate of normalized counts for the signal regions for the sample, preferably wherein the point estimate is the median.
[0425] Page 72 of 124
[0426] 13414664vlAttorney Docket No. : 2014191-0048
[0427] 21. The method of any one of embodiments 1-20, wherein producing the model comprises fitting a cTF regression for each of the signal regions.
[0428] 22. The method of any one of embodiments 1-21, wherein the cTF estimation model comprises an expectation maximization model.
[0429] 23. The method of embodiment 22, wherein the expectation maximization model is structured to estimate cTF by maximizing expectation across the signal regions independently.
[0430] 24. The method of embodiment 22 or embodiment 23, wherein the expectation maximization model is structured to estimate cTF by maximizing across the signal regions simultaneously.
[0431] 25. The method of any one of embodiments 1-24, wherein producing the model comprises producing (e.g., defining and / or fitting) a set of relationships (e.g., regressions) between signal regions and cTF such that an expectation maximization can be performed across the set of relationships (e.g., on a per-region basis).
[0432] 26. The method of any one of embodiments 1-25, wherein producing the model comprises defining a global loss function for a set of relationships (e.g., regressions) between cTF and normalized counts for a signal region for each of the signal regions.
[0433] 27. The method of any one of embodiments 1-26, wherein the model comprises a multivariate, regularized, logistic, and / or tree-based regression (e.g., that uses a feature matrix).
[0434] 28. The method of any one of embodiments 1-27, wherein the model comprises a feature aggregation and regression model that uses a per-sample-based point estimate (e.g., sum, mean, or median) (e.g., for the signal regions, optionally and the noise regions).
[0435] Page 73 of 124
[0436] 13414664vlAttorney Docket No. : 2014191-0048
[0437] 29. The method of any one of embodiments 1 -28, wherein the model comprises a deep learning model (e.g., an artificial neural network) and producing the model comprises training the deep learning model using tensors derived from the normalized counts for the signal regions.
[0438] 30. The method of any one of embodiments 1-29, wherein the cTF estimation model comprises an ensemble model (e.g., comprising two, three, or more constituent models).
[0439] 31. The method of any one of embodiments 1-30, wherein producing the cTF estimation model comprises producing a plurality of constituent models to produce an ensemble model (e.g., comprising two, three, or more constituent models).
[0440] 32. The method of embodiment 30 or embodiment 31, wherein the ensemble model uses a geometric mean of constituent models to estimate cTF.
[0441] 33. The method of any one of embodiments 30-32, wherein the ensemble model comprises (e g., the plurality of constituent models comprises) a first model that estimates cTF based on all of the signal regions together (e.g., a point estimate regression model) and a second model that estimates cTF based on each of the signal regions individually (e.g., an expectation maximization model).
[0442] 34. The method of any one of embodiments 30-33, wherein the ensemble model comprises (e.g., the plurality of constituent models comprises) a first model that estimates cTF based on a metric determined from all of the signal regions together [e.g., a point estimate (e.g., median) of normalized counts for all of the signal regions] and a second model that estimates cTF based on individual metrics for each of the signal regions (e.g., an expectation maximization model).
[0443] 35. The method of any one of embodiments 30-34, wherein the ensemble model comprises an expectation maximization model and (i) a signal-to-noise ratio (SNR) based model (e.g., an SNR regression model) and / or (ii) a point estimate regression based model (e.g., wherein producing the ensemble model comprises producing an expectation maximization model and (i) a signal-to-noise
[0444] Page 74 of 124
[0445] 13414664vlAttorney Docket No. : 2014191-0048
[0446] ratio (SNR) based model (e.g., an SNR regression model) and / or (ii) a point estimate regression based model).
[0447] 36. The method of any one of embodiments 1-35, wherein producing the model comprises independently fitting the signal regions (e.g., in an expectation maximization model).
[0448] 37. The method of any one of embodiments 1-36, wherein the model determines estimated cTF based on independent consideration of normalized counts for the signal regions in a sample.
[0449] 38. The method of any one of embodiments 1-37, wherein producing the model comprises selecting a type of model based on cancer type for ones of the samples.
[0450] 39. The method of any one of embodiments 1-38, wherein producing the model comprises producing a pan-cancer ensemble model comprising at least one constituent model for a plurality of cancer types (e.g., wherein each of the at least one constituent models comprises an ensemble model).
[0451] 40. The method of any one of embodiments 1-39, wherein selecting the candidate regions comprises (e.g., consists of) selecting a set of CpG islands (e.g., reference CpG islands) [e.g., CpG islands that have been identified (e g., named) by a reference source (e.g., UCSC defined CpG islands)].
[0452] 41. The method of any one of embodiments 1-40, wherein the candidate regions are CpG islands.
[0453] 42. The method of any one of embodiments 1-39, wherein selecting the candidate regions comprises (e.g., consists of) selecting (e.g., overlapping or non-overlapping) bins of uniform size within the genome.
[0454] 43. The method of any one of embodiments 1-39, wherein the candidate regions are bins of uniform size.
[0455] Page 75 of 124
[0456] 13414664vlAttorney Docket No. : 2014191-0048
[0457] 44. The method of any one of embodiments 1-39, wherein selecting the candidate regions comprises (e.g., consists of) determining regions having differential methylation between conditions (e.g., healthy state and cancer state).
[0458] 45. The method of embodiment 44, wherein the regions having differential methylation between conditions are determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples [e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% cTF (e.g., as determined by an ichor method)] (e.g., derived from one or more cell lines).
[0459] 46. The method of any one of embodiments 1-45, comprising determining noise regions in the genome (e.g., that indicate correlation only in healthy samples of the plurality of samples) (e.g., where there is correlation between normalized counts in healthy samples and no correlation between normalized counts in cancerous samples) (e.g., that are indicative of batch effects in the sequencing data, before or after normalization).
[0460] 47. The method of embodiment 46, wherein determining the noise regions comprises, for each of the noise regions, determining that a metric (e g., normalized counts) for samples having a known cTF at or below a low threshold (e.g., of 0) has a positive correlation and that the metric for samples having a known cTF above a threshold is uncorrelated.
[0461] 48. The method of embodiment 46 or embodiment 47, wherein producing the model is further based on normalized counts of the sequencing data for the noise regions.
[0462] 49. The method of any one of embodiments 46-48, comprising calculating a signal-to-noise ratio for each of the samples, wherein the signal-to-noise ratio is based on a point estimate (e.g., median) count value across each of the noise regions for a sample and a point estimate (e.g., median) count value across each of the signal regions for the sample (e.g., wherein producing the model is based on determining a relationship between the signal-to-noise ratio and the known cTF for each of the samples).
[0463] Page 76 of 124
[0464] 13414664vlAttorney Docket No. : 2014191-0048
[0465] 50. The method of any one of embodiments 1-49, wherein the plurality of samples are liquid biopsy samples.
[0466] 51. The method of any one of embodiments 1-50, wherein the plurality of samples are plasma samples (e.g., comprising empirical and / or simulated plasma samples).
[0467] 52. The method of any one of embodiments 1-51, wherein the plurality of samples comprise in silico diluted (e.g., titrated) samples (e.g., in silico plasma (ISP) samples).
[0468] 53. The method of any one of embodiments 1-52, wherein the sequencing data for the in silico diluted samples is a random sampling of fragments from a healthy (e.g., zero or below limit of detection cTF) sample and a sample with a known non-zero cTF in a predetermined ratio of amounts.
[0469] 54. The method of any one of embodiments 1-53, wherein the plurality of samples comprises samples spanning a range of known cTF (e.g., from 0% cTF to 100% cTF, from 1% cTF to 100% cTF, from 2% cTF to 100%, or from 3% cTF to 100% cTF).
[0470] 55. The method of any one of embodiments 1-54, wherein, for at least one of the plurality of samples, the known cTF is known from an orthogonal reference method (e.g., ichorCNA).
[0471] 56. The method of any one of embodiments 1-55, wherein, for at least one of the plurality of samples, the sample is a simulated dilution and the known cTF for the sample has been determined by linear combination of known cTFs of reference samples used to generate the sample.
[0472] 57. The method of any one of embodiments 1-56, wherein the samples correspond to a single cancer.
[0473] Page 77 of 124
[0474] 13414664vlAttorney Docket No. : 2014191-0048
[0475] 58. The method of any one of embodiments 1-57, wherein the samples comprise different samples corresponding to different cancers (e.g., wherein the cTF estimation model is a pan-cancer estimation model).
[0476] 59. A (e.g., computer-implemented) method of estimating cTF in a sample, the method comprising:
[0477] providing sequencing data for a liquid sample and a cTF estimation model that has been formed by a method according to any one of embodiments 1-58; and
[0478] estimating cTF in the sample using the cTF estimation model.
[0479] 60. A (e.g., computer-implemented) method of estimating cTF in a sample, the method comprising:
[0480] receiving (e.g., obtaining) sequencing data for a (e.g., liquid) sample (e g., a liquid plasma sample) from a subject;
[0481] producing normalized counts of the sequencing data for signal regions in the genome of the subject; and
[0482] estimating cTF for the sample with a cTF estimation model based at least on the normalized counts.
[0483] 61. The method of embodiment 60, wherein the sequencing data has been derived from an enrichment method.
[0484] 62. The method of embodiment 60 or embodiment 61, wherein the sequencing data is methyl-binding domain sequencing (MBD-seq) data.
[0485] 63. The method of any one of embodiments 60-62, wherein the signal regions are regions determined to be cancer-specific differentially methylated regions (cDMRs).
[0486] 64. The method of any one of embodiments 60-63, wherein the normalizing comprises, for each of the signal regions, normalizing the counts based on number of fragments in the sequencing data for the sample and based on length of the signal region.
[0487] Page 78 of 124
[0488] 13414664vlAttorney Docket No. : 2014191-0048
[0489] 65. The method of any one of embodiments 60-64, wherein producing the normalized counts comprises, for each of the signal regions, performing a normalization according to
[0490] f * (cr+ d) \
[0491] logCPKMr= log
[0492]
[0493] lr* LS )
[0494] where cris unique counts in the sequencing data for the sample in region r, d is a constant (e.g., a non-zero constant, e.g., 1), / is anon-zero constant (e g., 1000), LS is sample library size as defined by the number of unique fragments sequenced for the sample, and lris region length for region r.
[0495] 66. The method of any one of embodiments 60-65, wherein, for each of the signal regions, the counts for the signal region are a number of fragments in the sequencing data that overlap with signal region.
[0496] 67. The method of any one of embodiments 60-66, wherein estimating the cTF comprises estimating a per-region-based estimated cTF for each of the signal regions individually and then determining the estimated cTF based on the per-region-based estimated cTF using the cTF estimation model (e.g., using an expectation maximization model thereof).
[0497] 68. The method of any one of embodiments 60-67, wherein estimating the cTF comprises performing an expectation maximization of cTF for the sample with the cTF estimation model based on the normalized counts (e.g., with an expectation maximization model thereof that has been trained to maximize expectation on a per-region basis).
[0498] 69. The method of any one of embodiments 60-68, wherein estimating the cTF comprises determining a point estimate of the normalized counts for the signal regions for the sample and the cTF estimation model uses the point estimate to estimate the cTF for the sample (e.g., in one or more constituent models of the cTF estimation model).
[0499] 70. The method of embodiment 69, wherein the point estimate is the median.
[0500] Page 79 of 124
[0501] 13414664vlAttorney Docket No. : 2014191-0048
[0502] 71. The method of any one of embodiments 60-70, wherein estimating the cTF comprises estimating a plurality of preliminary estimated cTFs comprising a per-region-based estimated cTF estimated using a first model of the cTF estimation model and a per-sample-based estimated cTF estimated using a second model of the cTF estimation model, wherein the cTF is estimated by a combination (e.g., geometric mean) that includes (e.g., is of) the per-region-based estimated cTF and the per-sample-based estimated cTF.
[0503] 72. The method of any one of embodiments 60-71, comprising:
[0504] producing normalized counts of the sequencing data for noise regions in the genome of the subject; and
[0505] determining a signal-to-noise ratio (SNR) for the sample based on the normalized counts for the noise regions and the normalized counts for the signal regions,
[0506] wherein the cTF estimation model uses the SNR to estimate the cTF [e.g., produces an SNR-based preliminary estimated cTF that is combined with one or more other preliminary estimated cTFs (e.g., a per-region-based estimated cTF and / or per-sample-based estimated cTF) to estimate the cTF for the sample],
[0507] 73. The method of embodiment 72, wherein producing the normalized counts for the noise regions comprises, for each of the noise regions, normalizing counts for the noise region based on number of fragments in the sequencing data for the sample and based on length of the noise region.
[0508] 74. The method of embodiment 72 or embodiment 73, wherein producing the normalized counts for the noise regions comprises, for each of the noise regions, performing a normalization according to
[0509] (f * + d)\
[0510] logCPKMr= log
[0511]
[0512] \ * ZjiJ /
[0513] where cris unique counts in the sequencing data for the sample in region r, d is a constant (e.g., a non-zero constant, e.g., 1), / is anon-zero constant (e.g., 1000), LS is sample library size as defined by the number of unique fragments sequenced for the sample, and lris region length for region r.
[0514] Page 80 of 124
[0515] 13414664vlAttorney Docket No. : 2014191-0048
[0516] 75. The method of any one of embodiments 72-74, wherein the SNR is determined as a ratio of a point estimate of the normalized counts for the signal regions and a point estimate of the normalized counts for the noise regions.
[0517] 76. The method of embodiment 75, wherein the point estimate for the signal regions and the point estimate for the noise regions is the median.
[0518] 77. The method of any one of embodiments 68-76, wherein the cTF estimation model combines a plurality of preliminary estimated cTFs, each determined by a separate constituent model of the cTF estimation model, to estimate the cTF.
[0519] 78. The method of embodiment 77, wherein the combining comprises determining a geometric mean of the preliminary estimated cTFs.
[0520] 79. The method of embodiment 77 or embodiment 78, wherein the preliminary estimated cTFs comprise a per-region-based estimated cTF and a per-sample-based estimated cTF.
[0521] 80. The method of embodiment 79, wherein the per-sample-based estimated cTF is determined using a first constituent model that is a point estimate based model.
[0522] 81. The method of embodiment 79 or embodiment 80, wherein the per-region-based estimated cTF is determined using an expectation maximization model.
[0523] 82. The method of any one of embodiments 75-81, wherein the preliminary estimated cTFs comprise an SNR-based preliminary estimated cTF.
[0524] 83. The method of any one of embodiments 60-82, wherein estimating the cTF comprises using a first model of the cTF estimation model to detect if ctDNA is present in the sample and subsequently using a second model of the cTF estimation model to estimate the cTF if ctDNA is present or estimating the cTF at 0 if ctDNA is determined to be not present using the first model.
[0525] Page 81 of 124
[0526] 13414664vlAttorney Docket No. : 2014191-0048
[0527] 84. The method of any one of embodiments 60-83, wherein estimating the cTF comprises using a first model of the cTF estimation model to detect that ctDNA is present in the sample and, upon determining that ctDNA is present, using a second model of the cTF estimation model to estimate the cTF.
[0528] 85. The method of embodiment 83 or embodiment 84, wherein the first model comprises an ensemble model that uses a geometric mean of preliminary estimated cTFs and the second model is an expectation maximalization model (e.g., and the sample is of a subject suspected to have, that has, or that has had breast cancer or lung cancer) (e.g., and the cTF estimation model is a breast cancer cTF estimation model or lung cancer cTF estimation model).
[0529] 86. The method of embodiment 83 or embodiment 84, wherein the first model is a signal-to-noise ratio (SNR) based model and the second model is an expectation maximalization model (e.g., and the sample is of a subject suspected to have, that has, or that has had breast cancer or lung cancer) (e.g., and the cTF estimation model is a prostate cancer cTF estimation model).
[0530] 87. The method of any one of embodiments 83-86, wherein using the first model comprises estimating a preliminary estimated cTF and the second model used to estimate the cTF depends on the preliminary estimated cTF (e.g., using one model to estimate the cTF if the preliminary estimated cTF is above a threshold and using another model to estimate the cTF if the preliminary estimated cTF is below a threshold).
[0531] 88. The method of any one of embodiments 60-87, wherein the model comprises a multivariate, regularized, logistic, and / or tree-based regression (e.g., that uses a feature matrix).
[0532] 89. The method of any one of embodiments 60-88, wherein the model comprises a feature aggregation and regression model that uses a per-sample-based point estimate (e.g., sum, mean, or median) (e.g., for the signal regions).
[0533] Page 82 of 124
[0534] 13414664vlAttorney Docket No. : 2014191-0048
[0535] 90. The method of any one of embodiments 60-89, wherein the model comprises a deep learning model (e.g., an artificial neural network) that has been trained using tensors derived from normalized sequencing data for the signal regions.
[0536] 91. The method of any one of embodiments 60-90, wherein the cTF estimation model comprises an ensemble model.
[0537] 92. The method of any one of embodiments 60-91, wherein the model uses the signal regions independently to obtain the estimated cTF.
[0538] 93. The method of any one of embodiments 60-92, wherein the model determines estimated cTF based on independent consideration of normalized counts for the signal regions in a sample.
[0539] 94. The method of any one of embodiments 60-93, wherein the model has been trained using data from cancer-specific samples and the subject is suspected of having and / or known to have or have had the cancer to which the model is specific.
[0540] 95. The method of any one of embodiments 60-94, wherein the model has been trained using data from samples having different cancers (e g., wherein the model is a pan-cancer model).
[0541] 96. The method of any one of embodiments 60-95, wherein the normalizing comprises accounting for copy number alterations reflected in the sequencing data (e.g., using shallow whole genome sequencing data).
[0542] 97. The method of any one of embodiments 60-96, wherein the counts are pseudocounts [e.g., offset true counts that are offset (e.g., by uniform addition) in order to avoid errors subsequent in model fitting and / or normalization],
[0543] 98. The method of any one of embodiments 60-97, wherein the counts are per-region-based counts (e.g., MBD-seq counts) (e.g., not per-site counts) that indicate that a region is methylated.
[0544] Page 83 of 124
[0545] 13414664vlAttorney Docket No. : 2014191-0048
[0546] 99. The method of any one of embodiments 60-98, wherein the sample is a liquid biopsy sample.
[0547] 100. The method of any one of embodiments 60-99, wherein the sample is a plasma sample.
[0548] 101. A system comprising a processor and one or more non-transitory computer readable media (e.g., a memory) having instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising the method according to any one of embodiments 1-100.
[0549] 102. One or more non-transitory computer readable media having instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising the method according to any one of embodiments 1-100.
[0550] 103. A method of characterizing cancer recurrence and / or progression, the method comprising:
[0551] estimating cTF in a first sample for a subject taken at a first time point using a method according to any one of embodiments 59-100;
[0552] estimating cTF in a second sample for the subject taken at a second time point after the first time point using a method according to any one of embodiments 59-100; and determining a difference in the estimated cTF at the second time point and at the first time point.
[0553] 104. A method of monitoring cancer in a subject, the method comprising:
[0554] estimating cTF in a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of embodiments 59-100; and
[0555] determining whether there is a difference in cTF for the samples over time.
[0556] 105. The method of embodiment 103 or embodiment 104, comprising administering a therapy (e.g., initiating administration of the therapy) to the subject when the difference is determined to be at least as large as a threshold difference.
[0557] Page 84 of 124
[0558] 13414664vlAttorney Docket No. : 2014191-0048
[0559] 106. The method of any one of embodiments 103-105, comprising altering administration of a therapy to the subject when the difference is determined to be at least as large as a threshold difference.
[0560] 107. The method of embodiment 106, wherein altering administration comprises increasing a dosage and / or frequency of administration.
[0561] 108. A method of prognosing cancer, the method comprising:
[0562] estimating cTF in a sample for a subject using a method according to any one of embodiments 59-100; and
[0563] prognosing cancer (e.g., a stage of cancer, e.g., as early or late stage) in the subject based on the estimated cTF in the sample.
[0564] 109. The method of embodiment 108, comprising administering a therapy based on the prognosis.
[0565] 110. A method of diagnosing cancer in a subject, the method comprising:
[0566] estimating cTF in a sample for a subject using a method according to any one of embodiments 59-100; and
[0567] determining that the estimated cTF exceeds a threshold.
[0568] 111. The method of embodiment 110, comprising initiating administration of a therapy based on the estimated cTF.
[0569] 112. The method of embodiment 101, comprising selecting a dosing regimen for the therapy based on the estimated cTF.
[0570] 113. A method of determining whether a cancer has been removed from a subject, the method comprising, after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, estimating cTF in a sample for the subject using a method according to any one of embodiments 59-100.
[0571] Page 85 of 124
[0572] 13414664vlAttorney Docket No. : 2014191-0048
[0573] 114. The method of embodiment 113, comprising continuing administration of a therapy based on the estimated cTF (e.g., based on a determination that the cancer has not been sufficiently removed based on the estimated cTF).
[0574] 115. The method of embodiment 114, comprising ceasing administration of a therapy based on the estimated cTF (e.g., based on a determination that the cancer has been sufficiently removed based on the estimated cTF).
[0575] 116. A method of predicting a tumor of origin for ctDNA in a sample from a subj ect, the method comprising:
[0576] estimating cTF in a sample for a subject using a first method according to any one of embodiments 59-100, wherein the cTF estimation model for the first method has been trained for a first type of cancer [e.g., using only healthy samples and samples corresponding to the first type of cancer (e.g., pure or in silico diluted samples)];
[0577] estimating cTF in the using a second method according to any one of embodiments 59-100, wherein the cTF estimation model for the second method has been trained for a second type of cancer [e.g., using only healthy samples and samples corresponding to the second type of cancer (e.g., pure or in silico diluted samples)]; and
[0578] preliminarily determining a tumor of origin based on a difference in the cTF estimated using the first method and the cTF estimated using the second method (e.g., based on the cTF estimated using the first method exceeding a threshold difference with the cTF estimated using the second method, and optionally also exceeds a threshold value) (e.g., based on which of the cTF estimated using the first method and the cTF estimated using the second method is higher and which is lower).
[0579] 117. The method of embodiment 116, comprising selecting and performing a subsequent diagnostic assessment of the subject to determine the tumor of origin based on the preliminary determination of the tumor of origin.
[0580] 118. A method of making a preliminary diagnosis of cancer in a subject, the method comprising:
[0581] Page 86 of 124
[0582] 13414664vlAttorney Docket No. : 2014191-0048
[0583] estimating cTF in a sample for a subject using a first method according to any one of embodiments 59-100, wherein the cTF estimation model for the first method has been trained for a first type of cancer [e.g., using only healthy samples and samples corresponding to the first type of cancer (e.g., pure or in silica diluted samples)];
[0584] estimating cTF in the using a second method according to any one of embodiments 59-100, wherein the cTF estimation model for the second method has been trained for a second type of cancer [e.g., using only healthy samples and samples corresponding to the second type of cancer (e g., pure or in silica diluted samples)]; and
[0585] preliminarily diagnosing a cancer in the subject based on a difference in the cTF estimated using the first method and the cTF estimated using the second method (e.g., based on the cTF estimated using the first method exceeding a threshold difference with the cTF estimated using the second method, and optionally also exceeds a threshold value) (e.g., based on which of the cTF estimated using the first method and the cTF estimated using the second method is higher and which is lower).
[0586] 119. The method of embodiment 118, comprising selecting and performing a subsequent diagnostic assessment of the subject for a particular cancer, wherein the particular cancer is selected based on the preliminary diagnosis.
[0587] 120. A method of producing a circulating tumor fraction (cTF) estimation model for estimating an amount of circulating tumor DNA in a sample of a subject, the method comprising:
[0588] selecting candidate regions in a genome where epigenetic modification of the candidate regions may be indicative of cancer;
[0589] receiving epigenetic modification sequencing data for a plurality of samples corresponding to the genome, wherein each of the samples has a known cTF;
[0590] normalizing counts of the sequencing data for each of the candidate regions for each of the samples; determining signal regions in the genome where epigenetic modification is indicative of cancer using the normalized counts for the candidate regions and the known cTF for the samples; and
[0591] producing a cTF estimation model that determines an estimated cTF in a sample based on input sequencing data or data derived therefrom, wherein the producing comprises using at least Page 87 of 124
[0592] 13414664vlAttorney Docket No. : 2014191-0048
[0593] the normalized counts for the signal regions and the known cTF for the samples to produce the cTF estimation model.
[0594] 121. The method of embodiment 120, wherein the epigenetic modification comprises histone modification, DNA methylation, transcription factor binding and / or chromatin accessibility. 123. The method of embodiment 120 or 121, wherein the epigenetic modification comprises histone modification.
[0595] 124. The method of any one of embodiments 120-123, wherein the epigenetic modification comprises DNA methylation.
[0596] 125. The method of embodiment 123, wherein the histone modification comprises H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, or pan acetylation.
[0597] 126. The method aby one of embodiments 121-125, wherein the histone modification comprises H3K4me3.
[0598] 127. The method of any one of embodiments 121-126, wherein the histone modification comprises H3K27ac.
[0599] 128. The method of any one of embodiments 121-127, wherein the histone modification comprises H3K4me3 and H3K27ac.
[0600] 129. The method of any one of embodiments 120-128, wherein the signal regions are regions determined to be cancer-specific regions with differential histone modifications (cDHMRs).
[0601] 130. The method of any one of embodiments 120-129, wherein the candidate regions are regions determined to be cancer-specific regions with differential histone modifications (cDHMRs).
[0602] 131. The method of any one of embodiments 120-130, wherein the counts are per-region-based counts that indicate that a region has been bound by a differentially modified histone.
[0603] Page 88 of 124
[0604] 13414664vlAttorney Docket No. : 2014191-0048
[0605] 132. The method of any one of embodiments 120-131, wherein determining the signal regions comprises determining that ones of the candidate regions have been bound by a differentially modified histone over a range of cTF.
[0606] 133. The method of any one of embodiments 1-132, wherein determining the signal regions comprises determining that one or more of the candidate regions have differential histone modification observable down to a target cTF in a range of from 0.5% to 3% cTF.
[0607] 134. The method of any one of embodiments 1-133, wherein producing the model comprises producing a pan-cancer ensemble model comprising at least one constituent model for a plurality of cancer types.
[0608] 135. The method of embodiment 134, wherein the plurality of cancer types comprises lung cancer, breast cancer, colorectal cancer, gastric cancer (e.g., gastroesophageal cancer), and / or prostate cancer.
[0609] 136. The method of any one of embodiments 1-135, wherein selecting the candidate regions comprises determining regions having differential epigenetic modification between conditions.
[0610] 137. The method of any one of embodiments 1-136, wherein selecting the candidate regions comprises determining regions having differential histone modification between conditions.
[0611] 138. The method of embodiment 136 or embodiment 137, wherein the regions having differential epigenetic modification between conditions are determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples.
[0612] 139. The method of any one of embodiments 136-138, wherein the regions having differential histone modification between conditions are determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples.
[0613] 140. A method of estimating cTF in a sample, the method comprising:
[0614] receiving sequencing data for a liquid sample from a subject;
[0615] Page 89 of 124
[0616] 13414664vlAttorney Docket No. : 2014191-0048
[0617] producing normalized counts of the sequencing data for signal regions in the genome of the subject; and
[0618] estimating cTF for the sample with a cTF estimation model based at least on the normalized counts.
[0619] 141. The method of embodiment 140, wherein the sequencing data has been derived from an enrichment method.
[0620] 142. The method of embodiment 140 or embodiment 141, wherein the sequencing data is epigenetic modification sequencing data.
[0621] 143. The method of embodiment 142, wherein the epigenetic modification sequencing data comprises histone modification sequencing data, DNA methylation sequencing data, transcription factor binding sequencing data and / or chromatin accessibility sequencing data.
[0622] 144. The method of any one of embodiments 140-143, wherein the sequencing data is chromatin immunoprecipitation sequencing (ChlP-seq) data.
[0623] 145. The method of embodiment 144, wherein the ChlP-seq data comprises sequencing data for one or more histone modifications.
[0624] 146. The method of embodiment 145, wherein the one or more histone modifications comprise H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, orH3K4me3, or pan acetylation.
[0625] 147. The method of embodiment 145 or embodiment 146, wherein the one or more histone modifications comprise H3K4me3.
[0626] 148. The method of any one of embodiments 145-147, wherein the one or more histone modifications comprise H3K27ac.
[0627] Page 90 of 124
[0628] 13414664vlAttorney Docket No. : 2014191-0048
[0629] 149. The method of any one of embodiments 145-148, wherein the one or more histone modifications comprise H3K4me3 and H3K27ac.
[0630] 150. The method of any one of embodiments 140-149, wherein the sequencing data is methyl-binding domain sequencing (MBD-seq) data.
[0631] 151. The method of embodiment 150, wherein the MBD-seq data comprises sequencing data for DNA methylation.
[0632] 152. The method of any one of embodiments 145-151, wherein the signal regions are regions determined to be cancer-specific regions with differential histone modifications (DHMRs).
[0633] Definitions
[0634]
[0197] “A” or “An”: The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” refers to one element or more than one element.
[0635]
[0198] About: The term “about”, when used herein in reference to a value, refers to a value that is similar, in context, to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” can encompass a range of values that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or within a fraction of a percent, of the referenced value.
[0636]
[0199] Administration: As used herein, the term “administration” typically refers to the administration of a disease appropriate (e.g., appropriate for administration to a subject having a certain activity in a gene relevant for an indication (e.g., a gene overexpressed and / or overactive in subjects diagnosed with an indication)) treatment. In some embodiments, the disease appropriate treatment may include administering a composition to a subject, for example to achieve delivery of an agent that is, is included in, or is otherwise delivered by, the composition. In some embodiments, the disease appropriate treatment may include administering an appropriate surgical procedure or radiological procedure, optionally in combination with administration of a composition.
[0637] Page 91 of 124
[0638] 13414664vlAttorney Docket No. : 2014191-0048
[0639]
[0200] Agent: As used herein, the term “agent” may refer to any chemical or physical entity, including without limitation any of one or more of an atom, e.g., a radioactive atom, molecule, compound, conjugate, polypeptide, polynucleotide, polysaccharide, lipid, cell, or combination or complex thereof.
[0640]
[0201] Antibody: As used herein, the term “antibody” refers to a polypeptide that includes one or more canonical immunoglobulin sequence elements sufficient to confer specific binding to a particular antigen (e.g., a heavy chain variable domain, a light chain variable domain, and / or one or more CDRs). Thus, the term antibody includes, without limitation, human antibodies, nonhuman antibodies, synthetic and / or engineered antibodies, fragments thereof, and agents including the same. Antibodies can be naturally occurring immunoglobulins (e.g., generated by an organism reacting to an antigen). Synthetic, non-naturally occurring, or engineered antibodies can be produced by recombinant engineering, chemical synthesis, or other artificial systems or methodologies known to those of skill in the art.
[0641]
[0202] Associated with: Two events or entities are “associated” with one another, as that term is used herein, if the presence, level and / or form of one is correlated with that of the other. For example, a particular entity (e g., an epigenetic profile comprising one or more histone modifications at a set of genomic loci, etc.) is considered to be associated with a particular disease, disorder, or condition, if its presence, level and / or form correlates with incidence of and / or susceptibility to the disease, disorder, or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically “associated” with one another if they interact, directly or indirectly, so that they are and / or remain in physical proximity with one another. In some embodiments, two or more entities that are physically associated with one another are covalently linked to one another; in some embodiments, two or more entities that are physically associated with one another are not covalently linked to one another but are non-covalently associated, for example by means of hydrogen bonds, van der Waals interaction, hydrophobic interactions, magnetism, or a combination thereof.
[0642]
[0203] “Between” or “From”: As used herein, the term “between” refers to content that falls between indicated upper and lower, or first and second, boundaries, inclusive of the boundaries. Similarly, the term “from”, when used in the context of a range of values, indicates that the range includes content that falls between indicated upper and lower, or first and second, boundaries, inclusive of the boundaries.
[0643] Page 92 of 124
[0644] 13414664vlAttorney Docket No. : 2014191-0048
[0645]
[0204] Biological Sample: As used herein, the term “biological sample” typically refers to a sample obtained or derived from a biological source (e.g., a tissue or organism or cell) of interest, as described herein. In some embodiments, a biological source is or includes an organism, such as a human subject. In some embodiments, a biological sample is or includes a biological tissue or fluid. In some embodiments, a biological sample can be or include cells, tissue, or bodily fluid. “Bodily fluids” refer to fluids that are excreted or secreted from the body as well as fluids that are normally not (e.g., blood, serum, plasma, Cowper’s fluid or pre-ejaculate fluid, chyle, chyme, stool, interstitial fluid, intracellular fluid, lymph, menses, saliva, sebum, semen, serum, sweat, synovial fluid, tears, urine, vitreous humor, vomit). In some embodiments, a biological sample can be or include blood, blood components, cell-free DNA (cfDNA), circulating-tumor DNA (ctDNA), ascites, biopsy samples, surgical specimens, cell-containing body fluids, sputum, saliva, feces, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, lymph, gynecological fluids, secretions, excretions, skin swabs, vaginal swabs, oral swabs, nasal swabs, washings or lavages such as a ductal lavages or bronchoalveolar lavages, aspirates, scrapings, or bone marrow. In some embodiments, a biological sample is a liquid biopsy sample obtained from a bodily fluid. In some embodiments, a biological sample is or includes DNA obtained from a single subject or from a plurality of subjects. A biological sample can be a “primary sample” obtained directly from a biological source or can be a “processed sample”, i.e., a sample that was derived from a primary sample, e.g., via dilution, purification, mixing with one or more reagents, or any other processing step(s) as described herein. A biological sample may be referred to simply as a “sample.”
[0646]
[0205] Blood component: As used herein, the term “blood component” refers to any component of whole blood, including red blood cells, white blood cells, plasma, platelets, endothelial cells, mesothelial cells, epithelial cells, cell-free DNA (cfDNA), and circulating-tumor DNA (ctDNA). Blood components also include the components of plasma, including proteins, metabolites, lipids, nucleic acids, and carbohydrates, and any other cells that can be present in blood, e.g., due to pregnancy, organ transplant, infection, injury, or disease.
[0647]
[0206] Breast Cancer: As used herein, the term “breast cancer” refers to histologically or cytologically confirmed cancer of the breast. In some embodiments, the breast cancer is a carcinoma. In some embodiments, the breast cancer is an adenocarcinoma. In some embodiments, the breast cancer is a sarcoma. In some embodiments, the breast cancer is an HR+ breast cancer. In some embodiments, the HR+ breast cancer is an AR+ breast cancer. In some embodiments, the Page 93 of 124
[0648] 13414664vlAttorney Docket No. : 2014191-0048
[0649] AR+ breast cancer is luminal A breast cancer. Tn some embodiments, the AR+ breast cancer is luminal B breast cancer. In some embodiments, the breast cancer is a metastatic or a locally advanced breast cancer.
[0650]
[0207] As used herein, the term “locally advanced breast cancer” refers to cancer that has spread from where it started in the breast to nearby tissue or lymph nodes, but not to other parts of the body.
[0651]
[0208] As used herein, the term “metastatic breast cancer” refers to cancer that has spread from the breast to other parts of the body, such as the bones, liver, lungs, or brain. Metastatic breast cancer may also be referred to as stage IV breast cancer.
[0652]
[0209] As used herein, the term “ductal carcinoma in situ breast cancer” or (DCIS cancer) refers to breast cancers characterized as being intraductal, non-evasive, and pre-invasive primary tumors as understood in the art.
[0653]
[0210] In some embodiments, a cancer is associated with a expression or activity status of a certain gene, e.g., in breast cancer.
[0654]
[0211] Cancer: As used herein, the terms “cancer,” “malignancy,” “tumor,” and “carcinoma,” are used interchangeably to refer to a disease, disorder, or condition in which cells exhibit or exhibited relatively abnormal, uncontrolled, and / or autonomous growth, so that they display or displayed an abnormally elevated proliferation rate and / or aberrant growth phenotype. In some embodiments, a cancer can include one or more tumors. In some embodiments, a cancer can be or include cells that are precancerous (e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic. In some embodiments, a cancer can be or include a solid tumor.
[0655]
[0212] Examples of cancer include but are not limited to, carcinoma, lymphoma, blastoma, sarcoma, and leukemia or lymphoid malignancies. More particular examples of such cancers include, but are not limited to, breast cancer (e.g., an HR+ breast cancer (e.g., an AR+ breast cancer (e.g., luminal A breast cancer or luminal B breast cancer)), DCIS, and / or a metastatic or a locally advanced breast cancer)); lung cancer, including small-cell lung cancer, non-small cell lung cancer, adenocarcinoma of the lung, and squamous carcinoma of the lung; bladder cancer (e.g., urothelial bladder cancer (UBC), muscle invasive bladder cancer (MIBC), and BCG-refractory non-muscle invasive bladder cancer (NMIBC)); kidney or renal cancer (e.g., renal cell carcinoma (RCC)); cancer of the urinary tract; prostate cancer, such as castration-resistant prostate cancer (CRPC) and prostate-specific membrane antigen (PSMA) positive prostate cancer; cancer of the Page 94 of 124
[0656] 13414664vlAttorney Docket No. : 2014191-0048
[0657] peritoneum; hepatocellular cancer; gastric or stomach cancer, including gastrointestinal cancer, gastroesophageal cancer and gastrointestinal stromal cancer; pancreatic cancer; glioblastoma; cervical cancer; ovarian cancer; liver cancer; hepatoma; colon cancer; rectal cancer; colorectal cancer; endometrial or uterine carcinoma; salivary gland carcinoma; prostate cancer; vulval cancer; thyroid cancer; hepatic carcinoma; anal carcinoma; penile carcinoma; melanoma, including superficial spreading melanoma, lentigo maligna melanoma, acral lentiginous melanomas, and nodular melanomas; multiple myeloma and B-cell lymphoma (including low grade / follicular non-Hodgkin’s lymphoma (NHL); small lymphocytic (SL) NHL; intermediate grade / follicular NHL; intermediate grade diffuse NHL; high grade immunoblastic NHL; high grade lymphoblastic NHL; high grade small non-cleaved cell NHL; bulky disease NHL; mantle cell lymphoma; AIDS-related lymphoma; and Waldenstrom’s Macroglobulinemia); chronic lymphocytic leukemia (CLL); acute lymphoblastic leukemia (ALL); acute myologenous leukemia (AML); hairy cell leukemia; chronic myeloblastic leukemia (CML); post-transplant lymphoproliferative disorder (PTLD); and myelodysplastic syndromes (MDS), as well as abnormal vascular proliferation associated with phakomatoses, edema (such as that associated with brain tumors), Meigs’ syndrome, brain cancer, head and neck cancer, and associated metastases.
[0658]
[0213] Combination therapy: As used herein, the term “combination therapy” refers to administration to a subject of two or more therapeutic agents or therapeutic regimens such that the two or more therapeutic agents or therapeutic regimens together treat a disease, condition, or disorder of the subject. In some embodiments, the two or more therapeutic agents or therapeutic regimens can be administered simultaneously, sequentially, or in overlapping dosing regimens. Those of skill in the art will appreciate that combination therapy includes but does not require that the two therapeutic agents or therapeutic regimens be administered together in a single composition, nor at the same time.
[0659]
[0214] “Diagnosing”, “Detecting”, “Determining” or “Screening for” : As used herein, “diagnosing”, “detecting”, “determining”, “screening for” the presence of a condition or disease (e g., AR-positive cancer), or a related state (e.g., responsiveness of an AR-positive cancer to one or more AR-targeted therapies) includes the act, process, and / or outcome of determining whether, and / or the qualitative of quantitative probability that, a subject has or will develop the condition, disease, or related state. In some instances, diagnosing can include a determination relating to
[0660] Page 95 of 124
[0661] 13414664vlAttorney Docket No. : 2014191-0048
[0662] prognosis and / or likely response to one or more general or particular therapeutic agents or regimens.
[0663] [2151 Differentially modified: As used herein, the term “differentially modified” describes a genomic locus for which histone modification status and / or DNA methylation status differs between a first condition or sample and a second condition or sample (e.g., a standard or reference). A differentially modified genomic locus can include a greater or smaller number or frequency of histone modification and / or DNA methylations under a selected condition of interest, such as activated signaling of a gene with known high activity in a particular disease state (e.g., cancer), as compared to a reference state, such as a state in which a signaling pathway comprising activation of such gene has not been stimulated.
[0664]
[0216] It is contemplated that systems, devices, methods, and processes of the disclosure encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and / or modification of the systems, devices, methods, and processes described herein may be performed by those of ordinary skill in the relevant art.
[0665]
[0217] Throughout the description, where articles, devices, and systems are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are articles, devices, and systems according to certain embodiments of the present disclosure that consist essentially of, or consist of, the recited components, and that there are processes and methods according to certain embodiments of the present disclosure that consist essentially of, or consist of, the recited processing steps.
[0666]
[0218] “Improve” "increase." "inhibit" or "reduce": As used herein, the terms “improve”, “increase”, “inhibit”, and “reduce”, and grammatical equivalents thereof, indicate qualitative or quantitative difference from a reference.
[0667]
[0219] Methylation Status: As used herein, “methylation status” of a genomic locus refers to the frequency with which DNA sequences corresponding to the genomic locus are identified in an assay for detection of DNA methylated sequences and / or the density (e.g., the measured density) of DNA methylation corresponding to the genomic locus. Methylation status can be determined by various assays known in the art, including without limitation Bisulfite sequencing (BS-Seq), Whole Genome Bisulfite Sequencing (WGBS), Methylated DNA ImmunoPrecipitation sequencing (MeDIP-seq), or Methyl-CpG-Binding Domain sequencing (MBD-seq). Where two Page 96 of 124
[0668] 13414664vlAttorney Docket No. : 2014191-0048
[0669] samples are separately analyzed by the same assay or comparable assays for detection of DNA methylated sequences, differences in methylation status of genomic loci can be detected. Methylation status can be compared to a standard or reference. A sample that has a methylation status that differs from a standard or reference can be referred to as differentially modified.
[0670]
[0220] Subject: As used herein, the term “subject” refers to an organism, typically a mammal (e.g., a human). In some embodiments, a subject is suffering from a disease, disorder or condition (e.g., e.g. cancer). In some embodiments, a subject is susceptible to a disease, disorder, or condition. In some embodiments, a subject displays one or more symptoms or characteristics of a disease, disorder or condition. In some embodiments, a subject is not suffering from a disease, disorder or condition. In some embodiments, a subject does not display any symptom or characteristic of a disease, disorder, or condition. In some embodiments, a subject has one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition. In some embodiments, a subject is a subject that has been tested for a disease, disorder, or condition, and / or to whom therapy has been administered. A subject may be a patient. In some instances, a human subject may be referred to as a “patient” or “individual.”
[0671]
[0221] Therapeutic agent: As used herein, the term “therapeutic agent” refers to any agent that elicits a desired pharmacological effect when administered to a subject. In some embodiments, an agent is considered to be a therapeutic agent if it demonstrates a statistically significant effect across an appropriate population. In some embodiments, the appropriate population can be a population of model organisms or a human population. In some embodiments, an appropriate population can be defined by various criteria, such as a certain age group, gender, genetic background, preexisting clinical conditions, etc. In some embodiments, a therapeutic agent is a substance that can be used for treatment of a disease, disorder, or condition (e.g., e.g. cancer). In some embodiments, a therapeutic agent is an agent that has been or is required to be approved by a government agency before it can be marketed for administration to humans. In some embodiments, a therapeutic agent is an agent for which a medical prescription is required for administration to humans.
[0672]
[0222] Therapeutically effective amount: As used herein, “therapeutically effective amount” refers to an amount that produces the desired effect for which it is administered. In some embodiments, the term refers to an amount that is sufficient, when administered to a population suffering from or susceptible to a disease, disorder, and / or condition (e.g., e.g.cancer) in Page 97 of 124
[0673] 13414664vlAttorney Docket No. : 2014191-0048
[0674] accordance with a therapeutic dosing regimen, to treat the disease, disorder, and / or condition. In some embodiments, a therapeutically effective amount is one that reduces the incidence and / or severity of, and / or delays onset of, one or more symptoms of the disease, disorder, and / or condition. Those of ordinary skill in the art will appreciate that the term “therapeutically effective amount” does not in fact require successful treatment be achieved in a particular individual. Rather, a therapeutically effective amount may be that amount that provides a particular desired pharmacological response in a significant number of subjects when administered to patients in need of such treatment. In some embodiments, reference to a therapeutically effective amount may be a reference to an amount as measured in one or more specific tissues (e.g., a tissue affected by the disease, disorder or condition) or fluids (e.g., blood, saliva, serum, sweat, tears, urine, etc.). Those of ordinary skill in the art will appreciate that, in some embodiments, a therapeutically effective amount of a particular agent or therapy may be formulated and / or administered in a single dose. In some embodiments, a therapeutically effective amount of a particular agent or therapy may be formulated and / or administered in a plurality of doses, for example, as part of a dosing regimen.
[0675]
[0223] Treatment: As used herein, the term “treatment” (also “treat” or “treating”) refers to administration of a therapy that partially or completely alleviates, ameliorates, relieves, inhibits, delays onset of, reduces severity of, and / or reduces incidence of one or more symptoms, features, and / or causes of a particular disease, disorder, or condition, or is administered for the purpose of achieving any such result. In some embodiments, such treatment can be of a subject who does not exhibit signs of the relevant disease, disorder, or condition and / or of a subject who exhibits only early signs of the disease, disorder, or condition (e.g., cancere.g.). Alternatively, or additionally, such treatment can be of a subject who exhibits one or more established signs of the relevant disease, disorder and / or condition. In some embodiments, treatment can be of a subject who has been diagnosed as suffering from the relevant disease, disorder, and / or condition. In some embodiments, treatment can be of a subject known to have one or more susceptibility factors that are statistically correlated with increased risk of development of the relevant disease, disorder, or condition. A “prophylactic treatment” includes a treatment administered to a subject who does not display signs or symptoms of a condition to be treated or displays only early signs or symptoms of the condition to be treated such that treatment is administered for the purpose of diminishing, preventing, or decreasing the risk of developing the condition. Thus, a prophylactic treatment Page 98 of 124
[0676] 13414664vlAttorney Docket No. : 2014191-0048
[0677] functions as a preventative treatment against a condition. A “therapeutic treatment” includes a treatment administered to a subject who displays symptoms or signs of a condition and is administered to the subject for the purpose of reducing the severity or progression of the condition.
[0678] EXAMPLES
[0679]
[0224] The present Examples demonstrate the formation and validation of estimation models for use with plasma samples obtained from subjects. The present Examples show that methods of forming estimation models as disclosed herein may produce estimation models that can be validated and used to accurately determine cTF in different samples.
[0680] Example 1
[0681]
[0225] A cTF estimation model was trained and subsequently validated for determining cTF of plasma samples that included ctDNA from breast cancer. A general overview of the process used for training is that candidate regions were selected, region counts for MBD-seq sequencing data for a range of samples at different known cTF were normalized, signal regions were selected, noise regions were selected, and a cTF estimation model was produced using the signal regions, the noise regions, and the sequencing data.
[0682]
[0226] In this example, candidate regions were selected as predefined regions. The predefined regions were CpG islands that had been identified by a reference source of University of California Santa Cruz.
[0683]
[0227] MBD-seq sequencing data were obtained for a variety of samples. Healthy samples, cell line samples, and in silica plasma (ISP) samples were used. ISP samples were generated by random sampling of fragments from other samples with known cTF in a predetermined proportion to form a new sample with a known cTF. For example, an ISP sample may be made by randomly sampling 5 million fragments from a 50% ctDNA plasma and combining it with 5 million fragments from a healthy sample to form a new in silica sample with a 25% known (expected) cTF. In this way, a large plurality of samples having an even distribution of cTFs can be easily assembled.
[0684]
[0228] The MBD-seq sequencing data for samples were pre-processed to get unique fragments from aligned sequencing reads. That is, the data were deduplicated to obtain the unique Page 99 of 124
[0685] 13414664vlAttorney Docket No. : 2014191-0048
[0686] fragments. The counts in the sequencing data were taken as the number of fragments that overlap with a region, in this case a defined CpG island. Pseudocounts were used, where the pseudocounts were the actual count plus one, in order to prevent infinite or null values when doing future logarithmic operations. Because MBD-seq data were used, two normalizations were performed. First, the pseudocounts were normalized to the length of region as longer regions will naturally “emit” more fragments. This normalization generated a counts per kilobase (CPK) normalized counts value. CPK was calculated as 1000 * pseudocounts / region_length. Second, the CPK was normalized to how many fragments were sequenced in the library and that normalization was put on the log scale. This normalization was calculated as logCPKM = log(CPK) - log(# of fragments sequenced). In this way, a double normalization may account for positive enrichment bias in MBD-seq data. As a singular formula, the double normalization can be written as Formula 1 :
[0687] / 1000 * (cir+ 1 \
[0688] logCPKMi r= log - - yr* LO /
[0689] where a,ris unique counts for sample i and region r, LS is the library size of the sample as defined by the number of unique fragments sequenced, and lris the length of the candidate region. In this way the logCPKM is roughly the rate of captured fragments per base pair and fragment sequenced. logCPKM values were used as normalized counts in subsequent steps.
[0690]
[0229] Signal regions that were differentially methylated regions (DMRs) were selected based on correlation between cTF and logCPKM. In this example, a candidate region was selected as a signal region if it has a linear relationship with ctDNA and the logCPKM at 1% ctDNA (~ -2 continuous ctdna) is higher than the 95% quantile of healthy signal. FIG. 4 illustrates an example of such relationships existing for a particular exemplary region (chr7: 121956543-121957341) in this example. The median value of normalized counts (logCPKM) for all signal regions (signal median logcpk) was used as a per-sample-based metric and is plotted on the y-axis of FIG.
[0691] 4. Median performs better than mean for certain cancer types. The x-axis of FIG. 4 is cTF on the logit scale (log(cTF)Z(l-cTF)) (continuous ctdna). LogCPKM represent as per-region-based metric for each of the signal regions.
[0692]
[0230] Noise regions were also selected. In this example, a noise region was defined as a region that has a linear correlation with the signal metric (signal_median_logcpk) in heathy samples but not in cancer samples. In this example, signal-to-noise ratio (SNR) is noise_median_logcpk / signal_median_logcpk (the ratio has been inverted because the metrics are Page 100 of 124
[0693] 13414664vlAttorney Docket No. : 2014191-0048
[0694] in negative log space, thereby providing a positive relationship with ctDNA). FIG. 5 is a plot illustrating a noise region in this example (chrl: 171454517-171455239). The y-axis represents a logcpk value for the noise region selected. FIG. 5 illustrates that there is no correlation between the noise region logcpk and the signal median logcpk in cancer samples (black points), but it does show a correlation in healthy plasma (blue, orange, green, and red) points. This indicates a region that may measure technical noise in the data.
[0695]
[0231] Different model types were tested across training samples to produce a cTF estimation model. Regressions were fit to produce estimation models that estimate cTF. Linear, polynomial, sigmoid, tan fit, and other regressions were tested. Moreover, different combinations of models were tested as ensemble methods. FIG. 6 illustrates an example of an SNR-based regression model that was fit. The y-axis is cTF as represented by continuous ctdna and the x-axis is SNR. A linear fit was made but a non-linear fit (e.g., polynomial, sigmoid, or tan) could be made instead.
[0696]
[0232] FIGs. 7-10 illustrate performance of candidate models used to produce a cTF estimation model. In FIG. 7, performance of an ensemble model (model gmean) that uses a geometric mean of preliminary estimated cTFs determined by an expectation maximalization model, a point-estimate model that uses signal median logcpk as the point estimate, and an SNRbased model was compared against performance of the expectation maximalization model alone (em_pred). Expected fraction is the cTF expected based on an orthogonal method for the samples. MBTA estimate is the estimated cTF determined by the models. As can be seen, the model gmean performs better at low expected cTF and em_pred performs better at high cTF. FIG. 8 illustrates more detail demonstrating performance of the expectation maximalization model. As can be seen from the bottom two panels, the expectation maximalization model accurately predicted cTF across a wide range of cTF fractions as represented by expected cTFs determined using an orthogonal method. The top panels illustrate that variance of predictions across the range of cTFs is low, in particular less than 50% above the limit of detection for the expectation maximalization model. FIG. 9 illustrates that a LoD of about 0.5% is stable, across a wide range of sample fragment amounts (legend amounts are in millions of fragments), assuming a limit of blank (LoB) is about 0.3% for a breast cancer model. FIG. 10 illustrates that the model gmean can accurately detect presence of ctDNA at low cTF consistent with a 0.3% limit of blank (LoB) and 95% specificity in unseen healthies. The x-axis represents estimated cTF as determined by the Page 101 of 124
[0697] 13414664vlAttorney Docket No. : 2014191-0048
[0698] expectation maximalization model alone. Only a few false positives were produced and could be explained by sample processing differences and batch effects.
[0699] [2331 FIG. 11 illustrates a comparison of performance of different models, including ensemble methods, for breast cancer models. FIG. 11 compares the performance of different cTF model predictions across a range of simulated number of sequenced fragments (n sample frags). Ensemble models that used different combinations of preliminary estimated cTFs from the constituent models were tested including constituent models of an expectation maximization model, a point-estimate model that uses signal median logcpk as the point estimate, and an SNRbased model. Tested combinations of these three constituent models include geometric mean (model gmean), harmonic mean (model hmean), and arithmetic mean (model mean). (Other models tested (snr_polym_pred_ctdna, signal_polym_pred_ctdna, and adj model gmean (an ensemble model of an SNR-based model (snr_polym_pred_ctdna) and model gmean)) performed worse.) The x-axis is the number of fragments in the sample (range 2.5M to 25M). For plasma samples obtained from 1-2 mL sample volumes, on average, about 15M fragments have been obtained when we sequenced using MBD-seq. The y-axis is the average limit of detection (LoD) across five cross validation folds. So 0.5 corresponds to a LoD of 0.5% cTF. A lower LoD is better and means we can more sensitively detect cTF. While arithmetic mean showed the best performance as gauged by lowest LoD at 10M fragments, geometric mean showed best average performance across a wide range of fragment amounts (-0.55-0.6% cTF LoD in a range of -5-25M fragments). Performance across a wide range of fragment amounts indicates model robustness across a wide range of samples where number of sequenced fragments may vary between samples. For this reason, in this example, the model represented by model gmean was selected as the cTF estimation model. Based on the results shown in FIG. 7 where the expectation maximalization model alone was shown to be more predictive at higher cTF, another model was also produced where the model represented by model gmean was selected as a first model to determine whether ctDNA is present in a sample and the expectation maximalization model was selected as a second model to estimate cTF where ctDNA was determined to be present by the first model (e.g., where the first model determines a non-zero cTF).
[0700] Example 2
[0701] Page 102 of 124
[0702] 13414664vlAttorney Docket No. : 2014191-0048
[0703]
[0234] In a similar approach described in Example 1, the present Example provides a general framework for estimating cTF, specifically, using histone modification sequencing data. Briefly, candidate regions were selected, region counts for histone modification sequencing data for a range of samples at different known cTF were normalized, signal regions were selected, noise regions were selected, and a cTF estimation model was produced using the signal regions, the noise regions, and the sequencing data. Similar to Example 1, predefined regions of known CpG island from a University of California Santa Cruz reference source informed selection of candidate regions.
[0704]
[0235] Histone modification sequencing data for H3K4me3 histone modification from in silico plasma (ISP) diluted samples, samples with low ctDNA% estimated using the model described in Example 1, and samples from healthy subjects were used for training the a cTF estimation model.
[0705]
[0236] Sequencing data for ChlP-seq was deduplicated, overlapping sequence reads were counted, and converted to pseudocounts. CPK normalized values were obtained by normalizing pseudocounts to the length of region, CPK was further normalized to a number of sequenced fragments, and outputs were log transformed, as shown in Formula 1 of Example 1.
[0706]
[0237] Signal regions for differential histone modifications were selected based on correlation between cTF and logCPKM. Signal regions and noise regions were selected based on similar parameters described in Example 1.
[0707]
[0238] Ensemble models were used to generate cTF estimation models. First, an L2 regularized logistic regression model was trained on healthy samples, low ctDNA samples and ISP samples. Next, histogram-based gradient boosting regression tree (HGBR) was trained on ISP samples to predict cTF on a continuous scale.
[0708]
[0239] Data for signal regions are used as inputs for the L2-regularized logistic regression, If a predicted cTF for a sample is not zero, the data are then used to estimate cTF using HGBR.
[0709]
[0240] FIG. 12 shows a comparison of cTF prediction using H3K4me3 differential histone modification-based models described in the present Example, compared to ichorCNA. cfDNA from breast, prostate, colorectal, lung, and gastroesophageal cancer were used to train the models, in addition to healthy cfDNA samples. Approximately 95% specificity was observed for the presently described cTF prediction models. FIG. 12 demonstrates that cTF estimation using the presently described methods was accomplished for cTF down to 0.01%, and well below 3% limit Page 103 of 124
[0710] 13414664vlAttorney Docket No. : 2014191-0048
[0711] of detection for ichorCNA. Above 3% cTF, Pearson correlation was 0.862 (p=2.79e-54) for ichorCNA and H3K4me3-based cTF estimation.
[0712] OTHER EMBODIMENTS
[0713]
[0241] It will be appreciated that the scope of the present disclosure is to be defined by that which may be understood from the disclosure and claims rather than by the specific embodiments that have been presented by way of example. Elements described with respect to one aspect or embodiment of the present disclosure are also contemplated with respect to other aspects or embodiments of the present disclosure. For example, elements of claims that depend directly or indirectly from a certain independent claim presented herein serve as support for those elements being presented in additional dependent claims of one or more other independent claims. Throughout the description, where compositions or methods are described as having, including, or comprising specific elements, it is to be understood that compositions or methods that consist essentially of, consist of, or do not comprise the recited elements are likewise hereby disclosed. Moreover, it is to be understood that the features of the various embodiments described in the present disclosure were not mutually exclusive and can exist in various combinations and permutations, even if such combinations or permutations were not made express, without departing from the spirit and scope of the disclosure.
[0714]
[0242] It should further be understood that the order of steps or order for performing certain action is immaterial so long as operability is not lost. Moreover, two or more steps or actions may be conducted simultaneously. As is understood by those skilled in the art, the terms “over”, “under”, “above”, “below”, “beneath”, and “on” are relative terms and can be interchanged as appropriate, for example in reference to different orientations of the layers, elements, and substrates included in the present disclosure. In some embodiments, a first layer on a second layer means a first layer directly on and in contact with a second layer. In some embodiments, a first layer on a second layer may include another layer therebetween.
[0715]
[0243] Headers have been provided for the convenience of the reader and are not intended to be limiting with respect to the claimed subject matter.
[0716]
[0244] All references cited herein are hereby incorporated by reference.
[0717] Page 104 of 124
[0718] 13414664vl
Claims
Attorney Docket No. : 2014191-0048What is claimed is:
1. A method of estimating cTF in a sample, the method comprising:receiving sequencing data for a liquid sample from a subject;producing normalized counts of the sequencing data for signal regions in the genome of the subject; andestimating cTF for the sample with a cTF estimation model based at least on the normalized counts.
2. The method of claim 1, wherein the sequencing data has been derived from an enrichment method.
3. The method of claim 1 or claim 2, wherein the sequencing data is epigenetic modification sequencing data.
4. The method of claim 3, wherein the epigenetic modification sequencing data comprises histone modification sequencing data, DNA methylation sequencing data, transcription factor binding sequencing data and / or chromatin accessibility sequencing data.
5. The method of any one of claims 1-4, wherein the sequencing data is chromatin immunoprecipitation sequencing (ChlP-seq) data.
6. The method of claim 5, wherein the ChlP-seq data comprises sequencing data for one or more histone modifications.
7. The method of claim 6, wherein the one or more histone modifications comprise H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, or pan acetylation.
8. The method of claim 6 or 7, wherein the one or more histone modifications comprise H3K4me3.Page 105 of 12413414664vlAttorney Docket No. : 2014191-00489. The method of any one of claims 6-8, wherein the one or more histone modifications comprise H3K27ac.
10. The method of any one of claims 6-9, wherein the one or more histone modifications comprise H3K4me3 and H3K27ac.
11. The method of any one of claims 1-4, wherein the sequencing data is methyl-binding domain sequencing (MBD-seq) data.
12. The method of claim 11, wherein the MBD-seq data comprises sequencing data for DNA methylation.
13. The method of any one of claims 1-10, wherein the signal regions are regions determined to be cancer-specific regions with differential histone modifications (DHMRs).
14. The method of any one of claims 1-4, 11, or 12, wherein the signal regions are regions determined to be cancer-specific differentially methylated regions (cDMRs).
15. The method of any one of claims 1-14, wherein the normalizing comprises, for each of the signal regions, normalizing the counts based on number of fragments in the sequencing data for the sample and based on length of the signal region.
16. The method of any one of claims 1-15, wherein producing the normalized counts comprises, for each of the signal regions, performing a normalization according to(f * (cr+ d)\logCPKMr= log\ * ZjiJ / where cris unique counts in the sequencing data for the sample in region r, d is a constant, / is a non-zero constant, LS is sample library size as defined by the number of unique fragments sequenced for the sample, and lris region length for region r.Page 106 of 12413414664vlAttorney Docket No. : 2014191-004817. The method of any one of claims 1-16, wherein, for each of the signal regions, the counts for the signal region are a number of fragments in the sequencing data that overlap with signal region.
18. The method of any one of claims 1-17, wherein estimating the cTF comprises estimating a per-region-based estimated cTF for each of the signal regions individually and then determining the estimated cTF based on the per-region-based estimated cTF using the cTF estimation model.
19. The method of any one of claims 1-18, wherein estimating the cTF comprises performing an expectation maximization of cTF for the sample with the cTF estimation model based on the normalized counts.
20. The method of any one of claims 1-19, wherein estimating the cTF comprises determining a point estimate of the normalized counts for the signal regions for the sample and the cTF estimation model uses the point estimate to estimate the cTF for the sample.
21. The method of claim 20, wherein the point estimate is the median.
22. The method of any one of claims 1-21, wherein estimating the cTF comprises estimating a plurality of preliminary estimated cTFs comprising a per-region-based estimated cTF estimated using a first model of the cTF estimation model and a per-sample-based estimated cTF estimated using a second model of the cTF estimation model, wherein the cTF is estimated by a combination that includes the per-region-based estimated cTF and the per-sample-based estimated cTF.
23. The method of any one of claims 1-22, comprising:producing normalized counts of the sequencing data for noise regions in the genome of the subject; anddetermining a signal-to-noise ratio (SNR) for the sample based on the normalized counts for the noise regions and the normalized counts for the signal regions,Page 107 of 12413414664vlAttorney Docket No. : 2014191-0048wherein the cTF estimation model uses the SNR to estimate the cTF24. The method of claim 23, wherein producing the normalized counts for the noise regions comprises, for each of the noise regions, normalizing counts for the noise region based on number of fragments in the sequencing data for the sample and based on length of the noise region.
25. The method of claim 23 or claim 24, wherein producing the normalized counts for the noise regions comprises, for each of the noise regions, performing a normalization according to / f * (cr+ d)\logCPKMr= log ,r,c\ Lp * L / S Jwhere cris unique counts in the sequencing data for the sample in region r, d is a constant, / is a non-zero constant, LS is sample library size as defined by the number of unique fragments sequenced for the sample, and lris region length for region r.
26. The method of any one of claims 23-25, wherein the SNR is determined as a ratio of a point estimate of the normalized counts for the signal regions and a point estimate of the normalized counts for the noise regions.
27. The method of claim 26, wherein the point estimate for the signal regions and the point estimate for the noise regions is the median.
28. The method of any one of claims 19-27, wherein the cTF estimation model combines a plurality of preliminary estimated cTFs, each determined by a separate constituent model of the cTF estimation model, to estimate the cTF.
29. The method of claim 28, wherein the combining comprises determining a geometric mean of the preliminary estimated cTFs.
30. The method of claim 28 or claim 29, wherein the preliminary estimated cTFs comprise a per-region-based estimated cTF and a per-sample-based estimated cTF.Page 108 of 12413414664vlAttorney Docket No. : 2014191-004831. The method of claim 30, wherein the per-sample-based estimated cTF is determined using a first constituent model that is a point estimate-based model.
32. The method of claim 30 or claim 31, wherein the per-region-based estimated cTF is determined using an expectation maximization model.
33. The method of any one of claims 26-32, wherein the preliminary estimated cTFs comprise an SNR-based preliminary estimated cTF.
34. The method of any one of claims 1-33, wherein estimating the cTF comprises using a first model of the cTF estimation model to detect if ctDNA is present in the sample and subsequently using a second model of the cTF estimation model to estimate the cTF if ctDNA is present or estimating the cTF at 0 if ctDNA is determined to be not present using the first model.
35. The method of any one of claims 1-34, wherein estimating the cTF comprises using a first model of the cTF estimation model to detect that ctDNA is present in the sample and, upon determining that ctDNA is present, using a second model of the cTF estimation model to estimate the cTF.
36. The method of claim 34 or claim 35, wherein the first model comprises an ensemble model that uses a geometric mean of preliminary estimated cTFs and the second model is an expectation maximalization model.
37. The method of claim 34 or claim 35, wherein the first model is a signal-to-noise ratio (SNR) based model and the second model is an expectation maximalization model.
38. The method of any one of claims 34-37, wherein using the first model comprises estimating a preliminary estimated cTF and the second model used to estimate the cTF depends on the preliminary estimated cTF.Page 109 of 12413414664vlAttorney Docket No. : 2014191-004839. The method of any one of claims 1-38, wherein the model comprises a multivariate, regularized, logistic, and / or tree-based regression (e.g., that uses a feature matrix, e.g., a histogram-based gradient boosting regression tree).
40. The method of any one of claims 1-39, wherein the model comprises a feature aggregation and regression model that uses a per-sample-based point estimate.
41. The method of any one of claims 1-40, wherein the model comprises a deep learning model that has been trained using tensors derived from normalized sequencing data for the signal regions.
42. The method of any one of claims 1-41, wherein the cTF estimation model comprises an ensemble model.
43. The method of any one of claims 1-42, wherein the model uses the signal regions independently to obtain the estimated cTF.
44. The method of any one of claims 1-43, wherein the model determines estimated cTF based on independent consideration of normalized counts for the signal regions in a sample.
45. The method of any one of claims 1-44, wherein the model has been trained using data from cancer-specific samples and the subject is suspected of having and / or known to have or have had the cancer to which the model is specific.
46. The method of any one of claims 1-45, wherein the model has been trained using data from samples having different cancers.
47. The method of any one of claims 1-46, wherein the normalizing comprises accounting for copy number alterations reflected in the sequencing data.
48. The method of any one of claims 1-47, wherein the counts are pseudocounts.Page 110 of 12413414664vlAttorney Docket No. : 2014191-004849. The method of any one of claims 1-48, wherein the counts are per-region-based counts that indicate that a region is methylated.
50. The method of any one of claims 1-49, wherein the sample is a liquid biopsy sample.
51. The method of any one of claims 1-50, wherein the sample is a plasma sample.
52. A method of producing a circulating tumor fraction (cTF) estimation model for estimating an amount of circulating tumor DNA in a sample of a subject, the method comprising:selecting candidate regions in a genome where epigenetic modification of the candidate regions may be indicative of cancer;receiving epigenetic modification sequencing data for a plurality of samples corresponding to the genome, wherein each of the samples has a known cTF;normalizing counts of the sequencing data for each of the candidate regions for each of the samples; determining signal regions in the genome where epigenetic modification is indicative of cancer using the normalized counts for the candidate regions and the known cTF for the samples; andproducing a cTF estimation model that determines an estimated cTF in a sample based on input sequencing data or data derived therefrom, wherein the producing comprises using at least the normalized counts for the signal regions and the known cTF for the samples to produce the cTF estimation model.
53. The method of claim 52, wherein the epigenetic modification comprises histone modification, DNA methylation, transcription factor binding and / or chromatin accessibility.
54. The method of claim 52 or 53, wherein the epigenetic modification comprises histone modification.
55. The method of any one of claims 52-54, wherein the epigenetic modification comprises DNA methylation.Page 111 of 12413414664vlAttorney Docket No. : 2014191-004856. The method of claim 54, wherein the histone modification comprises H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K4mel, H3K4me2, or H3K4me3, or pan acetylation.
57. The method aby one of claims 53-56, wherein the histone modification comprises H3K4me3.
58. The method of any one of claims 53-57, wherein the histone modification comprises H3K27ac.
59. The method of any one of claims 53-58, wherein the histone modification comprises H3K4me3 and H3K27ac.
60. A method of producing a circulating tumor fraction (cTF) estimation model for estimating an amount of circulating tumor DNA in a sample of a subject, the method comprising:selecting candidate regions in a genome where methylation of the candidate regions may be indicative of cancer;receiving methylation sequencing data for a plurality of samples corresponding to the genome, wherein each of the samples has a known cTF;normalizing counts of the sequencing data for each of the candidate regions for each of the samples;determining signal regions in the genome where methylation is indicative of cancer using the normalized counts for the candidate regions and the known cTF for the samples; and producing a cTF estimation model that determines an estimated cTF in a sample based on input sequencing data or data derived therefrom, wherein the producing comprises using at least the normalized counts for the signal regions and the known cTF for the samples to produce the cTF estimation model.
61. The method of any one of claims 52-60, wherein the sequencing data has been derived from an enrichment method.Page 112 of 12413414664vlAttorney Docket No. : 2014191-004862. The method of any one of claims 52-61, wherein the sequencing data is methyl-binding domain sequencing (MBD-seq) data.
63. The method of any one of claims 52-62, wherein the signal regions are regions determined to be cancer-specific regions with differential histone modifications (cDHMRs).
64. The method of any one of claims 52-63, wherein the candidate regions are regions determined to be cancer-specific regions with differential histone modifications (cDHMRs).
65. The method of any one of claims 52-64, wherein the signal regions are regions determined to be cancer-specific differentially methylated regions (cDMRs).
66. The method of any one of claims 52-65, wherein the candidate regions are regions that are or may be cancer-specific differentially methylated regions (cDMRs).
67. The method of any one of claims 52-66, wherein the normalizing comprises, for each of the candidate regions for each of the samples, normalizing the counts based on number of fragments in the sequencing data for the sample and based on length of the candidate region.
68. The method of any one of claims 52-67, wherein the normalizing comprises performing a normalization according tologCPKMi r= loglr* LSwhere ct ris unique counts for sample z and region r, d is a constant, / is a non-zero constant, LS is sample library size as defined by the number of unique fragments sequenced, and lris region length for region r.
69. The method of any one of claims 52-68, wherein, for each of the candidate regions, the counts for the candidate region correspond to a number of fragments in the sequencing data that overlap with candidate region.Page 113 of 12413414664vlAttorney Docket No. : 2014191-004870. The method of any one of claims 52-69, wherein the normalizing comprises accounting for copy number alterations reflected in the sequencing data.
71. The method of any one of claims 52-70, wherein the counts are pseudocounts.
72. The method of any one of claims 52-71, wherein the counts are per-region-based counts that indicate that a region has been bound by a differentially modified histone.
73. The method of any one of claims 52-72, wherein the counts are per-region-based counts that indicate that a region is methylated.
74. The method of any one of claims 52-73, wherein the signal regions are determined based on the normalized counts for a candidate region having a linear relationship with cTF and the normalized counts for the candidate region for samples at a threshold cTF being higher than an upper quantile or lower than a lower quantile for healthy samples.
75. The method of claim 74, wherein the threshold cTF is in a range of from 0.5 to 2% cTF.
76. The method of claim 74 or claim 75, wherein the normalized counts for the candidate region for the samples at the threshold cTF being higher than an upper quantile for healthy samples with the upper quantile being in a range of 85%-99% is used as a condition for determining the signal regions or wherein the normalized counts for the candidate region for the samples at the threshold cTF being lower than a lower quantile for healthy samples with the lower quantile being in a range of 1%- 15% is used as a condition for determining the signal regions.
77. The method of any one of claims 52-76, wherein determining the signal regions comprises determining that ones of the candidate regions have been bound by a differentially modified histone over a range of cTF.Page 114 of 12413414664vlAttorney Docket No. : 2014191-004878. The method of any one of claims 52-77, wherein determining the signal regions comprises determining that ones of the candidate regions are differentially methylated over a range of cTF.
79. The method of any one of claims 52-78, wherein determining the signal regions comprises determining that one or more of the candidate regions have differential histone modification observable down to a target cTF in a range of from 0.5% to 3% cTF.
80. The method of any one of claims 52-79, wherein determining the signal regions comprises determining that one or more of the candidate regions have differential methylation observable down to a target cTF in a range of from 0.5% to 3% cTF.
81. The method of any one of claims 52-80, wherein determining the signal regions comprises determining that there exists a linear relationship between the normalized counts for a candidate region and cTF.
82. The method of any one of claims 52-81, wherein the signal regions correspond to CpG islands in the genome.
83. The method of any one of claims 52-82, wherein determining the signal regions comprises, for each of the candidate regions, determining a relationship between the normalized counts for the candidate region and the known cTF for the samples.
84. The method of any one of claims 52-83, wherein producing the model is based on, for each of the plurality of samples, a point estimate of normalized counts for the signal regions for the sample, preferably wherein the point estimate is the median.
85. The method of any one of claims 52-84, wherein producing the model comprises fitting a cTF regression for each of the signal regions.Page 115 of 12413414664vlAttorney Docket No. : 2014191-004886. The method of any one of claims 52-85, wherein the cTF estimation model comprises an expectation maximization model.
87. The method of claim 86, wherein the expectation maximization model is structured to estimate cTF by maximizing expectation across the signal regions independently.
88. The method of claim 86 or claim 87, wherein the expectation maximization model is structured to estimate cTF by maximizing across the signal regions simultaneously.
89. The method of any one of claims 52-88, wherein producing the model comprises producing (e.g., defining and / or fitting) a set of relationships (e.g., regressions) between signal regions and cTF such that an expectation maximization can be performed across the set of relationships (e.g., on a per-region basis).
90. The method of any one of claims 52-89, wherein producing the model comprises defining a global loss function for a set of relationships (e.g., regressions) between cTF and normalized counts for a signal region for each of the signal regions.
91. The method of any one of claims 52-90, wherein the model comprises a multivariate, regularized, logistic, and / or tree-based regression (e.g., that uses a feature matrix, e g., a histogram-based gradient boosting regression tree).
92. The method of any one of claims 52-91, wherein the model comprises a feature aggregation and regression model that uses a per-sample-based point estimate (e g., sum, mean, or median) (e.g., for the signal regions, optionally and the noise regions).
93. The method of any one of claims 52-92, wherein the model comprises a deep learning model (e.g., an artificial neural network) and producing the model comprises training the deep learning model using tensors derived from the normalized counts for the signal regions.Page 116 of 12413414664vlAttorney Docket No. : 2014191-004894. The method of any one of claims 52-93, wherein the cTF estimation model comprises an ensemble model.
95. The method of any one of claims 52-94, wherein producing the cTF estimation model comprises producing a plurality of constituent models to produce an ensemble model.
96. The method of claim 94 or claim 95, wherein the ensemble model uses a geometric mean of constituent models to estimate cTF.
97. The method of any one of claims 94-96, wherein the ensemble model comprises (e.g., the plurality of constituent models comprises) a first model that estimates cTF based on all of the signal regions together and a second model that estimates cTF based on each of the signal regions individually.
98. The method of any one of claims 94-97, wherein the ensemble model comprises a first model that estimates cTF based on a metric determined from all of the signal regions together and a second model that estimates cTF based on individual metrics for each of the signal regions.
99. The method of any one of claims 94-98, wherein the ensemble model comprises an expectation maximization model and (i) a signal-to-noise ratio (SNR) based model and / or (ii) a point estimate regression based model.
100. The method of any one of claims 52-99, wherein producing the model comprises independently fitting the signal regions.
101. The method of any one of claims 52-100, wherein the model determines estimated cTF based on independent consideration of normalized counts for the signal regions in a sample.
102. The method of any one of claims 52-101, wherein producing the model comprises selecting a type of model based on cancer type for ones of the samples.Page 117 of 12413414664vlAttorney Docket No. : 2014191-0048103. The method of any one of claims 52-102, wherein producing the model comprises producing a pan-cancer ensemble model comprising at least one constituent model for a plurality of cancer types.
104. The method of claim 103, wherein the plurality of cancer types comprises lung cancer, breast cancer, colorectal cancer, gastric cancer (e.g., gastroesophageal cancer), and / or prostate cancer.
105. The method of any one of claims 52-104, wherein selecting the candidate regions comprises selecting a set of CpG islands.
106. The method of any one of claims 52-105, wherein the candidate regions are CpG islands.
107. The method of any one of claims 52-106, wherein selecting the candidate regions comprises selecting bins of uniform size within the genome.
108. The method of any one of claims 52-107, wherein the candidate regions are bins of uniform size.
109. The method of any one of claims 52-108, wherein selecting the candidate regions comprises determining regions having differential epigenetic modification between conditions.
110. The method of any one of claims 52-109, wherein selecting the candidate regions comprises determining regions having differential histone modification between conditions.
111. The method of any one of claims 52-110, wherein selecting the candidate regions comprises determining regions having differential methylation between conditions.
112. The method of any one of claims 109-111, wherein the regions having differential epigenetic modification between conditions are determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples.Page 118 of 12413414664vlAttorney Docket No. : 2014191-0048113. The method of any one of claims 109-112, wherein the regions having differential histone modification between conditions are determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples.
114. The method of any one of claims 109-113, wherein the regions having differential methylation between conditions are determined using sequencing data for one or more healthy samples and sequencing data for one or more high cTF samples.
115. The method of any one of claims 52-114, comprising determining noise regions in the genome.
116. The method of claim 115, wherein determining the noise regions comprises, for each of the noise regions, determining that a metric for samples having a known cTF at or below a low threshold has a positive correlation and that the metric for samples having a known cTF above a threshold is uncorrelated.
117. The method of claim 115 or claim 116, wherein producing the model is further based on normalized counts of the sequencing data for the noise regions.
118. The method of any one of claims 115-117, comprising calculating a signal-to-noise ratio for each of the samples, wherein the signal-to-noise ratio is based on a point estimate count value across each of the noise regions for a sample and a point estimate count value across each of the signal regions for the sample.
119. The method of any one of claims 52-118, wherein the plurality of samples are liquid biopsy samples.
120. The method of any one of claims 52-119, wherein the plurality of samples are plasma samples.Page 119 of 12413414664vlAttorney Docket No. : 2014191-0048121. The method of any one of claims 52-120, wherein the plurality of samples comprise in silico diluted samples.
122. The method of any one of claims 52-121, wherein the sequencing data for the in silico diluted samples is a random sampling of fragments from a healthy sample and a sample with a known non-zero cTF in a predetermined ratio of amounts.
123. The method of any one of claims 52-122, wherein the plurality of samples comprises samples spanning a range of known cTF.
124. The method of any one of claims 52-123, wherein, for at least one of the plurality of samples, the known cTF is known from an orthogonal reference method.
125. The method of any one of claims 52-124, wherein, for at least one of the plurality of samples, the sample is a simulated dilution and the known cTF for the sample has been determined by linear combination of known cTFs of reference samples used to generate the sample.
126. The method of any one of claims 52-125, wherein the samples correspond to a single cancer.
127. The method of any one of claims 52-126, wherein the samples comprise different samples corresponding to different cancers.
128. A method of estimating cTF in a sample, the method comprising:providing sequencing data for a liquid sample and a cTF estimation model that has been formed by a method according to any one of claims 52-127; andestimating cTF in the sample using the cTF estimation model.
129. A system comprising a processor and one or more non-transitory computer readable media having instructions stored thereon that, when executed by the processor, cause the Page 120 of 12413414664vlAttorney Docket No. : 2014191-0048processor to perform operations comprising the method according to any one of claims 1-128.
130. One or more non-transitory computer readable media having instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising the method according to any one of claims 1-128.
131. A method of characterizing cancer recurrence and / or progression, the method comprising:estimating cTF in a first sample for a subject taken at a first time point using a method according to any one of claims 1-51 or 128;estimating cTF in a second sample for the subject taken at a second time point after the first time point using a method according to any one of claims 1-51 or 128; and determining a difference in the estimated cTF at the second time point and at the first time point.
132. A method of monitoring cancer in a subject, the method comprising:estimating cTF in a series of two or more samples for a subject, each taken at a different time point, using a method according to any one of claims 1-51 or 128; anddetermining whether there is a difference in cTF for the samples over time.
133. The method of claim 131 or claim 132, comprising administering a therapy to the subject when the difference is determined to be at least as large as a threshold difference.
134. The method of any one of claims 131-133, comprising altering administration of a therapy to the subject when the difference is determined to be at least as large as a threshold difference.
135. The method of claim 134, wherein altering administration comprises increasing a dosage and / or frequency of administration.
136. A method of prognosing cancer, the method comprising:Page 121 of 12413414664vlAttorney Docket No. : 2014191-0048estimating cTF in a sample for a subject using a method according to any one of claims 1-51 or 128; andprognosing cancer in the subject based on the estimated cTF in the sample.
137. The method of claim 136, comprising administering a therapy based on the prognosis.
138. A method of diagnosing cancer in a subject, the method comprising:estimating cTF in a sample for a subject using a method according to any one of claims 1-51 or 128; anddetermining that the estimated cTF exceeds a threshold.
139. The method of claim 138, comprising initiating administration of a therapy based on the estimated cTF.
140. The method of claim 139, comprising selecting a dosing regimen for the therapy based on the estimated cTF.
141. A method of determining whether a cancer has been removed from a subject, the method comprising, after a subject has been administered a therapy to remove cancer and / or had a surgical removal of cancer, estimating cTF in a sample for the subject using a method according to any one of claims 1-51 or 128.
142. The method of claim 141, comprising continuing administration of a therapy based on the estimated cTF.
143. The method of claim 142, comprising ceasing administration of a therapy based on the estimated cTF.
144. A method of making a preliminary determination of a tumor of origin for ctDNA in a sample from a subject, the method comprising:Page 122 of 12413414664vlAttorney Docket No. : 2014191-0048estimating cTF in a sample for a subject using a first method according to any one of claims 1-51 or 128, wherein the cTF estimation model for the first method has been trained for a first type of cancer;estimating cTF in the using a second method according to any one of claims 1-51 or 128, wherein the cTF estimation model for the second method has been trained for a second type of cancer; andpreliminarily determining a tumor of origin based on a difference in the cTF estimated using the first method and the cTF estimated using the second method.
145. The method of claim 144, comprising selecting and performing a subsequent diagnostic assessment of the subject to determine the tumor of origin based on the preliminary determination of the tumor of origin.
146. A method of making a preliminary diagnosis of cancer in a subject, the method comprising:estimating cTF in a sample for a subject using a first method according to any one of claims 1-51 or 128, wherein the cTF estimation model for the first method has been trained for a first type of cancer;estimating cTF in the using a second method according to any one of claims 1-51 or 128, wherein the cTF estimation model for the second method has been trained for a second type of cancer; andpreliminarily diagnosing a cancer in the subject based on a difference in the cTF estimated using the first method and the cTF estimated using the second method.
147. The method of claim 146, comprising selecting and performing a subsequent diagnostic assessment of the subject for a particular cancer, wherein the particular cancer is selected based on the preliminary diagnosis.Page 123 of 12413414664vl