Promoter methylation

The multi-modal assessment of promoter methylation patterns through an epigenomic, genomic liquid biopsy platform addresses the imprecision in cancer therapy selection by providing comprehensive diagnostic insights, enhancing the accuracy of treatment predictions.

WO2026112439A1PCT designated stage Publication Date: 2026-05-28GUARDANT HEALTH INC +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2025-11-21
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Current cancer therapy selection for patients is imprecise due to the separation and incompatibility of genomic and epigenomic analysis methods, leading to incomplete diagnostic information and suboptimal treatment decisions.

Method used

A multi-modal assessment using an epigenomic, genomic liquid biopsy platform for simultaneous detection and analysis of promoter methylation patterns, particularly in tumor suppressor genes like BRCA1 and BRCA2, to predict therapy response and guide personalized treatment choices.

Benefits of technology

Enhances the accuracy of predicting therapy effectiveness by incorporating comprehensive genomic and epigenomic information, enabling precise treatment recommendations such as PARP inhibitors for patients with biallelic loss of BRCA1 or BRCA2.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025056580_28052026_PF_FP_ABST
    Figure US2025056580_28052026_PF_FP_ABST
Patent Text Reader

Abstract

In implementations described herein, methylation information is determined with respect to classification regions of a reference genome that are related to the presence of a tumor in a subject. The methylation information can be analyzed using a number of computational techniques to provide metrics related to the presence or absence of a tumor in a given subject, including a determination of tumor fraction.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: GH0259WOPROMOTER METHYLATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. provisional patent application no. 63 / 723,337, filed November 21, 2024, which is incorporated by reference herein in its entirety.BACKGROUND

[0002] Therapy selection for cancer patients is imprecise. DNA, RNA and protein from patient samples are often analyzed for patterns that can predict response to particular treatments. These biomarkers can range from single genes (e.g. EGFR, via real-time PCR) and proteins (e.g. HER2, via immuno-histochemistry) to complex genomic signatures (e.g. Tumor Mutational Burden, via Next-Generation Sequencing). Testing workflows for multiple types of analytes are generally separate and cannot be combined due to incompatibility in separation, chemical, and quantification processes etc. As a result, diagnostic tests do not examine the whole range of informative biomarkers and multiple, separate tests must be performed if multiomic results are desired. In most cases multiple tests are not performed and clinical decisions are made based on incomplete information for multiple reasons including lack of sufficient patient samples.

[0003] There is a great need in the art for improvements in personalized medicine. By simultaneously interrogating the genomic and epigenomic status of the patient sample, more accurate prediction of therapy effectiveness can be accomplished.

[0004] Described herein is simultaneous testing for and incorporating information derived from both the genomic and epigenomic components of the patient sample, such a diagnostic test will be able to take into consideration additional information not available otherwise. Of interest is methylation status, particularly of promoter region, including methylation patterns associated with epigenetic allelic status, which may account for instances genomic alterations fail to account varying efficacy of treatments in patients including for example, parp inhibitors (PARPi). Thus, promoter methylation (PM) is a critical epigenetic mechanism that silences tumor suppressor genes and is associated with key pathways in cancer progression. Traditionally, PM detection has relied on single-gene testing of tissue biopsies, which are invasive, less clinically accessible, and impractical for tracking tumor evolution. Here, the Inventors developed a several approaches for characterization PMAttorney Docket No.: GH0259WO within clinical multi-modal assessments using an epigenomic, genomic liquid biopsy platform which could provide significant insights into cancer biology.SUMMARY OF THE INVENTION

[0005] Described herein is a method, comprising: detecting methylation in one or more promoter regions of at least one of a plurality of genes; and generating a plurality of methylation calls to quantify methylation of the one or more promoter regions. In other embodiments, the method includes obtaining a sample. In other embodiments, the method includes having obtained a sample. In other embodiments, the method includes processing the quantities of methylation of the one or more promoter regions to characterize a sample. In other embodiments, the method includes characterizing the sample includes HRD, cancer derived promoter methylation, familial forms of colorectal cancer, or Lynch syndrome tumor types. In other embodiments, the promoter includes a region of 5kb upstream of the transcription start site (TSS), wherein the 5kb region is further refined using one or more of: costume panel regions, methylation peaks found in clinical samples, and excluding peaks found in normal samples. In other embodiments, the TSS is defined at the transcript level. In other embodiments, the TSS is defined at the gene level. In other embodiments, the method includes determining the ratio of the number of molecules that overlap a target region normalized by total positive control molecules. In other embodiments, the method includes determining the ratio includes filtering of a molecule based at least on the number of overlapping CpGs. In other embodiments, the method includes quantifying of methylation of the one or more promoter regions is based on the number of methylated CpGs. In other embodiments, the method includes refining the one or more promoter regions based at least on literature annotations, common methylation peak positions, and / or public datasets. In other embodiments, the genes comprise tumor suppressor genes, HRR genes, and IO genes. In other embodiments, the HRR genes comprise at least BRCA1 and BRCA2.

[0006] In other embodiments, the method includes comparing to a minimum methylation threshold derived from a population of training samples. In other embodiments, the training samples comprise cancer-free samples. In other embodiments, the minimum methylation threshold for calling includes at least one of: a minimum molecule count of 1- 100 and a minimum methylation score per gene is the max of: 95 quantile in normal + 8X105 or Median + 5 * median absolute deviation. In other embodiments, the method includes quantifying methylation of the one or more promoter regions is predictive of therapy response. In other embodiments, a plurality of quantitative measure weighted in a logisticAttorney Docket No.: GH0259WO regression model, optionally included fixed global weights. In other embodiments, the logistic regression model is a gated classifier, optionally a monotonic gated classifier. In other embodiments, the status is a driver, and weighting is a function of the status, optionally applied for noise abatement. In other embodiments, the method includes application of a gate around the threshold, optionally comprising a probabilistic function such as a Gaussian, Bayesian or probabilistic function. In other embodiments, quantifying methylation of the one or more promoter regions is combined with status, optionally MSI, HRD, MMR, TMB, or other cancer status. In other embodiments, predicting therapy response comprises use of a classifier, optionally including a gated classifier, and threshold in causing selection of a therapy for administration selected from one or more of an immune checkpoint inhibitor, poly (ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor. In other embodiments, the therapy includes one or more of an immune checkpoint inhibitor, poly (ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor. In other embodiments, the immune checkpoint inhibitor is Pembrolizumab. In other embodiments, the method includes poly (ADP-ribose) polymerase (PARP) inhibitor Olaparib or Talazoparib. In other embodiments, the therapy is a combination of a PI3K and mTOR inhibitor and a poly (ADP- ribose) polymerase (PARP) inhibitor. In other embodiments, the PI3K and mTOR inhibitor is Gedatolisib and the poly (ADP-ribose) polymerase (PARP) inhibitor is Talazoparib.

[0007] Described herein is a method comprising determining promoter regions of at least one of a plurality of genes, each obtained from a plurality of samples, determining methylation scores for the promoter regions to generate a plurality of methylation calls and / or quantification of promoter methylation, processing the plurality of methylation calls to generate a prediction that a test sample exhibits a genomic state.

[0008] Described herein is a method comprising: determining promoter regions of a plurality of genes, each obtained from a plurality of samples; determining methylation scores for the promoter regions to generate a plurality of methylation calls and / or quantification of promoter methylation; processing the plurality of methylation calls to generate a prediction that a test sample exhibits a genomic state. In other embodiments, the genomic state includes HRD, cancer derived promoter methylation, familial forms of colorectal cancer, or Lynch syndrome tumor types. In other embodiments, the promoter includes a region of 5kb upstream of the TSS, wherein the 5kb region is further refined using one or more of costume panel regions, methylation peaks found in clinical samples, and excluding peaks found in normal samples. In other embodiments, the TSS is defined at the transcript level. In otherAttorney Docket No.: GH0259WO embodiments, the TSS is defined at the gene level. In other embodiments, the methylation score is determined as the ratio of the number of molecules that overlap a target region normalized by total positive control molecules. In other embodiments, the molecule supporting the methylation score is filtered based at least on the number of overlapping CpGs. In other embodiments, the promoter regions are refined based at least on literature annotations, common methylation peak positions, and / or public datasets. In other embodiments, the genes comprise tumor suppressor genes, HRR genes, and IO genes. In other embodiments, the HRR genes comprise at least BRCA1 and BRCA2. In other embodiments, the calling includes deriving a minimum methylation threshold from a population of training samples. In other embodiments, the training samples comprise cancer- free samples. In other embodiments, the minimum methylation threshold for calling includes:

[0009] Minimum molecule count of 1-100 minimum and / or the minimum methylation score per gene is the max of 95 quantile in normal + 8X105 or Median + 5 * median absolute deviation. In other embodiments, the promoter methylation call combined with an MSI-H status, is predictive of therapy response. In other embodiments, the therapy includes one or more of an immune checkpoint inhibitor, poly (ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor. In other embodiments, the immune checkpoint inhibitor is Pembrolizumab. In other embodiments, the poly (ADP-ribose) polymerase (PARP) inhibitor Olaparib or Talazoparib. In other embodiments, the therapy is a combination of a PI3K and mTOR inhibitor and a poly (ADP-ribose) polymerase (PARP) inhibitor. In other embodiments, the PI3K and mTOR inhibitor is Gedatolisib and the poly (ADP-ribose) polymerase (PARP) inhibitor is Talazoparib.

[0010] Described herein is a method comprising determining promoter regions of BRCA1 and BRCA2, each obtained from a plurality of samples, determining methylation scores for the promoter regions to generate a plurality of methylation calls, processing the plurality of methylation calls to generate a prediction that a patient exhibits biallelic loss of BRCA1 or BRCA2.

[0011] Described herein is a method comprising determining promoter regions of BRCA1 and BRCA2, each obtained from a plurality of samples, determining methylation scores for the promoter regions to generate a plurality of methylation calls, processing the plurality of methylation calls to generate a prediction that a patient exhibits biallelic loss of BRCA1 or BRCA2, determining that the patient is a candidate for treatment with a PARPi.Attorney Docket No.: GH0259WO

[0012] Described herein is a method comprising determining promoter regions of BRCA1 and BRCA2, each obtained from a plurality of samples, determining methylation scores for the promoter regions to generate a plurality of methylation calls, processing the plurality of methylation calls to generate a prediction that a patient exhibits biallelic loss of BRCA1 or BRCA2, determining that the patient is a candidate for treatment with Gedatolisib and Talazoparib. In other embodiments, the methods include, wherein gedatolisib sensitizes advanced TNBC or BRCA1 / 2 mutant breast cancers to PARP inhibition with talazoparib.

[0013] Described herein is a method including determining promoter regions for MLH1, each obtained from a plurality of samples, determining methylation scores for the promoter region to generate a plurality of promoter methylation calls, determining from genomic data that the patient is BRAF V600E positive, wherein detection of the promoter methylation in a BRAF V600E positive patient identifies the patient as one who may be at risk for genetic / familial forms of colorectal cancer or Lynch syndrome-associated tumor types.

[0014] Described herein is a method, comprising: obtaining, by a computing system having one or more hardware processors and memory, sequencing reads derived from a sample of a subject, determining one or more classification regions corresponding to a plurality of genes included in the sample; and determine a methylation level of the one or more classification regions by generating a quantitative measure derived from the sequencing reads in the sample of the subject. In other embodiments, the method includes obtaining a sample. In other embodiments, the method includes having obtained a sample. In other embodiments, the method includes processing the methylation level of the one or more classification regions to characterize the sample. In other embodiments, the method includes characterizing the sample comprises determining HRD status, promoter methylation associated with cancer. In other embodiments, the quantitative measure comprises determining the ratio of the number of molecules that overlap a classification region normalized by total positive control molecules, wherein the molecules exhibit a threshold amount of methylated cytosines. In other embodiments, the quantitative measure is compared to a predetermined threshold value to call methylation status of the one or more classification regions. In other embodiments, determining the ratio comprises filtering of a molecule based at least on a threshold amount of methylated cytosines. In other embodiments, determining a methylation level of the one or more classification regions is based on the number of methylated CpGs. In other embodiments, the e classification regions comprise promoter regions. In other embodiments, the one or more classification regions individually correspondAttorney Docket No.: GH0259WO to genomic regions in which a methylation rate of cytosines in the genomic regions of nucleic acids derived from cells obtained from subjects in which cancer is present is different from a methylation rate of cytosines in the genomic regions of nucleic acids derived from cells obtained from subjects in which cancer is not present. In other embodiments, the plurality of samples and the additional sample include cell free nucleic acids. In other embodiments, the method includes performing, by the computing system, a training process using the training data to generate the model, wherein the training process includes: determining, by the computing system, one or more additional weights of individual samples included in the training data based on the indication of cancer for the individual samples being within a threshold confidence level. In other embodiments, the indication of cancer for an individual sample is outside of the threshold confidence level and the method comprises: applying, by the computing system, a penalty to a weight of the individual sample during the training process. The method of any preceding claim, comprising: performing, by the computing system and using the one or more machine learning algorithms, one or more first iterations of the training process for the model using a portion of the training data; and generating, by the computing system, first output data for the model based on the one or more first iterations of the training process, the first output data corresponding to one or more first additional indications of cancer being present in first individual subjects of the plurality of subjects, the first individual subjects corresponding to the portion of the training data. In other embodiments, the method includes combining, by the computing system, the first output data and the training data to produce additional training data; performing, by the computing system, one or more second iterations of the training process for the model using a portion of the additional training data; and generating, by the computing system, second output data for the model based on the one or more second iterations of the training process, the second output data indicating one or more second additional indications of cancer being present in second individual subjects of the plurality of subjects, the second individual subjects corresponding to the portion of the additional training data. In other embodiments, the weights for the individual classification regions of the plurality of classification regions are determined based on the first output data and the second output data. In other embodiments, the method includes determining, by the computing system, that a number of indications of cancer being present that were determined during one or more iterations of the training process are at least a threshold value for one or more samples included in the training data; and determining, by the computing system, that modifications to one or more weights of the model are not modified or are modified by a minimal amount. In other embodiments, theAttorney Docket No.: GH0259WO method includes determining, by the computing system, that an additional number of indications of cancer being present that were determined during the one or more iterations of the training process are less than the threshold value for one or more additional samples included in the training data; and determining, by the computing system, that modifications to one or more additional weights of the model are modified by more than the minimal amount. In other embodiments, the method includes combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleic acid-MBD protein solution; and performing a plurality of washes of the nucleic acid-MBD protein solution with a salt solution to produce a number of nucleic acid fractions, individual nucleic acid fractions having a threshold number of methylated cytosines in regions of the plurality of nucleic acids having at least the threshold cytosine-guanine content. In other embodiments, the wash of the plurality of washes is performed with a solution having a concentration of sodium chloride (NaCl) and produces a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins. In other embodiments, the method includes determining that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding strengths to MBD proteins; attaching a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode being included in a first set of molecular barcodes associated with the first partition; determining that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding energies to MBD proteins different from the first range of binding strengths to MBD proteins; and attaching a second molecular barcode to nucleic acids of the second nucleic acid fraction, the second molecular barcode being included in a second set of molecular barcodes associated with the second partition.

[0015] In other embodiments, the method includes combining at least a portion of the number of nucleic acid fractions with an amount of restriction enzyme that cleaves molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads, wherein the threshold amount of methylated cytosines corresponds to a minimum frequency of methylated cytosines within a region having at least the threshold cytosine-guanine content.

[0016] Described herein is a method, comprising: obtaining, by a computing system having one or more hardware processors and memory, sequencing reads derived from a sample of a subject, determining one or more classification regions corresponding to aAttorney Docket No.: GH0259WO plurality of genes included in the sample, determine a methylation level of the one or more classification regions by generating a quantitative measure comprising the ratio of the number of molecules that overlap a classification region normalized by total positive control molecules, wherein the molecules exhibit a threshold amount of methylated cytosines; and comparing the quantitate measure to a predetermined threshold value to call methylation status of the one or more classification regions.

[0017] In various embodiments, determination of a quantitative measure can include combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleic acid-MBD protein solution; and performing a plurality of washes of the nucleic acid- MBD protein solution with a salt solution to produce a number of nucleic acid fractions. In some instances, individual nucleic acid fractions having a threshold number of methylated cytosines in regions of the plurality of nucleic acids having at least the threshold cytosine- guanine content. Thereafter, a wash of the plurality of washes is performed with a solution having a concentration of sodium chloride (NaCl) and produces a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins.

[0018] One may determine that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding strengths to MBD proteins; attach a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode being included in a first set of molecular barcodes associated with the first partition, and subsequently determine that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding energies to MBD proteins different from the first range of binding strengths to MBD proteins; and thereafter attach a second molecular barcode to nucleic acids of the second nucleic acid fraction, the second molecular barcode being included in a second set of molecular barcodes associated with the second partition.

[0019] In some instances, one may combine at least a portion of the number of nucleic acid fractions with an amount of restriction enzyme that cleaves molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads, wherein the threshold amount of methylated cytosines corresponds to a minimum frequency of methylated cytosines within a region having at least the threshold cytosine-guanine content.Attorney Docket No.: GH0259WO

[0020] Additionally, one can combine at least a portion of the number of nucleic acid fractions with an amount of a restriction enzyme that cleaves molecules with one or more methylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads, wherein the threshold amount of unmethylated cytosines corresponds to a maximum frequency of methylated cytosines that are not cleaved within a region having at least the threshold cytosine-guanine content.

[0050] Described herein is a method including obtaining training sequence data including training sequencing reads derived from a plurality of samples of a plurality of subjects, individual training sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and an amount of methylated cytosines included in regions of the nucleotide sequence having cytosine-guanine content, analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that have an amount of methylated cytosines in subjects in which cancer is detected, analyzing the training sequencing reads to determine a second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have an amount of methylated cytosines in which cancer is not detected, determining a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions, generating training data that includes the metric for the individual classification regions of the plurality of classification regions for the training sequence reads from samples of training subjects,

[0051] implementing one or more machine learning algorithms to generate a model, the model including weights for individual classification regions of the plurality of classification regions and at least a portion of the weights of the individual classification regions being different from one another. In other embodiments, the method includes obtaining testing sequence data from an additional subject that is not included in the plurality of subjects, the testing sequence data including testing sequencing reads derived from a sample of the additional subject, individual testing sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in the additional sample andAttorney Docket No.: GH0259WO individual testing sequencing reads corresponding to molecules having an amount of methylated cytosines included in regions of the nucleotide sequence having cytosine-guanine content, and determining, using the model and the additional sequence data, a measurement of tumor fraction in the additional subject, from one another.

[0052] In other embodiments, the method includes selecting a sub-set of the plurality of classification regions, from one another. In other embodiments, the sub-set of the plurality of classification regions comprise one or more cancer-specific regions, from one another. In other embodiments, metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions includes a sub-set of the plurality of classification regions, from one another.

[0053] In other embodiments, the method includes analyzing the testing sequencing reads to determine a first quantitative measure derived from the testing sequencing reads that correspond to the individual classification regions of the plurality of classification regions, analyzing the testing sequencing reads to determine a second quantitative measure derived from the testing sequencing reads that correspond to the individual control regions of the plurality of control regions, normalizing the second quantitative measure based on the corresponding individual control regions of the plurality of control regions, determining the metric for the individual classification regions based on the first quantitative measure for the individual classification regions and the normalized second quantitative measure for the plurality of control regions, and applying a machine learning algorithm to the metrics for the individual classification regions to determine a measurement of tumor fraction in the additional subject, from one another.

[0054] In other embodiments, the one or more machine learning algorithms include one or more classification algorithms, from one another. In other embodiments, the one or more machine learning algorithms include one or more regression algorithms. In other embodiments, the method includes applying a machine learning algorithm to the metrics for the individual classification regions to determine a measurement of tumor fraction in the additional subject includes selecting a sub-set of the plurality of classification regions. In other embodiments, the training data includes the individual measures of tumor fraction for the individual samples of the plurality of samples, and the model is generated based on the individual measures of tumor fraction for the individual samples of the plurality of samples. In other embodiments, the metric for the individual classification regions is determined based on a scaling factor and / or an error correction factor. In other embodiments, the plurality ofAttorney Docket No.: GH0259WO classification regions individually correspond to genomic regions in which a methylation rate of cytosines in the genomic regions of nucleic acids derived from cells obtained from subjects in which cancer is present is different from a methylation rate of cytosines in the genomic regions of nucleic acids derived from cells obtained from subjects in which cancer is not present. In other embodiments, the plurality of classification regions correspond to a first plurality of classification regions for a first cancer type and the model can be generated for a second cancer type based on a second plurality of classification regions that are different from the first plurality of classification regions.

[0055] Further described herein is a method including obtaining sequencing reads derived from a sample obtained from a subject, individual sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in the sample and corresponding to molecules having a threshold amount of methylated cytosines included in regions of the nucleotide sequence having cytosine-guanine content, determining, by the computing system, a first quantitative measure derived from the sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome with amount of methylated cytosines in subjects in which cancer is detected, analyzing, by the computing system, the sequencing reads to determine a second quantitative measure derived from the sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have cytosine-guanine content and an amount of methylated cytosines in additional subjects in which cancer is not detected, determining, by the computing system, a plurality of metrics with individual metrics of the plurality of metrics corresponding to individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions, and determining a measurement of tumor fraction in the additional subject.

[0056] In other embodiments, the method includes selecting a sub-set of the plurality of classification regions. In other embodiments, the sub-set of the plurality of classification regions comprise one or more cancer-specific regions. In other embodiments, the metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions includes a sub-set of the plurality of classificationAttorney Docket No.: GH0259WO regions. In other embodiments, the method includes determining an order of the values of the plurality of metrics, and determining a subset of classification regions from among the plurality of classification regions based on the order, wherein a portion of the plurality of metrics that correspond to the subset of the classification regions is used to determine a measurement of tumor fraction in the additional subject. In other embodiments, the method includes determining a measurement of tumor fraction in the additional subject includes applying a scaling factor. In other embodiments, the determined measurement of tumor fraction corresponds to an indication of cancer status in the subject. In other embodiments, the method includes determining a measurement of tumor fraction in the subject includes, applying a model generated from training data.

[0057] In other embodiments, the model generated from training data includes: obtaining training sequence data including training sequencing reads derived from a plurality of samples of a plurality of subjects, individual training sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and an amount of methylated cytosines included in regions of the nucleotide sequence having cytosine-guanine content, analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that have a threshold amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content, analyzing the training sequencing reads to determine a second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have at least the threshold cytosine-guanine content and that have a threshold amount of methylated cytosines in which cancer is not detected, determining a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions, generating training data that includes the metric for the individual classification regions of the plurality of classification regions for the training sequence reads from samples of training subjects, implementing one or more machine learning algorithms to generate the model.Attorney Docket No.: GH0259WO

[0058] Also described herein is method including: obtaining testing sequence data from a subject, the testing sequence data including testing sequencing reads derived from a sample of the subject, individual testing sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in the additional sample and individual testing sequencing reads corresponding to molecules having an amount of methylated cytosines included in regions of the nucleotide sequence, analyzing the testing sequencing reads to determine a first quantitative measure derived from the testing sequencing reads that correspond to individual classification regions of a plurality of classification regions at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that have an amount of methylated cytosines in subjects in which cancer is detected, analyzing the testing sequencing reads to determine a second quantitative measure derived from the testing sequencing reads that correspond to individual control regions a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have an amount of methylated cytosines in additional subjects in which cancer is not detected,

[0059] determining a metric for the individual classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions, and generating training data that includes the metric for the individual classification regions of the plurality of classification regions for the training sequence reads from samples of training subjects,

[0060] implementing one or more machine learning algorithms to generate a model, the model including weights for individual classification regions of the plurality of classification regions and at least a portion of the weights of the individual classification regions being different from one another to determine a measurement of tumor fraction in the subject. In other embodiments, the method includes obtaining training sequence data including training sequencing reads derived from a plurality of samples of a plurality of training subjects, individual training sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and individual training sequencing reads corresponding to molecules having a threshold amount of methylated cytosines included in regions of the nucleotide sequence having at least a threshold cytosine-guanine content, analyzing the training sequencing reads to determine an additional first quantitative measure derived from the training sequencing reads that corresponds to individual classification regions of the plurality of classification regions,Attorney Docket No.: GH0259WO analyzing the training sequencing reads to determine an additional second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions, determining an additional metric for the individual classification regions of the plurality of classification regions based on the additional first quantitative measure for the individual classification regions and the additional second quantitative measure for the plurality of control regions, generating training data that includes the additional metric for the individual classification regions of the plurality of classification regions for the training sequence reads from samples of the plurality of training subjects, implementing using the training data, one or more machine learning algorithms to generate the model to determine the indications of cancer status in subjects based on amounts of methylated cytosines in at least a portion of the plurality of classification regions. In other embodiments, the one or more machine learning algorithms include one or more classification algorithms. In other embodiments, the one or more machine learning algorithms include one or more regression algorithms, and the indication corresponds to an estimate of tumor fraction of the sample. In other embodiments, the method includes the training sequencing reads comprise a first portion of the training sequence data and additional training sequencing reads comprise a second portion of the training sequence data, wherein the additional training sequencing reads are different from the training sequencing reads, and the method including: analyzing at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine an individual frequency of a plurality of variants present in an individual sample of the plurality of samples, determining for the individual sample, a variant of the plurality of variants having a maximum frequency that corresponds to the individual frequency having a greatest value among individual frequencies derived from an individual sample, and determining individual measures of tumor fraction for an individual sample based on the greatest value of the individual frequencies derived from the individual sample. In other embodiments, the training data includes the individual measures of tumor fraction for the individual samples of the plurality of samples, and the model is generated based on the individual measures of tumor fraction for the individual samples of the plurality of samples. In other embodiments, the sample of the subject and the plurality of samples of the plurality of training subjects include cell free nucleic acids.

[0061] Further described herein is a method including: obtaining sequencing reads derived from one or more samples obtained from a subject, individual sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in the sample and corresponding to molecules having a threshold amount of methylated cytosines included inAttorney Docket No.: GH0259WO regions of the nucleotide sequence having at least a threshold cytosine-guanine content, determining a first quantitative measure derived from the sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that an amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content, analyzing the sequencing reads to determine a second quantitative measure derived from the sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have at least the threshold cytosine-guanine content and that have an amount of methylated cytosines in additional subjects in which cancer is not detected, determining a plurality of metrics with individual metrics of the plurality of metrics corresponding to individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions, and determining an indication of cancer status in the subject based on at least a portion of the plurality of metrics. In other embodiments, the method includes selecting a sub-set of the plurality of classification regions. In other embodiments, the sub-set of the plurality of classification regions comprise one or more cancer-specific regions. In other embodiments, the metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions includes a sub-set of the plurality of classification regions. In other embodiments, the method includes at least two samples obtained from a subject In other embodiments, the method includes determining a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions comprsises: selecting a sub-set of the plurality of classification regions based on a regression algorithm based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions. In other embodiments, the first quantitative measure is normalized based on second quantitative measure. In other embodiments, the plurality of samples and the additional sample include cell free nucleic acids. In other embodiments, the method includes combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleicAttorney Docket No.: GH0259WO acid-MBD protein solution, and performing a plurality of washes of the nucleic acid-MBD protein solution with a salt solution to produce a number of nucleic acid fractions, individual nucleic acid fractions having a threshold number of methylated cytosines in regions of the plurality of nucleic acids having at least the threshold cytosine-guanine content. In other embodiments, the method includes a wash of the plurality of washes is performed with a solution having a concentration of sodium chloride (NaCl) and produces a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins. In other embodiments, the method includes determining that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding strengths to MBD proteins, attaching a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode being included in a first set of molecular barcodes associated with the first partition, determining that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding energies to MBD proteins different from the first range of binding strengths to MBD proteins, and attaching a second molecular barcode to nucleic acids of the second nucleic acid fraction, the second molecular barcode being included in a second set of molecular barcodes associated with the second partition. In other embodiments, the method includes combining at least a portion of the number of nucleic acid fractions with an amount of restriction enzyme that cleaves molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads, wherein the threshold amount of methylated cytosines corresponds to a minimum frequency of methylated cytosines within a region having at least the threshold cytosine-guanine content. In other embodiments, the method includes combining at least a portion of the number of nucleic acid fractions with an amount of a restriction enzyme that cleaves molecules with one or more methylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads, wherein the threshold amount of unmethylated cytosines corresponds to a maximum frequency of methylated cytosines that are not cleaved within a region having at least the threshold cytosine-guanine content. In other embodiments, the method includes a limit of detection for the model to determine tumor fraction of samples is no greater than 0.05%. 0.05%.

[0062] In one or more aspects, a method includes obtaining, by a computing system having one or more hardware processors and memory, training sequence data including training sequencing reads derived from a plurality of samples of a plurality of subjects, individualAttorney Docket No.: GH0259WO training sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and individual training sequencing reads corresponding to molecules having a threshold amount of methylated cytosines included in regions of the nucleotide sequence having at least a threshold cytosine- guanine content. The method also includes analyzing, by the computing system, the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that have the threshold amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content. The method also includes analyzing, by the computing system, the training sequencing reads to determine a second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have at least the threshold cytosine- guanine content and that have at least the threshold amount of methylated cytosines in subjects in which cancer is detected and in additional subjects in which cancer is not detected. The method also includes determining, by the computing system, a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions. The method also includes generating, by the computing device, training data that includes the metric for the individual classification regions of the plurality of classification regions for the training sequence reads from samples of training subjects. The method also includes implementing, by the computing system and using the training data, one or more machine learning algorithms to generate a model to determine an indication of cancer status in subjects based on amounts of methylated cytosines in at least a portion of the plurality of classification regions, the model including weights for individual classification regions of the plurality of classification regions and at least a portion of the weights of the individual classification regions being different from one another.

[0063] In one or more aspects, the method includes obtaining, by the computing system, testing sequence data from an additional subject that is not included in the plurality of subjects, the testing sequence data including testing sequencing reads derived from a sample of the additional subject, individual testing sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in the additional sample andAttorney Docket No.: GH0259WO individual testing sequencing reads corresponding to molecules having at least the threshold amount of methylated cytosines included in regions of the nucleotide sequence having at least the threshold cytosine-guanine content, and determining, using the model and the additional sequence data, the indication of cancer status in the additional subject.

[0064] In one or more aspects, the method includes analyzing, by the computing system, the testing sequencing reads to determine a first quantitative measure derived from the testing sequencing reads that correspond to the individual classification regions of the plurality of classification regions, analyzing, by the computing system, the testing sequencing reads to determine a second quantitative measure derived from the testing sequencing reads that correspond to the individual control regions the plurality of control regions, determining, by the computing system, the metric for the individual classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions, and generating, by the computing system, an input vector that includes the metrics for the individual classification regions, where the model uses the input vector to determine the indication of cancer status in the additional subject. In one or more aspects, the one or more machine learning algorithms include one or more classification algorithms and the indication of cancer status corresponds to a probability of cancer status in the additional subject. In one or more aspects, the one or more machine learning algorithms include one or more regression algorithms and the indicator corresponds to an estimate of tumor fraction of the additional sample.

[0065] In one or more aspects, the training sequencing reads comprise a first portion of the training sequence data and additional training sequencing reads comprise a second portion of the training sequence data, where the additional training sequencing reads are different from the training sequencing reads and the method includes analyzing, by the computing system, at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine an individual frequency of a plurality of variants present in an individual sample of the plurality of samples, determining, by the computing system and for the individual samples, a variant of the plurality of variants having a maximum frequency that corresponds to the individual frequency having a greatest value among individual frequencies derived from an individual sample, and determining, by the computing system, individual measures of tumor fraction for an individual sample based on the greatest value of the individual frequencies derived from the individual sample. In one or more aspects, the training data includes the individual measures of tumor fraction for the individual samples of the plurality of samples and the model is generated based on the individual measures ofAttorney Docket No.: GH0259WO tumor fraction for the individual samples of the plurality of samples. In one or more aspects, the metric for the individual classification regions is determined based on a scaling factor and an error correction factor. In one or more aspects, the plurality of classification regions individually correspond to genomic regions in which a methylation rate of the genomic regions in nucleic acids derived from cells obtained from subjects in which cancer is present is different from a methylation rate of the genomic regions in nucleic acids derived from cells obtained from subjects in which cancer is not present.

[0066] In one or more aspects, the plurality of classification regions correspond to a first plurality of classification regions for a first cancer type and the model can be generated for a second cancer type based on a second plurality of classification regions that are different from the first plurality of classification regions. In one or more aspects, the plurality of samples and the additional sample include cell free nucleic acids. In one or more aspects, the method includes performing, by the computing system, a training process using the training data to generate the model, where the training process includes determining, by the computing system, one or more additional weights of individual samples included in the training data based on the indication of cancer for the individual samples being within a threshold confidence level. In one or more aspects, the indication of cancer for an individual sample is outside of the threshold confidence level and the method includes applying, by the computing system, a penalty to a weight of the individual sample during the training process. In one or more aspects, the method includes performing, by the computing system and using the one or more machine learning algorithms, one or more first iterations of the training process for the model using a portion of the training data, and generating, by the computing system, first output data for the model based on the one or more first iterations of the training process, the first output data corresponding to one or more first additional indications of cancer status in first individual subjects of the plurality of subjects, the first individual subjects corresponding to the portion of the training data.

[0067] In one or more aspects, the method includes combining, by the computing system, the first output data and the training data to produce additional training data, performing, by the computing system, one or more second iterations of the training process for the model using a portion of the additional training data, and generating, by the computing system, second output data for the model based on the one or more second iterations of the training process, the second output data indicating one or more second additional indications of cancer status in second individual subjects of the plurality of subjects, the second individual subjects corresponding to the portion of the additional training data. In one or more aspects, theAttorney Docket No.: GH0259WO weights for the individual classification regions of the plurality of classification regions are determined based on the first output data and the second output data. In one or more aspects, the method includes determining, by the computing system, that a number of indications of cancer status that were determined during one or more iterations of the training process are at least a threshold value for one or more samples included in the training data, and determining, by the computing system, that modifications to one or more weights of the model are not modified or are modified by a minimal amount. In one or more aspects, the method includes determining, by the computing system, that an additional number of indications of cancer status that were determined during the one or more iterations of the training process are less than the threshold value for one or more additional samples included in the training data, and determining, by the computing system, that modifications to one or more additional weights of the model are modified by more than the minimal amount.

[0068] In one or more aspects, the method includes combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleic acid-MBD protein solution, and performing a plurality of washes of the nucleic acid-MBD protein solution with a salt solution to produce a number of nucleic acid fractions, individual nucleic acid fractions having a threshold number of methylated cytosines in regions of the plurality of nucleic acids having at least the threshold cytosine-guanine content.

[0069] In one or more aspects, a wash of the plurality of washes is performed with a solution having a concentration of sodium chloride (NaCl) and produces a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins.

[0070] In one or more aspects, the method includes determining that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding strengths to MBD proteins, attaching a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode being included in a first set of molecular barcodes associated with the first partition, determining that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding strengths to MBD proteins different from the first range of binding strengths to MBD proteins, and attaching a second molecular barcode to nucleic acids of the second nucleic acid fraction, the second molecular barcode being included in a second set of molecular barcodes associated with the second partition.Attorney Docket No.: GH0259WO

[0071] In one or more aspects, the method includes combining at least a portion of the number of nucleic acid fractions with an amount of one or more methylation sensitive restriction enzymes that cleave molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads. In one or more aspects, the method includes combining at least a portion of the number of nucleic acid fractions with an amount of one or more methylation dependent restriction enzymes that cleaves molecules with one or more methylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads. In one or more aspects, a limit of detection for the model to determine tumor fraction of samples is no greater than 0.05%.

[0072] In one or more aspects, a computing system includes: one or more hardware processors, and one or more non-transitory computer-readable storage media including computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including: obtaining training sequence data including training sequencing reads derived from a plurality of samples of a plurality of subjects, individual training sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and individual training sequencing reads corresponding to molecules having a threshold amount of methylated cytosines included in regions of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that have the threshold amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have at least the threshold cytosine-guanine content and that have at least the threshold amount of methylated cytosines in subjects in which cancer is detected and in additional subjects in which cancer is not detected. The operations also include determining a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions andAttorney Docket No.: GH0259WO the second quantitative measure for the plurality of control regions. The operations also include generating training data that includes the metric for the individual classification regions of the plurality of classification regions for the training sequence reads from samples of training subjects. The operations also include implementing, using the training data, one or more machine learning algorithms to generate a model to determine an indication of cancer status in subjects based on amounts of methylated cytosines in at least a portion of the plurality of classification regions, the model including weights for individual classification regions of the plurality of classification regions and at least a portion of the weights of the individual classification regions being different from one another.

[0073] In one or more aspects, one or more computer-readable storage media comprise computer-readable instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations including: obtaining training sequence data including training sequencing reads derived from a plurality of samples of a plurality of subjects, individual training sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and individual training sequencing reads corresponding to molecules having a threshold amount of methylated cytosines included in regions of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads that corresponds to individual classification regions of a plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome that have the threshold amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions, individual control regions of the plurality of control regions corresponding to additional genomic regions of the reference genome that have at least the threshold cytosine-guanine content and that have at least the threshold amount of methylated cytosines in subjects in which cancer is detected and in additional subjects in which cancer is not detected. The operations also include determining a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions. The operations also include generating training data that includes the metric for the individual classificationAttorney Docket No.: GH0259WO regions of the plurality of classification regions for the training sequence reads from samples of training subjects. The operations also include implementing, using the training data, one or more machine learning algorithms to generate a model to determine an indication of cancer status in subjects based on amounts of methylated cytosines in at least a portion of the plurality of classification regions, the model including weights for individual classification regions of the plurality of classification regions and at least a portion of the weights of the individual classification regions being different from one another.

[0074] In one or more aspects, a method includes obtaining a first sample from a subject and a second sample from the subject, obtaining, by a computing system having one or more hardware processors and memory, sequence data including sequencing reads derived from a plurality of samples of a subject, individual sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples and individual training sequencing reads corresponding to molecules having a threshold amount of methylated cytosines included in regions of the nucleotide sequence having at least a threshold cytosine-guanine content. The method also includes analyzing, by the computing system, first sequencing reads included in the sequence data to determine first quantitative measures that correspond to individual first classification regions of a plurality of first classification regions, at least a portion of the individual first classification regions of the plurality of first classification regions corresponding to first genomic regions of a reference genome that have the threshold amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content. The method also includes analyzing, by the computing system, second sequencing reads included in the sequence data to determine second quantitative measures that correspond to individual second classification regions of a plurality of second classification regions, at least a portion of the individual second classification regions of the plurality of second classification regions corresponding to second genomic regions of a reference genome that have the threshold amount of methylated cytosines in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content. In addition, the method includes determining, by the computing system, one or more first classification regions that overlap with one or more second classification regions to produce third classification regions. Further, the method includes analyzing, by the computing system, at least one of the first quantitative measures or the second quantitative measures to determine an indication of cancer status in the subject. The indication of cancer status can include at least one of tumor fraction or minor allele frequency.Attorney Docket No.: GH0259WO

[0075] In one or more aspects, the method includes obtaining the first sample before at least one of a procedure or administration of a treatment for cancer and obtaining the second sample after at least one of the procedure or the administration of the treatment for cancer.

[0076] In one or more aspects, the method includes determining, by the computing system, the first quantitative measures by analyzing a number of first sequencing reads included in the sequence data that correspond to individual first classification regions in relation to a total number of the first sequencing reads that correspond to a group of the first classification regions. In one or more examples, the group of the first classification regions can include all of the first classification regions.

[0077] In one or more aspects, the method includes, determining, by the computing system, an individual first classification region by: determining, by the computing system, a number of the first sequencing reads that correspond to a genomic region of the reference genome, determining, by the computing system, a portion of the genomic region for which at least a threshold amount of the number of first sequencing reads overlap, and determining, by the computing system, that the portion of the genomic region corresponds to the individual first classification region. The genomic region can include a differentially methylated region. In one or more aspects, the threshold amount can include at least 70% of the number of first sequencing reads.

[0078] Described herein is a method for generating a tumor fraction estimate from a ceil-free deoxyribonucleic acid (cfDNA) sample of a subject, including receiving a dataset of methylation sequence reads from a cfDNA sample of a subject, determining at each of a plurality of one or more methylation levels and generating a methylation pattern over one or more CpG sites; comparing the plurality of variants to reference sequence reads to generate a subset of variants. In some embodiments the subset of variants is generated by comparison to reference sequence reads generated from non-cancer cfDNA samples. In some instances, the reference sequence reads are obtained from biopsy samples of a plurality of tissues of reference individuals. In various embodiments, for one or more variants, including in a subset a variant, a count of methylation sequence reads that include the variant; inputting the counts of methylation sequence reads for the variants of the subset to a model. In various embodiments, the model is trained based frequency rates of the plurality of variants. In some embodiments, the method includes generating a tumor fraction estimate of the cfDNA sample. In various embodiments, occurrence, recurrence or alterations rates of the plurality of variants are determined based on the reference sequence reads in the bank. In various embodiments, comparing the plurality of variants to reference sequence reads to generate aAttorney Docket No.: GH0259WO subset of variants comprises filtering out one or more variants whose rates of presence in the noncancer samples exceeds a threshold. In various embodiments, the particular occurrence, recurrence or alteration of a particular variant corresponds to a rate of observation of the particular variant among the reference sequence reads in the bank. In various embodiments, the tumor fraction prediction is a distribution of probability of a fraction of fragments in the cfDNA sample that are tumor derived. In various embdoiments, the tumor fraction prediction is a fraction of fragments in the cfDNA sample that is tumor derived. In various embodiments, the model comprises at least one probabilistic model, the probabilistic model comprising a Poisson distribution for a particular variant, and the Poisson distribution is weighted by the recurrence rate of the particular variant. In various embodiments, the method includes a plurality of probabilistic distributions, each probabilistic distribution corresponding to a particular variant and parameterized based on a site-specific noise rate of the particular variant and per-site sequencing depth of the particular variant. In various embodiments, the h probabilistic distribution corresponding to a particular variant is further adjusted based on at least one of: a depth of the cfDNA sample, a binding panel efficiency of the cfDNA sample, and an estimated tumor fraction of the cfDNA sample. In various embodiments, the count for each variant of the filtered subset comprises a count of methylation sequence reads of the cfDNA sample that include the methy lation pattern over the one or more CpG sites of the variant.

[0079] A system including instructions for processing the methods of any preceding embodiment.

[0080] 4 computer readable medium including instructions for processing the methods of any preceding embodiment.BRIEF DESCRIPTION OF FIGURES

[0021] Figure 1. Breast vs BRCA. The prevalence of S3 vs TCGA in each analogous cancer type is captured in scatterplots. Left plot shows the prevalence after suppression of subcl onal PM, right plot shows all PM+ with call=l by single region caller.

[0022] Figure 2. CRC vs COAD+READ.

[0023] Figure 3. NSCLC vs LUAD+LUSC.

[0024] Figure 4. MLH1 using epitracer tool from hyper methylated bams. These genomic regions are evenly spaced windows with no overlap. Counts used epi tracer tool (filters by quality, outputs counts per base). Includes the same cohort as slide 1-2 but I included -100 cancer free samples. This heatmap shows the relationship between MLH1 and MSI and the story of the region optimization using cancer free samples.Attorney Docket No.: GH0259WO

[0025] Figure 5. MLH1 and MSI Concordance (no suppression) MLH1 PM was significantly associated with microsatellite instability (MSI) high status in both colorectal cancer (CRC), 54%, and endometrial carcinomas (EC), 100% of MSI-H cases (p<0.001).

[0026] Figure 6. MLH1 and MSI Concordance

[0027] Figure 7. MLH1 and MSI Concordance with sub-clonal suppression

[0028] Figure 8. BRAF Pathogenic SNV (no suppression) MLH1 PM in our cohort of CRC showed significant enrichment for BRAF V600E (36%, p<0.001).

[0029] Figure 9. BRCA1 LOH (no suppression) High prevalence of BRCA1 PM was observed in breast (BC) and ovarian cancer (OC), with enrichment of PM in BC with loss of heterozygosity (LOH) (6%, p=0.02).

[0030] Figure 10. MGMT

[0031] Figure 11. MLH1 promoter region definition refinement. MLH1 original promoter region

[0032] chr3: 37033671-37035422 = 175 Ibp region. TRUE MSI-H has distinct MLH1 promoter methylation pattern (proximal region to TSS) as shown in the IGV. Similar between different cancer types.

[0033] Figure 12. MLH1 vs MSI-H association. MLH1 methyl / MSS or MSI-H cases

[0034] For MSI-H samples, -72% cases are MLH1 methylated In some instances,MLH1 methylation with MSS may be due to MLH1 silencing with no genomic instability.

[0035] Distal regions MLH1 promoter may be age related, proximal may be disease related

[0036] Partial methylation, tumor fraction.

[0037] Figure 13. Proposed optimized MLH1 promoter regions.

[0038] Figure 14. Mismatch Repair Mutations. MMR genes include 'POLD3', 'MLH3', 'MSH6', 'RPA4', 'LIGF, 'MLH1', 'MSH2', 'MSH3', 'PCNA', 'PMS2', 'POLDI', 'POLD2', 'POLD4', 'RFC1', 'RFC2', 'RFC3', 'RFC4', 'RFC5', 'RPA1', 'RPA2', 'RPA3', 'SSBP1', 'EXO1'. Variants in MLH1, MSH2, MSH6 were found to be pathogenic / likely pathogenic. If MSI status is truth & excluding samples with known path variants, PPA is - 70%. NPA is not a problem given high prevalence of MLH1 WT.

[0039] Figure 15. Enrichment of endometrial cancers in MLH1 cohort.

[0040] Figure 16. BRAF variants, MLH1 PM and MMR mutation status co-occurring events. Based on observations only the C region of the MLH1 promoter was considered for further analyses when assessing methylation as a predictor of MMR mutation status. Tumor methylation in the C region of the MLH1 promoter was significantly decreased in knownAttorney Docket No.: GH0259WOMLH1 mutation carriers compared to MSI-H tumors from MMR mutation-negative cases overall (6% vs 47%; p<0.00001), and also compared to the subset of MSI-H tumors from MMR mutation-negative cases demonstrating tumor MLH1 protein loss (6% vs 40%; p<0.00001). Methylation of the ‘C region’ was a predictor of MMR mutation-negative status in MSI-H CRC cases (47% vs 6% in MLH1 mutation carriers, p<0.0001). BRAF V600E mutation, and MLH1 promoter ‘C region’ methylation specifically, are strong predictors of negative MMR mutation status.

[0041] Figure 17. MLH1 calling with new MLH1 definitions. Test if new LoD change with the new definitions with same threshold (min score 8.3e-05, min molecule 10). Estimated LoD 95 from simulations on 3 technical replicates. Defined at the min level where MLH1 PM detected in all 3 replicates. #MLH1 molecules under different settings. Simulated from runs in 3 technical replicates.

[0042] Figure 18. STAT5A Promoter Methylation and related pathways in melanoma cell.

[0043] Figure 19. STAT5A Region Refinement. Methylation probes most likely associated with gene silencing when correlating TCGA expression and methylation. Current STAT5A promoter region covers both probes sufficiently. Current region also avoids methylation observed in cancer free samples

[0044] Figure 20. High concordance in cancer type prevalence compared to TCGA. Highly concordant cancer type prevalence between TCGA methylation and GH liquid STAT5A promoter methylation.

[0045] Figure 21. Lower expression with methylation observed in multiple cancer types. ACC: Adrenocortical Carcinoma. BRCA: Breast, GBM: Glioblastoma, STAD: Gastric, KIRP: Renal Papillary, KIRC: Renal Clear Cell. 18 cancer types were analyzed.

[0046] Figure 22. Current STAT5A overlaps 1 of 2 probes (2nd only 4 base away)

[0047] Figure 23. STAT5A hyper methylated heatmap in 263 liquid samples.

[0048] Figure 24. Concordant prevalence with TCGA cancer types and liquid.STAT5A hypermethylated samples were significantly enriched in the immune-cool group both in HNSCC and LSCC.

[0049] Figure 25. Concordant prevalence in tissue.

[0050] Figure 26. MSI enhanced with MLH1 promoter methylation

[0051] Figure 27. Logistic reg prediction relies on PM and MSI score equally. Insilico dilution around LoD to simulate borderline samples.

[0052] Figure 28. Improved LoD observed from PM in insilico dilutions.Attorney Docket No.: GH0259WO

[0053] Figure 29. Improved MSI Sensitivity: MLH1 Promoter rescued low TF MSI- H. Integrating PM signal using new monotonic gated classifier model. In our clinical data, 2 / 5693 (0.035%) were rescued with 2% and 40% tumor fraction.

[0054] No samples converted from MSI-H to MSS / MSI-L.

[0055] Figure 30. Sample Prediction

[0056] Figure 31. Cross validation for parameter optimization. The inventors applied a gating window width, with empirical window from +2sd from median of MSS and MSI-H = 0.73. Thereafter, iterations were perofrmd with iterations over [0.3, 0.4, 0.5, 0.65, 0.75, 0.9], iterations over [0.1, 0.2, 0.05, 0.01], Monotonicity constraint on features included (1,0,1): MSI and the gate monotonical increases above cutoff.

[0057] Figure 32. All samples predictionDETAILED DESCRIPTION

[0058] Here, the Inventors have established a hereto unachieved detection method for interrogating promoter methylation (PM) status, and quantification of methylation without or without epigenetic allelic status which can be combined with detection of genomic alterations. Here, detection of PM and the use of multimodal genomic alterations in various cancer using a combined epigenomic genomic detection platform including methyl binding domain partitioning, allows a liquid biopsy assay interrogating 800+ genes and genome-wide methylation detection. Pre-defined promoter regions of each covered gene can be analyzed, including new approaches for promoter definition. For each sample, methylation scores were calculated for each gene and used as the basis for making PM calls. A limit of detection (LoD) was determined through in silico and experimental titrations of ctDNA from clinical samples and cell lines with known gene PM into the plasma of cancer-free donors.

[0059] Additionally, establishing the aforementioned detection approach, allows epigenetic allelic status determination at a systematic-level. Allele-specific methylation patterns play an important role in controlling gene expression and maintaining normal cellular functions, and disruptions in these patterns can contribute to pathogenesis including oncogenesis. Imprinting is a form of allele-specific methylation pattern in which one allele of a gene is methylated and silenced depending on whether it is inherited from the mother or the father. The differential methylation and resulting monoallelic expression of imprinted genes are important for normal development and physiological functions and abnormal changes in these imprinting patterns (either loss or gain of methylation), can lead to developmental disorders and increased susceptibility to diseases, including cancer. For example, loss ofAttorney Docket No.: GH0259WO imprinting (LOI) can lead to the expression of both alleles of a gene that is normally imprinted, potentially doubling the expression of genes that promote cell growth, a common feature in various cancers, see for example Figure 10 panel A. In additional cases, tumor suppressor genes that are typically unmethylated and active can become methylated on one allele. This methylation can silence the gene’s expression from that allele, contributing to cancer progression if the other allele is lost or mutated. A well-known example is the pl6 gene (CDKN2A), which can undergo hypermethylation in various cancers such as melanoma, bladder cancer, and others, see for example Figure 10 panel B. Partial allele-specific methylation patterns (see Figure 10 panel C) may impact gene function more subtly compared to the complete methylation of an entire allele. This selective methylation can occur in specific regions of a gene, such as promoters, enhancers, or other regulatory elements, influencing the transcriptional activity of that gene in a cell-type specific manner. In cancer, partial methylation of promoter regions of tumor suppressor genes can downregulate gene expression without completely silencing the gene. This partial methylation might occur in only certain CpG islands within the promoter region.Additionally, methylation of enhancer regions can modulate the activity of enhancers, thus indirectly influencing the expression of genes associated with these enhancers. Partial methylation of enhancer regions can result in altered gene expression profiles that contribute to oncogenesis.

[0060] Current approaches are to omit testing both genomic and epigenomic attributes of the patient sample or to perform multiple tests separately. Omitting genomic or epigenomic information can result in prescription of cancer therapies that could be known to be ineffective or withholding cancer therapies that could be known to be effective, had both genomic and epigenomic information been available. For instance, patients with the KRASG12C biomarker are prescribed KRAS inhibitors but if epigenomic information showing the KRAS promoter was methylated and thus the gene silenced it would be apparent that KRAS inhibitors will not be effective. On the other hand, patients with no detected BRCA1 mutations may not be prescribed PARP inhibitors but if epigenomic information showing the BRCA1 promoter was methylated and thus the gene silenced the patient would be a good candidate for PARP inhibitors. Multiple tests are often not performed due to a variety of reasons including lack of sufficient patient samples. Other drawbacks include lack of reimbursement, inconvenience and lack of available commercial offerings etc.

[0061] Cancer can be indicated by epigenetic variations, such as methylation. Examples of methylation changes in cancer include local gains of DNA methylation in theAttorney Docket No.: GH0259WOCpG islands at the transcription start site (TSS) of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation can be associated with an aberrant loss of transcriptional capacity of involved genes and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression. DNA methylation profiling can be used to detect regions with different extents of methylation (“differentially methylated regions” or “DMRs”) of the genome that are altered during development or that are perturbed by disease, for example, cancer or any cancer- associated disease. The genome of cancer cells harbor imbalance in the above DNA methylation patterns, and therefore in functional packaging of the DNA. The abnormalities of chromatin organization are therefore coupled with methylation changes and may contribute to enhanced cancer profiling when analyzed jointly. Combining MBD-partitioning with fragmentomic data, such as fragment mapped starts and stops positions (correlated with nucleosome positions) , fragment length and associated nucleosome occupancy, can be used for chromatin structure analysis in hypermethylation studies with the aim to improve biomarker detection rate.

[0062] Methylation profiling can involve determining methylation patterns across different regions of the genome. For example, after partitioning molecules based on extent of methylation (e.g., relative number of methylated sites per molecule) and sequencing, the sequences of molecules in the different partitions can be mapped to a reference genome. This can show regions of the genome that, compared with other regions, are more highly methylated or are less highly methylated. In this way, genomic regions, in contrast to individual molecules, may differ in their extent of methylation.

[0063] A characteristic of nucleic acid molecules may be a modification, which may include various chemical or protein modifications (i.e. epigenetic modifications). Nonlimiting examples of chemical modification may include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation includes addition of a methyl group to a cytosine at a CpG site (a cytosine followed by a guanine in a nucleic acid sequence). In some embodiments, DNA methylation includes addition of a methyl group to adenine, such as in N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of the 6 carbon ring of cytosine). In some embodiments, 5-methylation includes addition of a methyl group to the 5C position of the cytosine to create 5-methylcytosine (m5c). In some embodiments, methylation includes a derivative of m5c. Derivatives of m5c include, but are not limited to, 5- hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC).Attorney Docket No.: GH0259WOIn some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of the 6 carbon ring of cytosine). In some embodiments, 3C methylation includes addition of a methyl group to the 3C position of the cytosine to generate 3 -methylcytosine (3mC). Other examples include N6-methyladenine or glycosylation. DNA methylation includes addition of methyl groups to DNA (e.g. CpG) and can change the expression of methylated DNA region.. Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.

[0064] A CpG dyad is the dinucleotide CpG (cytosine-phosphate-guanine, i.e. a cytosine followed by a guanine in a 5’ -> 3’ direction of the nucleic acid sequence) on the sense strand and its complementary CpG on the antisense strand of a double-stranded DNA molecule. CpG dyads can be either fully methylated or hemi-methylated (methylated on one strand only).

[0065] The CpG dinucleotide is underrepresented in the normal human genome, with the majority of CpG dinucleotide sequences being transcriptionally inert (e.g. DNA heterochromatic regions in pericentromeric parts of the chromosome and in repeat elements) and methylated. However, many CpG islands are protected from such methylation especially around transcription start sites (TSS).

[0066] Protein modifications include binding to components of chromatin, particularly histones including modified forms thereof, and binding to other proteins, such as proteins involved in replication or transcription. The disclosure provides methods of processing and analyzing nucleic acids with different extents of modification, such that the nature of their original modification is correlated with a nucleic acid tag and can be decoded by sequencing the tag when nucleic acids are analyzed. Genetic variation of sample nucleic acid modifications can then be associated with the extent of modification (epigenetic variation) of that nucleic acid in the original sample, include single stranded (e.g., ssDNA or RNA) or double stranded molecules (e.g., dsDNA).

[0067] The loss of DNA can reduce the presence of one or more types of DNA such that the presence of the one or more types of DNA such as cfDNA, is difficult to detect. In one or more additional scenarios, existing methods to measure DNA methylation, such asAttorney Docket No.: GH0259WO enrichment or depletion methods, can have a relatively high level of resolution, such as about 100 base pairs (bp) to about 200 bp that can make accurately determining an amount of methylation of DNA difficult. The accuracy with which DNA methylation is determined can impact the accuracy of estimates of tumor fraction for samples. Since tumor fraction can be used to determine whether a sample is derived from a subject in which a tumor is present or not, the accuracy of determinations of tumor fraction estimates can impact diagnosis and / or treatment decisions for individuals.

[0068] More specifically, the techniques described herein allow quantification of promoter region methylation. Gedatolisib is an intravenously administered PI3K and mTOR inhibitor which has been shown to be safe in patients with metastatic breast cancer, either alone or in combination with oral therapies. Previous research has shown that PI3K inhibitors lower nucleotide pools required for DNA synthesis and S-phase progression. Additionally, inhibition of PI3K / mT0R could impede PI3K interaction with the homologous recombination complex, increasing dependency on PARP enzymes for DNA repair. Based on this data, the combination of a PI3K inhibitor and PARP inhibitor could potentially lead to a new, non-chemotherapy treatment option for TNBC with wild-type BRCA and improve the modest PFS seen with the PARP inhibitors as single agents in BRCA1 / 2 mutant advanced breast cancer. The hypothesis for this trial is that the gedatolisib will sensitize advanced TNBC or BRCA1 / 2 mutant breast cancers to PARP inhibition with talazoparib. Of interesting is determining the recommended phase 2 dose of gedatolisib in combination with talazoparib and to evaluate the efficacy of this combination in advanced HER2 negative breast cancer that is triple negative or BRCA1 / 2 positive (mutated / deficient).Samples

[0069] A sample can be any biological sample isolated from a subject. A sample can be a bodily sample. Samples can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusions, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine. Samples are preferably body fluids, particularly blood and fractions thereof, and urine. A sample can be in the form originally isolated from a subject or can have been subjected to further processing to remove or add components, such as cells, or enrich for one component relative to another. Thus, a preferredAtorney Docket No.: GH0259WO body fluid for analysis is plasma or serum containing cell-free nucleic acids. A sample can be isolated or obtained from a subject and transported to a site of sample analysis. The sample may be preserved and shipped at a desirable temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. A sample can be isolated or obtained from a subject at the site of the sample analysis. The subject can be a human, a mammal, an animal, a companion animal, a service animal, or a pet. The subject may have a cancer. The subject may not have cancer or a detectable cancer symptom. The subject may have been treated with one or more cancer therapy, e.g., any one or more of chemotherapies, antibodies, vaccines or biologies. The subject may be in remission. The subject may or may not be diagnosed of being susceptible to cancer or any cancer-associated genetic mutations / disorders.

[0070] The volume of plasma can depend on the desired read depth for sequenced regions. Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For examples, the volume can be 0.5 mL, 1 mL, 5 mL 10 mL, 20 mL, 30 mL, or 40 mL. A volume of sampled plasma may be 5 to 20 mL.

[0071] A sample can comprise various amount of nucleic acid that contains genome equivalents. For example, a sample of about 30 ng DNA can contain about 10,000 (104) haploid human genome equivalents and, in the case of cfDNA, about 200 billion (2x1011) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents and, in the case of cfDNA, about 600 billion individual molecules.

[0072] A sample can comprise nucleic acids from different sources, e.g., from cells and cell-free of the same subject, from cells and cell-free of different subjects. A sample can comprise nucleic acids carrying mutations. For example, a sample can comprise DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations existing in germline DNA of a subject. Somatic mutations refer to mutations originating in somatic cells of a subject, e.g., cancer cells. A sample can comprise DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations). A sample can comprise an epigenetic variant (i.e. a chemical or protein modification), wherein the epigenetic variant associated with the presence of a genetic variant such as a cancer- associated mutation. In some embodiments, the sample includes an epigenetic variant associated with the presence of a genetic variant, wherein the sample does not comprise the genetic variant.

[0073] Exemplary amounts of cell-free nucleic acids in a sample before amplification range from about 1 fg to about 1 pg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng.Attorney Docket No.: GH0259WOFor example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can comprise obtaining 1 femtogram (fg) to 200 ng.

[0074] Cell-free nucleic acids are nucleic acids not contained within or otherwise bound to a cell or in other words nucleic acids remaining in a sample after removing intact cells. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or a hybrid thereof. A cell-free nucleic acid can be released into bodily fluid through secretion or cell death processes, e.g., cellular necrosis and apoptosis. Some cell-free nucleic acids are released into bodily fluid from cancer cells e.g., circulating tumor DNA, (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA) In some embodiments, cell free nucleic acids are produced by tumor cells. In some embodiments, cell free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.

[0075] Cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of molecules, with a mode of about 168 nucleotides and a second minor peak in a range between 240 to 440 nucleotides. Cell-free nucleic acids can be isolated from bodily fluids through a fractionation or partitioning step in which cell-free nucleic acids, as found in solution, are separated from intact cells and other non-soluble components of the bodily fluid. Partitioning may include techniques such as centrifugation or filtration. Alternatively, cells in bodily fluids can be lysed and cell-free and cellular nucleic acids processed together. Generally, after addition of buffers and wash steps, nucleic acids can be precipitated with an alcohol. Further clean up steps may be used such as silica based columns to remove contaminants or salts. Non-specific bulk carrier nucleic acids, such as Cot-1 DNA, DNA or protein for bisulfite sequencing, hybridization, and / or ligation, may be added throughout the reaction to optimize certain aspects of the procedure such as yield.Attorney Docket No.: GH0259WO

[0076] After such processing, samples can include various forms of nucleic acid including double stranded DNA, single stranded DNA and single stranded RNA. In some embodiments, single stranded DNA and RNA can be converted to double stranded forms so they are included in subsequent processing and analysis steps.Analytes

[0077] Analytes can include nucleic acid analytes, and non-nucleic acid analytes. The disclosure provides for detecting genetic variations in biological samples from a subject. Biological samples may include polynucleotides from cancer cells. Polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNAs), or any combination thereof. Biological samples may include tumor tissue, e.g., from a biopsy. In some cases, biological samples may include blood or saliva. In particular cases, biological samples may comprise cell free DNA (“cfDNA”) or circulating tumor DNA (“ctDNA”). Cell free DNA can be present in, e.g., blood.

[0078] Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphoproteins, specific phosphorylated or acetylated variants of proteins, amidation variants of proteins, hydroxylation variants of proteins, methylation variants of proteins, ubiquity lati on variants of proteins, sulfation variants of proteins, viral proteins (e.g., viral capsid, viral envelope, viral coat, viral accessory, viral glycoproteins, viral spike, etc.), extracellular and intracellular proteins, antibodies, and antigen binding fragments. This further includes receptor, an antigen, a surface protein, a transmembrane protein, a cluster of differentiation protein, a protein channel, a protein pump, a carrier protein, a phospholipid, a glycoprotein, a glycolipid, a cell-cell interaction protein complex, an antigen-presenting complex, a major histocompatibility complex, an engineered T-cell receptor, a T-cell receptor, a B-cell receptor, a chimeric antigen receptor, an extracellular matrix protein, a posttranslational modification (e.g., phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation or lipidation) state of a cell surface protein, a gap junction, and an adherens junction.

[0079] In general, the systems, apparatus, methods, and compositions can be used to analyze any number of analytes, further including both nucleic acid analytes and non-nucleic acid analytes. For example, the number of analytes that are analyzed can be at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least aboutAttorney Docket No.: GH0259WO13, at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, at least about 1,000, at least about 10,000, at least about 100,000 or more different analytes present in a region of the sample or within an individual feature of the substrate. Methods for performing multiplexed assays to analyze two or more different analytes will be discussed in a subsequent section of this disclosure.

[0080] One or more nucleic acid analytes and / or non-nucleic acid analytes constitute a set of molecular interactions in a biological system under study (e.g., cells), which may be regarded as “interactome” - the molecular interactions that occur between molecules belonging to different biochemical families (proteins, nucleic acids, lipids, carbohydrates, etc.) and also within a given family. In various embodiments, an interactome is a protein- DNA interactome (network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes. In other embodiments, interactome refers to protein-protein interaction network(PPI), or protein interaction network (PIN). The methods described herein allow for study and analysis of the interactome. Techniques such as proteogenomics (whole genome sequencing, whole exome sequencing and RNA-seq, and mass spectrometry as examples) can support study of the interactome.Analysis

[0081] The present methods can be used to diagnose presence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.

[0082] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, solid stateAttorney Docket No.: GH0259WO tumors, heterogeneous tumors, homogenous tumors and the like. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0083] Genetic and other analyte data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.

[0084] The present analyses are also useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.

[0085] The present methods can also be used for detecting genetic variations in conditions other than cancer. Immune cells, such as B cells, may undergo rapid clonal expansion upon the presence certain diseases. Clonal expansions may be monitored using copy number variation detection and certain immune states may be monitored. In this example, copy number variation analysis may be performed over time to produce a profile of how a particular disease may be progressing. Copy number variation or even rare mutation detection may be used to determine how a population of pathogens is changing during the course of infection. This may be particularly important during chronic infections, such as HIV / AIDS or Hepatitis infections, whereby viruses may change life cycle state and / or mutateAttorney Docket No.: GH0259WO into more virulent forms during the course of infection. The present methods may be used to determine or profile rejection activities of the host body, as immune cells attempt to destroy transplanted tissue to monitor the status of transplanted tissue as well as altering the course of treatment or prevention of rejection.

[0086] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.

[0087] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation and mutation analyses alone or in combination.

[0088] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0089] Determination of 5-methylcytosine pattern of nucleic acids

[0090] Bisulfite-based sequencing and variants thereof provides a means of determining the methylation pattern of a nucleic acid. In some embodiments, determining the methylation pattern comprises distinguishing 5-methylcytosine (5mC) from non-methylated cytosine. In some embodiments, determining methylation pattern comprises distinguishing N6-methyladenine from non-methylated adenine. In some embodiments, determining the methylation pattern comprises distinguishing 5-hydroxymethylcytosine (5hmC), 5- formylcytosine (5fC), and 5-carboxylcytosine (5caC) from non-methylated cytosine. Examples of bisulfite sequencing include, but are not limited to oxidative bisulfiteAttorney Docket No.: GH0259WO sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reduced bisulfite sequencing (redBS-seq).

[0091] Oxidative bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC, by first converting the 5hmC to 5fC, and then proceeding with bisulfite sequencing as previously described. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish 5mc and 5hmC. In TAB-seq, 5hmC is protected by glucosylation. A Tet enzyme is then used to convert 5mC to 5caC before proceeding with bisulfite sequencing, as previously described. Reduced bisulfite sequencing is used to distinguish 5fC from modified cytosines.

[0092] Generally, in bisulfite sequencing, a nucleic acid sample is divided into two aliquots and one aliquot is treated with bisulfite. The bisulfite converts native cytosine and certain modified cytosine nucleotides (e.g. 5-formylcytosine or 5-carboxylcytosine) to uracil whereas other modified cytosines (e.g., 5- methylcytosine, 5-hydroxylmethylcystosine) are not converted. Comparison of nucleic acid sequences of molecules from the two aliquots indicates which cytosines were and were not converted to uracils. Consequently, cytosines which were and were not modified can be determined. The initial splitting of the sample into two aliquots is disadvantageous for samples containing only small amounts of nucleic acids, and / or composed of heterogeneous cell / tissue origins such as bodily fluids containing cell- free DNA.

[0093] The present disclosure provides methods allowing bisulfite sequencing and variants thereof. These methods work by linking nucleic acids in a population to a capture moiety, i.e., a label that can be captured or immobilized. Capture moieties include, without limitation, biotin, avidin, streptavidin, a nucleic acid comprising a particular nucleotide sequence, a hapten recognized by an antibody, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety that is attached to an analyte is captured by its binding pair which is attached to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented through centrifugation. The capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin which allows affinity separation by binding to streptavidin linked or linkable to a solid phase or an oligonucleotide, which allows affinity separation through binding to a complementary oligonucleotide linked or linkable to a solid phase. Following linking of capture moieties to sample nucleic acids, the sample nucleic acids serve asAttorney Docket No.: GH0259WO templates for amplification. Following amplification, the original templates remain linked to the capture moieties but amplicons are not linked to capture moieties.

[0094] The capture moiety can be linked to sample nucleic acids as a component of an adapter, which may also provide amplification and / or sequencing primer binding sites. In some methods, sample nucleic acids are linked to adapters at both ends, with both adapters bearing a capture moiety. Preferably any cytosine residues in the adapters are modified, such as by 5methylcytosine, to protect against the action of bisulfite. In some instances, the capture moieties are linked to the original templates by a cleavable linkage (e.g., photocleavable desthiobiotin-TEG or uracil residues cleavable with USER™ enzyme, Chem. Commun. (Camb). 2015 Feb 21; 51(15): 3266-3269), in which case the capture moieties can, if desired, be removed.

[0095] The amplicons are denatured and contacted with an affinity reagent for the capture tag. Original templates bind to the affinity reagent whereas nucleic acid molecules resulting from amplification do not. Thus, the original templates can be separated from nucleic acid molecules resulting from amplification.

[0096] Following separation or partition, the respective populations of nucleic acids (i.e., original templates and amplification products) can be subjected to bisulfite treatment with the original template population receiving bisulfite treatment and the amplification products not. Alternatively, the amplification products can be subjected to bisulfite treatment and the original template population not. Following such treatment, the respective populations can be amplified (which in the case of the original template population converts uracils to thymines). The populations can also be subjected to biotin probe hybridization for enrichment. The respective populations are then analyzed and sequences compared to determine which cytosines were 5-methylated (or 5-hydroxylmethylated) in the original. Detection of a T nucleotide in the template population (corresponding to an unmethylated cytosine converted to uracil) and a C nucleotide at the corresponding position of the amplified population indicates an unmodified C. The presence of C's at corresponding positions of the original template and amplified populations indicates a modified C in the original sample.

[0097] In some embodiments, a method uses sequential DNA-seq and bisulfite-seq (BlS-seq) NGS library preparation of molecular tagged DNA libraries. This process is performed by labeling of adapters (e.g., biotin), DNA-seq amplification of whole library, parent molecule recovery (e.g. streptavidin bead pull down), bisulfite conversion and BIS- seq. In some embodiments, the method identifies 5-methylcytosine with single-baseAttorney Docket No.: GH0259WO resolution, through sequential NGS-preparative amplification of parent library molecules with and without bisulfite treatment. This can be achieved by modifying the 5-methyl-ated NGS-adapters (directional adapters; Y-shaped / forked with 5-methylcytosine replacing) used in BlS-seq with a label (e.g., biotin) on one of the two adapter strands. Sample DNA molecules are adapter ligated, and amplified (e.g., by PCR). As only the parent molecules will have a labeled adapter end, they can be selectively recovered from their amplified progeny by label-specific capture methods (e.g., streptavidin-magnetic beads). As the parent molecules retain 5-methylation marks, bisulfite conversion on the captured library will yield single-base resolution 5-methylation status upon BlS-seq, retaining molecular information to corresponding DNA-seq. In some embodiments, the bisulfite treated library can be combined with a non-treated library prior to enrichment / NGS by addition of a sample tag DNA sequence in standard multiplexed NGS workflow. As with BlS-seq workflows, bioinformatics analysis can be carried out for genomic alignment and 5-methylated base identification. In sum, this method provides the ability to selectively recover the parent, ligated molecules, carrying 5-methylcytosine marks, after library amplification, thereby allowing for parallel processing for bisulfite converted DNA. This overcomes the destructive nature of bisulfite treatment on the quality / sensitivity of the DNA-seq information extracted from a workflow. With this method, the recovered ligated, parent DNA molecules (via labeled adapters) allow amplification of the complete DNA library and parallel application of treatments that elicit epigenetic DNA modifications. The present disclosure discusses the use of BlS-seq methods to identify cytosine5-methylation (5-methylcytosine), but this should is not limiting. Variants of BlS-seq have been developed to identify hydroxymethylated cytosines (5hmC; OX- BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq) and carboxylcytosines. These methodologies can be implemented with the sequential / parallel library preparation described herein.Alternative Methods of Modified Nucleic Acid Analysis

[0098] The disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, linked to histones and other modifications discussed above). In some such methods, a population of nucleic acids bearing the modification to different extents (e.g., 0, 1, 2, 3, 4, 5 or more methyl groups per nucleic acid molecule) is contacted with adapters before fractionation of the population depending on the extent of the modification. Adapters attach to either one end or both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number ofAttorney Docket No.: GH0259WO combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Following attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites within the adapters. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites, but preferably adapters include the same primer binding site. Following amplification, the nucleic acids are contacted with an agent that preferably binds to nucleic acids bearing the modification (such as the previously described such agents). The nucleic acids are separated into at least two partitions differing in the extent to which the nucleic acids bear the modification from binding to the agents. For example, if the agent has affinity for nucleic acids bearing the modification, nucleic acids overrepresented in the modification (compared with median representation in the population) preferentially bind to the agent, whereas nucleic acids underrepresented for the modification do not bind or are more easily eluted from the agent. Following separation, the different partitions can then be subject to further processing steps, which typically include further amplification, and sequence analysis, in parallel but separately. Sequence data from the different partitions can then be compared.

[0099] Nucleic acids can be linked at both ends to Y-shaped adapters including primer binding sites and tags. The molecules are amplified. The amplified molecules are then fractionated by contact with an antibody preferentially binding to 5-methylcytosine to produce two partitions. One partition includes original molecules lacking methylation and amplification copies having lost methylation. The other partition includes original DNA molecules with methylation. The two partitions are then processed and sequenced separately with further amplification of the methylated partition. The sequence data of the two partitions can then be compared. In this example, tags are not used to distinguish between methylated and unmethylated DNA but rather to distinguish between different molecules within these partitions so that one can determine whether reads with the same start and stop points are based on the same or different molecules.

[0100] The disclosure provides further methods for analyzing a population of nucleic acid in which at least some of the nucleic acids include one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications described previously. In these methods, the population of nucleic acids is contacted with adapters including one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably all cytosine residues in such adapters are also modified, or all such cytosines in a primer binding region of the adapters are modified. Adapters attach to both ends of nucleic acid moleculesAttorney Docket No.: GH0259WO in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites of the adapters. The amplified nucleic acids are split into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data on molecules in the first aliquot is thus determined irrespective of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosines to uracils. The bisulfite treated nucleic acids are then subjected to amplification primed by primers to the original primer binding sites of the adapters linked to nucleic acid. Only the nucleic acid molecules originally linked to adapters (as distinct from amplification products thereof) are now amplifiable because these nucleic acids retain cytosines in the primer binding sites of the adapters, whereas amplification products have lost the methylation of these cytosine residues, which have undergone conversion to uracils in the bisulfite treatment. Thus, only original molecules in the populations, at least some of which are methylated, undergo amplification. After amplification, these nucleic acids are subject to sequence analysis. Comparison of sequences determined from the first and second aliquots can indicate among other things, which cytosines in the nucleic acid population were subject to methylation.Partitioning the Sample into a Plurality of Subsamples; Aspects of Samples; Analysis of Epigenetic Characteristics

[0101] In certain embodiments described herein, a population of different forms of nucleic acids (e.g., hypermethylated and hypom ethylated DNA in a sample, such as a captured set of cfDNA as described herein) can be physically partitioned based on one or more characteristics of the nucleic acids prior to further analysis, e.g., differentially modifying or isolating a nucleobase, tagging, and / or sequencing. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, hypermethylation variable epigenetic target regions are analyzed to determine whether they show hypermethylation characteristic of tumor cells and / or hypomethylation variable epigenetic target regions are analyzed to determine whether they show hypomethylation characteristic of tumor cells. Additionally, by partitioning aAttorney Docket No.: GH0259WO heterogeneous nucleic acid population, one may increase rare signals, e.g., by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, a genetic variation present in hyper-methylated DNA but less (or not) in hypomethylated DNA can be more easily detected by partitioning a sample into hypermethylated and hypo-methylated nucleic acid molecules. By analyzing multiple fractions of a sample, a multi-dimensional analysis of a single locus of a genome or species of nucleic acid can be performed and hence, greater sensitivity can be achieved.

[0102] In some instances, a heterogeneous nucleic acid sample is partitioned into two or more partitions (e.g., at least 3, 4, 5, 6 or 7 partitions). In some embodiments, each partition is differentially tagged. Tagged partitions can then be pooled together for collective sample prep and / or sequencing. The partitioning-tagging-pooling steps can occur more than once, with each round of partitioning occurring based on a different characteristics (examples provided herein), and tagged using differential tags that are distinguished from other partitions and partitioning means.

[0103] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. Resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments and longer DNA fragments. In some embodiments, partitioning based on a cytosine modification (e.g., cytosine methylation) or methylation generally is performed and is optionally combined with at least one additional partitioning step, which may be based on any of the foregoing characteristics or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include presence or absence of methylation; level of methylation; type of methylation (e.g., 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins, such as histones. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules devoid of nucleosomes. Alternatively or additionally, a heterogeneous population of nucleic acids may be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or additionally, a heterogeneous population of nucleic acids may be partitionedAttomey Docket No.: GH0259WO based on nucleic acid length (e.g., molecules of up to 160 bp and molecules having a length of greater than 160 bp).

[0104] In some instances, each partition (representative of a different nucleic acid form) is differentially labelled, and the partitions are pooled together prior to sequencing. In other instances, the different forms are separately sequenced. In some embodiments, a population of different nucleic acids is partitioned into two or more different partitions. Each partition is representative of a different nucleic acid form, and a first partition (also referred to as a subsample) comprises DNA with a cytosine modification in a greater proportion than a second subsample. Each partition is distinctly tagged. The first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The tagged nucleic acids are pooled together prior to sequencing. Sequence reads are obtained and analyzed, including to distinguish the first nucleobase from the second nucleobase in the DNA of the first subsample, in silico. Tags are used to sort reads from different partitions. Analysis to detect genetic variants can be performed on a partition- by-partition level, as well as whole nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants, such as CNV, SNV, indel, fusion in nucleic acids in each partition. In some instances, in silico analysis can include determining chromatin structure. For example, coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or nucleosome depleted region (NDR).

[0105] Samples can include nucleic acids varying in modifications including postreplication modifications to nucleotides and binding, usually noncovalently, to one or more proteins.

[0106] In an embodiment, the population of nucleic acids is one obtained from a serum, plasma or blood sample from a subject suspected of having neoplasia, a tumor, or cancer or previously diagnosed with neoplasia, a tumor, or cancer. The population of nucleic acids includes nucleic acids having varying levels of methylation. Methylation can occur from any one or more post-replication or transcriptional modifications. Post-replication modifications include modifications of the nucleotide cytosine, particularly at the 5-position of the nucleobase, e.g., 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-Attorney Docket No.: GH0259WO carboxylcytosine. The affinity agents can be antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected e.g., by phage display to have specificity to a given target.

[0107] Examples of capture moieties contemplated herein include methyl binding domain (MBDs) and methyl binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies preferentially binding to 5-methylcytosine. Likewise, partitioning of different forms of nucleic acids can be performed using histone binding proteins which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides. Although for some affinity agents and modifications, binding to the agent may occur in an essentially all or none manner depending on whether a nucleic acid bears a modification, the separation may be one of degree. In such instances, nucleic acids overrepresented in a modification bind to the agent at a greater extent that nucleic acids underrepresented in the modification. Alternatively, nucleic acids having modifications may bind in an all or nothing manner. But then, various levels of modifications may be sequentially eluted from the binding agent.

[0108] For example, in some embodiments, partitioning can be binary or based on degree / level of modifications. For example, all methylated fragments can be partitioned from unmethylated fragments using methyl-binding domain proteins (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning may involve eluting fragments having different levels of methylation by adjusting the salt concentration in a solution with the methyl-binding domain and bound fragments. As salt concentration increases, fragments having greater methylation levels are eluted. In some instances, the final partitions are representative of nucleic acids having different extents of modifications (overrepresentative or underrepresentative of modifications). Overrepresentation and underrepresentation can be defined by the number of modifications bom by a nucleic acid relative to the median number of modifications per strand in a population. For example, if the median number of 5-methylcytosine residues in nucleic acid in a sample is 2, a nucleic acid including more than two 5-methylcytosine residues is overrepresented in this modification and a nucleic acid with 1 or zero 5- methylcytosine residues is underrepresented. The effect of the affinity separation is to enrich for nucleic acids overrepresented in a modification in a bound phase and for nucleic acidsAttorney Docket No.: GH0259WO underrepresented in a modification in an unbound phase (i.e. in solution). The nucleic acids in the bound phase can be eluted before subsequent processing.

[0109] When using MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific) various levels of methylation can be partitioned using sequential elutions. For example, a hypom ethylated partition (e.g., no methylation) can be separated from a methylated partition by contacting the nucleic acid population with the MBD from the kit, which is attached to magnetic beads. The beads are used to separate out the methylated nucleic acids from the non- methylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids having different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is once again used to separate higher level of methylated nucleic acids from those with lower level of methylation. The elution and magnetic separation steps can repeat themselves to create various partitions such as a hypomethylated partition (representative of no methylation), a methylated partition (representative of low level of methylation), and a hyper methylated partition (representative of high level of methylation).

[0110] In some methods, nucleic acids bound to an agent used for affinity separation are subjected to a wash step. The wash step washes off nucleic acids weakly bound to the affinity agent. Such nucleic acids can be enriched in nucleic acids having the modification to an extent close to the mean or median (i.e., intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not binding to the solid phase on initial contacting of the sample with the agent). The affinity separation results in at least two, and sometimes three or more partitions of nucleic acids with different extents of a modification. While the partitions are still separate, the nucleic acids of at least one partition, and usually two or three (or more) partitions are linked to nucleic acid tags, usually provided as components of adapters, with the nucleic acids in different partitions receiving different tags that distinguish members of one partition from another. The tags linked to nucleic acid molecules of the same partition can be the same or different from one another. But if different from one another, the tags may have part of their code in common so as to identify the molecules to which they are attached as being of a particular partition. For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some embodiments, the nucleic acid molecules can beAttorney Docket No.: GH0259WO fractionated into different partitions based on the nucleic acid molecules that are bound to a specific protein or a fragment thereof and those that are not bound to that specific protein or fragment thereof.[oni] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on a specific property of a protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation) or enzymatic activity. Examples of proteins which may bind to DNA and serve as a basis for fractionation may include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate the nucleic acid molecules based on protein bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein bound regions include, but are not limited to, SDS-PAGE, chromatin- immuno-precipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).

[0112] In some embodiments, partitioning of the nucleic acids is performed by contacting the nucleic acids with a methylation binding domain (“MBD”) of a methylation binding protein (“MBP”). MBD binds to 5 -methylcytosine (5mC). MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by eluting fractions by increasing the NaCl concentration.

[0113] An exemplary method for molecular tag identification of MBD-bead partitioned libraries through NGS is as follows:

[0114] Physical partitioning of an extracted DNA sample (e.g., extracted blood plasma DNA from a human sample) using a methyl-binding domain protein-bead purification kit, saving all elutions from process for downstream processing.

[0115] Parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, the hypermethylated, residual methylation ('wash'), and hypomethylated partitions are ligated with NGS-adapters with molecular tags.

[0116] Re-combining all molecular tagged partitions, and subsequent amplification using adapter-specific DNA primer sequences.

[0117] Enrichment / hybridization of re-combined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).

[0118] Re-amplification of the enriched total DNA library, appending a sample tag. Different samples are pooled, and assayed in multiplex on an NGS instrument.Attomey Docket No.: GH0259WO

[0119] Bioinformatics analysis of NGS data, with the molecular tags being used to identify unique molecules, as well deconvolution of the sample into molecules that were differentially MBD-partitioned. This analysis can yield information on relative 5- methylcytosine for genomic regions, concurrent with standard genetic sequencing / variant detection.

[0120] Examples of MBPs contemplated herein include, but are not limited to:

[0121] (a) MeCP2 is a protein preferentially binding to 5-methyl-cytosine over unmodified cytosine.

[0122] (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind to 5- hydroxymethyl-cytosine over unmodified cytosine.

[0123] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind to 5- formyl-cytosine over unmodified cytosine (lurlaro et al., Genome Biol. 14: R119 (2013)).

[0124] (d) Antibodies specific to one or more methylated nucleotide bases.

[0125] In general, elution is a function of number of methylated sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCl concentration. Salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and comprising a molecule comprising a methyl binding domain, which molecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration a population of molecules will bind to the MBD and a population will remain unbound. The unbound population can be separated as a “hypom ethylated” population. For example, a first partition representative of the hypomethylated form of DNA is that which remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. A second partition representative of intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. This is also separated from the sample. A third partition representative of hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0126] The disclosure provides further methods for analyzing a population of nucleic acids in which at least some of the nucleic acids include one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications described previously. In these methods, after partitioning, the subsamples of nucleic acids are contacted with adapters including one or more cytosine residues modified at the 5C position, such as 5-Attorney Docket No.: GH0259WO methylcytosine. Preferably all cytosine residues in such adapters are also modified, or all such cytosines in a primer binding region of the adapters are modified. Adapters attach to both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites of the adapters. The amplified nucleic acids are split into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data on molecules in the first aliquot is thus determined irrespective of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase comprises a cytosine modified at the 5 position, and the second nucleobase comprises unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. The nucleic acids subjected to the procedure are then amplified with primers to the original primer binding sites of the adapters linked to nucleic acid. Only the nucleic acid molecules originally linked to adapters (as distinct from amplification products thereof) are now amplifiable because these nucleic acids retain cytosines in the primer binding sites of the adapters, whereas amplification products have lost the methylation of these cytosine residues, which have undergone conversion to uracils in the bisulfite treatment. Thus, only original molecules in the populations, at least some of which are methylated, undergo amplification. After amplification, these nucleic acids are subject to sequence analysis. Comparison of sequences determined from the first and second aliquots can indicate among other things, which cytosines in the nucleic acid population were subject to methylation.

[0127] Such an analysis can be performed using the following exemplary procedure. After partitioning, methylated DNA is linked to Y-shaped adapters at both ends including primer binding sites and tags. The cytosines in the adapters are modified at the 5 position (e.g., 5-methylated). The modification of the adapters serves to protect the primer binding sites in a subsequent conversion step (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but affects unmodified cytosine). After attachment of adapters, the DNA molecules are amplified. The amplification product is split into two aliquots for sequencing with and without conversion. The aliquot not subjected toAttorney Docket No.: GH0259WO conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase comprises a cytosine modified at the 5 position, and the second nucleobase comprises unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. Only primer binding sites protected by modification of cytosines can support amplification when contacted with primers specific for original primer binding sites. Thus, only original molecules and not copies from the first amplification are subjected to further amplification. The further amplified molecules are then subjected to sequence analysis. Sequences can then be compared from the two aliquots. As in the separation scheme discussed above, nucleic acid tags in adapters are not used to distinguish between methylated and unmethylated DNA but to distinguish nucleic acid molecules within the same partition.Subjecting the First Subsample to a Procedure that Affects a First Nucleobase in the DNA Differently from a Second Nucleobase in the DNA of the First Subsample

[0128] Methods disclosed herein comprise a step of subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, if the first nucleobase is a modified or unmodified adenine, then the second nucleobase is a modified or unmodified adenine; if the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine; if the first nucleobase is a modified or unmodified guanine, then the second nucleobase is a modified or unmodified guanine; and if the first nucleobase is a modified or unmodified thymine, then the second nucleobase is a modified or unmodified thymine (where modified and unmodified uracil are encompassed within modified thymine for the purpose of this step).

[0129] In some embodiments, the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine. For example, first nucleobase may comprise unmodified cytosine (C) and the second nucleobase may comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, as indicated, e.g., in the Summary aboveAttorney Docket No.: GH0259WO and the following discussion, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.

[0130] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g. 5-formyl cytosine (fC) or 5-carboxylcytosine (caC)) to uracil whereas other modified cytosines (e.g., 5-methylcytosine, 5-hydroxylmethylcystosine) are not converted. Thus, where bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, 5-formyl cytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of bisulfite-treated DNA identifies positions that are read as cytosine as being mC or hmC positions. Meanwhile, positions that are read as T are identified as being T or a bisulfite-susceptible form of C, such as unmodified cytosine, 5-formyl cytosine, or 5-carboxylcytosine. Performing bisulfite conversion on a first subsample as described herein thus facilitates identifying positions containing mC or hmC using the sequence reads obtained from the first sub sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068..

[0131] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises oxidative bisulfite (Ox-BS) conversion. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises Tet-assisted bisulfite (TAB) conversion. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises Tet-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises chemical-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises APOBEC-coupled epigenetic (ACE) conversion.Attomey Docket No.: GH0259WO

[0132] In some embodiments, procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM- seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692vl. For example, TET2 and T4-PGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosines converting them to uracils.

[0133] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase.

[0134] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-m ethyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).

[0135] Techniques comprising methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mA from other DNA. See, e.g., Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161 : 868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies for various modified nucleobases, such as forms of thymine / uracil including halogenated forms such as 5-bromouracil, are commercially available. Various modified bases can also be detected based on alterations in their base-pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read in sequencing as a G. See, e.g., US Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, “Mutation, Repair, and Recombination.”Enriching / Capturing Step, Amplification., Adaptors, Barcodes

[0136] In some embodiments, methods disclosed herein comprise a step of capturing one or more sets of target regions of DNA, such as cfDNA. Capture may be performed using any suitable approach known in the art. In some embodiments, capturing comprises contacting the DNA to be captured with a set of target-specific probes. The set of target-Attorney Docket No.: GH0259WO specific probes may have any of the features described herein for sets of target-specific probes, including but not limited to in the embodiments set forth above and the sections relating to probes below. Capturing may be performed on one or more subsamples prepared during methods disclosed herein. In some embodiments, DNA is captured from at least the first subsample or the second subsample, e.g., at least the first subsample and the second subsample. Where the first subsample undergoes a separation step (e.g., separating DNA originally comprising the first nucleobase (e.g., hmC) from DNA not originally comprising the first nucleobase, such as hmC-seal), capturing may be performed on any, any two, or all of the DNA originally comprising the first nucleobase (e.g., hmC), the DNA not originally comprising the first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.

[0137] The capturing step may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization. In some embodiments, complexes of target-specific probes and DNA are formed.

[0138] In some embodiments, a method described herein comprises capturing cfDNA obtained from a test subject for a plurality of sets of target regions. The target regions comprise epigenetic target regions, which may show differences in methylation levels and / or fragmentation patterns depending on whether they originated from a tumor or from healthy cells. The target regions also comprise sequence-variable target regions, which may show differences in sequence depending on whether they originated from a tumor or from healthy cells. The capturing step produces a captured set of cfDNA molecules, and the cfDNA molecules corresponding to the sequence-variable target region set are captured at a greater capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the epigenetic target region set. For additional discussion of capturing steps, capture yields, and related aspects, see W02020 / 160414, which is incorporated herein by reference for all purposes.

[0139] In some embodiments, a method described herein comprises contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to the sequence-variableAttorney Docket No.: GH0259WO target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set.

[0140] It can be beneficial to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequencevariable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test fsor perturbation of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated partitions) is generally less than the volume of data needed to determine the presence or absence of cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).

[0141] In various embodiments, the methods further comprise sequencing the captured cfDNA, e.g., to different degrees of sequencing depth for the epigenetic and sequence-variable target region sets, consistent with the discussion herein. In some embodiments, complexes of target-specific probes and DNA are separated from DNA not bound to target-specific probes. For example, where target-specific probes are bound covalently or noncovalently to a solid support, a washing or aspiration step can be used to separate unbound material. Alternatively, where the complexes have chromatographic properties distinct from unbound material (e.g., where the probes comprise a ligand that binds a chromatographic resin), chromatography can be used.

[0142] As discussed in detail elsewhere herein, the set of target-specific probes may comprise a plurality of sets such as probes for a sequence-variable target region set and probes for an epigenetic target region set. In some such embodiments, the capturing step is performed with the probes for the sequence-variable target region set and the probes for the epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence-variable target region set is greater that the concentration of the probes for the epigenetic target region set.

[0143] Alternatively, the capturing step is performed with the sequence-variable target region probe set in a first vessel and with the epigenetic target region probe set in a second vessel, or the contacting step is performed with the sequence-variable target regionAttorney Docket No.: GH0259WO probe set at a first time and a first vessel and the epigenetic target region probe set at a second time before or after the first time. This approach allows for preparation of separate first and second compositions comprising captured DNA corresponding to the sequence-variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately as desired (e.g., to fractionate based on methylation as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis such as sequencing.

[0144] In some embodiments, the DNA is amplified. In some embodiments, amplification is performed before the capturing step. In some embodiments, amplification is performed after the capturing step.

[0145] In some embodiments, adapters are included in the DNA. This may be done concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer, e.g., as described above. Alternatively, adapters can be added by other approaches, such as ligation.

[0146] In some embodiments, tags, which may be or include barcodes, are included in the DNA. Tags can facilitate identification of the origin of a nucleic acid. For example, barcodes can be used to allow the origin (e.g., subject) whence the DNA came to be identified following pooling of a plurality of samples for parallel sequencing. This may be done concurrently with an amplification procedure, e.g., by providing the barcodes in a 5’ portion of a primer, e.g., as described above. In some embodiments, adapters and tags / barcodes are provided by the same primer or primer set. For example, the barcode may be located 3’ of the adapter and 5’ of the target-hybridizing portion of the primer. Alternatively, barcodes can be added by other approaches, such as ligation, optionally together with adapters in the same ligation substrate.

[0147] Additional details regarding amplification, tags, and barcodes are discussed in the “General Features of the Methods” section below, which can be combined to the extent practicable with any of the foregoing embodiments and the embodiments set forth in the introduction and summary section.Captured Set

[0148] In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. With respect to the disclosed methods, the captured set of DNA may be provided, e.g., by performing a capturing step after a partitioning step as described herein. The captured set may comprise DNA corresponding to a sequence-variable target region set, an epigeneticAttorney Docket No.: GH0259WO target region set, or a combination thereof. In some embodiments the quantity of captured sequence-variable target region DNA is greater than the quantity of the captured epigenetic target region DNA, when normalized for the difference in the size of the targeted regions (footprint size).

[0149] Alternatively, first and second captured sets may be provided, comprising, respectively, DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set. The first and second captured sets may be combined to provide a combined captured set.

[0150] In some embodiments in which a captured set comprising DNA corresponding to the sequence-variable target region set and the epigenetic target region set includes a combined captured set as discussed above, the DNA corresponding to the sequence-variable target region set may be present at a greater concentration than the DNA corresponding to the epigenetic target region set, e.g., a 1.1 to 1.2-fold greater concentration, a 1.2- to 1.4-fold greater concentration, a 1.4- to 1.6-fold greater concentration, a 1.6- to 1.8-fold greater concentration, a 1.8- to 2.0-fold greater concentration, a 2.0- to 2.2-fold greater concentration, a 2.2- to 2.4-fold greater concentration a 2.4- to 2.6-fold greater concentration, a 2.6- to 2.8-fold greater concentration, a 2.8- to 3.0-fold greater concentration, a 3.0- to 3.5- fold greater concentration, a 3.5- to 4.0, a 4.0- to 4.5-fold greater concentration, a 4.5- to 5.0- fold greater concentration, a 5.0- to 5.5-fold greater concentration, a 5.5- to 6.0-fold greater concentration, a 6.0- to 6.5-fold greater concentration, a 6.5- to 7.0-fold greater, a 7.0- to 7.5- fold greater concentration, a 7.5- to 8.0-fold greater concentration, an 8.0- to 8.5-fold greater concentration, an 8.5- to 9.0-fold greater concentration, a 9.0- to 9.5-fold greater concentration, 9.5- to 10.0-fold greater concentration, a 10- to 11-fold greater concentration, an 11- to 12-fold greater concentration a 12- to 13 -fold greater concentration, a 13- to 14-fold greater concentration, a 14- to 15-fold greater concentration, a 15- to 16-fold greater concentration, a 16- to 17-fold greater concentration, a 17- to 18-fold greater concentration, an 18- to 19-fold greater concentration, a 19- to 20-fold greater concentration, a 20- to 30- fold greater concentration, a 30- to 40-fold greater concentration, a 40- to 50-fold greater concentration, a 50- to 60-fold greater concentration, a 60- to 70-fold greater concentration, a 70- to 80-fold greater concentration, a 80- to 90-fold greater concentration, a 90- to 100-fold greater concentration, a 10- to 20-fold greater concentration, a 10- to 40-fold greater concentration, a 10- to 50-fold greater concentration, a 10- to 70-fold greater concentration, or a 10- to 100-fold greater concentration. The degree of difference in concentrationsAttorney Docket No.: GH0259WO accounts for normalization for the footprint sizes of the target regions, as discussed in the definition section.Epigenetic Target Region Set

[0151] The epigenetic target region set may comprise one or more types of target regions likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells and from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The epigenetic target region set may also comprise one or more control regions, e.g., as described herein. In some embodiments, the epigenetic target region set has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has a footprint in the range of 100-1000 kb, e.g., 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1,000 kb.HyDermethylation Variable Target Regions

[0152] In some embodiments, the epigenetic target region set comprises one or more hypermethylation variable target regions. In general, hypermethylation variable target regions refer to regions where an increase in the level of observed methylation, e.g., in a cfDNA sample, indicates an increased likelihood that a sample (e.g., of cfDNA) contains DNA produced by neoplastic cells, such as tumor or cancer cells. For example, hypermethylation of promoters of tumor suppressor genes has been observed repeatedly. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and references cited therein. In an example, hypermethylation variable target regions can include regions that do not necessarily differ in methylation in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ in methylation (e.g., have more methylation) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such hypermethylation variable target regions. In some embodiments, hypermethylation variable target regions include one or more genomic regions, where the cfDNA molecules in those regions do not differ in methylation state in cancer subjects relative to cfDNA from healthy subjects, but the presence / increased quantity of hypermethylated cfDNA in those regions is indicative of a particular tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g. tumor shedding) into circulation.Attorney Docket No.: GH0259WO

[0153] Hypermethylation target regions may be obtained, e.g., from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017), describe construction of a probabilistic method called CancerLocator using hypermethylation target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylation target regions can be specific to one or more types of cancer. Accordingly, in some embodiments, the hypermethylation target regions include one, two, three, four, or five subsets of hypermethylation target regions that collectively show hypermethylation in one, two, three, four, or five of breast, colon, kidney, liver, and lung cancers.HvDomethylation Variable Target Regions

[0154] Global hypomethylation is a commonly observed phenomenon in various cancers. See, e.g., Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1 :239-259 (2009) (review article noting observations of hypomethylation in colon, ovarian, prostate, leukemia, hepatocellular, and cervical cancers). For example, regions such as repeated elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and intergenic regions that are ordinarily methylated in healthy cells may show reduced methylation in tumor cells. Accordingly, in some embodiments, the epigenetic target region set includes hypomethylation variable target regions, where a decrease in the level of observed methylation indicates an increased likelihood that a sample (e.g., of cfDNA) contains DNA produced by neoplastic cells, such as tumor or cancer cells. In an example, hypomethylation variable target regions can include regions that do not necessarily differ in methylation state in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ in methylation (e.g., are less methylated) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such hypomethylation variable target regions. In some embodiments, hypomethylation variable target regions include one or more genomic regions, where the cfDNA molecules in those regions do not differ in methylation state in cancer subjects relative to cfDNA from healthy subjects, but the presence / increased quantity of hypomethylated cfDNA in those regions is indicative of a particular tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g. tumor shedding) into circulation.Attorney Docket No.: GH0259WO

[0155] In some embodiments, hypomethylation variable target regions include repeated elements and / or intergenic regions. In some embodiments, repeated elements include one, two, three, four, or five of LINE 1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.

[0156] Exemplary specific genomic regions that show cancer-associated hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1. In some embodiments, the hypomethylation variable target regions overlap or comprise one or both of these regions.

[0157] In some embodiments, the probes for the epigenetic target region set comprise probes specific for one or more hypomethylation variable target regions. The hypomethylation variable target regions may be any of those set forth above. For example, the probes specific for one or more hypomethylation variable target regions may include probes for regions such as repeated elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and intergenic regions that are ordinarily methylated in healthy cells may show reduced methylation in tumor cells.

[0158] In some embodiments, probes specific for hypomethylation variable target regions include probes specific for repeated elements and / or intergenic regions. In some embodiments, probes specific for repeated elements include probes specific for one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.

[0159] Exemplary probes specific for genomic regions that show cancer-associated hypomethylation include probes specific for nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. In some embodiments, the probes specific for hypomethylation variable target regions include probes specific for regions overlapping or comprising nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome

[0160] Probes for detecting the panel of regions can include those for detecting genomic regions of interest (hotspot regions) as well as nucleosome-aware probes (e.g., KRAS codons 12 and 13) and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation impacted by nucleosome binding patterns and GC sequence composition. Regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models. SubjectsAttorney Docket No.: GH0259WO

[0161] In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject having a cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject having neoplasia. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having neoplasia. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject in remission from a tumor, cancer, or neoplasia (e.g., following chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the lung. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the breast. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.

[0162] In some embodiments, the sequence-variable target region probe set has a footprint of at least 0.5 kb, e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5-100 kb, e.g., 0.5-2 kb, 2-10 kb, 10-20 kb, 20-30 kb, 30-40 kb, 40-50 kb, 50-60 kb, 60-70 kb, 70-80 kb, 80-90 kb, and 90-100 kb.

[0163] In some embodiments, the probes specific for the sequence-variable target region set comprise probes specific for target regions from at least 10, 20, 30, or 35 cancer- related genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESRI, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.Compositions Comprising Captured DNA

[0164] Provided herein is a combination comprising first and second populations of captured DNA. The first population may comprise or be derived from DNA with a cytosine modification in a greater proportion than the second population. The first population may comprise a form of a first nucleobase originally present in the DNA with altered base pairingAttorney Docket No.: GH0259WO specificity and a second nucleobase without altered base pairing specificity, wherein the form of the first nucleobase originally present in the DNA prior to alteration of base pairing specificity is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the form of the first nucleobase originally present in the DNA prior to alteration of base pairing specificity and the second nucleobase have the same base pairing specificity. The second population does not comprise the form of the first nucleobase originally present in the DNA with altered base pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleobase is a modified or unmodified cytosine and the second nucleobase is a modified or unmodified cytosine. The first and second nucleobase may be any of those discussed herein in the Summary or with respect to subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub sample.

[0165] In some embodiments, the first population comprises a sequence tag selected from a first set of one or more sequence tags and the second population comprises a sequence tag selected from a second set of one or more sequence tags, and the second set of sequence tags is different from the first set of sequence tags. The sequence tags may comprise barcodes.

[0166] In some embodiments, the first population comprises protected hmC, such as glucosylated hmC. In some embodiments, the first population was subjected to any of the conversion procedures discussed herein, such as bisulfite conversion, Ox-BS conversion, TAB conversion, ACE conversion, TAP conversion, TAPSP conversion, or CAP conversion. In some embodiments, the first population was subjected to protection of hmC followed by deamination of mC and / or C. In some embodiments of the combination, the first population comprises or was derived from DNA with a cytosine modification in a greater proportion than the second population and the first population comprises first and second subpopulations, and the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the second population does not comprise the first nucleobase. In some embodiments, the first nucleobase is a modified or unmodified cytosine, and the second nucleobase is a modified or unmodified cytosine, optionally wherein the modified cytosine is mC or hmC. In some embodiments, the first nucleobase is a modified or unmodified adenine,Attorney Docket No.: GH0259WO and the second nucleobase is a modified or unmodified adenine, optionally wherein the modified adenine is mA.

[0167] In some embodiments, the first nucleobase (e.g., a modified cytosine) is biotinylated. In some embodiments, the first nucleobase (e.g., a modified cytosine) is a product of a Huisgen cycloaddition to P-6-azide-glucosyl-5-hydroxymethylcytosine that comprises an affinity label (e.g., biotin).

[0168] In any of the combinations described herein, the captured DNA may comprise cfDNA. The captured DNA may have any of the features described herein concerning captured sets, including, e.g., a greater concentration of the DNA corresponding to the sequence-variable target region set (normalized for footprint size as discussed above) than of the DNA corresponding to the epigenetic target region set. In some embodiments, the DNA of the captured set comprises sequence tags, which may be added to the DNA as described herein. In general, the inclusion of sequence tags results in the DNA molecules differing from their naturally occurring, untagged form.

[0169] The combination may further comprise a probe set described herein or sequencing primers, each of which may differ from naturally occurring nucleic acid molecules. For example, a probe set described herein may comprise a capture moiety, and sequencing primers may comprise a non-naturally occurring label.Cancer and Other Diseases

[0170] The present methods can be used to diagnose presence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy.

[0171] Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.

[0172] In some embodiments, the methods and systems disclosed herein may be used to identify customized or targeted therapies to treat a given disease or condition in patientsAttorney Docket No.: GH0259WO based on the classification of a nucleic acid variant as being of somatic or germline origin. Typically, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, gliomas, astrocytomas, breast carcinoma, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal carcinoma, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinomas, gastrointestinal stromal tumors (GISTs), endometrial carcinoma, endometrial stromal sarcomas, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder carcinomas, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinomas, Wilms tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, Lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphomas, nonHodgkin lymphoma, diffuse large B-cell lymphoma, Mantle cell lymphoma, T cell lymphomas, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T cell lymphomas, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral cavity squamous cell carcinomas, osteosarcoma, ovarian carcinoma, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasms, acinar cell carcinomas. Prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine carcinomas, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0173] Genetic data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject orAttorney Docket No.: GH0259WO practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.

[0174] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile comprises a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.

[0175] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation, epigenetic variation, and mutation analyses alone or in combination.

[0176] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0177] Non-limiting examples of other genetic-based diseases, disorders, or conditions that are optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha- 1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri du chat, Crohn's disease, cystic fibrosis, Dercum disease, down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome,Attorney Docket No.: GH0259WO osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson disease, or the like.

[0178] In some embodiments, a method described herein comprises detecting a presence or absence of DNA originating or derived from a tumor cell at a preselected timepoint following a previous cancer treatment of a subject previously diagnosed with cancer using a set of sequence information obtained as described herein. The method may further comprise determining a cancer recurrence score that is indicative of the presence or absence of the DNA originating or derived from the tumor cell for the test subject. Where a cancer recurrence score is determined, it may further be used to determine a cancer recurrence status. The cancer recurrence status may be at risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may be at low or lower risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. In particular embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either at risk for cancer recurrence or at low or lower risk for cancer recurrence.

[0179] In some embodiments, a cancer recurrence score is compared with a predetermined cancer recurrence threshold, and the test subject is classified as a candidate for a subsequent cancer treatment when the cancer recurrence score is above the cancer recurrence threshold or not a candidate for therapy when the cancer recurrence score is below the cancer recurrence threshold. In particular embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for a subsequent cancer treatment or not a candidate for therapy.

[0180] The methods discussed above may further comprise any compatible feature or features set forth elsewhere herein, including in the section regarding methods of determining a risk of cancer recurrence in a test subject and / or classifying a test subject as being a candidate for a subsequent cancer treatment.Methods of Determining a Risk of Cancer Recurrence in a Test Subject and / or Classifying a Test Subject as Being a Candidate for a Subsequent Cancer Treatment

[0181] In some embodiments, a method provided herein is a method of determining a risk of cancer recurrence in a test subject. In some embodiments, a method provided herein is a method of classifying a test subject as being a candidate for a subsequent cancer treatment.Attorney Docket No.: GH0259WO

[0182] Any of such methods may comprise collecting DNA (e.g., originating or derived from a tumor cell) from the test subject diagnosed with the cancer at one or more preselected timepoints following one or more previous cancer treatments to the test subject. The subject may be any of the subjects described herein. The DNA may be cfDNA. The DNA may be obtained from a tissue sample.

[0183] Any of such methods may comprise capturing a plurality of sets of target regions from DNA from the subject, wherein the plurality of target region sets comprises a sequence-variable target region set and an epigenetic target region set, whereby a captured set of DNA molecules is produced. The capturing step may be performed according to any of the embodiments described elsewhere herein. In any of such methods, the previous cancer treatment may comprise surgery, administration of a therapeutic composition, and / or chemotherapy.

[0184] Any of such methods may comprise sequencing the captured DNA molecules, whereby a set of sequence information is produced. The captured DNA molecules of the sequence-variable target region set may be sequenced to a greater depth of sequencing than the captured DNA molecules of the epigenetic target region set.

[0185] Any of such methods may comprise detecting a presence or absence of DNA originating or derived from a tumor cell at a preselected timepoint using the set of sequence information. The detection of the presence or absence of DNA originating or derived from a tumor cell may be performed according to any of the embodiments thereof described elsewhere herein.

[0186] Methods of determining a risk of cancer recurrence in a test subject may comprise determining a cancer recurrence score that is indicative of the presence or absence, or amount, of the DNA originating or derived from the tumor cell for the test subject. The cancer recurrence score may further be used to determine a cancer recurrence status. The cancer recurrence status may be at risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may be at low or lower risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. In particular embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either at risk for cancer recurrence or at low or lower risk for cancer recurrence.

[0187] Methods of classifying a test subject as being a candidate for a subsequent cancer treatment may comprise comparing the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold, thereby classifying the test subject as aAttorney Docket No.: GH0259WO candidate for the subsequent cancer treatment when the cancer recurrence score is above the cancer recurrence threshold or not a candidate for therapy when the cancer recurrence score is below the cancer recurrence threshold. In particular embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for a subsequent cancer treatment or not a candidate for therapy. In some embodiments, the subsequent cancer treatment comprises chemotherapy or administration of a therapeutic composition.

[0188] Any of such methods may comprise determining a disease-free survival (DFS) period for the test subject based on the cancer recurrence score; for example, the DFS period may be 1 year, 2 years, 3, years, 4 years, 5 years, or 10 years.

[0189] In some embodiments, the set of sequence information comprises sequencevariable target region sequences, and determining the cancer recurrence score may comprise determining at least a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in sequence-variable target region sequences.

[0190] In some embodiments, a number of mutations in the sequence-variable target regions chosen from 1, 2, 3, 4, or 5 is sufficient for the first subscore to result in a cancer recurrence score classified as positive for cancer recurrence. In some embodiments, the number of mutations is chosen from 1, 2, or 3.

[0191] In some embodiments, the set of sequence information comprises epigenetic target region sequences, and determining the cancer recurrence score comprises determining a second subscore indicative of the amount of molecules (obtained from the epigenetic target region sequences) that represent an epigenetic state different from DNA found in a corresponding sample from a healthy subject (e.g., cfDNA found in a blood sample from a healthy subject, or DNA found in a tissue sample from a healthy subject where the tissue sample is of the same type of tissue as was obtained from the test subject). These abnormal molecules (i.e., molecules with an epigenetic state different from DNA found in a corresponding sample from a healthy subject) may be consistent with epigenetic changes associated with cancer, e.g., methylation of hypermethylation variable target regions and / or perturbed fragmentation of fragmentation variable target regions, where “perturbed” means different from DNA found in a corresponding sample from a healthy subject.

[0192] In some embodiments, a proportion of molecules corresponding to the hypermethylation variable target region set and / or fragmentation variable target region set that indicate hypermethylation in the hypermethylation variable target region set and / or abnormal fragmentation in the fragmentation variable target region set greater than or equalAttorney Docket No.: GH0259WO to a value in the range of 0.001%-10% is sufficient for the second subscore to be classified as positive for cancer recurrence. The range may be 0.001%-l%, 0.005%-l%, 0.01%-5%, 0.01%-2%, or 0.01%-l%.

[0193] In some embodiments, any of such methods may comprise determining a fraction of tumor DNA from the fraction of molecules in the set of sequence information that indicate one or more features indicative of origination from a tumor cell. This may be done for molecules corresponding to some or all of the epigenetic target regions, e.g., including one or both of hypermethylation variable target regions and fragmentation variable target regions (hypermethylation of a hypermethylation variable target region and / or abnormal fragmentation of a fragmentation variable target region may be considered indicative of origination from a tumor cell). This may be done for molecules corresponding to sequence variable target regions, e.g., molecules comprising alterations consistent with cancer, such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA may be determined based on a combination of molecules corresponding to epigenetic target regions and molecules corresponding to sequence variable target regions.

[0194] Determination of a cancer recurrence score may be based at least in part on the fraction of tumor DNA, wherein a fraction of tumor DNA greater than a threshold in the range of 10-11 to 1 or 10-10 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, a fraction of tumor DNA greater than or equal to a threshold in the range of 10-10 to 10-9, 10-9 to 10-8, 10-8 to 10-7, 10-7 to 10-6, 10-6 to 10-5, 10-5 to 10-4, 10-4 to 10-3, 10-3 to 10-2, or 10-2 to 10-1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, the fraction of tumor DNA greater than a threshold of at least 10-7 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. A determination that a fraction of tumor DNA is greater than a threshold, such as a threshold corresponding to any of the foregoing embodiments, may be made based on a cumulative probability. For example, the sample was considered positive if the cumulative probability that the tumor fraction was greater than a threshold in any of the foregoing ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999. In some embodiments, the probability threshold is at least 0.95, such as 0.99.

[0195] In some embodiments, the set of sequence information comprises sequencevariable target region sequences and epigenetic target region sequences, and determining the cancer recurrence score comprises determining a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in sequence-variable target regionAttorney Docket No.: GH0259WO sequences and a second subscore indicative of the amount of abnormal molecules in epigenetic target region sequences, and combining the first and second subscores to provide the cancer recurrence score. Where the first and second subscores are combined, they may be combined by applying a threshold to each subscore independently (e.g., greater than a predetermined number of mutations (e.g., > 1) in sequence-variable target regions, and greater than a predetermined fraction of abnormal molecules (i.e., molecules with an epigenetic state different from the DNA found in a corresponding sample from a healthy subject; e.g., tumor) in epigenetic target regions), or training a machine learning classifier to determine status based on a plurality of positive and negative training samples.

[0196] In some embodiments, a value for the combined score in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.

[0197] In any embodiment where a cancer recurrence score is classified as positive for cancer recurrence, the cancer recurrence status of the subject may be at risk for cancer recurrence and / or the subject may be classified as a candidate for a subsequent cancer treatment.

[0198] In some embodiments, the cancer is any one of the types of cancer described elsewhere herein, e.g., colorectal cancer.Therapies and Related Administration

[0199] In certain embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients given the status of a nucleic acid variant as being of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, customized therapies include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.

[0200] In certain embodiments, the status of a nucleic acid variant from a sample from a subject as being of somatic or germline origin may be compared with a database of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the test subject and / or patients who are receiving, or who have received, the same therapy as the test subject. A customized or targeted therapy (or therapies) may beAttorney Docket No.: GH0259WO identified when the nucleic variant and the comparator results satisfy certain classification criteria (e.g., are a substantial or an approximate match).

[0201] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing an immunotherapeutic agent are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by methods such as, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, which administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.

[0202] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the invention. It is therefore contemplated that the disclosure shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

[0203] While the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be clear to one of ordinary skill in the art from a reading of this disclosure that various changes in form and detail can be made without departing from the true scope of the disclosure and may be practiced within the scope of the appended claims. For example, all the methods, systems, computer readable media, and / or component features, steps, elements, or other aspects thereof can be used in various combinations.Cancer Treatments, TherapiesAttorney Docket No.: GH0259WO

[0204] In some cases, the cancer treatment includes, without limitation, imatinib, gefatinib, afatinib, dacomitinib, sunitinib, sorafenib, vandetanib, brivanib, cabozantib, neratinib, tivantinib, bevacizumab, cixutumumab, dalotuzumab, figitumumab, rilotumumab, onartuzumab, ganitumab, ramucirumab, ridaforolimus, tensirolimus, everolimus, BMS- 690514, BMS-754807, EMD 525797, GDC-0973, GDC-0941, MK-2206, AZD6244, GSK1120212, PX-866, XL821, IMC-A12, MM-121, PF-02341066, RG7160, and Sym004. Antibodies suitable for use as anti -EGFR therapy include cetuximab (Trade Name: Erbitux) and panitumumab (Trade Name: Vectibex). In some cases. In some cases, the cancer treatment includes EGFR tyrosine kinase inhibitors such as gefitinib (Trade Name: Iressa), erlotinib (Trade Name: Tarceva), lapatinib, canertinib, and cetuximab.

[0205] In some instances, therapties may be used in combination, such as an anti- EGFR therapy and an anti-EGFR therapy. Anti-EGFR therapy may be used in combination with any combination of chemotherapeutic agents or chemotherapeutic regimens, for example, FOLFOX (fluorouracil [5-FU] / leucovorin / oxaliplatin), FOLFIRI (5- FU / leucovorin / irinotecan), and the like.

[0206] In some aspects, a cancer treatment ai administered to a subject. In some cases, the cancer treatment is administered in combination another therapy, such as a non- anti-EGFR therapy with anti-EGFR therapy.Genetic Analysis

[0207] Genetic analysis includes detection of nucleotide sequence variants and copy number variations. Genetic variants can be determined by sequencing. The sequencing method can be massively parallel sequencing, that is, simultaneously (or in rapid succession) sequencing any of at least 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules. Sequencing methods may include, but are not limited to: high- throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), Nextgeneration sequencing, Single Molecule Sequencing by Synthesis (SMSS)(Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms and any other sequencing methods known in the art.

[0208] Sequencing can be made more efficient by performing sequence capture, that is, the enrichment of a sample for target sequences of interest, e.g., sequences including theAttorney Docket No.: GH0259WOKRAS and / or EGFR genes or portions of them containing sequence variant biomarkers. Sequence capture can be performed using immobilized probes that hybridize to the targets of interest.

[0209] Cell free DNA can include small amounts of tumor DNA mixed with germline DNA. Sequencing methods that increase sensitivity and specificity of detecting tumor DNA, and, in particular, genetic sequence variants and copy number variation, can be useful in the methods of this invention. Such methods are described in, for example, in WO 2014 / 039556. These methods not only can detect molecules with a sensitivity of up to or greater than 0.1%, but also can distinguish these signals from noise typical in current sequencing methods. Increases in sensitivity and specificity from blood-based samples of cfDNA can be achieved using various methods. One method includes high efficiency tagging of DNA molecules in the sample, e.g., tagging at least any of 50%, 75% or 90% of the polynucleotides in a sample. This increases the likelihood that a low-abundance target molecule in a sample will be tagged and subsequently sequenced, and significantly increases sensitivity of detection of target molecules.

[0210] Another method involves molecular tracking, which identifies sequence reads that have been redundantly generated from an original parent molecule, and assigns the most likely identity of a base at each locus or position in the parent molecule. This significantly increases specificity of detection by reducing noise generated by amplification and sequencing errors, which reduces frequency of false positives.

[0211] Methods of the present disclosure can be used to detect genetic variation in non-uniquely tagged initial starting genetic material (e.g., rare DNA) at a concentration that is less than 5%, 1%, 0.5%, 0.1%, 0.05%, or 0.01%, at a specificity of at least 99%, 99.9%, 99.99%, 99.999%, 99.9999%, or 99.99999%. Sequence reads of tagged polynucleotides can be subsequently tracked to generate consensus sequences for polynucleotides with an error rate of no more than 2%, 1%, 0.1%, or 0.01%.

[0212] Copy number variation determination can involve determining a quantitative measure of polynucleotides in a sample mapping to a genetic locus, such as the EGFR gene or KRAS gene. The quantitative measure can be a number. Once the total number of polynucleotides mapping to a locus is determined, this number can be used in standard methods of determining Copy Number Variation at the locus. A quantitative measure can be normalized against a standard. In one method, a quantitative measure at a test locus can be standardized against a quantitative measure of polynucleotides mapping to a control locus in the genome, such as gene of known copy number. In another method, the quantitativeAttorney Docket No.: GH0259WO measure can be compared against the amount of nucleic acid in the original sample. For example, the quantitative measure can be compared against an expected measure for diploidy. In another method, the quantitative measure can be normalized against a measure from a control sample, and normalized measures at different loci can be compared. In another method, quantifying involves quantifying parent or original molecules in a sample mapping to a locus, rather than number of sequence reads. A copy number variation may be an amplification or a deletion or truncation of a gene. An amplification may be 3, 4, 5, 6, 7, 8, 9, 10, or 10 or more copies of a gene. A deletion or truncation may be 0 or 1 copies of a gene.

[0213] An example of a method for detecting copy number variation may include an array. The array may comprise a plurality of capture probes. The capture probes can be oligonucleotides that are bound to the surface of the array. The capture probes may bind to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 genes as set forth in Table 1. DNA derived from the subject may be labeled (e.g., with a fluorophore) prior to hybridization for detection.

[0214] In other examples, a gene of interest may be amplified using primers that recognize the gene of interest. The primers may hybridize to a gene upstream and / or downstream of a particular region of interest (e.g., upstream of a mutation site). A detection probe may be hybridized to the amplification product. Detection probes may specifically hybridize to a wild-type sequence or to a mutated / variant sequence. Detection probes may be labeled with a detectable label (e.g., with a fluorophore). Detection of a wild-type or mutant sequence may be performed by detecting the detectable label (e.g., fluorescence imaging). In examples of copy number variation, a gene of interest may be compared with a reference gene. Differences in copy number between the gene of interest and the reference gene may indicate amplification or deletion / truncation of a gene. Examples of platforms suitable to perform the methods described herein include digital PCR platforms such as e.g., Fluidigm Digital Array.EXAMPLES

[0215] The following examples are given for the purpose of illustrating various embodiments of the invention and are not meant to limit the present invention in any fashion. The present examples, along with the methods described herein are presently representative of preferred embodiments, are exemplary, and are not intended as limitations on the scope of the invention. Changes therein and other uses which are encompassed within the spirit of the invention as defined by the scope of the claims will occur to those skilled in the art.Attorney Docket No.: GH0259WOExample 1 - Promoter methylation, generally

[0216] The described epigenomic, genomic liquid biopsy platform detection possesses capabilities that provides an innovative approach for monitoring epigenetic alterations, with significant implications for early diagnosis, prognostic assessment, and treatment stratification in oncology.

[0217] Data across >17k samples, including >5k cancer free samples, and >100 cancer types, were used to define promoter regions associated with hypermethylation in cancer related genes. PM regions were refined to ensure (a) the methylation signal is present exclusively in cancer and absent in cancer-free samples, and (b) a negative correlation between gene expression and methylation is observed in TCGA cancer samples. PM regions are quantified relative to constitutively hyper-methylated control regions, measuring gene PM relative to total cfDNA. Further information is found in Inti. App. No. PCT / US2024 / 023278, PCT / US2024 / 024378 each of which which is fully incorporated by reference herein.

[0218] In one example, assessment for BRCA1 PM in ctDNA from 1016 patients with late-stage breast cancer was performed, along with genomic sequencing of 800+ genes and PM profiling of 398 cancer-related genes was performed by the epigenomic methylation detection assay.

[0219] Analytical accuracy was demonstrated using an orthogonal Enzymatic Methyl Sequencing (EM-seq) assay on a cohort of clinical cfDNA and contrived samples. Limit of detection, precision and limit of blank were established based on cfDNA samples, cell line dilutions, and healthy donor samples. Gene PM prevalence was assessed in a test cohort of >7k samples across >17 different cancer types analyzed on Guardant Infinity platform and compared to relative prevalence in published tissue data sets.

[0220] The described epigenomic, genomic liquid biopsy promoter platform detection is highly concordant with orthogonal EM-Seq, with 98.2% positive concordance and 98% negative concordance above the limit of detection. A limit of detection (>=95% detection) was established at PM MAF 1.6% across a panel of 47 promoter regions.

[0221] The observed prevalence of PM in liquid biopsy samples is consistent with tissue-based findings and correlates with underlying tumor characteristics, such as MSI-H and LOH events, thus demonstrating Guardant Infinity’s ability to identify potentially actionable PM patterns in a non-invasive manner. This analysis represents a fraction of the PM landscape observed, highlighting the broader potential of methylation as a novel biomarker for cancer detection and characterization.Attorney Docket No.: GH0259WOExample 2 - Epigenetic silencing ofMLHl in colorectal cancer and endometrial cancer through assessment of promoter methylation in cfDNA liquid biopsy assay

[0222] Aberrant methylation of the MLH1 promoter (PM) is a primary cause of mismatch repair (MMR) deficiency, especially in sporadic colorectal (CRC) and endometrial cancers (EC). This epigenetic alteration silences MLH1 expression, leading to the accumulation of DNA replication errors that manifest as microsatellite instability -high (MSI- H). MLH1 PM is a well-established biomarker of MMR dysfunction, associated with tumor progression, and predictive of responses to immunotherapies in MSI-H cancers. Traditionally, detecting MLH1 PM has relied on single-gene testing of tissue biopsies, which are invasive, less accessible in a clinical setting and less practical for tracking tumor evolution over time.

[0223] Described herein is a detailed analysis ofMLHl PM detected on the epigenomic, genomic liquid biopsy platform in a cohort of CRC and EC to assess the role of PM in genomic instability. 978 CRC and EC samples were analyzed on a epigenomic, genomic liquid biopsy platform for characterizing promoter methylation. MSI was quantified as previously described across ~ 2000 homopolymer sites (“MSI score”). Thereafter, one can quantify the capture of methylated molecules in the MLH1 promoter region in relation to the capture of methylated molecules in the constitutively hyper-methylated positive control regions using a combined genomic, epigenomic platform, which provides a metric for gene PM relative to total ctDNA.

[0224] MLH1 PM was significantly associated with MSI high status in both CRC and EC (42%, 100%) in MSI-H cases. MLH1 PM in CRC showed significant enrichment for pathogenic BRAF mutations, V600E being the most abundant at 36% MLH1 PM, and MMR somatic alterations (OR=5, p=0.001). Mutations primarily occurred in MSH6 and MLH3. Pearson correlation between MSI score and PM levels was 0.88 (p=2.03e-22). Significantly higher MSI scores are observed between MLH1 PM and unmethylated groups when MSI is not high (p=1.5e-08), while no significant difference between methylation groups exists when MSI high samples (p=0.1).

[0225] Our findings confirm an epigenomic, genomic liquid biopsy detection platform’s ability to detect MLH1 PM with liquid biopsy. The strong correlation between MLH1 PM and MSI levels highlights MLH1 PM as a reliable biomarker of MMR pathway dysfunction. Differences in MSI levels between methylated and unmethylated, particularly in samples where MSI is not detected, underscores the potential of PM to enhance the detectionAttorney Docket No.: GH0259WO rate of patients likely to benefit from immunotherapy. Additionally, higher MSI scores and enrichment of MMR mutation in MLH1 PM, non MSI high samples indicates MMR dysfunction may still be occurring at an inceptive or decreased level.Example 3 - Analytical validation and refinement of cancer gene promoter methylation detection in cfDNA liquid biopsy assay / Landscape of promoter gene hypermethylation in cfDNA liquid biopsy assay

[0226] As described, promoter methylation (PM) is a critical epigenetic mechanism that silences tumor suppressor genes, thereby impairing pathways involved in DNA repair, cell cycle regulation, and apoptosis which collectively drive cancer progression. Understanding the role of PM in these pathways highlights its utility both as a biomarker for early detection and as a predictor of treatment response in precision oncology.

[0227] In an additional study that employs epigenomic, genomic liquid biopsy platform to enable high-sensitivity detection of PM across key tumor suppressor genes in circulating cell-free DNA (cfDNA). The Inventors comprehensively profiled PM in genes across diverse cancers, allowing us to assess methylation status non-invasively. This platform’s detection capabilities provide an innovative approach for monitoring epigenetic alterations, with significant implications for early diagnosis, prognostic assessment, and treatment stratification in oncology.

[0228] Data across more than 17k samples, including 5k cancer free samples, and >100 cancer types, were used to define PM regions associated with hypermethylation in cancer related genes. Promoter methylation (PM) regions were refined to (a) ensure the methylation signal detected is cancer-specific by isolating signals present exclusively in cancer samples and absent in cancer-free samples, and (b) focus on genomic regions in which a negative correlation between gene expression and methylation could be demonstrated in TCGA cancer samples.

[0229] Additionally, we quantified the capture of methylated molecules in the MLH1 promoter region in relation to the capture of methylated molecules in the constitutively hyper-methylated positive control regions using Guardant Infinity platform, which provides a metric for gene PM relative to total ctDNA.

[0230] Analytical accuracy was validated using an orthogonal Enzymatic Methyl Sequencing (EM-seq) assay on 83 cfDNA samples and one contrived sample. Other analytical performance characteristics such as limit of detection, precision and limit of blankAttorney Docket No.: GH0259WO were established based on fDNA samples spike-in with background, synthetic cell line dilutions, and healthy donor samples.

[0231] Gene PM prevalence was assessed in a cohort of >7300 samples across >17 different cancer types analyzed on epigenomic, genomic liquid biopsy platform, and compared to relative prevalence in TCGA and other published data sets.

[0232] The described epigenomic, genomic liquid biopsy platform promoter methylation detection has high concordance with orthogonal EM-Seq with 88% positive concordance and 98% negative concordance, indicating high concordance of capturing and quantifying PM status, with a limit of detection (>=95% detection) established at PM MAF =I.6% across a panel of 47 promoter regions.

[0233] Described herein is evidence that the described epigenomic, genomic liquid biopsy platform can detect PM in cancer related genes and relative prevalence aligns with the landscape of PM in external data sets. MLH1 PM at 2.4% and high MGMT PM frequency,I I.5%, in our cohort of colorectal cancers (n=933), and BRCA1 PM at 2.3% in breast cancer (n=1330).

[0234] These findings underline the clinical significance of PM in identifying inactivation of tumor suppressor genes. The described epigenomic, genomic liquid biopsy platform has the ability to detect PM to guide personalized treatment strategies.Example 4 - STAT5A promoter hypermethylation as a biomarker of immune checkpoint inhibitor response in squamous cell carcinoma in cfDNA liquid biopsies

[0235] Another biomarker of interest, STAT5A, directly binds on the promoter of DNMT3 A result in loss of downstream targets that normally restrain DNMT expression, resulting in hyper-methylation in wider genome. STAT5A promoter observed in lung squamous carcinomas, and prostate cancers, highly prevalent in bladder cancers. STAT5A- mediated interferon signaling regulates the expression of CD274 (PD-L1) and PDCD1LG2 (encodes PD-L2), where hypermethylated STAT5A are significantly associated with immune cell depletion in squamous tumors. DNA methyltransferase inhibitors (decitabine, azacitidine) can restore STAT5A expression, and restore its tumor-suppressive function. These observations all support the notion that promoter methylation can be utilized as an important component of tumor profiling.

[0236] A combined genomic, epigenomic detection platform was used to detect the hypermethylated STAT5A promoter region, defined initially as 5kb upstream from theAttorney Docket No.: GH0259WO transcription start sites (TSS). Promoter hypermethylation (PM) regions were evaluated for correlation between methylation and expression using The Cancer Genome Atlas (TCGA) data to identify loci where PM was strongly associated with transcriptional silencing. These candidate regions were evaluated in a large cohort of cfDNA cancer-free samples (n=32k) to ensure high analytical specificity and minimal background methylation. PM status is distinguished relative to constitutively hypermethylated control regions, measuring methylation level relative to total cfDNA.

[0237] Cross-platform concordance and external validation were assessed using external tissue cohorts. Clinical utility was explored using the a real world evidence database to identify patients with squamous cell carcinomas (SCC) treated with immune checkpoint inhibitor (ICPI), stratified by STAT5A PM status (n=815, 91% LSCC, 9% HNSCC).

[0238] STAT5A PM prevalence in cohorts of interest was concordant with other published datasets across tumor types. The assay demonstrated high analytical specificity (negative percent agreement > 98%) and minimal background methylation in liquid platforms.

[0239] In the a real world evidence cohort, STAT5 A PM was associated with a shorter time to next treatment (TTNT) among ICPI-treated patients with SCC. Median TTNT for STAT5A PM+ patients (n=16) was 8.6 mo (95% CI, 4.6-10.4) versus 30.65 mo (95% CI, 25.6-38.1) for PM- patients (n=799). A log-rank test indicated significant separation between monotherapy ICPI treated PM+ and PM- groups (p = 0.01). In a Cox proportional hazards model, PM+ was associated with shorter TTNT (HR = 2.10; 95% CI 1.08-4.09; p=0.03), indicating that PM+ patients transition more rapidly to subsequent therapy or progression.

[0240] When analyzed by cancer type, STAT5A PM+ prevalence was low in both subtypes: 13 of 745 LSCC (1.7%) and 3 of 70 HNSCC (4.3%). The association between STAT5A PM and TTNT was primarily driven by LSCC (HR = 2.33; p = 0.01, Table 1). The limited number of PM+ HNSCC cases (n=3) precluded statistical significance in that subgroup.

[0241] Table 1. Hazards Ratio for Lung Squamous Cell CarcinomasAttorney Docket No.: GH0259WO

[0242] Optimized detection of STAT5A PM using the combined genomic, epigenomic platform demonstrates high analytical performance and concordance with external tissue prevalence. These findings support the investigation of STAT5A PM as a clinically relevant biomarker associated with reduced ICPI benefit in patients with LSCC, and underscore its potential utility for treatment stratification, broadening the role of methylation-based profiling in precision oncology.Example 5 - MSI enhancement using methylation

[0243] Further described herein is the enhancement of microsatellite instability (MSI) calling by leveraging promoter methylation, specifically targeting MLH1 promoter methylation. One can generate a more informed and accurate interpretation of MLH1 status, which could improve diagnostic sensitivity and specificity.

[0244] In particular, one can apply a monotonic gated classifier, wherein lologistic modeling is utilized to weigh MSS / low — MSI-H. Fixed global weights are applied where PM’s effect is constant and does not vary with MSI, although one of skill readily appreciate that additional interaction terms can be introduced.

[0245] Interestingly, status can flip with MLH1 PM alone. Here, complex interaction terms to modify relationship depending on feature space. Gated, monotonic classifier with MSI as driver, make PM’s weight a function of MSI. In this regard, the relationship is monotonic in MSI (higher MSI can only increase the MSI-H probability). This prevents incidental “flipping” obviously MSI-H or clearly MSS cases due to noise.

[0246] By applying a Gaussian gate around MSI threshold, PM should influence the call only when the MSI score is near the decision boundary. Far from the boundary (very low or very high MSI), PM has little to no effect. This prevents “MSS — MSI-H” conversions driven by methylation alone. Thus, by providing a MSI floor, MSI samples below this floor are never reverted regardless of PM amplitude.Attorney Docket No.: GH0259WOExample 6 - BRCA1 zygosity

[0247] BRCA1 promoter methylation (PM) is an early initiating event in cancer, occurring in 3 to 65.2% of all breast tumors depending on subtype, and 30 to 65% of triple negative tumors. BRCA1 promoter methylation has been associated with defective homologous recombination repair (HRR), early onset of breast and ovarian cancer, and improved clinical response to adjuvant chemotherapy. To date, there has been no diagnostic assay that comprehensively evaluates both BRCA1 promoter methylation and genomic alterations in cell-free circulating tumor DNA (ctDNA).

[0248] BRCA1 biallelic loss (both alleles inactivated) — often through homozygous deletion or a combination of loss of function mutation, like a frameshift or structural deletion and loss of the wild-type allele — results in a complete loss of BRCA1 function, rendering the cell highly sensitive to PARP inhibitors (PARPi) due to impaired homologous recombination repair. In contrast, monoallelic (single-allele) loss or partial dysfunction generally retains some repair capacity and may confer reduced PARPi sensitivity. BRCA1 promoter methylation can act as a “second hit,” silencing expression of the remaining functional wildtype allele, thus mimicking biallelic loss and increasing PARPi sensitivity. However, if methylation is heterogeneous or reversible, it may lead to incomplete BRCA1 silencing and variable PARPi response.

Claims

Attorney Docket No.: GH0259WOTHE CLAIMS1. A method, comprising: detecting methylation in one or more promoter regions of at least one of a plurality of genes; determining a ratio of the number of molecules in a region normalized by total positive control molecules; and generating a plurality of methylation calls in a region based on the ratio to quantify methylation of the one or more promoter regions, wherein the promoter comprises a region of 5kb upstream of the transcription start site (TSS).

2. The method of claim 1, comprising obtaining a sample.

3. The method of claim 1, comprising having obtained a sample.

4. The method of claim 1, wherein the 5kb region is further refined using one or more of: costume panel regions, methylation peaks found in clinical samples, and excluding peaks found in normal samples.

5. The method of claim 1, wherein the TSS is defined at the transcript level.

6. The method of claim 1, wherein the TSS is defined at the gene level.

7. The method of claim 1, comprising determining the ratio of the number of molecules that overlap a target region normalized by total positive control molecules.

8. The method of claim 1, wherein determining the ratio comprises filtering of a molecule based at least on the number of overlapping CpGs.

9. The method of claim 1, wherein the quantifying of methylation of the one or more promoter regions is based on the number of methylated CpGs.

10. The method of claim 1, comprising refining the one or more promoter regions based at least on literature annotations, common methylation peak positions, and / or public datasets.

11. The method of claim 1, comprising comparing to a minimum methylation threshold derived from a population of training samples.

12. The method of claim 11, wherein the training samples comprise cancer-free samples.

13. The method of claim 11, wherein the minimum methylation threshold for calling comprises at least one of: a minimum molecule count of 1-100 and a minimum methylation score per gene is the max of: 95 quantile in normal + 8X105 or Median + 5 * median absolute deviation.

14. A method comprising:Attorney Docket No.: GH0259WO determining promoter regions of at least one of a plurality of genes, each obtained from a plurality of samples; determining methylation scores for the promoter regions to generate a plurality of methylation calls and / or quantification of promoter methylation; processing the plurality of methylation calls to generate a prediction that a test sample exhibits a genomic state.

15. A method, comprising: obtaining, by a computing system having one or more hardware processors and memory, sequencing reads derived from a sample of a subject, determining one or more classification regions corresponding to a plurality of genes included in the sample; and determine a methylation level of the one or more classification regions by generating a quantitative measure derived from the sequencing reads in the sample of the subject.

16. The method of claim 15, comprising obtaining a sample.

17. The method of claim 15, comprising having obtained a sample.

18. The method of claim 15, comprising processing the methylation level of the one or more classification regions to characterize the sample.

19. The method of any preceding claim, wherein the quantitative measure comprises determining the ratio of the number of molecules that overlap a classification region normalized by total positive control molecules, wherein the molecules exhibit a threshold amount of methylated cytosines.

20. The method of any preceding claim, wherein the quantitativee measure is compared to a predetermined threshold value to call methylation status of the one or more classification regions.

21. The method of any preceding claim, wherein determining the ratio comprises filtering of a molecule based at least on a threshold amount of methylated cytosines.

22. The method of any preceding claim, wherein the determine a methylation level of the one or more classification regions is based on the number of methylated CpGs.

23. The method of any one of the preceding claims, wherein the classification regions comprise promoter regions.

24. The method of any preceding claim, wherein the one or more classification regions individually correspond to genomic regions in which a methylation rate of cytosines in the genomic regions of nucleic acids derived from cells obtained from subjects in which cancer is present is different from a methylation rate of cytosines in the genomic regionsAttorney Docket No.: GH0259WO of nucleic acids derived from cells obtained from subjects in which cancer is not present.

25. The method of any preceding claim, wherein the plurality of samples and the additional sample include cell free nucleic acids.

26. The method of any preceding claim, comprising: performing, by the computing system, a training process using the training data to generate the model, wherein the training process includes: determining, by the computing system, one or more additional weights of individual samples included in the training data based on the indication of cancer for the individual samples being within a threshold confidence level.

27. The method of any preceding claim, wherein the indication of cancer for an individual sample is outside of the threshold confidence level and the method comprises: applying, by the computing system, a penalty to a weight of the individual sample during the training process.

28. The method of any preceding claim, comprising: performing, by the computing system and using the one or more machine learning algorithms, one or more first iterations of the training process for the model using a portion of the training data; and generating, by the computing system, first output data for the model based on the one or more first iterations of the training process, the first output data corresponding to one or more first additional indications of cancer being present in first individual subjects of the plurality of subjects, the first individual subjects corresponding to the portion of the training data.

29. The method of any preceding claim, comprising: combining, by the computing system, the first output data and the training data to produce additional training data; performing, by the computing system, one or more second iterations of the training process for the model using a portion of the additional training data; and generating, by the computing system, second output data for the model based on the one or more second iterations of the training process, the second output data indicating one or more second additional indications of cancer being present in second individual subjects of the plurality of subjects, the second individual subjects corresponding to the portion of the additional training data.Attorney Docket No.: GH0259WO30. The method of any preceding claim, wherein the weights for the individual classification regions of the plurality of classification regions are determined based on the first output data and the second output data.

31. The method of any preceding claim, comprising: determining, by the computing system, that a number of indications of cancer being present that were determined during one or more iterations of the training process are at least a threshold value for one or more samples included in the training data; and determining, by the computing system, that modifications to one or more weights of the model are not modified or are modified by a minimal amount.

32. The method of any preceding claim, comprising: determining, by the computing system, that an additional number of indications of cancer being present that were determined during the one or more iterations of the training process are less than the threshold value for one or more additional samples included in the training data; and determining, by the computing system, that modifications to one or more additional weights of the model are modified by more than the minimal amount.

33. The method of any preceding claim, comprising: combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleic acid-MBD protein solution; and performing a plurality of washes of the nucleic acid-MBD protein solution with a salt solution to produce a number of nucleic acid fractions, individual nucleic acid fractions having a threshold number of methylated cytosines in regions of the plurality of nucleic acids having at least the threshold cytosine-guanine content.

34. The method of claim 20, wherein a wash of the plurality of washes is performed with a solution having a concentration of sodium chloride (NaCl) and produces a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins.

35. The method of any preceding claim, comprising: determining that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding strengths to MBD proteins;Attorney Docket No.: GH0259WO attaching a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode being included in a first set of molecular barcodes associated with the first partition; determining that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding energies to MBD proteins different from the first range of binding strengths to MBD proteins; and attaching a second molecular barcode to nucleic acids of the second nucleic acid fraction, the second molecular barcode being included in a second set of molecular barcodes associated with the second partition.

36. The method of any preceding claim, comprising: combining at least a portion of the number of nucleic acid fractions with an amount of restriction enzyme that cleaves molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads, wherein the threshold amount of methylated cytosines corresponds to a minimum frequency of methylated cytosines within a region having at least the threshold cytosine- guanine content.

37. A method, comprising: obtaining, by a computing system having one or more hardware processors and memory, sequencing reads derived from a sample of a subject, determining one or more classification regions corresponding to a plurality of genes included in the sample, determine a methylation level of the one or more classification regions by generating a quantitative measure comprising the ratio of the number of molecules that overlap a classification region normalized by total positive control molecules, wherein the molecules exhibit a threshold amount of methylated cytosines; and comparing the quantitate measure to a predetermined threshold value to call methylation status of the one or more classification regions.

38. The method of any preceding claim wherein quantifying methylation of the one or more promoter regions is predictive of therapy response.

39. The method of any preceding claim, wherein a plurality of quantitative measure weighted in a logistic regression model, optionally included fixed global weights.

40. The method of claim 39, wherein the logistic regression model is a gated classifier, optionally a monotonic gated classifier.Attorney Docket No.: GH0259WO41. The method of any preceding claim, wherein a status is a driver, and weighting is a function of the status, optionally applied for noise abatement.

42. The method of any preceding claim, comprising a gate around the threshold, optionally comprising a probabilistic function.

43. The method of any preceding claim, wherein quantifying methylation of the one or more promoter regions is combined with status, optionally MSI, HRD, MMR, TMB, or other cancer status.

44. The method of any preceding claim, wherein predicting therapy response comprises use of a classifier, optionally including a gated classifier, and threshold in causing selection of a therapy for administration selected from one or more of an immune checkpoint inhibitor, poly (ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor.

Citation Information

Patent Citations

  • Methods for accurate sequence data and modified base position determination

    US8486630B2

  • Systems and methods to detect rare mutations and copy number variation

    WO2014039556A1

  • Methods and systems for analyzing nucleic acid molecules

    WO2018119452A2

  • Detecting the presence of a tumor based on methylation status of cell-free nucleic acid molecules

    WO2024211717A1

  • Compositions and methods for isolating cell-free DNA

    WO2020160414A1