Methods for preparing and evaluating methylated nucleic acid molecules

The method enriches and analyzes cell-free DNA using oligonucleotide probes and discrimination agents to improve the sensitivity and specificity of methylation detection, addressing the limitations of existing technologies in early disease detection.

WO2026072590A1PCT designated stage Publication Date: 2026-04-02REITER JOHANNES G +8
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for detecting diseases or conditions through nucleic acid methylation analysis suffer from high molecular loss, technical noise, and low specificity and sensitivity due to chemical or enzymatic conversion processes, making it challenging to accurately detect early-stage diseases using minimally-invasive samples.

Method used

A method involving the extraction of cell-free DNA from a sample, followed by enrichment using oligonucleotide probes designed to hybridize to differentially methylated target regions with a tiling density of 0.9x to 2x, and treatment with agents to discriminate between methylated and unmethylated cytosines, followed by sequencing to determine methylation status.

Benefits of technology

This approach enhances the sensitivity and specificity of methylation analysis, allowing for early detection of diseases by reducing noise and improving the detection of methylation patterns in small amounts of nucleic acid molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000086_0001
    Figure IMGF000086_0001
  • Figure IMGF000088_0001
    Figure IMGF000088_0001
  • Figure IMGF000088_0002
    Figure IMGF000088_0002
Patent Text Reader

Abstract

The present disclosure provides methods, compositions, and kits for preparing preparations of DNA from a subject useful for determining a methylation status of a genomic region of interest.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS FOR PREPARING AND EVALUATING METHYLATED NUCLEIC ACID MOLECULES CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefits of U.S. Provisional Application No.63 / 698,440 filed September 24, 2024, the contents of which are hereby incorporated by reference in their entirety. TECHNICAL FIELD

[0002] The present disclosure relates generally to methods for preparing and evaluating nucleic acid molecules, such as deoxyribonucleic acid (DNA) molecules, and, more particularly, to methods for preparing and analyzing nucleic acid molecules that include one or more methylation sites. BACKGROUND

[0003] Detection of diseases or recurrence of diseases at earlier stages of progression can substantially improve patient outcome. Generally, less invasive tests for diseases and other conditions tend to be administered more often than more invasive tests, in part due to risk and compliance considerations. Accordingly, improvements in the ability of non-invasive (or minimally-invasive) tests increase the chances that such tests will be more routinely used, and that anomalous conditions will be detected more often and at earlier stages. For example, a test that requires a blood draw or cheek swab is generally less risky for patients than procedures that require a surgical procedure (e.g., a biopsy), and such a test would thus generally be preferable for patients and clinicians. However, physiological manifestations of a condition might be subtle, especially at earlier stages of a disease, making them more difficult to detect using samples obtained through less-invasive means.

[0004] Methylation is responsible for how cells differentiate and determines tissue types and disease state in many eukaryotic organisms. Gene methylation can silence signal pathways (such as tumor suppressor pathways). Genome-wide hypomethylation and focal hypermethylation of CpG blocks or islands located in gene promoters or gene bodies are correlated with aging and cancer. For example, between 0.2% to 80% of the human genome may be differentiallymethylated in cancers, and about 0.5-30% of the colorectal cancer (CRC) genome may be differentially methylated, some of which are recurrently differentially methylated with weak functional consequences; whereas only about 0.0003% of the CRC genome harbors somatic point mutations, most of which are not recurrent. Differential methylation signals are thus useful biomarkers for early detection of a disease or condition or recurrence of the disease or condition due to their frequency and homogeneity across patients and subpopulations of patients. However, methylation patterns are subject to higher technical noise. Chemical or enzymatic conversion process for distinguishing methylation status of a nucleotide damages DNA, leading to high molecular loss (typically 80-95%) and limiting effective sequencing depth, typically to about 1000x. Conversion and protection errors also contribute to the background noise, resulting in decreased specificity and / or sensitivity. For example, for hypermethylated regions, conversion errors increase background noise; for hypomethylated regions, protection errors increase background noise. In addition, biological noise from other cancers, diseases, or comorbidities (e.g., clonal hematopoiesis (CH), aging), for example, makes it challenging to achieve the desired specificity for the cancer, disease, or condition being evaluated or monitored.

[0005] Accordingly, there is a need for improved methods for preparing and analyzing small amounts of nucleic acid molecules in samples, such as circulating free DNA (cfDNA), carrying epigenetic alterations, such as changes in methylation status or patterns, useful for earl detection of diseases or conditions or their recurrence. There is a need for sensitive and efficient methods for preparing and quantifying small amounts of differentially methylated nucleic acid molecules of interest in a sample. SUMMARY

[0006] To overcome the above-mentioned and additional problems in the art, the present disclosure provides methods for preparing a preparation of nucleic acid molecules (such as DNA) from a subject useful for determining a methylation status of a genomic region of interest by enriching and evaluating methylated target loci.

[0007] Provided herein in one aspect is a method for preparing a composition of non-naturally occurring DNA, comprising: (a) extracting cell-free DNA molecules from a sample of a subject;and (b) contacting the extracted cell-free DNA molecules or their derivative with a panel of oligonucleotide probes designed to hybridize to 5-200,000 different target regions having CpG sites that are differentially methylated in a phenotype, thereby generating enriched DNA, wherein the oligonucleotide probes for at least one of the target regions are designed to cover the target region with a tiling density of between 0.9x and 2x.

[0008] In some embodiments of the methods described herein, the method further comprises treating (i) the extracted cell-free DNA, or (ii) the enriched DNA or its derivative, with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines.

[0009] In some embodiments of the methods described herein, the panel of oligonucleotide probes are designed to cover at least 50%, at least 80%, at least 90%, or at least 95% of the target regions with a tiling density of between 0.9x and 2x. In some embodiments, the panel of oligonucleotide probes are designed to cover at least 50%, at least 80%, at least 90%, or at least 95% of the target regions with a tiling density of 1x to 1.4x. In some embodiments, the panel of oligonucleotide probes are designed to cover at least 50%, at least 80%, at least 90%, or at least 95% of the target regions with a tiling density of 1x.

[0010] In some embodiments of the methods described herein, the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions are tiled at less than 25 bp, less than 20 bp, less than 15 bp, less than 10 bp, less than 5 bp. In some embodiments, the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for both top and bottom DNA strands. In some embodiments, the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for DNA strands that are fully unmethylated. In some embodiments, the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for DNA strands methylated at different levels. In some embodiments, the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for both fully methylated and fully unmethylated DNA strands.

[0011]

[0012] In some embodiments of the methods described herein, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 5 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 10 CpG sites. In some embodiments of the methods described herein, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 20 CpG sites.

[0013] In some embodiments of the methods described herein, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.02 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.03 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.04 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.05 CpG / bp.

[0014] In some embodiments of the methods described herein, the target regions have a combined footprint of 800 bp to 30 Mbp. In some embodiments, the target regions have a combined footprint of at least 15 Kb. In some embodiments, the target regions have a combined footprint of no more than 500 Kb.

[0015] In some embodiments of the methods described herein, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are selected based on whole genome sequencing of treated DNA samples from a population of subjects having a phenotype. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are selected based on targeted sequencing of treated DNA samples from a population of subjects having a phenotype. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are selected based on sequencing of one or more CpG islands in treated DNA samples from a population of subjects having a phenotype.

[0016] In some embodiments of the methods described herein, at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions selected comprise a linear association of methylation signal with VAF. In some embodiments of the methods described herein, at least80%, or at least 90%, or at least 95% of the target regions are unmethylated in subjects not having a phenotype.

[0017] In some embodiments of the methods described herein, the methos further comprises appending adapters to the extracted DNA before contacting the extracted cell-free DNA molecules or their derivative with the panel of oligonucleotide probes. In some embodiments, the adapters are Y-adaptors. In some embodiments, the adapters include a universal priming site. In some embodiments, the adapters include one or more methylated cytosines. In some embodiments, the adapters include one or more unmethylated cytosines. In some embodiments, the adapters include both methylated and unmethylated cytosines. In some embodiments, the adapters include a molecular barcode.

[0018] In some embodiments of the methods described herein, the method further comprises amplifying the enriched and / or treated DNA. In some embodiments, amplifying comprises targeted PCR. In some embodiments, amplifying comprises universal PCR.

[0019] In some embodiments of the methods described herein, the agent or the combination of agents comprises a deaminating agent. In some embodiments, the deaminating agent comprises sodium bisulfite. In some embodiments, the deaminating agent comprises a deaminase or its catalytic domain. In some embodiments, the combination of agents further comprises an oxidizing agent that oxidizes 5hmC and 5mC but not unmethylated cytosines thereby protecting 5hmC and 5mC from deamination. In some embodiments, the combination of agents further comprises a glycosylation agent that glycosylates 5hmC but not unmethylated cytosines thereby protecting 5hmC from deamination.

[0020] In some embodiments of the methods described herein, the agent or the combination of agents comprises one or more methylation sensitive restriction enzymes (MSREs) or methylation dependent restriction enzymes (MDREs). In some embodiments, the one or more MSREs is one or more of the MSRE is selected from HpaII, SalI, BbeI, NotI, SmaI, XmaI, MboI, BstUI, BstBI, ClaI, MluI, NaeI, NarI, PvuI, SacII, HpyCH41V, and HhaI. In some embodiments, the one or more MDREs is selected from AbaSI, AoxI, BisI, BlsI, DpnI, FspEI, GlaI, GluI, KroI, LpnPI, MalI, MspJI, MteI, PcsI, PkrI, and SgeI.

[0021] In some embodiments of the methods described herein, the agent or the combination of agents comprises a methylation binding domain (MBD) protein, a 5-methylcytosine antibody, a 5- hmC antibody, or a DNA methyltransferase.

[0022] In some embodiments of the methods described herein, the method further comprises performing a barcoding PCR on the treated DNA to add a sample barcode before the contacting step.

[0023] In some embodiments of the methods described herein, the method further comprises spiking in methylated pUC19 DNA and unmethylated Lambda DNA as control.

[0024] In some embodiments of the methods described herein, the panel of oligonucleotide probes comprises a plurality of primers and the contacting step comprises amplifying the treated DNA or its derivative using the plurality of primers to generate the enriched treated DNA. In some embodiments, the panel of oligonucleotide probes comprises a plurality of baits and the contacting step comprises hybridization capture of the extracted DNA or its derivative using the plurality of baits to generate the enriched treated DNA. In some embodiments, the treated DNA or its derivative is captured using an affinity-based capture method. In some embodiments, the capture method comprises click chemistry-based affinity capture. In some embodiments, the plurality of baits each includes a capture moiety comprising biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, or desthiobiotin-TEG, or biotin azide. In some embodiments, the treated DNA or its derivative is captured using a solid phase coated with streptavidin.

[0025] In some embodiments of the methods described herein, the method further comprises performing high-throughput sequencing on the treated DNA or its derivative to generate sequence reads. In some embodiments, the method further comprises determining a methylation status of one or more of the cell-free DNA molecules or one or more of the target regions based on the sequence reads. In some embodiments, determining the methylation status comprises determining a percentage of methylated CpG sites within one or more of the cell-free DNA molecules or one or more of the target regions based on the sequence reads.

[0026] In some embodiments of the methods described herein, the method further comprises filtering out sequence reads derived from treated DNA with incomplete conversion. In someembodiments, the filtering comprises filtering out sequence reads derived from cell-free DNA molecules having 1 or more CHG and / or CHH sites that appear to be methylated.

[0027] In some embodiments of the methods described herein, the sample is a liquid sample. In some embodiments, the liquid sample is a blood, serum, plasma, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample. In some embodiments, the liquid sample is a blood, plasma, serum, or urine sample. In some embodiments, the sample is a blood sample or a derivative sample thereof. In some embodiments, the sample is a plasma sample. In some embodiments, the sample comprises DNA from a tumor. In some embodiments, the sample is a sample from a cancer tissue.

[0028] In some embodiments of the methods described herein, the subject has, had, or is suspected or at risk of having a phenotype. In some embodiments, the phenotype includes age, gender, ethnicity, smoking status, or pregnancy status. In some embodiments, the phenotype is a disease. In some embodiments, the disease is a cancer.

[0029] In some embodiments, the cancer is selected from ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-celllymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.

[0030] In some embodiments, the cancer is selected from a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro- esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whipple resection.

[0031] In some embodiments, the cancer is selected from lung cancer, breast cancer, bladder cancer, and colorectal cancer. In some embodiments, the breast cancer is HER2 positive breast cancer, hormone receptor (HR) positive breast cancer, or triple negative breast cancer (TNBC). In some embodiments, the lung cancer is non-small cell lung cancer (NSCLC), adenocarcinoma NSCLC, squamous NSCLC, or small cell lung cancer (SCLC). In some embodiments, the bladder cancer is muscle invasive bladder cancer (MIBC).

[0032] In some embodiments of the methods described herein, the methylation status of one or more of the cell-free DNA molecules or one or more of the target regions is indicative of the presence or absence of the cancer.

[0033] In some embodiments of the methods described herein, the sample comprises DNA from a transplanted organ. In some embodiments, the subject is a transplant recipient.

[0034] In some embodiments of the methods described herein, the sample comprises DNA from a fetus. In some embodiments, the subject is a pregnant female.

[0035] In some embodiments of the methods described herein, the subject is an animal. In some embodiments, the subject is a domesticated animal. In some embodiments, the subject is a dog, a cat, a hamster, a pig, a cattle, a sheep, a goat, a horse, a deer, a buffalo, a monkey, a mouse, a rat, a guinea pig, a llama, a cow, a fish, a reptile.

[0036] In some embodiments of the methods described herein, the method further comprises enriching the sample DNA molecules for DNA molecules that are between 70 and 500 base pairs in length. In some embodiments, the method further comprises enriching the sample DNA molecules for DNA molecules that are between 100 and 200 base pairs in length. In some embodiments, the method further comprises enriching the sample DNA molecules for DNA molecules that are between 130 and 170 base pairs in length.

[0037] In some embodiments of the methods described herein, one or more steps of the method are performed in a microfluidic device.

[0038] In some embodiments of the methods described herein, the method further comprises performing the method on a second sample from the subject collected at a different time point.

[0039] In some embodiments of the methods described herein, the method further comprises calculating a differential methylation allele fraction (DMAF). In some embodiments, the method is capable of detecting a DMAF of a target region.

[0040] Provided herein in another aspect is a method of preparing a composition of non-naturally occurring DNA, comprising: (a) extracting DNA molecules from a sample of a subject; (b) contacting the extracted DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to 5-200,000 different target regions differentially methylated in a phenotype, thereby generating enriched treated DNA, wherein the oligonucleotide probes for at least one of the target regions are designed to cover the target region with a tiling density of between 0.9x and 2x, wherein at least 90% of the 5-200,000 different target regions each includes at least 3 CpG sitesand a CpG density of at least 0.02 CpG / bp; and (c) generating sequence reads by performing high- throughput sequencing on the enriched treated DNA or its derivative.

[0041] In some embodiments of the methods described herein, the method further comprises treating (i) the extracted DNA or its derivative, or (ii) the enriched DNA or its derivative, with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines.

[0042] In some embodiments of the methods described herein, the method further comprises determining a methylation status of one or more of the extracted DNA molecules or a portion thereof or one or more of the target regions based on the sequence reads.

[0043] In some embodiments of the methods described herein, determining the methylation status comprises determining a percentage of methylated CpG sites within one or more of the extracted DNA molecules or a portion thereof or one or more of the target regions based on the sequence reads. In some embodiments, the methylation status is quantified on a fragment level based on the co-methylation status of all CpGs.

[0044] In some embodiments of the methods described herein, the method further comprises identifying a sequence read mapped to a target region and having at least 3 CpG sites and a methylation fraction of at least 0.8 as a hypermethylated fragment, and determining a hypermethylated fragment fraction for the target region based on the number of hypermethylated fragments and the total number of sequence reads mapped to the target region and having at least 3 CpG sites. In some embodiments, the method further comprises filtering out sequence reads derived from treated DNA with incomplete conversion. In some embodiments, the filtering comprises filtering out sequence reads derived from extracted DNA molecules having 1 or more CHG and / or CHH sites that appear to be methylated. In some embodiments, the methylation status is determined using only sequence reads having one or more CHG or CHH sites.

[0045] In some embodiments of the methods described herein, the method further comprises classifying the sample as having or not having the phenotype / disease / cancer using a model-free algorithm. In some embodiments, the method further comprises classifying the sample as having or not having the phenotype / disease / cancer using a logistic regression classifier. In someembodiments, the classifier selects one or more targets based on a training set during cross- validation. In some embodiments, the classifier includes a transformed version of hypermethylated fraction of each selected target. In some embodiments, the classifier estimates posterior probability of a phenotype / disease / cancer-positive status.

[0046] Provided herein in another aspect is a method of preparing a composition of non-naturally occurring DNA, comprising: (a) extracting cell-free DNA molecules from a sample of a subject; (b) treating the extracted cell-free DNA molecules or their derivative with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; (c) contacting the treated DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to 5-200,000 different target regions having CpG sites that are differentially methylated in a phenotype, thereby generating enriched treated DNA, wherein the oligonucleotide probes for at least one of the target regions are designed to cover the target region with a tiling density of between 0.9x and 2x.

[0047] In some embodiments of the methods described herein, the method further comprises classifying the sample based on both the methylation status of the cell-free DNA molecules or their derivatives and one or more additional signals. In some embodiments, the one or more additional signals comprise fragmentomics, genomic alterations, proteins, extracellular vesicles, lipids, miRNA, or immune profiling. In some embodiments, the method further comprises classifying the sample based on one or more additional characteristics of the subject. In some embodiments, the one or more additional characteristics comprise age, gender, ethnicity, smoking status, polygenic risk score, or medical history.

[0048] Provided herein in another aspect is a method comprising: a) obtaining a first plurality of samples from healthy subjects and a second plurality of samples from subjects with a phenotype; b) generating a first set of sequencing data for the first plurality of samples and a second set of sequencing data for the second plurality of samples using a sequencing panel comprising 5- 200,000 different target regions having CpG sites that are differentially methylated in the phenotype; c) selecting, for subsequent use in detecting the phenotype in a subject, a subset of the target regions for the phenotype, wherein the subset of the target regions: i) are hypermethylated inthe sequencing data for the second plurality of samples; and ii) exhibits a linear relationship between the methylation signal and VAF.

[0049] In some embodiments of the methods described herein, the subset of the target regions from a panel of target regions are selected to have a linear relationship between differential methylation allele fraction (DMAF) as determined by methylation signals and variant allele frequence (VAF) as determined by genomic sequences, wherein the Pearson correlation between the DMAF and the VAF is at least 0.8, at least 0.85, at least 0.9, or at least 0.95.

[0050] In some embodiments, a sample is classified as having or not having a phenotype based on DMAF of a subset of the target regions selected from a panel of target regions having CpG sites that are differentially methylated in the phenotype by any of the methods described herein.

[0051] Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, and kits or functional elements therein across sections. Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, or other functional elements therein across sections. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] FIG.1 shows an exemplary, non-limiting workflow for preparing and enriching differentially methylated DNA molecules of interest.

[0053] FIG.2 shows a plot of some of the regions differentially methylated between healthy (blue) and CRC (red) cohorts identified using a machine learning model developed to discover differentially-methylated regions between healthy and disease cohorts as well as between subpopulations within the disease cohort, including age, sex, stage, histology, and MSI vs microsatellite stability (MSS), according to various embodiments. The plots illustrate a segmentation procedure using the machine learning model. For a region, FIG.3 plots the methylation load of the healthy cohort and tumor cohort and demonstrates the segmentationprocedure and scoring / ranking. The horizontal bars in the lower panel show the differential methylation load.

[0054] FIG.3 shows the characteristics of a preliminary targeted panel.

[0055] FIG.4 shows sequencing metrics for 1x and 2x probe tiling designs. Subsampled reads of 15M read pair were used.

[0056] FIG.5 shows stage and variant allele frequency (VAF) distribution of MRD-positive and MRD-negative samples from the Bespoke clinical trial used for the methylation performance study. (A) shows number of samples at each stage of CRC; (B) shows VAF distribution of the samples.

[0057] FIG.6 shows the distribution of (A) the mean per-Target Differentially-Methylated Fragments Fraction (TDMF) baseline (across healthy samples), (B) mean TDMF disease (across cancer samples), and (C) LR+ of hypermethylated targets in a panel.

[0058] FIG.7 shows a feature matrix for sample classification.

[0059] FIG.8(A) shows the distribution of differential methylation allele fraction (DMAF) as a function of input DNA amount. FIG.8(B) shows the distribution of DMAF as a function of total number of suitable fragments in selected targets.

[0060] FIG.9 shows correlation of DMAF with VAF and mean tumor molecules (MTM). (A) shows VAF vs. DMAF, (B) shows VAF vs. DMAF with relapse, and (C) shows MTM vs. DMAF.

[0061] FIG.10 shows VAF vs. DMAF correlations for MSI vs. MSS cancers. (A) shows correlation for MSI cancers and (B) shows correlation for MSS cancers.

[0062] FIG.11 shows VAF vs. DMAF correlations for colon vs. rectal cancers. (A) shows correlation for colon cancer and (B) shows correlation for rectal cancers.

[0063] FIG.12 shows VAF vs. DMAF correlation for (A) stage II, (B) stage II, and (C) stage IV cancers.

[0064] FIG.13 shows target level analysis for MSI vs. MSS cancers. (A) shows a Quantile- Quantile plot for p-values and (B) is a Volcano plot showing log fold change and p-value.

[0065] FIG.14 shows an overview of a cross-validation approach.

[0066] FIG.15 shows classification performance of a machine learning classifier tuned to 95% specificity. (A) shows the ROC curve of the classifier, (B) shows the distribution of sensitivity and specificity across the cross-validation folds, and (C) shows Signatera VAF vs. DMAF plot annotated by classification accuracy.

[0067] FIG.16 shows the details of another target panel including (A) percentage of targets by methylation status and (B) percentage of targets by CpG regions.

[0068] FIG.17 shows (A) tumor stage and (B)-(D) median target CpG coverage distribution after sample QC.

[0069] FIG.18 shows TDMF baseline mean comparison between ECD and MRD samples at (A) <100%, (B), <1%, (C) <0.1%, and (D) <10-3%.

[0070] FIG.19 shows (A) TDMF disease mean and (B) LR+ comparison between ECD and MRD samples.

[0071] FIG.20 shows cross-validation (CV) performance of a machine learning model based on best parameters selected by a nested grid-search. (A) shows ROC curve of the CV models on validation sets, and (B) shows 5-fold CV performance (20 repeats).

[0072] FIG.21 shows the performance of an MRD-trained model on ECD samples. (A) shows ROC curve of the CV models on validation sets, and (B) shows 5-fold CV performance (20 repeats).

[0073] FIG.22 shows the performance of an ECD-trained classifier on MRD samples. (A) shows ROC curve of the CV models on validation sets, and (B) shows 5-fold CV performance (20 repeats).

[0074] FIG.23 shows (A)-(B) correlation of DMAF with Signatera VAF, and (C) DMAF of the Signatera negatives.

[0075] FIG.24 shows distribution of logistic regression posterior probabilities of being MRD- positive on out-of-fold samples before (A, B) and after (C, D) learned posterior transformation.

[0076] FIG.25 shows performance improvement after combining hypermethylated fragment fraction (HFF) and posterior transformations. (A)-(B) shows distribution of logistic regression posteriors on out-of-fold samples after the transformations. (C)-(F) shows the performance of the classifier.

[0077] FIG.26 shows performance at different levels of tumor VAF (Signatera VAF subgroups) vs. target (desired) specificity. Thick lines represent overall performance, with the shaded blue region representing the 99% confidence interval for overall accuracy.

[0078] FIG.27 illustrates the characteristics of one of the candidate target evaluated using the methods herein. X-axis shows a segment of the human genome. The bracket indicates the intended target region designed using the methods herein. Y-axis is the posterior probability of methylation status.

[0079] FIG.28 illustrates the characteristics of another example target evaluated using the methods herein. X-axis shows a segment of the human genome. The bracket indicates the intended target region designed using the methods herein. Y-axis is the posterior probability of methylation status.

[0080] FIG.29 illustrates the characteristics of yet another example target evaluated using the methods herein. X-axis shows a segment of the human genome. The bracket indicates the intended target region designed using the methods herein. Y-axis is the posterior probability of methylation status.

[0081] FIG.30 shows the mean level of the background noise in healthy samples (Y-axis) and associated confidence intervals for all the targets selected at some time by the classifier (X-axis) in an illustrative embodiment using a cross-validation (CV) model, ordered by selection frequency.

[0082] FIG.31 shows DMAF and positivity rate by sample category of Cohort 1 and Cohort 2 samples.

[0083] FIG.32 shows DMAF comparison with and without CHH / CHG filter.

[0084] FIG.33 shows DMAF and positivity rate by sample category of Cohort 2 samples with CHH / CHG filter.

[0085] FIG.34 shows (A) ROC curve of the MRD CV models retrained with mCHH / CHG filter on ECD validation sets and (B) DMAF correlation with Signatera VAF and classification accuracy.

[0086] FIG.35 shows (A) ROC curve of the CV models on validation sets and (B) 5-fold CV performance (20 repeats) using the mCHH / CHG filtered data.

[0087] FIG.36 shows DMAF correlation with Signatera VAF for ECD-trained model on MRD data.

[0088] FIG.37 shows a method of methylation-based ctDNA quantification based on independent target selection and DMAF estimation. Independent target selection prioritizes regions with low background hypermethylation (TDMF baseline) in normal plasma samples, that are differentially hypermethylated between tumor-positive and tumor-negative / healthy plasma samples (high signal to noise ratio; SNR), and that have a methylation signal (logit-transformed hypermethylated fragment fraction, HFF) that exhibits a linear, dosage-based relationship with Signatera VAF. Targets are ranked by F-test p-value (ascending).

[0089] FIG.38 shows performance of a model-free DMAF method for colorectal cancer using a tfMRD-CRC panel. Y-axis is mean tumor molecules (MTM) from DMAF estimation across CV repeats and X-axis is MTM from observed Signatera VAF.

[0090] FIG.39 shows performance of a model-free DMAF method for colorectal cancer using a second tfMRD panel. Y-axis is median predicted DMAF across CV repeats and X-axis shows Signatera VAF.

[0091] FIG.40 shows performance of a model-free DMAF method for breast cancer (including subtypes HR+, HER2+, and TNBC) using a tfMRD panel. Y-axis is median predicted DMAF across CV repeats and X-axis shows Signatera VAF.

[0092] FIG.41 shows performance of a model-free DMAF method for lung cancer (including subtypes SCLC and NSCLC) using a tfMRD panel. Y-axis is median predicted DMAF across CV repeats and X-axis shows Signatera VAF.

[0093] FIG.42 shows performance of a model-free DMAF method for bladder cancer using a tfMRD panel. Y-axis is median predicted DMAF across CV repeats and X-axis shows Signatera VAF.

[0094] FIG.43 shows the effect of cell-free DNA input amount (left panels) and subject age (right panels) on DMAF predications in CRC, stratified by sample category (healthy (top panels) vs. MRD positive (PosMRD, bottom panels)). cfDNA input amount groups were 10-30 ng (healthy, N=97; CRC, N=32), 31-50 ng (healthy, N=87; CRC, N=30), and 51-100 ng (healthy, N=39; CRC, N=43). Age groups were 29-50 (healthy, N=113; CRC, N=11), 51-70 (healthy, N=105; CRC, N=56), and 71-100 (healthy, N=5; CRC, N=38) years old.

[0095] FIG.44 shows the effect of cell-free DNA input amount (left panels) and subject age (right panels) on DMAF predications in breast cancer, stratified by sample category (healthy (top panels) vs. MRD positive (PosMRD, bottom panels)). cfDNA input amount groups were 10-30 ng (healthy, N=71; breast cancer, N=49), 31-50 ng (healthy, N=62; breast, N=18), and 51-100 ng (healthy, N=25; breast cancer, N=40). Age groups were 29-50 (healthy, N=81; breast cancer, N=26), 51-70 (healthy, N=76; breast cancer, N=56), and 71-100 (healthy, N=1; breast cancer, N=30) years old.

[0096] FIG.45 shows the effect of cell-free DNA input amount (left panels) and subject age (right panels) on DMAF predications in lung cancer, stratified by sample category (healthy (top panels) vs. MRD positive (PosMRD, bottom panels)). cfDNA input amount groups were 10-30 ng (healthy, N=100; lung cancer, N=39), 31-50 ng (healthy, N=86; lung cancer, N=21), and 51-100 ng (healthy, N=39; lung cancer, N=48). Age groups were 29-50 (healthy, N=115; lung cancer,N=4), 51-70 (healthy, N=105; lung cancer, N=47), and 71-100 (healthy, N=5; lung cancer, N=62) years old.

[0097] FIG.46 shows the effect of cell-free DNA input amount (left panels) and subject age (right panels) on DMAF predications in bladder cancer, stratified by sample category (healthy (top panels) vs. MRD positive (PosMRD, bottom panels)). cfDNA input amount groups were 10-30 ng (healthy, N=100; bladder cancer, N=16), 31-50 ng (healthy, N=86; bladder cancer, N=15), and 51- 100 ng (healthy, N=39; bladder cancer, N=27). Age groups were 29-50 (healthy, N=115; bladder cancer, N=1), 51-70 (healthy, N=105; bladder cancer, N=27), and 71-100 (healthy, N=5; bladder cancer, N=32) years old.

[0098] FIG.47 depicts an overall system that includes a computing system that may interface with other components to enable various aspects of the disclosed approach, according to various embodiments.

[0099] While the above-identified drawings set forth some of the presently disclosed embodiments, other embodiments are also contemplated, as noted in the discussion. This disclosure presents illustrative embodiments by way of representation and not limitation. Numerous other modifications and embodiments can be devised by those skilled in the art which fall within the scope and spirit of the principles of the presently disclosed embodiments. DETAILED DESCRIPTION

[0100] Both global and local epigenetic changes are widely regarded as a hallmark of diseases, including cancer, or other phenotypes (such as age) and their progression. In cancer, for example, alterations include a global decrease in overall CpG methylation levels coupled with discrete regions of hypermethylation, typically in CpG islands located in the promoter regions of tumor suppressor genes. Hypermethylation has been associated with cancer progression and the silencing of growth regulating genes and tumor suppressor genes. As a result, a growing number of DNA methylation biomarkers are being utilized in the development of novel assays for monitoring cancer progression, treatment response and early detection. It is possible to detect the presence of cancer through the analysis of the methylation status of specific CpG sites that arepredominantly methylated in DNA from tumor cells, including circulating tumor DNA (ctDNA) released from tumor cells.

[0101] The present disclosure addresses many long-felt needs and long-standing problems in the art, such as, but not limited to, those mentioned in the Background section herein. For example, methods are provided herein for preparing DNA molecules useful for determining a methylation status of a genomic region of interest. In some embodiments, as illustrated in FIG.1, the methods provided herein comprise steps including (a) extracting cell-free DNA molecules from a sample of a subject; and (b) contacting the extracted cell-free DNA molecules or their derivative with a panel of oligonucleotide probes designed to hybridize to target regions having CpG sites that are differentially methylated in a phenotype, thereby generating enriched DNA, wherein the oligonucleotide probes for at least one of the target regions are designed to cover the target region with a tiling density of less than 2x, between 0.9x and 2x, and in illustrative embodiments 1x. The reduced tiling density, as compared to the 2x or more tiling in conventional hybrid-capture methods for sequencing library preparation using cfDNA samples, significantly reduces costs, especially where a large number of target regions are analyzed. The methods described herein may also comprise a step of treating the extracted or enriched DNA with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines. The libraries can then be sequenced and used for downstream analyses, for example methylation analysis. The order of the steps between sample extraction and sequencing may vary when using the methods disclosed herein. In some embodiments, the treatment step is performed before adapters are appended to the extract cfDNA molecules. In some embodiments, the treatment step is performed after adapters are appended to the extract cfDNA molecules. In some embodiments, the treatment step is performed before an optional amplification step. In some embodiments, the treatment step is performed after an optional amplification step provided that the amplification method preserves the methylation status of the amplified DNA molecules. In some embodiments, the capture step is performed before the treatment step. In some embodiments, the capture step is performed after the treatment step.

[0102] In some embodiments, adapters (such as Y-adapters) are added to the extracted DNA. In some embodiments, the adapter sequences comprise methylated cytosines. In someembodiments, the adapter sequences comprise at least one methylated nucleotide (such as methylated cytosine) and at least one unmethylated nucleotide (such as unmethylated cytosine).

[0103] In some embodiments, the methods herein include performing one or more steps to sense, detect, or “read out” the methylation signals of the DNA molecules or their derivatives. In some embodiments, the methylation signals are ascertained by treating the extracted DNA molecules or their methylation-status-preserved derivatives with an agent or a combination of agents that discriminates between methylated and non-methylated cytosines. In some embodiments, the agent comprises a chemical agent (such as bisulfite) or an enzyme, or a combination of chemicals and / or enzymes, that converts unmethylated cytosine to uracil, while methylated cytosine (such as 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC)) remain as cytosine. In some embodiments, the agent comprises one or more restriction enzymes that discriminate between methylated and unmethylated CpG dinucleotides, such as methylation sensitive restriction enzymes (MSREs), which can specifically cleave unmethylated CpG, but not methylated CpG, within the enzymes’ recognition sites, and methylation dependent restriction enzymes (MDREs), which can specifically cleave methylated CpG, but not unmethylated CpG, within the enzymes’ recognition sites. The methylation status of the genomic regions of interest can then be determined by sequencing (such as any next-generation sequencing methods) of the treated DNA molecules, and in some embodiments an enriched set of treated DNA molecules, and further analyzed using methods described herein.

[0104] In addition, to address the problem of the technical and biological noises, the methods herein leverage adjacent CpGs whose methylation states are correlated, focusing on regions with high CpG density and are differentially methylated between samples having or suspected of having a cancer, disease, phenotype, condition and normal samples (samples lacking the cancer, disease, phenotype, condition) and exhibit low to zero or nearly zero biological noise. Such differentially methylated regions (DMRs) having multiple adjacent CpGs are expected to be good methylation biomarkers. Accordingly, in some embodiments, target genomic regions selected to be included in a panel include two, three, four or more nucleic acid residues that are differentially methylated in a disease (such as cancer) or phenotype. For example, in some embodiments, the target regions may have 3, 4, 5, 6, 8, 10, 15, 20, 25, 30, 35, 40, 50, 100, 150 or more nucleic acid residues that are differentially methylated in a disease (such as cancer) orphenotype. In some embodiments, the differentially methylated nucleic acid residues are cytosine residues. In some embodiments, the target regions may have 3, 4, 5, 6, 8, 10, 15, 20, 25, 30, 35, 40, 50, 100, 150 or more CpG sites that are differentially methylated in a disease or phenotype.

[0105] In some embodiments, the methods herein include a classifier for disease or phenotype detection. In some embodiments, the classifier classifies a sample based at least in part on DNA methylation profiles.

[0106] In some embodiments, the methods herein analyze methylation profiles of cfDNA from a subject for a disease or phenotype detection using a machine-learning model that employs biology-informed target and feature selection methods. For example, in some embodiments, the error model selects targets having low or ultra-low background noise in health subjects. In some embodiments, background noise for each target region is used in an initial binary filtering. In some embodiments, the error model reduces false positives using informative priors from existing biological data. In some embodiments, the model selects targets that are mutually exclusive and can indicate a false-positive if observed in the same sample. In some embodiments, the classifier includes tissue classification. In some embodiments, the model identifies fully independent targets, cross-signal targets, and / or target families and determines the joint-target noise rate due to target interactions. In some embodiments, the error model reduces false positives by filtering out DNA fragments with likely incomplete conversion or MSRE digestion of unmethylated CpG sites. In some embodiments, the likely incomplete conversion or digestion of unmethylated CpG sites on a DNA fragment is determined based on the apparent methylation status of one or more CHG and / or CHH sites (where “H” refers to A, T, or C), which are usually not methylated in mammals (e.g. humans), on the same fragment. In some embodiments, fragments having one or more CHG and / or CHH sites that appear to be methylated are filtered before further analysis.

[0107] In some embodiments, the methods herein quantify tumor DNA in a cfDNA sample from a subject for a disease or phenotype detection by quantifying methylation using a model-free method. In this method, a subset of targets from a panel are first selected based on their association with Signatura VAF, their low background hyper-methylation (TDMF baseline) in normal plasma samples, and their differential hyper-methylation between tumor-positive and tumor-negative / healthy plasma, as described in Example 7 in further detail. DMAF can then becomputed per sample as a summary metric based on the ratio of hypermethylated target regions (e.g., regions having at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 CpGs, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or higher of which being methylated) to total target regions (e.g., regions having 3, 4, 5, 6, 7, 8, 9, 10 or more CpGs regardless of methylation status).

[0108] In some embodiments, the methods herein quantify tumor DNA in a cfDNA sample from a subject for a disease or phenotype detection by quantifying methylation using a regression model. Using the targets selected as described in Example 7, the model uses a ridge regression- based method for estimating DMAF. The model is trained using healthy samples as ctDNA negative controls and known positive samples (e.g., Signatera Positive) for ctDNA-positive ground truth. This approach leverages supervised learning to optimally weight per-target methylation signals for predicting ctDNA burden, using cancer VAF (e.g., Signatera VAF) as the dependent variable during training.

[0109] In some embodiments, the classifier is a logistic regression classifier (e.g. ridge, elastic net, lasso). In some embodiments, the classifier is a supportive vector machine (SVM). In some embodiments, the classifier is a random forest classifier. In some embodiments, the classifier uses hypermethylated fragment fraction per target as features. In some embodiments, the input to the classifier is a transformed version of the hypermethylated fragment fraction for each selected target. In some embodiments, the classifier uses a grid-search model during cross-validation. In some embodiments, the classifier selects the best N targets for training the machine-learning model based on a training set during cross-validation and then makes predictions in holdout test sets. In some embodiments, the classifier and the target and feature selection methods can be further adjusted based on one or more other characteristics of the subject, such as age, gender, ethnicity, smoking status, polygenic risk score, and medical history (including for example past diagnoses, current medications, allergies, and chronic conditions).

[0110] In some embodiments, the classifier and methods disclosed herein can combine analysis of methylation signals with analysis of one or more other analytes or biomarkers (for example, genomic or somatic mutations (identified for example using targeted or whole genome sequencing), proteins, extracellular vesicles, lipids, miRNA, fragmentomics, immune profiling),either in parallel or as a reflex strategy to improve sensitivity and / or specificity in detecting a disease, phenotype, or condition. In some embodiments, the disease / phenotype detection sensitivity threshold for the initial assay may be set low to allow maximum detection, followed by a second assay using the same or a different analyte / signal to ensure specificity. In some embodiments, the classifier and methods disclosed herein combine two or more sequential measurements of one or more signals (such as methylation signal) over multiple timepoints.

[0111] The methods disclosed herein can be used to diagnose, detect, or monitor a disease (including cancer, and whether symptomatic, pre-symptomatic, or asymptomatic), phenotype, or condition in any eukaryotic organism in which methylation is regulated in its genome (such as humans and domesticated animals, including pets (e.g. dogs, cats, birds, hamsters), livestock (e.g. pigs, sheep, cattle), and beasts of burden (e.g. horses, camels, elephants). The methods disclosed herein can be used in virtually any applications involving distinguishing methylation status of nucleic acid molecules in two or more states, including for example cancer or cancer-recurrence detection / diagnosis, other disease detection / diagnosis, therapy selection, non-invasive prenatal testing (NIPT), organ health or organ transplant evaluation, prediction and monitoring, and veterinary applications (e.g. aging or health status evaluation, prediction and monitoring of domesticated animals).

[0112] Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The singular terms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. “Comprising A or B” means including A, or B, or A and B. It is further to be understood that all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for DNA molecules or polypeptides are approximate, and are provided for description.

[0113] Further, ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 1 to 49, 1 to 25, 1.7 to 31.9, and so forth (as wellas fractions thereof unless the context clearly dictates otherwise). Any concentration range, percentage range, ratio range, or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), unless otherwise indicated. Also, any number range recited herein relating to any physical feature, such as polymer subunits, size or thickness, are to be understood to include any integer within the recited range, unless otherwise indicated. When multiple low and multiple high values for ranges are given that overlap, a skilled artisan will recognize that a selected range will include a low value that is less than the high value.

[0114] As used herein, “about” or “consisting essentially of” mean ± 10% of the indicated range, value, or structure, unless otherwise indicated. As used herein, the terms “include” and “comprise” are open ended and are used synonymously. As used herein, “comprising” is synonymous with “including,” “containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. As used herein, “consisting of” excludes any element, step, or ingredient not specified in the claim element. As used herein, “consisting essentially of” does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claim. In each instance herein any of the terms “comprising”, “consisting essentially of” and “consisting of” may be replaced with either of the other two terms. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein.

[0115] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entireties. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0116] It is appreciated that certain features of aspects and embodiments herein, which are, for clarity, discussed in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various aspects and embodiments, which are,for brevity, discussed in the context of a single aspect or embodiment, may also be provided separately or in any suitable sub-combination. All combinations of aspects and embodiments are specifically embraced herein and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various aspects and embodiments and elements thereof are also specifically disclosed herein even if each and every such sub-combination is not individually and explicitly disclosed herein. Subjects

[0117] Subjects in methods herein can be virtually any eukaryotes in which methylation is regulated in its genome, including an animal, in illustrative embodiments a mammal, and in further illustrative embodiments, a human. In some embodiments, the subject is a domesticated animal, in illustrative embodiments, a pet, and in further illustrative embodiments, a dog or a cat. In some embodiments, the subject has, had, or is suspected or at risk of having a disease, in some embodiments, cancer or precancerous lesion. In some embodiments, the subject has received or is receiving a treatment for cancer and is being tested for minimal / molecular residual disease (MRD). In some embodiments, the subject is a pregnant female. In some embodiments, the subject is a subject receiving or having received an organ from another subject.

[0118] In embodiments where the subject has, had, or is suspected or at risk of having cancer or precancerous lesion, the cancer can be any type of cancer provided that the genome of cancerous cells of the subject have portions of their genome that are differentially methylated compared to non-cancerous cells of the subject. Typically, some, most, almost all or all cancer cells of the subject have region(s) of their genome that are methylated that are not methylated in non-cancerous cells of the subject, or that are more methylated than non-cancerous cells. Thus, in some embodiments, the subject has one or more cancers (e.g., one cancer) including ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non- small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer,breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low- grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.

[0119] In certain embodiments of methods herein, the subject has, had, or is suspected or at risk of having a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, blood, bone, bone marrow, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro-esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovaries, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whipple resection. In illustrative embodiments, the cancer is selected from nasopharyngeal carcinoma, hepatocellular carcinoma, breast cancer, ovarian cancer, pancreatic cancer, colorectal cancer, lung cancer, oesophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. In some embodiments, the breast cancer may be human epidermal growth factor 2 (HER2) positive breast cancer, hormone receptor (HR) positive breast cancer, or triple negative breast cancer (TNBC). Insome embodiments, the lung cancer may be non-small cell lung cancer (NSCLC), NSCLC adenocarcinoma, NSCLC squamous cell carcinoma, or small cell lung cancer (SCLC). In some embodiments, the bladder cancer may be urothelial or muscle-invasive bladder cancer (MIBC). In illustrative embodiments, the cancer is selected from colorectal, breast, lung, and bladder cancer. Sample Extraction and Enrichment of DNA Molecules

[0120] Samples that are useful for methods herein can be virtually any nucleic acid sample. In illustrative embodiments, the nucleic acid sample is extracted or isolated from a subject. In illustrative examples, DNA molecules are extracted or isolated from a sample such as a tissue sample, and for illustrative embodiments herein, a liquid sample. Methods that are particularly useful in exemplary embodiments include methods for isolating circulating free DNA (cfDNA) from a liquid sample, and in illustrative embodiments from a blood or a blood product or any fraction of blood (e.g., whole blood, serum, plasma, buffy coat or the like), umbilical cord blood, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneal, ductal, ear, arthroscopic), washings of female reproductive tract, breast milk, breast fluid, nasal mucous, prostate fluid, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, semen sample or combinations thereof. Blood or fractions thereof often comprise nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes comprise nucleic acids and are sometimes cell-free or intracellular. Blood also comprises buffy coats. Buffy coats are sometimes isolated by utilizing a ficoll gradient. Buffy coats can comprise white blood cells (e.g., leukocytes, T-cells, B-cells, platelets, and the like). Blood plasma refers to the fraction of whole blood resulting from centrifugation of blood treated with anticoagulants. Blood serum refers to the watery portion of fluid remaining after a blood sample has coagulated. Fluid samples often are collected in accordance with standard protocols hospitals or clinics generally follow. For blood, an appropriate amount of peripheral blood (e.g., between 3-40 milliliters) often is collected and can be stored according to standard procedures prior to or after preparation. A fluid sample from which nucleic acid is extracted may be acellular (e.g., cell-free). In some embodiments, a fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancer cells may be included in the sample. In some embodiments the cells, cellular elements or cellular remnants are removed from the liquid sample prior to nucleic acid extraction.

[0121] A sample may be heterogeneous and include more than one type of nucleic acid species is present in the sample. For example, heterogeneous nucleic acids can include, but are not limited to, (i) fetal derived and maternal derived nucleic acids, (ii) nucleic acids from cancerous and non-cancerous cells, (iii) pathogen and host nucleic acids, (iv) transplant donor and recipient nucleic acids, or (v) mutated and wild-type nucleic acids.

[0122] To prepare serum or plasma from blood, blood can be placed for example in a tube containing EDTA or a specialized commercial product such as Streck BCT collection tubes (Streck). Serum may be obtained with or without centrifugation-following blood clotting. If centrifugation is used then it is typically, though not exclusively, conducted at an appropriate speed, e.g., 1,500-3,000 x g. Plasma or serum may be subjected to additional centrifugation steps before being transferred to a fresh tube for cfDNA extraction.

[0123] DNA methylation biomarkers are increasingly being utilized in the development of novel assays for use in, for example, cancer and other disease detection, women’s health, organ health, and veterinary health. For example, methylation profiling is being used in non-invasive prenatal testing (NIPT) for monitoring of placental and fetal epigenomic changes, pre- symptomatic detection of preterm birth, preeclampsia, placental insufficiency, and fetal growth restriction. Methylation markers can also be indicative of congenital diseases of the fetus. Furthermore, methylation biomarkers are also used for monitoring of organ health such as predicting and monitoring organ rejection in transplant patients, monitoring of immune changes in transplant rejection, and for monitoring of organ health in high risk or predisposed individuals. In addition, methylation biomarkers can be used to determine biological age or detect age-related diseases. Accordingly, the sample described herein may be a maternal sample, a fetal sample, a sample from a transplant patient, or a sample from a subject having, had, suspected of having, or at risk of having a certain phenotype.

[0124] In certain illustrative embodiments, isolation of cfDNA from a liquid (e.g., blood or blood derivative sample such as a serum or plasma sample) can involve binding DNA molecules from a sample to a matrix and isolating the DNA molecules in the presence of a solvent. In some embodiments, the method further comprises incubating the biological sample comprising DNA molecules with a protease, prior to contacting the DNA molecules to the matrix. In someembodiments, the method can further include the steps of washing the matrix with a wash buffer to remove impurities, and optionally, drying the matrix. Enriched nucleic acid samples can be eluted from the matrix with an elution buffer.

[0125] Other methods for nucleic acid isolation, for example cfDNA isolation, and optional enrichment of certain cfDNA can include ion exchange columns, or microfluidic devices, such as solid phase isolation, based on DNA capture by immobilized beads or functionalized surface. Additional methods include liquid phase isolation, utilizing an electric field, or chemical reagents, instead of a functionalized surface. In illustrative embodiments herein, isolation of cfDNA from a patient sample is performed using a DNA isolation kit (e.g., QIAamp Circulating Nucleic Acid kit (Qiagen)).

[0126] In some embodiments, cfDNA or their derivatives of certain sizes can be enriched or excluded before or after subjecting the cfDNA to methods herein. In some embodiments, size selection can be performed before the sequencing library preparation. In some embodiments, size selection or size exclusion can be performed after the sequencing library preparation and before sequencing. In some embodiments, size selection or size exclusion is performed on a sequencing- ready pool. Enriched cfDNA molecules can be, for example 50 to 1200 base pairs in length, 70 to 500 base pairs in length, 100 to 200 base pairs in length, or 130 to 170 base pairs in length. In some embodiments, the enriched cfDNA molecules are from 50 to 200 bp in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, or between 60 and 100 bp in length before the enriched cfDNA molecules, or derivatives thereof, are ligated to adapters in methods herein. In some embodiments, the enriched cfDNA molecules are less than 150, 100, 90, 75, or 50 bp in length before they are ligated to adapters. Such enrichment methods can be performed for example using the methods of WO2018156418 A1, Stray, et al. (incorporated herein by reference in its entirety).

[0127] In some embodiments, the sample is enriched for tumor DNA molecules, which are typically less than 160 bp and have a peak at about 145 bp in length. In illustrative embodiments, the enriched nucleic acid sample is circulating tumor DNA (ctDNA) or amplicons thereof. In such embodiments, the enriched nucleic acid is less than 160, 150, 145, 120, 100, 90, 75, or 50 bp in length. In some embodiments, size selection is used to filter out cfDNA molecules carryingmutations derived from clonal hematopoiesis (CH) but not tumor-derived mutations, which are typically longer than ctDNA molecules, at about 165bp.

[0128] In some embodiments, the sample is enriched for fetal DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is fetal circulating free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of fetal cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170 bp, from 160 to 190 bp, or from 170 bp to 220 bp.

[0129] In some embodiments, the sample is enriched for transplant donor DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is transplant donor circulating free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of transplant donor cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170bp, from 160 bp to 190 bp, or from 170 bp to 220 bp. Nucleic Acid Processing Appending Adapters

[0130] Typically, methods herein include a step of appending nucleic acid adapters to template DNA molecules or to nucleic acid derivatives generated therefrom. For example, adapters may be appended on to the DNA molecules by ligation or PCR. The template DNA molecules in illustrative embodiments are from a subject. In some embodiments, appending nucleic acid adapters is preformed after template DNA molecules are fragmented to form fragmented DNA molecules. Typically, methods include exposing template DNA molecules to one or more polymerases or kinases, such as Klenow Large Fragment Polymerase and T4 polynucleotide kinase (PNK), as well as a ligase, such as T4 ligase. In some embodiments, template DNA molecules or the fragmented DNA molecules are exposed to one or more polymerases and / or kinases to generate the nucleic acid derivatives. In some embodiments, the method further comprises appending adapters to the nucleic acid derivatives generated therefrom.In some embodiments, template DNA molecules are not fragmented prior to appending nucleic acid adapters thereto. In some embodiments, template DNA molecules are cfDNA molecules.

[0131] In some embodiments, adapters are ligated to template DNA molecules. Before such ligation, extracted or isolated DNA molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adapter ligation. For example, template DNA molecules can be blunt ended, nucleotides can be added to template DNA molecules or blunted-ended derivative therefrom, and / or phosphate moieties can be added or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, template DNA molecules may be blunt ended, and then a single adenosine base can be added to the 3’ end. Prior to ligation the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation the 3’ adenosine of the template fragments and the complementary 3’ thymidine overhang of an adapter can enhance ligation efficiency. In some embodiments, adapter ligation is performed using a T4 ligase.

[0132] In some embodiments, methylated adapters are utilized in methods herein. In some embodiments, adapters containing both one or more methylated Cs and one or more unmethylated Cs are utilized in methods herein. In some embodiments, adapters containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adapters are Y adapters, for example in methods in which target region amplicons are sequenced using NGS. In some embodiments, the adapters each comprises a universal priming site.

[0133] In some embodiments, the adapters each further comprises a sample barcode. Thus, multiple samples can be analyzed in the same sequencing reaction. The sample barcode can be used to process data according to the sample from which the data was generated.

[0134] In some embodiments, the adapters each further comprises an MIT. In some embodiments, the number of adapters having different MITs is between 10 to 1,000, and wherein the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction is at least 1,000:1. The number of different MITs in the ligation reaction, in certain embodiments, ranges from10 to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600, 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are atleast 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different MITs in the ligation reaction. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction is at least 10,000:1. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction ranges from 50,000: 1 to 50:1, from 25,000: 1 to 100:1, from 10,000:1 to 100:1, from 10:000:1 to 8,000:1 to 500:1, from 5,000:1 to 200:1, from 10,000:1 to 50:1. In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000 different MITs to each one template nucleic acid or cfDNA molecules. Methylation sensing

[0135] In some embodiments, the methods herein include performing one or more steps to sense, detect, or “read out” the methylation signals of the DNA molecules or their derivatives. In some embodiments, the methylation signals are ascertained by treating the extracted DNA molecules or their methylation-status-preserved derivatives with an agent or a combination of agents that discriminates between methylated and non-methylated cytosines. Any method suitable for detecting methylated cytosines at a single-base resolution may be used with the methods described herein. In some embodiments, the extracted DNA or derivative thereof is treated with a deaminating agent that converts (directly or indirectly through one or more other agents) unmethylated but not methylated cytosines to uracils before selective enrichment with a panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a deaminating agent (directly or indirectly through one or more other agents) that converts unmethylated but not methylated cytosines to uracils after selective enrichment with a panel of oligonucleotide probes. In some embodiments, the deaminating agent (e.g., a chemical reagent such as sodium bisulfite) directly converts unmethylated cytosines to uracils. In some embodiments, the deaminating agent converts unmethylated cytosines to uracils in combination with one or more other agents – for example, in order to convert unmethylated cytosines to uracils, the extracted DNA or derivative thereof may be first treated with an oxidizing agent (e.g., TET or its catalytic domain) that oxidizes one or more forms of methylated cytosines (e.g., 5hmC or 5mC), followed by treatment with a deaminase (e.g., APOBEC or its catalytic domain).

[0136] In some embodiments, the deaminating agent comprises a chemical reagent (e.g., sodium bisulfite). In some embodiments, the deaminating agent comprises a deaminase (e.g., APOBEC) or its catalytic domain. In some embodiments, the deaminating agent comprises APOBEC3A (A3A) or its catalytic domain.

[0137] In some embodiments, the extracted DNA or derivative thereof is treated with a reducing agent that converts certain forms of modified cytosines, but not unmodified cytosines, into dihydrouridine (DHU) before the contacting with the panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a reducing agent that converts certain forms of modified cytosines, but not unmodified cytosines, into DHU after the contacting with the panel of oligonucleotide probes. In some embodiments, the reducing agent comprises pyridine borane.

[0138] In some embodiments, the extracted DNA or derivative thereof is treated with an oxidizing agent that oxidizes one or more forms of methylated cytosines before being treated with a deaminating agent or a reducing agent. In some embodiments, the oxidizing agent oxidizes 5hmC and 5mC to 5caC (5-carboxylcytosine). In some embodiments, the oxidizing agent selectively oxidizes 5hmC, but not 5mC, to 5fC (5-formylcytosine). In some embodiments, the oxidizing agent comprises an enzyme of the ten-eleven translocation (TET) families or its catalytic domain. In some embodiments, the oxidizing agent comprises TET1 or its catalytic domain. In some embodiments, the oxidizing agent comprises TET2 or its catalytic domain. In some embodiments, the oxidizing agent comprises TET3 or its catalytic domain. In some embodiments, the oxidizing agent comprises potassium perruthenate (KRuO4). In some embodiments, the oxidizing agent comprises potassium ruthenate (K2RuO4).

[0139] In some embodiments, the extracted DNA or derivative thereof is treated with a glycosylation agent that glycosylates 5hmC to glucosyl-5hmC (5ghmC) to protect it from oxidation, deamination, or reduction before being treated with an oxidizing agent, a deaminating agent, or a reducing agent. In some embodiments, the glycosylation agent is a glycosyltransferase. In some embodiments, the glycosyltransferase is a glucosyltransferase. In some embodiments, the glucosyltransferase is β-glucosyltransferase (β-GT). In some embodiments, the glucosyltransferase is β-GT of T4 phase.

[0140] In some embodiments, the extracted DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, before selective enrichment with a panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5- hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, after selective enrichment with a panel of oligonucleotide probes. In some embodiments, other methylation detection methods such as methylated DNA immunoprecipitation (MeDIP) or direct detection by long read sequencers may also be used. (i) Conversion-based methylation detection

[0141] In some embodiments, the methods described herein include treating the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with one or more chemical or enzymatic agents or a combination of chemical and enzymatic agents (e.g. deaminating agents, reducing agents, oxidizing agents, glycosylation agents) that allow discrimination between methylated and unmethylated cytosines, or between 5mC and 5hmC. In some embodiments, a portion of a library of DNA molecules or their derivatives is treated with one or more such chemical and / or enzymatic agents. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more such chemical and / or enzymatic agents. In some embodiments, the entire library of DNA molecules or their derivatives is treated with one or more such chemical or enzymatic agents.

[0142] In some embodiments, the agent comprises a deaminating agent. In some embodiments, the deaminating agent comprises an enzyme such as a deaminase or its catalytic domain. In some embodiments, the deaminase may be APOBEC or its catalytic domain. In some embodiments, the deaminase may be APOBEC3A (A3A) or its catalytic domain, which deaminates unmethylated C, 5mC, and 5hmC, but not 5ghmC, 5fC, or 5caC. In some embodiments, the method described herein involves treating the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with TET2 / oxidation enhancer,which converts 5mC and 5hmC to 5caC and protects them from deamination by APOBEC, followed by treatment with APOBEC to deaminate the unmethylated cytosines to uracils. In some embodiments, the DNA sample is treated with a glycosylation agent (e.g. T4 β- glucosyltransferase) to glycosylates 5hmC to 5ghmC and protects it from deamination by APOBEC.

[0143] In some embodiments, the deaminating agent comprises a bisulfite reagent. Bisulfite reagents, such as sodium bisulfite, convert unmethylated cytosine to uracil and leave methylated cytosine (including 5mC and 5hmC) unchanged. Therefore, after bisulfite treatment, 5mC and 5hmC in the DNA remains as cytosine and unmodified or unmethylated cytosine will be changed to uracil. In some embodiments, the bisulfite reagent may be sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassium metabisulfite, ammonium metabisulfite, magnesium metabisulfite, and equivalents. In some embodiments, a combination of one or more chemical reagents and / or one or more enzymes can be used to discriminate between various forms methylated cytosines (e.g., mC or hmC) and unmethylated cytosines. For example, in some embodiments, the DNA sample is treated with an oxidizing agent (e.g. KRuO4or K2RuO4) that selectively oxidizes 5hmC, but not 5mC, to 5fC before bisulfite treatment, which converts unmethylated cytosine (including 5fC converted from 5hmC) to uracil but leaving 5mC unchanged. In some embodiments, a glycosylation agent (e.g. T4 β-glucosyltransferase or its catalytic domain) is used to glycosylate 5hmC to 5ghmC and an oxidizing agent (e.g. a TET enzyme or its catalytic domain) is used to convert 5mC to 5caC before bisulfite treatment, which converts unmethylated cytosine (including 5caC converted from 5mC) to uracil but leaving methylated cytosine (including 5ghmC converted from 5hmC) unchanged.

[0144] The bisulfite treatment can be performed by commercial kits such as the Imprint DNA Modification Kit (Sigma), EZ DNA Methylation- DirectTM Kit ( ZYMO), and the EZ DNA Methylation-Gold Kit (ZYMO). After DNA bisulfite conversion, single stranded DNA is captured, desulphonated and cleaned. The bisulfite-treated DNA can be captured by purification columns or magnetic beads and eluted. Bisulfite-treated single stranded DNA can be converted into dsDNA through an enzyme-catalyzed DNA strand synthesis with appropriate primers and polymerase. The polymerase will recognize the uracil in the ssDNA template as thymine and add an adenine to the complimentary strand. Further polymerase extension on the complimentarystrand will result in replication of the original bisulfite treated ssDNA template, substituting uracil with thymine. Identification of cytosine to thymine conversion and guanine to adenine conversions (complementary strand) through comparing to the reference genome, will determine all unmodified cytosines, while the remaining cytosines are considered to be methylated.

[0145] In some embodiments, the agent comprises a reducing agent (e.g. pyridine borane) that selectively reduces 5fC and 5caC to dihydrouridine (DHU), which, like uracil, is converted to T base following PCR amplification. In some embodiments, the method described herein involves treating the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with TET2 / oxidation enhancer, which converts 5mC and 5hmC to 5caC, followed by treatment with borane, which reduces 5fC and 5caC (including 5caC converted from 5mC and 5hmC) to DHU but leaves unmodified cytosine (dC) unchanged. In some embodiments, the DNA sample is treated with an oxidizing agent (e.g. KRuO4or K2RuO4) that selectively oxidizes 5hmC, but not 5mC, to 5fC, followed by treatment with borane, which reduces 5fC (including 5fC converted from 5hmC) and 5caC to DHU but leaves dC and 5mC unchanged. In some embodiments, the DNA sample is first treated with a glycosylation agent (e.g. T4 β- glucosyltransferase) and an oxidizing agent (e.g. TET) such that 5hmC is glycosylated to 5ghmC and only 5mC is oxidized and converted to 5caC. The subsequent treatment with borane reduces 5fC and 5caC (including 5caC converted from 5mC) to DHU but leaves dC and 5ghmC unchanged.

[0146] Other conversion-based methyl detection methods, including EM-seq, oxidative bisulfite sequencing (oxBS-seq), TET-assisted bisulfite sequencing (TAB-seq), TET-assisted pyridine borane sequencing (TAPS), TAPS with T4-βGT protection sequencing (TAPSβ), chemical-assisted pyridine borane sequencing (CAPS), and many variants of these that can discriminate between C, mC, and / or hmC. These methods can be performed with commercial kits such as NEBNext® Enzymatic Methyl-seq (EM-seq™) (New England Biolabs, Inc.), the 5hmC TAB-Seq Kit (WiseGene), and the EpiTect® Bisulfite Kit (Qiagen), for example. (ii) Restriction enzyme-based methylation detection methods

[0147] In some embodiments, the methods described herein include treating extracted or enriched DNA molecules, or their derivatives, with one or more restriction enzymes. In some embodiments, the one or more restriction enzymes are one or more methylation sensitive restriction enzymes (MSREs). In some embodiments, a portion of a library of extracted or enriched DNA molecules is treated with one or more MSREs. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more MSREs. In some embodiments, the entire library of DNA molecules is treated with one or more MSREs.

[0148] MSREs are useful in analyzing the methylation status of cytosine residues in CpG sites. As the name implies, these enzymes are not able to cleave their palindromic target sites when the cytosine residues are methylated. The size of the MSRE targets sites range from 4 bp up to 76bp, but typically are in the 4 to 8 bp range.

[0149] In one aspect, methods herein typically include contacting extracted or enriched DNA molecules, or their derivatives, with one or more MSREs. MSREs selectively cleave a sample nucleic acid when the MSRE recognition site is unmethylated, but not when the MSRE recognition site is methylated. An “isoschizomer” of an MSRE is a restriction enzyme that recognizes the same recognition site as a methylation sensitive restriction enzyme but cleaves both methylated CGs and unmethylated CGs. Isoschizomer of the selected MSREs may be used in control reactions. Non-limiting examples of methylation sensitive restriction enzyme include, and thus in some embodiments, the one or more MSREs can include, AatII, Acc65I, AccI, AciI, AclI, AfeI, AgeI, AgeI-HF®, AhdI, AleI-v2, ApaI, ApaLI ApeKI, AscI, AsiSI, AvaI, AvaII, BaeI, BanI, BbvCI, BceAI,, BcgI, BcoDI, BfuAI, BglI, BmgBI, BsaAI, BsaBI, BsaHI, BsaI-HF®v2, BseYI, BsiE, BsiWI, BsiWI-HF®, BslI, BsmAI, BsmBI-v2, BsmFI, BspDI, BspEI, BsrBI, BsrFI-v2, BssHII, BstAPI, BstBI, BstUI, BstZ17I-HF®, BtgZI, Cac8I, ClaI, DpnI, DraIII-HF®, DrdI, EaeI, EagI-HF®, EarI, EciI, Eco53kI, EcoRI, EcoRI-HF®,EcoRV, EcoRV-HF®, Esp3I, FauI, Fnu4HI, FokI, FseI, FspI, HaeII, HgaI, HhaI, HinP1I, HincII, HinfI, HpaI, HpaII, Hpy166II, Hpy188III, Hpy99I, HpyAV, HpyCH4IV, KasI, MboI, MluI, MluI-HF®, MmeI, MspA1I, MwoI, NaeI, NarI, NciI, NgoMIV, NheI-HF®, NlaIV, NotI, NotI-HF®, NruI, NruI-HF®, Nt.BbvCI, Nt.BsmAI, Nt.CviPII, PaeR7I, PaqCI, PleI, PluTI, PmeI, PmlI, PshAI, PspOMI, PspXI, PvuI, PvuI-HF®, RsaI, RsrII, SacI-HF®, SacII, SalI, SalI-HF®, Sau3AI, Sau96I, ScrFI, SfaNI, SfiI, SfoI, SgrAI,SmaI, SnaBI, SrfI, StyD4I, TfiI, TseI, TspMI, XhoI, XmaI, and / or ZraI. In illustrative embodiments, the one or more MSREs can include HhaI, HpaII, BstUI, and / or HpyCH4IV.

[0150] In some embodiments, the one or more restriction enzymes are one or more methylation dependent restriction enzymes (MDREs). MDREs selectively cleave a sample nucleic acid when one or more nucleotides in the MDRE recognition site is methylated. In some embodiments, one or more MDREs can be used in combination with one or more MSREs. In some embodiments, the one or more MDREs can be AbaSI, AoxI, BisI, BlsI, DpnI, FspEI, GlaI, GluI, KroI, LpnPI, MalI, MspJI, MteI, PcsI, PkrI, or SgeI.

[0151] The MSRE or MDRE can be selected based on differentially methylated CpG sites in target DNA molecules, such as tumors or ctDNA from specific cancer targets, or a diverse spectrum of tumors. Further criteria for selection may include low background methylation in normal tissues, size and number of cleavage fragments, presence and number of CHG and / or CHH sites, number of base pairs of recognition sequence, whether the cleavage results in blunt vs. tailed end fragments, and whether the enzymes have the same or similar reaction conditions such that contacting with the enzymes can be done under the same set of conditions and / or in a single reaction.

[0152] In some embodiments, more than one, a plurality, or a set of MSREs (and / or MDREs) can be used to contact extracted or enriched DNA molecules, or their derivatives, comprising one or more CpG sites of interest. The number of selected MSREs in certain embodiments is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; from 2 to 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; or from 5 to 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 10, 1 to 5, from 2 to 5, from 3 to 5, from 4 to 5, 3 to 7, from 5 to 10, from 6 to 10, or from 7 to 10. Thus, criteria such as target site, target site methylation status, number of base pairs in tailed end fragments, reaction conditions including but not limited to buffer conditions, incubation time and temperature for optimal activity, as well as deactivation time and temperature for each MSRE can be used for selecting the plurality or set of MSREs and / or MDREs to include in the method. As activity measured by units is specific to each enzyme, the target DNA molecules can be contacted with from 1 to 5 Units (U) of each MSREs inthe reaction sample. In some embodiments, the extracted or enriched cfDNA molecules, or their derivatives, are contacted with from 1 to 5 U, from 1.5 to 4 U, from 2 to 3 U, from 2.5 to 4 U, or from 3 to 5 U of each MSRE. In some embodiments, one or more of the MSRE is selected from HpaII, SalI, BbeI, NotI, SmaI, XmaI, MboI, BstUI, BstBI, ClaI, MluI, NaeI, NarI, PvuI, SacII, HpyCH41V, HhaI, and combinations thereof. In exemplary embodiments, the one or more MSREs is selected from one or more of HpaII, HhaI, HpyCH41V, and BstU1. In some embodiments, the one or more MSREs are selected from one or more of HpaII, HhaI, HpyCH41V, or BstU1. In some embodiments of the methods as described herein, the contacting comprises contacting with two or more MSREs. In some embodiments, the contacting comprises contacting with three or more MSREs. In some embodiments, the contacting comprises contacting with four or more MSREs. In some embodiments, the contacting comprises contacting with the two or more MSREs in a single reaction.

[0153] In some embodiments, MSREs and / or MDREs are included that are able to cleave the MSRE sites and / or MDRE sites in the same cleavage buffer conditions. In some embodiments, the cleavage buffer conditions include 5-500 mM potassium acetate, for example 25-100 mM potassium acetate. In some embodiments, the cleavage buffer conditions include 2-200 mM Tris- acetate, for example 10-40 mM Tris-acetate. In some embodiments, the cleavage buffer conditions include 1-100 mM magnesium acetate, for example 5-20 mM magnesium acetate. In some embodiments, the cleavage buffer conditions include 10-1000 μg / ml recombinant albumin, for example 50-200 μg / ml recombinant albumin. In some embodiments, the pH is between 6.9 and 8.9 at 25 °C, for example, between 7.4 and 8.4, 7.5 and 8.3, 7.6 and 8.2, 7.7 and 8.1, or 7.8 and 8, or about 7.9 at 25 °C. In some embodiments, the cleavage buffer conditions include 1-100 mM bis- tris-propane-HCl, for example 5-20 mM bis-tris-propane-HCl. In some embodiments, the cleavage buffer conditions include 1-100 mM MgCl2, for example 5-20 mM MgCl2. In some embodiments, the pH is between 6 and 8 at 25 °C, for example, between 6.5 and 7.5, 6.6 and 7.4, 6.7 and 7.3, 6.8 and 7.2, or 6.9 and 7.1, or about 7.0 at 25 °C. In some embodiments, the cleavage buffer conditions include 5-500 mM NaCl, for example 25-100 mM NaCl. In some embodiments, the cleavage buffer conditions include 1-100 mM Tris-HCl, for example 5-20 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 10-1000 mM NaCl, for example 50-200 mM NaCl. In some embodiments, the cleavage buffer conditions include 5-500 mM Tris-HCl, forexample 25-100 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 25- 100 mM potassium acetate, 10-40 mM Tris-acetate, 5-20 mM magnesium acetate, and 50-200 μg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 5-20 mM bis-tris-propane-HCl, 5-20 mM MgCl2, and 50- 200 μg / ml recombinant albumin, and the pH is between 6.7 and 7.3at 25 °C. In some embodiments, the cleavage buffer conditions include 25-100 mM NaCl, 5-20 mM Tris-HCl, 5-20 mM MgCl2, and 50-200 μg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 50-200 mM NaCl 25-100 mM Tris- HCl, 5-20 mM MgCl2, and 50-200 μg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. Amplification

[0154] In some embodiments, DNA is amplified before or after one or more of the steps of the methods disclosed herein. In some embodiments, converted or MERE / MDRE-treated DNA is amplified after treatment with an agent or a combination of agents that discriminates between methylated and unmethylated cytosine. In some embodiments, extracted DNA or its derivative is amplified before treatment with a methylation-sensing agent or a combination of agents using an amplification method that preserves the methylation status of the nucleotides. In some embodiments, the extracted and / or otherwise processed DNA is amplified before selective enrichment with a panel of oligonucleotides. In some embodiments, the processed DNA is amplified after selective enrichment with a panel of oligonucleotides. In some embodiments, the processed DNA is amplified both before and after selective enrichment with a panel of oligonucleotides.

[0155] Methods in some aspects herein include performing one, two, or more amplifications. In some embodiments, methods herein include one or more universal amplifications using primers that hybridize to universal priming sequences. In some embodiments, sample barcodes are added during one or more amplifications.

[0156] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g. recombinase polymeraseamplification (RPA) (Kersting et al., Microchim. Acta 181(13-14): 1715–1723, 2014; incorporated by reference in its entirety), a ligase-based amplification, PCR, or a combination thereof (e.g. ligation-mediated PCR).

[0157] In some methods herein, a universal amplification(s) can be performed before and / or after target enrichment, following bisulfite or enzymatic treatment of the DNA molecules. Such universal amplification can be performed for example using a primer pair that binds universal priming sites in the adapter. In some embodiments, the methods herein include performing a universal PCR using a plurality of adapted DNA molecules or a plurality of adapted cfDNA, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters, to generate amplified, adapted, DNA or cfDNA molecules, before performing target enrichment. In some embodiments, the methods herein include performing a universal PCR using a plurality of enriched, adapted DNA molecules or a plurality of enriched, adapted cfDNA, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters or introduced via an initial amplification, to generate amplified, enriched, adapted, DNA or cfDNA molecules, after performing target enrichment.

[0158] In some embodiments, performing a PCR further comprises using primers comprising a sequence that can be utilized in subsequent sequencing step, such as NGS. For example, such sequences can include NGS flow cell binding sites (e.g. Illumina P5 and P7 sequences) and / or NGS sequencing primer binding sites. In some embodiments, performing a PCR further comprises using primers comprising a sample index.

[0159] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g. multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more cycles. In some embodiments, amplification cycles can include at least 5, 6, 7, 8, 9, or 10 cycles. In some embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.

[0160] Typically, in embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template (e.g., adapted sample DNA or cfDNA) followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, deoxynucleotides (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is between from 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.

[0161] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgCl2), potassium chloride (KCl), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KCl can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgCl2 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgCl2 is 2.0 mM.

[0162] In some embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England BioLabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England BioLabs, Inc.).

[0163] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3´→ 5´ exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5´→ 3´exonuclease activity and strand displacement activity.

[0164] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5´→ 3´ direction and requires the presence of template and primer. This enzyme has a 3´→ 5´ exonuclease activity which is much more active than that found in DNA Polymerase I. T4 DNA polymerase lacks 5´→ 3´ exonuclease activity and strand displacement activity.

[0165] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 nucleotides in length, between 10 and 40 nucleotides in length, between 15 and 30 nucleotides in length, between 15 and 25 nucleotides in length, between 20 and 40 nucleotides in length, between 25 and 50 nucleotides in length, or between 30 and 50 nucleotides in length. In some embodiments, the primers are between 25 and 100 nucleotides in length, between 35 and 100 nucleotides in length, between 45 and 100 nucleotides in length, between 55 and 100 nucleotides in length, between 65 and 100 nucleotides in length, or between 75 and 100 nucleotides in length.

[0166] In some embodiments, the amplifying adds a sample barcode, and wherein a plurality of libraries of the selected DNA or derivative thereof each produced from a differentsample are sequenced together in one sequencing lane. In some embodiments, the methods herein further comprise performing size selection / exclusion before and / or after the amplifying step. Panel Target Selection

[0167] Methylation status of CpG loci in a genomic region proves to be spatially correlated, giving rise to a host of methods that leverage this pattern to increase the effective signal-to-noise ratio (SNR) in determining cell-type of origin and early disease detection (e.g., cancer or cancer recurrence detection). High local methylation correlation means that multiple CpGs can be integrated into differentially methylation regions (DMRs), which strongly increase SNR, in particular at low circulating tumor DNA (ctDNA) fractions. DMRs with multiple adjacent CpG sites that have different DNA methylation status among different disease states or different biological samples (such as different cells / tissues within the same individual, the same cells / tissues within the same individual at different times, cells / tissues from different populations representing different disease states, or different alleles in the same cell) are useful biomarkers for early detection of a disease or condition or recurrence of the disease or condition.

[0168] In some embodiments, differentially methylated genomic patterns are defined between patients having a disease or phenotype, such as colorectal cancer (CRC), and healthy individuals by utilizing patterns of co-methylation of adjacent CpG loci. Example embodiments seek to identify and characterize a phenotype through phenotype-specific methylation biomarkers (MBMs) via identification of phenotype-specific differentially methylated regions from a large set of wide genomic regions, further discussed below. To achieve a set of features which robustly differentiate disease from normal samples, for example, example embodiments use read-level methylation information to jointly model the methylation status of different CpGs in each genomic window. Joint modeling of methylation patterns of consecutive CpG loci leverages read-level information, which has the statistical power to distinguish between, for example, tumor-derived DNA and normal fragments even in presence of rare amounts of tumor fraction in low-coverage. Modeling helps identify new biomarkers that can be used in panels with target-specific primers or probes that increase detection sensitivity and specificity for cancer or other diseases. This approach can provide a set of biomarkers which can be used to detect a phenotype or a disease (for example, CRC and AA, and which can be sub-selected to represent samples of different ages,genders, ethnicities etc. This approach can also be used to identify and characterize biomarkers suitable for detecting other cancers or phenotypes, or for differentiating cancer types, origins, or states.

[0169] In various embodiments, potentially informative regions of the genome are pre- selected for downstream analysis, thereby significantly reducing the amount of data that need to be stored, analyzed, and otherwise manipulated. In some embodiments, to identify tightly and potentially co-methylated CpG blocks across the genome, whole genome methylation sequencing data is preliminarily segmented into smaller regions based on biological priors and the sequencing technology and a chosen processing pipeline before any further analysis or modeling of the data. In various embodiments, the genome may be split into “ridges” (stretches or blocks of CpGs in which the distance between adjacent CpGs are less than a threshold) and “trenches” (stretches devoid of any CpG). For example, the maximum allowable distance between adjacent CpG loci in a ridge can be set to no more than 50 bp, no more than 100 bp, no more than 150 bp.

[0170] The preliminary segmentation can be based on additional selected criteria, such as a threshold based on a certain primer or probe length (e.g.75 bp -150 bp), expected fragment lengths (e.g.50 bp – 140 bp), or processivity (in bp) of the DNA (cytosine-5)-methyltransferase 1 (DNMT1) enzyme (e.g.20 bp – 1,000 bp).

[0171] In various embodiments, a set of preliminary segments is constructed by constructing a reference methylome. In some embodiments, the preliminary segments are constructed using a common public reference methylome. In some embodiments, the preliminary segments are constructed using a set of individual methylomes from each sample in a study cohort. In some embodiments, the preliminary segments are constructed using reference methylomes of healthy / control samples, without using methylome data of tumor samples.

[0172] In some embodiments, any CpG block that is found in less than a threshold percentage (e.g. between about 10% and about 40%) of total samples is removed. In some embodiments, the ridges are filtered to contain at least a threshold number of CpG sites (e.g. at least 3 CpGs). In some embodiments, the ridges are filtered to have a CpG density of at least a density threshold (e.g. at least 0.02 CpG / bp). In various embodiments, preliminarily segmentingthe data comprises filtering the ridges (i) to autosomes only, (ii) to having at least 5 CpG sites, and (iii) to having a density of at least 0.03 CpG / bp.

[0173] In various embodiments, probabilistic modeling, such as hierarchical Bayesian modeling (e.g., fully variational Bayesian modeling), is used to further segment and define a narrower set of relevant genomic regions, such as DMRs, that can potentially be used as biomarkers for disease or phenotype detection. In various embodiments, filtered ridges are used as inputs to a segmentation model, which further segments the ridges into smaller regions where CpGs are co-methylated within each cohort. In various embodiments, a hierarchical Bayesian model is designed and implemented to infer and estimate change points, which indicate the boundaries of regions where the methylation load signal is relatively constant or almost constant in a given region across all samples within a cohort. In various embodiments, hierarchical Bayesian modeling is used to fit a piecewise constant function to beta-binomially distributed methylation load signals. The modeling takes into consideration not only the beta (expected) values of the signals but also the variance of the beta values, thereby naturally increasing segmentation robustness. In some embodiments, stochastic variational inference (SVI) is used.

[0174] In various embodiments, after regions where the methylation status is relatively constant or almost constant across all samples within a cohort are identified, regions where the methylation status are significantly different between samples from different cohorts are identified as biomarker candidates.

[0175] In various embodiments, candidate biomarkers are scored and prioritized based on their respective association with different clinical covariates to obtain a select set of biomarkers. In various embodiments, each candidate biomarker is scored on the signal-to-noise ratio (SNR) with the aim of prioritizing regions where the background noise (e.g., healthy methylation signal) is minimal (TDMF baseline) compared to the phenotype signal (e.g., the cancer signal). Scores for the candidate biomarkers can be used to prioritize the candidate biomarkers, such that the candidate biomarkers can be ranked based on their scores (e.g., from highest scores to lowest scores). For example, candidate biomarkers with relatively higher scores, or having scores above a threshold value, or having scores falling within a first range of values, may be deemed to be relatively more likely to distinguish between healthy subjects and subjects exhibiting a phenotype,and candidate biomarkers with relatively lower scores, or having scores below the threshold value, or having scores falling within a second range of values, may be deemed to be relatively less likely to distinguish between the healthy subjects and the subjects exhibiting the phenotype. The candidate biomarkers can also be prioritized based on a number of criteria, such as designability of probes or primers for the biomarkers. In some embodiments, biomarkers in repeat or low- complexity regions are deprioritized. Table A lists exemplary biomarkers that may be generated by the pre-segmentation, segmentation, scoring, and prioritization process according to various embodiments.

[0176] In some embodiments, a further narrowed set of candidate biomarkers comprising recurrently differentially methylated regions is included in a sequencing panel for targeted methylation sequencing. In illustrative examples, criteria for selection of suitable DMRs include factors such as high correlation of methylation status of adjacent CpGs, high CpG density such that sufficient CpGs are captured with a typical cfDNA fragment, and virtually zero background signal in health / normal samples to reveal ctDNA at VAF of about 0.01%. Table A illustrates some of the candidate biomarkers for CRC identified using the methods disclosed herein.

[0177] In some embodiments, instead of whole genome methylation sequencing (e.g., whole genome bisulfite sequencing), a targeted sequencing panel comprising preselected genomic regions, such as CpG ridges filtered according to the methods herein, can be used for methylation biomarker discovery.

[0178] Methylation biomarker discovery and identification are described in, e.g., International patent application No. PCT / US2025 / 010630, filed January 7, 2025 and titled “Methylation Biomarker Generation and Analysis”, and International patent application No. PCT / US2025 / 042194, filed August 15, 2025 and titled “Improved Method For Methylation Biomarker Generation and Analysis”, which are incorporated herein by reference in their entireties.

[0179] In some embodiments, the panel targets a plurality of differentially methylated regions having for example at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 100, or at least 150differentially methylated CpG sites. In some embodiments, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% of the target regions each includes at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 100, or at least 150 differentially methylated CpG sites. In some embodiments, 75% or more of the target regions each includes 10 or more CpGs per target, with a median of 15 to 17 CpGs and a mean of 18 to 20 CpGs per target. In some embodiments, 75% or more of the target regions each includes 5 or more CpGs per target, with a median of 9 to 11 CpGs and a mean of 12 to 14 CpGs per target.

[0180] In some embodiments, the panel targets a plurality of differentially methylated regions having for example a CpG density of between 0.02 to 0.8 CpG / base. In some embodiments, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of the target regions each has a CpG density of between 0.02 to 0.7 CpG / base, between 0.03 to 0.6 CpG / base, between 0.04 to 0.5 CpG / base, or between 0.05 to 0.4 CpG / base.

[0181] In some embodiments, the panel targets a plurality of differentially methylated regions having for example a length of between 10 bp to 1500 bp. In some embodiments, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% of the target regions each has a length of between 20 bp to 1000 bp, between 30 bp to 800 bp, between 40 bp to 600 bp, between 40 bp to 500 bp, between 40 bp to 400 bp, between 40 bp to 300 bp, between 40 bp to 200 bp, or between 40 bp to 120 bp.

[0182] In some embodiments, the panel targets a plurality of differentially methylated regions having for example a differential methylation load of between 0.12 to 0.80. In some embodiments, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% of the target regions each has a differential methylation load of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8.

[0183] In some embodiments, the panel targets up to 100 (e.g., 10-100, or 20-100, or 50- 100) differentially methylated regions. In some embodiments, the panel targets up to 200 (e.g., 20- 200, or 50-200, or 100-200) differentially methylated regions. In some embodiments, the panel targets up to 500 (e.g., 50-500, or 100-500, or 200-500) differentially methylated regions. In someembodiments, the panel targets up to 800 (e.g., 50-800, or 100-800, or 200-800) differentially methylated regions. In some embodiments, the panel targets up to 1,000 (e.g., 50-1,000, or 100- 1,000, or 200-1,000, or 500-1,000) differentially methylated regions. In some embodiments, the panel targets up to 2,000 (e.g., 50-2,000, or 100-2,000, or 200-2,000, or 500-2,000) differentially methylated regions. In some embodiments, the panel targets up to 5,000 (e.g., 50-5,000, or 100- 5,000, or 200-5,000, or 500-5,000) differentially methylated regions. In some embodiments, the panel targets up to 10,000 (e.g., 50-10,000, or 100-10,000, or 200-10,000, 500-10,000, or 1,000- 10,000) differentially methylated regions. In some embodiments, the panel targets up to 50,000 (e.g., 50-50,000, or 100-50,000, or 200-50,000, 500-50,000, 1,000-50,000, or 10,000-50,000) differentially methylated regions. In some embodiments, the panel targets up to 100,000 (e.g., 50- 100,000, or 100-100,000, or 200-100,000, 500-100,000, 1,000-100,000, or 10,000-100,000) differentially methylated regions. In some embodiments, the panel targets up to 150,000 (e.g., 50- 150,000, or 100-150,000, or 200-150,000, 500-150,000, 1,000-150,000, or 10,000-150,000) differentially methylated regions. The panel may optionally include probes targeting methylated and / or unmethylated sequences from a control plasmid or phage (e.g. pUC19, lambda). In some embodiments, the panel targets some or all of the differentially methylated regions listed in Table A.

[0184] In some embodiments, the panel may target at least 50 bp, at least 100 bp, at least 250 bp, at least 500 bp, at least 1 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 10 kb, at least 15 kb, at least 20 kb, at least 50 kb, at least 100 kb, at least 150 kb, at least 200 kb, at least 500 kb, at least 1 Mb, at least 5 Mb, at least 10 Mb, or at least 20 Mb total genomic regions. In some embodiments, the panel may target no more than 20 Mb, no more than 10 Mb, no more than 5 Mb, no more than 1 Mb, no more than 500 kb, no more than 200 kb, no more than 150 kb, no more than 100 kb, no more than 50 kb, no more than 20 kb, no more than 15 kb, no more than 10 kb, no more than 5 kb total genomic regions. Enrichment

[0185] In some embodiments, DNA (e.g. cfDNA) fragments or amplicons derived therefrom comprising genomic regions of interest (e.g. DMRs) may be selectively enriched using a custom panel containing oligonucleotide probes or primers designed to hybridize to the genomicregions of interest or regions flanking the genomic regions of interest. In some embodiments, the selective enrichment is after extraction of DNA and before treating the DNA with one or more methylation-sensing agents. In some embodiments, the selective enrichment is after treating the DNA with one or more methylation-sensing agents. In some embodiments, the selective enrichment is after amplification of methylation status transferred or preserved DNA.

[0186] In some embodiments, an optional PCR can be performed to amplify the selected DNA molecules comprising all or part of the genomic regions of interest (e.g. DMRs).

[0187] In some embodiments, additional methods, including but not limited to MSRE methods, can be used to further enrich DNA fragments that are fully methylated. In some embodiments, such further enrichment can be performed before subsequent enrichment by an oligonucleotide panel method disclosed herein. In some embodiments, such further enrichment can be performed after an initial enrichment by an oligonucleotide panel method disclosed herein. Hybrid Capture

[0188] In some embodiments, selective enrichment comprises fragment or amplicon capture by hybridization. In some embodiments, the panel of oligonucleotide probes are hybrid capture probes. In some embodiments described herein, such selective enrichment steps can be performed after the DNA (such as cfDNA) molecules are extracted. In some embodiments described herein, such selective enrichment steps can be performed following the addition of adapters to the DNA molecules. In some embodiments described herein, such selective enrichment steps can follow treatment of the DNA with an agent or combination of agents that discriminates between methylated and unmethylated cytosines, and optionally and typically following one or more universal amplifications.

[0189] In capture by hybridization, hybrid capture oligonucleotide probes complementary to a specific nucleic acid sequence are utilized to capture nucleic acid fragments, such as DNA or cfDNA fragments in a sample or DNA or cfDNA derived therefrom, comprising at least a portion of the sequence. In some embodiments, hybrid capture oligonucleotide probes may be strand specific. In some embodiments, probes complementary to one or both strands of specific target DNA sequences may be utilized. In some embodiments, the probes may be complementary to thetarget sequences in either a 100% or 0% methylated state. In some embodiments, both probes complementary to the target sequences in a 100% methylated state and probes complementary to the target sequences in a 0% methylated state are included in the panel. Accordingly, each strand specific probe may have two designs, assuming both a completely methylated and a completely unmethylated state of each target DNA sequence, yielding a total of four unique probes per target- sequence pair (top and bottom strands). The specific target DNA sequence for a capture probe in illustrative embodiments overlaps with or is found within a preselected target region of a sample DNA molecule such as a cfDNA. Thus, hybrid capture probes when used in methods herein can be designed to hybridize to a DNA molecule that contains at least one target region or a portion thereof. In some embodiments, the hybrid capture probes can be designed to bind to a target DNA sequence within or overlapping a target region, which in illustrative embodiments contains one or more methylated nucleic acids. In other examples, the hybrid capture probes can be designed to bind to a common region that is flanking but not overlapping the target region and that can be a common region that was added to some, most, almost all or all of the DNA in a sample, or added to all amplicons using a common sequence on at least one primer of a primer pair. In some embodiments, a hybrid capture probe or set thereof, are designed to bind to a target DNA sequence within target region, or set of target regions, respectively.

[0190] Hybrid capture probes may be added to a prepared sample and hybridized through a denature-reannealing process to form duplexes of exogenous-endogenous fragments (e.g., hybrid capture probes bound to sample DNA molecules, or DNA derived therefrom). These duplexes may then be physically separated from the sample by various means. In some embodiments, once the hybrid capture probes are removed, the sample DNA molecules, or DNA derived therefrom can be amplified. Some ways to physically remove the hybrid capture probes are by covalently bonding the hybrid capture probes to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the hybrid capture probes is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in SURE SELECT (Agilent). Thus, hybrid capture probes, for example that bind to a target DNA sequence within or overlapping a target region of a DNA molecule obtained or derived from a sample, can be covalently attached to a biotin molecule, and after hybridization with sample DNA or DNA derived therefrom, a solidsupport with streptavidin affixed can be used to pull down the biotinylated hybrid capture probes, which are hybridized to DNA molecules obtained or derived from a sample that include a target region that includes the target DNA sequence recognized by the hybrid capture probes. Thus, in some embodiments, the hybrid capture probes are immobilized, directly or indirectly to a solid support. In some embodiments, the hybrid capture probes include a binding partner, for example biotin. Hybridization conditions, including for example probe concentrations, can be adjusted and optimized to compensate for bias and prevent fragment loss.

[0191] In some embodiments, a hybrid capture panel includes one or more probe sets each having at least one hybrid capture probe, wherein each set is designed to bind to a different genomic region of interest (e.g. a DMR or a control region). In some embodiments, one or more hybrid capture probes in the one or more probe sets can bind to two or more target regions (e.g. DMRs.) In illustrative embodiments disclosed herein, probe sets are designed to ensure full coverage of the entire target region (e.g. DMR). In some embodiments, the set includes at least one hybrid capture probe or one subset of hybrid capture probes for each target region. In some embodiments, the set includes two or more hybrid capture probes or two or more subsets of hybrid capture probes for each target region. In illustrative embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more subsets of probes (each subset can include, for example four unique probes per target-sequence pair) may be designed to hybridize to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30 or more target sequences within or overlapping a target region. For example, in some embodiments, a single hybrid capture probe in a subset is designed to bind the top strand of the target sequence (1T). In some embodiments, the subset includes two hybrid capture probes that bind each target region on either the top strand (2T) or the bottom strand (2B) of the target sequences. In some embodiments, the subset includes at least one hybrid capture probe that binds the top strand of the target sequence and at least one hybrid capture probe that binds the bottom strand of the target sequence (1T1B). In some embodiments, the subset includes at least two hybrid capture probes that bind the top strand of the target sequences and at least two hybrid capture probes that bind the bottom strand of the target sequences (2T2B). Sets or subsets can also be designed as 3T3B, 4T4B, etc. In some embodiments, the subset includes both hybrid capture probe(s) designed to bind the methylated version of the target sequence(s) and hybrid capture probe(s) designed to bind the unmethylatedversion of the target sequence(s). In some embodiments, the subset includes both hybrid capture probe(s) designed to bind the 100% methylated version of the target sequence(s) and hybrid capture probe(s) designed to bind the 0% methylated version of the target sequence(s).

[0192] In some embodiments, at least two hybrid capture probes are tiled to ensure coverage of the target sequence. If two or more hybrid capture probes or two or more sets of hybrid capture probes are used for targeting two or more target sequences in or overlapping with a target region, the probes may be designed to hybridize to overlapping or separate sequences in the target region. For example, in illustrative examples, the probes may be tiled at a density between 0.8x and 3x, 0.85x and 3x, 0.88x and 3x, 0.9x and 3x, between 0.95x and 3x, between 1x and 2x, between 1x and 1.5x, between 0.94x and 1x, or at a density of 0.8x 0.85x, 0.88x, 0.94x, 0.95x, 0.96x, 0.97x, 0.98x, 0.99x, 1x, 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 2x, 2.5x, or 3x. For example, in some embodiments, the probes in a subset of probes for a target region may overlap by about 5 nucleotides, by about 10 nucleotides, by about 15 nucleotides, by about 20 nucleotides, by about 25 nucleotides, by about 30 nucleotides, by about 40 nucleotides, or by about 50 nucleotides or more, or may have no overlap, abut one another, or have a gap of between 1 to 90 bp, or between 5 to 80 bp, or between 10 to 70 bp, or between 20 to 50 bp. In illustrative embodiments, probes in a subset of probes for a target region can be designed and placed leveraging nucleosome positioning, which can be disease and tissue specific, and gaps can be introduced between the probes to achieve sub-1x tiling while minimizing downstream fragment loss during capture.

[0193] In some embodiments of any of the aspects herein, the hybrid capture probes can have a length in the range of 20 bases to 170 bases, 20 bases to 160 bases, 20 bases to 150 bases, 20 bases to 140 bases, 20 bases to 130 bases, 20 bases to 120 bases, 20 bases to 110 bases, 20 bases to 100 bases, 20 bases to 90 bases, 20 bases to 80 bases, 20 bases to 70 bases, 20 bases to 60 bases, 20 bases to 50 bases, 20 bases to 40 bases, 30 bases to 170 bases, 30 bases to 160 bases, 30 bases to 150 bases, 30 bases to 140 bases, 30 bases to 130 bases, 30 bases to 120 bases, 30 bases to 110 bases, 30 bases to 100 bases, 30 bases to 90 bases, 30 bases to 80 bases, 30 bases to 70 bases, 30 bases to 60 bases, 30 bases to 50 bases, 40 bases to 160 bases, 40 bases to 150 bases, 40 bases to 140 bases, 40 bases to 130 bases, 40 bases to 120 bases, 40 bases to 110 bases, 40 bases to 100 bases, 40 bases to 90 bases, 40 bases to 80 bases, 40 bases to 70 bases, 40 bases to 60 bases, 50 bases to 150 bases, 50 bases to 140 bases, 50 bases to 130 bases, 50 bases to 120 bases,50 bases to 110 bases, 50 bases to 100 bases, 50 bases to 90 bases, 50 bases to 80 bases, 50 bases to 70 bases, 60 bases to 140 bases, 60 bases to 130 bases, 60 bases to 120 bases, 60 bases to 110 bases, 60 bases to 100 bases, 60 bases to 90 bases, 60 bases to 80 bases, 70 bases to 130 bases, 70 bases to 120 bases, 70 bases to 110 bases, 70 bases to 100 bases, 70 bases to 90 bases, 80 bases to 120 bases, 80 bases to 110 bases, 80 bases to 100 bases, 90 bases to 120 bases, 90 bases to 110 bases, 100 bases to 165 bases, 100 bases to 150 bases, 100 bases to 140 bases, 100 bases to 130 bases, 100 bases to 120 bases, 110 bases to 150 bases, 110 bases to 140 bases, 110 bases to 130 bases, 120 bases to 150 bases, or 130 bases to 160 bases.

[0194] In embodiments where hybrid capture is used upstream of a next-generation sequencing reaction, one way to increase the number of reads that interrogate the position of interest is to decrease the length of the hybrid capture probe, as long as it does not result in bias in the underlying enriched alleles. The length of the hybrid capture probe should be long enough such that two hybrid capture probes designed to bind to two different target DNA sequences within the same target region (e.g. DMR) hybridize with near equal affinity to the target sequences. In certain embodiments, the use of shorter probes results in a greater chance that the hybrid capture probes bind to DNA molecular fragments from liquid samples, such as cfDNA. Thus, using hybrid capture, DNA molecules that include all or part of target regions (e.g. targeted DMRs) in the DNA sample can be selectively enriched.

[0195] In some embodiments, target enrichment may involve Primer Extension Target Enrichment (PETE), an NGS hybridization capture technology that uses primer extension reactions to specifically capture and release target library molecules for sequencing. Commercial kits are available for performing PETE, including the KAPA HyperPETE kit form Roche.

[0196] In some embodiments, an optional PCR can be performed to amplify the captured DNA molecules comprising sequence in target regions (e.g. DMRs).

[0197] In some embodiments, additional methods, including but not limited to MSRE methods, can be used to further enrich DNA fragments that are fully (i.e.100%) methylated. In some embodiments, such further enrichment can be performed before a target enrichment bycapture method disclosed herein. In some embodiments, such further enrichment can be performed after a target enrichment by capture method disclosed herein. Targeted amplification

[0198] In some embodiments, selected enrichment can be by targeted amplification using a panel of target-specific primers.

[0199] Methods in some aspects herein include performing one or, in some embodiments, two or more amplifications. Such amplifications in certain illustrative embodiments include at least one targeted amplification wherein at least one primer and in certain embodiments both primers of a primer pair, one or more primer pairs, or a set of primer pairs used for the amplification are each designed to bind to a specific nucleic acid sequence at, near, or within, a genomic region of interest (i.e., are target-specific primers) to generate target region (or a part thereof) amplicons. In some embodiments, methods herein include one or more universal amplifications before or after or both before and after selective enrichment.

[0200] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g., recombinase polymerase amplification (RPA) (Kersting et al.2014 Microchim Acta 181 (13–14), 1715–1723, (incorporated by reference in its entirety)), a ligase-based amplification, PCR, or a combination thereof (e.g., ligation-mediated PCR). In some illustrative embodiments, the targeted amplification is a targeted PCR(s) that is performed using a PCR reaction mixture that includes one primer pair, or in illustrative embodiments a set of primer pairs, having at least one target-specific primer in a pair, and extracted, adapted, or methylation-sensing agent treated DNA molecules, or universally amplified methylation-status preserved or transferred DNA molecules.

[0201] Typically, at least one primer of a primer pair used for targeted amplification herein, is a target-specific primer designed to bind to a specific nucleic acid sequence in or near, typically within, a genomic region of interest, which in illustrative examples can be genomic regions where epigenetic changes, such as changes in DNA methylation, are associated with or indicative of formation or presence of cancer, and can, for example, include promoter regions of tumor suppressor genes, and in some embodiments DNA methylation markers for a specificcancer type. A target-specific primer can be designed to bind to any sequence within or near a target region for amplification of the target region or a portion of the target region. In some embodiments, a target-specific primer can be designed to bind to a sequence that includes a CpG site. In other embodiments, a target-specific primer can be designed to bind to a sequence that is upstream or downstream to one or more CpG sites. One of the advantages of the methods described herein is increased flexibility in primer / probe design for targeted amplification or enrichment. In some embodiments, one primer of the one or more primer pairs or the set of primer pairs in the reaction mixture used for a targeted amplification is a universal primer and binds to a primer binding site on an adapter. Thus, for example, in such embodiments a universal primer that binds an adapter primer binding site can be used for an amplification reaction along with a target- specific primer that binds a primer binding site on a sample DNA region.

[0202] In some embodiments, a selective enrichment panel includes a set of target-specific primer pairs each having at least one target-specific primer. Target-specific primer pairs typically define the ends of target region sequence amplicons, which typically encompass at least a portion of the target region sequence(s). In some embodiments, to selectively enrich a target region or a portion thereof, a PCR can be performed using a single primer pair comprising two target-specific primers. The target region amplicon in such embodiment would extend from the sample DNA region bound by target-specific primer on a 5’ end to the sample DNA region bound by primer on the 3’ end. In some embodiments, to selectively enrich a target region or a portion thereof, a PCR can be performed using a universal primer and a target-specific primer. The target region amplicon in such embodiment would extend from the sample DNA region bound by target-specific primer on a 3’ end of one strand to the end of the sample DNA fragment on the 5’ end of that strand. In some embodiments, a PCR can be performed using a set of primer pairs to selectively enrich a set of target regions. In some embodiments, the set includes at least one primer pair for each target region. In illustrative embodiments disclosed herein, primer pair(s) for a target region are designed to ensure full coverage of the entire target region (e.g. DMR). In some embodiments, the set includes two or more primer pairs for each target region. In some embodiments, the set includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more primer pairs for each target region. If two or more primer pairs are used for a target region, the primer pairs may be designed to hybridize to overlapping or separate sequences in the target region. For example, the primerpairs may be tiled at a density between 0.8x and 3x, 0.85x and 3x, 0.88x and 3x, 0.9x and 3x, between 0.95x and 3x, between 1x and 2x, between 1x and 1.5x, between 0.88x and 1x, or at a density of 0.8x, 0.85x, 0.88x, 0.94x, 0.95x, 0.96x, 0.97x, 0.98x, 0.99x, 1x, 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 2x, 2.5x, or 3x.

[0203] In some methods herein, a universal amplification of the library of DNA treated with an agent or combination of agents that discriminates between methylated and unmethylated cytosines can be performed before the targeted amplification. Such universal amplification can be performed for example using a universal primer pair that binds primer binding sites in the adapter. Thus, in some embodiments, the methods herein include performing a universal PCR using the plurality of adapted treated cfDNA molecules, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters, before performing one or more targeted PCRs. In some embodiments, a universal amplification of selectively enriched DNA can be performed before sequencing.

[0204] In some embodiments, at least one of the primer pairs comprises a universal primer and a target-specific primer. In some embodiments, at least one of the primer pairs comprises two target-specific primers. In some embodiments, at least one of the primers comprises a sequencing tag. In some embodiments, at least one of the primers comprises a sample index. In some embodiments, at least one of the primers comprises biotin modification. In some embodiments, performing a PCR further comprises using primers comprising a sequencing tag. In some embodiments, performing a PCR further comprises using primers comprising a sample index. In some embodiments, the primers of the primer pairs are probe-dependent primers and the amplification is a target capture polymerase chain reaction. In some embodiments, the target regions each comprises one or more CpG sites differentially methylated in one or more cancers.

[0205] In some embodiments, one or both of the primer binding sites of a primer pair can include at least a portion of one of the adapter sequences. In some embodiments, one of the primer binding sites of a primer pair can include at least a portion of the adapter sequences. In some embodiments, both of the primer binding sites of a primer pair can include at least a portion of the adapter sequences. In some embodiments, neither of the primer binding sites of a primer pair include any of the adapter sequences.

[0206] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g., multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cycles. In some embodiments, amplification cycles can include at least 7, 8, 9, or 10 cycles. In illustrative embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.

[0207] Typically, in embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template (e.g., methylation-transferred DNA molecules) followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, deoxynucleotides (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In illustrative embodiments, the final concentration of each dNTP in the reaction mixture is between 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.

[0208] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgCl2), potassium chloride (KCl), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KCl can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgCl2concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM,3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgCl2 is 2.0 mM.

[0209] In some embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England Biolabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England Biolabs, Inc.).

[0210] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3´→ 5´ exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5´→ 3´exonuclease activity and strand displacement activity.

[0211] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5´→ 3´ direction and requires the presence of template and primer. This enzyme has a 3´→ 5´ exonuclease activity which is much more active than that found in DNA Polymerase I. T4 DNA polymerase lacks 5´→ 3´ exonuclease activity and strand displacement activity.

[0212] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length, between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length.

[0213] In some embodiments of any of the aspects or embodiments herein, the number of primer pairs can range from 1 to 200,000 primer pairs that each bind to one or more primer binding sequences. In some embodiments, the primer pairs are a part of a set of primer pairs. In some embodiments, the set of primer pairs comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 50, at least 100, at least 1,000, at least 5,000, or at least 10,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, or at least 1,000,000 primer pairs. In some embodiments, the set of primers range from 2 to 200,000, range from 2 to 100,000, from 2 to 10,000, from 2 to 1,000, from 2 to 100, from 2 to 50, from 10 to 100, from 50 to 100, from 100 to 200, from 100 to 500, from 100 to 1,000, from 100 to 10,000, from 100 to 100,000, from 1,000 to 100,00, or from 10,000 to 100,000 primer pairs. In some embodiments, the number of primer pairs can range from 10 to 10,000, 10 to 1,000, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 15 to 30, or 15 to 25 primer pairs.

[0214] In some embodiments, PCR is used to generate very short amplicons. cfDNA (such as fetal cfDNA in maternal serum or necrotically- or apoptotically-released cancer cfDNA) is highly fragmented. For fetal cfDNA, the fragment sizes are distributed in approximately a Gaussian fashion with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of about 100 bp, and a maximum size of about 220 bp. Methylation site(s) of interest may occupy any position from the start to the end among the various fragments originating from a particular locus. Because cfDNA fragments are short, the likelihood of both primer sites being present the likelihood of a fragment of length L comprising both the forward and reverse primers sites is the ratio of the length of the amplicon to the length of the fragment. Under ideal conditions, assays in which the amplicon is 45, 50, 55, 60, 65, or 70 bp will successfully amplify from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of available template fragment molecules. Thus, in some embodiments target amplicons generated in method herein are between 40 and 100, 40 and 75, or 45 and 70 bp in length. In certain embodiments that relate most preferably to cfDNA from samples of individuals suspected of having cancer, the cfDNA is amplified using primers that yield amaximum amplicon length of 85, 80, 75 or 70 bp, and in certain preferred embodiments 75 bp, and that have a melting temperature between 50 and 65°C, and in certain preferred embodiments, between 54–60.5°C. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. Amplicon length that is shorter than typically used by those known in the art may result in more efficient measurements of the desired methylation sites by only requiring short sequence reads. In an embodiment, a substantial fraction of the amplicons are between 25 on the low end of the range, and 100 bp, 90 bp, 80 bp, 70 bp, 65 bp, 60 bp, 55 bp, 50 bp, or 45 bp on the high end of the range. Probe-dependent primers

[0215] In some embodiments, selective enrichment can involve one or more probe- dependent primers. Probe-dependent primers (PDPs) have been disclosed (Pel, et al. “Rapid and highly-specific generation of targeted DNA sequencing libraries enabled by linking capture probes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO 2017168332A1 “Linked duplex target capture”, which are hereby incorporated by reference in their entirety). Such embodiments can be considered LTC methods. Briefly, in an LTC method, PDPs are designed to incorporate non-extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. Typically, probes of PDPs can be between 30 to 70 nucleotides in length, and include or comprise a 3’ inverted dT base to inhibit polymerase extension. In some embodiments, probes are designed to cover the desired region with zero gap between forward and reverse probes. In some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60, between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the probes of a PDP pair comprises a sample index.

[0216] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a genomic region of interest on a sample DNA molecule to enrich nucleic acid molecules comprising the genomic region of interest or copies thereof. In some embodiments, at least one of the probe binding regions can include one or more CpG sites. Insome embodiments, both of the probe binding sites of the probe binding regions can include one or more CpG sites.

[0217] Typically, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on the appended adapter. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desired melting temperature. In illustrative embodiments, the primer is a low melting temperature universal primer complementary to a portion of the ligated adapter.

[0218] The primer can be tailed or untailed depending on the specific requirements. In some embodiments, the universal primer comprises an A tail. In some embodiments, the universal primer is blunt ended. The length of the primers of the PDP can range from 5 to 40 nucleotides in length. In certain embodiments, the PDP primers are between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.

[0219] Typically, probe dependent primers comprise a linker between the probe and the primer. Probe and primer portions of the PDP are typically linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, the linker is an oligosaccharide. In some embodiments, the linker is a lipid. In some embodiments, the linker is a hydrocarbon. In some embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof. Sequencing

[0220] The methylation status of the original sample DNA can be determined by sequencing the selected DNA or its derivatives. In some embodiments, the sequencing is next- generation sequencing or high-throughput sequencing.

[0221] DNA sequencing techniques, particularly high throughput next-generation sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed NOVASEQ (Illumina), MISEQ (Illumina), HISEQ (Illumina), ION TORRENT (Life Technologies), GENOME ANALYZER ILX (Illumina), GS FLEX+(Roche / 454 Life Sciences) etc., can be used for determining the methylation status, somatic mutation status, fragmentomics, and / or quantifying nucleic acids prepared by the methods described herein. Other short-read sequencers, including for example BGI T7, Element Biosciences AVITI, PacBio ONSO, Ultima Genomics UG100, and Singular Genomics G4 platforms, can also be used. High throughput genetic sequencers are amenable to the use of barcoding (such as sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. Methods as described herein that utilize NGS detection, in some embodiments can have an average depth of read of at least 0.1, 0.5, 1, 10, 50, 100, 200, 500, 1000, 2000, 2900, 3000, 3500, 4000, 5000, 10,000, 50,000, 75,000, 100,000, 130,000, 150,000, 175,000, or 200,000.

[0222] Methods herein can include analyzing data obtained from next-generation sequencing techniques. In some embodiments of methods herein, deaminated or MSRE treated cfDNA molecules can be subjected to sequencing using next-generation sequencing techniques. Nucleic acid sequencing data can be generated for amplicons of target sequences created by PCR, for example a multiplex targeted PCR (mPCR) or a universal PCR before or after target-sequence capture or enrichment. For a skilled artisan, algorithm design tools are available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.

[0223] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows-Wheeler alignment software, Bwa mem function (BWA, Burrows-WheelerAlignment Software (see, Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics) on single end mode using pear merged reads to the hg19 genome. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.

[0224] Methods herein can include a background error model that can be constructed using normal, or healthy liquid samples, in illustrative embodiments, normal, or healthy plasma samples, which are sequenced on the same sequencing run to account for run-specific artifacts. In some embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal, or healthy liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to 400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g., plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10 are removed from the data analysis. For each base substitution of every genomic loci, the DOR weighted mean and standard deviation of the error can be calculated.

[0225] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200- nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotidesis 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).

[0226] In some embodiments, the uniformity in DOR can be measured using standard methods such as, but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90th-95thpercentile. For this measurement, the loci are sorted in descending DOR order. In illustrative embodiments, a DOR distribution using the 90th-95thpercentile contains 5 percent of reads. The reads of all loci between the 90thpercentile and 95thpercentile can be counted and divided by the total reads for all loci.

[0227] In some embodiments, the magnitude of the DOR slope can be less than 0.005, 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90th-95thpercentile can be between 0.2 and 9 percent, such as between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent, 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to 6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent, 3 to 7 percent, 3 to 6 percent, 0.2 to 1.0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent, 3 to 4 percent, 4 to 5 percent, 5 to 6 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method or the amplification steps in the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, or 100,000 non- identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90th-95thpercentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15,000 non-identical amplicons.

[0228] In some embodiments of methods herein, in addition to analyzing an altered (increased or decreased) methylation levels in a sample, one or more other analytes, biomarkers, or signals (e.g., fragmentomics or genomic alterations such as SNVs, insertions, deletions, fusions, rearrangements, aneuploidies, or other copy number variations) can be analyzed if desired. One or more of the additional analytes, biomarkers, or signals can be used to increase the sensitivity and / or specificity of the sample calling (such as determining the presence or absence of cancer or an increased risk for cancer, classifying the cancer type or tissue of origin, or the staging or prognosis of the cancer). Classifier

[0229] In some embodiments, the methods herein include a classifier for disease or phenotype detection. In some embodiments, the classifier classifies a sample based at least in part on DNA methylation signals. In some embodiments, the methods herein analyze methylation signals of cfDNA from a subject for disease or phenotype detection using a machine-learning model that employs biology-informed target and feature selection and feature transformation methods. In some embodiments, the methods herein analyze methylation signals of cfDNA from a subject for disease or phenotype detection using a model-free method or a regression model that relies on a ratio of hypermethylated target regions to total target regions.

[0230] To address the problem of technical and biological noises often associated with methylation biomarkers, the methods herein leverage adjacent CpGs whose methylation states are correlated, focusing on regions with high CpG density and are recurrently differentially methylated between samples having or suspected of having a cancer, disease, phenotype, condition and normal samples (samples lacking the cancer, disease, phenotype, condition) and exhibit low to zero or nearly zero biological noise. Such recurrently differentially methylated regions having multiple adjacent CpGs are expected to be good methylation biomarkers.

[0231] In illustrative embodiments, criteria for selection of suitable DMRs for inclusion in a classifier for disease or phenotype detection include factors such as high correlation of methylation status of adjacent CpGs, high CpG density such that sufficient CpGs are capturedwith a typical cfDNA fragment, linearity of methylation signal with Signatera VAF, and virtually zero background signal in health / normal samples to reveal ctDNA at VAF of about 0.01%.

[0232] In some embodiments, cumulative probability of fragments with a certain fraction of methylated CpGs occur in healthy samples (such as plasma or cfDNA samples) vis-a-vis in diseased samples (such as CRC tissue samples) can be analyzed to understand the signal-to-noise ratio of fragments with different number of CpGs and different methylation fractions. Biological or technical background noise is reduced with hypermethylated or hypomethylated fragments with more CpGs. In some embodiments, a fragment is considered “hypermethylated” if it has at least 3, 4, 5, 6, 8, 10, 15, or 20 CpGs, and at least 55%, 60%, 70%, 80%, 90%, or 95% the CpGs are methylated. In some embodiments, two or more different methylated CpG percentage thresholds may be applied to two or more different targets.

[0233] Sufficient target coverage is typically necessary for feature extraction. Target coverage increases with DNA input amount. In some embodiments, the number of on-target fragments per target for a panel of targets is between 0 and 10,000, with a median of 1300 to 1500 on-target fragments per target. In some embodiments, targets with at least 1000, at least 1500, or at least 2000 on-target fragments per target, or a median of at least 1300, at least 1350, at least 1400, at least 1450, or at least 1500 on-target fragments per target are used for feature extraction. In some embodiments, MITs can be added to the starting DNA molecules and incorporated in the analyses to provide increased target coverage.

[0234] Methods herein can be used to increase molecular recovery and target coverage. In some embodiments, molecular recovery can be improved by molecular barcoding (e.g. using MIT- adapters). In some embodiments, strand-aware deduplication can be performed to remove only true PCR duplicates. For example, as a result of conversion, the top and bottom strands of an original DNA molecule can be distinguished and both strands can be considered instead of having one of the strands removed as a PCR duplicate. In some embodiments, methylation-aware deduplication can be performed such that fragments can be distinguished by methylation content to further improve molecular recovery. The increased recovery will increase the number of fragments per target and hence improve the sensitivity of the classifier. In illustrative embodiments, methylation-aware deduplication is performed by, for example, computing theprobability that two DNA fragments have different methylation patterns due to protection and conversion errors. In some embodiments, the methods herein account for multiple testing across the total number of reads per sample by calculating a q-value to control the false discovery rate due to classifying fragments as originating from two distinct DNA molecules while in reality they are PCR duplicates with methylation conversion / protection errors.

[0235] In illustrative embodiments, a noise model is used (for example during the initial target selection or a training phase) to select a subset of targets from a sequencing panel that have low or ultra-low background noise in healthy subjects. In some embodiments, background noise for each target region is used in an initial binary filtering. In some embodiments, target performance is evaluated based on mean baseline of target differentially-methylated fragments (TDMF). In some embodiments, the TDMF mean baseline for a panel of targets is between 1e-6 and 1e-1. In some embodiments, the TDMF mean baseline for one or more of the selected targets is below 1e-4, and in some embodiments below 1e-5. In some embodiments, target performance is evaluated based on the target’s positive likelihood ratio (LR+), computed per target as TPR / FPR. In some embodiments, LR+ of a panel of targets is between about 1e-1 to about 1e+3. In some embodiments, LR+ for one or more of the selected targets is above 1e+1, and in some embodiments above 1e+2.

[0236] In some embodiments, an error model reduces false positives using informative priors from existing biological data. In some embodiments, the error model reduces false positives by filtering out DNA fragments with likely incomplete conversion or MSRE digestion of unmethylated CpG sites. In some embodiments, the likely incomplete conversion or digestion of unmethylated CpG sites on a DNA fragment is determined based on the apparent methylation status of one or more CHG and / or CHH sites (where “H” refers to A, T, or C), which are usually not methylated in mammals (e.g. humans), on the same fragment. In some embodiments, fragments having one or more CHG and / or CHH sites that appear to be methylated are filtered before further analysis.

[0237] FIG.27 illustrates the characteristics of an example target evaluated using the methods herein. The target in this example has a length of 71 bp, 12 CpGs, median coverage (on- target fragments) of 857, median fragment count of 778, baseline / healthy TDMF mean of 2.8e-5,CRC TDMF mean of 3.6e-3, LR+ score of 130, and selection frequency (frequency of selection by the classifier in training phase) of 1.3%. X-axis shows a segment of the human genome. The bracket indicates the intended target region designed using the methods herein. Y-axis is the posterior probability of methylation status. Red (top) is the average methylation signal (methylation beta value) observed in CRC tissue samples. Blue (bottom) shows the average methylation signal observed in healthy plasma samples. As shown in FIG.27, there is good LR+, and good differentiation between average background / baseline methylation signal and average CRC methylation signal in the target region, with the average baseline signal nearly zero. However, in this example, the classifier only selected the target 1.3% of the time using the methods herein, likely due to the relatively low median fragment count and median coverage, which can contribute to more uncertainty in estimating the signal-to-noise ratio.

[0238] FIG.28 illustrates the characteristics of another example target evaluated using the methods herein. The target in this example has a length of 267 bp, 26 CpGs, median coverage (on- target fragments) of 1141, median fragment count of 2153, baseline / healthy TDMF mean of 5.0e- 6, CRC TDMF mean of 1.5e-3, LR+ score of 332, and selection frequency of 100%. The bracket indicates the intended target region designed using the methods herein. Y-axis is the posterior probability of methylation status. Red (top) is the average methylation signal (methylation beta value) observed in CRC tissue samples. Blue (bottom) shows the average methylation signal observed in healthy plasma samples. As this example target illustrates, the larger targeted region resulted in more on-target fragments. The target has 5-fold lower TDMF mean baseline, slightly lower TDMF mean CRC, and much higher LR+ as compared to the example target shown in FIG. 27, and was selected 100% of the time by the classifier using the methods herein.

[0239] FIG.29 illustrates the characteristics of yet another example target evaluated using the methods herein. The target in this example has a length of 270 bp, 29 CpGs, median coverage (on-target fragments) of 1359, median fragment count of 2727, baseline / healthy TDMF mean of 2.8e-4, CRC TDMF mean of 1.1e-3, LR+ score of 4.1, and selection of 10%. The bracket indicates the intended target region designed using the methods herein. Y-axis is the posterior probability of methylation status. Red (top) is the average methylation signal (methylation beta value) observed in CRC tissue samples. Blue (bottom) shows the average methylation signal observed in healthy plasma samples. As shown in FIG.29, there is elevated background noise in parts of the target,resulting in low LR+, and was selected by the classifier using the methods herein only 10% of the time despite having good coverage median and fragment count median. As this example indicates, hypermethylation may also occur in healthy colon tissue or other organs.

[0240] As illustrated above, background noise level in healthy samples is the most crucial aspect of what makes a biomarker useful for the classifier. Accordingly, background noise level is a key element of how the classifier selects different targets for features being used for the classification problem. FIG.30 shows the mean level of the background noise in healthy samples (Y-axis) and associated confidence intervals for all the targets selected at some time by the classifier (X-axis) in an illustrative embodiment using a cross-validation (CV) model as described herein, ordered by selection frequency. Higher coverage typically allows better estimation of the background noise level. Targets with low coverage may appear to have good features but can contribute to high variability. For example, the reason why it appears to have good features in the training data set may just be because only very few on-target fragments have been analyzed. Accordingly, in some embodiments, as shown in FIG.30, the noise model computes TDMF mean baseline confidence intervals to reduce variability from coverage.

[0241] In some embodiments, the classifier is a logistic regression classifier (e.g. ridge, elastic net, lasso). In some embodiments, the classifier is a supportive vector machine (SVM). In some embodiments, the classifier is a random forest classifier.

[0242] In some embodiments, the classifier uses hypermethylated fragment fraction per target as features. In some embodiments, the input to the classifier is a transformed version of the hypermethylated fragment fraction for each selected target. In some embodiments, the classifier uses a grid-search during cross-validation. In some embodiments, the classifier selects all targets for training the machine-learning model. In some embodiments, the classifier selects the best N targets for training the machine-learning model based on a training set during cross-validation and then makes predictions in holdout test sets. In some embodiments, targets are selected from ranking by the upper confidence interval bound. In some embodiments, the number of selected targets is a hyper-parameter chosen within inner grid search. As discussed above, limited coverage can lead to low hypermethylated fragment count and challenging feature matrix sparsity. Toaddress this problem, in some embodiments, the logistic regression classifier uses logit- transformed hypermethylated fragment fraction per target as features.

[0243] In some embodiments, a classifier independently trained on one cohort is applied to data from another cohort and vice versa. The methylation classifier using one or more of the methods described herein generalizes robustly across sample batches, time points (e.g. early cancer detection vs cancer recurrence detection), and signal source (e.g. primary tumor vs metastases). As such, in some embodiments, the same classifier can be used for early cancer detection and cancer recurrence monitoring products.

[0244] In some embodiments, the classifier and the target and feature selection methods can be further adjusted based on one or more other characteristics of the subject, such as age, gender, ethnicity, smoking status, polygenic risk score, and medical history (including for example past diagnoses, current medications, allergies, and chronic conditions). In some embodiments, the classifier may use additional target-specific feature parameters to enhance performance.

[0245] In some embodiments, the classifier and methods disclosed herein can combine analysis of methylation signals with analysis of one or more other analytes or biomarkers (for example, genomic or somatic mutations (identified for example using targeted or whole genome sequencing), proteins, extracellular vesicles, lipids, miRNA, fragmentomics, immune profiling), either in parallel or as a reflex strategy to improve sensitivity and / or specificity in detecting a disease, phenotype, or condition. In some embodiments, the disease / phenotype detection sensitivity threshold for the initial assay may be set low to allow maximum detection, followed by a second assay using the same or a different analyte / signal to ensure specificity. In some embodiments, the classifier and methods disclosed herein combine two or more sequential measurements of one or more signals (such as methylation signal) over multiple timepoints.

[0246] The methods disclosed herein can be used to diagnose, detect, or monitor a disease (including cancer, and whether symptomatic, pre-symptomatic, or asymptomatic), phenotype, or condition in any eukaryotic organism in which methylation is regulated in its genome (such as humans and domesticated animals, including pets (e.g. dogs, cats, birds, hamsters), livestock (e.g. pigs, sheep, cattle), and beasts of burden (e.g. horses, camels, elephants). The methods disclosedherein can be used in virtually any applications involving distinguishing methylation status of nucleic acid molecules in two or more states, including for example cancer or cancer-recurrence detection / diagnosis, other disease detection / diagnosis, therapy selection, non-invasive prenatal testing (NIPT), organ health or organ transplant evaluation, prediction and monitoring, and veterinary applications (e.g. aging or health status evaluation, prediction and monitoring of domesticated animals). Controls & Standards

[0247] In various embodiments and typically, controls and standards are used, for example, for quality control and optimization of the assays described herein and normalization of the readout signals, and to troubleshoot issues that may arise. For example, in some embodiments, fully methylated and / or fully unmethylated sequences, or sequences with and / or without MSRE / MDRE recognition sites, are used as controls and standards with the methods herein. In some embodiments, the control sequences comprise endogenous sequences (e.g. DNA sequences from the genome of the subject). In some embodiments, the control sequences comprise exogenous sequences (e.g. plasmid (e.g. pUC19), bacteriophage (e.g. lambda phage), in vitro methylated cfDNA molecules, or synthetic molecules). In some embodiments, one or more of the control sequences are spiked-in and added to one or more reactions disclosed herein. In some embodiments, control sequences having one or more methylated CGs and / or one or more unmethylated CGs are incorporated in one or more adapters appended to the DNA (e.g. cfDNA) fragments or their derivatives to be prepared and analyzed.

[0248] In some embodiments, exogenous sequences (e.g. in vitro methylated cfDNA, methylated pUC19, unmethylated lambda) are spiked in as controls for methylation-sensing assays such as chemical (e.g. bisulfite) or enzymatic (e.g. EM-seq) conversion and MSRE / MDRE digestion. In some embodiments, adapters having one or methylated CGs and / or one or more unmethylated CGs can be added to the DNA molecules before subjecting them to methylation assays. In some embodiments, regions in a subject’s genome having a known methylation pattern or with no MSRE sites can be used as controls for methylation-sensing assays.

[0249] In some embodiments, quantifying target coverage for at least some of the targets includes normalizing the target coverage with control sequences. Depth of read for each of the target can be normalized relative to a depth of read for a normalization sequence. DNA sequences used for normalization can be derived from the subject’s genomic DNA or from control plasmids, synthetic DNA molecules, etc.

[0250] In some embodiments, the normalization sequence can be a 100% methylated contrived DNA molecule derived from a control genomic DNA sample, in some embodiments a pUC19 plasmid or a lambda phage. A fully methylated (100% methylated) synthetic sequence, for example, can also be considered as a normalization sequence. In some embodiments, the normalization sequence can be a non-MSRE DNA molecule having no MSRE recognition site for any of the MSRE or MSREs used. In some embodiments, the normalization sequence can be a non-MDRE DNA molecule having no MDRE recognition site for any of the MDRE or MDREs used. In some embodiments, non-MSRE / non-MDRE sequences for normalization have no MSRE or MDRE recognition site on the prospective amplicons. In some embodiments, non-MSRE / non- MDRE sequences for normalization have no MSRE or MDRE recognition site on and + / - 200bp, + / - 150bp, + / - 100bp, + / - 50bp, + / - 40bp, or + / - 40bp of the prospective amplicons. Such a non- MSRE / non-MDRE DNA molecule can be derived from a first subject. A non-MSRE / non-MDRE DNA molecule as per methods herein can be a synthetic sequence.

[0251] A spike-in control DNA sample can be a fully methylated genomic, plasmid, bacteriophage, or synthetic DNA sample, a fully unmethylated genomic, plasmid, bacteriophage, or synthetic DNA sample, or a combination of both. A spike-in control sample used for normalization, in some embodiments, can be 0.05%, 0.1%, 0.2%, 0.5%, 0.7%, 0.8%, 1.0%, 1.2%, 1.4%, 1.6%, 1.8, 2%, or more fully methylated DNA by mass. A spike-in control sample can be fully methylated, fully unmethylated, non-MSRE DNA, or non-MDRE DNA that range between 0.005 to 2%, 0.01 to 2%, 0.05 to 2%, 0.1 to 2%, 0.5 to 2%, 1 to 2%, 0.005 to 1.5%, 0.005 to 1.2%, 0.005 to 1%, 0.005 to 0.8%, 0.005 to 0.5%, 0.005 to 0.3%, or 0.005 to 0.2% by weight. Methods herein can include more than one, for example 2, 3, 4, 5, or more, control or spike-in control sample.

[0252] In most instances, in vitro methylation may not result in 100% methylation. Commercially available methylated controls typically are only up to 95% methylated. In some embodiments, methylated controls (e.g. in vitro methylated cfDNA, fully methylated pUC19 or lambda) are treated with MSRE to remove unmethylated fragments before use.

[0253] In some embodiments, one or more sequences from human genome are added to a sequencing panel as controls. For example, in some embodiments, SNP tracers are added to the panel for sample tracking. Quantification & Limit of Detection

[0254] Exemplary methods herein can be used to detect a small amount of differentially methylated DNA (e.g. ctDNA) molecules in a sample. The methods herein are particularly useful for methylation fraction estimation and hence quantitative.

[0255] Measurement of ctDNA has been established as an indicator for the presence of cancer cells in a body and as a prognostic biomarker of cancer disease burden. One way to measure or estimate the amount of ctDNA in a sample is by determining variant allele frequency (VAF), typically characterized as the measurement of the specific variant allele proportion within a genomic locus, of one or more DNA variants (e.g. SNVs) present in tumor cells but not in normal cells. (See e.g., WO2019200228, incorporated by reference herein in its entirety). In NGS- based detection methods, VAF is typically calculated as the percentage of sequence reads or fragments mapped to a target variant locus divided by all the sequence reads or fragments mapped to that locus. The VAF for a sample is typically calculated as the mean VAF of all the targeted variant loci (e.g. targeted SNVs). VAF can be used as a surrogate measure of the percentage of ctDNA in the cfDNA sample. In certain embodiments described herein, Signatera VAF is used. Signatera VAF is VAF determined by the Signatera method as described in, for example, Coombes et al., Clinical Cancer Research 25(14):4255-4263 (2019); Christensen et al., J. Clin. Oncology 37(18):1547-1557 (2019); Shaw et al., JCO Precis Oncol.8:e2300456 (2024); Chen et al., BMC Cancer 24:1016 (2024), which are incorporated herein by reference in their entireties.

[0256] As described herein, differential methylation patterns can serve as an accurate biomarker both for cancer detection and for predicting the tissue of origin. As shown inembodiments herein, differential methylation allele fraction (DMAF), which estimates the fraction of ctDNA using differentially methylated DNA fragments in a plasma sample, strongly correlates with sample VAF regardless of clinicopathologic features such as disease stage, histology, age, or sex and can be used as an alternative tool to select biomarkers or predict ctDNA fraction. In some embodiments, DMAF estimates the fraction of differentially methylated DNA fragments across two or more target regions within a DNA sample. In some embodiments, DMAF estimates the fraction of differentially methylated DNA fragments across all target regions within a DNA sample. Exemplary methods herein, in some embodiments, can detect as low as 2.0%, 1.0%, 0.5%, 0.1%, 0.05%, 0.02%, 0.01%, 0.005%, 0.002%, or 0.001% DMAF. In some embodiments, the VAF-DMAF correlation is between 0.8 to 0.98 or between 0.85 to 0.95. In some embodiments, the VAF-DMAF correlation is about 0.90. DMAF is also highly correlated with mean tumor molecules (MTM), another metric that can be used to estimate or quantify the fraction of ctDNA in a sample.

[0257] In some embodiments, methods herein can detect fully methylated DNA molecules present at 2.0%, 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, or 0.001% (by mass) or more in a mixture of DNA molecules. In some embodiments, methods herein are capable of differentiating samples having 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, or 0.001% (by mass) or more contrived fully methylated DNA molecules from samples having 0% of contrived fully methylated DNA molecules.

[0258] In certain embodiments, the methods herein are capable of detecting ctDNA when it is present in 5%, 2.5%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.002%, 0.001%, or more of the total cfDNA in a sample. In some embodiments, methods herein detect or are capable of detecting ctDNA from a sample when it is present at a range between 0.1%, 0.05%, 0.02%, 0.01%, 0.005%, 0.002%, or 0.001% on the low end and 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2.5%, 1%, or 0.5% on the high end of the range of total cfDNA in the sample. The percentage of ctDNA in a cfDNA sample can be determined or approximated by mean sample VAF or by DMAF.

[0259] In some embodiments, measurements can be adjusted for bias, such as bias due to differences in amplification efficiency or adjusted for sequencing errors. In some embodiments,differentiation between methylated or hypermethylated and unmethylated target sequences can be analyzed using a control sequence that is fully methylated, for example a fully methylated plasmid control (e.g. pUC19), for normalization of quantitative results. In some embodiments, the differentiation between the methylated or hypermethylated and unmethylated target sequences can be achieved after normalization of a detected and typically quantified signal using one or more control sequences. Comparison of Two or More Samples

[0260] Methods herein, in some embodiments, can include preparing 2 or more libraries of cfDNA molecules. In some embodiments, methods as described herein can include preparing 3, 4, 5, 6, 7, 8, 9, 10 or more libraries of cfDNA molecules. In some embodiments, methods as described herein can include preparing 2 libraries of cfDNA molecules, such methods further comprise performing the method on a second sample. In some embodiments, the second sample is from a second subject. In some embodiments, the second sample is derived from a cell line. In some embodiments, the method is performed on the first and second samples simultaneously. In some embodiments, methods as described herein can include comparing 2 libraries of cfDNA molecules. In some embodiments, such methods further comprise determining and comparing the methylation status of a target region of interest in each of the libraries. In some embodiments, such methods further comprise determining and comparing the somatic mutation status in each of the libraries. In some embodiments, a first library is derived from a subject suspected or at risk of having a disease, and a second library is derived from a subject that is not suspected or at risk of having the disease. In some embodiments, the first and second libraries are derived from the same subject. In some embodiments, the library is derived from a sample collected at a first timepoint and the second library is derived from a sample collected at a second timepoint. In some embodiments, the first timepoint precedes the second timepoint. In some embodiments, the method comprises comparing the methylation status and / or somatic mutation status of a genomic region of interest derived from the first sample to that from a second sample using the same method. In some embodiments, the determining comprises comparing the methylation status and / or somatic mutation status of a target region of interest to a preset threshold amount. Kits

[0261] Additional embodiments of the invention described herein relate to kits for preparing nucleic acids from a biological sample using the methods described herein. In some embodiments, the kits may comprise one or more of nucleotides, e.g., dNTPs, polymerase, e.g., BST polymerase, and a methyl transfer agent, e.g., DNMT1, necessary to carry out the claimed invention. In some embodiments, the kit may comprise primers, such as universal primers, and / or probes or primers specific for target regions.

[0262] In some embodiments, the kit may further comprise reagents necessary for nucleic acid processing. For example, the kit may comprise adapters, barcodes, molecular indexes, sample indexes. In some embodiments, the kit may comprise ligases and or primers sued to append the barcodes and indexes to the template DNA.

[0263] In some embodiments, the kit may further comprise reagents necessary for downstream applications. For example, in some embodiments, the kits may comprise reagents necessary for enrichment of target regions, such as hybrid capture probes or target specific oligonucleotides. In some embodiments, the kit may comprise reagents required for methylation analysis such as deaminating reagents or MSREs. In some embodiments, the kits may comprise reagents necessary for somatic mutation analysis, such as probes and primers. In some embodiments, the kits may comprise reagents necessary for sequencing reactions.

[0264] In some embodiments, the kit may further comprise tubes, tools, or devices necessary for performing the methods described herein. In some embodiments, the kit may further comprise instructions for performing the methods described herein. In some embodiments, the kits may include one or more devices, such as a microfluidic device, on which the method can be performed. In some embodiments, the kits may be used in combination with the diagnostic box, as disclosed herein. Diagnostic Box

[0265] In an embodiment, the present disclosure comprises a diagnostic box that is capable of partly or completely carrying out any of the methods described in this disclosure. In an embodiment, the diagnostic box may be located at a physician’s office, a hospital laboratory, or any suitable location reasonably proximal to the point of patient care. The box may be able to runthe entire method in a wholly automated fashion, or the box may require one or a number of steps to be completed manually by a technician. In an embodiment, the box may be able to analyze at least the genotypic data measured on the samples from the subject. In an embodiment, the box may be linked to means to transmit the genotypic data measured on the diagnostic box to an external computation facility which may then analyze the genotypic data, and possibly also generate a report. The diagnostic box may include a robotic unit that is capable of transferring aqueous or liquid samples from one container to another. It may comprise a number of reagents, both solid and liquid. It may comprise a high throughput sequencer. It may comprise a computer.

[0266] Referring to FIG.47, in various embodiments, a system 100 may include a computing system 110 (which may be or may include one or more computing devices, co-located or remote to each other), a sequencing system 160 (capable of, e.g., generating DNA sequencing data), an information system 170 (such as a management or clinical information system that may record and / or provide information regarding samples or patients), vessels 175 (e.g., for holding samples), sample processing system 180 (e.g., a system that may include robotic components 182 for preparing, moving, and / or processing samples), and imaging and / or therapy systems 190 (which may include e.g., treatment unit 192 and imager / sensors 194 for imaging patients using various imaging modalities, providing treatments such as radiotherapy, and / or performing other evaluation or therapeutic procedures based on what is learned from patient samples). The computing system 110 (e.g., one or more computing devices) may be used to control and / or exchange signals and / or data with sequencing system 160, sample processing system 180, information system 170, and / or imaging and / or therapy systems 190, directly (e.g., through wireless and / or wired communication) or indirectly via another component of system 100 (e.g., via any combination of wireless and / or wired communication). In certain embodiments, computing system 110 may be used to control and / or exchange data or other signals with sequencing system 160, information system 170, sample processing system 180, and / or imaging and / or therapy systems 190. The computing system 110 may include one or more processors and one or more volatile and / or non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated.

[0267] The computing system 110 may include a controller 112 that is configured to exchange control signals with sequencing system 160, information system 170, sample processingsystem 180, and / or imaging and / or therapy systems 190, and / or any components thereof, allowing the computing system 110 to be used to control, for example, processing of samples, capture of images, acquisition of signals by sensors, positioning or repositioning of samples being tested or imaged, recording or obtaining other information, and / or running assays or tests.

[0268] A transceiver 114 allows the computing system 110 to exchange readings, control commands, and / or other data or signals, wirelessly or via wires, directly or indirectly via networking protocols, with, for example, other components of system 100. One or more user interfaces 116 allow the computing system 110 to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, biometric scanners, etc.) and provide outputs (e.g., via a display screen, audio speakers, light emitters, etc.) with users. The computing system 110 may additionally include one or more databases 118 for storing, for example, data acquired from one or more systems or devices, signals acquired via one or more sensors, sequencing data, biomarkers, images, etc. In some implementations, database 118 (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote (e.g., via “cloud computing”) and in communication with computing system 110, sequencing system 160, information system 170, sample processing system 180, imaging and / or therapy systems 190, and / or components thereof.

[0269] Imaging and / or therapy systems 190 may include a treatment unit 192, which may include any systems and / or devices employed to administer a treatment to a patient, such as a radiation therapy system. Imager and / or sensors 194 may be or may include, for example, any system or device that is involved in, for example, capturing images of patients and / or samples. Imagers may include detectors for visible light and / or light in any frequencies of interest, such as (but not limited to) the spectrum from infrared to ultraviolet. Imagers may employ any suitable optical components (e.g., lenses, mirrors, filters, beam splitters, prisms, diffusers, diffraction gratings, etc.), digital components (e.g., charge-coupled devices (CCDs), as well as an area for placement of samples, computing components (e.g., one or more processors, such as digital signal processors) to process and / or pre-process images, etc. Imagers may be, or may employ, any microscopes or microscopy systems (e.g., confocal microscopy), tools, and / or techniques that will provide the desired imaging data. The imager and / or sensors 194 may have the capability of receiving control signals from a computing device or system to, for example, initiate or ceaseimage capture, and / or return images or other imaging data or status signals to the computing device or system (e.g., computing system 1110). Imaging and / or therapy systems 190 may include any tools for obtaining additional data about samples and / or patients (such as spectrometers, chromatographs, etc.). Sensors may be used to detect, for example, other aspects of samples and / or patients, such as temperature, humidity, location, etc.

[0270] Sample processing system 180 may include any components used to automate preparation of samples and running tests on samples. Robotics 182, for example, may include any combination of actuators, stepper motors, servomotors, control system, robot arms, end effectors like grippers and manipulators, etc. Robotics 182 may be, or may comprise, for example, an automated robotics system with process automation software. In various embodiments, robotics 182 may employ vessels 175, such as vials or plates, that can be moved and manipulated, for example, for positioning of samples to be evaluated using sequencing system 160 or components thereof. Sample processing system 180 may include components that may be used to maintain the integrity of samples during storage, transport, and testing, such as heaters, passive or active coolers, etc. Sensors may be used by sample processing system 180 to, for example, monitor, guide, and evaluate the progress of tests (e.g., by detecting temperature, fill level of vessels, or other states or conditions). Sequencing system 160 may include sequencing platforms or devices as well as various analytical tools used to process data from the sequencing platforms. Analytical tools may include software and / or hardware used to generate computer files with raw and / or processed sequencing data.

[0271] In various implementations, components of system 100 may be rearranged or integrated in other configurations. For example, computing system 110 (or components thereof) may be integrated with one or more of the sequencing system 160, sample processing system 180, and / or components thereof. The sequencing system 160, sample processing system 180, and / or components thereof may be directed to a vessel 175 in which a sample can be situated (e.g., so as to test or evaluate biological samples). In various embodiments, the vessel 175 may be movable (e.g., using any combination of motors, magnets, etc.) to allow for positioning and repositioning of samples (such as micro-adjustments for positioning of samples to be sequenced). It is also noted that not all components of system 100 are required to implement the disclosed approach, and in various embodiments, only a subset of the components of system 100 may be employed. Forexample, in various embodiments, computing system 110 may obtain and process data that was obtained via sequencing system 160 or another system that is or is not in direct communication with the computing system 110. In certain embodiments, data may be obtained for samples that were not processed in an automated fashion but rather manually by one or more users.

[0272] A data acquisition unit 124 may retrieve, acquire, or otherwise obtain various data, such as sequencing data. The data acquisition unit 124 may, for example, obtain data stored in database 118, sequencing system 160, information system 170, sample processing system 180, and / or imaging and / or therapy systems 190. An interaction unit 126 may interact (e.g., via user interfaces 116) with users (e.g., laboratory technicians, data scientists, clinicians, etc.) to obtain information or commands needed for system 100 or components thereof to function, or to start operations, repeat processes, and / or stop operations. In certain embodiments, the data acquisition unit 124 may obtain data from users via interaction unit 126.

[0273] A preprocessor 128 may perform any number of steps disclosed herein in preprocessing sequencing data. Machine learning platform 130 may be configured to train and update machine learning models, as further discussed herein. Machine learning platform 130 may, for example, employ certain machine learning techniques and algorithms to train and update predictive models. Machine learning platform 130 may include a training engine 132 which may, for example, fit models to sequencing data. The training engine 132 may, for example, obtain sequencing data from or via data acquisition unit 124, sequencing system 160 or components thereof, and / or information system 170. Inference engine 134 may apply models (e.g., models obtained from machine learning platform 130 or components thereof) to generate outputs as disclosed herein. Biomarker generator 140 may perform analyses on data from machine learning platform 130 to, for example, generate biomarkers to be used for various conditions or sub- populations.

[0274] A reporter 142 may generate reports that include, for example, information on sample sequencing (e.g., via sequencing system 160 or components thereof), imaging or administration of treatments (e.g., via imaging and / or therapy systems 190 or components thereof and / or via information system 170), model training (e.g., via machine learning platform 130), and biomarker generation (e.g., via biomarker generator 140). For example, reporter 142 may provideor identify trained and / or updated models, and / or biomarkers. In certain embodiments, reporter 142 may obtain information to be reported from, for example, database 118 and / or information system 170. Data that may be reported may be, for example, transmitted to another system or device (which may or may not be part of system 100), saved in non-transitory computer-readable storage media (e.g., database 118, information system 170, and / or elsewhere), and / or presented or otherwise provided to users (e.g., via a display device that is part of user interfaces 116).

[0275] While the embodiments of the present disclosure are amenable to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail below. The intention, however, is not to limit the disclosure to the particular embodiments described. On the contrary, the disclosure is intended to cover all modifications, equivalents, and alternatives falling within the scope of the disclosure as defined by the appended claims.

[0276] The terms and expressions which have been employed herein are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by illustrative aspects, exemplary aspects and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims. The specific aspects provided herein are examples of useful aspects of the present invention and it will be apparent to one skilled in the art that the present invention may be carried out using a large number of variations of the devices, device components, methods steps set forth in the present description. As will be obvious to one of skill in the art, methods and devices useful for the present methods can include a large number of optional composition and processing elements and steps.

[0277] All patents and publications mentioned in the specification are indicative of the levels of skill of those skilled in the art to which the invention pertains. References cited herein are incorporated by reference herein in their entirety to indicate the state of the art as of theirpublication or filing date and it is intended that this information can be employed herein, if needed, to exclude specific aspects that are in the prior art. For example, when composition of matter are claimed, it should be understood that compounds known and available in the art prior to Applicant's invention, including compounds for which an enabling disclosure is provided in the references cited herein, are not intended to be included in the composition of matter claims herein.

[0278] One of ordinary skill in the art will appreciate that starting materials, biological materials, reagents, synthetic methods, purification methods, analytical methods, assay methods, and biological methods other than those specifically exemplified can be employed in the practice of the invention without resort to undue experimentation. All art-known functional equivalents of any such materials and methods are intended to be included in this invention. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by illustrative aspects and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.

[0279] The disclosed embodiments, examples and experiments are not intended to limit the scope of the disclosure or to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. It should be understood that variations in the methods as described may be made without changing the fundamental aspects that the experiments are meant to illustrate.

[0280] Those skilled in the art can devise many modifications and other embodiments within the scope and spirit of the present disclosure. Indeed, variations in the materials, methods, drawings, experiments, examples, and embodiments described may be made by skilled artisans without changing the fundamental aspects of the present disclosure. Any of the disclosed embodiments can be used in combination with any other disclosed embodiment.

[0281] In some instances, some concepts have been described with reference to specific embodiments. However, one of ordinary skill in the art appreciates that various modifications and changes can be made without departing from the scope of the invention as set forth in the claims below. Accordingly, the specification and figures are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of invention.

[0282] The following non-limiting examples are provided purely by way of illustration of exemplary embodiments, and in no way limit the scope and spirit of the present disclosure. Furthermore, it is to be understood that any inventions disclosed or claimed herein encompass all variations, combinations, and permutations of any one or more features described herein. Any one or more features may be explicitly excluded from the claims even if the specific exclusion is not set forth explicitly herein. It should also be understood that disclosure of a reagent for use in a method is intended to be synonymous with (and provide support for) that method involving the use of that reagent, according either to the specific methods disclosed herein, or other methods known in the art unless one of ordinary skill in the art would understand otherwise. In addition, where the specification and / or claims disclose a method, any one or more of the reagents disclosed herein may be used in the method, unless one of ordinary skill in the art would understand otherwise. EXAMPLES Example 1. Identifying, scoring, and prioritization of DMRs

[0283] Deep whole genome methylation sequencing was conducted on converted DNA samples from 100 healthy individuals and from tumor specimens of 100 CRC patients. Following library preparation, sequencing was conducted on converted DNA. Reads were mapped and aligned to a reference genome.

[0284] The two cohorts were chosen to be balanced in terms of age and to represent major ethnicities in the USA, and the CRC cohort was chosen to represent all stages, histologies and morphologies, and microsatellite instability. Differentially methylated genomic patterns were defined between patients with CRC and healthy individuals by utilizing patterns of co-methylationof adjacent CpG loci. These differential methylation profiles are strong biomarkers for targeted sequencing panels for early cancer detection.

[0285] Further segmentation was then performed using a machine learning model as described herein to identify regions differentially methylated between the healthy and CRC cohorts as well as between subpopulations within the CRC cohort, including age, sex, ethnicity, stage, histology, and microsatellite instability (MSI) vs microsatellite stability (MSS). FIG.2 illustrates such a segmentation procedure. For a given region, the methylation load of the healthy and CRC cohorts were plotted and the differential methylation load between healthy and CRC for the region were calculated. The horizontal bars show the differential methylation load. Following feature scoring and prioritization as described herein, a panel comprising the selected set of methylation biomarkers (86% of which hypermethylated and 11% of which hypomethylated) is designed. As shown in FIG.3, the target regions include a high proportion of promoter regions in addition to intronic and exonic targets. The target regions also include CpG islands, shelf and shore.

[0286] To evaluate the discriminative power of panel targets, the differentially-methylated fragments fraction per target was calculated for both baseline (e.g. across healthy samples) and disease (e.g. across cancer samples). For this example, differentially methylated fragment fraction for a target tiwas defined as:

[0287] Hypermethylated fragments were considered only in hypermethylation targets and hypomethylated fragments were considered only in hypomethylation targets.

[0288] In this example, for a given target, Target Differentially Methylated Fragments Fraction (TDMF) mean baseline was defined as the mean differentially-methylated fragment fraction across healthy samples, and TDMF mean disease was defined as the mean differentially- methylated fragment fraction across diseased (e.g. CRC) samples. LR+ score (positive likelihood ratio) per target was also computed as TDMF mean disease / TDMF mean baseline . The distribution of TDMF mean baseline, TDMF mean disease and LR+ across hypermethylated targets are shown in panels A-C of FIG.6.

[0289] The targets for the panel can be further adjusted or narrowed using the noise models disclosed herein. The target selection strategy includes selecting targets that have good coverage (e.g. median on-target coverage of at least 1000x), low TDMF mean baseline, and high TDMF mean disease. Appropriate filters such as those described herein can be implemented for selecting good targets. Example 2. Analysis of hybrid capture probe tiling densities

[0290] Targeted hybrid capture probes are typically designed across a target region to achieve the desired amount of coverage of the target region by the probes. Enrichment of target regions from cfDNA by hybrid capture is typically performed using overlapping probes tiled at a density of 2x (each nucleotide in the target region is covered twice using a series of overlapping probes) or more to ensure adequate coverage of the entire target region. Previous studies have generally used 2x or more tiling of hybrid capture probes for targeted methylation profiling of cfDNA fragments. To evaluate whether lower tiling density (and hence fewer probes necessary) is sufficient for assessing methylation status of cfDNA fragments, 1x (each nucleotide in the target region is covered once by the probes) and 2x probe tiling densities were tested.

[0291] cfDNA from healthy and CRC samples were end-repaired and A-tailed prior to adapter ligation. Enzymatic conversion was performed, followed by barcoding PCR. The libraries were amplified for 13 cycles, normalized, and pooled. The pooled libraries were then captured using either 1x or 2x tiling panels under standard hybridization conditions according to the manufacturer’s hybridization protocol Each tiling condition had 7 CRC samples with Signatera VAF from 0.1%-5%. Each tiling condition also had 4 (1x) or 5 (2x) healthy samples. The post- capture libraries were amplified using 11 cycles and sequenced following basic QC.15M subsampled read pairs were used for the analyses.

[0292] The results showed that 1x tiling and 2x tiling designs had similar capture performance. As shown in FIG.4, both 1x and 2x tiling panels resulted in comparable mapping efficiency, on target rate, mean target coverage, and uniformity.1x and 2x tiling panels are expected to have comparable performance (e.g. DMAF estimations) using the methods disclosed herein.1x or less probe tiling will reduce cost without affecting assay performance. Example 3. Feature engineering and model development

[0293] A feature matrix (as illustrated in FIG.7) used for the classification of healthy and cancer samples was built by taking the hypermethylated fragment fraction of each selected target as a feature.

[0294] Classifiers used so far include: linear Logistic Regression and random forest. The limitations of the current approach are as follows:

[0295] Logistic Regression: while there is a theoretical guarantee that the sum of the training samples' posterior probability estimates equals the class-weighted sum of the labelsis no control for the stability of theestimate. In a class-balanced optimization, the posterior estimates can be expected to cluster around 0.5, but a small change in the distribution of feature values (here, target regions' hyper- methylation fraction) of a new sample set can substantially alter the optimal probability threshold relative to the variance of the posteriors in the training sample. Thus, naively using an empirical threshold onived from cross-validation is prone to large variation in resulting specificity and sensitivity values in a new cohort. Note that while standard regularization methods (e.g. L1, L2, elastic-net) may stabilize the classifier itself, it does not generally stabilize the distribution of the posteriors (except to stabilize their weighted mean).

[0296] Random Forest: RF's led to even fewer theoretical guarantees with respect to the distribution of the posterior estimate, as each sample's posterior estimate is based on a weighted sum of training sample label distribution in each tree's terminal node to which the given new sample is assigned.

[0297] In both cases, features that are highly correlated with each other (i.e. when the covariance has a very high condition number) can lead to very clustered posterior probability estimates - when all samples of the same label have nearly the same posterior estimate.

[0298] One exemplary way to control TDMF baseline involves the following:

[0299] (1) Assume classifier estimates posterior probability of MRD-positive status; (2) bootstrap an empirical distribution of posterior probabilities from the classifier; and (3) find optimal posterior distribution transform and threshold for the desired TDMF baseline.Learning a posterior threshold on the transformed posteriors for a desired specificity then leads to more predictable results.

[0300] As shown in FIG.24(A)-(B), before posterior transformation, posteriors were distributed very tightly around p = 0.5, with very little variability. Maximum and minimum posteriors differed by less than 1e-4. Finding a reliable threshold for such a distribution is challenging. As shown in FIG.24(C)-(D), applying the learned posterior transformation approach, the 99% training specificity threshold was found to be p=0.53.

[0301] The learned posterior transformation was combined with a feature transformation approach to further improve the classifier performance. Hypermethylated fragment fractions (HFF) was transformed to be approximately normally distributed. As an example,was used, and the resulting target-level features were used without any new linear transformations. Classification performance was notably more stable with respect to the padding parameter ^ over several orders of magnitude. FIG.25 shows the results of combining the feature transformation and the learned posterior transformation described above. FIG.25(A)-(B) shows the distribution of logistic regression posteriors on out-of-fold samples after the combined transformations. FIG.25(C) shows classification accuracy based on the fraction of the 30 repeats that a given validation sample was correctly classified. FIG.25(D) shows ROC curve for the model. FIG.25(E)-(F) shows violin plots of performance measures based on 30 repeats of 5-fold cross-validation, with the desired specificity set at 97% and 98%, respectively. As shown is FIG. 25(E)-(F), the desired specificity was achieved fairly precisely with the two transformations combined: 96% out-of-fold specificity was achieved at 97% desired specificity.97% out-of-fold specificity was achieved at 98% desired specificity.

[0302] The performance of the model was evaluated using 5-fold cross-validation. In order to get a more robust estimation of the performance, the 5-fold cross-validation was repeated 20 times. A general overview of the performance evaluation process is depicted in FIG.14.

[0303] In each fold, model selection was performed only on the training set through an independent 5-fold cross-validation, to obtain the best model parameters. Then the whole trainingset was used to train a model with the parameters obtained in the model selection step. At the end of each fold, the sensitivity, specificity, and the area on the ROC (AUC) was calculated for the validation set. After finishing all the repeats / folds, the mean ROC curve, mean AUC, mean sensitivity, and mean specificity were computed and reported as the final performance estimation. Conclusion

[0304] A straightforward method was developed to represent the likelihood of CRC observation as an intuitive score ranging from 0 to 1. This was done by transforming the methylation-based features into a Gaussian distribution as well as modifying the distribution of the classifier posterior probabilities. In this way, the methylation score was stabilized.

[0305] Combining two separate transformation functions (one on the data, one on the modeled posterior distributions) allowed achieving the desired specificity more closely and more reliably. Table 1 shows actual vs. target specificity when using the two-transformation method described above.

[0306] Table 1: Actual out-of-fold specificity achieved with 5-fold cross-validation repeated over 30 random shuffles vs. target (desired) specificity, using the two- transformation method.

[0307] The two-transformation approach also achieved 83% sensitivity at 96% specificity, and 80% sensitivity at 97% specificity (FIG.25(E)-(F)). FIG.26 shows performance of the two- transformation model across tumor VAF levels for different target specificity levels. Example 4. Performance of classifier on MRD samples

[0308] The independently trained classifier described above was used to predict MRD status in a cohort of 247 patients enrolled in the Bespoke CRC trial (NCT04264702) who previously had MRD testing for ctDNA completed using Signatera™. The cohort included 163 MRD-negative patients without clinical progression and 84 MRD-positive patients, of whom 48 had radiographic recurrent disease.

[0309] The cohort distribution of stage and Signatera VAF are shown in FIG.5. The median VAF across all samples was 1.2e-3 (Fig.5B). Good uniformity of coverage across samples was observe. The uniformity metric, defined as 0.8x of the mean was 76% across both MRD positive and negative samples.

[0310] cfDNA extracted from the samples were blunt-end repaired, A-tailed, adapter- ligated, and enzymatically converted. The enzyme-treated samples were universally amplified with barcoding PCR before enriched for target-region amplicons using a hybrid capture panel. Post-capture PCR was performed on enriched samples, and the amplified, enriched samples were pooled for sequencing after QC. DMAF

[0311] In order to investigate “methylation signal” of a cohort or perform exploratory analysis, it is useful to define a measure that summarizes this “methylation signal” at a high-level for each sample. Here, the Differential Methylation Allele Fraction (DMAF, also referred to sometimes as Differential Methylation Abundance Fraction) was used as a proxy for the methylation signal. DMAF can generally be calculated as:

[0312] For example, where a set or subset of target regions are hypermethylated in a phenotype sample, DMAF can be calculated as:

[0313] Since a per-sample measure is needed, DMAF aggregates the methylation data from different targets of the panel, and since it has to be a discriminative measure, the aggregation is done only on targets with sufficient signal. TDMF mean baseline, LR+, and coverage thresholds can be established to ensure that only good targets are considered for DMAF calculation. The required minimum number of CpGs in fragments can be target-specific (e.g., different thresholds for different targets).

[0314] FIG.8 shows the distribution of DMAF as a function of input DNA amount and total number of suitable fragments in selected targets.

[0315] Good correlation of DMAF with the Signatera ctDNA VAF (spearman rho = 0.77 and slope of 0.68) was observed, as shown in panels A and B of FIG.9. Panel C shows the correlation of DMAF with Mean Tumor Molecules (MTM) which has identical spearman rho and slope as VAF.

[0316] The correlation of DMAF and Signatera ctDNA VAF by sub-groups was next analyzed. FIG.10 shows the correlation plots for MSI vs. MSS cancers. FIG.11 shows the correlation plots for colon vs. rectal cancers. FIG.12 shows the correlation plots for stage II, III and IV (stage I not shown due to too few samples). While slight differences were observed in the correlation across subtypes, it is believed that these differences are explained by the sample size within individual groups and are not otherwise significant. Target enrichment within molecular and histological subtypes

[0317] To identify targets that might be enriched in one subtype vs. others, a differential methylation analysis was performed by comparing the methylation status per group per target. A target was considered to be methylated in a given sample if any hyper-methylated fragments were seen in that target for that sample. To identify targets specific to MSI vs. MSS, a fisher's exact test was performed per target comparing the number of samples that are hypermethylated vs. not hypermethylated for MSI and MSS samples. The quantile-quantile plots of the observed p-values across all targets compared to the expected p-values (based on a normal distribution) is shown in panel A of FIG.13. Panel B shows the volcano plot - log2 of fold change vs. negative log 10 of p- value. Each point shows the fold change and p-value for a particular target. Points on the right size of x = 0 show targets that are enriched in MSI compared to MSS and points to the left of x = 0 show targets that are enriched in MSS compared to MSI. After correction for multiple hypothesis testing, none of the targets were significant. However, targets with p-values less than 0.005 are shown in the plot. Looking at the top target EPM2AIP1, prior studies have reported that EPM2AIP1 is closely related to MLH1 promoter methylation. MLH1 has been shown to be a bi- directional promoter with a second gene, EPM2AIP1, located in a head-to-head orientation on theopposite strand 321 base pairs away from the MLH1 transcription start site. The methylation of this promoter region results in the transcriptional silencing of both the MLH1 and EPM2AIP1 genes. Furthermore, MSI-H tumors show a significant reduction in MLH1 / EPM2AIP1 expression compared to normal colonic mucosa and microsatellite stable tumors, indicating that EPM2AIP1 is concurrently turned off in the majority of sporadic tumors with MLH1 hyper-methylation. This confirms that EPM2AIP1 promoter methylation is related to promoter hyper-methylation of MLH1 which occurs in approximately 20% of sporadic MSI cancers.

[0318] Target enrichment was also assessed between colon and rectal cancers as well as stage 4 cancers vs. stages I, II and III. No significant associations were observed from these analyses. Classification performance

[0319] To determine the classification performance of MRD-positive samples, a logistic regression classifier trained on MRD positive vs. MRD negative samples was used as described above. The model was tuned to report sensitivity at 95% and 97% specificity.

[0320] The ROC curve of the classifier tuned to 95% specificity is shown in panel A of FIG.22. The distribution of sensitivity and specificity across the cross-validation folds are shown in panel B of FIG.22. The Signatera VAF vs. DMAF correlation plot annotated by classification accuracy is shown in FIG.22.

[0321] Positive percent agreement (PPA) of the methylation classifier at different VAF bins (based on median Signatera VAF of the sample) and at 95%, 97% specificity was also assessed as shown in Table 2. Table 2: Sensitivity of methylation assay at different VAF bins >0 0 < TExample 5. Performance of MRD-trained and ECD-trained classifiers on ECD samples

[0322] The performance of the classifiers developed using the methods described above were evaluated on ECD samples. A panel as described in Example 1 above plus additional target regions for advanced adenomas (AA) was used in this analysis. This modified panel contains about 87.9% hyper and 10.5% hypo-methylated targets (shown in panel A of FIG.29). The number of targets by source and CpG annotation are shown in panels B and C of FIG.29.

[0323] The ECD readout included 383 samples in total with 146 colorectal cancers, 31 advanced adenomas (AA) and 190 healthy samples in addition to 16 sequencing controls. The breakdown of samples by type is shown in Table 3 below. There were 133 samples in total from Circulate trial and 100 samples from Stem Express, the two groups that are used for the primary readout. Additional samples were sourced from other biobanks. Table 3: Sample counts by category S CRC Healt Cont Adva (AA) Colo

[0324] The distribution of cancer stage, median target CpG coverage and coverage across input DNA is shown in FIG.17.Cohort description of Primary readout samples

[0325] As described earlier, only Circulate and Stem Express samples were used for the primary readout. The breakdown of stage and MSI status in the 129 samples within the Circulate cohort is shown in Table 4 below. Of these samples, 30 / 129 were screen-detected of which 67% (20 / 30) are stage I or II. Of the 30 screen detected samples, 10 were screened using colonoscopy and 20 had a positive stool test.

[0326] Of the remaining samples, 53 / 129 were symptom-detected of which 38% (20 / 52) were stage I or II. For the remaining samples it was unknown whether they were screen or symptom detected. Table 5 shows the breakdown of eligibility criteria across all stages. Table 4: stage and MSI status breakdown in Circulate cohort Cance I II III IV MSI st MSI-H MSSTable 5: Sample breakdown by eligibility arm Stag I II IIA IIC IIIAIIIB IIIC IV IVA IVC

[0327] cfDNA extracted from the samples were blunt-end repaired, A-tailed, adapter- ligated, and enzymatically converted. The enzyme-treated samples were universally amplified with barcoding PCR before enriched for target-region amplicons using a hybrid capture panel. Post-capture PCR was performed on enriched samples, and the amplified, enriched samples were pooled for sequencing after QC. Comparison of target metrics between ECD and MRD samples

[0328] Target metrics were evaluated as described above. When compared to TDMF baseline mean computed from Bespoke MRD samples described above, a strong correlation was observed across all targets (shown in panel A of FIG.18). The targets with low TDMF baseline mean (<1e-4) in either experiment and the confidence intervals on these estimates were also observed (shown in panels B and C of FIG.18). Taken together, these comparisons show that TDMF baseline mean values are very similar between MRD and ECD. There is a high level of variability in the low TDMF baseline mean range due to the extremely small number of fragments that would contribute to the estimate.

[0329] A strong correlation of the TDMF disease mean between MRD and ECD samples was also observed (shown in panel A of FIG.19). The LR+ between MRD and ECD samples are more variable as it combines the variability in both TDMF disease mean and TDMF baseline mean (shown in panel B of FIG.19). CRC vs. Healthy classification (Circulate vs. Stem Express)

[0330] The CRC vs. Healthy primary readout based on Circulate and Stem Express data was reported with different approaches.Cross-validated performance based on best parameters from nested grid-search

[0331] A nested cross-validation was performed using a gird search to identify the optimal number of targets and regularization parameter for the logistic regression classifier. Based on the best parameters chosen by the grid-search, the logistic regression model was retrained and then applied to the samples in validation fold of the outer CV. Data showed: 91% (CI 80%-100%) sensitivity at 90% (CI 75%-100%) specificity Stage I sensitivity close to Signatera pre-surgical: 74% (71 / 96) 79% (24 / 30) sensitivity for screen-detected CRCs 97% (51 / 53) sensitivity for symptom-detected CRCs

[0332] FIG.20 shows the ROC curve and confidence intervals for the CV performance. Table 6 below shows the performance breakdown by stage and Table 7 shows the performance breakdown by eligibility arm. Table 6: Sensitivity by stage and MSI type for ECD trained classifier S I I I I (Table 7: Sensitivity by eligibility for ECD trained classifierE s s d E s s dPerformance based on model trained on MRD samples

[0333] In order to test the generalizability of the methylation approach, a model was used trained on MRD data to predict on ECD samples. Cross-validation was performed on the MRD data using a gird search to identify the optimal number of targets and regularization parameter for the logistic regression classifier. Based on the best parameters chosen by the grid-search, a final model was trained on all MRD data and predictions were made on the ECD data (as a hold-out dataset). Data showed: 93% (CI 86%-100%) sensitivity at 93% (CI 82%-100%) specificity ; 90% (27 / 30) sensitivity for screen-detected CRCs; 94% (50 / 53) sensitivity for symptom-detected CRCs.

[0334] FIG.21 shows the ROC curve of the classifier and the confidence intervals around the sensitivity and specificity. Table 8 shows the performance breakdown by stage and Table 9 shows the performance breakdown by eligibility arm. Table 8: Sensitivity by stage for TF-MRD trained classifierTable 9: Sensitivity by eligibility arm for TF-MRD trained classifierPerformance of ECD-trained model on MRD samples

[0335] The performance of the ECD-trained classifier on the Bespoke MRD data was evaluated to test the model’s generalizability.

[0336] Cross-validation was performed on the ECD data restricted to the targets from Example 1 using a gird search to identify the optimal number of targets and regularization parameter for the logistic regression classifier. Based on the best parameters chosen by the grid- search, a final model was trained on all ECD data and predictions were made on the MRD data. The performance of the ECD-trained model on the MRD data is shown in FIG.22.

[0337] Compared to the performance of the MRD 5x5 CV cross-validation model described in Example 4 above, the ECD-trained classifier had comparable performance at 95% specificity, and better performance at 97% specificity on the same MRD data, as shown in Table 10: Table 10: Sensitivity of ECD-trained classifier (ECD CLF) vs. MRD cross-validation model (MRD 5x5 CV) by VAF binsCorrelation of DMAF with Signatera VAF

[0338] The correlation of Signatera VAF for Circulate samples was compared to the Differential Methylation Allele Fraction (DMAF). A good correlation was observed with Signatera VAF, suggesting that the estimate of DMAF based on the targets from the TF-MRD model is a good estimate of underlying ctDNA fraction.

[0339] Panel A and B of FIG.23 show the correlation of DMAF with VAF with the classification labels denoted in panel B. In total there were 18 Signatera false negatives. The input volume available for Signatera testing for these samples was only one tube of blood compared to the usual input of two tubes. As a result, the sensitivity of Signatera for this cohort is lower than the typical sensitivity of Signatera testing. Panel C shows the DMAF of the Signatera negatives, 8 of which were SOVATS (sample with one variant above threshold).15 / 18 were detected with the TF-MRD trained classifier. With the ECD classifier, 14 / 18 Signatera false negatives were detect. Conclusions

[0340] Results for primary readout from Circulate and Stem Express exceed readout goal of 90% sensitivity at 90% specificity. With the ECD trained classifier, the performance from 5- fold cross-validation was 91% sensitivity at 91% specificity and with MRD trained classifier, 93% sensitivity at 93% specificity was observed. The strong performance of MRD trained classifier on ECD data suggests that it generalizes well and that the classification approach described above is quite robust. The TDMF mean baseline rates are also very similar between the MRD and ECD cohorts. Example 6. CHG / CHH filtering

[0341] Compared to the cohort described in Example 5 (Cohort 1), a significantly higher TDMF baseline mean was observed in a new cohort of ECD samples (Cohort 2), which includes samples from the same sources as Cohort 1, including pre-surgical CIRCULATE-Japan samples. Cohort 2 also includes samples from PROCEED-CRC, an ongoing, decentralized, prospective sample collection clinical study. The high background noise in Cohort 2 led to an increase in positivity rate (defined as proportion of samples within a group that are called positive) in healthy samples. As shown in FIG.31, a nearly 9 fold increase in positivity rate in healthy samples from Cohort 2 compared to Cohort 1 was also observed, whereas positivity rate in CRC samples only increased slightly in Cohort 2. indicating high background noise in the Cohort 2 data. About 2-3 fold more CHH or CHG (H correspond to A, T or C) methylated reads were observed in Cohort 2 data compared to the same samples processed and analyzed in Cohort 1.

[0342] DNA methylation is found in three different sequence contexts: CG (or CpG), CHG or CHH. In mammals (e.g. humans), with the exception of embryonic stem cells, DNA methylation is almost exclusively found in CpG dinucleotides, with the cytosines being methylated usually on both strands. CHG and CHH are typically not methylated in mammals. Apparent CHH / CHG methylation can indicate incomplete conversion, which can lead to seemingly hypermethylated fragments.

[0343] A mCHH / CHG filter was developed and implemented to remove likely incompletely converted DNA fragments based on the number of methylated cytosines in non-CG context (CHH and / or CHG). The threshold can be set as an absolute number (e.g.1, 2, 3, 4, 5, 6, or more) of such methylated cytosines or as a percent of the total cytosines in the fragment.

[0344] Following hybrid capture and sequencing, DMAF comparison was done on healthy and CRC samples with and without the mCHH / CHG filter. As shown in FIG.32, DMAF in healthy samples was reduced after mCHH / CHG filtering. DMAF in CRC samples was impacted less by the filter, suggesting that the filter appropriately removed false positive fragments.

[0345] The locked MRD classifier described in Example 4 was retrained on mCHH / CHG filtered data and applied to Cohort 2. As shown in FIG.33, DMAF and positivity rates were well differentiated between CRC and healthy samples across different sample sources. The desiredspecificity of about 90% was achieved across vendors, with 91% combined specificity (n = 313) and 93% (147 / 158) specificity for PROCEED samples analyzed so far. Sensitivity at 91% specificity was 95.3% (CI 88% - 100%), with 85.3% (29 / 34) sensitivity for screen-detected CRCs and 100% (80 / 80) sensitivity for symptom-detected CRCs. Table 11 below shows performance of MRD-trained models with and without mCHH / CHG filter by clinical stage. Sensitivity for MSI-Hi and MSS with mCHH / CHG filter was 100% (14 samples) and 95% (113 samples), respectively. FIG.34 shows (A) ROC curve of the retrained classifier on validation sets and (B) DMAF correlation with Signatera VAF and classification accuracy. Table 11. MRD-trained classifier performance estimation on ECD data with and without mCHG / CHH filter

[0346] Table 12 shows sensitivity by VAF bin and positive percent agreement (PPA) of the MRD-trained classifier with mCHG / CHH filtering on ECD data. Table 12. Sensitivity & PPA of MRD-trained classifier with mCHG / CHH filtering

[0347] 13 of the 127 samples did not have Signatera results and were excluded from both the sensitivity and PPA analyses. All 18 Signatera negative samples were excluded in the PPA analysis. The sensitivity analysis included 8 Signatera negative samples with a single variant above threshold (SOVATS) and excluded the rest. As shown in FIG.34(B), there was a strong correlation between DMAF and Signatera ctDNA VAF.

[0348] mCHH / CHG filter also improved the performance of the MRD classifier on the BESPOKE cohort described in Example 4, as shown below in Table 13. Table 13. MRD-trained classifier performance estimation on MRD data with and without mCHG / CHH filter

[0349] Likewise, mCHH / CHG filtering also improved the overall performance of an independently trained ECD classifier on MRD data, especially at VAF < 0.04% (compare Table 14 below with Table 11, and FIG.35 with FIG.22). Table 14 shows performance data by VAF bins using the updated mCHG / CHH filtered data. FIG.35 shows (A) ROC curve of the CV models on validation sets and (B) 5-fold CV performance (20 repeats) using the mCHH / CHGfiltered data. A strong correlation between DMAF Signatera ctDNA VAF was observed (see FIG. 36). Table 14: Sensitivity of ECD-trained classifier (ECD CLF) vs. MRD cross-validation model (MRD 5x5 CV) by VAF bins with mCHG / CHH filtered data

[0350] After mCHG / CHH filtering, calling thresholds became identical between ECD- trained and MRD-trained classifiers. The strong performance of ECD-trained classifier on MRD data and TF-MRD trained classifier on ECD data suggests that the methylation classifiers developed using the methods herein are robust and generalize robustly across sample batches, time points (e.g. early cancer detection vs cancer recurrence detection), and signal source (e.g. primary tumor vs metastases). Example 7. Methylation-based ctDNA quantification

[0351] As illustrated in FIG.37, the model-free methylation-based ctDNA quantification method described herein included two core components: (a) independent target selection, and (b) DMAF estimation. Independent Target Selection

[0352] In this example, instead of using the same targets that optimized classification performance, DMAF was computed by selecting targets with minimal background noise and linear association with Signatera VAF. Specifically, target selection prioritized regions that met the following criteria: (1) low background hypermethylation (TDMF baseline) in normal plasma samples to minimize false-positive signal which will translate to lower LoB and LoD; (2) high signal-to-noise ratio (SNR), i.e., high differential hypermethylation between tumor-positive and tumor-negative / healthy plasma samples; and (3) linearity with Signatera VAF, i.e., a methylation signal that exhibits a linear, dosage-based relationship with Signatera VAF, to support the use of these regions for quantitative ctDNA burden estimation.

[0353] To identify methylation targets that exhibit a linear relationship with ctDNA burden based on Signatera VAF, an F-test was performed to assess the association between each candidate target’s methylation signal and Signatera VAF. For each candidate target region, the fraction of hyper-methylated fragments (HFF) was calculated and applied a padded logit transformation (pseudocount 1e-8).

[0354] A univariate F-test was then performed, modeling the association between the padded logit transformed methylation signal (HFF) and log10(Signatera VAF) across tumor- positive samples. Targets with p-value > 0.1 were excluded and the remaining targets were ranked by F-test p-value (ascending), and the top targets (tuned via cross-validation) were selected for inclusion in the final target panel.

[0355] This approach prioritizes targets whose methylation levels increase proportionally with ctDNA burden, supporting the use of these regions in a quantitative, tumor-naive quantification metric. DMAF Estimation Model-free DMAF

[0356] Using the final set of targets selected based on their association with Signatera VAF, DMAF was computed per sample as a summary metric. For each sample, DMAF was calculated using the following formula:∑9 ^^^^^ − ^^^ℎ^^^^^^ ^^^^^^^^^^^6

[0357] Here, N is the number of selected target regions, prioritized based on differential methylation (SNR), low background in normals (TDMF baseline), and linearity with Signatera VAF; hyper-methylated fragments are fragments with a predefined minimum number of CpGs (e.g. at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 CpGs) overlapping selected target regions, where a predefined minimum percentage (e.g. at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) of overlapping CpGs are methylated; and on-target fragments are all fragments with the predefined minimum number of CpGs overlapping selected target regions, regardless of methylation status.

[0358] This yields a tumor-naive and model-free estimate of methylation-based tumor burden that reflects the proportion of cfDNA fragments across selected targets exhibiting tumor- associated hyper-methylation patterns. Regression model based DMAF

[0359] As a complementary strategy to the model-free approach, a preliminary evaluation of a Ridge regression-based method for estimating DMAF was also performed. This approach leverages supervised learning to optimally weight per-target methylation signals for predicting ctDNA burden, using Signatera VAF as the dependent variable during training. Ridge regression optimizes for the following objective function: M=> =

[0360] Here, γiis the log10-transformed Signatera VAF for sample i, xijis the hypermethylated fragment fraction for target jjj in sample iii, βjare the regession coefficients for each target, λ is the regularization parameter (tuned via cross-validation). The penalty termN J Mscourages large weights and helps prevent overfitting.Using the coefficients ( β )jestimated during model training with a Ridge penalty, the regression model predicts the outcome ^U for each sample as:

[0361] The predictions from a regression model are in log10space. These can be converted to linear space as 10y.Model training setup

[0362] Prior model training used samples belonging to three categories based on ctDNA status: 1. Signatera Positive – Samples that were called positive by the Signatera assay. 2. Signatera Negative – Samples that were called negative by the Signatera assay. 3. Healthy – Self-declared healthy individuals.

[0363] While Signatera negatives indicate the absence of detected ctDNA by a tumor- informed assay, they are not nested negatives and may include false negatives or signal from secondary primaries that Signatera is not designed to detect. Therefore in this example, these samples were excluded from model training to avoid introducing labeling noise.

[0364] For both the model-free and model-based DMAF methods, training and target selection were conducted using only Signatera Positive samples (as ctDNA-positive ground truth), and healthy samples (as ctDNA-negative controls).

[0365] For each cancer type, the downstream pipeline including fragment counting was run using a cancer-type-specific sub-panel, refined based on the correlation of methylation status between neighboring CpGs within that cancer type.

[0366] Sample-level quality control (QC) thresholds, including median CpG coverage, cfDNA input amount, conversion rate, protection rate, and fragment length were applied to the data. Samples failing any of these thresholds were excluded from downstream analysis to maintain data integrity and consistency across the cohort.Cross-validation setup

[0367] To robustly assess model performance and tune hyperparameters, a nested cross- validation framework was implemented. The model was trained and evaluated using 10-fold cross- validation, repeated 10 times to account for variance in fold assignments. Inner folds were used for hyperparameter tuning (e.g., number of targets selected), while outer folds were used for unbiased performance evaluation. This nested setup helped ensure that hyperparameter optimization does not leak into the final performance estimate, providing a more reliable assessment of generalizability.

[0368] Cross-validation folds were stratified by Signatera VAF, age, and sample category to ensure balanced representation of samples from these groups.

[0369] For each outer fold, hyperparameters were selected by optimizing the Pearson correlation between predicted and observed values on the inner-fold validation sets. Hyperparameters that were tuned during cross-validation include the number of targets with low TDMF baseline, SNR threshold, and final number of targets selected. Cross-validation based predictions

[0370] Predictions for each sample were generated as the median DMAF value across 10 repeated cross-validation runs. This approach reduces variability arising from fold assignments and can be a more stable estimate: X

[0371] Le cated DMAF in sample ^ in cross-validation repeat ^, and R = 10is the total number of cross-validation repeats.

[0372] The final prediction for sample is:

[0373] where agg is an aggregation function (e.g., mean, median, max) of sample . Metrics used for performance assessment

[0374] Model performance was evaluated using the following metrics:

[0375] Mean Squared Error (MSE): Quantifies the average squared difference between predicted log10(DMAF) and observed log10(Signatera VAF), measuring overall prediction accuracy.

[0376] Pearson Correlation Coefficient (): Assesses the strength of the linear relationship between predicted log10(DMAF) and observed values log10(Signatera VAF). Performance of the model-free DMAF for CRC

[0377] The performance of the model-free DMAF method for CRC was first evaluated with a tfMRD-CRC panel using 10-fold, 10-repeat nested cross-validation. As shown in FIG.38, the model-free DMAF method demonstrated increased linearity and Pearson correlation with observed Signatera VAF in CRC (N=54) and reduced underestimation. MTM / ml (methylation) may be calculated by: DMAF * cfDNA mass * 1000 pg per ng / (3.3 pg * plasma volume), which is similar to the MTM / ml calculation based on VAF.

[0378] The performance of the model-free DMAF method for CRC was also evaluated with a second tfMRD panel using 10-fold, 10-repeat nested cross-validation. As shown in FIG. 39, in CRC samples (N=105), the model-free DMAF method demonstrated a strong correlation between predicted methylation-based quantification, DMAF, and observed Signatera VAF. The median number of targets selected across CV folds was 20. Mean squared error (MSE) was 0.064, and Pearson correlation (ρ) was 0.955, indicating a strong linear association between DMAF predictions and Signatera VAF. Performance of the model-free DMAF for breast cancer

[0379] The performance of the model-free DMAF approach for breast cancer was evaluated using 10-fold, 10-repeat nested cross-validation. As shown in FIG.40, in breast cancer samples, the model-free DMAF method again demonstrated a strong correlation between predicted DMAF and observed Signatera VAF. As breast cancer predominantly occurs in females and the training cohort contained only female cases, negative samples were restricted to female participants to ensure sex-matched controls during model training. The median number of targetsselected across CV folds was 150. For all breast cancer samples (N=112), Mean squared error (MSE) was 0.161 and Pearson correlation (ρ) was 0.863. For HR+ subtype (N=45), MSE was 0.132 and Pearson ρ was 0.888. For HER2+ subtype (N=29), MSE was 0.272 and Pearson ρ was 0.815. For TNBC subtype (N=38), MSE was 0.109 and Pearson ρ was 0.897. Performance of the model-free DMAF for lung cancer

[0380] The performance of the model-free DMAF approach for lung cancer was evaluated using 10-fold, 10-repeat nested cross-validation. As shown in FIG.41, in lung cancer samples, the model-free DMAF method again demonstrated a strong correlation between predicted DMAF and observed Signatera VAF. The median number of targets selected across CV folds was 40. For all lung cancer samples (N=113), Mean squared error (MSE) was 0.179 and Pearson correlation (ρ) was 0.829. For NSCLC subtype (N=90), MSE was 0.207 and Pearson ρ was 0.802. For SCLC subtype (N=23), MSE was 0.069 and Pearson ρ was 0.941. Performance of the model-free DMAF for bladder cancer

[0381] The performance of the model-free DMAF approach for bladder cancer was evaluated using 10-fold, 10-repeat nested cross-validation. As shown in FIG.42, in bladder cancer samples (N=60), the model-free DMAF method again demonstrated a strong correlation between predicted DMAF and observed Signatera VAF. The median number of targets selected across CV folds was 50. Mean squared error (MSE) was 0.079, and Pearson correlation (ρ) was 0.954, indicating a strong linear association between DMAF predictions and Signatera VAF. Conclusion

[0382] The tumor-naïve, model-free methylation-based metric DMAF effectively estimated ctDNA burden in CRC, breast, lung, and bladder cancers. Across cancer-types with Signatera positive MRD samples, the model-free DMAF quantification approach exhibited highest correlations (ρ ≥ 0.95) and lowest MSE in CRC and bladder cancer samples. The performance of the model-free DMAF quantification method was also robust across breast cancer subtypes, with slightly reduced correlation in HER2+ cases. In lung cancer, the model-free DMAF approachshowed a higher correlation with Signatera VAF in SCLC than NSCLC. Healthy controls exhibited consistently low DMAF values. Example 8. Effect of input and age on DMAF predictions across cancer types

[0383] This example examines DMAF values across categories of (i) input cfDNA amount (ng) and (ii) age, stratified by sample category (Healthy vs. PosMRD). The four boxplots in each of FIGs.43-46 correspond to DMAF predictions from cancer-type specific DMAF quantification models. The y-axis for all plots is on a log10scale. The top panels of each of the figures corresponds to healthy samples and the bottom panels are the MRD positive samples. The panels on the left correspond to input amount and the panels on the right correspond to age. CRC

[0384] Tables 15 and 16 show the number of CRC and healthy samples in each category for the analysis: Table 15. CRC and healthy sample numbers for different input cfDNA amounts Inp CR HeaTable 16. CRC and healthy sample numbers for different age groups Age CRC Healt

[0385] As shown in the top left panel of FIG.43, healthy samples with higher input amounts (51-100 ng) show slightly elevated median DMAF compared to 10-30 ng and 31-50 ng groups, but all remain below , close to background levels. The top right panel demonstrates thathealthy samples across the 29-50 and 50-70 age groups show similar median DMAF values, with some outliers. The 70-100 age group shows minimal variability and very low DMAF, possibly due to limited sample size.

[0386] For the MRD positive samples, the bottom left panel of FIG.43 demonstrates that median DMAF increases with higher input amounts, with the 51-100 ng group showing the highest median and upper spread. Biologically, this is expected, as ctDNA-positive cases particularly those with higher tumor burden are more likely to yield greater cfDNA input amounts. Across age groups (bottom right panel), median DMAF values appear similar, though the 29-50 group has a slightly higher median than the oldest group (70-100). Variability is substantial across all age groups, with notable high-value outliers. Breast cancer

[0387] Tables 17 and 18 show the number of breast cancer and healthy samples in each category for the analysis: Table 17. Breast cancer and healthy sample numbers for different input cfDNA amounts Inp Bre HeaTable 18. Breast cancer and healthy sample numbers for different age groups Age Breas Healt

[0388] As shown in the top left panel of FIG.44, for healthy samples median DMAF remains low across all input groups (10-30 ng, 31-50 ng, 51-100 ng), with slightly lower median and reduced spread in the highest input group (51-100 ng). Additionally, healthy samples from the29-50 and 50-70 age groups (top right panel) have similar median DMAF, with slightly greater variability in the younger group.

[0389] For the MRD positive samples, the bottom left panel of FIG.44 demonstrates that Median DMAF increases with higher input amounts, with the 51-100 ng group showing the highest median and upper spread. This is consistent with the biological expectation that samples from patients with higher tumor burden yield higher cfDNA input amounts. Median DMAF appears similar across the 29-50, 50-70, and 70-100 age groups (bottom right panel), though variability is slightly greater in the youngest group. No strong age-related trend is observed in tumor-positive samples. Lung cancer

[0390] Tables 19 and 20 show the number of lung cancer and healthy samples in each category: Table 19. Lung cancer and healthy sample numbers for different input cfDNA amounts Inp Lun HeaTable 20. Lung cancer and healthy sample numbers for different age groups Age Lung Healt

[0391] As shown in the top left panel of FIG.45, for healthy samples median DMAF is similar across input ranges (10-30 ng, 31-50 ng, 51-100 ng), with overlapping distributions and no strong trend. For input age (top right panel), healthy samples from the 29-50 and 50-70 age groupsshow comparable median DMAF values. Only one or very few samples are present in the 70-100 age group, limiting interpretation.

[0392] For the MRD positive samples, median DMAF remains relatively stable across input ranges, though the 10-30 ng group shows a slightly higher median than the other two groups (bottom left panel). For input age (bottom right panel), DMAF medians are similar across all age categories, with modest variability. The 29-50 group shows slightly higher spread compared to the other groups. Bladder cancer

[0393] Tables 21 and 22 show the number of bladder cancer and healthy samples in each category: Table 21. Bladder cancer and healthy sample numbers for different input cfDNA amounts Inp Bla HeaTable 22. Bladder cancer and healthy sample numbers for different age groups Age Blad Healt

[0394] FIG.46 top left panel shows that for healthy samples, median DMAF is similar across all input ranges (10-30 ng, 31-50 ng, 51-100 ng), with overlapping variability and no clear trend. For input age (top right panel), DMAF medians for the 29-50 and 50-70 age groups are comparable.

[0395] For the MRD positive samples, median DMAF is broadly consistent across input ranges, though the 51-100 ng group shows a slightly higher upper spread (bottom left panel). For input age (bottom right panel), median DMAF appears similar between the 50-70 and 70-100 age groups, with the oldest group (70-100) showing greater variability and several high-value outliers. Conclusion

[0396] Higher cfDNA input amounts were associated with increased DMAF in MRD positive cases in CRC and breast cancer samples, consistent with biological expectations. Age showed minimal impact on DMAF in any cohort. Overall, this suggests that while higher cfDNA input can enhance DMAF signal in some cancers, age has limited impact on assay performance.

[0397] Accordingly, the model-free DMAF provided herein provides a reliable, quantitative, methylation based tissue-free approach for ctDNA quantification, particularly useful where tumor-informed assays are not feasible. * * * *

Claims

What is claimed is:

1. A method of preparing a composition of non-naturally occurring DNA, comprising: (a) extracting cell-free DNA molecules from a sample of a subject; and (b) contacting the extracted cell-free DNA molecules or their derivative with a panel of oligonucleotide probes designed to hybridize to 5-200,000 different target regions having CpG sites that are differentially methylated in a phenotype, thereby generating enriched DNA, wherein the oligonucleotide probes for at least one of the target regions are designed to cover the target region with a tiling density of between 0.9x and 2x.

2. The method of claim 1, further comprising treating (i) the extracted cell-free DNA, or (ii) the enriched DNA or its derivative, with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines.

3. The method of claim 1 or 2, wherein the panel of oligonucleotide probes are designed to cover at least 50%, at least 80%, at least 90%, or at least 95% of the target regions with a tiling density of between 0.9x and 2x.

4. The method of claim 1 or 2, wherein the panel of oligonucleotide probes are designed to cover at least 50%, at least 80%, at least 90%, or at least 95% of the target regions with a tiling density of 1x to 1.4x.

5. The method of claim 1 or 2, wherein the panel of oligonucleotide probes are designed to cover at least 50%, at least 80%, at least 90%, or at least 95% of the target regions with a tiling density of 1x.

6. The method of claim 1 or 2, wherein the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions are tiled at less than 25 bp, less than 20 bp, less than 15 bp, less than 10 bp, less than 5 bp.

7. The method of claim 1 or 2, wherein the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for both top and bottom DNA strands.

8. The method of claim 1 or 2, wherein the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for DNA strands that are fully unmethylated.

9. The method of claim 1 or 2, wherein the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for DNA strands methylated at different levels.

10. The method of claim 1 or 2, wherein the oligonucleotide probes for at least 50%, at least 80%, at least 90%, or at least 95% of the target regions include probes for both fully methylated and fully unmethylated DNA strands.

11. The method of claim 1 or 2, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 3 CpG sites.

12. The method of claim 1 or 2, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 6 CpG sites.

13. The method of claim 1 or 2, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 10 CpG sites.

14. The method of claim 1 or 2, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has at least 20 CpG sites.

15. The method of any of claims 1-14, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.02 CpG / bp.

16. The method of any of claims 1-14, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.03 CpG / bp.

17. The method of any of claims 1-14, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.04 CpG / bp.

18. The method of any of claims 1-14, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions each has a CpG density of at least 0.05 CpG / bp.

19. The method of any of claims 1-18, wherein the target regions have a combined footprint of 800 bp to 30 Mbp.

20. The method of any of claims 1-18, wherein the target regions have a combined footprint of at least 15 Kb.

21. The method of any of claims 1-18, wherein the target regions have a combined footprint of no more than 500 Kb.

22. The method of any of claims 2-21, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are selected based on whole genome sequencing of treated DNA samples from a population of subjects having a phenotype.

23. The method of any of claims 2-21, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are selected based on targeted sequencing of treated DNA samples from a population of subjects having a phenotype.

24. The method of any of claims 2-21, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are selected based on sequencing of one or more CpG islands in treated DNA samples from a population of subjects having a phenotype.

25. The method of any of claims 2-21, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions selected comprise a linear association of methylation signal with VAF.

26. The method of any of claims 2-21, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the target regions are unmethylated in subjects not having a phenotype.

27. The method of any of claims 1-26, further comprising appending adapters to the extracted DNA before contacting the extracted cell-free DNA molecules or their derivative with the panel of oligonucleotide probes.

28. The method of claim 27, wherein the adapters are Y-adaptors.

29. The method of claim 27 or 28, wherein the adapters include a universal priming site.

30. The method of any of claims 27-29, wherein the adapters include one or more methylated cytosines.

31. The method of any of claims 27-30, wherein the adapters include one or more unmethylated cytosines.

32. The method of any of claims 27-31, wherein the adapters include both methylated and unmethylated cytosines.

33. The method of any of claims 27-32, wherein the adapters include a molecular barcode.

34. The method of any one of claims 2-33, further comprising amplifying the enriched and / or treated DNA.

35. The method of claim 33, wherein amplifying comprises targeted PCR.

36. The method of claim 33, wherein amplifying comprises universal PCR.

37. The method of any of claims 2-36, wherein the agent or the combination of agents comprises a deaminating agent.

38. The method of claim 37, wherein the deaminating agent comprises sodium bisulfite.

39. The method of claim 37, wherein the deaminating agent comprises a deaminase or its catalytic domain.

40. The method of claim 39, wherein the combination of agents further comprises an oxidizing agent that oxidizes 5hmC and 5mC but not unmethylated cytosines thereby protecting 5hmC and 5mC from deamination.

41. The method of claim 39, wherein the combination of agents further comprises a glycosylation agent that glycosylates 5hmC but not unmethylated cytosines thereby protecting 5hmC from deamination.

42. The method of any one of claims 1-36, wherein the agent or the combination of agents comprises one or more methylation sensitive restriction enzymes (MSREs) or methylation dependent restriction enzymes (MDREs).

43. The method of claim 42, wherein the one or more MSREs is one or more of the MSRE is selected from HpaII, SalI, BbeI, NotI, SmaI, XmaI, MboI, BstUI, BstBI, ClaI, MluI, NaeI, NarI, PvuI, SacII, HpyCH41V, and HhaI.

44. The method of claim 42, wherein the one or more MDREs is selected from AbaSI, AoxI, BisI, BlsI, DpnI, FspEI, GlaI, GluI, KroI, LpnPI, MalI, MspJI, MteI, PcsI, PkrI, and SgeI.

45. The method of any one of claims 1-36, wherein the agent or the combination of agents comprises a methylation binding domain (MBD) protein, a 5-methylcytosine antibody, a 5-hmC antibody, or a DNA methyltransferase.

46. The method of any of claims 2-45, further comprising performing a barcoding PCR on the treated DNA to add a sample barcode before the contacting step.

47. The method of any of claims 1-46, further comprising spiking in methylated pUC19 DNA and unmethylated Lambda DNA as control.

48. The method of any of claims 1-47, wherein the panel of oligonucleotide probes comprises a plurality of primers and the contacting step comprises amplifying the treated DNA or its derivative using the plurality of primers to generate the enriched treated DNA.

49. The method of any of claims 1-33, wherein the panel of oligonucleotide probes comprises a plurality of baits and the contacting step comprises hybridization capture of the extracted DNA or its derivative using the plurality of baits to generate the enriched treated DNA.

50. The method of claim 49, wherein the treated DNA or its derivative is captured using an affinity-based capture method.

51. The method of claim 49, wherein the capture method comprises click chemistry-based affinity capture.

52. The method of claim 48, wherein the plurality of baits each includes a capture moiety comprising biotin, biotin dT, biotin-TEG, photocleavable (PC) biotin, or desthiobiotin-TEG, or biotin azide.

53. The method of claim 50, wherein the treated DNA or its derivative is captured using a solid phase coated with streptavidin.

54. The method of any of claims 2-53, further comprising performing high-throughput sequencing on the treated DNA or its derivative to generate sequence reads.

55. The method of claim 54, further comprising determining a methylation status of one or more of the cell-free DNA molecules or one or more of the target regions based on the sequence reads.

56. The method of claim 55, wherein determining the methylation status comprises determining a percentage of methylated CpG sites within one or more of the cell-free DNA molecules or one or more of the target regions based on the sequence reads.

57. The method of any of claims 2-56, wherein the method further comprises filtering out sequence reads derived from treated DNA with incomplete conversion.

58. The method of claim 57, wherein the filtering comprises filtering out sequence reads derived from cell-free DNA molecules having 1 or more CHG and / or CHH sites that appear to be methylated.

59. The method of any of claims 1-58, wherein the sample is a liquid sample.

60. The method of claim 59, wherein the liquid sample is a blood, serum, plasma, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample.

61. The method of claim 59 or 60, wherein the liquid sample is a blood, plasma, serum, or urine sample.

62. The method of any of claims 1-58, wherein the sample is a blood sample or a derivative sample thereof.

63. The method of any of claims 1-59, wherein the sample is a plasma sample.

64. The method of any of claims 1-58, wherein the sample comprises DNA from a tumor.

65. The method of any of claims 1-58, wherein the sample is a sample from a cancer tissue.

66. The method of any of claims 1-63, wherein the subject has, had, or is suspected or at risk of having a phenotype.

67. The method of claim 66, wherein the phenotype includes age, gender, ethnicity, smoking status, or pregnancy status.

68. The method of claim 66, wherein the phenotype is a disease.

69. The method of claim 68, wherein the disease is a cancer.

70. The method of claim 69, wherein the cancer is selected from ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma,neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.

71. The method of claim 69, wherein the cancer is selected from a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro- esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whipple resection.

72. The method of claim 69, wherein the cancer is selected from lung cancer, breast cancer, bladder cancer, and colorectal cancer.

73. The method of claim 72, wherein the breast cancer is HER2 positive breast cancer, hormone receptor (HR) positive breast cancer, or triple negative breast cancer (TNBC).

74. The method of claim 72, wherein the lung cancer is non-small cell lung cancer (NSCLC), adenocarcinoma NSCLC, squamous NSCLC, or small cell lung cancer (SCLC).

75. The method of claim 72, wherein the bladder cancer is muscle invasive bladder cancer (MIBC).

76. The method of any of claims 55-75, wherein the methylation status of one or more of the cell-free DNA molecules or one or more of the target regions is indicative of the presence or absence of the cancer.

77. The method of any of claims 1-63, wherein the sample comprises DNA from a transplanted organ.

78. The method of any of claims 1-63, wherein the subject is a transplant recipient.

79. The method of any of claims 1-63, wherein the sample comprises DNA from a fetus.

80. The method of any of claims 1-63, wherein the subject is a pregnant female.

81. The method of any of claims 1-80, wherein the subject is an animal.

82. The method of any of claims 1-80, wherein the subject is a domesticated animal.

83. The method of any of claims 1- 80, wherein the subject is a dog, a cat, a hamster, a pig, a cattle, a sheep, a goat, a horse, a deer, a buffalo, a monkey, a mouse, a rat, a guinea pig, a llama, a cow, a fish, a reptile.

84. The method of any of claims 1-83, wherein the method further comprises enriching the sample DNA molecules for DNA molecules that are between 70 and 500 base pairs in length.

85. The method of any of claims 1-83, wherein the method further comprises enriching the sample DNA molecules for DNA molecules that are between 100 and 200 base pairs in length.

86. The method of any of claims 1-83, wherein the method further comprises enriching the sample DNA molecules for DNA molecules that are between 130 and 170 base pairs in length.

87. The method of any of claims 1-86, wherein one or more steps of the method are performed in a microfluidic device.

88. The method of any of claims 1-87, further comprising performing the method on a second sample from the subject collected at a different time point.

89. The method of any one of claims 1-88, further comprising calculating a differential methylation allele fraction (DMAF).

90. The method of any one of claims 1-89, wherein the method is capable of detecting a DMAF of a target region.

91. A method of preparing a composition of non-naturally occurring DNA, comprising:(a) extracting DNA molecules from a sample of a subject; (b) contacting the extracted DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to 5-200,000 different target regions differentially methylated in a phenotype, thereby generating enriched treated DNA, wherein the oligonucleotide probes for at least one of the target regions are designed to cover the target region with a tiling density of between 0.9x and 2x, wherein at least 90% of the 5-200,000 different target regions each includes at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp; and (c) generating sequence reads by performing high-throughput sequencing on the enriched treated DNA or its derivative.

92. The method of claim 91, further comprising treating (i) the extracted DNA or its derivative, or (ii) the enriched DNA or its derivative, with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines.

93. The method of claim 91 or 92, further comprising determining a methylation status of one or more of the extracted DNA molecules or a portion thereof or one or more of the target regions based on the sequence reads.

94. The method of claim 93, wherein determining the methylation status comprises determining a percentage of methylated CpG sites within one or more of the extracted DNA molecules or a portion thereof or one or more of the target regions based on the sequence reads.

95. The method of claim 94, wherein the methylation status is quantified on a fragment level based on the co-methylation status of all CpGs.

96. The method of claim 91 or 92, wherein the method further comprises identifying a sequence read mapped to a target region and having at least 3 CpG sites and a methylation fraction of at least 0.8 as a hypermethylated fragment, and determining a hypermethylated fragment fraction for the target region based on the number of hypermethylated fragments and the total number of sequence reads mapped to the target region and having at least 3 CpG sites.

97. The method of any of claims 91-96, wherein the method further comprises filtering out sequence reads derived from treated DNA with incomplete conversion.

98. The method of claim 97, wherein the filtering comprises filtering out sequence reads derived from extracted DNA molecules having 1 or more CHG and / or CHH sites that appear to be methylated.

99. The method of any of claims 91-98, wherein the methylation status is determined using only sequence reads having one or more CHG or CHH sites.

100. The method of any of claims 1-99, wherein the method further comprises classifying the sample as having or not having the phenotype / disease / cancer using a model-free algorithm.

101. The method of any of claims 1-99, wherein the method further comprises classifying the sample as having or not having the phenotype / disease / cancer using a logistic regression classifier.

102. The method of claim 101, wherein the classifier selects one or more targets based on a training set during cross-validation.

103. The method of claim 102, wherein the input to the classifier includes a transformed version of hypermethylated fraction of each selected target.

104. The method of claim 103, wherein the classifier estimates posterior probability of a phenotype / disease / cancer-positive status.

105. A method of preparing a composition of non-naturally occurring DNA, comprising: (a) extracting cell-free DNA molecules from a sample of a subject; (b) treating the extracted cell-free DNA molecules or their derivative with an agent or a combination of agents to generate treated DNA, wherein the agent or the combination of agents discriminates between methylated and unmethylated cytosines; (c) contacting the treated DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to 5-200,000 different target regions having CpG sites that are differentially methylated in a phenotype, thereby generating enriched treated DNA, wherein the oligonucleotideprobes for at least one of the target regions are designed to cover the target region with a tiling density of between 0.9x and 2x.

106. The method of any of claims 1-105, wherein the method further comprises classifying the sample based on both the methylation status of the cell-free DNA molecules or their derivatives and one or more additional signals.

107. The method of claim 106, wherein the one or more additional signals comprise fragmentomics, genomic alterations, proteins, extracellular vesicles, lipids, miRNA, or immune profiling.

108. The method of any of claims 1-107, wherein the method further comprises classifying the sample based on one or more additional characteristics of the subject.

109. The method of claim 108, wherein the one or more additional characteristics comprise age, gender, ethnicity, smoking status, polygenic risk score, or medical history.

110. A method comprising: a) obtaining a first plurality of samples from healthy subjects and a second plurality of samples from subjects with a phenotype; b) generating a first set of sequencing data for the first plurality of samples and a second set of sequencing data for the second plurality of samples using a sequencing panel comprising 5- 200,000 different target regions having CpG sites that are differentially methylated in the phenotype; c) selecting, for subsequent use in detecting the phenotype in a subject, a subset of the target regions for the phenotype, wherein the subset of the target regions: i) are hypermethylated in the sequencing data for the second plurality of samples; and ii) exhibits a linear relationship between the methylation signal and VAF.

111. The method of 110, wherein the subset of the target regions are selected to have a linear relationship between differential methylation allele fraction (DMAF) as determined by methylation signals and variant allele frequence (VAF) as determined by genomic sequences,wherein the Pearson correlation between the DMAF and the VAF is at least 0.8, at least 0.85, at least 0.9, or at least 0.

95.

112. The method of claims 1, 91, 105, further comprising classifying the sample as having or not having the phenotype based on DMAF of a subset of the target regions selected by the method of claim 110 or 111.

Citation Information

Patent Citations

  • Linked duplex target capture

    WO2017168332A1

  • Compositions, methods, and kits for isolating nucleic acids

    WO2018156418A1

  • Methods for cancer detection and monitoring by means of personalized detection of circulating tumor DNA

    WO2019200228A1

  • Methylation biomarker generation and analysis

    WO2025151453A1

  • Improved method for methylation biomarker generation and analysis

    WO2026043738A1