Improved method for methylation biomarker generation and analysis

A targeted panel for methylation biomarker discovery, focusing on specific CpG regions, addresses the inefficiencies of existing methods by reducing costs and improving throughput and information relevance in methylation biomarker detection.

WO2026043738A1PCT designated stage Publication Date: 2026-02-26NATERA INC +4

Patent Information

Application Number
PCT/US2025/042194
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-20
Filing Date
2025-08-15
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing methods for methylation biomarker discovery, such as whole genome methylation sequencing and CpG microarrays, suffer from low sequencing depth, high cost, and limited sample throughput, while commercial methylome panels have non-specific and incomplete coverage of informative CpG regions.

Method used

A targeted panel for methylation biomarker discovery is used, focusing on CpG regions with at least 3 CpG sites and a CpG density of at least 0.02 CpG/bp, covering 10-500 Mb of sequence space, and employing high-throughput sequencing to generate sequencing reads.

Benefits of technology

This approach reduces sequencing cost per sample, increases sample throughput, and provides more relevant information than conventional methods, enabling better methylation biomarker discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000065_0001
    Figure IMGF000065_0001
  • Figure IMGF000065_0002
    Figure IMGF000065_0002
  • Figure IMGF000065_0003
    Figure IMGF000065_0003
Patent Text Reader

Abstract

Disclosed are methods of preparing a composition of non-naturally occurring DNA, comprising: (a) extracting DNA from a sample of a subject; (b) contacting the extracted DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to a plurality of preselected genomic regions, thereby generating selected DNA, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome, wherein the extracted DNA or its derivative or the selected DNA or its derivative is further treated with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines; and (c) performing high-throughput sequencing on the selected DNA or its derivative and generating sequencing reads.
Need to check novelty before this filing date? Find Prior Art

Description

IMPROVED METHOD FOR METHYLATION BIOMARKER GENERATION AND ANALYSIS CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 685,006 filed August 20, 2024, the contents of which are hereby incorporated by reference in their entirety. BACKGROUND

[0002] Detection of diseases at earlier stages of progression can substantially improve patient outcome. Generally, less invasive tests for diseases and other conditions tend to be administered more often than more invasive tests, in part due to risk and compliance considerations. Accordingly, improvements in the ability of non-invasive (or minimally- invasive) tests increase the chances that such tests will be more routinely used, and that conditions will be more often detected (and at earlier stages). For example, a test that requires a blood draw or cheek swab is generally less risky for patients than procedures that require a surgical procedure (e.g., a biopsy), and such a test would thus generally be preferable for patients and clinicians. However, physiological manifestations of a condition might be subtle, especially at earlier stages of a disease, making them more difficult to detect using samples obtained through less-invasive means.

[0003] Previously, whole genome methylation sequencing (e.g., whole genome bisulfite sequencing) of negative (healthy or non-diseased) and positive (diseased) samples were used for methylation biomarker discovery, which suffers from issues such as low sequencing depth, high sequencing cost, and low sample throughputs. On the other hand, CpG microarrays, while much less expensive, only measure DNA methylation state at a single CpG site per molecule level and therefore can only provide limited information. Commercially available methylome panels are typically based on published database and have limited and non-specific coverage of themethylome, leaving out a large portion of informative CpG regions while including less informative or non-informative CpG regions. SUMMARY OF THE DISCLOSURE

[0004] Using a targeted panel for methylation biomarker discovery is more efficient than whole genome methylation sequencing and provides more relevant information than conventional CpG microarrays or methylome panels. Targeting only the CpG regions likely to be of interest (e.g., CpG regions each comprising at least 3, at least 4, at least 5, or at least 6 CpG sites and a CpG density of at least 0.02, at least 0.03, at least 0.04, or at least 0.05 CpG / bp) lower the amount of sequence data per sample needed, which in turn leads to (1) lower sequencing cost per sample, (2) higher samples throughput, and (3) better performance in methylation biomarker discovery.

[0005] Provided herein in one aspect is a method of preparing a composition of non- naturally occurring DNA, comprising: (a) extracting DNA from a sample of a subject; (b) contacting the extracted DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to a plurality of preselected genomic regions, thereby generating selected DNA, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome, wherein the extracted DNA or its derivative or the selected DNA or its derivative is further treated with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines; and (c) performing high-throughput sequencing on the selected DNA or its derivative and generating sequencing reads.

[0006] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 4 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 5 CpGsites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 6 CpG sites.

[0007] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.03 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.04 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

[0008] In some embodiments, at least 80%, or at least 90%, or at least 95% of the preselected genomic regions are the CpG regions.

[0009] In some embodiments, the preselected genomic regions cover 20-300 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 50-200 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 100-150 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 130-140 Mb of sequence space in human genome.

[0010] In some embodiments, the preselected genomic regions cover 1-5% of sequence space in human genome. In some embodiments, the preselected genomic regions cover 2-3% of sequence space in human genome.

[0011] In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 80 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 60 bp. In some embodiments, in at least 50%, atleast 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 50 bp.

[0012] In some embodiments, the method further comprises fragmenting the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes.

[0013] In some embodiments, the method further comprises appending adaptors to the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes. In some embodiments, the adaptors are Y-adaptors. In some embodiments, the adaptors comprise a universal priming site, a molecular barcode, and / or methylated cytosines.

[0014] In some embodiments, the extracted DNA or derivative thereof is treated with a deaminating agent that converts (directly or indirectly through one or more other agents) unmethylated cytosines to uracils before the contacting with the panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a deaminating agent that converts (directly or indirectly through one or more other agents) unmethylated cytosines to uracils after the contacting with the panel of oligonucleotide probes. In some embodiments, the deaminating agent (e.g., a chemical reagent such as sodium bisulfite) directly converts unmethylated cytosines to uracils. In some embodiments, the deaminating agent converts unmethylated cytosines to uracils in combination with one or more other agents – for example, in order to convert unmethylated cytosines to uracils, the extracted DNA or derivative thereof may be first treated with an oxidizing agent (e.g., TET or its catalytic domain) that oxidizes one or more forms of methylated cytosines (e.g., 5hmC or 5mC), followed by treatment with a deaminase (e.g., APOBEC or its catalytic domain).

[0015] In some embodiments, the deaminating agent comprises sodium bisulfite. In some embodiments, the deaminating agent comprises a deaminase such as APOBEC or its catalytic domain. In some embodiments, the deaminating agent comprises APOBEC3A (A3A) or its catalytic domain.

[0016] In some embodiments, the extracted DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, before the contacting with the panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5- hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, after the contacting with the panel of oligonucleotide probes.

[0017] In some embodiments, the method further comprises amplifying the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes.

[0018] In some embodiments, the method further comprises amplifying the selected DNA or derivative thereof after the contacting with the panel of oligonucleotide probes.

[0019] In some embodiments, the amplifying adds a sample barcode, and wherein a plurality of libraries of the selected DNA or derivative thereof each produced from a different sample are sequenced together in one sequencing lane.

[0020] In some embodiments, the method further comprises performing size selection before or after the amplifying step.

[0021] In some embodiments, the oligonucleotide probe is a hybrid capture probe. In some embodiments, the oligonucleotide probe comprises a probe section covalently linked to a primer section.

[0022] In some embodiments, the CpG regions do not comprise CpG sites located in repeat elements excluding GC-rich regions.

[0023] In some embodiments, the sample of DNA comprises cellular DNA derived from a fresh-frozen tissue or an FFPE tissue. In some embodiments, the sample of DNA comprisescell-free DNA derived from a blood, plasma, serum, or urine sample. In some embodiments, the sample of DNA comprises tumor DNA.

[0024] In some embodiments, the method further comprises producing a set of libraries of the selected DNA from a set of samples from a population of subjects, and using the sequencing reads of the set of libraries of the selected DNA to identify preselected genomic regions of which the co-methylation patterns of the CpG sites are substantially consistent.

[0025] In some embodiments, the method further comprises producing a first set of libraries of the selected DNA from a first set of samples from a population of diseased subjects, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-diseased subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of diseased subjects and the population of healthy or non-diseased subjects.

[0026] In some embodiments, the method further comprises producing a first set of libraries of the selected DNA from a first set of samples from a population of subjects having cancer, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-cancerous subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of subjects having cancer and the population of healthy or non-cancerous subjects.

[0027] In some embodiments, the method further comprises identifying a plurality of differentially methylated regions (DMRs) between the first set of samples and the second set of samples.

[0028] Provided herein in another aspect is a composition comprising a sample of DNA that has been selectively enriched at a plurality of preselected genomic regions, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sitesand a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10- 500 Mb of sequence space in human genome.

[0029] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 4 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 5 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 6 CpG sites.

[0030] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.03 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.04 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

[0031] In some embodiments, at least 80%, or at least 90%, or at least 95% of the preselected genomic regions are the CpG regions.

[0032] In some embodiments, the preselected genomic regions cover 20-300 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 50-200 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 100-150 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 130-140 Mb of sequence space in human genome.

[0033] In some embodiments, the preselected genomic regions cover 1-5% of sequence space in human genome. In some embodiments, the preselected genomic regions cover 2-3% of sequence space in human genome.

[0034] In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp.In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 80 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 60 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 50 bp.

[0035] In some embodiments, the CpG regions substantially do not comprise CpG sites located in repeat elements excluding GC-rich regions.

[0036] In some embodiments, the sample of DNA comprises uracils or thymines converted or derived from unmethylated cytosines.

[0037] In some embodiments, the sample of DNA comprises an adaptor section linked to a genomic DNA section. In some embodiments, the adaptor section comprises a universal priming site, a molecular barcode, and / or methylated cytosines.

[0038] In some embodiments, the sample of DNA comprises cellular DNA derived from a fresh-frozen tissue or an FFPE tissue. In some embodiments, the sample of DNA comprises cell-free DNA derived from a blood, plasma, serum, or urine sample. In some embodiments, the sample of DNA comprises tumor DNA.

[0039] Provided herein in a further aspect is a composition comprising a panel of oligonucleotide probes designed for selectively enriching a plurality of preselected genomic regions from a sample of DNA by hybrid capture or linked target capture, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome.

[0040] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 4 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 5 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 6 CpG sites.

[0041] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.03 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.04 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

[0042] In some embodiments, at least 80%, or at least 90%, or at least 95% of the preselected genomic regions are the CpG regions.

[0043] In some embodiments, the preselected genomic regions cover 20-300 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 50-200 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 100-150 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 130-140 Mb of sequence space in human genome.

[0044] In some embodiments, the preselected genomic regions cover 1-5% of sequence space in human genome. In some embodiments, the preselected genomic regions cover 2-3% of sequence space in human genome.

[0045] In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regionsthe maximum distance between any two adjacent CpG sites is 80 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 60 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 50 bp.

[0046] In some embodiments, the CpG regions do not comprise CpG sites located in repeat elements excluding GC-rich regions.

[0047] In some embodiments, the oligonucleotide probe is a hybrid capture probe. In some embodiments, the oligonucleotide probe comprises a probe section covalently linked to a primer section.

[0048] Provided herein in an additional aspect is a composition comprising a panel of oligonucleotide probes designed for selectively enriching a plurality of preselected differentially methylated regions (DMRs) from a sample of DNA by hybrid capture or linked target capture, wherein the DMRs are identified and selected by: producing a first set of libraries of the selected DNA from a first set of samples from a population of diseased subjects, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-diseased subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of diseased subjects (e.g., subjects having cancer) and the population of healthy or non-diseased subjects.

[0049] In some aspects, the disclosed approach involves determining a set of features which robustly differentiate diseased samples from healthy or non-diseased samples. Determining the set of features may comprise identifying sets of methylation biomarkers. In example embodiments, these biomarkers are selected from genomic regions containing multiple potentially differentially methylated sites, termed methylation haplotype blocks (MHBs), such as potentially differentially or variably methylated CpG loci. As used herein, CpG loci are regionsof deoxyribonucleic acid (DNA) in which a cytosine nucleotide is linearly followed by a guanine nucleotide in the 5' → 3' direction. Read-level information can be leveraged to identify regions in which patterns of co-methylation are consistently and significantly different between diseased and healthy / non-diseased samples. For example, read-level methylation information can be used to jointly model the methylation status of different CpGs in each genomic window. Various embodiments generate a set of biomarkers which can be used to detect or determine a likelihood of a disease or other condition. These biomarkers can also be used to locate the anatomic site of the disease. Subsets of biomarkers can be selected to represent samples of patients of different ages, genders, ethnicities, etc.

[0050] In various embodiments, computational modules are provided for in-depth characterization of genomic regions with methylated cytosines (e.g., CpG di-nucleotides) for appropriate marker selection. For example, tools are provided for full characterization of genomic regions with respect to CpG methylation and potential as methylation biomarker (MBM). Example tools specifically address the needs for efficient manipulation of a large set of methylation sequencing data, focusing the biomarker discovery search space to relevant parts of the genome, ranking candidate marker regions based on a comprehensive knowledge-base of their characteristics. In various embodiments, a module is configured to pre-select potentially informative regions of the genome for downstream analysis. This may involve, for example, using available data to identify methylation cluster regions with high CpG density and co-high methylation, followed by genome segmentation to find MHBs for consideration in methylation biomarker selection.

[0051] In one aspect, various embodiments relate to a method comprising: obtaining, from a cohort of subjects, a first plurality of samples from subjects with a phenotype and a second plurality of samples from subjects without the phenotype; generating a first set of sequencing data for the first plurality of samples and a second set of sequencing data for the second plurality of samples; constructing, using the first set of sequencing data, the second set of sequencing data, or a combination thereof, one or more sequencing files; preliminarily segmenting data in the one or more sequencing files to obtain one or more sets of preliminaryDNA segments; and generating, for subsequent use in developing one or more tools for detecting the phenotype in a subject, a set of candidate methylation biomarkers for the phenotype, the generating the set of candidate methylation biomarkers comprising applying a model to the one or more sets of preliminary DNA segments.

[0052] In various embodiments, generating the first and second sets of sequencing data comprises performing methylation sequencing of DNA in the first and second pluralities of samples, respectively. In various embodiments, the first plurality of samples are cell-free DNA (cfDNA) samples and the second plurality of samples are tissue samples. In various embodiments, the one or more files are generated based on one or more Binary Alignment Map (BAM) files and / or Sequence Alignment Map (SAM) files. In various embodiments, constructing the one or more sequencing files comprises generating one or more files in which each line corresponds to a DNA fragment. In various embodiments, for each DNA fragment, the one or more files identify a chromosome and start and end coordinates. In various embodiments, for each DNA fragment, the one or more files identify a methylation pattern. In various embodiments, for each DNA fragment, the one or more files identify a methylation score. In various embodiments, for each DNA fragment, the one or more files identify at least one coordinate corresponding to one or more methylation sites. In various embodiments, preliminarily segmenting the data in the one or more sequencing files comprises constructing a set of preliminary segments based at least on a reference methylome. In various embodiments, preliminarily segmenting the data in the one or more sequencing files comprises constructing a set of preliminary segments based at least on a set of individual methylomes from one or more samples from the cohort of subjects. In various embodiments, preliminarily segmenting the data in the one or more sequencing files is based on one or more criteria. In various embodiments, the one or more criteria comprises a primer length threshold for primer length. In various embodiments, the primer length threshold is between 75 bp and 150 bp. In various embodiments, the one or more criteria comprises a probe length threshold for probe length. In various embodiments, the probe length threshold is between 75 bp and 150 bp. In various embodiments, the one or more criteria comprises a fragment length threshold based on expectedfragment lengths. In various embodiments, the fragment length threshold is between 50 to 1200 bp, between 70 to 800 bp, between 100 to 200 bp, between 130 to 180 bp, between 50 to 200 bp, between 60 and 200 bp, between 60 and 150 bp, or between 60 and 100 bp. In some embodiments, the fragment length threshold ranges from 100 to 200 bp, from 120 to 180 bp, from 140 to 160 bp, from 150 to 170 bp, from 160 to 190 bp, or from 170 to 220 bp. In some embodiments, the fragment length threshold is at or less than 500, 400, 200, 150, 100, 90, 75, 70, 60, or 50 bp in length. In various embodiments, the one or more criteria comprises a processivity threshold for the DNA (cytosine-5)-methyltransferase 1 (DNMT1) enzyme. In various embodiments, the processivity threshold is at least between 20 bp and 1000 bp. In various embodiments, the one or more criteria comprises a maximum allowable distance between adjacent CpG loci. In various embodiments, the maximum allowable distance is 150 bp. In various embodiments, preliminarily segmenting the data comprises removing one or more CpGs that are found in less than a threshold percentage of total samples. In various embodiments, the threshold percentage is between about 15% and about 35%.

[0053] In various embodiments, preliminarily segmenting the data comprises filtering CpG ridges to contain at least a threshold number of CpG sites. In various embodiments, the threshold number is at least 3 CpG sites. In various embodiments, preliminarily segmenting the data comprises filtering ridges to have a CpG density of at least a density threshold. In various embodiments, the density threshold is at least 0.02 CpG / bp. In various embodiments, preliminarily segmenting the data comprises filtering ridges (i) to autosomes only, (ii) to having at least 5 CpG sites, and (iii) to having a density of at least 0.03 CpG / bp. However, such preliminarily segmenting would not be necessary when a targeted panel of oligonucleotide probes as described herein are used for selective enrichment of only the CpG regions of interest (e.g., CpG regions each comprising at least 3, at least 4, at least 5, or at least 6 CpG sites and a CpG density of at least 0.02, at least 0.03, at least 0.04, or at least 0.05 CpG / bp) for targeted methylation sequencing.

[0054] In various embodiments, generating the set of candidate methylation biomarkers for the phenotype comprises further segmenting the data. In various embodiments, the model isa Bayesian model. In various embodiments, the model is used to infer and / or estimate change points, which indicate boundaries of regions where a methylation load signal has a variation below a threshold in a given region across all samples within the cohort. In various embodiments, the model fits a piecewise constant function to beta-binomially distributed methylation load signals. In various embodiments, the model employs stochastic variational inference (SVI). In various embodiments, further comprising generating a select set of methylation biomarkers from the set of candidate methylation biomarkers. In various embodiments, generating the select set of methylation biomarkers comprises scoring the methylation biomarkers in the set of candidate methylation biomarkers. In various embodiments, candidate biomarkers are scored based on a signal-to-noise ratio (SNR) so as to prioritize regions where background noise is minimal compared to phenotype signal. In various embodiments, generating the select set of methylation biomarkers comprises prioritizing candidate methylation biomarkers based on one or more criteria. In various embodiments, the one or more criteria comprises designability of probes for the biomarkers. In various embodiments, methylation biomarkers in repeat or low-complexity regions are deprioritized. In various embodiments, further comprising prioritizing the methylation biomarkers for predictive powers across subpopulations. In various embodiments, the select set of methylation biomarkers is a first select set, and wherein the method further comprises generating a second select set of methylation biomarkers. In various embodiments, the first select set of methylation biomarkers corresponds to a first disease subtype and the second select set of methylation biomarkers corresponds to a second disease subtype. In various embodiments, the first select set of methylation biomarkers corresponds to a first population and the second select set of methylation biomarkers corresponds to a second population. In various embodiments, the first select set of methylation biomarkers corresponds to a first gender and the second select set of methylation biomarkers corresponds to a second gender. In various embodiments, the first select set of methylation biomarkers corresponds to a first ethnicity and the second select set of methylation biomarkers corresponds to a second ethnicity. In various embodiments, the first select set of methylation biomarkers corresponds to a first disease subtype and the second select set of methylation biomarkers corresponds to a second disease subtype.

[0055] In another aspect, various embodiments relate to a method comprising: obtaining a first plurality of samples from subjects with a phenotype and a second plurality of samples from subjects without the phenotype; generating methylation sequencing data for the first plurality of samples and the second plurality of samples; constructing, based on the generated methylation sequencing data, one or more sequencing files; preliminarily segmenting data in the one or more sequencing files to obtain a set of preliminary DNA segments; further segmenting the set of preliminary DNA segments to generate candidate methylation biomarkers; and generating, for subsequent use in developing one or more tools for detecting the phenotype in a subject, a select set of methylation biomarkers from the candidate methylation biomarkers.

[0056] In another aspect, various embodiments relate to a method comprising: generating a first set of methylation sequencing data for a first plurality of samples from subjects with a phenotype and a second set of methylation sequencing data for a second plurality of samples from subjects without the phenotype; constructing, using the first set of methylation sequencing data, the second set of methylation sequencing data, or a combination thereof, one or more sequencing files; preliminarily segmenting data in the one or more sequencing files to obtain one or more sets of preliminary DNA segments; generating a set of candidate biomarkers for the phenotype, the generating the set of candidate biomarkers comprising applying a model to the one or more sets of preliminary DNA segments; and generating a select set of biomarkers from the set of candidate biomarkers for subsequent detection of the phenotype in one or more subjects.

[0057] In another aspect, various embodiments relate to a method for developing, for detection of the phenotype in a subject, a panel based on the set of candidate methylation biomarkers generated by any of the above methods.

[0058] In another aspect, various embodiments relate to a method of detecting a phenotype in a subject using a panel developed based on a select set of methylation biomarkers generated by any of the above methods.

[0059] In another aspect, various embodiments relate to a panel for detecting a phenotype in a subject, the panel being based on a select set of methylation biomarkers generated by any of the above methods.

[0060] In another aspect, various embodiments relate to a method of detecting colorectal cancer in a subject using a panel developed based on a select set of methylation biomarkers generated by any of the above methods.

[0061] In another aspect, various embodiments relate to a panel for detecting colorectal cancer in a subject, the panel being based on a select set of methylation biomarkers generated by any of the above methods.

[0062] In another aspect, various embodiments relate to a computing system comprising one or more processors configured to perform any of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] FIG. 1 depicts an overall system that includes a computing system that may interface with other components to enable various aspects of the disclosed approach, according to various embodiments.

[0064] FIG. 2 depicts an example process for generating biomarkers, according to various embodiments.

[0065] FIG. 3A depicts an example workflow, according to various embodiments. FIG. 3B depicts cohort characteristics, with a histogram of age distribution among healthy individuals and CRC patients (left), self-reported ethnicity among healthy individuals and CRC patients (middle), and proportion of CRC patients by stage of disease progression, further stratified by microsatellite status (right), according to various embodiments.

[0066] FIG. 4 provides cohort level distributions of coverage, methylation level, and protection rate, according to various embodiments.

[0067] FIG. 5 provides a cohort CpG coverage plot, according to various embodiments. The number of CpGs covered is plotted against the fraction of samples in each group covering those CpGs.

[0068] FIG. 6 provides length (left panel), number of CpG sites (middle panel), and density distribution (right panel) of the filtered ridge set, according to various embodiments. Red line shows the median, and shaded region is the interquartile range (IQR).

[0069] FIG. 7 depicts a comparison of different baselines from the literature, according to various embodiments.

[0070] FIG. 8 illustrates an example CpGBlocks latent variable model of processive methylation of CpG sites by DNA (cytosine-5)-methyltransferase 1 (DNMT1), according to various embodiments.

[0071] FIG. 9 includes plots corresponding to a machine learning model developed to discover differentially-methylated regions between healthy and disease cohorts as well as between subpopulations within the disease cohort, including age, sex, stage, histology, and MSI vs microsatellite stability (MSS), according to various embodiments. The plots illustrate a segmentation procedure using the machine learning model. For a region, FIG.9 plots the methylation load of the healthy cohort and tumor cohort and demonstrates the segmentation procedure and scoring / ranking. The horizontal bars in the lower panel show the differential methylation load.

[0072] FIG. 10 depicts differential methylation load and CpG density distribution of biomarker candidates, according to various embodiments, colored by their differential methylation status: hyper-methylated, hypo-methylated, and not significantly differentially methylated according to a Wald test of significance.

[0073] FIG. 11 depicts number of biomarkers and the differential methylation score cutoff based on a preset 2Mbp total panel size (left) and CpG density, length and methylation load distribution of biomarker candidates (right), according to various embodiments.

[0074] FIGs.12A and 12B depict extracting fragment information and combining CpG methylation status of two paired reads, with 12A illustrating overlapping reads and 12B illustrating non-overlapping reads, according to various embodiments.

[0075] FIG. 13 provides an example of a CpG file (referred to interchangeably as “cpg file” or “.cpg file”) containing fragments mapped to a segment of a chromosome, according to various embodiments.

[0076] FIG. 14 depicts an expanded methylation pattern in an example CpG file, according to various embodiments.

[0077] FIG. 15 illustrates “CpGReads,” showing extended CpG status of reads in a genomic window from the CpGReads module, according to various embodiments.

[0078] FIG. 16 illustrates CpGReads and compact CpG block visualization along with inferred methylation states, according to various embodiments.

[0079] FIG. 17 depicts population level plots of CpG matrices, according to various embodiments, in which “H” corresponds to “healthy” subjects in a cohort and “D” corresponds to “disease” subjects in the cohort.

[0080] FIG. 18 depicts a triangle heatmap showing pairwise correlation of methylation levels between CpGs, according to various embodiments. CpG coordinates are listed in genomic order on the right, and positions of interest are highlighted in red. The bars correspond to fraction methylated or beta-value for each CpG in this region of interest.

[0081] FIG. 19 depicts triangle heat maps plotted according to various embodiments. Both regions are hyper-methylated in CRC samples (left column) compared to healthy samples(right column) as depicted by the gray bars. However, the level of co-methylation is different between the two targets, which can be used as a way of ranking genomic regions for marker selection.

[0082] FIG. 20 depicts overall findings of an example analysis according to various embodiments. Red line shows the median, and the shaded region is the interquartile range (IQR). In total, 184,265 differentially methylated regions between healthy and CRC cohorts were identified, spanning more than 1.71 million CpGs and more than 18.76 Mbp.76% of all identified regions were hypermethylated (70% CpG islands, 1.3% CpG shelves, 21% CpG shores) and the rest were hypomethylated (7.3% CpG islands, 8.6% CpG shelves, 25% CpG shores). Out of which, 11,579 biomarkers spanning 1.69Mbp were selected for the final panel based on various filters and criteria.

[0083] FIG. 21 depicts a subpopulation analysis according to various embodiments, with the right panel representing microsatellite status and the left panel representing mucinous adenocarcinoma. Red line shows the median, and shaded region is the interquartile range (IQR). 220 hypermethylated regions were differentially methylated between CRC MSI samples and CRC MSS samples and healthy cfDNA samples. Subtype-specific differentially methylated regions, such as 62 hyper-methylated regions that were differentially methylated between mucinous adenocarcinoma samples and the rest of the cohort, were identified.

[0084] FIG. 22 depicts a block diagram of a representative server system and client computer system usable to implement various embodiments of the present disclosure.

[0085] FIG. 23 shows CpG distribution across autosomes.

[0086] FIG. 24 shows CpG distribution after removing overlaps with repeat regions.

[0087] FIG. 25 shows histograms of distributions of length, CpG number, and CpG density after removing singletons (CpG number =1).

[0088] FIG. 26 shows survival function of number of CpG loci and total length of DNA covered given a minimum CpG number and a minimum CpG density per ridge.

[0089] FIG. 27 shows an exemplary experimental workflow for testing the methylation discovery hybrid capture panel.

[0090] Fig.28 shows raw pairs, mapping efficiency, read duplicate fraction, and mean insert size, as grouped by sample type.

[0091] Fig.29 shows CpG, CHG, CHH methylation levels, as well as conversion and protection rates, as grouped by sample type.

[0092] Fig.30 shows on-target rate, bait coverage, target coverage, and fold-80 base penalty, as grouped by sample type.

[0093] Fig.31 shows raw pairs, mapping efficiency, read duplicate fraction, mean insert size, CpG, CHG, CHH methylation levels, conversion and protection rates, on-target rate, bait coverage, target coverage, and fold-80 base penalty, as grouped by probe concentration level.

[0094] Fig.32 shows raw pairs, mapping efficiency, read duplicate fraction, and mean insert size, as grouped by sample plexing level.

[0095] Fig.33 shows CpG, CHG, CHH methylation levels, as well as conversion and protection rates, as grouped by sample plexing level.

[0096] Fig.34 shows on-target rate, bait coverage, target coverage, and fold-80 base penalty, as grouped by sample plexing level. DETAILED DESCRIPTION

[0097] Epigenetic alterations have been shown to govern many phenotype or disease states, including cancer. A key epigenetic marker in humans is the DNA methylation of CpGs. DNA methylation is fundamental to gene regulation and cell identity. DNA methylation (e.g.,patterns of CpG methylation) can directly alter gene regulation and cell identity through alteration of chromatin accessibility. Methylation can positively or negatively be associated with transcription factor binding events, which in turn alter the cell’s transcription programs.

[0098] Provided herein in one aspect is a method of preparing a composition of non- naturally occurring DNA, comprising: (a) extracting DNA from a sample of a subject; (b) contacting the extracted DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to a plurality of preselected genomic regions, thereby generating selected DNA, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome, wherein the extracted DNA or its derivative or the selected DNA or its derivative is further treated with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines; and (c) performing high-throughput sequencing on the selected DNA or its derivative and generating sequencing reads.

[0099] Provided herein in another aspect is a composition comprising a sample of DNA that has been selectively enriched at a plurality of preselected genomic regions, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10- 500 Mb of sequence space in human genome.

[0100] Provided herein in a further aspect is a composition comprising a panel of oligonucleotide probes designed for selectively enriching a plurality of preselected genomic regions from a sample of DNA by hybrid capture or linked target capture, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome.

[0101] Preselected genomic regions for methylation biomarker discovery

[0102] Using a targeted panel for methylation biomarker discovery according to the methods described herein is more efficient and cost-effective than whole genome methylation sequencing (e.g., whole genome bisulfite sequencing). In addition, using a targeted panel for methylation biomarker discovery according to the methods described herein provides more relevant information than conventional CpG microarrays and better and more relevant methylome coverage than commercial methylome panels, allowing discovery of biologically informative biomarkers. Accordingly, in some embodiments, the method described herein involves a panel of oligonucleotide probes designed to hybridize to a plurality of preselected genomic regions, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome.

[0103] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 4 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 5 CpG sites. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 6 CpG sites.

[0104] In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.03 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.04 CpG / bp. In some embodiments, at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

[0105] In some embodiments, the preselected genomic regions cover 20-300 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 50-200 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover 100-150 Mb of sequence space in human genome. In some embodiments, thepreselected genomic regions cover 130-150 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover about 110 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover about 136 Mb of sequence space in human genome. In some embodiments, the preselected genomic regions cover about 147 Mb of sequence space in human genome.

[0106] In some embodiments, the preselected genomic regions cover 1-5% of sequence space in human genome. In some embodiments, the preselected genomic regions cover 2-3% of sequence space in human genome.

[0107] In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 80 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 60 bp. In some embodiments, in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 50 bp.

[0108] In some embodiments, the CpG regions do not comprise CpG sites located in repeat elements excluding GC-rich regions.

[0109] In some embodiments, the preselected genomic regions are obtained by applying a minimum CpG density of 0.03 and a minimum CpG number of 5, which is exemplified in detail in Example 1. In addition, Table A include(s) preselected genomic regions (e.g., ridges) on chromosomes 21 and 22 that may be generated by applying a minimum CpG density of 0.03 and a minimum CpG number of 5.

[0110] Sample Extraction and Enrichment of DNA Molecules

[0111] Samples that are useful for methods herein can be virtually any nucleic acid sample. In illustrative embodiments, the nucleic acid sample is extracted or isolated from a subject or a population of subjects. In illustrative examples, DNA molecules are extracted or isolated from a sample such as a tissue sample (e.g., fresh-frozen tissue, FFPE tissue), and for illustrative embodiments herein, a liquid sample. Methods that are particularly useful in exemplary embodiments include methods for isolating circulating free DNA (cfDNA) from a liquid sample, and in illustrative embodiments from a blood, serum, urine, vitreous, sputum, saliva, tears, perspiration, ascites fluid, feces, bile, lymph, cervical mucus, or semen sample, and in further illustrative embodiments, a plasma sample.

[0112] DNA methylation biomarkers are increasingly being utilized in the development of novel assays for use in, for example, cancer and other disease detection, women’s health, organ health, and veterinary health. For example, methylation profiling is being used in non- invasive prenatal testing (NIPT) for monitoring of placental and fetal epigenomic changes, pre- symptomatic detection of preterm birth, preeclampsia, placental insufficiency, and fetal growth restriction. Methylation markers can also be indicative of congenital diseases of the fetus. Furthermore, methylation biomarkers are also used for monitoring of organ health such as predicting and monitoring organ rejection in transplant patients, monitoring of immune changes in transplant rejection, and for monitoring of organ health in high risk or predisposed individuals. In addition, methylation biomarkers can be used to determine biological age or detect age-related diseases. Accordingly, the sample described herein may be a maternal sample, a fetal sample, a sample from a transplant patient, or a sample from a subject having, had, suspected of having, or at risk of having a certain phenotype or a certain disease (e.g., cancer).

[0113] In certain illustrative embodiments, isolation of cfDNA from a liquid (e.g., blood or blood derivative sample such as a serum or plasma sample) can involve binding DNA molecules from a sample to a matrix and isolating the DNA molecules in the presence of a solvent. In some embodiments, the method further comprises incubating the biological sample comprising DNA molecules with a protease, prior to contacting the DNA molecules to the matrix. In some embodiments, the method can further include the steps of washing the matrixwith a wash buffer to remove impurities, and optionally, drying the matrix. Enriched nucleic acid samples can be eluted from the matrix with an elution buffer.

[0114] Other methods for nucleic acid isolation, for example cfDNA isolation, and optional enrichment of certain cfDNA can include ion exchange columns, or microfluidic devices, such as solid phase isolation, based on DNA capture by immobilized beads or functionalized surface. Additional methods include liquid phase isolation, utilizing an electric field, or chemical reagents, instead of a functionalized surface. In illustrative embodiments herein, isolation of cfDNA from a patient sample is performed using a DNA isolation kit (e.g., QIAamp Circulating Nucleic Acid kit (Qiagen)).

[0115] In some embodiments, cfDNA or their derivatives of certain sizes can be enriched before or after subjecting the cfDNA to methods herein. In some embodiments, size selection can be performed before the sequencing library preparation. In some embodiments, size selection can be performed after the sequencing library preparation and before sequencing. In some embodiments, size selection is performed on a sequencing-ready pool. Enriched cfDNA molecules can be, for example 50 to 1200 base pairs in length, 70 to 500 base pairs in length, 100 to 200 base pairs in length, or 130 to 170 base pairs in length. In some embodiments, the enriched cfDNA molecules are from 50 to 200 bp in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, or between 60 and 100 bp in length before the enriched cfDNA molecules, or derivatives thereof, are ligated to adaptors in methods herein. In some embodiments, the enriched cfDNA molecules are less than 150, 100, 90, 75, or 50 bp in length before they are ligated to adaptors. Such enrichment methods can be performed for example using the methods of WO2018156418 A1, Stray, et al. (incorporated herein by reference in its entirety).

[0116] In some embodiments, the sample is enriched for tumor DNA molecules, which are typically less than 160 bp and have a peak at about 145 bp in length. In illustrative embodiments, the enriched nucleic acid sample is circulating tumor DNA (ctDNA) or derivatives thereof. In such embodiments, the enriched nucleic acid is less than 160, 150, 145, 120, 100, 90,75, or 50 bp in length. In some embodiments, size selection is used to filter out cfDNA molecules carrying mutations derived from clonal hematopoiesis of indeterminate potential (CHIP) but not tumor-derived mutations, which are typically longer than ctDNA molecules, at about 165bp.

[0117] In some embodiments, the sample is enriched for fetal DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is fetal circulating free DNA or derivatives thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of fetal cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170 bp, from 160 to 190 bp, or from 170 bp to 220 bp.

[0118] In some embodiments, the sample is enriched for transplant donor DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is transplant donor circulating free DNA or derivatives thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of transplant donor cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170bp, from 160 bp to 190 bp, or from 170 bp to 220 bp.

[0119] Library Preparation

[0120] In some embodiments, the method further comprises fragmenting the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes. In some embodiments, the method further comprises appending adaptors to the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes. In some embodiments, the adaptors are Y-adaptors. In some embodiments, the adaptors comprise a universal priming site, a molecular barcode, and / or methylated cytosines.

[0121] Typically, methods herein include a step of appending nucleic acid adaptors to extracted DNA molecules or to nucleic acid derivatives generated therefrom. For example,adaptors may be appended on to the DNA molecules by ligation. The extracted DNA molecules in illustrative embodiments are extracted from a sample of a subject. In some embodiments, appending nucleic acid adaptors is preformed after the extracted DNA molecules are fragmented to form fragmented DNA molecules. Typically, methods include exposing the extracted DNA molecules to one or more polymerases or kinases, such as Klenow Large Fragment Polymerase and T4 polynucleotide kinase (PNK), as well as a ligase, such as T4 ligase. In some embodiments, the extracted DNA molecules or the fragmented DNA molecules are exposed to one or more polymerases and / or kinases to generate the nucleic acid derivatives. In some embodiments, the method further comprises appending adaptors to the nucleic acid derivatives generated therefrom. In some embodiments, the extracted DNA molecules are not fragmented prior to appending nucleic acid adaptors thereto. In some embodiments, the extracted DNA molecules are cellular DNA molecules extracted from a tissue sample (e.g., fresh-frozen tissue, FFPE tissue). In some embodiments, the extracted DNA molecules are cfDNA molecules.

[0122] In some embodiments, adaptors are ligated to the extracted DNA molecules. Before such ligation, the extracted DNA molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adaptor ligation. For example, extracted DNA molecules can be blunt ended, nucleotides can be added to the extracted DNA molecules or blunted-ended derivative therefrom, and / or phosphate moieties can be added or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, the extracted DNA molecules may be blunt ended, and then a single adenosine base can be added to the 3’ end. Prior to ligation the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation the 3’ adenosine of the DNA fragments and the complementary 3’ thymidine overhang of an adaptor can enhance ligation efficiency. In some embodiments, adaptor ligation is performed using a T4 ligase.

[0123] In some embodiments, adaptors containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adaptors are Y adaptors, for example in methods involving NGS sequencing. In some embodiments, the adaptors eachcomprises a universal priming site. The adaptors may or may not include methylated cytosine residues.

[0124] In some embodiments, the adaptor may comprise a sample barcode. Thus, multiple samples can be analyzed in the same sequencing reaction. The sample barcode can be used to process data according to the sample from which the data was generated.

[0125] In some embodiments, the adaptor may comprise a molecular barcode or molecular index tag (MIT). In some embodiments, the number of adaptors having different MITs is between 10 to 1,000, and wherein the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction is at least 1,000:1. The number of different MITs in the ligation reaction, in certain embodiments, ranges from10 to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600, 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are at least 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different MITs in the ligation reaction. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction is at least 10,000:1. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction ranges from 50,000: 1 to 50:1, from 25,000: 1 to 100:1, from 10,000:1 to 100:1, from 10:000:1 to 8,000:1 to 500:1, from 5,000:1 to 200:1, from 10,000:1 to 50:1. In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000 different MITs to each one template nucleic acid or cfDNA molecules.

[0126] Treatment with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines

[0127] Any method suitable for detecting methylated cytosines at a single-base resolution may be used with the target panel described herein. In some embodiments, the extracted DNA or derivative thereof is treated with a deaminating agent that converts (directly or indirectly throughone or more other agents) unmethylated but not methylated cytosines to uracils before selective enrichment with a panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a deaminating agent (directly or indirectly through one or more other agents) that converts unmethylated but not methylated cytosines to uracils after selective enrichment with a panel of oligonucleotide probes. In some embodiments, the deaminating agent (e.g., a chemical reagent such as sodium bisulfite) directly converts unmethylated cytosines to uracils. In some embodiments, the deaminating agent converts unmethylated cytosines to uracils in combination with one or more other agents – for example, in order to convert unmethylated cytosines to uracils, the extracted DNA or derivative thereof may be first treated with an oxidizing agent (e.g., TET or its catalytic domain) that oxidizes one or more forms of methylated cytosines (e.g., 5hmC or 5mC), followed by treatment with a deaminase (e.g., APOBEC or its catalytic domain).

[0128] In some embodiments, the deaminating agent comprises a chemical reagent (e.g., sodium bisulfite). In some embodiments, the deaminating agent comprises a deaminase (e.g., APOBEC) or its catalytic domain. In some embodiments, the deaminating agent comprises APOBEC3A (A3A) or its catalytic domain.

[0129] In some embodiments, the extracted DNA or derivative thereof is treated with a reducing agent that converts certain forms of modified cytosines, but not unmodified cytosines, into dihydrouridine (DHU) before the contacting with the panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a reducing agent that converts certain forms of modified cytosines, but not unmodified cytosines, into DHU after the contacting with the panel of oligonucleotide probes. In some embodiments, the reducing agent comprises pyridine borane.

[0130] In some embodiments, the extracted DNA or derivative thereof is treated with an oxidizing agent that oxidizes one or more forms of methylated cytosines before being treated with a deaminating agent or a reducing agent. In some embodiments, the oxidizing agent oxidizes 5hmC and 5mC to 5caC (5-carboxylcytosine). In some embodiments, the oxidizingagent selectively oxidizes 5hmC, but not 5mC, to 5fC (5-formylcytosine). In some embodiments, the oxidizing agent comprises an enzyme of the ten-eleven translocation (TET) families or its catalytic domain. In some embodiments, the oxidizing agent comprises TET1 or its catalytic domain. In some embodiments, the oxidizing agent comprises TET2 or its catalytic domain. In some embodiments, the oxidizing agent comprises TET3 or its catalytic domain. In some embodiments, the oxidizing agent comprises potassium perruthenate (KRuO4). In some embodiments, the oxidizing agent comprises potassium ruthenate (K2RuO4).

[0131] In some embodiments, the extracted DNA or derivative thereof is treated with a glycosylation agent that glycosylates 5hmC to glucosyl-5hmC (5ghmC) to protect it from oxidation, deamination, or reduction before being treated with an oxidizing agent, a deaminating agent, or a reducing agent. In some embodiments, the glycosylation agent is a glycosyltransferase. In some embodiments, the glycosyltransferase is a glucosyltransferase. In some embodiments, the glucosyltransferase is β-glucosyltransferase (β-GT). In some embodiments, the glucosyltransferase is β-GT of T4 phase.

[0132] In some embodiments, the extracted DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, before selective enrichment with a panel of oligonucleotide probes. In some embodiments, the selected DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5- hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, after selective enrichment with a panel of oligonucleotide probes. In some embodiments, other methylation detection methods such as methylated DNA immunoprecipitation (MeDIP) or direct detection by long read sequencers may also be used.

[0133] Conversion-based methylation detection

[0134] In some embodiments, the methods described herein include treating the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with one or more chemical or enzymatic agents or a combination of chemical and enzymatic agents (e.g. deaminating agents, reducing agents, oxidizing agents, glycosylation agents) that allow discrimination between methylated and unmethylated cytosines, or between 5mC and 5hmC. In some embodiments, a portion of a library of DNA molecules or their derivatives is treated with one or more such chemical and / or enzymatic agents. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more such chemical and / or enzymatic agents. In some embodiments, the entire library of DNA molecules or their derivatives is treated with one or more such chemical or enzymatic agents.

[0135] In some embodiments, the agent comprises a deaminating agent. In some embodiments, the deaminating agent comprises an enzyme such as a deaminase or its catalytic domain. In some embodiments, the deaminase may be APOBEC or its catalytic domain. In some embodiments, the deaminase may be APOBEC3A (A3A) or its catalytic domain, which deaminates unmethylated C, 5mC, and 5hmC, but not 5ghmC, 5fC, or 5caC. In some embodiments, the method described herein involves treating the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with TET2 / oxidation enhancer, which converts 5mC and 5hmC to 5caC and protects them from deamination by APOBEC, followed by treatment with APOBEC to deaminate the unmethylated cytosines to uracils. In some embodiments, the DNA sample is treated with a glycosylation agent (e.g. T4 β- glucosyltransferase) to glycosylates 5hmC to 5ghmC and protects it from deamination by APOBEC.

[0136] In some embodiments, the deaminating agent comprises a bisulfite reagent. Bisulfite reagents, such as sodium bisulfite, convert unmethylated cytosine to uracil and leave methylated cytosine (including 5mC and 5hmC) unchanged. Therefore, after bisulfite treatment, 5mC and 5hmC in the DNA remains as cytosine and unmodified or unmethylated cytosine will be changed to uracil. In some embodiments, the bisulfite reagent may be sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassiummetabisulfite, ammonium metabisulfite, magnesium metabisulfite, and equivalents. In some embodiments, a combination of one or more chemical reagents and / or one or more enzymes can be used to discriminate between various forms methylated cytosines (e.g., mC or hmC) and unmethylated cytosines. For example, in some embodiments, the DNA sample is treated with an oxidizing agent (e.g. KRuO4 or K2RuO4) that selectively oxidizes 5hmC, but not 5mC, to 5fC before bisulfite treatment, which converts unmethylated cytosine (including 5fC converted from 5hmC) to uracil but leaving 5mC unchanged. In some embodiments, a glycosylation agent (e.g. T4 β-glucosyltransferase or its catalytic domain) is used to glycosylate 5hmC to 5ghmC and an oxidizing agent (e.g. a TET enzyme or its catalytic domain) is used to convert 5mC to 5caC before bisulfite treatment, which converts unmethylated cytosine (including 5caC converted from 5mC) to uracil but leaving methylated cytosine (including 5ghmC converted from 5hmC) unchanged.

[0137] The bisulfite treatment can be performed by commercial kits such as the Imprint DNA Modification Kit (Sigma), EZ DNA Methylation- DirectTM Kit ( ZYMO), and the EZ DNA Methylation-Gold Kit (ZYMO). After DNA bisulfite conversion, single stranded DNA is captured, desulphonated and cleaned. The bisulfite-treated DNA can be captured by purification columns or magnetic beads and eluted. Bisulfite-treated single stranded DNA can be converted into dsDNA through an enzyme-catalyzed DNA strand synthesis with appropriate primers and polymerase. The polymerase will recognize the uracil in the ssDNA template as thymine and add an adenine to the complimentary strand. Further polymerase extension on the complimentary strand will result in replication of the original bisulfite treated ssDNA template, substituting uracil with thymine. Identification of cytosine to thymine conversion and guanine to adenine conversions (complementary strand) through comparing to the reference genome, will determine all unmodified cytosines, while the remaining cytosines are considered to be methylated.

[0138] In some embodiments, the agent comprises a reducing agent (e.g. pyridine borane) that selectively reduces 5fC and 5caC to dihydrouridine (DHU), which, like uracil, is converted to T base following PCR amplification. In some embodiments, the method described herein involves treating the extracted DNA molecules or their derivatives, or the selected DNAmolecules or their derivatives, with TET2 / oxidation enhancer, which converts 5mC and 5hmC to 5caC, followed by treatment with borane, which reduces 5fC and 5caC (including 5caC converted from 5mC and 5hmC) to DHU but leaves unmodified cytosine (dC) unchanged. In some embodiments, the DNA sample is treated with an oxidizing agent (e.g. KRuO4or K2RuO4) that selectively oxidizes 5hmC, but not 5mC, to 5fC, followed by treatment with borane, which reduces 5fC (including 5fC converted from 5hmC) and 5caC to DHU but leaves dC and 5mC unchanged. In some embodiments, the DNA sample is first treated with a glycosylation agent (e.g. T4 β-glucosyltransferase) and an oxidizing agent (e.g. TET) such that 5hmC is glycosylated to 5ghmC and only 5mC is oxidized and converted to 5caC. The subsequent treatment with borane reduces 5fC and 5caC (including 5caC converted from 5mC) to DHU but leaves dC and 5ghmC unchanged.

[0139] Other conversion-based methyl detection methods, including EM-seq, oxidative bisulfite sequencing (oxBS-seq), TET-assisted bisulfite sequencing (TAB-seq), TET-assisted pyridine borane sequencing (TAPS), TAPS with T4-βGT protection sequencing (TAPSβ), chemical-assisted pyridine borane sequencing (CAPS), and many variants of these that can discriminate between C, mC, and / or hmC. These methods can be performed with commercial kits such as NEBNext® Enzymatic Methyl-seq (EM-seq™) (New England Biolabs, Inc.), the 5hmC TAB-Seq Kit (WiseGene), and the EpiTect® Bisulfite Kit (Qiagen), for example.

[0140] Restriction enzyme-based methylation detection

[0141] In some embodiments, the methods described herein include treating the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with one or more restriction enzymes. In some embodiments, the one or more restriction enzymes are one or more methylation sensitive restriction enzymes (MSREs). In some embodiments, a portion of a library of DNA molecules is treated with one or more MSREs. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more MSREs. In some embodiments, the entire library of DNA molecules is treated with one or more MSREs.

[0142] MSREs are useful in analyzing the methylation status of cytosine residues in CpG sites. As the name implies, these enzymes are not able to cleave their palindromic target sites when the cytosine residues are methylated. The size of the MSRE targets sites range from 4 bp up to 76bp, but typically are in the 4 to 8 bp range.

[0143] In one aspect, methods herein typically include contacting the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, with one or more MSREs. MSREs selectively cleave a sample nucleic acid when the MSRE recognition site is unmethylated, but not when the MSRE recognition site is methylated. An “isoschizomer” of an MSRE is a restriction enzyme that recognizes the same recognition site as a methylation sensitive restriction enzyme but cleaves both methylated CGs and unmethylated CGs. Isoschizomer of the selected MSREs may be used in control reactions. Non-limiting examples of methylation sensitive restriction enzyme include, and thus in some embodiments, the one or more MSREs can include, AatII, Acc65I, AccI, AciI, AclI, AfeI, AgeI, AgeI-HF®, AhdI, AleI- v2, ApaI, ApaLI ApeKI, AscI, AsiSI, AvaI, AvaII, BaeI, BanI, BbvCI, BceAI,, BcgI, BcoDI, BfuAI, BglI, BmgBI, BsaAI, BsaBI, BsaHI, BsaI-HF®v2, BseYI, BsiE, BsiWI, BsiWI-HF®, BslI, BsmAI, BsmBI-v2, BsmFI, BspDI, BspEI, BsrBI, BsrFI-v2, BssHII, BstAPI, BstBI, BstUI, BstZ17I-HF®, BtgZI, Cac8I, ClaI, DpnI, DraIII-HF®, DrdI, EaeI, EagI-HF®, EarI, EciI, Eco53kI, EcoRI, EcoRI-HF®,EcoRV, EcoRV-HF®, Esp3I, FauI, Fnu4HI, FokI, FseI, FspI, HaeII, HgaI, HhaI, HinP1I, HincII, HinfI, HpaI, HpaII, Hpy166II, Hpy188III, Hpy99I, HpyAV, HpyCH4IV, KasI, MboI, MluI, MluI-HF®, MmeI, MspA1I, MwoI, NaeI, NarI, NciI, NgoMIV, NheI-HF®, NlaIV, NotI, NotI-HF®, NruI, NruI-HF®, Nt.BbvCI, Nt.BsmAI, Nt.CviPII, PaeR7I, PaqCI, PleI, PluTI, PmeI, PmlI, PshAI, PspOMI, PspXI, PvuI, PvuI-HF®, RsaI, RsrII, SacI- HF®, SacII, SalI, SalI-HF®, Sau3AI, Sau96I, ScrFI, SfaNI, SfiI, SfoI, SgrAI, SmaI, SnaBI, SrfI, StyD4I, TfiI, TseI, TspMI, XhoI, XmaI, and / or ZraI. In illustrative embodiments, the one or more MSREs can include HhaI, HpaII, BstUI, and / or HpyCH4IV.

[0144] In some embodiments, the one or more restriction enzymes are one or more methylation dependent restriction enzymes (MDREs). MDREs selectively cleave a sample nucleic acid when one or more nucleotides in the MDRE recognition site is methylated. In someembodiments, one or more MDREs can be used in combination with one or more MSREs. In some embodiments, the one or more MDREs can be AbaSI, AoxI, BisI, BlsI, DpnI, FspEI, GlaI, GluI, KroI, LpnPI, MalI, MspJI, MteI, PcsI, PkrI, or SgeI.

[0145] The MSRE or MDRE can be selected based on differentially methylated CpG sites in target DNA molecules, such as tumors or ctDNA from specific cancer targets, or a diverse spectrum of tumors. Further criteria for selection may include low background methylation in normal tissues, size and number of cleavage fragments, number of base pairs of recognition sequence, whether the cleavage results in blunt vs. tailed end fragments, and whether the enzymes have the same or similar reaction conditions such that the contacting step can be done under the same set of conditions and / or in a single reaction.

[0146] In some embodiments, more than one, a plurality, or a set of MSREs (and / or MDREs) can be used to contact the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives. The number of selected MSREs in certain embodiments is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; from 2 to 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; or from 5 to 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 10, 1 to 5, from 2 to 5, from 3 to 5, from 4 to 5, 3 to 7, from 5 to 10, from 6 to 10, or from 7 to 10. Thus, criteria such as target site, target site methylation status, number of base pairs in tailed end fragments, reaction conditions including but not limited to buffer conditions, incubation time and temperature for optimal activity, as well as deactivation time and temperature for each MSRE can be used for selecting the plurality or set of MSREs and / or MDREs to include in the method. As activity measured by units is specific to each enzyme, the target DNA molecules can be contacted with from 1 to 5 Units (U) of each MSREs in the reaction sample. In some embodiments, the extracted DNA molecules or their derivatives, or the selected DNA molecules or their derivatives, are contacted with from 1 to 5 U, from 1.5 to 4 U, from 2 to 3 U, from 2.5 to 4 U, or from 3 to 5 U of each MSRE. In some embodiments, one or more of the MSRE is selected from HpaII, SalI, BbeI, NotI, SmaI, XmaI, MboI, BstUI, BstBI, ClaI, MluI, NaeI, NarI, PvuI, SacII, HpyCH41V, HhaI, and combinationsthereof. In exemplary embodiments, the one or more MSREs is selected from one or more of HpaII, HhaI, HpyCH41V, and BstU1. In some embodiments, the one or more MSREs are selected from one or more of HpaII, HhaI, HpyCH41V, or BstU1. In some embodiments of the methods as described herein, the contacting comprises contacting with two or more MSREs. In some embodiments, the contacting comprises contacting with three or more MSREs. In some embodiments, the contacting comprises contacting with four or more MSREs. In some embodiments, the contacting comprises contacting with the two or more MSREs in a single reaction.

[0147] In some embodiments, MSREs and / or MDREs are included that are able to cleave the MSRE sites and / or MDRE sites in the same cleavage buffer conditions. In some embodiments, the cleavage buffer conditions include 5-500 mM potassium acetate, for example 25-100 mM potassium acetate. In some embodiments, the cleavage buffer conditions include 2- 200 mM Tris-acetate, for example 10-40 mM Tris-acetate. In some embodiments, the cleavage buffer conditions include 1-100 mM magnesium acetate, for example 5-20 mM magnesium acetate. In some embodiments, the cleavage buffer conditions include 10-1000 μg / ml recombinant albumin, for example 50-200 μg / ml recombinant albumin. In some embodiments, the pH is between 6.9 and 8.9 at 25 °C, for example, between 7.4 and 8.4, 7.5 and 8.3, 7.6 and 8.2, 7.7 and 8.1, or 7.8 and 8, or about 7.9 at 25 °C. In some embodiments, the cleavage buffer conditions include 1-100 mM bis-tris-propane-HCl, for example 5-20 mM bis-tris-propane-HCl. In some embodiments, the cleavage buffer conditions include 1-100 mM MgCl2, for example 5- 20 mM MgCl2. In some embodiments, the pH is between 6 and 8 at 25 °C, for example, between 6.5 and 7.5, 6.6 and 7.4, 6.7 and 7.3, 6.8 and 7.2, or 6.9 and 7.1, or about 7.0 at 25 °C. In some embodiments, the cleavage buffer conditions include 5-500 mM NaCl, for example 25-100 mM NaCl. In some embodiments, the cleavage buffer conditions include 1-100 mM Tris-HCl, for example 5-20 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 10- 1000 mM NaCl, for example 50-200 mM NaCl. In some embodiments, the cleavage buffer conditions include 5-500 mM Tris-HCl, for example 25-100 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 25-100 mM potassium acetate, 10-40 mMTris-acetate, 5-20 mM magnesium acetate, and 50-200 μg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 5- 20 mM bis-tris-propane-HCl, 5-20 mM MgCl2, and 50-200 μg / ml recombinant albumin, and the pH is between 6.7 and 7.3at 25 °C. In some embodiments, the cleavage buffer conditions include 25-100 mM NaCl, 5-20 mM Tris-HCl, 5-20 mM MgCl2, and 50-200 μg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 50-200 mM NaCl 25-100 mM Tris-HCl, 5-20 mM MgCl2, and 50-200 μg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C.

[0148] In some embodiments, the method further comprises amplifying the converted DNA or derivative thereof after treatment with an agent or a combination of agents that discriminates between methylated and unmethylated cytosine. In some embodiments, the method further comprises amplifying the extracted DNA or derivative thereof before selective enrichment with the panel of oligonucleotide probes. In some embodiments, the method further comprises amplifying the selected DNA or derivative thereof after selective enrichment with the panel of oligonucleotide probes.

[0149] In some embodiments, the amplifying adds a sample barcode, and wherein a plurality of libraries of the selected DNA or derivative thereof each produced from a different sample are sequenced together in one sequencing lane. In some embodiments, the method further comprises performing size selection before or after the amplifying step.

[0150] Panel of oligonucleotide probes to generate selected DNA

[0151] In some embodiments, a panel of oligonucleotide probes are designed to hybridize to a plurality of preselected genomic regions, thereby generating selected DNA. In some embodiments, the panel of oligonucleotide probes are hybrid capture probes. In some embodiments, the panel of oligonucleotide probes each comprises a probe section covalently linked to a primer section.

[0152] In some embodiments, the preselected genomic regions may be selectively enriched either before or after treating the extracted DNA with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines. For example, in some embodiments, after the DNA is extracted from a sample of the subject, the DNA may be selectively enriched for the preselected genomic regions using oligonucleotide probes specific for the preselected genomic regions. In some embodiments, after treating the extracted DNA with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines, the DNA library may be selectively enriched for the preselected genomic regions using oligonucleotide probes specific for the preselected genomic regions.

[0153] Hybrid Capture

[0154] In some embodiments, the selective enrichment technique can involve fragment capture by hybridization (i.e., hybrid capture). Although any hybrid capture method can be used to perform methods herein that include a selective enrichment step, in some embodiments, a method of the present disclosure may involve using any of the hybrid capture methods disclosed herein to selectively enrich DNAA. In some embodiments described herein, such selective enrichment steps can be performed after DNA molecules are extracted. In some embodiments described herein, such selective enrichment steps can be performed following the addition of adaptors to the extracted DNA molecules. In some embodiments described herein, such selective enrichment steps can follow treatment of the extracted DNA with an agent that discriminates between methylated and unmethylated cytosines.

[0155] In capture by hybridization, hybrid capture oligonucleotide probes complementary to one or both strands of specific target DNA sequences, or DNA derived therefrom in a sample, are utilized, i.e., the probes may be strand specific. In some embodiments, the probes may be complementary to the target sequence in either a 100% or 0% methylated state. Accordingly, each strand specific probe may have two designs, assuming both a completely methylated and a completely unmethylated status of each target DNA sequence, yielding a total of four unique probes per target. The specific target DNA sequence in illustrativeembodiments overlaps with or is found within a target region of a sample DNA molecule such as a cfDNA. Thus, hybrid capture probes when used in methods herein can be designed to bind to a DNA molecule that contains at least one target region or a portion thereof. In some embodiments, the hybrid capture probes can be designed to bind to a target DNA sequence within or overlapping a target region, which in illustrative embodiments contains one or more methylated nucleic acids. In other examples, the hybrid capture probes can be designed to bind to a common region that is flanking but not overlapping the target region and that can be a common region that was added to some, most, almost all or all of the DNA in a sample, or added to all amplicons using a common sequence on at least one primer of a primer pair. In illustrative embodiments, a hybrid capture probe or set thereof, are designed to bind to a target DNA sequence within target region, or set of target regions, respectively.

[0156] Hybrid capture probes may be added to a prepared sample and hybridized through a denature-reannealing process to form duplexes of exogenous-endogenous fragments (e.g., hybrid capture probes bound to sample DNA molecules, or DNA derived therefrom). These duplexes may then be physically separated from the sample by various means. In some embodiments, once the hybrid capture probes are removed, the sample DNA molecules, or DNA derived therefrom can be amplified. Some ways to physically remove the hybrid capture probes are by covalently bonding the hybrid capture probes to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the hybrid capture probes is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in SURE SELECT (Agilent). Thus, hybrid capture probes, for example that bind to a target DNA sequence within or overlapping a target region of a DNA molecule obtained or derived from a sample, can be covalently attached to a biotin molecule, and after hybridization with sample DNA or DNA derived therefrom, a solid support with streptavidin affixed can be used to pull down the biotinylated hybrid capture probes, which are hybridized to DNA molecules obtained or derived from a sample that include a target region that includes the target DNA sequence recognized by the hybrid capture probes. Thus, in some embodiments, the hybrid capture probes areimmobilized, directly or indirectly to a solid support. In some embodiments, the hybrid capture probes include a binding partner, for example biotin.

[0157] In some embodiments of any of the aspects herein, the hybrid capture probes can be a part of a set of at least two hybrid capture probes. In some embodiments, the set includes at least one hybrid capture probe for each target region. In some embodiments, the set includes two or more hybrid capture probes for each target region.

[0158] In some embodiments of any of the aspects herein, the hybrid capture probes can have a length in the range of 30 bases to 170 bases, 30 bases to 160 bases, 30 bases to 150 bases, 30 bases to 140 bases, 30 bases to 130 bases, 30 bases to 120 bases, 30 bases to 110 bases, 30 bases to 100 bases, 30 bases to 90 bases, 30 bases to 80 bases, 30 bases to 70 bases, 30 bases to 60 bases, 30 bases to 50 bases, 40 bases to 160 bases, 40 bases to 150 bases, 40 bases to 140 bases, 40 bases to 130 bases, 40 bases to 120 bases, 40 bases to 110 bases, 40 bases to 100 bases, 40 bases to 90 bases, 40 bases to 80 bases, 40 bases to 70 bases, 40 bases to 60 bases, 50 bases to 150 bases, 50 bases to 140 bases, 50 bases to 130 bases, 50 bases to 120 bases, 50 bases to 110 bases, 50 bases to 100 bases, 50 bases to 90 bases, 50 bases to 80 bases, 50 bases to 70 bases, 60 bases to 140 bases, 60 bases to 130 bases, 60 bases to 120 bases, 60 bases to 110 bases, 60 bases to 100 bases, 60 bases to 90 bases, 60 bases to 80 bases, 70 bases to 130 bases, 70 bases to 120 bases, 70 bases to 110 bases, 70 bases to 100 bases, 70 bases to 90 bases, 80 bases to 120 bases, 80 bases to 110 bases, 80 bases to 100 bases, 90 bases to 120 bases, 90 bases to 110 bases, 100 bases to 165 bases, 100 bases to 150 bases, 100 bases to 140 bases, 100 bases to 130 bases, 100 bases to 120 bases, 110 bases to 150 bases, 110 bases to 140 bases, 110 bases to 130 bases, 120 bases to 150 bases, or 130 bases to 160 bases.

[0159] Probe-dependent primers

[0160] In some embodiments, the selective enrichment technique can involve probe- dependent primers. Probe-dependent primers (PDPs) have been disclosed (Pel, et al. “Rapid and highly-specific generation of targeted DNA sequencing libraries enabled by linking captureprobes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO 2017168332A1 “Linked duplex target capture”, which are hereby incorporated by reference in their entirety). Such embodiments can be considered LTC methods. Briefly, in an LTC method, PDPs are designed to incorporate non-extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. Typically, probes of PDPs can be between 30 to 70 nucleotides in length, and include or comprise a 3’ inverted dT base to inhibit polymerase extension. In some embodiments, probes are designed to cover the desired region with zero gap between forward and reverse probes. In some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60, between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the probes of a PDP pair comprises a sample index.

[0161] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a genomic region of interest on a sample DNA molecule to enrich nucleic acid molecules comprising the genomic region of interest or copies thereof. In some embodiments, at least one of the probe binding regions can include one or more CpG sites. In some embodiments, both of the probe binding sites of the probe binding regions can include one or more CpG sites.

[0162] Typically, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on the appended adaptor. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desiredmelting temperature. In illustrative embodiments, the primer is a low melting temperature universal primer complementary to a portion of the ligated adaptor.

[0163] The primer can be tailed or untailed depending on the specific requirements. In some embodiments, the universal primer comprises an A tail. In some embodiments, the universal primer is blunt ended. The length of the primers of the PDP can range from 5 to 40 nucleotides in length. In certain embodiments, the PDP primers are between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.

[0164] Typically, probe dependent primers comprise a linker between the probe and the primer. Probe and primer portions of the PDP are typically linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, the linker is an oligosaccharide. In some embodiments, the linker is a lipid. In some embodiments, the linker is a hydrocarbon. In some embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof.

[0165] Sequencing to determine methylation status

[0166] The methylation status of the original sample DNA can be determined by sequencing the selected DNA or its derivatives. In some embodiments, the sequencing is next- generation sequencing or high-throughput sequencing.

[0167] DNA sequencing techniques, particularly high throughput next-generation sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed NOVASEQ (ILLUMINA), MISEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), GS FLEX+(ROCHE 454) etc., can be used for determining the sequences of the selected DNA or its derivatives to elucidate the methylation status of the original sample DNA. High throughput genetic sequencers are amenable to the use of barcoding (i.e., sample tagging with distinctivenucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. Methods as described herein that utilize NGS detection, in some embodiments can have an average depth of read of at least 0.1, 0.5, 1, 10, 50, 100, 200, 500, 1000, 2000, 2900, 3000, 3500, 4000, 5000, 10,000, 50,000, 75,000, 100,000, 130,000, 150,000, 175,000, or 200,000.

[0168] Methods herein can include analyzing data obtained from next-generation sequencing techniques. In some embodiments of methods herein, deaminated, or MSRE treated, methylation-transferred DNA molecules can be subjected to sequencing using next-generation sequencing techniques. For a skilled artisan, algorithm design tools are available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.

[0169] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows-Wheeler alignment software, Bwa mem function (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics.) on single end mode using pear merged reads to the hg19 genome. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.

[0170] Methods herein can include a background error model that can be constructed using normal, healthy, or non-diseased liquid samples, in illustrative embodiments, normal, healthy, or non-diseased plasma samples, which are sequenced on the same sequencing run to account for run-specific artifacts. In some embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal, healthy, or non-diseased liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g., plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10 are removed from the data analysis. For each base substitution of every genomic loci, the DOR weighted mean and standard deviation of the error can be calculated.

[0171] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200-nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotides is 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).

[0172] In some embodiments, the uniformity in DOR can be measured using standard methods such as, but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90th- 95th percentile. For this measurement, the loci are sorted in descending DOR order. In illustrative embodiments, a DOR distribution using the 90th-95th percentile contains 5 percent ofreads. The reads of all loci between the 90th percentile and 95th percentile can be counted and divided by the total reads for all loci.

[0173] In some embodiments, the magnitude of the DOR slope can be less than 0.005, 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90th-95th percentile can be between 0.2 and 9 percent, such as between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent, 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to 6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent, 3 to 7 percent, 3 to 6 percent, 0.2 to 1.0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent, 3 to 4 percent, 4 to 5 percent, 5 to 6 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, or 100,000 non-identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90th-95th percentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15,000 non-identical amplicons.

[0174] Discovery of methylation biomarkers

[0175] In some embodiments, the method described herein further comprises producing a set of libraries of the selected DNA from a set of samples from a population of subjects, and using the sequencing reads of the set of libraries of the selected DNA to identify preselected genomic regions of which the co-methylation patterns of the CpG sites are substantially consistent.

[0176] In some embodiments, the method described herein further comprises producing a first set of libraries of the selected DNA from a first set of samples from a population of diseased subjects, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-diseased subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of diseased subjects and the population of healthy or non-diseased subjects.

[0177] In some embodiments, the method described herein further comprises producing a first set of libraries of the selected DNA from a first set of samples from a population of subjects having cancer, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-cancerous subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of subjects having cancer and the population of healthy or non-cancerous subjects.

[0178] In some embodiments, the method described herein further comprises identifying a plurality of differentially methylated regions (DMRs) between the first set of samples and the second set of samples.

[0179] Aberrations in patterns of DNA methylation (e.g., aberrations in patterns of co- methylation of adjacent CpG loci) are an indicator of a certain phenotype, including a disease, disease type, and / or disease progression. In the case of cancer, for example, aberrations are associated with different aspects of cancer, from tumor initiation to cancer progression and metastasis. For example, between 0.2% to 80% of the genome may be differentially methylated in cancers, and methylation thus is uniquely suited both for cancer detection and for tissue of origin (TOO) prediction. Differentially methylated regions are ideal biomarkers for early cancer detection due to their frequency and homogeneity across patients and subpopulations of patients. Likewise, differences in methylation patterns between maternal and fetal DNA, for example, canbe used in various non-invasive prenatal testing applications. In addition, cell-specific methylation changes can be used to detect, monitor, or predict transplant rejection.

[0180] Methylation status of CpG loci in a genomic region proves to be spatially correlated, giving rise to a host of methods that leverage this pattern to increase the effective signal-to-noise ratio (SNR) in determining cell-type of origin and early diagnostics (e.g., cancer diagnostics), especially considering the limited depth of whole genome bisulfite sequencing (WGBS) datasets. WGBS is a next-generation sequencing technology used to determine DNA methylation status of single cytosines, and involves treating DNA with sodium bisulfite before high-throughput DNA sequencing. The treatment with sodium bisulfite converts unmethylated cytosines into uracil, enabling methylation detection by helping distinguish methylated cytosines, which resist bisulfite treatment, from uracil. Other deamination agents (e.g. APOBEC) or reducing agents (e.g. borane), alone or in combination with an oxidizing agent (e.g. TET, potassium ruthenate) and / or a glycosylation agent (e.g. T4 β-glucosyltransferase), that specifically discriminate between methylated and unmethylated cytosines and / or different forms of methylated cytosines may also be used. During polymerase chain reaction (PCR) amplification, the uracils or dihydrouridines are converted into thymines, and unconverted cytosine forms are recognized as cytosines. Locations of methylated cytosines can be identified by comparison of, for example, a bisulfite-treated DNA sequence with an original DNA sequence. Non-conversion based methylation detection methods (e.g. MSRE / MDRE, MBD, MeDIP, or direct detection by long read sequencers) may also be used.

[0181] High local methylation correlation means that multiple CpGs can be integrated into methylation haplotypes which strongly increase SNR, in particular at low circulating tumor DNA (ctDNA) fractions. The rationale for methylation haplotypes is similar to phased single- nucleotide variants (SNVs), with the difference that sequencing errors for bisulfite sequencing are approximately two orders of magnitude higher (due to conversion and protection errors), and hence it becomes useful to leverage adjacent CpGs.

[0182] Concordant hypo-methylation of consecutive CpG loci within large genomic regions has been used in cancer diagnosis and is associated with cancer initiation and chromatin and genomic instability. Similarly, hyper-methylation of CpG loci in the promoter region of tumor suppressor genes occurs in tumor cells which is associated with gene suppression and tumor cell proliferation. These observations support the idea of considering concomitant changes in CpG methylation patterns over an entire genomic region as a single biomarker in cancer diagnosis and cell-type of origin problems.

[0183] Differentially methylated regions (DMRs) are genomic regions, typically with multiple adjacent CpG sites, that have different DNA methylation status among different biological samples, such as different cells / tissues within the same individual, the same cell / tissue at different times, different cells / tissues from different individuals, or different alleles in the same cell.

[0184] In some embodiments, differentially methylated genomic patterns are defined between patients with colorectal cancer (CRC) and healthy or non-CRC individuals by utilizing patterns of co-methylation of adjacent CpG loci. Example embodiments seek to identify and characterize CRC and advanced adenomas (AA) that can develop into CRC through CRC / AA- specific MBM via identification of tumor-specific differentially methylated regions from a large set of wide genomic regions, further discussed below. To achieve a set of features which robustly differentiate tumor from normal samples, for example, example embodiments use read- level methylation information to jointly model the methylation status of different CpGs in each genomic window. Joint modeling of methylation patterns of consecutive CpG loci leverages read-level information, which has the statistical power to distinguish between, for example, tumor-derived DNA and normal fragments even in presence of rare amounts of tumor fraction in low-coverage. Modeling helps identify new biomarkers that can be used in panels with target- specific primers or probes that increase detection sensitivity and specificity for cancer or other diseases. The disclosed approach can provide a set of biomarkers which can be used to detect a phenotype or a disease (for example, in example embodiments, CRC and AA), and which can be sub-selected to represent samples of different ages, genders, ethnicities etc. The methodsdisclosed herein can also be used to identify and characterize biomarkers suitable for detecting other cancers or phenotypes, or for differentiating cancer types, origins, or states.

[0185] To allow efficient manipulation of a large set of sequencing data, such as whole genome methylome data, various embodiments disclosed herein incorporate and use a novel file format that improves upon and overcomes the shortcomings of other available file formats. For example, regular bam files are too large and typically contain more information than necessary. Other available binary formats (such as pat file from WGBSTools) do not faithfully conserve fragment level information and are not portable; that is, in order to read, annotate, or otherwise manipulate these files, one needs to utilize specific toolbox or libraries, which may require developing new interfaces for manipulating these files. The file format embodiments disclosed herein allow preservation of fragment level information, including mapping quality information, by representing each fragment on a separate line and scoring each fragment based on the phred score of the CpGs in the fragment and the MAPQ score. The disclosed file format embodiments accommodate efficient file storage without loss of relevant information for downstream processes for all DNA methylation sequencing data, and allow archiving of all fastq or bam files on any cloud platform, e.g. DNANexus, Google Cloud Platform or Amazon AWS, while directly providing convenient input to downstream analyses. In various embodiments, with the disclosed file format, files can be read, annotated, or otherwise manipulated using commonly used bedtools and basic table editing software. In some embodiments, standard visualizations for read-level representation of CpG haplotype block methylation status are implemented.

[0186] To identify tightly co-methylated blocks across the genome, various embodiments preliminarily segment the whole genome into smaller regions based on biological priors and the sequencing technology and a chosen processing pipeline before any further analysis or modeling of the data. There are at least several main benefits to this preliminary segmentation. First, the preliminary segmentation drastically reduces the computational complexity, which in turn results in reduced computational cost and time. Second, the preliminary segmentation guides the biomarker discovery towards more valid segments which satisfy the biological and sequencingtechnology priors. Third, the preliminary segmentation reduces the number of probes or primers needed to cover the designed panel by removing parts of the genome which are less informative.

[0187] In example embodiments, the preliminary segmentation can be based on certain selected criteria. For example, as a first criterion, the preliminary segmentation can employ a threshold based on a certain primer or probe length, such as 75 bp to 150 bp (e.g., about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 105 bp, about 110 bp, about 115 bp, about 120 bp, about 125 bp, about 130 bp, about 135 bp, about 140 bp, about 145 bp, or about 150 bp). As a second criterion, the preliminary segmentation can employ a threshold based on expected fragment lengths, such as between 60 bp to 140 bp (e.g., about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 105 bp, about 110 bp, about 115 bp, about 120 bp, about 125 bp, about 130 bp, about 135 bp, or about 140 bp). As a third criterion, the processivity (in bp) of the DNA (cytosine-5)-methyltransferase 1 (DNMT1) enzyme may be set from tens of base pairs to a few hundred base pairs.

[0188] In various embodiments, probabilistic modeling, such as hierarchical Bayesian modeling (e.g., fully variational Bayesian modeling), is used to further segment and define a narrower set of relevant genomic regions, such as DMRs that can potentially be used as biomarkers for disease or phenotype detection. In various embodiments, a hierarchical Bayesian model is designed and implemented to infer and estimate change points, which indicate the boundaries of regions where the methylation load signal is relatively constant or almost constant in a given region across all samples within a cohort. In various embodiments, hierarchical Bayesian modeling is used to fit a piecewise constant function to beta-binomially distributed methylation load signals. The modeling takes into consideration not only the beta (expected) values of the signals but also the variance of the beta values. In some embodiments, stochastic variational inference (SVI) is used. The variational probabilistic modeling methods disclosed herein have several advantages. For example, a generative model can be trained as close as possible to the biological explanations of the CpG methylation patterns. Publicly available information can be leveraged to construct priors, which can be injected at different levels of themodel. Hypothesis testing can be performed very naturally as these models can be fully investigated. The probability distribution of the unknown parameters (such as CpG loci methylation status) can be inferred, informed by uncertainty estimation. Once the probability distributions are inferred, the probability of two distributions being different can be easily estimated. In various embodiments, the disclosed approach improves and expands on prior approaches at least by: (i) considering each read’s mapping quality in joint modeling of CpG loci; (ii) inferring the methylation state of a tentative biomarker through modeling of the processivity of DNMT1 via jointly modeling adjacent CpG loci methylation patterns; and (iii) benchmarking different models of varying complexity.

[0189] In various embodiments, after regions where the methylation status is relatively constant or almost constant across all samples within a cohort are identified, regions where the methylation status are significantly different between samples from different cohorts are identified as biomarker candidates, and the identified regions are scored and / or prioritized based on their respective association with different clinical covariates to obtain a select set of the biomarker candidates for clinical use.

[0190] Referring to FIG.1, in various embodiments, a system 100 may include a computing system 110 (which may be or may include one or more computing devices, co- located or remote to each other), a sequencing system 160 (capable of, e.g., generating DNA sequencing data), an information system 170 (such as a management or clinical information system that may record and / or provide information regarding samples or patients), vessels 175 (e.g., for holding samples), sample processing system 180 (e.g., a system that may include robotic components 182 for preparing, moving, and / or processing samples), and imaging and / or therapy systems 190 (which may include e.g., treatment unit 192 and imager / sensors 194 for imaging patients using various imaging modalities, providing treatments such as radiotherapy, and / or performing other evaluation or therapeutic procedures based on what is learned from patient samples). The computing system 110 (e.g., one or more computing devices) may be used to control and / or exchange signals and / or data with sequencing system 160, sample processing system 180, information system 170, and / or imaging and / or therapy systems 190, directly (e.g.,through wireless and / or wired communication) or indirectly via another component of system 100 (e.g., via any combination of wireless and / or wired communication). In certain embodiments, computing system 110 may be used to control and / or exchange data or other signals with sequencing system 160, information system 170, sample processing system 180, and / or imaging and / or therapy systems 190. The computing system 110 may include one or more processors and one or more volatile and / or non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated.

[0191] The computing system 110 may include a controller 112 that is configured to exchange control signals with sequencing system 160, information system 170, sample processing system 180, and / or imaging and / or therapy systems 190, and / or any components thereof, allowing the computing system 110 to be used to control, for example, processing of samples, capture of images, acquisition of signals by sensors, positioning or repositioning of samples being tested or imaged, recording or obtaining other information, and / or running assays or tests.

[0192] A transceiver 114 allows the computing system 110 to exchange readings, control commands, and / or other data or signals, wirelessly or via wires, directly or indirectly via networking protocols, with, for example, other components of system 100. One or more user interfaces 116 allow the computing system 110 to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, biometric scanners, etc.) and provide outputs (e.g., via a display screen, audio speakers, light emitters, etc.) with users. The computing system 110 may additionally include one or more databases 118 for storing, for example, data acquired from one or more systems or devices, signals acquired via one or more sensors, sequencing data, biomarkers, images, etc. In some implementations, database 118 (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote (e.g., via “cloud computing”) and in communication with computing system 110, sequencing system 160, information system 170, sample processing system 180, imaging and / or therapy systems 190, and / or components thereof.

[0193] Imaging and / or therapy systems 190 may include a treatment unit 192, which may include any systems and / or devices employed to administer a treatment to a patient, such as a radiation therapy system. Imager and / or sensors 194 may be or may include, for example, any system or device that is involved in, for example, capturing images of patients and / or samples. Imagers may include detectors for visible light and / or light in any frequencies of interest, such as (but not limited to) the spectrum from infrared to ultraviolet. Imagers may employ any suitable optical components (e.g., lenses, mirrors, filters, beam splitters, prisms, diffusers, diffraction gratings, etc.), digital components (e.g., charge-coupled devices (CCDs), as well as an area for placement of samples, computing components (e.g., one or more processors, such as digital signal processors) to process and / or pre-process images, etc. Imagers may be, or may employ, any microscopes or microscopy systems (e.g., confocal microscopy), tools, and / or techniques that will provide the desired imaging data. The imager and / or sensors 194 may have the capability of receiving control signals from a computing device or system to, for example, initiate or cease image capture, and / or return images or other imaging data or status signals to the computing device or system (e.g., computing system 1110). Imaging and / or therapy systems 190 may include any tools for obtaining additional data about samples and / or patients (such as spectrometers, chromatographs, etc.). Sensors may be used to detect, for example, other aspects of samples and / or patients, such as temperature, humidity, location, etc.

[0194] Sample processing system 180 may include any components used to automate preparation of samples and running tests on samples. Robotics 182, for example, may include any combination of actuators, stepper motors, servomotors, control system, robot arms, end effectors like grippers and manipulators, etc. Robotics 182 may be, or may comprise, for example, an automated robotics system with process automation software. In various embodiments, robotics 182 may employ vessels 175, such as vials or plates, that can be moved and manipulated, for example, for positioning of samples to be evaluated using sequencing system 160 or components thereof. Sample processing system 180 may include components that may be used to maintain the integrity of samples during storage, transport, and testing, such as heaters, passive or active coolers, etc. Sensors may be used by sample processing system 180 to,for example, monitor, guide, and evaluate the progress of tests (e.g., by detecting temperature, fill level of vessels, or other states or conditions). Sequencing system 160 may include sequencing platforms or devices as well as various analytical tools used to process data from the sequencing platforms. Analytical tools may include software and / or hardware used to generate computer files with raw and / or processed sequencing data.

[0195] In various implementations, components of system 100 may be rearranged or integrated in other configurations. For example, computing system 110 (or components thereof) may be integrated with one or more of the sequencing system 160, sample processing system 180, and / or components thereof. The sequencing system 160, sample processing system 180, and / or components thereof may be directed to a vessel 175 in which a sample can be situated (e.g., so as to test or evaluate biological samples). In various embodiments, the vessel 175 may be movable (e.g., using any combination of motors, magnets, etc.) to allow for positioning and repositioning of samples (such as micro-adjustments for positioning of samples to be sequenced). It is also noted that not all components of system 100 are required to implement the disclosed approach, and in various embodiments, only a subset of the components of system 100 may be employed. For example, in various embodiments, computing system 110 may obtain and process data that was obtained via sequencing system 160 or another system that is or is not in direct communication with the computing system 110. In certain embodiments, data may be obtained for samples that were not processed in an automated fashion but rather manually by one or more users.

[0196] A data acquisition unit 124 may retrieve, acquire, or otherwise obtain various data, such as sequencing data. The data acquisition unit 124 may, for example, obtain data stored in database 118, sequencing system 160, information system 170, sample processing system 180, and / or imaging and / or therapy systems 190. An interaction unit 126 may interact (e.g., via user interfaces 116) with users (e.g., laboratory technicians, data scientists, clinicians, etc.) to obtain information or commands needed for system 100 or components thereof to function, or to start operations, repeat processes, and / or stop operations. In certain embodiments, the data acquisition unit 124 may obtain data from users via interaction unit 126.

[0197] A preprocessor 128 may perform any number of steps disclosed herein in preprocessing sequencing data. Machine learning platform 130 may be configured to train and update machine learning models, as further discussed herein. Machine learning platform 130 may, for example, employ certain machine learning techniques and algorithms to train and update predictive models. Machine learning platform 130 may include a training engine 132 which may, for example, fit models to sequencing data. The training engine 132 may, for example, obtain sequencing data from or via data acquisition unit 124, sequencing system 160 or components thereof, and / or information system 170. Inference engine 134 may apply models (e.g., models obtained from machine learning platform 130 or components thereof) to generate outputs as disclosed herein. Biomarker generator 140 may perform analyses on data from machine learning platform 130 to, for example, generate biomarkers to be used for various conditions or sub-populations.

[0198] A reporter 142 may generate reports that include, for example, information on sample sequencing (e.g., via sequencing system 160 or components thereof), imaging or administration of treatments (e.g., via imaging and / or therapy systems 190 or components thereof and / or via information system 170), model training (e.g., via machine learning platform 130), and biomarker generation (e.g., via biomarker generator 140). For example, reporter 142 may provide or identify trained and / or updated models, and / or biomarkers. In certain embodiments, reporter 142 may obtain information to be reported from, for example, database 118 and / or information system 170. Data that may be reported may be, for example, transmitted to another system or device (which may or may not be part of system 100), saved in non- transitory computer-readable storage media (e.g., database 118, information system 170, and / or elsewhere), and / or presented or otherwise provided to users (e.g., via a display device that is part of user interfaces 116).

[0199] Referring to FIG.2, an example process 200 is illustrated, according to various example embodiments. Various elements of process 200 may be implemented by or via system 100 or components thereof. Process 200 may begin at block 210 with obtaining DNA samplesfrom healthy subjects (i.e. subjects not exhibiting a phenotype of interest) and subjects exhibiting the phenotype (e.g., a disease of interest), and perform methylation sequencing of the DNA.

[0200] At block 220, process 200 includes constructing one or more files (e.g., one or more .cpg files) based on sequencing data obtained from block 210. In various embodiments, the input data is the output of any methylation mapping / alignment tool (e.g., Bismark), usually in a BAM / SAM (Binary Alignment Map / Sequence Alignment Map) format. CpG files capture the methylation status of all CpG sites within a mapped DNA fragment. Each line of a cpg file corresponds to a DNA fragment that is specified by chromosome, start and end coordinates and the CpG methylation information within that fragment is encoded via a novel encoding method. In some embodiments, given a set of CpG sites captured by a fragment, the methylation status of each site is encoded as the number of base pairs from the beginning of the fragment, followed by the methylation status. Given a fragment with coordinates chrZ:X-Y and CpG sites located on this fragment with coordinates chrZ:x_1, chrZ:x_2, chrZ:x_3, the encoding will be (x_1- X)s_1(x_2-x1)s_2(x_3-x_2)s_3 where s_i is the methylation status of the CpG site x_i, i.e. C for methylated, T for unmethylated and N for undetermined. For example: Consider a fragment with the chromosome, start and end coordinates as chr1:1-100, and CpG sites chr1:20, chr1:40, and chr1:60 captured by this fragment where the first site is methylated and the rest are unmethylated. In this case, the encoding will be 19C20T20T. This allows conservation of storage space and thus reduces the amount of computational resources required.

[0201] At block 230, process 200 includes preliminarily segmenting the data in the generated files. In this step, blocks of CpG sites where the maximum distance between adjacent CpG sites are less than a threshold are identified. In various embodiments, the threshold is generally or typically chosen based on biological priors (e.g. chromosome boundaries, cytobands, etc.) and sequencing technology (e.g. maximum read length). This step guides further inference models and helps ensure that discovered biomarkers are designable (for example, the designability of probes or primers used for enrichment of DNA fragments comprising the biomarkers) and make biological sense. This also reduces the computational time and cost of the inference by transforming a relatively complex regression problem into a numberof smaller ones. Table A include(s) ridges (chr 21, 22) that may be generated by this preliminary segmentation process according to various embodiments.

[0202] At block 240, process 200 includes generating candidate biomarkers using a suitable model. In various embodiments, regions resulting from block 230 are used as inputs to a segmentation model, which further segments the regions into smaller regions where CpGs are co-methylated within each cohort. For example, a region where CpG sites are strongly co- methylated are segmented into fewer segments compared to a region with the same number of CpG sites where the CpG sites are not as strongly co-methylated. In example embodiments, the segmentation model is a Bayesian segmentation model, where segmentation robustness is increased naturally via modeling methylation signal variance as well as the expected value.

[0203] At block 250, process 200 includes analyzing candidate biomarkers to obtain a select set of biomarkers. Analyzing the candidate biomarkers may include scoring the candidate biomarkers. Scores for the candidate biomarkers can be used to prioritize the candidate biomarkers, such that the candidate biomarkers can be ranked based on their scores (e.g., from highest scores to lowest scores). For example, candidate biomarkers with relatively higher scores, or having scores above a threshold value, or having scores falling within a first range of values, may be deemed to be relatively more likely to distinguish between healthy subjects (i.e. subjects not exhibiting a phenotype of interest) and subjects exhibiting the phenotype, and candidate biomarkers with relatively lower scores, or having scores below the threshold value, or having scores falling within a second range of values, may be deemed to be relatively less likely to distinguish between the healthy subjects and the subjects exhibiting the phenotype.

[0204] In various embodiments, each candidate biomarker is scored on the signal-to- noise ratio (SNR) with the aim of prioritizing regions where the background noise (e.g., healthy methylation signal) is minimal compared to the phenotype signal (e.g., the cancer signal). The candidate biomarkers can be prioritized based on a number of criteria, such as designability of probes or primers for the biomarkers. In some embodiments, biomarkers in repeat or low- complexity regions are deprioritized.

[0205] At block 260, process 200 includes using the ranked / prioritized / selected biomarkers from block 250 to develop a panel to detect the phenotype (e.g., colorectal cancer) in a patient or other subject. For example, a sample (e.g. a blood, plasma, serum, or urine sample containing cell-free DNA) from a patient who may have a phenotype can be analyzed as disclosed herein, with primers or probes designed based on the selected biomarkers and used to evaluate likelihood of having the phenotype. Preliminary Selection of Epigenetic Regions of Interest

[0206] In various embodiments, potentially informative regions of the genome are pre- selected for downstream analysis, thereby significantly reducing the amount of data that need to be stored, analyzed, and otherwise manipulated. Using whole genome methylation sequencing data, example embodiments identify methylation cluster regions with high CpG density and co- high methylation, followed by genome segmentation to find methylation haplotype blocks for consideration in methylation biomarker selection. This preliminary segmentation drastically reduces the computational complexity, which in turn results in reduction in computational cost and time. It also guides the biomarker discovery towards more valid segments that satisfy the biological and sequencing technology priors. In addition, it reduces the number of probes needed to cover the designed panel.

[0207] As an initial step in example embodiments, a set of preliminary segments may be constructed by constructing a reference methylome, according to various embodiments. In some embodiments, this step can be performed using a common public reference methylome. In other embodiments, the preliminary segments can be constructed using a set of individual methylomes from each sample in a study cohort. In some embodiments, the preliminary segments can be constructed using reference methylomes of healthy / non-diseased samples, without using methylome data of tumor samples.

[0208] Example embodiments identify tightly and potentially co-methylated blocks across the genome. In example embodiments, the preliminary segmentation can be based oncertain selected criteria. For example, as a first criterion, the preliminary segmentation can employ a threshold based on a certain primer or probe length, such as 75 bp to 150 bp (e.g., about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 105 bp, about 110 bp, about 115 bp, about 120 bp, about 125 bp, about 130 bp, about 135 bp, about 140 bp, about 145 bp, or about 150 bp). As a second criterion, the preliminary segmentation can employ a threshold based on expected fragment lengths, such as between 60 bp to 140 bp (e.g., about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 105 bp, about 110 bp, about 115 bp, about 120 bp, about 125 bp, about 130 bp, about 135 bp, or about 140 bp). In some embodiments, the fragment length threshold is between 50 to 1200 bp, between 70 to 800 bp, between 100 to 200 bp, between 130 to 180 bp, between 50 to 200 bp, between 60 and 200 bp, between 60 and 150 bp, or between 60 and 100 bp. In some embodiments, the fragment length threshold ranges from 100 to 200 bp, from 120 to 180 bp, from 140 to 160 bp, from 150 to 170 bp, from 160 to 190 bp, or from 170 to 220 bp. In some embodiments, the fragment length threshold is less than 500, 400, 200, 150, 100, 90, 75, 70, 60, or 50 bp in length. As a third criterion, the processivity of the DNA (cytosine-5)-methyltransferase 1 (DNMT1) enzyme may be set from tens of base pairs to a few hundred base pairs.

[0209] For example, a threshold can be selected for the maximum allowable distance between adjacent CpG loci. The genome may be split into stretches in which the distance between adjacent CpGs are less than the threshold (coined as a ridge) and stretches of the genome devoid of any CpG (coined as a trench). In one embodiment, a threshold of about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, or about 150 bp is selected for the maximum allowable distance between adjacent CpG loci, and ridges in which the distance between adjacent CpGs are less than 50 bp, less than 60 bp, less than 70 bp, less than 80 bp, less than 90 bp, less than 100 bp, less than 110 bp, less than 120 bp, less than 130 bp, less than 140 bp, or less than 150 bp are identified and selected for subsequent processing.

[0210] In some embodiments, any CpG that is available in less than a threshold (e.g., about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, or about 40% of total samples) is removed. In some embodiments, the ridges are filtered to contain at least 3 CpG sites, at least 4 CpG sites, at least 5 CpG sites, at least 6 CpG sites, or at least 7 CpG sites. In some embodiments, the ridges are filtered to have a CpG density of at least 0.02 CpG / bp, at least 0.03 CpG / bp, at least 0.04 CpG / bp, at least 0.05 CpG / bp, at least 0.06 CpG / bp, or at least 0.07 CpG / bp. However, such preliminarily segmenting would not be necessary when a targeted panel of oligonucleotide probes as described herein are used for selective enrichment of only the CpG regions of interest (e.g., CpG regions each comprising at least 3, at least 4, at least 5, or at least 6 CpG sites and a CpG density of at least 0.02, at least 0.03, at least 0.04, or at least 0.05 CpG / bp) for targeted methylation sequencing.

[0211] To make the output of the block finding module more versatile, example embodiments use a lenient threshold for r-squared p-value, total number of CpGs in block, and distance between adjacent CpGs in block, but for each block also report the median r-squared, median Pearson r, median Euclidean distance, and median pairwise distance among all adjacent CpG pairs within the block as well as its CpG number. This information can later be used to further filter and narrow down the MHBs reported for each sample if desired. Annotation of putative DMRs / MBM

[0212] Pre-selected regions may be annotated with relevant information to aid in marker selection and prioritization as further discussed below. The annotations may cover three main categories:

[0213] (1) Structure / function: Genomic regions may be structurally and functionally annotated using a comprehensive collection of information containing the existing knowledge on the human genome. For each region, one or more of the following annotations may be included: distance to CpG regions (e.g., CpG island, shore, or shelf); position regarding genes (e.g., exon, intron, 3’UTR, 5’UTR, intergenic, TSS, first exon); overlap with regulatory elements (e.g.,promoter, enhancer, super-enhancer); enrichment for histone marks (e.g., histone methylation and acetylation marks associated with expression state); and multiple additional tracks for annotating transcription factor binding, lncRNAs, chromatin accessibility (DHS), etc. In some embodiments, this information may be pre-calculated for the human genome and imported from external resources (e.g., UCSC genome browser, annotationHub, chromHMM). Example embodiments may employ a module that runs bedtools in python to find overlaps for input regions of interest. Customizable parameters may include: minimum overlap in terms of fraction of input region lengths; minimum overlap in terms of fraction of genomic element lengths; total number of base overlaps reported; etc. As shown in the FIG. 10 schematic, the “intersect” function with the -wa and -wb parameters will return origin entries in A (the input CpG regions bed file) and report specifically the type and size of overlaps (-wao) with additional bed files (e.g., B1 and B2), which would be the annotation tables in our module.

[0214] (2) Fragmentation: Example embodiments include a module that uses plasma cfDNA data to characterize the fragmentation pattern of a given region in healthy / non-diseased cfDNA. This module may further add annotation columns for fragmentation and report overall characteristics and distribution of fragment lengths, end motifs, and other fragmentation metrics. This module may be used for downstream marker selection and prioritization, as discussed above, potentially combining fragmentation information with methylation in targeted sequencing of patient plasma.

[0215] (3) Literature DMRs: In addition to the above functional and structural annotations, previously identified regions of the human genome harboring CpG clusters and cancer DMRs may be added as separate columns in the output. The goal is to specify if each candidate biomarker region overlaps with independently identified regions previously reported in the literature as DMR or other relevant association.

[0216] Example embodiments include: a script to receive genomic regions as input in .bed format and output an annotated .bed file with multiple columns added based on user- selected annotations from the list of available annotations; a script to use available sequencingdata for characterization of fragmentation profiles in genomic regions of interest; a confidence level assigned to each region based on existing evidence (e.g., known CRC markers) to help in marker ranking and prioritization; and visualization summary plots for overlaps with known annotations. Methylation Biomarker Selection

[0217] Biomarkers generally are informative regions that contain differential signals between different cohorts (e.g., between individuals with cancer and healthy or non-cancerous individuals). In terms of methylation biomarkers, informative regions typically have two main characteristics: (i) CpG loci are co-methylated in these regions, and (ii) they correspond to distinct functional domains (i.e., each region captures one or similar functional features). How to identify and score such biomarker regions will be further discussed below, including an example modeling approach that identifies putative biomarkers by comparing the methylation and co- methylation patterns between different cohort samples in different regions and identifying differentially methylated genomic regions. The various illustrative approaches and embodiments disclosed herein improve and expand on existing methods by, for example, inferring the methylation state of a tentative biomarker through modeling the processivity of DNMT1 by jointly modeling adjacent CpG loci methylation patterns; considering each read’s mapping quality in such joint modeling; and benchmarking different models of varying complexity. (i) Methylation Haplotype Measures

[0218] In various embodiments, with reference to FIG.7, the analysis may optionally implement and compare the methods developed in the literature (e.g., Methylation Entropy from Xie et al. (2011), Epi-Polymorphism from Landan et al. (2012), and Methylation Haplotype Load from Guo et al. (2017)) to score the methylation status of a genomic region as a function of the observed methylation patterns of the fragments in that region. These models and methods may be used as a baseline for the method described herein to infer the methylation status of the CpG loci in the WGBS data across samples. Example embodiments identify regions where themethylation status is almost constant in a given region across samples, and identify regions which are significantly differential between samples. In example embodiments, this approach also prioritizes each region based on its association with different clinical covariates.

[0219] To achieve robustness in the methylation inference model and to guide the inference, example embodiments may leverage publicly available methylation datasets to construct CpG methylation status prior distributions. A unified framework may be employed to fetch methylation status of CpG loci in a given genomic region. As examples, the Human Methylome Atlas, Roadmap Epigenetics, and Blueprint Epigenome may be used for tissue and cell-type specific methylation prior distributions, and Fox-Fisher (2021) may be used for bulk plasma and white blood cell methylation prior distributions.

[0220] Example embodiments consider a Markovian process to infer the methylation status of a CpG locus and capture the patterns of CpG co-methylation in a methylation haplotype block by leveraging read-level information. Various embodiments implement probabilistic modeling of the DNMT1 methylation mechanism in order to accurately infer the methylation state of a tentative biomarker (see FIG. 8). The processivity and substrate-specificity of methyltransferases may be modeled via a Hidden Markov Model.

[0221] In various embodiments, a methylation score is used to capture the level of co- methylation in a given methylation haplotype block. This score offers an estimate for the total methylation level as well as the level of co-methylation in a given window and forms part of a module called CpGBlocks. In example embodiments, the CpGBlocks module is packaged in ProcessSegmentation, as further described below. (ii) Process Segmentation of Methylation Haplotype Blocks

[0222] In various embodiments, one of the main steps of identifying methylation biomarkers is to define the notion of a “region.” For example, a biomarker may be a region where the methylation score (e.g. the output of the methylation haplotype measure) is as uniform as possible under a loss function (e.g., maximize a likelihood or minimize a negative loglikelihood) within each sample type or population yet different across different sample types or populations. In some embodiments, a biomarker region may be a region where the methylation scores of all CpGs in that region can be approximated as a constant function within each sample type or population.

[0223] In example embodiments, to further refine the identified sample-specific MHBs into smaller regions where CpGs are co-methylated under the defined MHM across multiple samples, the CpGBlocks module is packaged in ProcessSegmentation, which is a generalized module for fitting a piecewise constant function to methylation load signals which are exponential family distributed. Example embodiments apply the model to the observed MHM values on all MHBs across all samples to identify regions of constant methylation state across multiple samples. The Methylation Haplotype Measure is applied to all samples and then the Process Segmentation model applied to the signal to identify the genomic blocks. This step provides as output the number of blocks, their lengths, and their scores.

[0224] For example, suppose that the observed methylation statuses are realizations of independent random variables

[0225] Further assume that the support for each parameter is divided into ^contiguous, non-overlapping intervals Exampleembodiments model the distribution of as independent and identicallydistributed (“IID”):

[0226] Under this model the log-likelihood may be written aswhich can be solved via dynamic programming.

[0227] In example embodiments, a probabilistic model is constructed to infer the methylation status of the CpG loci. There are multiple advantages to the probabilistic modeling of methylation status rather than using deterministic functions that are tailored to estimate methylation load: (i) various embodiments make the generative model as close as possible to the biological explanation of the CpG methylation patterns; (ii) various embodiments can leverage publicly available information to construct priors and inject these priors at different levels of the model; (iii) hypothesis testing can be performed very naturally as these models can be fully investigated; and (iv) inference comes with uncertainty estimation. Example embodiments first infer the methylation status of all CpGs in a genomic window by assuming that the CpGs are independent then add the dependence structure between these loci and perform inference.

[0228] Example embodiments design and implement a model to fit a piecewise constant function to simulated values. A Bayesian approach (e.g., a fully variational Bayesian approach or an exact inference Bayesian method) to the process segmentation problem may be employed; for example, a Bayesian method for gaussian distributed signal. Example embodiments extend such a model to beta-binomially distributed signal where Stochastic Variational Bayes (SVI) is used instead of exact inference.

[0229] In example embodiments, the prior distribution of the break-points can be constructed from publicly available DMRs (e.g., Guo et al.(2017), Benelli et al.(2021), Loyfer et al.(2022), Roadmap Epigenomics Project). In order to estimate the “optimal” number of breakpoints (i.e., segments), multiple methods may be employed. Example embodiments formalize this model selection criterion as a penalty term in a penalized log-likelihood maximization:where &1is the penalized likelihood and ^ is the optimal number of breakpoints. Table 1 depicts different criteria for model selection:

[0230] Example embodiments fit a change point estimation model to the methylation load signal. The inferred change points indicate the boundaries of the regions where the signal is almost constant. These regions will in turn be passed to the next part of the discovery pipeline for scoring and prioritization.

[0231] Given a sequence of methylation levels of n consecutive CpG sites,assuming each is independently distributed according to a normal distribution, the data likelihood is:

[0232] Example embodiments model the segmentation problem as a regression where the function is piecewise constant. That is, given k boundariesthe true signal is generated from a function I which is constant onTherefore y is IID distributed within each segment. An objective is to estimate the number of distinct regions N, the location of the boundaries of these regions and regionmethylation levels

[0233] With respect to the model, example embodiments start by assuming priors over parameters of interest. Assuming a uniform prior over the number of segments N, a uniform prior over all segmentations into N segments, and a Gaussian distribution over the methylation level with parameters O and 5 , the prior distribution takes the following form:(iii) Identifying, scoring, and prioritization of DMRs

[0234] Having identified the differentially methylated regions and inferred the methylation status of each CpG in each of the regions, example embodiments compare scores from above (e.g., differential methylation score) across cohorts and test for significant differences, which will result in a set of differentially methylated biomarkers.

[0235] Example embodiments develop baseline models based on available methods from the literature (e.g., Benelli (2021), Loyfer (2022), and Guo (2017)). For example, WGBS Tools, simple t-tests or Area Under the Curve (AUC) from the Receiver Operator Characteristic (ROC) curves may be used to call a biomarker as significantly differential.

[0236] One of the main advantages of the disclosed Bayesian probabilistic modeling is that as part of the inference, example embodiments infer the probability distribution of the unknown parameters, in this case CpG methylation status. Having access to the distributions, estimating the probability of the two distributions being different is relatively not computationally demanding.

[0237] In various embodiments, sets of published DMRs from the literature can be used as inputs to the disclosed MHM measure and the level of concordance between the disclosed score and hypo / hyper methylation status of these regions can be compared. ROC curves of predicting hypo- / hyper-methylated status of these biomarkers can be plotted. [0238ple embodiments construct control regions from GenomeInABottle (GIAB) Stratifications, an application programming interface (API) for the Genome In A BottleStratifications project. These stratification files were developed as a standard resource of Browser Extensible Data (BED) files (a text file format used to store genomic regions as coordinates and associated annotations) for use in stratifying true positive, false positive, and false negative variant calls in challenging and targeted regions of the genome. Example embodiments annotate biomarker regions in terms of Genome In A Bottle stratification. Difficult-to-map regions should not have a statistically significant overlap with biomarkers discovered according to the disclosed embodiments. Example embodiments add some of the regions to the final panel as negative controls, which should not be differentially methylated if other parts of the mapping / processing is correctly performed. Example embodiments create a list of genomic loci / regions with high noise, high repeat, or containing no CpGs. This list of blacklist regions can be used as negative controls.

[0239] Example embodiments prioritize the identified candidate biomarkers for inclusion in a final panel based on certain criteria. In some embodiments, the identified candidate biomarkers are prioritized based on clinical data (e.g., sex, age, ethnicity, etc.). Example embodiments use feature scoring methods to score methylation biomarkers set by accounting for the methylation signal correlations of DMRs and prioritize them for their varying predictive powers across subpopulations (e.g., cancer subtypes, high-risk populations, age, gender, ethnicity, ctDNA fractions, etc.).

[0240] In various embodiments, a Constrained Probability Sampling (CPS) module is used for sampling from a population under multiple probability constraints. Example embodiments score each biomarker region based on their predictive power. Then the prioritization problem is stated as a constrained optimization problem where the objective function is the sum of these scores such that the cohort constraints are met; for example, the ethnic diversity of the cohort is a good representative of the target population.

[0241] Various embodiments assign a differential power,where N is the total number of biomarkers, J is the number of samples incohort k, and K is the total number of cohorts, as the AUC of the ROC curve of differentiatingbetween the cohorts. The problem can then be formalized as the following optimization problem:

[0242] For different constraint thresholds, this problem can be solved to prioritize the top M biomarkers under length constraints, ethnicity, sex, age, etc.

[0243] In example embodiments, an output of this analysis is a list of target regions that is balanced in terms of predictive power across subpopulations along with functional annotation and GIAB classification. Fragment Level Methylation Data Manipulation and Representation ToolBox

[0244] Example embodiments provide a file format to concisely represent CpG methylation patterns in a human-readable format – e.g., text instead of binary – such that mapping quality information is preserved and provides portability in a sense that it is manipulatable using commonly-used tools, such as BEDtools, and basic table editing software. Example embodiments implement standard visualizations for read level representation of CpG haplotype block methylation status.

[0245] In order to identify, annotate, and interpret DMRs, obtained as discussed above, example embodiments use fragment-level methylation information to infer the methylation status in a genomic region. Regular BAM files are too large and also contain more information than necessary. Other available formats (e.g., .pat from WGBSTools) are binary formats which have two main shortcomings: (1) these files are not portable – that is, in order to read, manipulate, or annotate these files, one needs to utilize specific toolbox / libraries which necessitate developing new interfaces for manipulating these files; and (2) these formats do not faithfully conserve fragment level information. In example embodiments of the disclosed file format, this information can be preserved by, for example, representing each fragment on a separate line.Each fragment is also scored based on the phred quality score of the CpGs in the fragment and the MAPQ score.

[0246] Example embodiments accommodate efficient file storage without loss of relevant information for downstream processes for all DNA methylation sequencing experiments in the future, such as deep WGBS samples from the DNA methylation biomarker discovery discussed above. Embodiments of the disclosed file format will allow archiving of all FASTQ or BAM files on any cloud platform, such as DNANexus, Google Cloud Platform, or Amazon AWS, while directly providing convenient input to downstream analyses. Also, example embodiments provide the ability to present the innovative analyses algorithms and resulting summarization developed through methylation biomarker discovery.

[0247] .cpg Constructor: Example embodiments provide for a .cgp file format that is used to represent methylation patterns in a given genomic region at the fragment / read level. Example embodiments provide multiple visualization tools to represent the methylation patterns of CpG loci in a given genomic interval. For example, “CpG Heatmap” provides a heatmap of fragment / read level methylation patterns in a given genomic window, and “CpG Tracks” provides stacked barplots representing methylation status in each CpG loci in a given interval. This allows for the implementation of tools and visualizations to represent methylation haplotype blocks in such a way that captures patterns of co-methylation within genomic windows.

[0248] Table 2 illustrates example embodiments of the .cpg file format (see FIGs. 13 and 14), with each row representing a fragment. In Table 2, chrom: chromosome, start: start of the fragment, end: end of the fragment – the beginning and the end of the fragment are inferred from the BAM file, cpg methylation pattern: methylation state of cpgs in the fragment as well as their loci in a CIGAR string compliant format, score: -10log10(p) where p is 1 – accuracy of each fragment in calling each base correctly.

[0249] In the case of paired-reads, example embodiments may summarize the fragment information using the following conventions: “start” of the fragment is the minimum of the two reads’ start coordinates; “end” of the fragment is set to the end of the two reads; “cpg methylation pattern” of the fragment is set to the methylation state if the two reads match, and in case they do not, example embodiments set it to the state to N. If the CpG is covered only by one of the reads in a pair, the CpG methylation status is set according to that one read. Previous approaches to paired end read alignment merger did not combine reads if they do not overlap, which can be very prevalent in WGBS data.

[0250] Using the .cpg file format as disclosed herein enables efficient extraction of CpG methylation information from bismark output. Example embodiments of the .cpg file format provide numerous advantages over the existing formats, such as:

[0251] (1) Preservation of fragment level methylation information: .cpg files are lossless representations of the sample CpG methylation information. In effect, each line represents a single fragment.

[0252] (2) Portability: unlike other file formats (such as .pat from WGBSTools) CpG files are viewable using any text viewer / editor (such as vim, emacs, and less), or spreadsheet software (e.g., Google Spreadsheet or Microsoft Excel). No special tool is needed for reading the .cpg files.

[0253] (3) Compatibility: CpG files are tab-delimited table format with headers and thus compatible with most common table manipulation software (e.g., pandas, R) and genomic interval manipulation packages (e.g., BedTools or intervaltree).

[0254] (4) Compressibility: Because only relevant information (e.g. information related to CpG methylation) is retained, example embodiments of the CpG files offer at least 10x, 20x, 30x, 50x, 100x, 150x, or 200x compression relative to the BAM files.

[0255] Example embodiments allow for conversion of a bam file, which is the output of the bismark software, to a .cpg file in one step using the CpGReads package. For example: CpGReads.convert_bismark_bam_to_cpg(input_bam_fpath, output_cpg_fpath, compress=True)

[0256] If compress = True, the output will be compressed and saved as a .cpg.gz which can be loaded the same way as a regular .cpg file.

[0257] In example embodiments, once CpG files are constructed, they can be loaded using the following command, whether the input is a compressed .gz file or not. For example: cpg_reads = CpGReads.load_from_cpg(output_cpg_fpath)

[0258] Having instantiated a cpg_reads object, one can parse the methylation patterns using, for example:cpg_reads.expand_cpg_pattern()

[0259] Now the user has the pandas machinery to efficiently query, sort, search these datasets. For example: cpg_reads.query("chrom == @region.chrom").query("start >= @region.start and end <= @region.end").sample(frac=.01).sort_values(['start', 'end'])

[0260] Example embodiments of this module are able to handle both single-end and paired-end reads. Paired-end reads are especially very useful in cell-free DNA analysis, but processing paired-end runs require much greater care and effort.

[0261] CpG Loci Visualization Tool: Example embodiments provide a CpG loci visualization tool that uses the .cpg file format as an input and retrieves the CpG methylation pattern of the fragments in a .cpg file as a color coded heatmap. These plots facilitate exploratory data analysis and quality control as well as publication quality visualizations for posters / publications. These plots provide the ability to visualize discrete as well as continuous heatmaps such that they can be used for visualization of methylation level (β) values data.

[0262] Example embodiments provide user-friendly, publication quality visualization tools (e.g., included in CpGReads) to plot fragment-level methylation patterns. Example embodiments enable obtaining plots by:

[0263] Fragment level methylation pattern plot of CpG loci only: cpg_reads_region.plot(kind='heatmap', compact=True, annotate=True, ax=ax)

[0264] Plotting methylation information of all loci in a given region: cpg_reads = CpGReads.load_from_cpg(output_cpg_fpath) cpg_reads_region = cpg_reads.query("chrom == @region.chrom").query("start >= @region.start and end <= @region.end").sample(frac=.01).sort_values(['start', 'end'])cpg_reads_region.plot(kind='heatmap', compact=False, annotate=False, ax=ax)

[0265] Code to plot population level CpG matrices (see FIG. 17): = HybridCaptureSampleManifest.load_hybrid_capture() sample_ids = hc_sample_manifest.sample_id.values # Loading target regions as a RegionManifest region_manifest = RegionManifest.from_bed(regions_bed_fpath, assembly='hg19').attach_region_ids() # Loads fragment information region_manifest = hc_sample_manifest.load_fragments_in_regions(region_manifest, n_jobs) cpg_matrices = CpGReads.cpgreads_list_to_cpg_matrices(region_manifest [sample_ids], sampling_frac) CpGReads.plot_cpg_matrices(cpg_matrices, sample_ids, out_fpath)

[0266] Methylation Block Visualization: Example embodiments provide visual representation of DNA co-methylation levels across genomic regions of interest. It has previously been established that due to processivity of DNA methyltransferase enzymes (DNMTs), adjacent CpGs present concordant methylation states (typically within 100 bp distance). Therefore, regions of tightly correlated methylation called co-methylation blocks are discovered for pre-selection. When correlation matrices are measured and plotted, these block regions can be visualized in the form of triangles (e.g., FIGs. 18 and 19) of high correlation as shown above.

[0267] Example embodiments provide functionality that is compatible with other methylation modules in python (e.g., as discussed above). More specifically, exampleembodiments of this function will receive the .cpg file as input and process the file using functions to load, query regions, and convert methylation haplotypes to numeric matrices developed as part of the CpGreads module. Next, the function calculates and plots the correlation matrix (or linkage disequilibrium, controlled by the “method” parameter) for all CpGs within the given region of interest purely using python functionalities.

[0268] Example embodiments provide tools for visualizing CpG co-methylation (LD or Pearson correlation) and fraction methylation (beta-value) as part of the CpGBlocks module. An example usage of this on a CRC sample in a given region of interest (target 675 in hybrid capture (HC) panel):

[0269] Following is a code snippet for creating triangle plots in an exemplary embodiment: from ecd.methylation.cpgreads import CpGReads from ecd.methylation.cpgblocks import CpGBlocks cpgreads = CpGReads.load_from_cpg(path_to_cpg_file).query("chrom == @region.chrom").query("start >= @region.start and end <= @region.end") CpGBlocks(cpg_reads_region).plot_triangle_from_cpg(output_path=". / ", clustering=1, method='corr', highlights=list_of_cpg_loci_to_highlight)

[0270] One application of the Methylation Block Visualization tool is for comparing multiple regions for marker selection. In an example, 627 marker regions were included in a hybrid capture (HC) targeted panel. By quantifying and visualizing the level of co-methylation among CpGs in different target regions, it is possible to filter and prioritize target regions. For instance, FIG.19 shows two example regions (targets 1 and 2). Both targets show decent co-methylation in CRC samples which is desirable. However, only target 2 also shows some level of correlation in cfDNA controls.

[0271] As illustrated, CpGReads provides the capability to process BAM files and convert them to a considerably compressed format, .cpg, which is highly portable compared to the s.o.t.a in linear time O(LL). This library also offers publication quality plots for performing EDA on methylome sequencing data. Exemplary Identification of Candidate Methylation Biomarkers for CRC Detection

[0272] The various methods and their embodiments disclosed herein can be used to detect the presence of a certain phenotype, such as a disease or cancer, in a subject. For example, various embodiments seek to identify sets of DMRs for detecting a phenotype by methylation sequencing of converted DNA from phenotypic samples (e.g. formalin-fixed paraffin-embedded (FFPE) tumor specimens or circulating tumor DNA (ctDNA) samples) and converted gDNA or cfDNA from healthy / non-phenotypic samples. An example experimental design may involve use of deep, whole genome or panel-based sequencing on DNA from healthy / non-phenotypic samples (e.g., gDNA or cfDNA from normal tissues or healthy / non- phenotypic individuals) and from phenotypic samples (e.g., DNA from tumor biopsies or ctDNA from cancer patients) following conversion of unmethylated cytosine (C) in CpG di-nucleotides to uracil (U). In some embodiments, methylome sequencing data of cfDNA from healthy / non- phenotypic individuals are compared to methylome sequencing data of DNA from tumor biopsies. In some embodiments, methylome sequencing data of DNA from tumor biopsies are compared to methylome sequencing data of DNA from adjacent normal tissues.

[0273] In one example embodiment, deep whole genome methylation sequencing was conducted on converted DNA samples from 100 healthy individuals (median effective coverage of 352x±38x) and from tumor specimens of 75 CRC patients (median effective coverage of 48x±18x). An exemplary experimental workflow is shown in Fig.3A. DNA was extracted from samples sourced from 100 healthy individuals and 75 patients with CRC. Following librarypreparation, sequencing was conducted on converted DNA. Reads were mapped and aligned to a reference genome.

[0274] As shown in FIG.3B, the two cohorts were chosen to be balanced in terms of age (CRC: 58±11 years; healthy: 57±8 years) and to represent major ethnicities in the USA, and the CRC cohort was chosen to represent all stages, histologies and morphologies, and microsatellite instability (MSI, 23%). FIG. 3B shows histograms of age distribution among healthy individuals and CRC patients (left), self-reported ethnicity among healthy individuals and CRC patients (middle), and proportion of CRC patients by stage, further stratified by microsatellite status (right). FIG. 4 depicts the cohort level distributions of effective coverage, methylation levels, and protection rates. Differentially methylated genomic patterns were defined between patients with CRC and healthy individuals by utilizing patterns of co-methylation of adjacent CpG loci. These differential methylation profiles are strong biomarkers for targeted sequencing panels for early cancer detection.

[0275] In the exemplary embodiment, individualized sample-level methylomes were created and combined with hg19 reference using the methods disclosed herein, which resulted in 51M total CpGs. Any CpG available in less than 25% of total samples were removed, resulting in 28.5M CpGs. FIG. 5 shows the CpG distribution per cohort. The number of CpGs covered is plotted against the fraction of samples in each group covering those CpGs.

[0276] Using the combined methylome and a threshold of 100bp, 8.1M ridges were found using the preliminary segmentation methods disclosed herein. The ridges were filtered to autosomes only, and to regions that contain at least 5 CpGs at a density of at least 0.03 CpG / bp, which resulted in 482K ridges containing 8.26M CpGs and 147.5Mbp. The distribution of these ridges is shown in FIG.6.

[0277] Further segmentation was then performed using a machine learning model as described herein to identify regions differentially methylated between the healthy / non-CRC and CRC cohorts as well as between subpopulations within the CRC cohort, including age, sex,ethnicity, stage, histology, and microsatellite instability (MSI) vs microsatellite stability (MSS). FIG. 9 illustrates such a segmentation procedure. For a given region, the methylation load of the healthy / non-CRC and CRC cohorts were plotted (top), and the segmentation procedure identifies and calculates the differential methylation load between the region (bottom). The horizontal bars in the lower panel show the differential methylation load.

[0278] Candidate biomarkers are scored and prioritized using methods disclosed herein. FIG. 10 shows differential methylation score and CpG density distribution of biomarker candidates. FIG. 11 shows number of biomarker candidates and the differential methylation score cutoff based on a preset 2Mbp total panel size or number of biomarkers (top), and CpG density, length, and methylation load distribution of biomarker candidates (bottom).

[0279] Using the methods disclosed herein, 184,265 differentially methylated regions between healthy / non-CRC and CRC cohorts, spanning more than 1.71 million CpGs and more than 18.76 Mbp, were discovered. 76% of all discovered regions in the example embodiment were hypermethylated (70% CpG islands, 1.3% CpG shelves, 21% CpG shores) and the rest were hypomethylated (7.3% CpG islands, 8.6% CpG shelves, 25% CpG shores). FIGs.20 and 21 show the density distributions of the biomarkers selected using the methods described herein.

[0280] Based on deep whole genome methylation sequencing data from heterogeneous cohorts (see Fig. 3), representative of the US population eligible for CRC screening, the example embodiment identified many differentially methylated regions as strong biomarker candidates for the early detection of cancer (see Fig.20). Compared to recurrently mutated driver genes in CRC, it was found that differential methylation profiles of CRC samples are generally much more homogenous and are therefore ideal biomarkers for targeted sequencing panels with the purpose of early cancer detection. Further, analysis of these patterns based on clinicopathologic factors enabled discovery of subtype specific markers (see Fig. 21), which can be incorporated to ensure equivalent performance for underrepresented subtypes or subpopulations.

[0281] Various operations described herein can be implemented on computer systems having various configurations. FIG. 22 shows a simplified block diagram of a representative server system 2200 (e.g., computing system 110 or other components in FIG. 1) and client computer system 2214 (e.g., computing system 110, sequencing system 160, information system 170, and / or sample processing system 180) usable to implement various embodiments of the present disclosure. In various embodiments, server system 2200 or similar systems can implement services or servers described herein or portions thereof. Client computer system 2214 or similar systems can implement clients described herein.

[0282] Server system 2200 can have a modular design that incorporates a number of modules 2202 (e.g., blades in a blade server embodiment); while two modules 2202 are shown, any number can be provided. Each module 2202 can include processing unit(s) 2204 and local storage 2206.

[0283] Processing unit(s) 2204 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 2204 can include a general-purpose primary processor as well as one or more special-purpose co-processors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 2204 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 2204 can execute instructions stored in local storage 2206. Any type of processors in any combination can be included in processing unit(s) 2204.

[0284] Local storage 2206 can include volatile storage media (e.g., conventional DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 2206 can be fixed, removable or upgradeable as desired. Local storage 2206 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 2204 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 2204. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 2202 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections.

[0285] In some embodiments, local storage 2206 can store one or more software programs to be executed by processing unit(s) 2204, such as an operating system and / or programs implementing various server functions or any system or device described herein.

[0286] “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 2204 cause server system 2200 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 2204. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 2206 (or non-local storage described below), processing unit(s) 2204 can retrieve program instructions to execute and data to process in order to execute various operations described above.

[0287] In some server systems 2200, multiple modules 2202 can be interconnected via a bus or other interconnect 2208, forming a local area network that supports communication between modules 2202 and other components of server system 2200. Interconnect 2208 can be implemented using various technologies including server racks, hubs, routers, etc.

[0288] A wide area network (WAN) interface 2210 can provide data communication capability between the local area network (interconnect 2208) and a larger network, such as the Internet. Conventional or other activities technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.11 standards).

[0289] In some embodiments, local storage 2206 is intended to provide working memory for processing unit(s) 2204, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 2208. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 2212 that can be connected to interconnect 2208. Mass storage subsystem 2212 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 2212. In some embodiments, additional data storage resources may be accessible via WAN interface 2210 (potentially with increased latency).

[0290] Server system 2200 can operate in response to requests received via WAN interface 2210. For example, one of modules 2202 can implement a supervisory function and assign discrete tasks to other modules 2202 in response to received requests. Conventional work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 2210. Such operation can generally be automated. Further, in some embodiments, WAN interface 2210 can connect multiple server systems 2200 to each other, providing scalable systems capable of managing high volumes of activity. Conventional or other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation.

[0291] Server system 2200 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 22 as client computing system 2214. Client computing system 2214 can be implemented,for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on.

[0292] Client computing system 2214 can communicate via WAN interface 2210. Client computing system 2214 can include conventional computer components such as processing unit(s) 2216, storage device 2218, network interface 2220, user input device 2222, and user output device 2224. Client computing system 2214 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like.

[0293] Processor 2216 and storage device 2218 can be similar to processing unit(s) 2204 and local storage 2206 described above. Suitable devices can be selected based on the demands to be placed on client computing system 2214; for example, client computing system 2214 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 2214 can be provisioned with program code executable by processing unit(s) 2216 to enable various interactions with server system 2200 of a message management service such as accessing messages, performing actions on messages, and other interactions described above. Some client computing systems 2214 can also interact with a messaging service independently of the message management service.

[0294] Network interface 2220 can provide a connection to a wide area network (e.g., the Internet) to which WAN interface 2210 of server system 2200 is also connected. In various embodiments, network interface 2220 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, 5G, LTE, etc.).

[0295] User input device 2222 can include any device (or devices) via which a user can provide signals to client computing system 2214; client computing system 2214 can interpret the signals as indicative of particular user requests or information. In various embodiments, userinput device 2222 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on.

[0296] User output device 2224 can include any device via which client computing system 2214 can provide information to a user. For example, user output device 2224 can include a display to display images generated by or delivered to client computing system 2214. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to- analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 2224 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.

[0297] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 2204 and 2216 can provide various functionality for server system 2200 and client computing system 2214, including any of the functionality described herein as being performed by a server or client, or other functionality associated with message management services.

[0298] It will be appreciated that server system 2200 and client computing system 2214 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 2200 and client computing system 2214 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software.

[0299] While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. For instance, although specific examples of rules (including triggering conditions and / or resulting actions) and processes for generating suggested rules are described, other rules and processes can be implemented. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to specific examples described herein.

[0300] Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and softwarecomponents, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa.

[0301] Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium).

[0302] Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims.

[0303] As utilized herein, the terms “approximately,” “about,” “substantially”, and similar terms are intended to have a broad meaning in harmony with the common and accepted usage by those of ordinary skill in the art to which the subject matter of this disclosure pertains. It should be understood by those of skill in the art who review this disclosure that these terms are intended to allow a description of certain features described and claimed without restricting the scope of these features to the precise numerical ranges provided. Accordingly, these terms should be interpreted as indicating that insubstantial or inconsequential modifications or alterations of the subject matter described and claimed are considered to be within the scope of the disclosure as recited in the appended claims. In some examples, these terms allow for a plus- or-minus deviation of 10 percent.

[0304] It should be noted that the terms “exemplary,” “example,” “potential,” and variations thereof, as used herein to describe various embodiments, are intended to indicate that such embodiments are possible examples, representations, or illustrations of possibleembodiments (and such terms are not intended to connote that such embodiments are necessarily extraordinary or superlative examples).

[0305] The term “coupled” and variations thereof, as used herein, means the joining of two members directly or indirectly to one another. Such joining may be stationary (e.g., permanent or fixed) or moveable (e.g., removable or releasable). Such joining may be achieved with the two members coupled directly to each other, with the two members coupled to each other using a separate intervening member and any additional intermediate members coupled with one another, or with the two members coupled to each other using an intervening member that is integrally formed as a single unitary body with one of the two members. If “coupled” or variations thereof are modified by an additional term (e.g., directly coupled), the generic definition of “coupled” provided above is modified by the plain language meaning of the additional term (e.g., “directly coupled” means the joining of two members without any separate intervening member), resulting in a narrower definition than the generic definition of “coupled” provided above. Such coupling may be mechanical, electrical, or fluidic.

[0306] The term “or,” as used herein, is used in its inclusive sense (and not in its exclusive sense) so that when used to connect a list of elements, the term “or” means one, some, or all of the elements in the list. Conjunctive language such as the phrase “at least one of X, Y, and Z,” unless specifically stated otherwise, is understood to convey that an element may be either X, Y, Z; X and Y; X and Z; Y and Z; or X, Y, and Z (i.e., any combination of X, Y, and Z). Thus, such conjunctive language is not generally intended to imply that certain embodiments require at least one of X, at least one of Y, and at least one of Z to each be present, unless otherwise indicated.

[0307] References herein to the positions of elements (e.g., “top,” “bottom,” “above,” “below”) are merely used to describe the orientation of various elements in the Figures. It should be noted that the orientation of various elements may differ according to other exemplary embodiments, and that such variations are intended to be encompassed by the present disclosure.

[0308] The embodiments described herein have been described with reference to drawings. The drawings illustrate certain details of specific embodiments that implement the systems, methods and programs described herein. However, describing the embodiments with drawings should not be construed as imposing on the disclosure any limitations that may be present in the drawings.

[0309] It is important to note that the construction and arrangement of the devices, assemblies, and steps as shown in the various exemplary embodiments is illustrative only. Additionally, any element disclosed in one embodiment may be incorporated or utilized with any other embodiment disclosed herein. Although only one example of an element from one embodiment that can be incorporated or utilized in another embodiment has been described above, it should be appreciated that other elements of the various embodiments may be incorporated or utilized with any of the other embodiments disclosed herein.

[0310] The foregoing description of embodiments has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form disclosed, and modifications and variations are possible in light of the above teachings or may be acquired from this disclosure. The embodiments were chosen and described in order to explain the principles of the disclosure and its practical application to enable one skilled in the art to utilize the various embodiments and with various modifications as are suited to the particular use contemplated. Other substitutions, modifications, changes and omissions may be made in the design, operating conditions and arrangement of the embodiments without departing from the scope of the present disclosure as expressed in the appended claims. WORKING EXAMPLES.

[0311] Example 1

[0312] Deep whole genome methylation sequencing was conducted on converted DNA samples from 100 healthy individuals (median effective coverage of 352x±38x) and from tumor specimens of 75 CRC patients (median effective coverage of 48x±18x). An exemplaryexperimental workflow is shown in Fig.3A. DNA was extracted from samples sourced from 100 healthy individuals and 75 patients with CRC. Following library preparation, sequencing was conducted on converted DNA. Reads were mapped and aligned to a reference genome.

[0313] As shown in FIG.3B, the two cohorts were chosen to be balanced in terms of age (CRC: 58±11 years; healthy: 57±8 years) and to represent major ethnicities in the USA, and the CRC cohort was chosen to represent all stages, histologies and morphologies, and microsatellite instability (MSI, 23%). FIG. 3B shows histograms of age distribution among healthy individuals and CRC patients (left), self-reported ethnicity among healthy individuals and CRC patients (middle), and proportion of CRC patients by stage, further stratified by microsatellite status (right). FIG. 4 depicts the cohort level distributions of effective coverage, methylation levels, and protection rates. Differentially methylated genomic patterns were defined between patients with CRC and healthy individuals by utilizing patterns of co-methylation of adjacent CpG loci. These differential methylation profiles are strong biomarkers for targeted sequencing panels for early cancer detection.

[0314] Individualized sample-level methylomes were created and combined with hg19 reference using the methods disclosed herein, which resulted in 51M total CpGs. Any CpG available in less than 25% of total samples were removed, resulting in 28.5M CpGs. FIG. 5 shows the CpG distribution per cohort. The number of CpGs covered is plotted against the fraction of samples in each group covering those CpGs. Using the combined methylome and a threshold of 100bp, 8.1M ridges were found using the preliminary segmentation methods disclosed herein. The ridges were filtered to autosomes only, and to regions that contain at least 5 CpGs at a density of at least 0.03 CpG / bp, which resulted in 482K ridges containing 8.26M CpGs and 147.5Mbp. The distribution of these ridges is shown in FIG. 6.

[0315] Example 2: Methylation Biomarker Discovery Panel

[0316] A new methylated biomarker discovery engine according to the methods disclosed herein includes (i) a custom hybrid capture panel that targets all the regions of thehuman genome that have a high enough CpG count and CpG density, (ii) an end to end workflow for sequencing the high CpG density regions from negative (e.g. cancer free) and positive (for example, cancer) samples, and (iii) an automated analysis pipeline that uses the sequence data to identify the candidate differentially methylated biomarkers (DMR) that will differentiate the positive from negative samples. The DMR discovery engine is compatible with all sample types, including for example fresh-frozen tissue, FFPE tissue, fresh blood, serum, buffy coat, plasma, urine, vitreous, sputum, saliva, tears, perspiration, ascites fluid, feces, bile, lymph, cervical mucus, semen, gDNA, and cfDNA. The protocol is also sequencing platform agnostic and is compatible with any short-read sequencer, including for example Illumina NovaSeq, BGI T7, Element Biosciences AVITI, PacBio Onso, Ultima Genomics UG100, and Singular Genomics G4 platforms. In addition, this DMR discovery engine is application agnostic and can be used for the discovery of DMRs that distinguish any pathological condition from the healthy condition, including in oncology (e.g., early cancer detection, tissue-free MRD), women’s health (e.g., NIPT), and organ health (e.g., transplant rejection).

[0317] A prior version of the methylation biomarker discovery workflow is based on whole genome bisulfite sequencing and includes the steps of: DNA extraction, DNA fragmentation, size selection, library prep, cytosine conversion, library amplification, and whole genome bisulfite sequencing. The sequence data from negative (non-cancer) and positive (cancer) samples is analyzed to determine the genomic regions that are differentially methylated between the two sample sets (positive and negative). A targeted panel consisting of these differentially methylated biomarkers is then generated for non-invasive detection of the condition of interest (e.g., cancer).

[0318] In order to leverage co-methylation patterns for plasma cfDNA analysis, it is desirable that the selected DMRs contain several (e.g. >3) CpG sites per cfDNA molecule. Because of this, only genomic regions with a high enough CpG count and CpG density are analyzed for DMR discovery. These genomic regions add up to ~136 Mb of sequence space, and the rest of the genomic regions are discarded in the analysis. Therefore, an improved version of the methylation biomarker discovery workflow described herein uses a panel of hybrid captureprobes for enrichment of all genomic regions with sufficiently high CpG number and CpG density (filtered for repeat regions), wherein only these genomic regions with sufficiently high CpG number and CpG density are sequenced for methylation biomarker discovery. In addition to the target regions, the panel of hybrid capture probes can also include probes for: (i) endogenous methylation control regions (e.g. constitutively methylated and unmethylated regions) and / or exogenous control sequences (e.g. pUC19 and lambda sequences) for an orthogonal measurement of the conversion and protection rates, (ii) age and gender targets for sample integrity checks, and (iii) SNP tracers for controlling for sample swaps. Further, a hybrid capture step prior to sequencing can improve mapping rate by only selecting library molecules that are expected to map.

[0319] Specifically, we extracted all CpG loci in the hg19 reference genome (n=26.75M when limited to Autosomes). Fig. 23 shows the distribution of CpG loci as well as their annotation across autosomes. Many of these CpG sites are located on repeat regions where the sequence complexity is very low, which makes them not suitable for panel design. Designing probes in these regions would result in high off-target rate and a decrease in effective coverage. Fig.24 shows the distribution of CpG loci after removing CpG sites located in repeat regions except the “GC rich” category. “GC rich” in this context refers to repeat regions that typically have 60% or more GC content. About 13.8 million CpG sites are located in a repeat region.

[0320] In order to leverage the co-methylation patterns of adjacent CpG loci, we segmented the genome into smaller stretches by placing boundaries between any pair of adjacent CpGs if their distance is greater than a predefined threshold T. We call such stretch of DNA where CpGs are located on as a “ridge”. The threshold T is chosen based on the length of the probes and expected fragment lengths. Setting T=100bp, we find 2.0M ridges covering 260Mbp after removing CpGs overlapping repeat regions. Fig. 25 depicts the distribution of the resulting ridges. Depending on minimum LOD, sensitivity and specificity, not all these ridges may be useful. In order to reduce the estimation variance of any statistical model applied to the observed fragment level methylation pattern, we choose to apply minimum thresholds on the number of CpGs and the CpG density per ridge. Fig. 26 shows the survival functions of number of CpGscovered and total length of all ridges as a function of minimum number of CpGs plotted for several minimum CpG densities. Applying a minimum number of CpGs of 5 and a minimum CpG density of 0.03, we initially arrived at 482K ridges covering about 8.26M CpGs and 147.5Mbp. After filtering the ridges to mappable regions, we arrived at 328K ridges covering about 110Mbp and 5.66M CpG sites. This has coverage of the lower density regions outside CpG islands, shores and shelves, where recent studies show that most cell- / cancer-specific biomarkers reside. Table A include(s) preselected genomic regions (or ridges) on chromosomes 21 and 22 (minimum CpG density of 0.03 and minimum CpG number of 5).

[0321] A methylation biomarker discovery hybrid capture panel that covers 136 Mb of sequence space was subsequently created, including: (i) all human genomic regions with CpG density >3% and that contain at least 5 CpGs and that are not in sequence repeats except for GC- rich regions, (ii) endogenous and exogenous control sequences for measuring methylated / unmethylated cytosine protection and conversion rates, (iii) regions for estimating age, and (iv) SNP tracers for detecting sample swaps.

[0322] One example workflow using this hybrid capture panel may include the following steps: Library prep → conversion → amplification → hybrid capture → sequencing. Another example workflow using this hybrid capture panel may include the following steps: Library prep → hybrid capture → conversion → amplification → sequencing.

[0323] Using the panel of hybrid capture probes targeting all genomic regions with sufficiently high CpG number and CpG density (filtered for repeat elements other than GC-rich regions), the following workflow metrics and hybrid capture metrics were used for performance evaluation and workflow and hybridization condition optimization. Workflow metrics: (i) conversion rate >99.5%; (ii) protection rate >96%; (iii) CHG and CHH methylation < 1%; (iv) mean insert size >250 bp for tissue samples (FFPE and fresh-frozen); (v) raw coverage >200x. Hybrid capture metrics: (i) mapping rate > 70%; (ii) on-target rate >70%; (iii) target coverage uniformity: >70% of targets with a sample fold coverage of 0.8 (80% of sample mean); (iv) median target coverage >100x. These metrics may be further adjusted or improved.

[0324] Example 3

[0325] 32 samples, including cfDNA, fresh-frozen gDNA, buffy DNA and FFPE DNA (including CRC, lung, breast, and AA, see Table 3) were used for testing the methylation discovery end-to-end workflow and for the initial QC and optimization of the custom methylation discovery hybrid capture pool and the workflow.

[0326] Table 3

[0327] Figure 27 provides an overview of an exemplary experimental workflow. Three conditions were tested: (i) baseline condition (1x probe concentration, 8-plex sample plexing); (ii) lower probe concentration (0.5x probe concentration, 8-plex sample plexing); and (iii) different sample plexing levels (1x probe concentration, 4-, 6-, 10-, and 12-plex sample plexing). All samples except cfDNA were processed by UltraShear DNA fragmentation (fragment size target 250 bp). The following spike-in controls were added: 0.001% (DNA mass) unmethylated lambda and 0.00005% (DNA mass) methylated pUC19 based on 1:1 copy number.

[0328] Following QC, the samples underwent library preparation (including for example end preparation and adaptor ligation), conversion, and amplification. The amplified material from each sample was purified and QCed (LabChip and Quant-It), and used in 3 separate sample barcoding PCR reactions (BC-PCR1, BC-PCR2, BC-PCR3) with distinct condition barcodes. The BC-PCR product was purified and QCed (LabChip and Quant-It).

[0329] For the baseline condition, the 32 barcoded samples from BC-PCR1 (each carrying distinct sample barcodes and the condition barcode for the baseline condition) were combined into 4 pools of 8 samples each, using 187.5 ng barcoded product per sample. Samples of the same type are kept in each pool when possible. Specifically, pool 1 contains gDNA from 8 FFPE samples (4 CRC, 4 lung cancer); pool 2 contains gDNA from 8 FFPE samples (4 breast cancer, 4 advanced adenoma (AA)); pool 3 contains gDNA from 6 fresh-frozen (FF) samples (2 AA, 4 blood cancer) and 2 buffy coat / PBMC blood cancer samples; and pool 4 contains gDNA from 2 buffy coat / PBMC blood cancer samples and 2 cell line controls, and cfDNA from 4 healthy samples. The 4 pools were subjected to hybrid capture using the custom methylation discovery panel at 1x (baseline) probe concentration.

[0330] For the lower probe concentration condition, the 32 barcoded samples from BC- PCR2 (each carrying distinct sample barcodes and the condition barcode for the lower probe concentration condition) were combined into 4 pools of 8 samples each as described above, using 187.5 ng barcoded product per sample. The 4 pools were subjected to hybrid capture using the custom methylation discovery panel at 0.5x probe concentration.

[0331] For the sample plexing variation condition, the 32 barcoded samples from BC- PCR3 (each carrying distinct sample barcodes and the condition barcode for the sample plexing variation condition) were combined into 4 pools with variable number of samples per pool (4, 6, 10 and 12 samples respectively), using 375 ng barcoded product per sample for 4-plex, 250 ng barcoded product per sample for 6-plex; 200 ng barcoded product per sample for 10-plex; and 200 ng barcoded product per sample for 12-plex.

[0332] For the sample plexing variation experiment, samples of the same type were kept in each pool when possible. FFPE samples were tested in the 4-plex and 6-plex pools, as these are the lowest quality samples and could benefit the most from the lower plex hybrid capture condition. cfDNA (5) and fresh-frozen (5) samples were tested in the 10-plex condition, as these are the best quality samples with the highest chance of success in a high plex capture. The remaining samples were placed in the 12-plex pool (6 FFPE, 3 fresh-frozen and 3 cfDNA). The 4pools were subjected to hybrid capture using the custom methylation discovery panel at 1x probe concentration.

[0333] The captured products from the three different test conditions were amplified, pooled, and sequenced. The sequencing data were analyzed to determine Basic Sequencing QC and Hybrid Capture QC.

[0334] Basic Sequencing QC: Pre-Alignment (total number of raw sequencing reads, per base sequence quality, per base sequence content, adapter contamination); Post Alignment (mapping efficiency, read duplicate percentage, mean insert size, insert size distribution); Methylation Conversion (conversion efficiency using lambda, protection rate using pUC19, methylation profiles for all CX context including CpG, CHH, CHG).

[0335] Hybrid Capture QC: Library Complexity; Mean Target Coverage; On-Target Rate; Fraction of Duplicated Reads; Uniformity; Custom On-Target Analysis (on / off target CpG methylation levels, per probe / bait read coverage before and after read deduplication); Methylation Target Custom Analysis (effective coverage distribution / target and per sample, fraction of hyper-methylated / hypo-methylated fragments per target and per sample).

[0336] Fig.28 shows raw pairs, mapping efficiency, read duplicate fraction, and mean insert size, as grouped by sample type. Similar overall sequencing metrics were observed among all sample types.

[0337] Fig.29 shows CpG, CHG, CHH methylation levels, as well as conversion and protection rates, as grouped by sample type. As expected, more variation in human CpG methylation was observed in FFPE samples. Low CHG and CHH methylation levels were observed in all samples except FFPE, indicating good conversion rates. Higher CHG and CHH methylation levels observed in FFPE was expected due to lower sample quality compared to other sample types. cfDNA conversion and protection rates were lower due to low pUC19 and lambda coverage (spike-ins amount).

[0338] Fig.30 shows on-target rate, bait coverage, target coverage and fold-80 base penalty, as grouped by sample type. The on-target rates were very high, which relate to the large panel size of ~136 Mb.

[0339] As shown in Figs.28-30, the results are generally in line with the workflow and hybrid capture metrics outlined in Example 2 across all sample types, with the exception of fold 80 base penalty. The observed fold 80 base penalty was 2.3-3 for FFPE samples and 2.1-2.5 for other sample types, compared to the target of <2. If desired, further optimization (e.g. adjusting probe concentrations or other hybrid capture parameters) can be performed to increase coverage uniformity.

[0340] Compared to 1x probe concentrations, 0.5x probe concentrations resulted in higher coverage and better mapping efficiency but lower on-target rates (Fig.31). In terms of sample plexing levels, 8-plex, 10-plex, and 12-plex showed better mapping efficiency, higher on- target rate, but lower coverage compared to 4-plex and 6-plex.12-plex had the best mapping efficiency, 10-plex had the highest on-target rate, and 4-plex had the highest bait and target coverage compared to others (Figs.32-34). * * * *

Claims

1. WHAT IS CLAIMED IS:

1. A method of preparing a composition of non-naturally occurring DNA, comprising: (a) extracting DNA from a sample of a subject; (b) contacting the extracted DNA or its derivative with a panel of oligonucleotide probes designed to hybridize to a plurality of preselected genomic regions, thereby generating selected DNA, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome, wherein the extracted DNA or its derivative or the selected DNA or its derivative is further treated with an agent or a combination of agents that discriminates between methylated and unmethylated cytosines; and (c) performing high-throughput sequencing on the selected DNA or its derivative and generating sequencing reads.

2. The method of claim 1, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 4 CpG sites.

3. The method of claim 1, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 5 CpG sites.

4. The method of claim 1, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises at least 6 CpG sites.

5. The method of any of claims 1-4, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.03 CpG / bp.

6. The method of any of claims 1-4, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.04 CpG / bp.

7. The method of any of claims 1-4, wherein at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

8. The method of any of claims 1-7, wherein at least 80% of the preselected genomic regions are CpG regions.

9. The method of any of claims 1-7, wherein at least 90% of the preselected genomic regions are CpG regions.

10. The method of any of claims 1-7, wherein at least 95% of the preselected genomic regions are CpG regions.

11. The method of any of claims 1-10, wherein the preselected genomic regions cover 20-300 Mb of sequence space in human genome.

12. The method of any of claims 1-10, wherein the preselected genomic regions cover 50-200 Mb of sequence space in human genome.

13. The method of any of claims 1-10, wherein the preselected genomic regions cover 1-5% of sequence space in human genome.

14. The method of any of claims 1-10, wherein the preselected genomic regions cover 2-3% of sequence space in human genome.

15. The method of any of claims 1-14, wherein in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp.

16. The method of any of claims 1-14, wherein in at least 50%, at least 80%, or at least 90%, or at least 95% of the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp, or 80 bp, or 60 bp, or 50 bp.

17. The method of any of claims 1-16, further comprising fragmenting the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes.

18. The method of any of claims 1-17, further comprising appending adaptors to the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes.

19. The method of claim 18, wherein the adaptors are Y-adaptors.

20. The method of claim 18 or 19, wherein the adaptors comprise a universal priming site, a molecular barcode, and / or methylated cytosines.

21. The method of any of claims 1-20, wherein the extracted DNA or derivative thereof is treated with a deaminating agent or a combination of an oxidizing agent and a deaminating agent, that converts unmethylated cytosines to uracils before the contacting with the panel of oligonucleotide probes.

22. The method of any of claims 1-20, wherein the selected DNA or derivative thereof is treated with a deaminating agent or a combination of an oxidizing agent and a deaminating agent that converts unmethylated cytosines to uracils after the contacting with the panel of oligonucleotide probes.

23. The method of claim 21 or 22, wherein the deaminating agent comprises sodium bisulfite.

24. The method of claim 21 or 22, wherein the deaminating agent comprises a deaminase.

25. The method of any of claims 1-24, wherein the extracted DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependent restriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, before the contacting with the panel of oligonucleotide probes.

26. The method of any of claims 1-24, wherein the selected DNA or derivative thereof is treated with a methylation sensitive restriction enzyme (MSRE), methylation dependentrestriction enzymes (MDRE), a 5-methylcytosine (5mC) antibody, a 5-hydroxymethylcytosine (5hmC) antibody, a methyl-CpG-binding domain (MBD) protein, or a DNA methyltransferase, after the contacting with the panel of oligonucleotide probes.

27. The method of any of claims 1-26, further comprising amplifying the extracted DNA or derivative thereof before the contacting with the panel of oligonucleotide probes.

28. The method of any of claims 1-26, further comprising amplifying the selected DNA or derivative thereof after the contacting with the panel of oligonucleotide probes.

29. The method of claim 27 or 28, wherein the amplifying adds a sample barcode, and wherein a plurality of libraries of the selected DNA or derivative thereof each produced from a different sample are sequenced together in one sequencing lane.

30. The method of any of claims 27-29, further comprising performing size selection before or after the amplifying step.

31. The method of any of claims 1-30, wherein the oligonucleotide probe is a hybrid capture probe.

32. The method of any of claims 1-30, wherein the oligonucleotide probe comprises a probe section covalently linked to a primer section.

33. The method of any of claims 1-32, wherein the CpG regions do not comprise CpG sites located in repeat elements excluding GC-rich regions.

34. The method of any of claims 1-33, wherein the sample of DNA comprises cellular DNA derived from a fresh-frozen tissue or an FFPE tissue.

35. The method of any of claims 1-33, wherein the sample of DNA comprises cell-free DNA derived from a blood, plasma, serum, or urine sample.

36. The method of claim 34 or 35, wherein the sample of DNA comprises tumor DNA.

37. The method of any of claims 1-36, further comprising producing a set of libraries of the selected DNA from a set of samples from a population of subjects, and using the sequencing reads of the set of libraries of the selected DNA to identify preselected genomic regions of which the co-methylation patterns of the CpG sites are substantially consistent.

38. The method of any of claims 1-37, further comprising producing a first set of libraries of the selected DNA from a first set of samples from a population of diseased subjects, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-diseased subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of diseased subjects and the population of healthy or non-diseased subjects.

39. The method of any of claims 1-37, further comprising producing a first set of libraries of the selected DNA from a first set of samples from a population of subjects having cancer, producing a second set of libraries of the selected DNA from a second set of samples from a population of healthy or non-cancerous subjects, and using the sequencing reads of the first and second sets of libraries of the selected DNA to identify preselected genomic regions of which the methylation loads of the CpG sites are significantly different between the population of subjects having cancer and the population of healthy or non-cancerous subjects.

40. The method of claim 38 or 39, further comprising identifying a plurality of differentially methylated regions (DMRs) between the first set of samples and the second set of samples.

41. A composition comprising a sample of DNA that has been selectively enriched at a plurality of preselected genomic regions, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome.

42. The composition of claim 41, wherein the CpG regions each comprises at least 4 CpG sites.

43. The composition of claim 42, wherein the CpG regions each comprises at least 5 CpG sites.

44. The composition of claim 43, wherein the CpG regions each comprises at least 6 CpG sites.

45. The composition of any of claims 41-44, wherein the CpG regions each comprises a CpG density of at least 0.03 CpG / bp.

46. The composition of any of claims 41-44, wherein the CpG regions each comprises a CpG density of at least 0.04 CpG / bp.

47. The composition of any of claims 41-44, wherein the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

48. The composition of any of claims 41-47, wherein at least 80% of the preselected genomic regions are CpG regions.

49. The composition of any of claims 41-47, wherein at least 90% of the preselected genomic regions are CpG regions.

50. The composition of any of claims 41-47, wherein at least 95% of the preselected genomic regions are CpG regions.

51. The composition of any of claims 41-50, wherein the preselected genomic regions cover 20-300 Mb of sequence space in human genome.

52. The composition of any of claims 41-50, wherein the preselected genomic regions cover 50-200 Mb of sequence space in human genome.

53. The composition of any of claims 41-50, wherein the preselected genomic regions cover 1-5% of sequence space in human genome.

54. The composition of any of claims 41-50, wherein the preselected genomic regions cover 2-3% of sequence space in human genome.

55. The composition of any of claims 41-54, wherein in the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp.

56. The composition of any of claims 41-54, wherein in the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp.

57. The composition of any of claims 41-56, wherein the CpG regions substantially do not comprise CpG sites located in repeat elements excluding GC-rich regions.

58. The composition of any of claims 41-57, wherein the sample of DNA comprises uracils or thymines converted or derived from unmethylated cytosines.

59. The composition of any of claims 41-48, wherein the sample of DNA comprises a adaptor section linked to a genomic DNA section.

60. The composition of claim 59, wherein the adaptor section comprises a universal priming site, a molecular barcode, and / or methylated cytosines.

61. The composition of any of claims 41-60, wherein the sample of DNA comprises cellular DNA derived from a fresh-frozen tissue or an FFPE tissue.

62. The composition of any of claims 41-60, wherein the sample of DNA comprises cell-free DNA derived from a blood, plasma, serum, or urine sample.

63. The composition of claim 61 or 62, wherein the sample of DNA comprises tumor DNA.

64. A composition comprising a panel of oligonucleotide probes designed for selectively enriching a plurality of preselected genomic regions from a sample of DNA by hybrid capture or linked target capture, wherein at least 50% of the preselected genomic regions are CpG regions that each comprises at least 3 CpG sites and a CpG density of at least 0.02 CpG / bp, wherein the preselected genomic regions cover 10-500 Mb of sequence space in human genome.

65. The composition of claim 64, wherein the CpG regions each comprises at least 4 CpG sites.

66. The composition of claim 64, wherein the CpG regions each comprises at least 5 CpG sites.

67. The composition of claim 64, wherein the CpG regions each comprises at least 6 CpG sites.

68. The composition of any of claims 64-67, wherein the CpG regions each comprises a CpG density of at least 0.03 CpG / bp.

69. The composition of any of claims 64-67, wherein the CpG regions each comprises a CpG density of at least 0.04 CpG / bp.

70. The composition of any of claims 64-67, wherein the CpG regions each comprises a CpG density of at least 0.05 CpG / bp.

71. The composition of any of claims 64-70, wherein at least 80% of the preselected genomic regions are CpG regions.

72. The composition of any of claims 64-70, wherein at least 90% of the preselected genomic regions are CpG regions.

73. The composition of any of claims 64-70, wherein at least 95% of the preselected genomic regions are CpG regions.

74. The composition of any of claims 64-73, wherein the preselected genomic regions cover 20-300 Mb of sequence space in human genome.

75. The composition of any of claims 64-73, wherein the preselected genomic regions cover 50-200 Mb of sequence space in human genome.

76. The composition of any of claims 64-73, wherein the preselected genomic regions cover 1-5% of sequence space in human genome.

77. The composition of any of claims 64-73, wherein the preselected genomic regions cover 2-3% of sequence space in human genome.

78. The composition of any of claims 64-77, wherein in the CpG regions the maximum distance between any two adjacent CpG sites is 150 bp.

79. The composition of any of claims 64-77, wherein in the CpG regions the maximum distance between any two adjacent CpG sites is 100 bp.

80. The composition of any of claims 64-79, wherein the CpG regions substantially do not comprise CpG sites located in repeat elements excluding GC-rich regions.

81. The composition of any of claims 64-80, wherein the oligonucleotide probe is a hybrid capture probe.

82. The composition of any of claims 64-80, wherein the oligonucleotide probe comprises a probe section covalently linked to a primer section.

83. A composition comprising a panel of oligonucleotide probes designed for selectively enriching a plurality of preselected differentially methylated regions (DMRs) from a sample of DNA by hybrid capture or linked target capture, wherein the DMRs are selected according to the method of claim 40.

Citation Information

Patent Citations

  • Linked duplex target capture

    WO2017168332A1

  • Compositions, methods, and kits for isolating nucleic acids

    WO2018156418A1

  • Methylation markers and targeted methylation probe panel

    WO2020069350A1

  • Detection of genetic and epigenetic information in a single workflow

    WO2023129965A2

  • Methods for assaying circulating tumor DNA

    WO2025019370A1

Cited By

  • Methods for preparing and evaluating methylated nucleic acid molecules

    WO2026072590A1