Framework to identify genomic regions indicative of one or more biological conditions

A framework for analyzing CpG dinucleotide methylation patterns in sequencing data improves cancer detection sensitivity by identifying tumor-specific clusters and applying machine learning, addressing the challenges of low nucleic acid amounts and heterogeneity in liquid biopsies.

WO2026073140A1PCT designated stage Publication Date: 2026-04-02GUARDANT HEALTH INC +1
View PDF 24 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing cancer detection methods using liquid biopsies face challenges due to the low amount and heterogeneity of nucleic acids in body fluids, making it difficult to accurately classify samples for tumor-derived DNA with high sensitivity.

Method used

A framework is developed for analyzing sequencing data that involves calculating background and tumor signal rates, identifying CpG dinucleotide clusters with specific methylation criteria, and applying machine learning to provide indications of tumor-related conditions such as cancer type, tissue of origin, and tumor burden.

Benefits of technology

Enhances the sensitivity and accuracy of cancer detection by identifying tumor-specific methylation patterns in cell-free DNA, enabling non-invasive diagnosis with improved classification of tumor-related biological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025048500_02042026_PF_FP_ABST
    Figure US2025048500_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A method may produce a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition. A method may obtain a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, and calculate a background signal rate using the background data set. A method may obtain a tumor data set based on the predetermined methylation criteria, and may calculate tissue signal rate based on the tumor data set. A method may identify a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tissue signal cutoff value and produce the framework accordingly. A method may apply the framework to the sample data set and may provide an indication of a tumor-related biological condition.
Need to check novelty before this filing date? Find Prior Art

Description

FRAMEWORK TO IDENTIFY GENOMIC REGIONS INDICATIVE OF ONE OR MORE BIOLOGICAL CONDITIONS PRIORITY CLAIM AND CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application is related to and claims priority to U.S. provisional patent application number 63 / 700,411 filed September 27, 2024, and entitled AGGREGATING SINGLE MOLECULE METHYLATION STATES FROM CYTOSINE-GUANINE LOCI TO INCREASE SIGNAL TO NOISE RATIO, which is incorporated by reference herein in its entirety. BACKGROUND

[0002] Cancer is a major cause of disease worldwide. Each year, tens of millions of people are diagnosed with cancer around the world, and more than half eventually die from it. In many countries, cancer ranks the second most common cause of death following cardiovascular diseases. Early detection is associated with improved outcomes for many cancers.

[0003] Cancer can be caused by the accumulation of genetics variations within an individual's normal cells, at least some of which result in improperly regulated cell division. Such variations commonly include copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), epigenetic variations including 5-methylation of cytosine (5-methylcytosine) and association of DNA with chromatin and transcription factors.

[0004] Cancers are often detected by biopsies of tumors followed by analysis of cells, markers or DNA extracted from cells. But more recently it has been proposed that cancers can also be detected from cell-free nucleic acids in body fluids, such as blood or urine. Such tests have the advantage that they are noninvasive and can be performed without identifying suspected cancer cells in biopsy. However, such tests are complicated by the fact that the amount of nucleic acids in body fluids is very low and what nucleic acids are present are heterogeneous in form (e.g., RNA and DNA, single-stranded and double-stranded, and various states of post-replication modification and association with proteins, such as histones).

[0005] Thus, there is a need for improved systems and methods for improved cancer detection using liquid biopsy assays. Therefore, it is an object of the disclosure to provide computer-implemented systems and methods that have improved capability to classify a sample as containing tumor-derived DNA with heightened sensitivity.SUMMARY

[0006] In some aspects, the techniques described herein relate to a method of methylation analysis for detection of a tumor-related biological condition, the method including: producing a framework for analyzing a sample data set including sequencing data derived from a test subject for the tumor-related biological condition, producing the framework including: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data sest originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and producing the framework including the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters; applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0007] In some aspects, the techniques described herein relate to a method, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0008] In some aspects, the techniques described herein relate to a method, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0009] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0010] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0011] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0012] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0013] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0014] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0015] In some aspects, the techniques described herein relate to a method, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0016] In some aspects, the techniques described herein relate to a method, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0017] In some aspects, the techniques described herein relate to a method, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0018] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is six.

[0019] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is four.

[0020] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is three.

[0021] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0022] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0023] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0024] In some aspects, the techniques described herein relate to a method, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0025] In some aspects, the techniques described herein relate to a method, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0026] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0027] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer- associated fibroblasts present in the test subject.

[0028] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0029] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0030] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0031] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0032] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0033] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0034] In some aspects, the techniques described herein relate to a method, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0035] In some aspects, the techniques described herein relate to a method, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0036] In some aspects, the techniques described herein relate to a method, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0037] In some aspects, the techniques described herein relate to a method comprising determining that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0038] In some aspects, the techniques described herein relate to a method comprising determining a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0039] In some aspects, the techniques described herein relate to a method of methylation analysis for detection of a tumor-related biological condition, the method including: producing a framework for analyzing a sample data set including sequencing data derived from a test subject for the tumor-related biological condition, producing the framework including: obtaining a background data set including sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set including sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating a tumor signal rate based on the tumor data set; comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tumor signal rate above a tumor signal cutoff value; and producing the framework includingthe subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

[0040] In some aspects, the techniques described herein relate to a method, including: applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor- related biological condition present in the test subject based on application of the framework.

[0041] In some aspects, the techniques described herein relate to a method, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0042] In some aspects, the techniques described herein relate to a method, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0043] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0044] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0045] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0046] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0047] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0048] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0049] In some aspects, the techniques described herein relate to a method, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0050] In some aspects, the techniques described herein relate to a method, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0051] In some aspects, the techniques described herein relate to a method, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0052] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is six.

[0053] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is four.

[0054] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is three.

[0055] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0056] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0057] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0058] In some aspects, the techniques described herein relate to a method, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0059] In some aspects, the techniques described herein relate to a method, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0060] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0061] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer- associated fibroblasts present in the test subject.

[0062] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0063] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0064] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0065] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0066] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0067] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0068] In some aspects, the techniques described herein relate to a method, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0069] In some aspects, the techniques described herein relate to a method, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0070] In some aspects, the techniques described herein relate to a method, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0071] In some aspects, the techniques described herein relate to a method including determining that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0072] In some aspects, the techniques described herein relate to a method including determining a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0073] In some aspects, the techniques described herein relate to a method of methylation analysis for detection of a tumor-related biological condition, the method including: applying a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides in the sample data set at each of a subset of CpG clusters wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold; wherein applying the framework includes aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject; and providing the indication of the tumor-related biological condition present in the test subject based on application of the framework.

[0074] In some aspects, the techniques described herein relate to a method, wherein the framework is produced by: analyzing a sample data set including the sequencing data derived from the test subject for the tumor-related biological condition: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and wherein the framework includes the subset of the CpG dinucleotide clusters and methylation criteria associated with each of the clusters.

[0075] In some aspects, the techniques described herein relate to a method, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0076] In some aspects, the techniques described herein relate to a method, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0077] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0078] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0079] In some aspects, the techniques described herein relate to a method, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0080] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0081] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0082] In some aspects, the techniques described herein relate to a method, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0083] In some aspects, the techniques described herein relate to a method, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0084] In some aspects, the techniques described herein relate to a method, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0085] In some aspects, the techniques described herein relate to a method, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0086] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is six.

[0087] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is four.

[0088] In some aspects, the techniques described herein relate to a method, wherein the predetermined number of CpG dinucleotides is three.

[0089] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0090] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0091] In some aspects, the techniques described herein relate to a method, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0092] In some aspects, the techniques described herein relate to a method, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0093] In some aspects, the techniques described herein relate to a method, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0094] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0095] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer- associated fibroblasts present in the test subject.

[0096] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0097] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0098] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0099] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0100] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0101] In some aspects, the techniques described herein relate to a method, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0102] In some aspects, the techniques described herein relate to a method, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0103] In some aspects, the techniques described herein relate to a method, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0104] In some aspects, the techniques described herein relate to a method, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0105] In some aspects, the techniques described herein relate to a method including determining that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0106] In some aspects, the techniques described herein relate to a method including determining a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0107] In some aspects, the techniques described herein relate to a system to perform methylation analysis for detection of a tumor-related biological condition, the system including one or more hardware processors and memory comprising computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: producing a framework for analyzing a sample data set including sequencing data derived from a test subject for the tumor-related biological condition, producing the framework including: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and producing the framework including the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters; applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0108] In some aspects, the techniques described herein relate to a system, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0109] In some aspects, the techniques described herein relate to a system, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0110] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0111] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0112] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0113] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0114] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0115] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0116] In some aspects, the techniques described herein relate to a system, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0117] In some aspects, the techniques described herein relate to a system, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0118] In some aspects, the techniques described herein relate to a system, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0119] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is six.

[0120] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is four.

[0121] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is three.

[0122] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0123] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0124] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0125] In some aspects, the techniques described herein relate to a system, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0126] In some aspects, the techniques described herein relate to a system, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0127] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0128] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer- associated fibroblasts present in the test subject.

[0129] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0130] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0131] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0132] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0133] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0134] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0135] In some aspects, the techniques described herein relate to a system, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0136] In some aspects, the techniques described herein relate to a system, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0137] In some aspects, the techniques described herein relate to a system, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0138] In some aspects, the techniques described herein relate to a system that determines that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0139] In some aspects, the techniques described herein relate to a system that determines a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0140] In some aspects, the techniques described herein relate to a system to perform methylation analysis for detection of a tumor-related biological condition, the system including one or more hardware processors and memory comprising computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: producing a framework for analyzing a sample data set including sequencing data derived from a test subject for the tumor-related biological condition, producing the framework including: obtaining a background data set including sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one ormore first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set including sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating a tumor signal rate based on the tumor data set; comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tumor signal rate above a tumor signal cutoff value; and producing the framework including the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

[0141] In some aspects, the techniques described herein relate to a system, including: applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0142] In some aspects, the techniques described herein relate to a system, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0143] In some aspects, the techniques described herein relate to a system, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0144] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0145] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0146] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0147] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0148] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0149] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0150] In some aspects, the techniques described herein relate to a system, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0151] In some aspects, the techniques described herein relate to a system, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0152] In some aspects, the techniques described herein relate to a system, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0153] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is six.

[0154] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is four.

[0155] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is three.

[0156] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0157] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0158] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0159] In some aspects, the techniques described herein relate to a system, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0160] In some aspects, the techniques described herein relate to a system, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0161] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0162] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer- associated fibroblasts present in the test subject.

[0163] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0164] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0165] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0166] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0167] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0168] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0169] In some aspects, the techniques described herein relate to a system, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0170] In some aspects, the techniques described herein relate to a system, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0171] In some aspects, the techniques described herein relate to a system, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0172] In some aspects, the techniques described herein relate to a system that determines that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0173] In some aspects, the techniques described herein relate to a system that determines a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0174] In some aspects, the techniques described herein relate to a system to perform methylation analysis for detection of a tumor-related biological condition, the system including one or more hardware processors and memory comprising computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: applying a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides in the sample data set at each of a subset of CpG clusters wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold; wherein applying the framework includes aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject; and providing the indication of the tumor-related biological condition present in the test subject based on application of the framework.

[0175] In some aspects, the techniques described herein relate to a system, wherein the framework is produced by: analyzing a sample data set including the sequencing data derived from the test subject for the tumor-related biological condition: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and wherein the framework includes the subset of the CpG dinucleotide clusters and methylation criteria associated with each of the clusters.

[0176] In some aspects, the techniques described herein relate to a system, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0177] In some aspects, the techniques described herein relate to a system, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0178] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0179] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0180] In some aspects, the techniques described herein relate to a system, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0181] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0182] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0183] In some aspects, the techniques described herein relate to a system, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0184] In some aspects, the techniques described herein relate to a system, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0185] In some aspects, the techniques described herein relate to a system, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0186] In some aspects, the techniques described herein relate to a system, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0187] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is six.

[0188] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is four.

[0189] In some aspects, the techniques described herein relate to a system, wherein the predetermined number of CpG dinucleotides is three.

[0190] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0191] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0192] In some aspects, the techniques described herein relate to a system, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0193] In some aspects, the techniques described herein relate to a system, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0194] In some aspects, the techniques described herein relate to a system, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0195] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0196] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer- associated fibroblasts present in the test subject.

[0197] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0198] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0199] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0200] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0201] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0202] In some aspects, the techniques described herein relate to a system, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0203] In some aspects, the techniques described herein relate to a system, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0204] In some aspects, the techniques described herein relate to a system, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0205] In some aspects, the techniques described herein relate to a system, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0206] In some aspects, the techniques described herein relate to a system that determines that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0207] In some aspects, the techniques described herein relate to a system that determines a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0208] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable storage media storing computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: producing a framework for analyzing a sample data set including sequencing data derived from a test subject for the tumor-related biological condition, producing the framework including: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and producing the framework including the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters; applyingthe framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor- related biological condition present in the test subject based on application of the framework.

[0209] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0210] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0211] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0212] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0213] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0214] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0215] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0216] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0217] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0218] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0219] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0220] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is six.

[0221] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is four.

[0222] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is three.

[0223] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0224] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0225] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0226] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0227] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0228] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0229] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer-associated fibroblasts present in the test subject.

[0230] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0231] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0232] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0233] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0234] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0235] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0236] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0237] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0238] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0239] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable media that determine that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0240] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable media that determine a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0241] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media to perform methylation analysis for detection of a tumor-related biological condition, the one or more non-transitory computer readable storage media storing computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: producing a framework for analyzing a sample data set including sequencing data derived from a test subject for the tumor-related biological condition, producing the framework including: obtaining a background data set including sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one or more first subjects in which a tumor is not detected;calculating a background signal rate using the background data set; obtaining a tumor data set including sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating a tumor signal rate based on the tumor data set; comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tumor signal rate above a tumor signal cutoff value; and producing the framework including the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

[0242] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, including: applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0243] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0244] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0245] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0246] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0247] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0248] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0249] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0250] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0251] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0252] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0253] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0254] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is six.

[0255] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is four.

[0256] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is three.

[0257] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0258] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0259] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0260] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0261] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework includes aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0262] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0263] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer-associated fibroblasts present in the test subject.

[0264] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0265] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0266] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0267] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0268] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0269] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0270] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0271] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0272] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0273] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable media that determine that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0274] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable media that determine a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0275] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media to perform methylation analysis for detection of a tumor-related biological condition, the one or more non-transitory computer readable storage media storing computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: applying a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides in the sample data set at each of a subset of CpG clusters wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold; wherein applying the framework includes aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject; and providing the indication of the tumor-related biological condition present in the test subject based on application of the framework.

[0276] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the framework is produced by: analyzing a sample data set including the sequencing data derived from the test subject for the tumor-related biological condition: obtaining a background data set indicating a measure of methylated cytosine- guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and wherein the framework includes the subset of the CpG dinucleotide clusters and methylation criteria associated with each of the clusters.

[0277] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein calculating the background signal rate includes: determining, based on the background data set, a first quantitative measurement of CpGdinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0278] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein determining the tumor signal rate includes: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0279] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

[0280] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes sequence representations that indicate a gain in methylation.

[0281] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined methylation criteria includes sequence representations that indicate a loss in methylation.

[0282] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes three CpG dinucleotides.

[0283] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes four CpG dinucleotides.

[0284] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the number of CpG dinucleotides that satisfy the methylation state includes five CpG dinucleotides.

[0285] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the background signal cutoff value includes 1e-6 to 0.01.

[0286] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the tumor signal cutoff value includes 0.05 to 1.00.

[0287] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein each of the subset of the CpG dinucleotide clusters includes no more than a predetermined number of CpG dinucleotides.

[0288] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is six.

[0289] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is four.

[0290] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the predetermined number of CpG dinucleotides is three.

[0291] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

[0292] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

[0293] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein producing the framework further includes classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0294] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework includes determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

[0295] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework includesaggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

[0296] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

[0297] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer-associated fibroblasts present in the test subject.

[0298] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

[0299] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

[0300] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

[0301] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

[0302] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

[0303] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

[0304] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

[0305] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

[0306] In some aspects, the techniques described herein relate to one or more non- transitory computer readable storage media, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters includes using a machine learning model.

[0307] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable media that determine that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

[0308] In some aspects, the techniques described herein relate to one or more non- transitory computer-readable media that determine a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

[0309] In some aspects, results of the methods disclosed herein are used as an input to generate a report. The report may be in a paper or electronic format. For example, the determined likelihood of whether a subject has a disease, or information derived therefrom, can be displayed directly in such a report. Alternatively, or additionally, diagnostic information or therapeutic recommendations which are at least in part based on the methods disclosed herein can be included in the report.

[0310] The various steps of the methods disclosed herein may be carried out at the same or different times, in the same or different geographical locations, e.g. countries, and / or by the same or different people.

[0311] The disclosed methods can be combined with analysis of one or more additional biomarkers. In some embodiments, the disclosed methods are combined with one or more methods, such as but not limited to, methods for assessing DNA methylation patterns, DNA mutations (such as somatic mutations), nucleic acid fragmentation patterns, non-coding RNA (such as micro RNAs (miRNAs), ribosomal RNAs, transfer RNAs, small nucleolar RNAs (snow RNAs), and / or small nuclear RNAs (snRNAs)) levels, and / or cell type proportions / levels, cellular locations, and / or structural modifications of one or more proteins (such as in a sample from a subject), and / or levels or abundance or proportions of one or more metabolites or metabolitesignatures, and / or levels or abundance or proportions of one or more lipid molecules or lipid signatures, and / or the levels or proportion or abundance of one or more proteins and / or nucleic acids associated extracellular vesicles (e.g., exosomes). In some embodiments, the disclosed methods are combined with one or more analyses of genetic variations including mutations, rare mutations, indels, rearrangements, copy number variations, transversions, translocations, recombinations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and / or abnormal changes in nucleic acid 5-methylcytosine.

[0312] Additional advantages will be set forth in part in the description which follows or may be learned by practice. The advantages will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0313] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain implementations, and together with the written description, serve to explain certain principles of the methods, computer readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings which are included by way of example and not by way of limitation. It will be understood that like reference numerals identify like components throughout the drawings, unless the context indicates otherwise. It will also be understood that some or all of the Figures may be schematic representations for purposes of illustration and do not necessarily depict the actual relative sizes or locations of the elements shown.

[0314] FIG.1 comprises a block diagram of a system for determining an indication of the amount or type of tumor molecules in an example.

[0315] FIG. 2 comprises a block diagram of a method of a method of determining an indication of the amount or type of tumor molecules in an example.

[0316] FIG.3 comprises a diagram of selecting CpG clusters in an example.

[0317] FIG.4 comprises a flow chart of a method of providing an indication of a tumor- related biological condition in an example.

[0318] FIG.5 illustrates a framework for generating data that is analyzed by one or more computational models to determine one or more biological condition indicators

[0319] FIG.6 is a block diagram illustrating components of a machine, according to some example implementations, able to read instructions from a machine-readable medium (e.g., a machine-readable storage medium) and perform any one or more of the methodologies discussed herein.

[0320] FIG.7 is a block diagram illustrating system that includes an example software architecture, which may be used in conjunction with various hardware architectures herein described.

[0321] FIG. 8 is a graphical representation of background rate of CpG clusters in an example.

[0322] FIG. 9 is a graphical representation of background rate of CpG clusters in an example.

[0323] FIG.10 is a graphical representation of CpG cluster selection in an example.

[0324] FIG.11 is a graphical representation of signal molecules observed as a function of background rate cutoff.

[0325] FIG.12 is a graphical representation of region selection and molecule counting in an example.

[0326] FIG.13 shows the LOD (limit of detection) 98 that was determined for selection frameworks 1, 2, 3, and 4.

[0327] FIG.14 shows the receiver operating characteristic (ROC) curves with respect to specificity and sensitivity for selection frameworks 1, 2, 3, and 4.

[0328] FIG.15 shows scaled 50-percentile scores of batches using selection frameworks 1, 2, 3, and 4.

[0329] FIG.16 shows scaled 80-percentile scores of batches using selection frameworks 1, 2, 3, and 4. DETAILED DESCRIPTION

[0330] Discussed herein is a method of detecting tumor molecules based on methylation data. In this method, cytosine-guanine (CpG) clusters are identified in cell-free DNA (cfDNA)samples and used to determine methylation data, which can be correlated to tumor detection. This method of using CpG clusters helps select regions with wide-spread gain of methylation in tumor tissue but also have low background signal rate. Using CpG clusters helps select the best regions for methylation analysis.

[0331] Specifically, methods involve defining a CpG cluster as a group of consecutive CpG loci (denoted as “K”), and classifying molecules based on the majority methylation state within these clusters. For instance, a cluster might be considered a signal molecule if at least three out of four CpGs are in a methylated state. This method can be useful in distinguishing between nearly fully-methylated states and background noise, which improves as the number of CpGs in the cluster increases. However, larger clusters also result in reduced coverage in less CpG-dense regions since only molecules containing the entire CpG cluster are used.

[0332] This method can be employed in the detection of tumor-derived molecules in cfDNA, particularly focusing on regions that gain methylation in tumor tissue. This approach allows for the precise selection of genomic regions with significant methylation changes in tumors, which can be crucial for early cancer detection and monitoring

[0333] The methods discussed herein address the challenges associated with accurately detecting and classifying tumor molecules based on methylation patterns. For instance, the methods discussed herein help in differentiation of methylation states; that is, the methods can be leveraged to help distinguish between fully-methylated, partially-methylated, and unmethylated states. The methods can improve the signal-to-noise ratio in such methylation analysis, which is beneficial to accurately identify tumor-derived molecules against background data.

[0334] The methods discussed herein can also help in analysis and coverage of less CpG dense regions. As the number of CpGs in a cluster increases (e.g., the larger the “K” value), the methods described herein can help optimize the size of those CpG clusters to balance clarity and coverage in the analysis. The method also addresses the variability of methylation signal across different normal cfDNA samples. By analyzing clusters of CpG loci, the method allows for a more robust and reliable identification of abnormal methylation patterns indicative of cancer.

[0335] Another challenge addressed herein is selecting genomic regions that are most indicative of cancer when they exhibit methylation changes. The method uses a strategic approach to select CpG clusters that show a significant gain in methylation in tumor tissues butmaintain low background signals in normal conditions, enhancing the specificity and sensitivity of cancer detection.

[0336] Additionally, cancer and other diseases can cause complex changes in DNA methylation patterns. The method provides a flexible framework that can be adapted to detect various methylation changes, not just gains in methylation, thereby broadening its applicability to different types of cancer and potentially other diseases. Overall, the methods herein provide a sophisticated method for analyzing methylation patterns in cfDNA, addressing key technical challenges in the field of non-invasive cancer diagnostics and improving the reliability and accuracy of such tests.

[0337] The discussed methods provide several advantages in the context of cancer detection and monitoring, some of which are unexpected. First, the methods discussed herein provide an improved signal-to-noise ratio. By using larger clusters of CpG loci (higher “K” values), the method enhances the ability to distinguish almost fully-methylated states from background noise. This is helpful for accurately identifying tumor-derived molecules in cfDNA.

[0338] The method also allows for the precise classification of molecules based on their methylation status at multiple consecutive CpG sites. This precision is vital for detecting subtle methylation patterns that are indicative of tumor presence. Although the primary focus is on detecting regions that gain methylation in tumor tissues, the method is adaptable to other methylation patterns. This includes regions that lose methylation in cancer tissues, making it versatile for various types of cancer and potentially other diseases involving methylation changes.

[0339] Moreover, the method facilitates the selection of genomic regions with widespread methylation changes in tumor tissue while maintaining low background signal rates in normal conditions. This selective capability is helpful for identifying robust biomarkers for cancer detection. The method can be applied to a variety of sample types, including cancer-free, AA (Advanced Adenoma), and CRC (colorectal cancer) samples. This demonstrates the method's utility across different clinical scenarios and sample conditions. By detecting methylation changes in cfDNA, the method offers a non-invasive approach to cancer detection and monitoring. This is particularly valuable for early detection and for tracking the effectiveness of treatments.

[0340] Overall, the methods discussed herein allow for a sophisticated take on methylation analysis in cfDNA, which helps in enhancing overall accuracy and effectiveness of cancer diagnostics and associated medical strategies.Definitions

[0341] In order for the present disclosure to be more readily understood, certain terms are first defined below. Additional definitions for the following terms and other terms may be set forth through the specification. If a definition of a term set forth below is inconsistent with a definition in an application or patent that is incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.

[0342] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, a reference to “a method” includes one or more methods, and / or steps of the type described herein and / or which will become apparent to those persons of ordinary skill in the art upon reading this disclosure and so forth.

[0343] It is also to be understood that the terminology used herein is for the purpose of describing particular implementations only, and is not intended to be limiting. Further, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer readable media, and systems, the following terminology, and grammatical variants thereof, will be used in accordance with the definitions set forth below.

[0344] About: As used herein, “about” or “approximately” as applied to one or more values or elements of interest, refers to a value or element that is similar to a stated reference value or element. In certain implementations, the term “about” or “approximately” refers to a range of values or elements that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value or element unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value or element).

[0345] Administer: As used herein, “administer” or “administering” a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means to give, apply or bring the composition into contact with the subject. Administration can be accomplished by any of a number of routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal and intradermal.

[0346] Adapter: As used herein, “adapter” refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that can be at least partially double-stranded and used to link to either or both ends of a given sample nucleic acid molecule. Adapters can include nucleic acid primer binding sites to permit amplification of a nucleic acid molecule flanked by adapters at both ends, and / or a sequencing primer binding site, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. Adapters can also include binding sites for capture probes, such as an oligonucleotide attached to a flow cell support or the like. Adapters can also include a nucleic acid tag as described herein. Nucleic acid tags can be positioned relative to amplification primer and sequencing primer binding sites, such that a nucleic acid tag is included in amplicons and sequence reads of a given nucleic acid molecule. The same or different adapters can be linked to the respective ends of a nucleic acid molecule. In some implementations, the same adapter is linked to the respective ends of the nucleic acid molecule except that the nucleic acid tag differs. In some implementations, the adapter is a Y-shaped adapter in which one end is blunt ended or tailed as described herein, for joining to a nucleic acid molecule, which is also blunt ended or tailed with one or more complementary nucleotides. In still other example implementations, an adapter is a bell-shaped adapter that includes a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other examples of adapters include T- tailed and C-tailed adapters.

[0347] Alignment: As used herein, “alignment” or “align” refers to determining whether at least two sequence representations have at least a threshold amount of homology. In one or more examples, the threshold amount of homology can be at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9%. In situations where two sequence representations have at least the threshold amount of homology, the two sequence representations can be referred to as being “aligned.”

[0348] Amplify: As used herein, “amplify” or “amplification” in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of the polynucleotide, starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification products or amplicons are generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.

[0349] Barcode: As used herein, “barcode” or “molecular barcode” in the context of nucleic acids refers to a nucleic acid molecule comprising a sequence that can serve as a molecular identifier. For example, individual "barcode" sequences can be added to each DNA fragment during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted before the final data analysis.

[0350] Cancer Type: As used herein, “cancer type” refers to a type or subtype of cancer defined, e.g., by histopathology. Cancer type can be defined by any conventional criterion, such as on the basis of occurrence in a given tissue (e.g., blood cancers, central nervous system (CNS), brain cancers, lung cancers (small cell and non-small cell), skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, breast cancers, prostate cancers, ovarian cancers, lung cancers, intestinal cancers, soft tissue cancers, neuroendocrine cancers, gastroesophageal cancers, head and neck cancers, gynecological cancers, colorectal cancers, urothelial cancers, solid state cancers, heterogeneous cancers, homogenous cancers), unknown primary origin and the like, and / or of the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma) and / or cancers exhibiting cancer markers, such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptor and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether of primary or secondary origin.

[0351] Carrier Signal: As used herein, “carrier signal” refers to any intangible medium that is capable of storing, encoding, or carrying transitory or non-transitory instructions 702 for execution by the machine 700, and includes digital or analog communications signals or other intangible medium to facilitate communication of such instructions 702. Instructions 702 may be transmitted or received over the network 734 using a transitory or non-transitory transmission medium via a network interface device and using any one of a number of well-known transfer protocols.

[0352] Cell-Free Nucleic Acid: As used herein, “cell-free nucleic acid” refers to nucleic acids not contained within or otherwise bound to a cell or, in some implementations, nucleic acids remaining in a sample following the removal of intact cells. Cell-free nucleic acids can include, for example, all non-encapsulated nucleic acids sourced from a bodily fluid (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acids include DNA(cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or a hybrid thereof. A cell-free nucleic acid can be released into bodily fluid through secretion or cell death processes, e.g., cellular necrosis, apoptosis, or the like. Some cell-free nucleic acids are released into bodily fluid from cancer cells, e.g., circulating tumor DNA (ctDNA). Others are released from healthy cells. CtDNA can be non-encapsulated tumor-derived fragmented DNA. A cell-free nucleic acid can have one or more epigenetic modifications, for example, a cell-free nucleic acid can be acetylated, 5-methylated, ubiquitylated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.

[0353] Cell Type: As used herein, a “cell type” is a set of cells having a shared characteristic. For example, immune cell types can include immune cells of different origins, differentiation types, different activation types, or any combination of different origins, different differentiation types, and different activation types. Indeed, differentiation status and activation status can overlap and often change together in a given immune cell. For example, activation of an immune cell may induce differentiation of the cell. Immune cells of different activation types can include activated cells (such as cells activated by inflammatory cytokines or antigens), suppressive cells (such as T regulatory cells (Tregs), M2 macrophages, and others, or their subsets), or suppressed cells, such as cells suppressed by Tregs. Exemplary immune cell types include, but are not limited to, macrophages (including M1 macrophages and M2 macrophages); activated B cells (including regulatory B cells, memory B cells, and plasma cells); T cell subsets, such as CD4 central memory T cells, CD8 central memory T cells, naïve-like T cells, naïve T cells, and activated T cells (including cytotoxic T cells, regulatory T cells (Tregs), CD4 effector memory T cells, and CD8 effector memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), low-density neutrophils, immature neutrophils, and immature granulocytes); and natural killer (NK) cells. Additional exemplary immune cell types include neutrophils, lymphocytes, plasma cells, monocytes, macrophages, dendritic cells, mast cells, eosinophils, T cells, CD4+ T cells, B cells, megakaryocytes, CD8+ central memory cells, CD4+ central memory cells, precursor B cells, plasma cells, memory-switched B cells, plasma cells, basophils, naïve B cells, memory B cells, CD8+ T cells, naïve CD4+ T cells, resting CD4+ memoryT cells, activated CD4+ memory T cells, follicular helper T cells, gamma delta T cells, resting NK cells; activated NK cells, M0 macrophages, resting dendritic cells, activated dendritic cells, resting mast cells, and activated mast cells. In some embodiments, cell types may be distinguished based on characteristics such as one or more cell surface markers, a genetic signature (such as expression (or expression level) of a particular gene or set of genes). Cell types can also correspond to cells extracted from the same organ, tissue, or body site. For example, lung cells may include lung alveolar epithelial cells, lung bronchial epithelial cells, or a combination of both alveolar and bronchial epithelial cells”.

[0354] Cellular Nucleic Acids: As used herein, “cellular nucleic acids” means nucleic acids that are disposed within one or more cells at least at the point a sample is taken or collected from a subject, even if those nucleic acids are subsequently removed as part of a given analytical process.

[0355] Classification Region: As used herein, “classification region” refers to a genomic region that may show sequence-independent changes in neoplastic cells (e.g., tumor cells and cancer cells) or that may show sequence-independent changes in cfDNA from subjects having cancer relative to cfDNA from subjects in which cancer is not present. Examples of sequence- independent changes include, but are not limited to, changes in methylation rate (increases or decreases), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. In one or more examples, sequence-independent changes in a classification region can indicate the presence of a single form of cancer in a subject. In one or more additional examples, sequence-independent changes in a classification region can correspond to the presence of multiple forms in a subject. The classification region can be enriched by one or more probes. In addition, the classification region can be defined by a pair of primer binding sites. Further, the classification region can be defined by a predetermined beginning genomic locus and a predetermined ending genomic locus. The classification region can include from about 25 nucleotides to about 250 nucleotides, from about 50 nucleotides to about 200 nucleotides, or from about 75 nucleotides to about 150 nucleotides. For instance, classification region can be a differentially methylated region. “Differentially methylated region” or “DMR” refers to a region of DNA having a detectably different degree of methylation in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type; or having a detectably different degree of methylation in at leastone cell or tissue type obtained from a subject having a disease or disorder relative to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region / hypermethylated target region) in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region / hypomethylated target region) in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, the classification regions comprise hypermethylated target regions and / or hypomethylated target regions.

[0356] Communications Network: As used herein, “communications network” refers to one or more portions of a network 114, 1034 that may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi® network, another type of network, or a combination of two or more such networks. For example, a network 114, 1034 or a portion of a network may include a wireless or cellular network and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) standard, others defined by various standard setting organizations, other long range protocols, or other data transfer technology.

[0357] Confidence Interval: As used herein, “confidence interval” means a range of values so defined that there is a specified probability that the value of a given parameter lies within that range of values.

[0358] Control Sample: As used herein, “control sample” or “reference sample” refers to a sample obtained from individuals without known copy number variation.

[0359] Coverage: As used herein, “coverage” or “coverage metrics” refer to the number of nucleic acid molecules or sequencing reads that correspond to a particular genomic region of a reference sequence.

[0360] Deoxyribonucleic Acid or Ribonucleic Acid: As used herein, “deoxyribonucleic acid” or “DNA” refers to a natural or modified nucleotide which has a hydrogen group at the 2′- position of the sugar moiety. DNA can include a chain of nucleotides comprising four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, “ribonucleic acid” or “RNA” refers to a natural or modified nucleotide which has a hydroxyl group at the 2′-position of the sugar moiety. RNA can include a chain of nucleotides comprising four types of nucleotides: A, uracil (U), G, and C. As used herein, the term “nucleotide” refers to a natural nucleotide or a modified nucleotide. Certain pairs of nucleotides specifically bind to one another in a complementary fashion (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand made up of nucleotides that are complementary to those in the first strand, the two strands bind to form a double strand. As used herein, “nucleic acid sequencing data”, “nucleic acid sequencing information”, “sequence information”, “sequence representation”, “nucleic acid sequence”, “nucleotide sequence”, “genomic sequence”, “genetic sequence”, “fragment sequence”, “sequencing read”, or “nucleic acid sequencing read” denotes any information or data that is indicative of the order and identity of the nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment) of a nucleic acid such as DNA or RNA. It should be understood that the present teachings contemplate sequence information obtained using all available varieties of techniques, platforms or technologies, including, but not limited to: capillary electrophoresis, microarrays, ligation-based systems, polymerase-basedsystems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.

[0361] Differentially Methylated Region: As used herein, differentially methylated region” refers to a region of DNA having a detectably different degree of methylation in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type; or having a detectably different degree of methylation in at least one cell or tissue type obtained from a subject having a disease or disorder relative to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type, such as at least one immune cell type, relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region) in at least one cell or tissue type, such as at least one immune cell type, relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject.

[0362] Driver Mutation: As used herein, “driver mutation” means a mutation that drives cancer progression.

[0363] Epigenetic Target Regions: As used herein, “epigenetic target regions” refers to target regions that may show sequence-independent differences in different cell or tissue types (e.g., different types of immune cells) or in neoplastic cells (e.g., tumor cells and cancer cells) relative to normal cells; or that may show sequence- independent differences (i.e., in which there is no change to the nucleotide sequence, e.g., differences in methylation, nucleosome distribution, or other epigenetic features) in DNA, such as cfDNA, from different cell types or from subjects having cancer relative to DNA, such as cfDNA, from healthy subjects, or in cfDNA originating from different cell or tissue types that ordinarily do not substantially contribute to cfDNA (e.g., immune, lung, colon, etc.) relative to background cfDNA (e.g., cfDNA that originated from hematopoietic cells). Examples of sequence-independent changes include, but are not limited to, changes in methylation (increases or decreases), nucleosome distribution, cfDNA fragmentation patterns,CCCTC-binding factor (“CTCF”) binding, transcription start sites (e.g., with respect to any one of more of binding of RNA polymerase components, binding of regulatory proteins, fragmentation characteristics, and nucleosomal distribution), and regulatory protein binding regions. Epigenetic target region sets thus include, but are not limited to, hypermethylation target region sets, hypomethylation target region sets, and fragmentation variable target region sets, such as CTCF binding sites and transcription start sites. For present purposes, loci susceptible to neoplasia-, tumor-, or cancer-associated focal amplifications and / or gene fusions may also be included in an epigenetic target region set because detection of a change in copy number by sequencing or a fused sequence that maps to more than one locus in a reference genome tends to be more similar to detection of exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions, or deletions, e.g., in that the focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing because their detection does not depend on the accuracy of base calls at one or a few individual positions. An epigenetic target region set is a set of epigenetic target regions.

[0364] Hypermethylation: As used herein, “hypermethylation” refers to an increased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypermethylated DNA can include DNA molecules comprising at least 1 methylated cytosine, at least 2 methylated cytosines, at least 3 methylated cytosines, at least 5 methylated cytosines, or at least 10 methylated cytosines.

[0365] Hypomethylation: As used herein, “hypomethylation” refers to a decreased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules comprising 0 methylated cytosine, at most 1 methylated cytosine, at most 2 methylated cytosines, at most 3 methylated cytosines, at most 4 methylated cytosines, or at most 5 methylated cytosines.

[0366] Immunotherapy: As used herein, “immunotherapy” refers to treatment with one or more agents that act to stimulate the immune system so as to kill or at least to inhibit growth of cancer cells, and preferably to reduce further growth of the cancer, reduce the size of the cancer and / or eliminate the cancer. Some such agents bind to a target present on cancer cells; somebind to a target present on immune cells and not on cancer cells; some bind to a target present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of pathways of the immune system that maintain self-tolerance and modulate the duration and amplitude of physiological immune responses in peripheral tissues to minimize collateral tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252–264 (2012)). Example agents include antibodies against any of PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. Other example agents include proinflammatory cytokines, such as IL-1β, IL-6, and TNF- α. Other example agents are T-cells activated against a tumor, such as T-cells activated by expressing a chimeric antigen targeting a tumor antigen recognized by the T-cell.

[0367] Indel: As used herein, “indel” refers to a mutation that involves the insertion or deletion of nucleotides in the genome of a subject.

[0368] Limit of Detection (LoD): As used herein, “limit of detection” means the smallest amount of a substance (e.g., a nucleic acid) in a sample that can be measured by a given assay or analytical approach.

[0369] Machine-Readable Medium: As used herein, “machine-readable medium” refers to a component, device, or other tangible media able to store instructions 702 and data temporarily or permanently and may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage (e.g., erasable programmable read-only memory (EEPROM)) and / or any suitable combination thereof. The term "machine-readable medium" may be taken to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) able to store instructions 702. The term "machine-readable medium" shall also be taken to include any medium, or combination of multiple media, that is capable of storing instructions 702 (e.g., code) for execution by a machine 700, such that the instructions 702, when executed by one or more processors 704 of the machine 700, cause the machine 700 to perform any one or more of the methodologies described herein. Accordingly, a "machine-readable medium" refers to a single storage apparatus or device, as well as "cloud-based" storage systems or storage networks that include multiple storage apparatus or devices. The term "machine- readable medium" excludes signals per se.

[0370] Maximum MAF: As used herein, “maximum MAF” or “max MAF” refers to the maximum MAF (mutant allele fraction) of all somatic variants in a sample.

[0371] Methylation: As used herein, “methylation” or “DNA methylation” refers to addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to addition of a methyl group to a cytosine at a CpG site (cytosine-phosphate- guanine site (i.e., a cytosine followed by a guanine in a 5’ ^ 3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to addition of a methyl group to adenine, such as in N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of the 6-carbon ring of cytosine). In some embodiments, 5- methylation refers to addition of a methyl group to the 5C position of the cytosine to create 5- methylcytosine (5mC). In some embodiments, methylation comprises a derivative of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5- formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation comprises addition of a methyl group to the 3C position of the cytosine to generate 3-methylcytosine (3mC). Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.

[0372] Methylation-Dependent Nuclease: As used herein, “methylation-dependent nuclease” refers to a nuclease that preferentially cuts methylated DNA relative to unmethylated DNA. For example, a methylation-dependent nuclease may cut at or near a recognition sequence such as a restriction site in a manner dependent on methylation of at least one of the nucleobases in the recognition sequence, such as a cytosine. In some embodiments, the nucleolytic activity of the methylation-dependent nuclease is at least 10, 20, 50, or 100-fold higher on a methylated recognition site relative to an unmethylated control in a standard nucleolysis assay. Methylation- dependent nucleases include methylation-dependent restriction enzymes.

[0373] Methylation-Dependent Restriction Enzyme: As used herein, “methylation- dependent restriction enzyme” or “MDRE” refers to a restriction enzyme that is dependent on methylation of the DNA (e.g. cytosine methylation) i.e., the presence or absence of methyl group in a nucleotide base alters the rate at which the enzyme cleaves the target DNA. In some embodiments, the methylation dependent restriction enzymes do not cleave the DNA if a particular nucleotide base is unmethylated at the recognition sequence. For example, MspJI is a methylation dependent restriction enzyme with a recognition sequence “mCNNR(N9)” and it does not cleave DNA if the absence of the methylated cytosine (mC) in the recognition sequence.

[0374] Methylation-Sensitive Nuclease: As used herein, “methylation-sensitive nuclease” refers to a nuclease that preferentially cuts unmethylated DNA relative to methylated DNA. For example, a methylation-sensitive nuclease may cut at or near a recognition sequence such as a restriction site in a manner dependent on lack of methylation of at least one of the nucleobases in the recognition sequence, such as a cytosine. In some embodiments, the nucleolytic activity of the methylation-sensitive nuclease is at least 10, 20, 50, or 100-fold higher on an unmethylated recognition site relative to a methylated control in a standard nucleolysis assay. Methylation-sensitive nucleases include methylation- sensitive restriction enzymes.

[0375] Methylation Sensitive Restriction Enzyme: As used herein, “methylation sensitive restriction enzyme” or “MSRE” refers to a restriction enzyme that is sensitive to the methylation status of the DNA (e.g. cytosine methylation) i.e., the presence or absence of methyl group in a nucleotide base alters the rate at which the enzyme cleaves the target DNA. In some embodiments, the methylation sensitive restriction enzymes do not cleave the DNA if a particular nucleotide base is methylated at the recognition sequence. For example, HpaII is a methylation sensitive restriction enzyme with a recognition sequence “CCGG” and it does not cleave DNA if the second cytosine in the recognition sequence is methylated.

[0376] Methylation rate: As used herein, “methylation rate” refers to the probability, likelihood, or percentage that a given base (for example: cytosine residue in a CpG) is methylated on a DNA molecule at a particular genomic region analyzed in the sample. In some embodiments, the methylation rate may be applied to a defined region that comprises one or more potentially methylated bases. In some embodiments, the methylation rate refers to the percentage of CpG residues methylated in a DNA molecule. In some embodiments, the methylation rate refers to the percentage of CpG residues methylated in molecules aligned to particular genomic position orgenomic region. Methylation rate can be measured by a variety of methods including, but not limited to, either using bisulfite sequencing (any single base resolution like TAPS, EM-SEQ, etc.) or using partitioning (DNA molecule resolution). Methylation rate can be measured in different ways. One estimation can be by counting how many DNA fragments end up in each methylation dependent partition or by counting the number of converted CpGs per fragment in the case of bisulfite sequencing or any other base-level resolution sequencing methods. In addition, in the case of methylation dependent partitioning, the rate calculation can be normalized using a set of predefined regions with known methylation state (i.e., positive control regions and / or negative control regions) or spiked- in synthetic DNA with known methylation state, deriving rate- parametrized partition distributions and estimating the rate using a maximum likelihood approach. In one or more examples, the methylation rate can be determined by determining an abundance of sequencing reads that correspond to a portion of a genomic region. The portion of the genomic region can include a number of genomic locations of the genomic region for which at least a threshold number of sequencing reads overlap.

[0377] Methylation Status: As used herein, “methylation status” or “methylation state” can refer to the presence or absence of methyl group on a DNA base (e.g. cytosine) at a particular genomic position in a nucleic acid molecule. It can also refer to the degree of methylation in a nucleic acid sequence (e.g., highly methylated, low methylated, intermediately methylated or unmethylated nucleic acid molecules). The methylation status can also refer to the number of nucleotides methylated in a particular nucleic acid molecule.

[0378] Modified Nucleotide Specific Binding Reagent: As used herein, refers to a binding reagent that is specific for, or targets, modified nucleotides. For example, a modified nucleotide can be a nucleotide that has been methylated, thus, the binding reagent can be specific for a methylated nucleotide. Examples of binding reagents include, but are not limited to, a methyl binding domain (MBD) of a methylation binding protein (“MBP”) or variants thereof, an antibody (and antibody variants e.g., single chain antibodies), aptamers, or combinations thereof. Thus, as disclosed throughout, the use of MBD can be exchanged for any other modified nucleotide specific binding reagent, provided the modified nucleotide specific binding reagent has the desired specificity and affinity for the specific modified base of interest in the selected implementation.

[0379] Mutant Allele Fraction: As used herein, “mutant allele fraction”, “mutation dose,” or “MAF” refers to the fraction of nucleic acid molecules harboring an allelic alteration or mutationat a given genomic position in a given sample. MAF is generally expressed as a fraction or a percentage. For example, an MAF can be less than about 0.5, 0.1, 0.05, or 0.01 (i.e., less than about 50%, 10%, 5%, or 1%) of all somatic variants or alleles present at a given locus.

[0380] Mutation: As used herein, “mutation” refers to a variation from a known reference sequence and includes mutations such as, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / aberrations, insertions or deletions (indels), gene fusions, transversions, translocations, frame shifts, duplications, repeat expansions, and epigenetic variants. A mutation can be a germline or somatic mutation. In some examples, a reference sequence for purposes of comparison is a wildtype genomic sequence of the species of the subject providing a test sample, typically the human genome.

[0381] Mutation Caller: As used herein, “mutation caller” means an algorithm (embodied in software or otherwise computer implemented) that is used to identify mutations in test sample data (e.g., sequence information obtained from a subject).

[0382] Mutation Count: As used herein, “mutation count” or “mutational count” refers to the number of somatic mutations in a whole genome or exome or targeted regions of a nucleic acid sample.

[0383] Negative Control Region: As used herein, “negative control region”, refers to a genomic region that is expected to be unmethylated or hypomethylated in essentially all samples, regardless of whether the DNA is derived from a cancer cell or a normal cell.

[0384] Neoplasm: As used herein, the terms “neoplasm” and “tumor” are used interchangeably. They refer to abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is referred to as a cancer or a cancerous tumor.

[0385] Next Generation Sequencing: As used herein, “next generation sequencing” or “NGS” refers to sequencing technologies having increased throughput as compared to traditional Sanger- and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequencing reads at a time. Some examples of next generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.

[0386] Nucleic Acid Tag: As used herein, “nucleic acid tag” refers to a short nucleic acid (e.g., less than about 500 nucleotides, about 100 nucleotides, about 50 nucleotides, or about 10nucleotides in length), used to distinguish nucleic acids from different samples (e.g., representing a sample index), or different nucleic acid molecules in the same sample (e.g., representing a molecular barcode), of different types, or which have undergone different processing. The nucleic acid tag comprises a predetermined, fixed, non-random, random or semi-random oligonucleotide sequence. Such nucleic acid tags may be used to label different nucleic acid molecules or different nucleic acid samples or sub-samples. Nucleic acid tags can be single-stranded, double- stranded, or at least partially double-stranded. Nucleic acid tags optionally have the same length or varied lengths. Nucleic acid tags can also include double-stranded molecules having one or more blunt-ends, include 5’ or 3’ single-stranded regions (e.g., an overhang), and / or include one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one end or to both ends of the other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample of origin, form, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples comprising nucleic acids bearing different molecular barcodes and / or sample indexes in which the nucleic acids are subsequently being deconvolved by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g. molecular identifier, sample identifier). Additionally, or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or amplicons of different parent molecules in the same sample or sub-sample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample, or non-uniquely tagging such molecules. In the case of non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their endogenous sequence information (for example, start and / or stop positions where they map to a selected reference sequence, a sub-sequence of one or both ends of a sequence, and / or length of a sequence) in combination with at least one molecular barcode. A sufficient number of different molecular barcodes are used such that there is a low probability (e.g., less than about a 10%, less than about a 5%, less than about a 1%, or less than about a 0.1% chance) that any two molecules may have the same endogenous sequence information (e.g., start and / or stop positions, subsequences of one or both ends of a sequence, and / or lengths) and also have the same molecular barcode.

[0387] Partitioning: As used herein, “partitioning” refers to physically separating or fractionating a mixture of nucleic acid molecules in a sample based on a characteristic of the nucleic acid molecules. The partitioning can be physical partitioning of molecules. Partitioning can involve separating the nucleic acid molecules into groups or sets based on the level of epigenetic feature (for e.g., methylation). For example, the nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, the methods and systems used for partitioning may be found in PCT Patent Application No. PCT / US2017 / 068329, which is hereby incorporated by reference in its entirety.

[0388] Partitioned set: As used herein, “partitioned set” or “partition” refers to a set of nucleic acid molecules partitioned into a set or group based on the differential binding affinity of the nucleic acid molecules or proteins associated with the nucleic acid molecules to a binding agent. A partitioned set may also be referred to as a subsample. The binding agent binds preferentially to the nucleic acid molecules comprising nucleotides with epigenetic modification. For example, if the epigenetic modification is methylation, the binding agent can be a methyl binding domain (MBD) protein. In some embodiments, a partitioned set can comprise nucleic acid molecules belonging to a particular level or degree of epigenetic feature (for e.g., methylation). For example, the nucleic acid molecules can be partitioned into three sets – one set for highly methylated nucleic acid molecules (first subsample, hyper partition, hyper partitioned set or hypermethylated partitioned set), a second set for low methylated nucleic acid molecules (second subsample, hypo partition, hypo partitioned set or hypomethylated partitioned set), and a third set for intermediate methylated nucleic acid molecules (third subsample, intermediate partitioned set, intermediately methylated partitioned set, residual partition, or residual partitioned set). In another example, the nucleic acid molecules can be partitioned based on the number of methylated nucleotides - one partitioned set can have nucleic acid molecules with nine methylated nucleotides, and another partitioned set can have unmethylated nucleic acid molecules (zero methylated nucleotides).

[0389] Polynucleotide: As used herein, “polynucleotide”, “nucleic acid”, “nucleic acid molecule”, “polynucleotide molecule”, or “oligonucleotide” refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleosidic linkages. A polynucleotide can comprise at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g. 3-4, to hundreds of monomeric units. Whenever apolynucleotide is represented by a sequence of letters, such as “ATGCCTG,” it will be understood that the nucleotides are in 5’ ^ 3’ order from left to right and that in the case of DNA, “A” denotes deoxyadenosine, “C” denotes deoxycytidine, “G” denotes deoxyguanosine, and “T” denotes deoxythymidine, unless otherwise noted. The letters A, C, G, and T may be used to refer to the bases themselves, to nucleosides, or to nucleotides comprising the bases, as is standard in the art.

[0390] Positive Control Region: As used herein, As used herein, “positive control region”, refers to a genomic region that is expected to be methylated or hypermethylated in essentially all samples, regardless of whether the DNA is derived from a cancer cell or a normal cell.

[0391] Probe: As used herein, “probe” refers to a polynucleotide comprising a functionality. The functionality can be a detectable label (fluorescent), a binding moiety (biotin), or a solid support (a magnetically attractable particle or a chip). Probes can include single- stranded DNA / RNA polynucleotides or double stranded DNA polynucleotides that hybridize to target nucleic acid sequences (e.g., SureSelect®probes, Agilent Technologies). Sequence capture using probes generally depends, in part, on the number of consecutive nucleotides in at least a portion of the target nucleic acid sequence that is complementary (or nearly complementary) to the sequence of the probe. In some examples, probes can correspond to driver mutations.

[0392] Processing: As used herein, the terms “processing”, “calculating”, and “comparing” can be used interchangeably. In certain applications, the terms refer to determining a difference, e.g., a difference in number or sequence. For example, gene expression, copy number variation (CNV), indel, and / or single nucleotide variant (SNV) values or sequences can be processed.

[0393] Processor: As used herein, “processor” refers to any circuit or virtual circuit (a physical circuit emulated by logic executing on an actual processor) that manipulates data values according to control signals (e.g., "commands," "op codes," "machine code," etc.) and which produces corresponding output signals that are applied to operate a machine. A processor may, for example, be a CPU, a RISC processor, a CISC processor, a GPU, a DSP, an ASIC, a RFIC or any combination thereof. A processor may further be a multi-core processor having two ormore independent processors (sometimes referred to as "cores") that may execute instructions contemporaneously.

[0394] Promoter Region As used herein, “promoter region” refers to a DNA sequence recognized by the synthetic machinery of the cell, or introduced synthetic machinery, required to initiate the specific transcription of a gene.

[0395] Quantitative Measures: As used herein, “quantitative measures” refers to an absolute or relative measure. A quantitative measure can be, without limitation, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or a relative quantity (e.g., high, medium, and low). A quantitative measure can be a ratio of two quantitative measures. A quantitative measure can be a linear combination of quantitative measures. A quantitative measure may be a normalized measure.

[0396] Reference Sequence: As used herein, “reference sequence” refers to a known sequence used for purposes of comparison with experimentally determined sequences. For example, a known sequence can be an entire genome, a chromosome, or any segment thereof. A reference sequence can include at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more nucleotides. A reference sequence can align with a single contiguous sequence of a genome or chromosome or can include non- contiguous segments that align with different regions of a genome or chromosome. Example reference sequences, include, for example, human genome reference sequences, such as, hG19 and hG38.

[0397] Sample: As used herein, “sample” means anything capable of being analyzed by the methods and / or systems disclosed herein.

[0398] Sensitivity: As used herein, “sensitivity” means the probability of detecting the presence of a single nucleotide variant, an insertion, and a deletion at a given MAF and coverage and the probability of detecting the presence of a copy number variant at a given tumor fraction and coverage.

[0399] Sequencing: As used herein, “sequencing” refers to any of a number of technologies used to determine the sequence (e.g., the identity and order of monomer units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Example sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exomesequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole-genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature-PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and a combination thereof. In some implementations, sequencing can be performer by a gene analyzer such as, for example, gene analyzers commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.

[0400] Single Nucleotide Variant: As used herein, “single nucleotide variant” or “SNV” means a mutation or variation in a single nucleotide that occurs at a specific position in the genome.

[0401] Somatic Mutation: As used herein, “somatic mutation” means a mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and accordingly, are not passed on to progeny.

[0402] Specifically binds: As used herein, “specifically binds” in the context of an probe or other oligonucleotide and a target sequence means that under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence, or replicates thereof, to form a stable probe:target hybrid, while at the same time formation of stable probe:non-target hybrids is minimized. Thus, a probe hybridizes to a target sequence or replicate thereof to a sufficiently greater extent than to a non-target sequence, to enable capture or detection of the target sequence. Appropriate hybridization conditions are well-known in the art, may be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) at §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, incorporated by reference herein).

[0403] Subject: As used herein, “subject” refers to an animal, such as a mammalian species (e.g., human) or avian (e.g., bird) species, or other organism, such as a plant. More specifically, a subject can be a vertebrate, e.g., a mammal such as a mouse, a primate, a simian or a human. Animals include farm animals (e.g., production cattle, dairy cattle, poultry, horses, pigs, and the like), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual that has or is suspected of having a disease or a predisposition to the disease, or an individual that is in need of therapy or suspected of needing therapy. The terms “individual” or “patient” are intended to be interchangeable with “subject.”

[0404] For example, a subject can be an individual who has been diagnosed with having a cancer, is going to receive a cancer therapy, and / or has received at least one cancer therapy. The subject can be in remission of a cancer. As another example, the subject can be an individual who is diagnosed of having an autoimmune disease. As another example, the subject can be a female individual who is pregnant or who is planning on getting pregnant, who may have been diagnosed of or suspected of having a disease, e.g., a cancer, an auto-immune disease.

[0405] Target Region: As used herein, “target region” refers to a genomic locus targeted for identification and / or capture, for example, by using probes (e.g., through sequence complementarity). A “target region set” or “set of target regions” refers to a plurality of genomic loci targeted for identification and / or capture, for example, by using a set of probes (e.g., through sequence complementarity)..

[0406] Threshold: As used herein, “threshold” refers to a predetermined value used to characterize experimentally determined values of the same parameter for different samples depending on their relation to the threshold.

[0407] Tumor Fraction: As used herein, “tumor fraction” refers to the estimate of the fraction of nucleic acid molecules derived from a tumor in a given sample. For example, the tumor fraction of a sample can be a measure derived from the max MAF of the sample or pattern of sequencing coverage of the sample or length of the cfDNA fragments in the sample or any other selected feature of the sample. In some instances, the tumor fraction of a sample is equal to the max MAF of the sample.

[0408] Variant: As used herein, a “variant” can be referred to as an allele. A variant is usually presented at a frequency of 50% (0.5) or 100% (1), depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. Somatic variants; however, are acquired variants and usually have a frequency of < 0.5. Major and minor alleles of a genetic locus refer to nucleic acids harboring the locus in which the locus is occupied by a nucleotide of a reference sequence, and a variant nucleotide different than the reference sequence respectively. Measurements at a locus can take the form of allelic fractions (AFs), which measure the frequency with which an allele is observed in a sample. Method of Aggregating Single Molecule Methylation States

[0409] FIG.1 comprises a block diagram of a system 100 for determining an indication of the amount or type of tumor molecules in an example. The system 100 can include a wet lab environment 110, a bioinformatics environment 140, and a user interface 170. The wet lab environment 110 can include collection tools 120, for collection of background data set 122, tumor data set 124, and collection tools 130 for collection of target genetic data 134 from the test subject 132. The bioinformatics environment 140 can include computational tools 150 for CpG selection 152 and production of a framework 154, in addition to the produced framework 160. These can be used to produce an indication of tumor on the user interface 170.

[0410] The wet lab environment 110 can be a laboratory setting configured for receipt and processing of biological samples. The wet lab environment 110 can be, for example, a commercial, research, medical, or academic laboratory for receiving and processing genetic samples. Such samples, such as biological samples, can be processed to produce the background data set 122, the tumor data set 124, and the sample data set 134. In the case of a sample data set 134, the data can originate from a patient 132. Examples such a wet lab environment 110 samples, and associated techniques are described in detail below.

[0411] The collected background data set 122 and tumor data set 124 from the wet lab environment 110 can be used in the bioinformatic tool 150, through methods of CpG selection 152 and framework production 154, to produce the framework 160. The sample data set 134 can then be applied to the framework 160. A determination of a tumor can then be indicated on the user interface 170.

[0412] FIG.2 comprises a block diagram of a method 200 of determining an indication of the amount or type of tumor molecules in an example. The method 200 can include blocks 220 to 250. Blocks 210 to 250 involved the production of a framework for application to a sample data set 240. The method 200 can be used to produce a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, and providing that indication of the amount or type of tumor molecules.

[0413] At block 210, a background data set can be collected. The background data set can include both sequencing data and methylation data. The methylation data can indicate an amount of methylation of CpG nucleotides in the background data set. The background data set can be produced in a bioinformatics environment. The background data set can originate from a subject, or multiple subjects, in which a tumor is not detected. For example, a large number of background subjects can be used to create a larger background data set. The techniques used to collect and process such samples and production of such a data set are discussed in more detail below. At block 212, the background signal rate can be calculated based on the background data set.

[0414] At block 220, a tumor data set can be collected. The tumor data set can include both sequencing data and methylation data. The methylation data can indicate an amount of methylation of CpG nucleotides in the tumor data set. The tumor data set can be produced in a bioinformatics environment. The tumor data set can originate from a subject, or multiple subjects, in which a tumor is detected. The techniques used to collect and process such samples and production of such a data set are discussed in more detail below. At block 222, the tissue signal rate can be calculated based on the tumor data set.

[0415] At block 232, the calculated background signal rate and tissue signal rate can be compared together. They can be used to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tissue signal rate above a tissue signal cutoff value. At block 234, the framework can be produced based on the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters. The identification of such a subset of CpG clusters and production of the framework are discussed in more detail with reference to FIG.3 below.

[0416] At block 240, a sample data set can be collected and the framework applied. At block 250, an indication of a tumor-related biological condition present in a test subject based on application of the framework can be presented, such as on a user interface.

[0417] FIG. 3 comprises a diagram of a method 300 of selecting CpG clusters in an example. The method 300 can use sequencing data 310 for analysis of CpG clusters 320 and CpG clusters 330. They can be analyzed according to signal rate 325.

[0418] The CpG clusters shown here are an extension of a single CpG locus to a group of (“K”) consecutive CpG loci. At this loci, the combined methylation state of the comprising CpG loci can be used to classify molecules. This can be applied to detect tumor molecules in cfDNA and regions that gain methylation in tumor tissue. In some cases, the method can be adapted to detect signal in regions that loose methylation in cancer tissue or other more complex methylation patterns relevant for other applications.

[0419] The number of signal molecules can be identified. Signal molecules are molecules with majority of K CpG loci in methylated state (“mCpG”). FIG.3 depicts an example of a CpG cluster of size four (e.g., K=4), where at least three out of four CpGs to be in methylated state (e.g., mCpG>=3).

[0420] Larger clusters (e.g., larger “K” values) allow for better separation of fully- methylated states from background noise and thus have better signal-to-noise properties. However, as “K” increases coverage can be lost at less CpG dense regions. This can occur if only using molecules that fully contain CpG clusters are used. Thus, relatively small values of K can be used for the methods herein, e.g., a K of about 1, 4, 5 or 6. Examples of selected such CpG clusters, with the associated K value and mCpG values, are shown and discussed below with reference to Examples 1 to 3.

[0421] FIG.4 comprises a flow chart of a method 400 of providing an indication of a tumor- related biological condition in an example. Method 400 can be a method of methylation analysis for detection of a tumor-related biological condition.

[0422] At block 410, the method can include receiving a sample data set to be tested by a framework. Producing a framework can include blocks 420 to 470. Three data sets are used in the method 400: a sample data set to be tested for tumor detection, the background data set that originated from one or more samples not having tumor tissue, and the tumor data set that originated from one or more samples having tumor tissue. The sample data set, the backgrounddata set, and the tumor data set can be derived from cell free deoxyribonucleic acid (DNA) molecules, from plasma samples, or tissue samples.

[0423] At block 420, the method can include obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria. The the background data set can originate from one or more first subjects in which a tumor is not detected.

[0424] The predetermined methylation criteria can include a methylation state and a number of CpG dinucleotides that satisfy the methylation state. In an example, the predetermined methylation criteria can include sequence representations that indicate a gain in methylation. In an example, the predetermined methylation criteria can include sequence representations that indicate a loss in methylation. In an example, the number of CpG dinucleotides that satisfy the methylation state can be three CpG dinucleotides, four CpG dinucleotides, or five CpG dinucleotides.

[0425] At block 430, the method can include calculating a background signal rate using the background data set. Calculating the background signal rate can include determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

[0426] At block 440, the method can include obtaining a tumor data set based on the predetermined methylation criteria. The tumor data set can originate from one or more second subjects in which a tumor is detected. At block 450, the method can include calculating tissue signal rate based on the tumor data set. Determining the tissue signal rate can include determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

[0427] At block 460, the method can include identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tissue signal cutoff value. In an example, the background signal cutoff value can be about 1e-6to 0.01. In an example, the tissue signal cutoff value can be about.0.05 to 1.00. The subset of the CpG dinucleotide clusters can be no more than a predeterminednumber of CpG dinucleotides. In an example, the predetermined number of CpG dinucleotides can be about six, four, or three.

[0428] At block 470, the method can include producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters. In an example, producing the framework can include ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate. In an example, producing the framework further can include selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters. In an example, producing the framework can include classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

[0429] At block 480, the method can include applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters. In an example, applying the framework can include determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold. In an example, applying the framework can include aggregating quantitative measures that are above the cutoff at each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition. In an example, applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters can include using a machine learning model.

[0430] At block 490, the method can include providing an indication of a tumor-related biological condition present in the test subject based on application of the framework. In an example, the indication of the tumor-related biological condition can correspond to an amount of activated T-cells present in the test subject. In an example, the indication of the tumor-related biological condition can correspond to an amount of cancer-associated fibroblasts present in the test subject. In an example, the indication of the tumor-related biological condition can be determined using one or more machine learning techniques. In an example, the indication of the tumor-related biological condition can include an indication of the type of cancer. In an example, the indication of the tumor-related biological condition can include an indication of the type of a tumor fraction. In an example, the indication of the tumor-related biological condition can include an indication of a tissue of origin. In an example, the indication of the tumor-related biologicalcondition can include an indication of tumor burden. In an example, the indication of the tumor- related biological condition can include an indication of tumor recurrence.

[0431] FIG.5 illustrates a framework 500 for generating data that is analyzed by one or more computational models to determine one or more biological condition indicators, in accordance with one or more example implementations. The framework 500 can include obtaining a sample 502 from a subject 504. In one or more examples, the sequencing data 506 and epigenomics data 508 can be produced from the sample 502. In at least some examples, the sample 502 can include at least one of one or more tissue samples or one or more fluid samples. The one or more fluid samples can include at least one of a whole blood sample, a buffy coat sample, or a plasma sample. In various examples, the sequencing data 506 can correspond to sequence representations derived from the sample 502 that correspond to one or more classification regions. In one or more illustrative examples, the one or more classification regions can be related to one or more diagnostic tests. The one or more diagnostic tests can include one or more cancer detection assays.

[0432] In various examples, the sample 502 used to produce sequencing data 506 and the epigenomics data 508 can include a number of cell-free nucleic acid molecules. For example, the sample 502 can include a number of cell-free DNA molecules. Nucleic acid molecules can be extracted from the sample 502 by implementing one or more cell lysis techniques to cleave the membranes of cells included in the sample 502 and applying one or more proteases to break down proteins included in the sample 502. The extraction of nucleic acids from the sample 502 can also include a number of washing and / or elution techniques to separate the nucleic acids from other components included in the sample 502. In various examples, thousands, up to millions, up to billions of nucleic acids can be extracted from the sample 502.

[0433] The sequencing data 506 and epigenomics data 508 can be produced by performing one or more molecule separation and sequencing processes with respect to the sample 502. The one or more molecule separation and sequencing processes can be performed by one or more sequencing machines that can produce the sequence data. In one or more examples, sequencing data 506 can include sequencing reads that correspond to nucleic acid molecules derived from the sample 502. In various examples, the sequencing data 506 can include alphanumeric representations of nucleic acids included in an amplification product produced by implementing the one or more molecule separation and sequencing processes withrespect to the sample 502. To illustrate, the sequencing data 506 can include, for individual nucleic acids of the amplification product, data that corresponds to a string of letters that represent the respective chains of nucleotides that correspond to the individual nucleic acids derived from the sample 502.

[0434] In one or more examples, the nucleic acids extracted from at least a portion of the sample 502 can also be enriched by causing hybridization between the extracted polynucleotides and probes that correspond to classification regions of a reference sequence. The enrichment process can identify thousands, hundreds of thousands, up to millions of polynucleotides that correspond to classification regions associated with the probes. In one or more illustrative examples, the classification regions can correspond to genomic regions that are part of one or more diagnostic assays that can be performed to identify one or more types of cancer present in subjects. In various examples, the classification regions can include one or more differentially methylated regions that correspond to a number of cancer types or subtypes. In one or more illustrative examples, the nucleic acids extracted from the sample 502 can be hybridized with respect to at least 10 classification regions, at least 25 classification regions, at least 50 classification regions, at least 100 classification regions, at least 250 classification regions, at least 500 classification regions, at least 1000 classification regions, at least 2500 classification regions, at least 5000 classification regions, or at least 10,000 classification regions.

[0435] Subsequent and / or prior to the enrichment process, the nucleic acids derived from the sample 502 can be amplified according to one or more amplification processes. The one or more amplification processes can produce thousands, up to millions of copies of individual nucleic acid molecules. In one or more examples, a portion of the unenriched polynucleotides can be amplified, in some instances, but not to the extent that the enriched polynucleotides are amplified. The one or more amplification processes can generate an amplification product that undergoes one or more sequencing operations to produce the sequence data and / or the epigenomic data derived from the sample 502.

[0436] The epigenomic data 508 can indicate characteristics of nucleic acids derived from the sample 502 that are in addition to the nucleotide sequences of the nucleic acids. For example, the epigenomic data 508 can indicate methylation states of cytosines present in cytosine-guanine dinucleotides of the nucleic acids derived from the sample 502. The nucleic acid molecules extracted from the sample 502 can include molecules having varying levels of methylation.Methylation can occur from any one or more post-replication or transcriptional modifications. Post- replication modifications include modifications of the nucleotide cytosine, including, but not limited to, 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine.

[0437] The methylation states of the cytosines can be determined by producing subsamples from the sample 502. To illustrate, the one or more molecule separation and sequencing processes used to produce the epigenomic data 508 can correspond to separating nucleic acid molecules into a number of partitions based on the characteristics of the nucleic acid molecules. Examples of characteristics that can be used for partitioning nucleic acid molecules include multiple different nucleotide modifications, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. In one or more illustrative examples, a heterogeneous population of nucleic acids can be partitioned into nucleic acids with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include, but are not limited to, presence or absence of methylation; level of methylation, hydroxymethylation, and type of methylation (5′ cytosine or 6 methyladenine).

[0438] In one or more illustrative examples, the one or more molecule separation and sequencing processes used to generate the epigenomic data 508 can separate nucleic acid molecules extracted from the sample 502 into a number of partitions with individual partitions corresponding to different levels of methylation. For example, the molecule separation and sequencing processes used to generate the epigenomic data 508 can produce a first partition of nucleic acid molecules having first levels of methylation, a second partition of nucleic acid molecules having second levels of methylation, and a third partition of nucleic acid molecules having third levels of methylation. In various examples, the second levels of methylation can be greater than the first levels of methylation and the third levels of methylation can be greater than the first levels of methylation and the second levels of methylation. In at least some examples, partitioning the sample into a plurality of subsamples can include contacting the number of nucleic acids with a methyl binding reagent immobilized on a solid support.

[0439] In one or more illustrative examples, the molecule separation and sequencing processes implemented to produce the epigenomic data 508 can include combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleic acid-MBD protein solution.A plurality of washes can then be performed of the nucleic acid-MBD protein solution with a salt solution to produce a number of nucleic acid fractions. Individual nucleic acid fractions can have a threshold number of molecules with a methylated cytosine in regions of the plurality of nucleic acids having at least the threshold cytosine-guanine content. In various examples, a wash of the plurality of washes can be performed with a solution having a concentration of sodium chloride (NaCl) and can produce a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins.

[0440] In at least some examples, a first nucleic acid fraction can be determined is associated with a first partition of a plurality of partitions of nucleic acids. The first partition corresponding to a first range of binding strengths to MBD proteins. Further, a first molecular barcode can be attached to nucleic acids of the first nucleic acid fraction. The first molecular barcode can be associated with the first partition. In addition, a second nucleic acid fraction can be determined that is associated with a second partition of the plurality of partitions of nucleic acids. The second partition can correspond to a second range of binding strengths to MBD proteins different from the first range of binding strengths to MBD proteins. A second molecular barcode can be attached to nucleic acids of the second nucleic acid fraction. The second molecular barcode being associated with the second partition.

[0441] In addition, the molecule separation and sequencing processes implemented to produce the epigenomic data 508 can include combining at least a portion of the number of nucleic acid fractions with an amount of restriction enzyme that cleaves molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads. In these scenarios, the threshold amount of molecules with a methylated cytosine corresponds to a minimum frequency of molecules with a methylated cytosine within a region having at least the threshold cytosine-guanine content. In one or more further examples, at least a portion of the number of nucleic acid fractions are combined with an amount of a restriction enzyme that cleaves molecules with a methylated cytosine to produce at least a portion of the plurality of samples used to produce the sequencing reads. In these situations, the threshold amount of molecules with a methylated cytosine corresponds to a maximum frequency of molecules with a methylated cytosine within a region having at least the threshold cytosine- guanine content.

[0442] In one or more additional illustrative examples, the one or more molecule separation and sequencing processes implemented to produce the sequence data and / or the epigenomic data can include (i) sodium bisulfite conversion and sequencing, (ii) Tet-assisted bisulfite sequencing (TAB-Seq), differential enzymatic cleavage, or (iii) other conversion procedures using enzymes – e.g., Enzymatic Methyl Sequencing (EM-Seq), Direct Methylation Sequencing (DM-Seq) or single-enzyme 5-methylcytosine sequencing (SEM-seq)-. In still other illustrative examples, the one or more molecule separation and sequencing processes implemented to produce the sequencing data 506 and / or the epigenomic data 508 derived from the sample 502 can include one or more single molecule sequencing methods, such as nanopore DNA sequencing or those described in Eid, J., et al. (2009) Real-time DNA sequencing from single polymerase molecules. Science, 323(5910), 133–138.

[0443] The framework 500 can include, at 510, aligning sequence representations included in the sequencing data 506 with a number of genomic subregions. For example, a reference genome can include a number of genomic regions. A number of genomic regions included in the reference genome can be indicative of the presence or absence of one or more biological conditions. For example, a number of genomic regions can be susceptible to changes that are indicative of the presence of at least one of one or more cancer types or one or more cancer subtypes. To illustrate, methylation patterns relating to changes of cytosine moieties included in nucleic acids derived from subjects that correspond to a number of genomic regions can be indicative of the presence of at least one of one or more cancer types or one or more cancer subtypes. In various examples, one or more diagnostic tests can be performed to detect the presence of nucleic acid molecules that have methylation patterns that indicate the presence of at least one of one or more cancer types or one or more cancer subtypes in subjects. In one or more illustrative examples, genomic regions in which nucleic acids derived from subjects have methylation patterns in cytosine-guanine dinucleotides that correspond to the presence of at least one of one or more cancer types or one or more cancer subtypes can be referred to as classification regions or differentially methylated regions. In the illustrative example of FIG.5, two example differentially methylated regions, a first differentially methylated region 512 and a second differentially methylated region 514 are shown.

[0444] In one or more examples, the accuracy, precision, and / or efficiency of computational models executed to determine the presence or absence of at least one of one ormore cancer types or one or more cancer subtypes in subjects can be improved when genomic subregions within the differentially methylated regions are used to identify nucleic acids having methylation patterns that are indicative of at least one of one or more cancer types or one or more cancer subtypes. In the illustrative example of FIG.5, the first differentially methylated region 512 includes a first genomic subregion 516 and the second differentially methylated region 514 includes a second genomic subregion 518 that can be used to identify nucleic acid molecules that satisfy one or more methylation criteria that are more likely to be indicative of the presence of at least one of one or more cancer types or one or more cancer subtypes than nucleic acid molecules that satisfy the one or more methylation criteria that are outside of the first genomic subregion 516 or outside of the second genomic subregion 518.

[0445] The first genomic subregion 516 and the second genomic subregion 518 can be determined by identifying one or more clusters of cytosine-guanine dinucleotides (CpGs) that have one or more methylation characteristics and that provide a relatively large signal in relation to a background signal in relation to the presence of at least one of one or more cancer types or one or more cancer subtypes. For example, the first genomic subregion 516 can be determined by identifying a first group of CpG clusters 520 that includes a number of individual first CpG clusters 522. Additionally, the second genomic subregion 518 can be determined by identifying a second group of CpG clusters 524 that includes a number of individual second CpG clusters 526. The individual first CpG clusters 522 and / or the individual second CpG clusters 526 can be arranged in consecutive genomic positions within the first genomic subregion 516 or the second genomic subregion 518. In one or more additional examples, one or more gaps can be present between the individual first CpG clusters 522 that comprise the first genomic subregion 516 or between the individual second CpG clusters 526 that comprise the second genomic subregion 518. The one or more gaps can be from 5 nucleotides to 500 nucleotides, from 10 nucleotides to 250 nucleotides, from 20 nucleotides to 100 nucleotides, from 10 nucleotides to 100 nucleotides, or from 50 nucleotides to 150 nucleotides.

[0446] Additionally, in various examples, at least a threshold number of individual first CpG clusters 522 define the first genomic subregion 516 and at least threshold number of individual second CpG clusters 526 define the second genomic subregion 518. To illustrate, genomic subregions can be determined within a given differentially methylated region and / or a given classification region in situations where at least 2 CpG clusters are present, at least 3 CpGclusters are present, at least 4 CpG clusters are present, at least 5 CpG clusters are present, at least 6 CpG clusters are present, at least 7 CpG clusters are present, at least 8 CpG clusters are present, at least 9 CpG clusters are present, at least 10 CpG clusters are present, at least 11 CpG clusters are present, or at least 12 CpG clusters are present. In one or more illustrative examples, genomic subregions can be determined within a given differentially methylated region and / or a given classification region in situations where from 2 to 20 CpG clusters are present, from 4 to 15 CpG clusters are present, from 8 to 12 CpG clusters are present, from 4 to 10 CpG clusters are present, from 6 to 12 CpG clusters are present, or from 5 to 15 CpG clusters are present.

[0447] Further, the individual first CpG clusters 522 and the individual second CpG clusters 526 can have at least a threshold number of nucleotides. For example, the individual first CpG clusters 522 and the individual CpG clusters 526 can have from 2 to 1000 nucleotides, from 50 to 500 nucleotides, from 10 to 100 nucleotides, from 50 to 200 nucleotides, from 100 to 300 nucleotides, or from 200 to 500 nucleotides. In still other examples, the individual first CpG clusters 522 and the individual second CpG clusters 526 can have at least a threshold number of CpGs having one or more specified methylation characteristics. In addition, the individual first CpG clusters 522 and the individual second CpG clusters 526 can have at least a threshold methylation rate. In various examples, the one or more specified methylation characteristics can correspond to a number of methylated cytosines or a number of unmethylated cytosines present in a number of CpG clusters. In one or more examples, the threshold methylation rate can correspond to a percentage of methylated cytosines or a percentage of unmethylated cytosines present in a number of CpG clusters with respect to a total number of cytosines present in the individual CpG clusters.

[0448] In at least some examples, first individual CpG clusters 522 and the individual second CpG clusters 526 can be determined at 528 where correlated CpG clusters are determined. Correlated CpG clusters can correspond to CpG clusters located within one or more differentially methylated regions or one or more classification regions that have at least a threshold tumor signal in relation to a background signal. In various examples, correlated CpGs can have a minimum signal to noise ratio with respect to at least one of one or more cancer types or one or more cancer subtypes. In one or more examples, a signal can correspond to a number of nucleic acids derived from a sample that are aligned with a given genomic region, such as aspecified CpG cluster. In one or more additional examples, a signal can correspond to a ratio of a number of nucleic acids derived from a sample that are aligned with a given genomic region in relation to a number of nucleic acids that are aligned with one or more control regions.

[0449] In one or more examples, the individual first CpG clusters 522 and the individual second CpG clusters 526 can be determined by applying CpG framework data 530 to features of a number of nucleic acids derived from a sample. In one or more additional examples, the CpG framework data 530 can be determined by analyzing nucleic acids derived from one or more training samples in relation to one or more selection criteria. In one or more illustrative examples, the CpG framework data can include the one or more selection criteria can correspond to alignment with at least a portion of one or more classification regions and / or one or more differentially methylated regions. In one or more additional illustrative examples, the one or more selection criteria can correspond to alignment with at least a portion of one or more probes that are used to amplify nucleic acids corresponding to one or more classification regions and / or one or more differentially methylated regions. In various examples, the one or more probes can be included in a diagnostic assay that is administered to detect the presence or absence of at least one of one or more cancer types or one or more cancer subtypes in subjects. In still other examples, the one or more selection criteria included in the CpG framework data 530 can include the presence of at least one, at least two, at least three, at least four, or at least five restriction enzyme cut sites, such as one or more methylation sensitive restriction enzyme (MSRE) cut sites and / or one or more methylation dependent restriction enzyme (MDRE) cut sites. In one or more examples, the CpG framework data 530 can indicate a maximum number of nucleotides for CpG clusters and / or indicate that CpG clusters that are not fully located within a classification region are not included in the individual first CpG clusters 522 or the individual second CpG clusters 526. The maximum number of nucleotides included in the CpG framework data 530 can include no greater than 500 nucleotides, no greater than 400 nucleotides, no greater than 300 nucleotides, no greater than 250 nucleotides, or no greater than 200 nucleotides. In various additional examples, the CpG framework data 530 can include normalization factors related to identifying control regions having at least a threshold number of methylated or unmethylated CpGs, such as at least one, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 methylated or unmethylated CpGs.

[0450] Determining the corelated CpG clusters at 528 can include determining background rates of methylation with respect to a number of CpG clusters. The background methylation rates can be determined by analyzing nucleic acids derived from training subjects in which a tumor has not been detected. The background rates can be determined by determining a number of control region nucleic acids present in one or more training samples. The background rates can also be determined by determining a number of target region nucleic acids present in the one or more training samples, where the number of target regions can include one or more classification regions or one or more differentially methylated regions. In at least some examples, the target region coverage can be determined by calculating the geometric mean of individual nucleic acids that are aligned with the individual target regions of number of target regions. A normalization factor included in the CpG framework data 530 can include a ratio of target coverage in relation to control coverage. A number of normalized molecules can be determined by taking the product of the normalization factor and the number of molecules that are aligned with at least one CpG cluster included in at least one target region. A total number of normalized molecules can be determined by taking the product of the number of training samples and the target coverage. In at least some examples, the background rate can be determined by Equation (1) below:

[0451] pseudo count (Equation 1).

[0452] The determination of the CpG framework data 530 can also include filtering CpG clusters for training sample in which a tumor corresponding to at least one of one or more cancer types or one or more cancer subtypes has been detected. In these scenarios, the CpG framework data 530 can indicate that the CpG clusters taken from the tumor positive training samples have at least a threshold number of molecules aligned with the respective CpG clusters and pass a p- value filter. The p-value filter can be at least 1e-4, at least 1e-5, at least 5e-5, at least 1e-6, at least 5e-6, or at least 1e-7.

[0453] CpG clusters that pass this initial set of selection criteria can also be subjected to an additional set of selection criteria. The additional set of selection criteria can include identifying CpG clusters having no greater than a threshold number of the 80thquantile of molecule counts for the training samples. In these scenarios, the threshold number can be from 30% to 70%, from40% to 60%, from 40% to 50%, or from 50% to 60%. The one or more selection criteria can also include a number of predefined background rate cutoffs. The predefined background rate cutoffs can be from 1e-06 to 5e-02, 6e-05 to 1e-03, 3e-05 to 1e-04, or from 3e-04 to 1e-03. In at least some examples, the CpG clusters that satisfy the one or more sets of selection criteria that are within at least a threshold number of genomic locations can be aggregated to form a genomic subregion, such as the first genomic subregion 516 or the second genomic subregion 518. The process for selecting CpG regions according to one or more selection criteria can be performed for a number of background rate cutoff values, such as at least 2 background cutoff values, at least 3 background cutoff values, at least 4 background cutoff values, or at least 5 background cutoff values. In this way, further filtering of CpG regions to be considered for the determination of the genomic subregions can take place. In one or more examples, at least a portion of the CpG clusters included in the genomic subregions can be determined at a least restrictive cutoff value. In at least some examples, a single CpG cluster can map to multiple classification regions or multiple differentially methylated regions.

[0454] In one or more additional examples, the background rate can be optimized. In one or more examples, the background rate can be optimized by identifying samples in which a maximum tumor fraction value is present, such as no greater than a 5% tumor fraction, no greater than a 4% tumor fraction, no greater than a 3% tumor fraction, no greater than a 2% tumor fraction, or no greater than a 1% tumor fraction. Additionally, a signal to noise ratio can be determined to identify a background rate that maximizes the signal to noise ratio. The signal to noise ratio can be determined by calculating a 95thquantile molecule counts for samples in which no tumor was detected for a number of classification regions and dividing the number of molecules per sample for individual classification regions by the 95thquantile of normal. The 80thquantile signal to noise ratio can then be determined for individual classification regions. The classification regions can then be ranked by a maximum signal to noise ratio.

[0455] Aligning the sequence representations with the genomic subregions can produce genomic region sequence representation data 532. The genomic region sequence representation data 532 can include sequence representations included in the sequencing data 506 that are aligned with at least a portion of the genomic subregions, such as the first genomic subregion 512 and the second genomic subregion 514. In at least some examples, the genomic region sequence representation data 532 can include sequence representations that correspond to one or moremethylation criteria, such as a minimum number of methylated CpGs or a maximum number of unmethylated CpGs.

[0456] At 534, the framework 500 can include determining genomic region quantitative measures 536 based on the genomic region sequence representation data 532. The genomic region quantitative measures 536 can be provided as input to one or more computational models 538 to generate one or more biological condition indicators 540. In one or more examples, the quantitative measures can correspond to sequence representations derived from or included in the sequence data derived from the sample 502 that satisfy one or more epigenomic characteristics and that are aligned with one or more classification regions. In one or more illustrative examples, the sequence representations used to generate the quantitative measures can correspond to sequencing reads included in the sequencing data 506. In one or more additional illustrative examples, the sequence representations used to generate the quantitative measures can be related to a group of sequence representations that correspond to a single nucleic acid molecule derived from the sample 502. For example, in scenarios where multiple sequence representations are present in the sequence data that correspond to a single nucleic acid molecule, the multiple sequence representations can be grouped together. In various examples, the groups of sequence representations that correspond to a single nucleic acid molecule can be referred to herein as “families.” Additionally, start and stop positions with respect to the reference sequence of sequence representations that have been aligned with respect to the reference genome can have a common molecular barcode that can be used to group the sequence representations that correspond to individual nucleic acids. In one or more illustrative examples, an individual sequence representation that represents a family of sequence representations that corresponds to a single nucleic acid molecule can be referred to herein as a “consensus sequence representation.”

[0457] The one or more epigenomic criteria used to identify the sequence representations that produce the quantitative measures can indicate a threshold number of CpGs present in individual sequence representations. For example, sequence representations derived from the sample 502 that are used to produce the quantitative measures can include at least a threshold number of CpGs. In one or more examples, the threshold number of CpGs can be at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10.

[0458] In one or more additional examples, the one or more epigenomic criteria applied to determine the sequence representations that are used to produce the quantitative measures can include one or more methylation characteristics. The one or more methylation characteristics can indicate a partition of nucleic acids that is produced by performing a number of washes with respect to nucleic acids derived from the sample 502 using solutions that include MBD and a number of different salt concentrations. To illustrate, the one or more methylation characteristics can indicate a first partition including first nucleic acids that is produced when a number of nucleic acids derived from the sample 502 is subjected to one or more first solutions having a concentration of MBD and one or more first salt concentrations. The one or more methylation characteristics can also indicate a second partition that includes second nucleic acids that is produced when the number of nucleic acids derived from the sample 502 is subjected to one or more second solutions having the concentration of MBD that is the same as or similar to the MBD concentration of the one or more first solutions and one or more second salt concentrations that is greater than the first salt concentration. In still other examples, the one or more methylation characteristics can indicate a third partition that includes third nucleic acids that are produced when the number of nucleic acids derived from the sample 502 is subjected to one or more third solutions having the concentration of MBD that is the same as or similar to the MBD concentration of the one or more first solutions and / or the one or more second solutions and one or more third salt concentrations that are greater than the one or more first salt concentrations and the one or more second salt concentrations. In one or more illustrative examples, the first partition can be referred to as a hypomethylation partition, the second partition can be referred to as a residual partition, and the third partition can be referred to as a hypermethylated partition. In at least some examples, the nucleic acids present in the third partition can, on average, have a greater number of methylation cytosines than nucleic acids present in the first partition and the second partition and the nucleic acids present in the second partition can have a greater number of methylated cytosines than the nucleic acids present in the first partition. Although the illustrative example of FIG.5 is described in relation to methylation characteristics corresponding to three partitions, in other implementations, more partitions or fewer partitions can be used.

[0459] In one or more further examples, the epigenomic criteria can indicate a threshold number of methylated cytosines present in sequence representations used to generate the quantitative measures. For example, sequence representations included in sequencing data 506that are used to determine the quantitative measures provided as input to the one or more computational models 538 can include at least a first threshold number of methylated cytosines. The first threshold number of methylated cytosines can correspond to a minimum number of methylated cytosines. In these scenarios, the first threshold number of methylated cytosines can correspond to at least 3 methylated cytosines, at least 4 methylated cytosines, at least 5 methylated cytosines, at least 6 methylated cytosines, at least 7 methylated cytosines, at least 8 methylated cytosines, at least 9 methylated cytosines, or at least 10 methylated cytosines. In various examples, the methylation criteria can also correspond to no greater than a second threshold number of methylated cytosines. In these situations, the second threshold number of methylated cytosines can correspond to a maximum number of methylated cytosines. To illustrate, the methylation criteria can correspond to no greater than 6 methylated cytosines, no greater than 5 methylated cytosines, no greater than 4 methylated cytosines, no greater than 3 methylated cytosines, no greater than 2 methylated cytosines, or no greater than 1 methylated cytosine.

[0460] After determining sequence representations derived from the sequencing data 506 that are aligned with at least a portion of one or more classification regions or genomic subregions and that correspond to the one or more epigenomic criteria quantitative measures for the individual classification regions or genomic subregions can be determined. In one or more examples, for the individual classification regions or genomic subregions, quantitative measures can be generated by determining a number of sequence representations satisfying the epigenomic criteria that align with the individual classification regions or genomic subregions. In at least some examples, counts of the sequence representations can be determined that satisfy the one or more epigenomic criteria and that align with the individual classification regions or genomic subregions. In one or more illustrative examples, a count of aligned sequence representations per classification region or per genomic subregion can be determined.

[0461] Additionally, the quantitative measures can include normalized quantitative measures for individual classification regions or individual genomic subregions that are based on the number of sequence representations aligned with the individual classification regions or genomic subregions and based on a number of sequence representations that are aligned with one or more control regions. In at least some examples, the normalized quantitative measures can be determined based on a number of sequence representations that are aligned with the oneor more control regions and that satisfy the one or more epigenomic criteria. In various examples, the quantitative measures can be generated by determining a ratio of counts of sequence representations that satisfy one or more epigenomic criteria and that are aligned with the individual classification regions or individual genomic subregions in relation to counts of sequence representations that are aligned with the one or more control regions. In one or more further examples the quantitative measures can be determined for individual classification regions by adding a pseudocount to the ratio of counts of sequence representations that satisfy one or more epigenomic criteria and that are aligned with the individual classification regions or individual genomic subregions in relation to counts of sequence representations that are aligned with the one or more control regions. In still other examples, normalized quantitative measures can be determined by performing one or more log-based operations with respect to the ratio of counts of sequence representations that satisfy one or more epigenomic criteria and that are aligned with the individual classification regions or individual genomic subregions in relation to counts of sequence representations that are aligned with the one or more control regions.

[0462] The genomic region quantitative measures 536 generated based on the genomic region sequence representation data 532 can be provided as input to the one or more computational models 538. The one or more computational models 538 can generate one or more biological condition indicators 540 corresponding to one or more cancer types or one or more cancer subtypes including one or more blood cancers, one or more central nervous system (CNS) cancers, one or more brain cancers, one or more skin cancers, one or more nose cancers, one or more throat cancers, one or more liver cancers, one or more bone cancers, one or more lymphomas, one or more pancreatic cancers, one or more bowel cancers, one or more rectal cancers, one or more thyroid cancers, one or more bladder cancers, one or more kidney cancers, one or more lung cancers, one or more mouth cancers, one or more stomach cancers, one or more breast cancers, one or more prostate cancers, one or more ovarian cancers, one or more intestinal cancers, one or more soft tissue cancers, one or more neuroendocrine cancers, one or more gastroesophageal cancers, one or more head and neck cancers, one or more gynecological cancers, one or more colorectal cancers, one or more urothelial cancers, one or more solid state cancers, one or more heterogeneous cancers, one or more homogenous cancers, one or more unknown primary origin cancers and the like, and / or of the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma)and / or one or more cancers exhibiting cancer markers, such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptor and NMP-22.

[0463] In one or more illustrative examples, the one or more biological condition indicators 540 can include a tumor fraction corresponding to the subject 504 in relation to at least one of one or more cancer types or one or more cancer subtypes. In one or more additional illustrative examples, the one or more biological condition indicators 504 can indicate a probability of at least one of one or more cancer types or one or more cancer subtypes being present in the subject 504. In one or more further illustrative examples, the one or more biological condition indicators 504 can indicate a presence or absence of at least one of one or more cancer types or one or more cancer subtypes in the subject 504. In still other examples, the one or more biological condition indicators 540 can indicate an amount of one or more cell types present in the subject 504. The one or more biological condition indicators 504 can also correspond to a responsiveness to one or more treatments by the subject 504, side effects present in the subject 504 with respect to one or more treatments, a recurrence of one or more cancer types and / or one or more cancer types with respect to the subject 504, an indicator of minimum residual disease with respect to the subject 504, or one or more combinations thereof.

[0464] In various examples, the one or more computational models 538 can include one or more classification machine learning models. Additionally, the one or more computational models 538 can include one or more computational models described in patent application number PCT / US2024 / 047535, filed September 19, 2024, and entitled DETECTING THE PRESENCE OF A TUMOR BASED ON METHYLATION STATUS OF CELL-FREE NUCLEIC ACID MOLECULES, which is incorporated by reference herein in its entirety. Further, the one or more computational models 538 can include one or more computational models described in patent application number 18 / 907,227, filed October 4, 2024, and entitled DETECTING THE PRESENCE OF A TUMOR BASED ON METHYLATION STATUS OF CELL-FREE NUCLEIC ACID MOLECULES, which is incorporated by reference herein in its entirety. In still other examples, the one or more computational models 538 can include one or more computational models described in patent application number PCT / US25 / 22042, filed March 28, 2025, and entitled METHODS FOR CANCER DETECTION USING MOLECULAR PATTERNS, which is incorporated by reference herein in its entirety. The one or more computational models 538 can also include one or more computational models described in patent application numberPCT / US2025 / 031219 filed May 28, 2025, and entitled MACHINE LEARNING CLASSIFICATION MODEL FOR CANCER DETECTION, which is incorporated by reference herein in its entirety. EXEMPLARY METHODS A. Determining an indication of a biological condition in a sample

[0465] In some embodiments a method includes producing a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected, calculating a background signal rate using the background data set, obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected, calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value, and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters, applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters, and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0466] In some embodiments a method includes producing a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set comprising sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one or more first subjects in which a tumor is not detected, calculating a background signal rate using the background data set; obtaining a tumor data set comprising sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected, calculating a tumor signal rate based on the tumor data set, comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff valueand a tumor signal rate above a tumor signal cutoff value, and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

[0467] In some embodiments a method includes applying a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides in the sample data set at each of a subset of CpG clusters, wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold, wherein applying the framework comprises aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject, and providing the indication of the tumor-related biological condition present in the test subject based on application of the framework.

[0468] In some embodiments, a computing apparatus may include a processor; and memory storing instructions that, when executed by the processor, configure the apparatus to produce a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected, calculating a background signal rate using the background data set, obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected, calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value, and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters, apply the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters, and provide an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0469] In some embodiments, a computing apparatus may include a processor; and memory storing instructions that, when executed by the processor, configure the apparatus to produce a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set comprising sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one or more first subjects in which a tumor is not detected, calculating a background signal rate using the background data set; obtaining a tumor data set comprising sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected, calculating a tumor signal rate based on the tumor data set, comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tumor signal rate above a tumor signal cutoff value, and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

[0470] In some embodiments, a computing apparatus may include a processor; and memory storing instructions that, when executed by the processor, configure the apparatus to apply a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides in the sample data set at each of a subset of CpG clusters, wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold, wherein applying the framework comprises aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject, and provide the indication of the tumor-related biological condition present in the test subject based on application of the framework.

[0471] In some embodiments, a non-transitory computer-readable storage medium, may include instructions that when executed by a computer, cause the computer to produce a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaininga background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected, calculating a background signal rate using the background data set, obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected, calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value, and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters, apply the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters, and provide an indication of a tumor-related biological condition present in the test subject based on application of the framework.

[0472] In some embodiments, a non-transitory computer-readable storage medium, may include instructions that when executed by a computer, cause the computer to produce a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set comprising sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one or more first subjects in which a tumor is not detected, calculating a background signal rate using the background data set; obtaining a tumor data set comprising sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected, calculating a tumor signal rate based on the tumor data set, comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tumor signal rate above a tumor signal cutoff value, and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

[0473] In some embodiments, a non-transitory computer-readable storage medium, may include instructions that when executed by a computer, cause the computer to apply a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides inthe sample data set at each of a subset of CpG clusters, wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold, wherein applying the framework comprises aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject, and provide the indication of the tumor-related biological condition present in the test subject based on application of the framework. B. Partitioning the sample into a plurality of subsamples

[0474] In some embodiments described herein, different forms of DNA (e.g., hypermethylated and hypomethylated DNA) are physically partitioned based on one or more characteristics of the DNA. This approach can be used to determine, for example, whether certain sites or regions are hypermethylated or hypomethylated. Partitioning can be performed before attaching adapters to DNA molecules in the sample, e.g., so as to facilitate including partition tags in the adapters. Partition tags can be used to identify which partition a molecule was found in. Following partitioning (and attachment of adapters if applicable), further steps such as amplification, target capture, and sequencing may be performed.

[0475] Methylation profiling can involve determining methylation patterns across different regions of the genome. For example, after partitioning molecules based on extent of methylation (e.g., relative number of methylated nucleobases per molecule) and further steps as discussed above including sequencing, the sequences of molecules in the different partitions can be mapped to a reference genome. This can show regions of the genome that, compared with other regions, are more highly methylated or are less highly methylated. In this way, genomic regions, in contrast to individual molecules, may differ in their extent of methylation.

[0476] Partitioning nucleic acid molecules in a sample can increase a rare signal, e.g., by enriching rare nucleic acid molecules that are more prevalent in one partition of the sample. For example, a genetic variation present in hypermethylated DNA but less (or not) present in hypomethylated DNA can be more easily detected by partitioning a sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple partitions of a sample, a multi-dimensional analysis of a single molecule can be performed and hence, greater sensitivity can be achieved. Partitioning may include physically partitioning nucleic acid molecules into partitions or subsamples based on the presence or absence of one or more methylated nucleobases. A sample may be partitioned into partitions or subsamples based on a characteristic that is indicative of differential gene expression or a disease state. A sample may be partitioned based on a characteristic, or combination thereof that provides a difference in signal between a normal and diseased state during analysis of nucleic acids, e.g., cell free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA) and cell free nucleic acids (cfNA).

[0477] In some embodiments, hypermethylation and / or hypomethylation variable epigenetic target regions are analyzed to determine whether they show differential methylation characteristic of particular immune cell types, such as rare immune cell types, tumor cells or cells of a type that does not normally contribute to the DNA sample being analyzed (such as cfDNA).

[0478] In some instances, heterogeneous DNA in a sample is partitioned into two or more partitions (e.g., at least 3, 4, 5, 6 or 7 partitions). In some embodiments, each partition is differentially tagged. Tagged partitions can then be pooled together for collective sample prep and / or sequencing. The partitioning-tagging-pooling steps can occur more than once, with each round of partitioning occurring based on a different characteristic (examples provided herein), and tagged using differential tags that are distinguished from other partitions and partitioning means. In other instances, the differentially tagged partitions are separately sequenced.

[0479] In some embodiments, sequence reads from differentially tagged and pooled DNA are obtained and analyzed in silico. Tags are used to sort reads from different partitions. Analysis to detect genetic variants can be performed on a partition-by-partition level, as well as whole nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants, such as CNV, SNV, indel, fusion in nucleic acids in each partition. In some instances, in silico analysis can include determining chromatin structure. For example, coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or nucleosome depleted region (NDR).

[0480] In some embodiments, partitioning is on the basis of one or more characteristics such as methylation. Molecules can be sorted according to other characteristics, such as sequence length, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteinsthat bind to DNA, using appropriate techniques as part of data analysis or partitioning as applicable. Resulting partitions can include one or more of the following nucleic acid forms: single- stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments and longer DNA fragments. In some embodiments, partitioning based on a cytosine modification (e.g., cytosine methylation) or methylation generally is performed and is optionally combined with at least one additional partitioning step, which may be based on any of the foregoing characteristics or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include presence or absence of methylation; level of methylation; type of methylation (e.g., 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins, such as histones. Alternatively, or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules devoid of nucleosomes. Alternatively, or additionally, a heterogeneous population of nucleic acids may be partitioned into single- stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or additionally, a heterogeneous population of nucleic acids may be partitioned based on nucleic acid length (e.g., molecules of up to 160 bp and molecules having a length of greater than 160 bp).

[0481] The agents used to partition populations of nucleic acids within a sample can be affinity agents, such as antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected e.g., by phage display to have specificity to a given target. In some embodiments, the agent used in the partitioning is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, such as a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is a product of a procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA of the sample. In some embodiments, the modified nucleobase may be a “converted nucleobase,” meaning that its base pairing specificity was changed by the procedure. For example, certain procedures convert unmethylated or unmodified cytosine to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modifiednucleobase in the context of DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, such as antibodies that recognize a modified nucleobase, which may be a modified cytosine, such as a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes a modified cytosine other than 5- methylcytosine, such as 5-carboxylcytosine (5caC). Alternative partitioning agents include methyl binding domain (MBDs) and methyl binding proteins (MBPs) as described herein, including proteins such as MeCP2.

[0482] Additional, non-limiting examples of partitioning agents are histone binding proteins which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides.

[0483] The binding of partitioning agents to particular nucleic acids and the partitioning of the nucleic acids into subsamples may occur to a certain extent or may occur in an essentially binary manner. In some instances, nucleic acids comprising a greater proportion of a certain modification bind to the agent at a greater extent than nucleic acids comprising a lesser proportion of the modification. Similarly, the partitioning may produce subsamples comprising greater and lesser proportions of nucleic acids comprising a certain modification. Alternatively, the partitioning may produce subsamples comprising essentially all or none of the nucleic acids comprising the modification. In all instances, various levels of modifications may be sequentially eluted from the partitioning agent.

[0484] In some embodiments, partitioning can comprise both binary partitioning and partitioning based on degree / level of modifications. For example, methylated fragments can be partitioned by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be partitioned from unmethylated fragments using methyl binding domain proteins (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific). Subsequently, additional partitioning may involve eluting fragments having different levels of methylation by adjusting the salt concentration in a solution with the methyl binding domain and bound fragments. As salt concentration increases, fragments having greater methylation levels are eluted.

[0485] In some instances, the final partitions are enriched in nucleic acids having different extents of modifications (overrepresentative or underrepresentative of modifications). Overrepresentation and underrepresentation can be defined by the number of modifications bornby a nucleic acid relative to the median number of modifications per strand in a population. For example, if the median number of 5-methylcytosine residues in nucleic acid in a sample is 2, a nucleic acid including more than two 5-methylcytosine residues is overrepresented in this modification and a nucleic acid with 1 or zero 5-methylcytosine residues is underrepresented. The effect of the affinity separation is to enrich for nucleic acids overrepresented in a modification in a bound phase and for nucleic acids underrepresented in a modification in an unbound phase (i.e., in solution). The nucleic acids in the bound phase can be eluted before subsequent processing.

[0486] When using MeDIP or MethylMiner®Methylated DNA Enrichment Kit (ThermoFisher Scientific) various levels of methylation can be partitioned using sequential elutions. For example, a hypomethylated partition (no methylation) can be separated from a methylated partition by contacting the nucleic acid population with the MBD from the kit, which is attached to magnetic beads. The beads are used to separate out the methylated nucleic acids from the non- methylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids having different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is once again used to separate higher level of methylated nucleic acids from those with lower level of methylation. The elution and magnetic separation steps can be repeated to create various partitions such as a hypomethylated partition (enriched in nucleic acids comprising no methylation), a methylated partition (enriched in nucleic acids comprising low levels of methylation), and a hyper methylated partition (enriched in nucleic acids comprising high levels of methylation).

[0487] In some methods, nucleic acids bound to an agent used for affinity separation- based partitioning are subjected to a wash step. The wash step washes off nucleic acids weakly bound to the affinity agent. Such nucleic acids can be enriched in nucleic acids having the modification to an extent close to the mean or median (i.e., intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not binding to the solid phase on initial contacting of the sample with the agent).

[0488] The affinity separation results in at least two, and sometimes three or more partitions of nucleic acids with different extents of a modification. While the partitions are still separate, the nucleic acids of at least one partition, and usually two or three (or more) partitions are linked to nucleic acid tags, usually provided as components of adapters, with the nucleic acids in different partitions receiving different tags that distinguish members of one partition from another. The tags linked to nucleic acid molecules of the same partition can be the same or different from one another. But if different from one another, the tags may have part of their code in common so as to identify the molecules to which they are attached as being of a particular partition.

[0489] For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.

[0490] In some embodiments, the nucleic acid molecules can be fractionated into different partitions based on the nucleic acid molecules that are bound to a specific protein or a fragment thereof and those that are not bound to that specific protein or fragment thereof.

[0491] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein- DNA complexes can be fractionated based on a specific property of a protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation) or enzymatic activity. Examples of proteins which may bind to DNA and serve as a basis for fractionation may include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate the nucleic acid molecules based on protein bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein bound regions include, but are not limited to, SDS-PAGE, chromatin-immuno-precipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).

[0492] In some embodiments, the partitioning of the sample into a plurality of subsamples is performed by contacting the nucleic acids with an antibody that recognizes a modified nucleobase in the DNA, which may be is a modified cytosine or a product of the procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA of the sample. In some embodiments, the modified nucleobase is 5mC. In some embodiments, the modified nucleobase is 5caC. In some embodiments, the modified nucleobase is dihydrouracil(DHU). In some embodiments, the antibody that recognizes a modified nucleobase in the DNA is used to partition single-stranded DNA.

[0493] In some embodiments, the partitioning is performed by contacting the nucleic acids with a methyl binding domain (“MBD”) of a methyl binding protein (“MBP”). In some such embodiments, the nucleic acids are contacted with an entire MBP. In some embodiments, an MBD binds to 5-methylcytosine (5mC), and an MBP comprises an MBD and is referred to interchangeably herein as a methyl binding protein or a methyl binding domain protein. In some embodiments, an MBD binds to 5mC and 5hmC. In some embodiments, MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by eluting fractions by increasing the NaCl concentration.

[0494] In some embodiments, bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This may be performed instead of or in addition to elution steps using NaCl as discussed above.

[0495] Examples of agents that recognize a modified nucleobase contemplated herein include, but are not limited to:

[0496] (a) MeCP2 is a protein that preferentially binds to 5-methyl-cytosine over unmodified cytosine.

[0497] (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind to 5- hydroxymethyl-cytosine over unmodified cytosine.

[0498] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind to 5-formyl cytosine over unmodified cytosine (Iurlaro et al., Genome Biol.14: R119 (2013)).

[0499] (d) Antibodies specific to one or more methylated or modified nucleobases or conversion products thereof, such as 5mC, 5caC, or DHU.

[0500] In general, elution is a function of the number of modifications, such as the number of methylated sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCl concentration. Salt concentration can range from about 100 nm to about 2500 mM NaCl. In one embodiment, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and comprising a molecule comprising an agent that recognizes a modified nucleobase, whichmolecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration a population of molecules will bind to the agent and a population will remain unbound. The unbound population can be separated as a “hypomethylated” population. For example, a first partition enriched in hypomethylated form of DNA is that which remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. A second partition enriched in intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. This is also separated from the sample. A third partition enriched in hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0501] In some embodiments, a monoclonal antibody raised against 5-methylcytidine (5mC) is used to purify methylated DNA. DNA is denatured, e.g., at 95°C in order to yield single- stranded DNA fragments. Protein G coupled to standard or magnetic beads as well as washes following incubation with the anti-5mC antibody are used to immunoprecipitate DNA bound to the antibody. Such DNA may then be eluted. Partitions may comprise unprecipitated DNA and one or more partitions eluted from the beads.

[0502] In some embodiments, sample DNA (e.g., between 5 and 200 ng) is mixed with methyl binding domain (MBD) buffer and magnetic beads conjugated with MBD proteins and incubated overnight. Methylated DNA (hypermethylated DNA) binds the MBD protein on the magnetic beads during this incubation. Non-methylated (hypomethylated DNA) or less methylated DNA (intermediately methylated) is washed away from the beads with buffers containing increasing concentrations of salt. For example, one, two, or more fractions containing non-methylated, hypomethylated, and / or intermediately methylated DNA may be obtained from such washes. Finally, a high salt buffer is used to elute the heavily methylated DNA (hypermethylated DNA) from the MBD protein. In some embodiments, these washes result in three partitions (hypomethylated partition, intermediately methylated fraction and hypermethylated partition) of DNA having increasing levels of methylation.

[0503] In some embodiments, partitioning procedures may result in imperfect sorting of DNA molecules among the subsamples. For example, a minority of the molecules in an unmethylated or hypomethylated subsample may be highly modified (e.g., hypermethylated), and / or a minority of the molecules in a hypermethylated subsample may be unmodified or mostlyunmodified (e.g., unmethylated or mostly unmethylated). Such molecules are considered nonspecifically partitioned.

[0504] In some embodiments, nonspecifically partitioned molecules are removed using a methylation-dependent nuclease, e.g., a methylation dependent restriction enzyme (MDRE), digesting / cleaving the DNA where the restriction enzyme (RE) recognition site contains a methylated nucleotide but not cleaving the DNA where the restriction enzyme (RE) recognition site contains an unmethylated nucleotide. In some embodiments, nonspecifically partitioned molecules are removed using a methylation sensitive nuclease, e.g., a methylation sensitive restriction enzyme (MSRE), digesting / cleaving the DNA where the restriction enzyme (RE) recognition site contains an unmethylated nucleotide but not cleaving the DNA where the restriction enzyme (RE) recognition site contains a methylated nucleotide. For example, in some embodiments, a hypomethylated subsample is contacted with a methylation-dependent nuclease, such as a methylation-dependent restriction enzyme, thereby degrading nonspecifically partitioned DNA, e.g., methylated DNA, in the subsample. Alternatively, or in addition, a hypermethylated subsample is contacted with a methylation-sensitive endonuclease, such as a methylation-sensitive restriction enzyme, thereby degrading nonspecifically partitioned DNA in the subsample.

[0505] Degradation of nonspecifically partitioned DNA in one or more partitioned subsamples may improve the performance of methods that rely on accurate partitioning of DNA on the basis of a cytosine modification. For example, such degradation may provide improved sensitivity and / or simplify downstream analyses. In some embodiments, partitioning DNA on the basis of a modification, such as methylation, then removing nonspecifically partitioned DNA using MDREs and / or MSREs as described herein provides improved efficiency and / or cost over DNA analysis methods comprising procedures that affect a first nucleobase differently from a second nucleobase, such as bisulfite sequencing or bisulfite conversion.

[0506] In some embodiments, one or more nucleases are used to degrade nonspecifically partitioned DNA molecules. In some embodiments, a subsample is contacted with a plurality of nucleases. The subsample may be contacted with the nucleases sequentially or simultaneously. Simultaneous use of nucleases may be advantageous when the nucleases are active under similar conditions (e.g., buffer composition) to avoid unnecessary sample manipulation. Contacting a subsample with more than one methylation-dependent restriction enzyme can morecompletely degrade nonspecifically partitioned hypermethylated DNA. Contacting a subsample with more than one methylation-sensitive restriction enzyme can more completely degrade nonspecifically partitioned hypomethylated and / or unmethylated DNA.

[0507] In some embodiments, a methylation-dependent nuclease comprises one or more of MspJI, LpnPI, FspEI, or McrBC. In some embodiments, at least two methylation-dependent nucleases are used. In some embodiments, at least three methylation-dependent nucleases are used.

[0508] In some embodiments, a methylation-sensitive nuclease comprises one or more of AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, and SnaBI. In some embodiments, at least two methylation-sensitive nucleases are used. In some embodiments, at least three methylation-sensitive nucleases are used. In some embodiments, the methylation-sensitive nucleases comprise BstUI and HpaII. In some embodiments, the two methylation-sensitive nucleases comprise HhaI and AccII. In some embodiments, the methylation-sensitive nucleases comprise BstUI, HpaII and Hin6I.

[0509] In some embodiments, the partitions of DNA are desalted and concentrated in preparation for enzymatic steps of library preparation. C. Adapter Ligation

[0510] In some embodiments, adapters are added to the DNA. This may be done concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer (where PCR is used, this can be referred to as library prep-PCR or LP-PCR). In some embodiments, adapters are added by other approaches, such as ligation. In some such methods, prior to partitioning or prior to capturing, first adapters are added to the nucleic acids by ligation to the 3’ ends thereof, which may include ligation to single-stranded DNA. The adapter can be used as a priming site for second-strand synthesis, e.g., using a universal primer and a DNA polymerase. A second adapter can then be ligated to at least the 3’ end of the second strand of the now double-stranded molecule. In some embodiments, the first adapter comprises an affinity tag, such as biotin, and nucleic acid ligated to the first adapter is bound to a solid support (e.g., bead), which may comprise a binding partner for the affinity tag such as streptavidin. For further discussion of a related procedure, see Gansauge et al., Nature Protocols 8:737-748 (2013).Commercial kits for sequencing library preparation compatible with single-stranded nucleic acids are available, e.g., the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, nucleic acids are amplified.

[0511] Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites, but preferably adapters include the same primer binding site.

[0512] In some embodiments, following attachment of adapters, the nucleic acids are subject to amplification. The amplification can use, e.g., universal primers that recognize primer binding sites in the adapters.

[0513] In some embodiments, following attachment of adapters, the DNA is partitioned, comprising contacting the DNA with an agent that preferentially binds to nucleic acids bearing an epigenetic modification. The nucleic acids are partitioned into at least two subsamples differing in the extent to which the nucleic acids bear the modification from binding to the agents. For example, if the agent has affinity for nucleic acids bearing the modification, nucleic acids overrepresented in the modification (compared with median representation in the population) preferentially bind to the agent, whereas nucleic acids underrepresented for the modification do not bind or are more easily eluted from the agent. The nucleic acids can then be amplified from primers binding to the primer binding sites within the adapters. Partitioning may be performed instead before adapter attachment, in which case the adapters may comprise differential tags that include a component that identifies which partition a molecule occurred in.

[0214] In some embodiments, the nucleic acids are linked at both ends to Y-shaped adapters including primer binding sites and tags. The molecules are amplified. D. Tagging

[0514] “Tagging” DNA molecules is a procedure in which a tag is attached to or associated with the DNA molecules. Tags can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. For example, molecules can bear a sample tag (which distinguishes molecules in one sample from those in a different sample) or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from one another (in both unique and non-unique tagging scenarios). For methods that involve apartitioning step, a partition tag (which distinguishes molecules in one partition from those in a different partition) may be included. In some embodiments, adapters added to DNA molecules comprise tags. In certain embodiments, a tag can comprise one or a combination of barcodes. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or can have sequences having a certain hamming distance, as desired for the specific purpose. So, for example, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally, or alternatively, for different partitions and / or samples, different sets of molecular barcodes, or molecular tags can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve to identify the partition and / or sample to which they correspond based the set of which they are a member.

[0515] In some embodiments, two or more partitions, e.g., each partition, is / are differentially tagged. Tags can be used to label the individual polynucleotide population partitions so as to correlate the tag (or tags) with a specific partition. Alternatively, tags can be used in embodiments that do not employ a partitioning step. In some embodiments, a single tag can be used to label a specific partition. In some embodiments, multiple different tags can be used to label a specific partition. In embodiments employing multiple different tags to label a specific partition, the set of tags used to label one partition can be readily differentiated for the set of tags used to label other partitions. In some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations, for example as in Kinde et al., Proc Nat’l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS ONE, 11 : e0146638 (2016)) or used as non-unique molecule identifiers, for example as described in US Pat. No.9,598,731. Similarly, in some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as non-unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).

[0516] In some embodiments, partition tagging comprises tagging molecules in each partition with a partition tag. After re-combining partitions (e.g., to reduce the number ofsequencing runs needed and avoid unnecessary cost) and sequencing molecules, the partition tags identify the source partition. In some embodiments, the partition tags can serve as identifiers of the source partition and the molecule, i.e., different partitions are tagged with different sets of molecular tags, e.g., comprised of a pair of barcodes. In this way, the one or more molecular barcodes attached to the molecule indicates the source partition as well as being useful to distinguish molecules within a partition. For example, a first set of 35 barcodes can be used to tag molecules in a first partition, while a second set of 35 barcodes can be used tag molecules in a second partition.

[0517] In some embodiments, after partitioning and tagging with partition tags, the molecules may be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, e.g., in a step subsequent to addition of partition tags and pooling. Sample tags can facilitate pooling material generated from multiple samples for sequencing in a single sequencing run.

[0518] Alternatively, in some embodiments, partition tags may be correlated to the sample as well as the partition. As a simple example, a first tag can indicate a first partition of a first sample; a second tag can indicate a second partition of the first sample; a third tag can indicate a first partition of a second sample; and a fourth tag can indicate a second partition of the second sample.

[0519] While tags may be attached to molecules already partitioned based on one or more characteristics, the final tagged molecules in the library may no longer possess that characteristic. For example, while single stranded DNA molecules may be partitioned and tagged, the final tagged molecules in the library are likely to be double stranded. Similarly, while DNA may be subject to partition based on different levels of methylation, in the final library, tagged molecules derived from these molecules are likely to be unmethylated. Accordingly, the tag attached to molecule in the library typically indicates the characteristic of the “parent molecule” from which the ultimate tagged molecule is derived, not necessarily to characteristic of the tagged molecule, itself.

[0520] As an example, barcodes 1, 2, 3, 4, etc. are used to tag and label molecules in the first partition; barcodes A, B, C, D, etc. are used to tag and label molecules in the second partition; and barcodes a, b, c, d, etc. are used to tag and label molecules in the third partition. Differentially tagged partitions can be pooled prior to sequencing. Differentially tagged partitions can beseparately sequenced or sequenced together concurrently, e.g., in the same flow cell of an Illumina sequencer.

[0521] After sequencing, analysis of reads can be performed on a partition-by-partition level, as well as a whole DNA population level. Tags are used to sort reads from different partitions. Analysis can include in silico analysis to determine genetic and epigenetic variation (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinates length, coverage, and / or copy number. In some embodiments, higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or a nucleosome depleted region (NDR). E. Enriching / Capturing step; Amplification

[0522] Methods disclosed herein can comprise capturing DNA, such as cfDNA target regions. In some embodiments, the capturing comprises contacting the DNA with probes (e.g., oligonucleotides) specific for the target regions. Enrichment or capture may be performed on any sample or subsample described herein using any suitable approach known in the art.

[0523] In some embodiments, enrichment or capture is performed after attachment of adapters to sample molecules. In some embodiments, enrichment or capture is performed after a partitioning step. In some embodiments, enrichment or capture is performed after an amplification step. In some embodiments, sample molecules are partitioned, then adapters are attached, then sample molecules are amplified, and then the amplified molecules are subjected to enrichment or capture. The enriched or captured molecules may then be subjected to another amplification and then sequenced.

[0524] In some embodiments, the probes specific for the target regions comprise a capture moiety that facilitates the enrichment or capture of the DNA hybridized to the probes. In some embodiments, the capture moiety is biotin. In some such embodiments, streptavidin attached to a solid support, such as magnetic beads, is used to bind to the biotin. Nonspecifically bound DNA that does not comprise a target region is washed away from the captured DNA. In some embodiments, DNA is then dissociated from the probes and eluted from the solid support using salt washes or buffers comprising another DNA denaturing agent. In some embodiments, the probes are also eluted from the solid support by, e.g., disrupting the biotin-streptavidin interaction. In some embodiments, captured DNA is amplified following elution from the solidsupport. In some such embodiments, DNA comprising adapters is amplified using PCR primers that anneal to the adapters. In some embodiments, captured DNA is amplified while attached to the solid support. In some such embodiments, the amplification comprises use of a PCR primer that anneals to a sequence within an adapter and a PCR primer that anneals to a sequence within a probe annealed to the target region of the DNA.

[0525] In some embodiments, the methods herein comprise enriching for or capturing DNA comprising epigenetic and / or sequence-variable target regions. Such regions may be captured from an aliquot of a sample (e.g., a sample that has undergone attachment of adapters and amplification), while the step of partitioning the DNA with an agent that recognizes a modified cytosine, such as methyl cytosine, is performed on a separate aliquot of the sample. Enriching for or capturing DNA comprising epigenetic and / or sequence-variable target regions may comprise contacting the DNA with a first or second set of target-specific probes. Such target-specific probes may have any of the features described herein for sets of target-specific probes, including but not limited to in the embodiments set forth above and the sections relating to probes below. Capturing may be performed on one or more subsamples prepared during methods disclosed herein. In some embodiments, DNA is captured from the first subsample or the second subsample, e.g., the first subsample and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture. Exemplary methods for capturing DNA comprising epigenetic and / or sequence-variable target regions can be found in, e.g., WO 2020 / 160414, which is hereby incorporated by reference.

[0526] The capturing step may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization. In some embodiments, complexes of target-specific probes and DNA are formed.

[0527] In some embodiments, methods described herein comprise capturing a plurality of sets of target regions of cfDNA obtained from a subject. The target regions may comprise differences depending on whether they originated from a tumor or from healthy cells or from a certain cell type. The capturing step produces a captured set of cfDNA molecules. In some embodiments, cfDNA molecules corresponding to a sequence-variable target region set are captured at a greater capture yield in the captured set of cfDNA molecules than cfDNA moleculescorresponding to an epigenetic target region set. In some embodiments, a method described herein comprises contacting cfDNA obtained from a subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set. For additional discussion of capturing steps, capture yields, and related aspects, see W02020 / 160414, which is incorporated herein by reference for all purposes.

[0528] It can be beneficial to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequence- variable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test for perturbation of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated partitions) is generally less than the volume of data needed to determine the presence or absence of cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).

[0529] In some embodiments, the DNA is amplified. In some embodiments, amplification is performed before the capturing step. In some embodiments, amplification is performed after the capturing step. In some embodiments, amplification is performed before and after the capturing step. In various embodiments, the methods further comprise sequencing the captured DNA, e.g., to different degrees of sequencing depth for the epigenetic and sequence-variable target region sets, consistent with the discussion herein.

[0530] In some embodiments, a capturing step is performed with probes for a sequence- variable target region set and probes for an epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence-variable target region set is greater that the concentration of the probes for the epigenetic target region set.

[0531] Alternatively, a capturing step is performed with a sequence-variable target region probe set in a first vessel and with an epigenetic target region probe set in a second vessel, or acontacting step is performed with a sequence-variable target region probe set at a first time and a first vessel and an epigenetic target region probe set at a second time before or after the first time. This approach allows for preparation of separate first and second compositions comprising captured DNA corresponding to a sequence-variable target region set and captured DNA corresponding to an epigenetic target region set. The compositions can be processed separately as desired (e.g., to partition based on methylation as described herein) and pooled in appropriate proportions to provide material for further processing and analysis such as sequencing.

[0532] In some embodiments, adapters are included in the DNA as described herein. In some embodiments, tags, which may be or include barcodes, are included in the DNA. In some embodiments, such tags are included in adapters. Tags can facilitate identification of the origin of a nucleic acid. For example, barcodes can be used to allow the origin (e.g., subject) whence the DNA came to be identified following pooling of a plurality of samples for parallel sequencing. This may be done concurrently with an amplification procedure, e.g., by providing the barcodes in a 5’ portion of a primer, e.g., as described herein. In some embodiments, adapters and tags / barcodes are provided by the same primer or primer set. For example, the barcode may be located 3’ of the adapter and 5’ of the target-hybridizing portion of the primer. Alternatively, barcodes can be added by other approaches, such as ligation, optionally together with adapters in the same ligation substrate.

[0533] Additional details regarding amplification, tags, and barcodes are discussed herein, which can be combined to the extent practicable with any of these embodiments. F. Procedures that affect a first nucleobase in the DNA differently from a second nucleobase in the DNA or methylation-sensitive conversion methods

[0534] In some embodiments, methods disclosed herein comprise a step of subjecting DNA, or a subsample thereof, to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase such that the base pairing specificity of the converted nucleobase is altered. In some embodiments, DNA is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA before library preparation using the DNA, beforea first amplification of the DNA, before dividing the DNA into a plurality of subsamples, or any combination thereof. In certain embodiments, the DNA is subjected to the procedure before or after contacting the DNA with a methylation-sensitive nuclease.

[0535] In some embodiments, the procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA is performed prior to the sequencing and / or (a) prior to or after the selectively depleting the target nucleic acid comprising the wild-type sequence, the target nucleic acid comprising the converted nucleotide, or the target nucleic acid that does not comprise the converted nucleotide; (b) prior to the amplifying the selectively digested population of target nucleic acids; (c) prior to or after the partitioning the population of target nucleic acids into a plurality of subsamples; and / or (d) prior to or after a step of enriching for one or more sets of target regions of DNA.

[0536] In some embodiments, if the first nucleobase is a modified or unmodified adenine, then the second nucleobase is a modified or unmodified adenine; if the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine; if the first nucleobase is a modified or unmodified guanine, then the second nucleobase is a modified or unmodified guanine; and if the first nucleobase is a modified or unmodified thymine, then the second nucleobase is a modified or unmodified thymine (where modified and unmodified uracil are encompassed within modified thymine for the purpose of this step).

[0537] In some embodiments, the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine. For example, first nucleobase may comprise unmodified cytosine (C) and the second nucleobase may comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.

[0538] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g. 5-formyl cytosine (fC) or 5-carboxylcytosine (caC)) to uracil whereas other modified cytosines (e.g., 5- methylcytosine, 5-hydroxylmethylcystosine) are not converted. Thus, where bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, 5-formyl cytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of bisulfite- treated DNA identifies positions that are read as cytosine as being mC or hmC positions. Meanwhile, positions that are read as T are identified as being T or a bisulfite-susceptible form of C, such as unmodified cytosine, 5-formyl cytosine, or 5-carboxylcytosine. Performing bisulfite conversion, such as on a DNA sample as described herein, facilitates identifying positions containing mC or hmC using the sequence reads obtained from the exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun.2018; 9: 5068.

[0539] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite susceptible, followed by bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises mC. Sequencing of Ox-BS converted DNA identifies positions that are read as cytosine as being mC positions. Meanwhile, positions that are read as T are identified as being T, hmC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or hmC. Performing Ox-BS conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing mC using the sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.

[0540] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized in advance of bisulfite treatment, so that positions originally occupied by mC are converted to U while positions originally occupied by hmC remain as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, β-glucosyl transferase can be used to protect hmC (forming 5-glucosylhydroxymethylcytosine (ghmC)), then a TET protein such as mTet1 can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while ghmC remains unaffected.

[0541] Alternatively, a carbamoyltransferase enzyme, such as 5-hydroxymethylcytosine carbamoyltransferase as described in Yang et al., Bio-protocol, 2023; 12(17): e4496, can be usedto protect hmC (by converting hmC to 5-carbamoyloxymethylcytosine (5cmC)), then a TET protein such as mTet1 or a TET2 comprising a T1372S mutation, can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while 5cmC remains unaffected. Thus, when TAB conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, mC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of TAB-converted DNA identifies positions that are read as cytosine as being hmC positions. Meanwhile, positions that are read as T are identified as being T, mC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or caC. Performing TAB conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing hmC using the sequence reads obtained from the sample.

[0542] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2- picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic- borane conversion with a substituted borane reducing agent conversion, a TET protein is used to convert mC and hmC to caC, without affecting unmodified C. caC, and fC if present, are then converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting unmodified C. See, e.g., Liu et al., Nature Biotechnology 2019; 37:424–429 (e.g., at Supplementary Fig.1 and Supplementary Note 7). Thus, when this type of conversion is used, the first nucleobase comprises one or more of 5mC, 5fC, 5caC, or 5hmC, and the second nucleobase comprises unmodified cytosine. DHU is read as a T in sequencing. Thus, when this type of conversion is used, the first nucleobase comprises one or more of mC, fC, caC, or hmC, and the second nucleobase comprises unmodified cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T, mC, fC, caC, or hmC. Performing TAP conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing unmodified C using the sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), described in further detail in Liu et al.2019, supra.

[0543] Alternatively, protection of hmC (e.g., using βGT or 5-hydroxymethylcytosine carbamoyltransferase) can be combined with Tet-assisted conversion with a substituted borane reducing agent, e.g. as described above. In this method (TAPS-β), 5hmC can be protected from conversion, for example through glucosylation using β-glucosyl transferase (βGT), forming 5- glucosylhydroxymethylcytosine (5ghmC), or through carbamoylation using 5- hydroxymethylcytosine carbamoyltransferase, forming 5cmC. This is described in Yu et al., Cell 2012; 149: 1368-80. Treatment with a TET protein, such as mTet1 or a TET2 comprising a T1372S mutation, then converts mC to caC but does not convert C, 5ghmC, or 5cmC.5caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting ghmC, 5cmC, or unmodified C. Thus, when Tet-assisted conversion with a substituted borane reducing agent is used, the first nucleobase comprises mC, and the second nucleobase comprises one or more of unmodified cytosine or hmC, such as unmodified cytosine and optionally hmC, fC, and / or caC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either hmC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or mC. Performing TAPSβ conversion, such as on a DNA sample as described herein, thus facilitates distinguishing positions containing unmodified C or hmC on the one hand from positions containing mC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424–429.5-hydroxymethylcytosine carbamoyltransferase is described in Yang et al., Bio-protocol, 2023; 12(17): e4496.

[0544] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises chemical-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2- picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In chemical- assisted conversion with a substituted borane reducing agent, an oxidizing agent such as potassium perruthenate (KRuO4) (also suitable for use in ox-BS conversion) is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane converts fC and caC to DHU but does not affect mC or unmodified C. Thus, when this type of conversion is used, the first nucleobase comprises one or more of hmC, fC, and caC, and the second nucleobase comprisesone or more of unmodified cytosine or mC, such as unmodified cytosine and optionally mC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either mC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or hmC. Performing this type of conversion, such as on a DNA sample as described herein, thus facilitates distinguishing positions containing unmodified C or mC on the one hand from positions containing hmC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424–429.

[0545] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme such as APOBEC3A (A3A) is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when ACE conversion is used, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of ACE-converted DNA identifies positions that are read as cytosine as being hmC, fC, or caC positions. Meanwhile, positions that are read as T are identified as being T, unmodified C, or mC. Performing ACE conversion on a DNA sample as described herein thus facilitates distinguishing positions containing hmC from positions containing mC or unmodified C using the sequence reads obtained from the sample. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083–1090.

[0546] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT or 5- hydroxymethylcytosine carbamoyltransferase (described in Yang et al., Bio-protocol, 2023; 12(17): e4496) can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosines converting them to uracils.

[0547] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using a non-specific, modification-sensitive double-stranded DNA deaminase, e.g., as in SEM-seq. See, e.g., Vaisvila et al. (2023) Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA. bioRxiv; DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047v1. SEM-Seq employs a non- specific, modification-sensitive double-stranded DNA deaminase (MsddA) in a nondestructive single-enzyme 5-methylctyosine sequencing (SEM-seq) method that deaminates unmodified cytosines. Accordingly, SEM-seq does not require the TET2 and T4-βGT or 5- hydroxymethylcytosine carbamoyltransferase protection and denaturing steps that are of use, e.g., in APOEC3A-based protocols. Additionally, MsddA does not deaminate 5-formylated cytosines (5fC) or 5-carboxylated cytosines (5caC). In SEM-seq, unmodified cytosines in the DNA are deaminated to uracil and is read as “T” during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as “C” during sequencing. Cytosines that are read as thymines are identified as unmodified (e.g., unmethylated) cytosines or as thymines in the DNA. Performing SEM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using MsddA.

[0548] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample converts a modified nucleoside. In some embodiments, the conversion procedure which converts a modified nucleosides comprises enzymatic conversion, such as DM-seq, for example, as described in WO2023 / 288222A1. In DM-seq, unmodified cytosines in the DNA are enzymatically protected from a subsequent deamination step wherein 5mC in 5mCpG is converted to T. The enzymatically protected unmodified (e.g., unmethylated) cytosines are not converted and are read as “C” during sequencing. Cytosines that are read as thymines (in a CpG context) are identified as methylated cytosines in the DNA. Thus, when this type of conversion is used, the first nucleobase comprises unmodified (such as unmethylated) cytosine, and the second nucleobase comprises modified(such as methylated) cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained.

[0549] Exemplary cytosine deaminases for use herein include APOBEC enzymes, for example, APOBEC3A. Generally, AID / APOBEC family DNA deaminase enzymes such as APOBEC3A (A3A) are used to deaminate (unprotected) unmodified cytosine and 5mC. For an exemplary description of APOBEC conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083–1090.

[0550] The enzymatic protection of unmodified cytosines in the DNA comprises addition of a protective group to the unmodified cytosines. Such protective groups can comprise an alkyl group, an alkyne group, a carboxyl group, a carboxyalkyl group, an amino group, a hydroxymethyl group, a glucosyl group, a glucosylhydroxymethyl group, an isopropyl group, or a dye. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds the protective group to unmodified cytosines. The term methyltransferase is used broadly herein to refer to enzymes capable of transferring a methyl or substituted methyl (e.g.,carboxymethyl) to a substrate (e.g., a cytosine in a nucleic acid). In some embodiments, the DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, e.g., WO2021 / 236778A2. In particular embodiments, the CxMTase can facilitate the addition of a protective carboxymethyl group to an unmethylated cytosine. In some embodiments, the unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of the cytosine during a deamination step (such as a deamination step using an APOBEC enzyme, such as A3A). Substituted methyl or carboxymethyl donors useful in the disclosed methods include but are not limited to, S-adenosyl-L-methionine (SAM) analogs, optionally wherein the SAM analog is carboxy-S-adenosyl-L-methionine (CxSAM). SAM analogs are described, for example, in WO2022 / 197593A1. The MTase may be, for example, a CpG methyltransferase from Spiroplasma sp. strain MQ1 (M.SssI), DNA-methyltransferase 1 (DNMT1), DNA- methyltransferase 3 alpha (DNMT3A), DNA-methyltransferase 3 beta (DNMT3B), or DNA adenine methyltransferase (Dam). The CxMTase may be a CpG methyltransferase fromMycoplasma penetrans (M.MpeI). In a particular embodiment, the methyltransferase enzyme is a variant of M.MpeI, wherein the amino acid corresponding to position 374 is R or K.

[0551] In one embodiment, the methyltransferase enzyme is a variant of M.MpeI having an N374R substitution or an N374K substitution. The methyltransferase variant can further comprise one or more amino acid substitutions selected from a) substitution of one or both residues T300 and E305 with S, A, G, Q, D, or N; b) substitution of one or more residues A323, N306, and Y299 with a positively charged amino acid selected from K, R or H; and / or c) substitution of S323 with A, G, K, R or H, which may enhance the activity of the enzyme.

[0552] Optionally, the conversion procedure further includes enzymatic protection of 5hmCs, such as by glucosylation of the 5hmCs (e.g., using βGT) or by carbamoylation of the 5hmCs (e.g., using 5-hydroxymethylcytosine carbamoyltransferase), in the DNA prior to the deamination of unprotected modified cytosines. In this method, 5hmC can be protected from conversion, for example through glucosylation using β-glucosyl transferase (βGT), forming (5- glucosylhydroxymethylcytosine) 5ghmC, or through carbamoylation using 5- hydroxymethylcytosine carbamoyltransferase, forming 5cmC. This is described, for example, in Yu et al., Cell 2012; 149: 1368-80, and in Yang et al., Bio-protocol, 2023; 12(17): e4496. Glucosylation or carbamoylation of 5hmC can reduce or eliminate deamination of 5hmC by a deaminase such as APOBEC3A. Treatment with an MTase or CxMTase then adds a protecting group to unmodified (unmethylated) cytosines in the DNA.5mC (but not protected, unmodified cytosine and not 5ghmC or 5cmC) is then deaminated (converted to T in the case of 5mC) by treatment with a deaminase, for example, an APOBEC enzyme (such as APOBEC3A). Sequencing of the converted DNA identifies positions that are read as cytosine as being either 5hmC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion with glucosylation of 5hmC on a sample as described herein thus facilitates distinguishing positions containing unmodified C or 5hmC on the one hand from positions containing 5mC using the sequence reads obtained.

[0553] Also provided herein are methods in which alternative base conversion schemes are used. For example, unmethylated cytosines can be left intact while methylated cytosines and hydroxymethylcytosines are converted to a base read as a thymine (e.g., uracil, thymine, or dihydrouracil).

[0554] In some embodiments, methylating a cytosine in at least one first complementary strand or second complementary strand comprises contacting the cytosine with a methyltransferase such as DNMT1 or DNMT5. In such embodiments, the step of oxidizing a 5- hydroxymethylated cytosine to 5-formylcytosine (such as by contacting the 5-hydroxymethyl cytosine in a first strand and a second strand with KRuO4) can be optional.

[0555] In some embodiments, converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine comprises oxidizing a hydroxymethyl cytosine, e.g., the hydroxymethyl cytosine is oxidized to formylcytosine. In some embodiments, oxidizing the hydroxymethyl cytosine to formylcytosine comprises contacting the hydroxymethyl cytosine with a ruthenate, such as potassium ruthenate (KRuO4).

[0556] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil. In any such embodiments, amplification methods may comprise uracil- and / or dihydrouracil-tolerant amplification methods, such as PCR using a uracil- and / or dihydrouracil- tolerant DNA polymerase.

[0557] In some embodiments, the method comprises converting a formylcytosine and / or a methylcytosine to carboxylcytosine as part of converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine. For example, converting the formylcytosine and / or the methylcytosine to carboxylcytosine can comprise contacting the formylcytosine and / or the methylcytosine with a TET enzyme, such as TET1, TET2, TET3, or a TET2 comprising a T1372S mutation. In some embodiments, the method comprises reducing the carboxylcytosine as part of converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine, and / or the carboxylcytosine is reduced to dihydrouracil. In some embodiments, reducing the carboxylcytosine comprises contacting the carboxylcytosine with a borane or borohydride reducing agent.

[0558] In some embodiments, the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBH3CN), lithium borohydride (LiBH4), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine,diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof.

[0559] Various TET enzymes may be used in the disclosed methods as appropriate. In some embodiments, the one or more TET enzymes comprise TETv. TETv is described in US Patent 10,260,088. In some embodiments, the one or more TET enzymes comprise TETcd. TETcd is described in US Patent 10,260,088. In some embodiments, the one or more TET enzymes comprise TET1. In some embodiments, the one or more TET enzymes comprise TET2. TET2 may be expressed and used as a fragment comprising TET2 residues 1129-1480 joined to TET2 residues 1844-1936 by a linker as described, e.g., in US Patent 10,961,525. In some embodiments, the one or more TET enzymes comprise TET1 and TET2. In some embodiments, the one or more TET enzymes comprise a V1900 TET mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET mutant. In some embodiments, the one or more TET enzymes comprise a V1900 TET2 mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET2 mutant. It can be beneficial to use a TET enzyme that maximizes formation of 5-carboxylcytosine (5-caC) relative to less oxidized modified cytosines, particularly 5-formylcytosine, because 5-caC is not a substrate for enzymatic deamination, e.g., by APOBEC enzymes such as APOBEC3A. Maximizing formation of 5-caC thus reduces the risk of false calls in which a base is identified as unmethylated because it underwent deamination even though it was methylated (or hydroxymethylated) in the original sample. Accordingly, in some embodiments, the TET enzyme comprises a mutation that increases formation of 5-caC. In some embodiments, the one or more TET enzymes comprise a TET2 enzyme comprising a T1372S mutation, such as TET2-CS- T1372S and TET2-CD-T1372S. A TET2 comprising a T1372S mutation is described in US Patent 10,961,525 and may be expressed and used as a fragment comprising TET2 residues 1129-1480 joined to TET2 residues 1844-1936 by a linker. Position 1372 of TET2 corresponds to position 258 of SEQ ID NO: 21 (wild type TET2 catalytic domain) of US Patent 10,961,525. Thus, the sequence of a T1372S TET2 catalytic domain may be obtained by changing the threonine at position 258 of SEQ ID NO: 21 of US Patent 10,961,525 to serine. TET2 comprising a T1372S mutation is also described in Liu et al., Nat Chem Biol. 2017 February; 13(2): 181–187. As demonstrated in Liu et al., TET2 comprising a T1372S mutation can more efficiently oxidize 5mC to produce 5-carboxylcytosine (5caC) than other versions of TET2 such as TET2 lacking a T1372S mutation. In some embodiments, the TET2 enzyme is a human TET2 enzymecomprising a T1372S mutation. Exemplary mutations are set forth above. “A mutation that increases formation of 5-caC” means that the TET enzyme having the mutation produces more 5-caC than a TET enzyme that lacks the mutation but is otherwise identical.5-caC production can be measured as described, e.g., in Liu et al., Nat Chem Biol 13:181-187 (2017) (see Online Methods section, TET reactions in vitro subsection, “driving” conditions). Any variants and / or mutants described in Liu et al. (2017) can be used in the disclosed methods as appropriate.

[0560] Provided herein is a method comprising contacting DNA contacting DNA with a mutant TET2 enzyme (e.g. comprising a V1900A, V1900C, V1900G, V1900I, V1900P, or T1372S mutation) to oxidize 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) present in the DNA to 5-carboxycytosine (5caC), subsequently contacting at least a portion of the DNA with a substituted borane reducing agent, thereby converting 5-caC in the DNA to dihydrouracil (DHU), thereby producing treated DNA, and sequencing at least a portion of the treated DNA.

[0561] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase. In some such embodiments, the first nucleobase is hmC. DNA originally comprising the first nucleobase may be separated from other DNA using a labeling procedure comprising biotinylating positions that originally comprised the first nucleobase. In some embodiments, the first nucleobase is first derivatized with an azide-containing moiety, such as a glucosyl-azide containing moiety. The azide-containing moiety then may serve as a reagent for attaching biotin, e.g., through Huisgen cycloaddition chemistry. Then, the DNA originally comprising the first nucleobase, now biotinylated, can be separated from DNA not originally comprising the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), or streptavidin. An example of a procedure for separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase is hmC-seal, which labels hmC to form β-6-azide-glucosyl-5-hydroxymethylcytosine and then attaches a biotin moiety through Huisgen cycloaddition, followed by separation of the biotinylated DNA from other DNA using a biotin-binding agent. For an exemplary description of hmC-seal, see, e.g., Han et al., Mol. Cell 2016; 63: 711-719. This approach is useful for identifying fragments that include one or more hmC nucleobases.

[0562] In some embodiments, following such a separation, the method further comprises differentially tagging each of the DNA originally comprising the first nucleobase, the DNA not originally comprising the first nucleobase. The method may further comprise pooling the DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase following differential tagging. The DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase may then be used in downstream analyses. For example, the pooled DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase may be sequenced in the same sequencing cell (such as after being subjected to further treatments, such as those described herein) while retaining the ability to resolve whether a given read came from a molecule of DNA originally comprising the first nucleobase or DNA not originally comprising the first nucleobase using the differential tags.

[0563] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).

[0564] Techniques comprising partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mC, mA, caC (which may be generated by oxidation of mC or hmC with Tet2, e.g., before enzymatic conversion of unmodified C to U, e.g., using a deaminase such as APOBEC3A), or dihydrouracil from other DNA. See, e.g., Kumar et al., Frontiers Genet.2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015; 37:1155-62. Antibodies for various modified nucleobases, such as mC, caC, and forms of thymine / uracil including dihydrouracil or halogenated forms such as 5-bromouracil, are commercially available. Various modified bases can also be detected based on alterations in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read in sequencing as a G. See, e.g., US Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, “Mutation, Repair, and Recombination.” G. Captured Set; Target Regions

[0565] In some embodiments, nucleic acids captured or enriched using a method described herein comprise captured DNA, such as one or more captured sets of DNA. In some embodiments, the captured DNA comprise target regions that are differentially methylated in different immune cell types. In some embodiments, the immune cell types comprise rare or closely related immune cell types, such as activated and naive lymphocytes or myeloid cells at different stages of differentiation.

[0566] In some embodiments, a captured epigenetic target region set captured from a sample or first subsample comprises hypermethylation target regions. In some embodiments, the hypermethylation target regions are differentially or exclusively hypermethylated in one cell type or in one immune cell type, or in one immune cell type within a cluster. In some embodiments, the hypermethylation target regions are hypermethylated to an extent that is distinguishably higher or exclusively present in one cell type or one immune cell type or one immune cell type within a cluster. Such hypermethylation target regions may be hypermethylated in other cell types but not to the extent observed in the one cell type. In some embodiments, the hypermethylation target regions show lower methylation in healthy cfDNA than in at least one other tissue type.

[0567] In some embodiments, a captured epigenetic target region set captured from a sample or second subsample comprises hypomethylation target regions. In some embodiments, the hypomethylation target regions are exclusively hypomethylated in one cell type or in one immune cell type or in one immune cell type within a cluster. In some embodiments, the hypomethylation target regions are hypomethylated to an extent that is exclusively present in one cell type or one immune cell type or in one immune cell type within a cluster.

[0568] Such hypomethylation target regions may be hypomethylated in other cell types but not to the extent observed in the one cell type. In some embodiments, the hypomethylation target regions show higher methylation in healthy cfDNA than in at least one other tissue type.

[0248] Without wishing to be bound by any particular theory, in an individual with cancer, proliferating or activated immune cells (and potentially also cancer cells) may shed more DNA into the bloodstream than immune cells in a healthy individual (and healthy cells of the same tissue type, respectively). As such, the distribution of cell type and / or tissue of origin of cfDNA may change upon carcinogenesis. For example, the distribution of immune cell type of origin may change in a subject having cancer, precancer, infection, transplant rejection, or other disease or disorder directly or indirectly affecting the immune system. The status of epigenetic target regionsof certain immune cell types likewise may change in a subject having such a disease relative to a healthy subject or relative to the same subject prior to having the disease or disorder. Thus, variations in hypermethylation and / or hypomethylation can be an indicator of disease. For example, an increase in the level of hypermethylation target regions and / or hypomethylation target regions in a subsample following a partitioning step can be an indicator of the presence (or recurrence, depending on the history of the subject) of cancer.

[0569] Exemplary hypermethylation target regions and hypomethylation target regions useful for distinguishing between various cell types, including but not limited to immune cell types, have been identified by analyzing DNA obtained from various cell types via whole genome bisulfite sequencing, as described, e.g., in Stunnenberg, H. G. et. al., “The International Human Epigenome Consortium: A Blueprint for Scientific Collaboration and Discovery,” Cell 167, 1145 (2016) (doi.org / 10.1186 / sl3059-020-02065-5). Whole-genome bisulfite sequencing data is available from the Blueprint consortium, available on the internet at dcc.blueprint-epigenome.eu.

[0570] In some embodiments, first and second captured target region sets comprise, respectively, DNA co...

Claims

CLAIMS What is Claimed Is:

1. A method of methylation analysis for detection of a tumor-related biological condition, the method comprising: producing a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation criteria associated with each of the clusters; applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

2. The method of claim 1, wherein calculating the background signal rate comprises: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

3. The method of claim 2, wherein determining the tumor signal rate comprises: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

4. The method of claim 1, wherein the predetermined methylation criteria comprises a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

5. The method of claim 4, wherein the predetermined methylation criteria comprises sequence representations that indicate a gain in methylation.

6. The method of claim 4, wherein the predetermined methylation criteria comprises sequence representations that indicate a loss in methylation.

7. The method of claim 4, wherein the number of CpG dinucleotides that satisfy the methylation state comprises three CpG dinucleotides.

8. The method of claim 4, wherein the number of CpG dinucleotides that satisfy the methylation state comprises four CpG dinucleotides.

9. The method of claim 4, wherein the number of CpG dinucleotides that satisfy the methylation state comprises five CpG dinucleotides.

10. The method of claim 1, wherein the background signal cutoff value comprises 1e-6to 0.

01.

11. The method of claim 1, wherein the tumor signal cutoff value comprises 0.05 to 1.

00.

12. The method of claim 1, wherein each of the subset of the CpG dinucleotide clusters comprises no more than a predetermined number of CpG dinucleotides.

13. The method of claim 12, wherein the predetermined number of CpG dinucleotides is six.

14. The method of claim 13, wherein the predetermined number of CpG dinucleotides is four.

15. The method of claim 14, wherein the predetermined number of CpG dinucleotides is three.

16. The method of claim 1, wherein producing the framework further comprises ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

17. The method of claim 1, wherein producing the framework further comprises selecting non- overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

18. The method of claim 1, wherein producing the framework further comprises classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

19. The method of claim 1, wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

20. The method of claim 1, wherein applying the framework comprises aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor- related biological condition.

21. The method of claim 1, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

22. The method of claim 1, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer-associated fibroblasts present in the test subject.

23. The method of claim 1, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

24. The method of claim 1, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

25. The method of claim 1, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

26. The method of claim 1, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

27. The method of claim 1, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

28. The method of claim 1, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

29. The method of claim 1, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

30. The method of claim 1, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

31. The method of claim 1, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters comprises using a machine learning model.

32. The method of claim 1, comprising determining that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

33. The method of claim 1, determining a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

34. A method of methylation analysis for detection of a tumor-related biological condition, the method comprising: producing a framework for analyzing a sample data set comprising sequencing data derived from a test subject for the tumor-related biological condition, producing the framework comprising: obtaining a background data set comprising sequencing data and methylation data indicating an amount of methylation for cytosine-guanine (CpG) nucleotides in the background data set, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set; obtaining a tumor data set comprising sequencing data and methylation data, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating a tumor signal rate based on the tumor data set; comparing the background signal rate and the tumor signal rate to identify a subset of CpG dinucleotide clusters that have a background signal rate below a background signal cutoff value and a tumor signal rate above a tumor signal cutoff value; and producing the framework comprising the subset of the CpG dinucleotide clusters and a methylation threshold associated with each of the clusters.

35. The method of claim 34, comprising: applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters; and providing an indication of a tumor-related biological condition present in the test subject based on application of the framework.

36. The method of claim 34, wherein calculating the background signal rate comprises:determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

37. The method of claim 36, wherein determining the tumor signal rate comprises: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

38. The method of claim 34, wherein the predetermined methylation criteria comprises a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

39. The method of claim 38, wherein the predetermined methylation criteria comprises sequence representations that indicate a gain in methylation.

40. The method of claim 38, wherein the predetermined methylation criteria comprises sequence representations that indicate a loss in methylation.

41. The method of claim 38, wherein the number of CpG dinucleotides that satisfy the methylation state comprises three CpG dinucleotides.

42. The method of claim 38, wherein the number of CpG dinucleotides that satisfy the methylation state comprises four CpG dinucleotides.

43. The method of claim 38, wherein the number of CpG dinucleotides that satisfy the methylation state comprises five CpG dinucleotides.

44. The method of claim 34, wherein the background signal cutoff value comprises 1e-6to 0.

01.

45. The method of claim 34, wherein the tumor signal cutoff value comprises 0.05 to 1.

00.

46. The method of claim 34, wherein each of the subset of the CpG dinucleotide clusters comprises no more than a predetermined number of CpG dinucleotides.

47. The method of claim 46, wherein the predetermined number of CpG dinucleotides is six.

48. The method of claim 47, wherein the predetermined number of CpG dinucleotides is four.

49. The method of claim 48, wherein the predetermined number of CpG dinucleotides is three.

50. The method of claim 34, wherein producing the framework further comprises ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

51. The method of claim 34, wherein producing the framework further comprises selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

52. The method of claim 34, wherein producing the framework further comprises classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

53. The method of claim 34, wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

54. The method of claim 34, wherein applying the framework comprises aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

55. The method of claim 34, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

56. The method of claim 34, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer-associated fibroblasts present in the test subject.

57. The method of claim 34, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

58. The method of claim 34, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

59. The method of claim 34, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

60. The method of claim 34, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

61. The method of claim 34, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

62. The method of claim 34, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

63. The method of claim 34, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

64. The method of claim 34, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

65. The method of claim 34, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters comprises using a machine learning model.

66. The method of claim 34, comprising determining that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

67. The method of claim 34, determining a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

68. A method of methylation analysis for detection of a tumor-related biological condition, the method comprising: applying a framework to a sample data set to analyze amounts of methylation of cytosine-guanine (CpG) nucleotides in the sample data set at each of a subset of CpG clusters wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above a threshold; wherein applying the framework comprises aggregating sequence representations including in sequencing data derived that satisfies one or more methylation criteria included in the framework with respect to each of the subset of the CpG dinucleotide clusters to determine an indication of the tumor-related biological condition, wherein the sequencing data is derived from a sample obtained from a test subject; and providing the indication of the tumor-related biological condition present in the test subject based on application of the framework.

69. The method of claim 68, wherein the framework is produced by: analyzing a sample data set comprising the sequencing data derived from the test subject for the tumor-related biological condition: obtaining a background data set indicating a measure of methylated cytosine-guanine (CpG) nucleotides in the background data set based on a predetermined methylation criteria, the background data set originating from one or more first subjects in which a tumor is not detected; calculating a background signal rate using the background data set;obtaining a tumor data set based on the predetermined methylation criteria, the tumor data set originating from one or more second subjects in which a tumor is detected; calculating tumor signal rate based on the tumor data set; identifying a subset of CpG dinucleotide clusters that each have a background signal rate below a background signal cutoff value and a tissue signal rate above a tumor signal cutoff value; and wherein the framework comprises the subset of the CpG dinucleotide clusters and methylation criteria associated with each of the clusters.

70. The method of claim 68, wherein calculating the background signal rate comprises: determining, based on the background data set, a first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions; and determining, based on a second quantitative measurement of methylated CpG dinucleotides, a cutoff amount of methylated CpG dinucleotides at each of the plurality of genomic regions.

71. The method of claim 70, wherein determining the tumor signal rate comprises: determining a quantitative measurement of sequence representations based on the first quantitative measurement of CpG dinucleotides located in a plurality of genomic regions and the second quantitative measurement of methylated CpG dinucleotides.

72. The method of claim 68, wherein the predetermined methylation criteria comprises a methylation state and a number of CpG dinucleotides that satisfy the methylation state.

73. The method of claim 72, wherein the predetermined methylation criteria comprises sequence representations that indicate a gain in methylation.

74. The method of claim 72, wherein the predetermined methylation criteria comprises sequence representations that indicate a loss in methylation.

75. The method of claim 72, wherein the number of CpG dinucleotides that satisfy the methylation state comprises three CpG dinucleotides.

76. The method of claim 72, wherein the number of CpG dinucleotides that satisfy the methylation state comprises four CpG dinucleotides.

77. The method of claim 72, wherein the number of CpG dinucleotides that satisfy the methylation state comprises five CpG dinucleotides.

78. The method of claim 68, wherein the background signal cutoff value comprises 1e-6to 0.

01.

79. The method of claim 68, wherein the tumor signal cutoff value comprises 0.05 to 1.00.

80. The method of claim 68, wherein each of the subset of the CpG dinucleotide clusters comprises no more than a predetermined number of CpG dinucleotides.

81. The method of claim 80, wherein the predetermined number of CpG dinucleotides is six.

82. The method of claim 81, wherein the predetermined number of CpG dinucleotides is four.

83. The method of claim 82, wherein the predetermined number of CpG dinucleotides is three.

84. The method of claim 68, wherein producing the framework further comprises ranking each of the subset of the CpG dinucleotide clusters based on the background signal rate.

85. The method of claim 68, wherein producing the framework further comprises selecting non-overlapping CpG dinucleotide clusters for the subset of the CpG dinucleotide clusters.

86. The method of claim 68, wherein producing the framework further comprises classifying CpG dinucleotide clusters to generate the subset of the CpG dinucleotide clusters.

87. The method of claim 68, wherein applying the framework comprises determining if methylation of CpG dinucleotides at each of the subset of the CpG dinucleotide clusters is above the threshold.

88. The method of claim 68, wherein applying the framework comprises aggregating sequence representations included in the sequencing data that satisfy the methylation criteria with respect to each of the subset of the CpG dinucleotide clusters to provide the indication of the tumor-related biological condition.

89. The method of claim 68, wherein the indication of the tumor-related biological condition corresponds to an amount of activated T-cells present in the test subject.

90. The method of claim 68, wherein the indication of the tumor-related biological condition corresponds to an amount of cancer-associated fibroblasts present in the test subject.

91. The method of claim 68, wherein the indication of the tumor-related biological condition is determined using one or more machine learning techniques.

92. The method of claim 68, wherein the indication of the tumor-related biological condition includes an indication of the type of cancer.

93. The method of claim 68, wherein the indication of the tumor-related biological condition includes an indication of the type of a tumor fraction.

94. The method of claim 68, wherein the indication of the tumor-related biological condition includes an indication of a tissue of origin.

95. The method of claim 68, wherein the indication of the tumor-related biological condition includes an indication of tumor burden.

96. The method of claim 68, wherein the indication of the tumor-related biological condition includes an indication of tumor recurrence.

97. The method of claim 68, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from cell-free deoxyribonucleic acid (DNA) molecules.

98. The method of claim 68, wherein the sample data set, the background data set, the tumor data set, or combinations thereof are derived from at least one of plasma samples or tissue samples.

99. The method of claim 68, wherein applying the framework to the sample data set to analyze amounts of methylation of CpG nucleotides in the sample data set at each of the subset of CpG clusters comprises using a machine learning model.

100. The method of claim 68, comprising determining that a number of CpG clusters are mapped to a genomic region that corresponds to a classification region included in a diagnostic assay.

101. The method of claim 68, determining a plurality of overlapping CpG clusters that correspond to a genomic subregion within a classification region.

Citation Information

Patent Citations

  • Compositions and methods for analyzing modified nucleotides

    US10260088B2

  • Hyperactive AID / APOBEC and hmC dominant TET enzymes

    US10961525B2

  • Oligonucleotides

    US20010053519A1

  • Method and apparatus for imaging a sample on a device

    US20030152490A1

  • Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels

    US20110160078A1