DETECTING THE PRESENCE OF TUMOR BASED ON THE METHYLATION STATUS OF CELL-FREE NUCLEIC ACID MOLECULES - Patent application

JP2025513786A5Pending Publication Date: 2026-04-07GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Current cancer detection methods using liquid biopsies face challenges due to the low amount and heterogeneous forms of nucleic acids in body fluids, making it difficult to accurately classify samples as containing tumor-derived DNA.

Method used

A computer-implemented system and method that analyzes methylation status of cell-free nucleic acid molecules using machine learning models to determine cancer metrics, specifically focusing on regions with a threshold number of CpGs to improve sensitivity and classification accuracy.

Benefits of technology

The system achieves enhanced sensitivity and accuracy in detecting cancer by effectively classifying samples based on methylation patterns, even with low amounts of nucleic acids, thereby improving early detection and diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In one or more embodiments, the method includes obtaining, by a computing system having one or more hardware processors and memory, training sequence data comprising training sequencing reads derived from a plurality of samples of a plurality of subjects, wherein each training sequencing read comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosines contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 328,602, filed April 7, 2022, and U.S. Provisional Patent Application No. 63 / 336,852, filed April 29, 2022, both of which are incorporated by reference in their entirety for all purposes. [Background technology]

[0002] background Cancer is a leading cause of disease worldwide. Each year, tens of millions of people worldwide are diagnosed with cancer, and more than half ultimately die from it. In many countries, cancer is the second leading cause of death after cardiovascular disease. Early detection is often associated with improved cancer outcomes.

[0003] Cancer can be caused by the accumulation of genetic variations in an individual's normal cells, resulting in improperly regulated cell division in at least some of them. Such variations generally include copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), epigenetic variations including 5-methylation of cytosine (5-methylcytosine), and association of DNA with chromatin and transcription factors.

[0004] Cancer is often detected by tumor biopsy followed by analysis of cells, markers, or DNA extracted from the cells. However, it has recently been proposed that cancer can also be detected from cell-free nucleic acids in bodily fluids such as blood or urine. Such tests have the advantage of being non-invasive and can be performed without the need to identify suspected cancer cells through a biopsy. However, such tests are complicated by the fact that the amount of nucleic acid in bodily fluids is very low and that nucleic acids exist in heterogeneous forms (e.g., RNA and DNA, single-stranded and double-stranded, and various states of post-replicative modification and association with proteins such as histones).

[0005] Thus, there is a need for improved systems and methods for improved cancer detection using liquid biopsy assays. It is therefore an object of the present disclosure to provide computer-implemented systems and methods with increased sensitivity that have an improved ability to classify samples as containing tumor-derived DNA.

[0006] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain implementations and, together with the written description, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and not by way of limitation. It will be understood that like reference numerals identify like components throughout the drawings, unless the context indicates otherwise. It will also be understood that some or all of the figures may be schematic for purposes of illustration and do not necessarily indicate the actual relative size or location of the depicted elements. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagrammatic representation of an example environment 100 for identifying nucleic acids corresponding to classification regions of a reference sequence, where the classification regions have at least a threshold number of CpGs.

[0008] [Figure 2] FIG. 2 is a diagrammatic representation of an example architecture for determining tumor metrics based on one or more models that analyze the methylation status of cell-free nucleic acid molecules, according to one or more implementations.

[0009] [Figure 3]FIG. 3 is a diagrammatic representation of an example architecture for training one or more machine learning models to determine cancer metrics based on the methylation status of cell-free nucleic acid molecules, according to one or more implementations.

[0010] [Figure 4] FIG. 4 is a flow diagram of an example process for determining tumor metrics associated with the level of methylation of classification regions of a reference sequence, according to one or more implementations.

[0011] [Figure 5] FIG. 5 is a block diagram illustrating components of a machine in the form of a computer system that can read and execute instructions from one or more machine-readable media to perform any one or more methodologies described herein, according to one or more example implementations.

[0012] [Figure 6] FIG. 6 is a block diagram illustrating a representative software architecture that may be used in conjunction with one or more hardware architectures described herein, according to one or more example implementations.

[0013] [Figure 7] 7A, 7B, and 7C are graphical representations showing promoter-methylation calls in training and test samples among 88 TSG+HRD genes.

[0014] [Figure 8] Figure 8A is a graphical representation showing cancer prediction scores in cancer-free samples with a call of >1. Figure 8B is a table showing the genes most frequently called for promoter methylation in cancer-free donors. Figure 8C is a table showing in-silico LoD estimates for selected genes from cell line KM12.

[0015] [Figure 9] Figure 9 is a graphical representation showing the prevalence of promoter methylation of TSG and HRD genes in different cancer types in our test dataset (total N = 559) and the TCGA public data (total N = 2,380). In TCGA, we limited the data to genes with promoter methylation.

[0016] [Figure 10] Figure 10A is a graphical representation showing MLH1 promoter methylation (MSI-H and MSS) in cancer-free donors and CRC patients. Figure 10B is a table showing MLH1 promoter methylation and BRAF-V600E calls in CRC patients.

[0017] [Figure 11] FIG. 11 is a table outlining the training and test datasets for Example 5.

[0018] [Figure 12] Figure 12A is a graphical representation of a graph showing model performance for predicting CRC / cancer-free status in the training set. Shading indicates variation in replicates. Figure 12B is a table showing the performance of the cancer prediction model on an independent test dataset.

[0019] [Figure 13] Figure 13A is a table showing the CV of TF estimates from genome calls and methylation in an in vitro dataset. Figure 13B is a graphical representation showing TF model performance (diagonal black line) in a training set of CRC and cancer-free samples (cross-validation). Figure 13C is a graphical representation showing an in-silico dataset for lower true TFs.

[0020] [Figure 14]FIG. 14 is a graphical representation showing the distribution of predictive TFs for CRC patients with and without driver mutations in the training set.

[0021] [Figure 15] FIG. 15A is a graphical representation showing individual positivity rates for lung cancer detection in stage I / II and stage III / IV patients.

[0022] FIG. 15B is a graphical representation showing individual positivity rates for multiple cancer detection (bladder, stomach, ovarian, pancreatic, and liver cancer) in stage I / II and stage III / IV patients.

[0023] [Figure 16] FIG. 16 is a graphical representation showing individual positivity rates for multiple cancer detection (bladder, stomach, ovarian, pancreatic, and liver cancer) in stage I, stage II, stage III, and stage IV patients.

[0024] [Figure 17] FIG. 17 is a graphical representation of the association of epigenomic MAFs with target MAFs for colorectal, lung, and breast cancer.

[0025] [Figure 18] FIG. 18 is a table showing that the quantitative accuracy of epigenomic cTFs can reach an LoQ of less than 0.1% in clinical samples of CRC, lung, and breast.

[0026] [Figure 19] FIG. 19A is a graphical representation showing that somatic mutation-based cTF is robust to repeated experiments within the same cTF level, especially at cTF levels of 0.5% or higher.

[0027] FIG. 19B is a graphical representation showing that epigenomic cTFs can maintain 100% validation rate and have LoQs down to 0.1% of cTFs.

[0028] [Figure 20] FIG. 20A is a graphical representation of methylation signals and somatic mutations for the first replicate of clinical dose-finding.

[0029] FIG. 20B is a graphical representation of methylation signals and somatic mutations for the second replicate of clinical titration.

[0030] [Figure 21] FIG. 21 is a table showing the change in ctDNA levels for the first and second replicates, calculated using the genomic-only and methylation methods.

[0031] [Figure 22] FIG. 22 is a graphical representation of epigenomic vs. genomic cTFs for clinical samples (one point per sample).

[0032] [Figure 23] FIG. 23 is a graphical representation of epiMAF distribution in early and late stage cancer patients for breast, colorectal, lung, and other cancer groups.

[0033] [Figure 24] FIG. 24 is a graphical representation showing probability distributions indicating the number of methylated cytosines in the three distributions.

[0034] [Figure 25] FIG. 25A includes a graphical representation showing the change in metrics for a first classification region for a first group of samples treated with the MBD using a first set of reagents and a second group of samples treated with the MBD using a second set of reagents.

[0035] FIG. 25B includes a graphical representation showing the change in metrics for the second classification region for the first classification region for a first group of samples treated with the MBD using the first set of reagents and a second group of samples treated with the MBD using the second set of reagents. DETAILED DESCRIPTION OF THE INVENTION

[0036] overview In one or more embodiments, the method includes obtaining, by a computing system having one or more hardware processors and memory, training sequence data including training sequencing reads from a plurality of samples from a plurality of subjects, wherein each training sequencing read comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The method also includes analyzing, by the computing system, the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosine in the subject in which cancer is detected and having at least the threshold cytosine-guanine content. The method also includes analyzing the training sequencing reads by a computing system to determine second quantitative measures derived from the training sequencing reads corresponding to a plurality of control regions, each control region of the plurality of control regions having at least a threshold cytosine-guanine content and corresponding to additional genomic regions of the reference genome having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The method also includes determining, by the computing system, metrics for each classification region of the plurality of classification regions based on the first quantitative measures for each classification region and the second quantitative measures for the plurality of control regions. The method also includes generating, by the computing device, training data including metrics for each classification region of the plurality of classification regions for the training sequence reads from the sample of the training subject.The method also includes implementing, by a computing system, one or more machine learning algorithms using the training data to generate a model for determining an indication of the presence of cancer in a subject based on the amount of methylated cytosines in at least a portion of a plurality of classification regions, wherein the model includes weights for each classification region of the plurality of classification regions, and at least a portion of the weights for the individual classification regions are different from one another.

[0037] In one or more embodiments, the method includes obtaining, by a computing system, test sequence data from an additional subject not included in the plurality of subjects, wherein the test sequence data includes test sequencing reads from a sample of the additional subject, wherein each test sequencing read includes a nucleotide sequence corresponding to a fragment of nucleic acid included in the additional sample, and wherein each test sequencing read corresponds to a molecule having at least a threshold amount of methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content; and determining an indication of the presence of cancer in the additional subject using the model and the additional sequence data.

[0038] In one or more embodiments, the method includes analyzing, by a computing system, the test sequencing reads to determine a first quantitative measure derived from the test sequencing reads corresponding to each individual classification region of the plurality of classification regions; analyzing, by the computing system, the test sequencing reads to determine a second quantitative measure derived from the test sequencing reads corresponding to each individual control region of the plurality of control regions; determining, by the computing system, metrics for each individual classification region based on the first quantitative measure for each individual classification region and the second quantitative measure for the plurality of control regions; and generating, by the computing system, an input vector including the metrics for each individual classification region, wherein the input vector is used in a model to determine an indication of the presence of cancer in the additional subject.

[0039] In one or more embodiments, the one or more machine learning algorithms include one or more classification algorithms, and the indication that cancer is present corresponds to a probability that cancer is present in the additional subject.

[0040] In one or more embodiments, the one or more machine learning algorithms comprise one or more regression algorithms, and the indicator corresponds to an estimate of the tumor fraction of the additional sample.

[0041] In one or more embodiments, the training sequencing reads comprise a first portion of the training sequence data and the additional training sequencing reads comprise a second portion of the training sequence data, where the additional training sequencing reads are different from the training sequencing reads, and the method includes analyzing, by a computing system, at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine individual frequencies of a plurality of variants present in each sample of the plurality of samples; determining, by the computing system, for the each sample, one variant of the plurality of variants having a highest frequency that corresponds to an individual frequency having a maximum value among the individual frequencies from the individual samples; and determining, by the computing system, an individual measure of tumor fraction for the each sample based on the maximum of the individual frequencies from the individual samples.

[0042] In one or more embodiments, the training data includes individual measures of tumor fraction for each sample of the plurality of samples, and the model is generated based on the individual measures of tumor fraction for each sample of the plurality of samples.

[0043] In one or more embodiments, the metrics for each classification region are determined based on a scaling factor and an error correction factor.

[0044] In one or more embodiments, the plurality of classification regions individually correspond to genomic regions in which the methylation rate of a genomic region in nucleic acid derived from a cell obtained from a subject having cancer differs from the methylation rate of a genomic region in nucleic acid derived from a cell obtained from a subject not having cancer.

[0045] In one or more embodiments, the plurality of classification regions corresponds to a first plurality of classification regions for a first cancer type, and a model for a second cancer type can be generated based on a second plurality of classification regions that are different from the first plurality of classification regions.

[0046] In one or more embodiments, the plurality of samples and the additional sample comprise cell-free nucleic acid.

[0047] In one or more embodiments, the method includes performing, by a computing system, a training process using the training data to generate a model, the training process including determining, by the computing system, one or more additional weights for individual samples included in the training data based on a cancer indicator for the individual sample falling within a threshold confidence level.

[0048] In one or more embodiments, the cancer indicator for the individual sample is outside a threshold confidence level, and the method includes applying, by the computing system, a penalty to the weight of the individual sample during the training process.

[0049] In one or more aspects, the method includes: performing, by a computing system, one or more first iterations of a training process for a model using one or more machine learning algorithms using a portion of the training data; and generating, by the computing system, first output data for the model based on the one or more first iterations of the training process, the first output data corresponding to one or more first additional indicators of the presence of cancer in a first individual subject of the plurality of subjects, the first individual subject corresponding to the portion of the training data.

[0050] In one or more embodiments, the method includes combining, by a computing system, the first output data and the training data to produce additional training data; performing, by the computing system, one or more second iterations of the training process for the model using a portion of the additional training data; and generating, by the computing system, second output data for the model based on the one or more second iterations of the training process, the second output data indicating one or more second additional indicators of the presence of cancer in a second individual subject of the plurality of subjects, the second individual subject corresponding to the portion of the additional training data.

[0051] In one or more embodiments, a weight for each classification region of the plurality of classification regions is determined based on the first output data and the second output data.

[0052] In one or more embodiments, the method includes determining, by a computing system, that some of the indicators of the presence of cancer determined during one or more iterations of the training process are thresholds for at least one or more samples included in the training data, and determining, by the computing system, that a modification to one or more weights of the model is to be no modification or a minimal modification.

[0053] In one or more embodiments, the method includes determining, by a computing system, that the number of additional indicators of the presence of cancer determined during one or more iterations of the training process is less than a threshold for one or more additional samples included in the training data, and determining, by the computing system, that a modification to one or more additional weights of the model is modified by more than a minimum amount.

[0054] In one or more embodiments, the method includes combining a plurality of nucleic acids from at least one of a subject's blood or tissue with a solution containing an amount of a methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution, and performing multiple washes of the nucleic acid-MBD protein solution with a salt solution to produce several nucleic acid fractions, each nucleic acid fraction having a threshold number of methylated cytosines within a region of the plurality of nucleic acids having at least a threshold cytosine-guanine content.

[0055] In one or more embodiments, some of the washes are performed with a solution having a concentration of sodium chloride (NaCl) to produce one of several nucleic acid fractions having a range of binding strengths to the MBD protein.

[0056] In one or more embodiments, the method comprises determining that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding strengths to MBD proteins; attaching a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode being included in a first set of molecular barcodes associated with the first partition; determining that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding strengths to MBD proteins that is different from the first range of binding strengths to MBD proteins; and attaching a second molecular barcode to nucleic acids of the second nucleic acid fraction, the second molecular barcode being included in a second set of molecular barcodes associated with the second partition.

[0057] In one or more embodiments, the method includes combining at least a portion of some nucleic acid fractions with a quantity of one or more methylation-sensitive restriction enzymes that cleave molecules with one or more unmethylated cytosines to produce at least a portion of a plurality of samples used to generate sequencing reads.

[0058] In one or more embodiments, the method includes combining at least a portion of some nucleic acid fractions with a quantity of one or more methylation-dependent restriction enzymes that cleave molecules having one or more methylated cytosines to produce at least a portion of a plurality of samples used to generate sequencing reads.

[0059] In one or more embodiments, the detection limit of the model for determining the tumor fraction of a sample is 0.05% or less.

[0060] In one or more embodiments, a computing system includes one or more hardware processors and one or more non-transitory computer-readable storage media containing computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including obtaining training sequence data including training sequencing reads from a plurality of samples of a plurality of subjects, wherein each training sequencing read includes a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosines in the subject in which cancer is detected and having at least the threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine second quantitative measures derived from the training sequencing reads corresponding to a plurality of control regions, each of the plurality of control regions corresponding to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in an additional subject in which cancer is not detected. The operations also include determining metrics for each of the plurality of classification regions based on the first quantitative measures for each of the classification regions and the second quantitative measures for the plurality of control regions. The operations also include generating training data including metrics for each of the plurality of classification regions for the training sequence reads from a sample of the training subject.The operations also include using the training data to implement one or more machine learning algorithms to generate a model for determining an indication of the presence of cancer in a subject based on the amount of methylated cytosines in at least a portion of the plurality of classification regions, wherein the model includes weights for each classification region of the plurality of classification regions, and at least a portion of the weights for the individual classification regions are different from one another.

[0061] In one or more embodiments, the computing system can obtain test sequence data from an additional subject not included in the plurality of subjects. The test sequence data can include test sequencing reads from a sample from the additional subject, each test sequencing read comprising a nucleotide sequence corresponding to a fragment of nucleic acid included in the additional sample, each test sequencing read corresponding to a molecule having at least a threshold amount of methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations can also include using the model and the additional sequence data to determine an indication that cancer is present in the additional subject.

[0062] In one or more embodiments, the computing system can analyze the test sequencing reads to determine a first quantitative measure derived from the test sequencing reads corresponding to each classification region of the plurality of classification regions. The computing system can also analyze the test sequencing reads to determine a second quantitative measure derived from the test sequencing reads corresponding to each control region of the plurality of control regions. The computing system can also determine metrics for each classification region based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. The computing system can also generate an input vector including the metrics for each classification region, where the input vector is used in the model to determine an indication that cancer is present in the additional subject.

[0063] In one or more embodiments of the computing system, the one or more machine learning algorithms include one or more classification algorithms, and the indication that cancer is present corresponds to a probability that cancer is present in the additional subject.

[0064] In one or more embodiments of the computing system, the one or more machine learning algorithms include one or more regression algorithms, and the index corresponds to an estimate of the tumor fraction of the additional sample.

[0065] In one or more embodiments of the computing system, the training sequencing reads include a first portion of the training sequence data, and the additional training sequencing reads include a second portion of the training sequence data, where the additional training sequencing reads are different from the training sequencing reads, and the computing system can analyze at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine individual frequencies of a plurality of variants present in each sample of the plurality of samples. The computing system can also determine, for each sample, one variant of the plurality of variants having a highest frequency corresponding to the individual frequency having the maximum value among the individual frequencies from the individual samples. The computing system can also determine an individual measure of tumor fraction for each sample based on the maximum individual frequency from the individual samples.

[0066] In one or more embodiments of the computing system, the training data includes individual measures of tumor fraction for each of the plurality of samples, and the model is generated based on the individual measures of tumor fraction for each of the plurality of samples.

[0067] In one or more aspects of the computing system, metrics for each classification region are determined based on a scaling factor and an error correction factor.

[0068] In one or more embodiments of the computing system, the plurality of classification regions individually correspond to genomic regions in which the methylation rate of a genomic region in nucleic acid derived from a cell obtained from a subject having cancer differs from the methylation rate of a genomic region in nucleic acid derived from a cell obtained from a subject not having cancer.

[0069] In one or more embodiments of the computing system, the plurality of classification regions corresponds to a first plurality of classification regions for a first cancer type, and a model for a second cancer type can be generated based on a second plurality of classification regions that are different from the first plurality of classification regions.

[0070] In one or more embodiments, the computing system can perform a training process using the training data to generate the model, where the training process includes determining one or more additional weights for individual samples included in the training data based on a cancer indicator for the individual sample that is within a threshold confidence level.

[0071] In one or more aspects of the computing system, the cancer indicator for the individual sample is outside a threshold confidence level, and the method includes applying, by the computing system, a penalty to the weight of the individual sample during the training process.

[0072] In one or more embodiments, the computing system may use one or more machine learning algorithms to perform one or more first iterations of a training process on the model using a portion of the training data. The computing system may also generate first output data for the model based on the one or more first iterations of the training process, the first output data corresponding to one or more first additional indicators of the presence of cancer in a first individual subject of the plurality of subjects, the first individual subject corresponding to the portion of the training data.

[0073] In one or more embodiments, the computing system can combine the first output data and the training data to produce additional training data, perform one or more second iterations of the training process on the model using a portion of the additional training data, and generate second output data for the model based on the one or more second iterations of the training process, the second output data indicating one or more second additional indicators of the presence of cancer in a second individual subject of the plurality of subjects, the second individual subject corresponding to the portion of the additional training data.

[0074] In one or more aspects of the computing system, weights for individual classification regions of the plurality of classification regions are determined based on the first output data and the second output data.

[0075] In one or more embodiments, the computing system can determine that some of the cancer presence indicators determined during one or more iterations of the training process are thresholds for at least one or more samples included in the training data. The computing system can also determine that one or more weights of the model are not modified or are modified only a minimal amount.

[0076] In one or more embodiments, the computing system can determine that the number of additional indicators of the presence of cancer determined during one or more iterations of the training process is less than a threshold for one or more additional samples included in the training data, and determine that a correction to one or more additional weights of the model is corrected by more than a minimum amount.

[0077] In one or more embodiments of the computing system, the limit of detection of the model for determining the tumor fraction of the sample is 0.05% or less.

[0078] In one or more embodiments, one or more computer-readable storage media include computer-readable instructions that, when executed by one or more processors of the computing system, cause the computing system to perform operations including obtaining training sequence data including training sequencing reads from a plurality of samples of a plurality of subjects, wherein each training sequencing read includes a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosines in the subject in which cancer is detected and having at least the threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine second quantitative measures derived from the training sequencing reads corresponding to a plurality of control regions, each of the plurality of control regions corresponding to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in an additional subject in which cancer is not detected. The operations also include determining metrics for each of the plurality of classification regions based on the first quantitative measures for each of the classification regions and the second quantitative measures for the plurality of control regions. The operations also include generating training data including metrics for each of the plurality of classification regions for the training sequence reads from a sample of the training subject.The operations also include using the training data to implement one or more machine learning algorithms to generate a model for determining an indication of the presence of cancer in a subject based on the amount of methylated cytosines in at least a portion of the plurality of classification regions, wherein the model includes weights for each classification region of the plurality of classification regions, and at least a portion of the weights for the individual classification regions are different from one another.

[0079] In one or more embodiments, the method includes obtaining, by a computing system having one or more hardware processors and memory, sequencing reads from a sample obtained from the subject, wherein each sequencing read comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in the sample and corresponds to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The method may also include determining, by the computing system, a first quantitative measure derived from the sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of each classification region of the plurality of classification regions corresponds to a genomic region of the reference genome having the threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in which cancer is detected. The method may also include analyzing, by the computing system, the sequencing reads to determine a second quantitative measure derived from the sequencing reads corresponding to a plurality of control regions, wherein each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least the threshold cytosine-guanine content and having at least the threshold amount of methylated cytosine in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The method may also include determining, by the computing system, a plurality of metrics, each of the plurality of metrics corresponding to a respective one of the plurality of classification regions, based on a first quantitative measure for the respective classification region and a second quantitative measure for the plurality of control regions. The method may also include determining, by the computing system, an indication that cancer is present in the subject based at least in part on the plurality of metrics.

[0080] In one or more embodiments, the method may also include determining, by a computing system, a distribution of sequence representations for the differentially methylated region using the sequencing data, and determining, by the computing system, that at least a threshold amount of sequence representations included in the distribution overlap with a subregion of the differentially methylated region. The method may also include determining, by the computing system, that the subregion of the differentially methylated region is one classification region of a plurality of classification regions.

[0081] In one or more embodiments, the method may also include determining, by a computing system, an order of values ​​of the plurality of metrics; and determining, by the computing system, a subset of classification regions from among the plurality of classification regions based on the order, wherein a portion of the plurality of metrics corresponding to the subset of classification regions is used to determine an indication that cancer is present in the subject.

[0082] In one or more embodiments, the indication that cancer is present in the subject is an initial indication that cancer is present in the subject, and the method may also include applying, by a computing system, a scaling factor to the initial indication that cancer is present in the subject to determine a modified indication that cancer is present in the subject.

[0083] In one or more embodiments, the indication that cancer is present in a subject corresponds to the tumor fraction.

[0084] In one or more embodiments, the sample is a first sample collected at least one of before and at the start of cancer treatment, and the method may also include obtaining, by a computing system having one or more hardware processors and memory, additional sequencing reads from a second sample obtained from the subject, wherein each additional sequencing read comprises an additional nucleotide sequence corresponding to a fragment of a nucleic acid contained in the second sample and corresponds to an additional molecule having a threshold amount of methylated cytosine contained within a region of the additional nucleotide sequence having at least a threshold cytosine-guanine content. The method may also include determining, by the computing system, additional first quantitative measures derived from the additional sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having a threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in whom cancer is detected. The method may also include analyzing, by a computing system, additional sequencing reads to determine additional second quantitative measures derived from the additional sequencing reads corresponding to a plurality of control regions, where each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in the additional subject in which cancer is not detected. The method may also include determining, by the computing system, a plurality of additional metrics, each of which corresponds to an individual classification region of the plurality of classification regions, based on the additional first quantitative measure for each classification region and the additional second quantitative measure for the plurality of control regions. The method may also include determining, by the computing system, an additional indication that the subject has cancer based at least in part on the plurality of additional metrics.

[0085] In one or more embodiments, the method includes obtaining test sequence data from a subject by a computing system having one or more hardware processors and memory, the test sequence data including test sequencing reads derived from a sample from the subject, each test sequencing read including a nucleotide sequence corresponding to a fragment of a nucleic acid contained in an additional sample, each test sequencing read corresponding to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The method may also include analyzing the test sequencing reads by the computing system to determine a first quantitative measure derived from the test sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in which cancer is detected. The method may also include a step of analyzing the test sequencing reads by a computing system to determine a second quantitative measure derived from the test sequencing reads corresponding to each control region of the plurality of control regions, wherein each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in an additional subject in which cancer is not detected. The method may also include a step of determining, by the computing system, metrics for each classification region based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. The method may also include a step of generating, by the computing system, an input vector including the metrics for each classification region.The method also includes determining, by a computing system, an indication that cancer is present in the subject by providing an input vector to a model implementing one or more machine learning techniques to generate an indication that cancer is present in the subject, wherein the model includes weights for each classification region of the plurality of classification regions, and at least a portion of the weights for the individual classification regions differ from one another.

[0086] In one or more embodiments, the method may also include obtaining, by a computing system having one or more hardware processors and memory, training sequence data including training sequencing reads from a plurality of samples of a plurality of training subjects, wherein each training sequencing read comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosines contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The method may also include analyzing, by the computing system, the training sequencing reads to determine an additional first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions. The method may also include analyzing, by the computing system, the training sequencing reads to determine an additional second quantitative measure derived from the training sequencing reads corresponding to the plurality of control regions. The method may also include determining, by the computing system, an additional metric for each classification region of the plurality of classification regions based on the additional first quantitative measure for each classification region and the additional second quantitative measure for the plurality of control regions. The method may also include generating, by the computing device, training data including additional metrics for each of the plurality of classification regions for training sequence reads from samples of the plurality of training subjects. The method may also include implementing, by the computing system, one or more machine learning algorithms using the training data to generate a model for determining an indication of the presence of cancer in a subject based on an amount of methylated cytosine in at least a portion of the plurality of classification regions.

[0087] In one or more embodiments, the one or more machine learning algorithms comprise one or more classification algorithms, and the indication that cancer is present corresponds to a probability that cancer is present in the subject.

[0088] In one or more embodiments, the one or more machine learning algorithms comprise one or more regression algorithms, and the index corresponds to an estimate of the tumor fraction of the sample.

[0089] In one or more embodiments, the subject's sample and the plurality of samples of the plurality of training subjects comprise cell-free nucleic acid.

[0090] In one or more embodiments, the threshold amount of sequence representations is at least about 70% of the sequence representations included in the distribution.

[0091] In one or more embodiments, the second sample is obtained at least one week after the subject has been administered the cancer treatment.

[0092] In one or more embodiments, the method may also include analyzing, by a computing system, the indication that cancer is present in the subject in association with additional indications that cancer is present in the subject to determine a response to treatment for the subject.

[0093] In one or more embodiments, the training sequencing reads include a first portion of the training sequence data, and the additional training sequencing reads include a second portion of the training sequence data, where the additional training sequencing reads are different from the training sequencing reads, and the method may also include analyzing, by a computing system, at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine individual frequencies of a plurality of variants present in each sample of the plurality of samples, and determining, by the computing system, for each sample, one variant of the plurality of variants having a highest frequency that corresponds to an individual frequency having a maximum value among the individual frequencies from the individual samples. The method may also include determining, by the computing system, individual measures of tumor fraction for each sample based on the maximum of the individual frequencies from the individual samples.

[0094] In one or more embodiments, the training data includes individual measures of tumor fraction for each sample of the plurality of samples, and the model is generated based on the individual measures of tumor fraction for each sample of the plurality of samples.

[0095] In one or more embodiments, a computing system includes a processor and a memory storing instructions that, when executed by the processor, configure the computing system to obtain sequencing reads from a sample obtained from the subject, each sequencing read comprising a nucleotide sequence corresponding to a fragment of nucleic acid contained in the sample and corresponding to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The computing system can also determine a first quantitative measure derived from the sequencing reads corresponding to each classification region of the plurality of classification regions, at least a portion of each classification region of the plurality of classification regions corresponding to a genomic region of the reference genome having the threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in which cancer is detected. The computing system can also analyze the sequencing reads to determine a second quantitative measure derived from the sequencing reads corresponding to a plurality of control regions, each control region of the plurality of control regions corresponding to an additional genomic region of the reference genome having at least the threshold cytosine-guanine content and having at least the threshold amount of methylated cytosine in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The computing system can also determine a plurality of metrics, each of which corresponds to a respective one of the plurality of classification regions, based on a first quantitative measure for each of the classification regions and a second quantitative measure for the plurality of control regions. The computing system can also determine an indication that cancer is present in the subject based at least in part on the plurality of metrics.

[0096] In one or more embodiments of the computing system, the threshold amount of sequence representations is at least about 70% of the sequence representations included in the distribution.

[0097] In one or more embodiments, the computing system can use the sequencing data to determine a distribution of sequence representations for the differentially methylated region. The computing system can also determine that at least a threshold amount of sequence representations included in the distribution overlap with a subregion of the differentially methylated region. The computing system can also determine that the subregion of the differentially methylated region is one classification region among a plurality of classification regions.

[0098] In one or more embodiments, the computing system can determine an order of values ​​of the plurality of metrics and determine a subset of classification regions from among the plurality of classification regions based on the order, where a portion of the plurality of metrics corresponding to the subset of classification regions is used to determine an indication that cancer is present in the subject.

[0099] In one or more embodiments of the computing system, the indication that cancer is present in the subject is an initial indication that cancer is present in the subject, and the computing system can apply a scaling factor to the initial indication that cancer is present in the subject to determine a modified indication that cancer is present in the subject.

[0100] In one or more embodiments of the computing system, the indication that cancer is present in the subject corresponds to the tumor fraction.

[0101] In one or more embodiments of the computing system, the sample is a first sample collected at least one of before and at the start of cancer treatment, and the computing system can obtain additional sequencing reads from a second sample obtained from the subject, each additional sequencing read comprising an additional nucleotide sequence corresponding to a fragment of nucleic acid contained in the second sample and corresponding to an additional molecule having a threshold amount of methylated cytosine contained within a region of the additional nucleotide sequence having at least a threshold cytosine-guanine content. The computing system can also determine additional first quantitative measures derived from the additional sequencing reads corresponding to individual classification regions of the plurality of classification regions, at least a portion of each classification region of the plurality of classification regions corresponding to a genomic region of the reference genome having a threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in whom cancer is detected. The computing system can also analyze the additional sequencing reads to determine additional second quantitative measures derived from the additional sequencing reads corresponding to a plurality of control regions, where each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in the additional subject in which cancer is not detected. The computing system can also determine a plurality of additional metrics, where each additional metric of the plurality of additional metrics corresponds to an individual classification region of the plurality of classification regions, based on the additional first quantitative measure for each classification region and the additional second quantitative measure for the plurality of control regions. The computing system can also determine an additional indication that cancer is present in the subject based at least in part on the plurality of additional metrics.

[0102] In one or more embodiments of the computing system, the second sample is obtained at least one week after the subject is administered the cancer treatment.

[0103] In one or more embodiments, the computing system can analyze an indication of the presence of cancer in a subject in association with additional indications of the presence of cancer in the subject to determine a response to treatment for the subject.

[0104] In one or more embodiments, a computing system includes a processor and a memory storing instructions that, when executed by the processor, configure the system to obtain test sequence data from a subject, the test sequence data including test sequencing reads from the subject's sample, individual test sequencing reads including nucleotide sequences corresponding to fragments of nucleic acids contained in an additional sample, and individual test sequencing reads corresponding to molecules having a threshold amount of methylated cytosines contained within regions of the nucleotide sequences having at least a threshold cytosine-guanine content. The computing system can also analyze the test sequencing reads to determine a first quantitative measure derived from the test sequencing reads corresponding to individual classification regions of the plurality of classification regions, at least a portion of the individual classification regions of the plurality of classification regions corresponding to genomic regions of a reference genome having the threshold amount of methylated cytosines and having at least the threshold cytosine-guanine content in the subject in whom cancer is detected. The computing system can also analyze the test sequencing reads to determine a second quantitative measure derived from the test sequencing reads corresponding to each control region of the plurality of control regions, where each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in the additional subject in which cancer is not detected. The computing system can also determine metrics for each classification region based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. The computing system can also generate an input vector including the metrics for each classification region. The computing system can also determine an indication that cancer is present in the subject by providing the input vector to a model implementing one or more machine learning techniques to generate an indication that cancer is present in the subject, the model including weights for each classification region of the plurality of classification regions, at least a portion of the weights for the individual classification regions being different from one another.

[0105] In one or more embodiments, the computing system can obtain training sequence data including training sequencing reads from a plurality of samples of a plurality of training subjects, each training sequencing read including a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one sample of the plurality of samples, and each training sequencing read corresponding to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The computing system can also analyze the training sequencing reads to determine an additional first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions. The computing system can also analyze the training sequencing reads to determine an additional second quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions. The computing system can also determine an additional metric for each classification region of the plurality of classification regions based on the additional first quantitative measure for each classification region and the additional second quantitative measure for the plurality of control regions. The computing system can also generate training data including an additional metric for each classification region of the plurality of classification regions for the training sequence reads from the samples of the plurality of training subjects. The computing system may also use the training data to implement one or more machine learning algorithms to generate a model for determining an indication that a subject has cancer based on the amount of methylated cytosines in at least a portion of a plurality of classification regions.

[0106] In one or more embodiments of the computing system, the one or more machine learning algorithms include one or more classification algorithms, and the indication that cancer is present corresponds to a probability that cancer is present in the subject.

[0107] In one or more embodiments of the computing system, the one or more machine learning algorithms include one or more regression algorithms, and the index corresponds to an estimate of the tumor fraction of the sample.

[0108] In one or more embodiments of the computing system, the training sequence reads include a first portion of the training sequence data, the additional training sequencing reads include a second portion of the training sequence data, and the additional training sequencing reads are different from the training sequencing reads, and the computing system can analyze at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine individual frequencies of a plurality of variants present in each sample of the plurality of samples. The computing system can also determine, for each sample, one variant of the plurality of variants having a highest frequency corresponding to the individual frequency having the maximum value among the individual frequencies from the individual samples. The computing system can also determine an individual measure of tumor fraction for each sample based on the maximum individual frequency from the individual samples.

[0109] In one or more embodiments of the computing device, the training data includes individual measures of tumor fraction for each of the plurality of samples, and the model is generated based on the individual measures of tumor fraction for each of the plurality of samples.

[0110] In one or more embodiments, a method includes obtaining sequencing data from a plurality of subjects by a computing system having one or more hardware processors and memory, the sequencing data including sequencing reads from a plurality of samples from the plurality of subjects, each of the sequencing reads including a nucleotide sequence corresponding to a fragment of a nucleic acid included in an additional sample, each of the sequencing reads corresponding to a molecule having a threshold amount of methylated cytosines included within a promoter region of a nucleotide sequence having at least a threshold cytosine-guanine content. The method may also include analyzing the sequencing reads by the computing system to determine a first quantitative measure derived from the sequencing reads corresponding to the promoter region. The method may also include analyzing the sequencing reads by the computing system to determine a second quantitative measure derived from the sequencing reads corresponding to individual control regions of a plurality of control regions, each of the control regions corresponding to additional genomic regions of a reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosines in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The method may also include determining, by a computing system, a metric for the promoter region based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. The method may also include generating, by the computing system, an indicator of the methylation state of the promoter region based on the metric having at least a threshold value.

[0111] In one or more embodiments, the computing system may include one or more hardware processors and one or more non-transitory computer-readable storage media containing computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including obtaining sequencing data from a plurality of subjects, the sequencing data including sequencing reads from a plurality of samples from the plurality of subjects, each sequencing read including a nucleotide sequence corresponding to a fragment of nucleic acid contained in an additional sample, each sequencing read corresponding to a molecule having a threshold amount of methylated cytosines contained within a promoter region of a nucleotide sequence having at least a threshold cytosine-guanine content. The system may also analyze the sequencing reads to determine a first quantitative measure derived from the sequencing reads corresponding to the promoter region. The system may also analyze the sequencing reads to determine a second quantitative measure derived from the sequencing reads corresponding to individual control regions of the plurality of control regions, each control region of the plurality of control regions corresponding to additional genomic regions of a reference genome having at least the threshold cytosine-guanine content and having at least the threshold amount of methylated cytosines in the subject in whom cancer is detected and in additional subjects in whom cancer is not detected. The system can also analyze the sequencing reads to determine a second quantitative measure derived from the sequencing reads corresponding to each control region of the plurality of control regions, where each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in the additional subject in which cancer is not detected. The system can also determine a metric for the promoter region based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. The system can also generate an index of the methylation state of the promoter region based on a metric having at least a threshold. (Mode for Carrying Out the Invention)

[0112] definition In order to more readily understand this disclosure, certain terms are first defined below. Additional definitions for these and other terms may be found throughout the specification. In the event that a definition of a term below conflicts with a definition in an application or patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.

[0113] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or that will become apparent to those skilled in the art upon reading this disclosure, and so forth.

[0114] It should also be understood that the terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer-readable media, and systems, the following terminology and grammatical variations thereof will be used in accordance with the definitions set forth below.

[0115] About: As used herein, "about" or "approximately," when applied to one or more values ​​or elements of interest, refers to a value or element that is similar to the specified reference value or element. In certain implementations, the term "about" or "approximately," unless otherwise specified or otherwise clear from the context, refers to a set of values ​​or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or fewer percent in either direction (greater or less) of the specified reference value or element (except where such number exceeds 100% of the possible values ​​or elements).

[0116] Administer: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means giving, applying, or contacting the composition to the subject. Administration may be accomplished by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.

[0117] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that can be used to ligate one or both ends of a given sample nucleic acid molecule that is at least partially double-stranded. The adapter may include nucleic acid primer binding sites to enable amplification of the nucleic acid molecule flanked by the adapters at both ends, and / or sequencing primer binding sites, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter may also include a binding site for a capture probe, such as an oligonucleotide attached to a flow cell support or the like. The adapter may also include a nucleic acid tag, as described herein. The nucleic acid tag can be positioned relative to the amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequence read of a given nucleic acid molecule. The same or different adapters can be ligated to each end of a nucleic acid molecule. In some implementations, the same adapter is ligated to each end of a nucleic acid molecule, except for the different nucleic acid tags. In some implementations, the adapters are Y-shaped adapters with one end blunt or tailed for joining to a similarly blunt or tailed nucleic acid molecule with one or more complementary nucleotides, as described herein. In yet another exemplary implementation, the adapters are bell-shaped adapters with a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other examples of adapters include T-tailed adapters and C-tailed adapters.

[0118] Alignment: As used herein, "alignment" or "aligning" refers to determining whether at least two sequence representations have at least a threshold amount of homology. In one or more examples, the threshold amount of homology can be at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9%. In situations where two sequence representations have at least a threshold amount of homology, the two sequence representations can be referred to as "aligned."

[0119] Amplify: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of a polynucleotide, starting from a small amount of polynucleotide (e.g., a single polynucleotide molecule), where an amplification product or amplicon is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes.

[0120] Barcode: As used herein, "barcode" or "molecular barcode" in the context of nucleic acids refers to a nucleic acid molecule that contains a sequence that can serve as a molecular identifier. For example, an individual "barcode" sequence can be added to each DNA fragment during next-generation sequencing (NGS) library preparation, thus allowing each read to be identified and sorted prior to final data analysis.

[0121] Cancer type: As used herein, "cancer type" refers to the type or subtype of cancer as defined, for example, by histopathology. Cancer type can be determined by any conventional criteria, for example, by occurrence in a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, etc. Cancers can be defined as cancers of unknown primary origin, etc., and / or based on the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or exhibiting cancer markers such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether primary or secondary.

[0122] Carrier Signal: As used herein, "carrier signal" refers to any intangible medium capable of storing, encoding, or carrying transitory or non-transitory instructions 502 for execution by machine 500, including digital or analog communication signals or other intangible media for facilitating communication of such instructions 502. Instructions 502 may be transmitted or received over network 534 using transitory or non-transitory transmission media, via network interface devices, and using any one of several well-known transfer protocols.

[0123] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acid that is not contained within or otherwise associated with cells, or, in some implementations, nucleic acid that remains in a sample after removal of intact cells. Cell-free nucleic acid can include, for example, all unencapsulated nucleic acids originating from a subject's bodily fluid (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acid can be double-stranded, single-stranded, or a hybrid thereof. Cell-free nucleic acid can be released into bodily fluids by secretion or cell death processes, such as cell necrosis, apoptosis, etc. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released from cancer cells into body fluids. Others are released from healthy cells. ctDNA can be unencapsulated tumor-derived fragmented DNA. Cell-free nucleic acids can have one or more epigenetic modifications, for example, cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.

[0124] Cellular nucleic acid: As used herein, "cellular nucleic acid" refers to nucleic acid that is located within one or more cells, at least at the time the sample is taken or collected from a subject, even if the nucleic acid is later removed as part of a given analytical process.

[0125] Classification region: As used herein, "classification region" refers to a genomic region that may exhibit sequence-independent alterations in neoplastic cells (e.g., tumor cells and cancer cells) or that may exhibit sequence-independent alterations in cfDNA from subjects with cancer compared to cfDNA from subjects without cancer. Examples of sequence-independent alterations include, but are not limited to, alterations in methylation rates (increase or decrease), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. In one or more examples, sequence-independent alterations in a classification region may indicate the presence of a single form of cancer in a subject. In one or more additional examples, sequence-independent alterations in a classification region may correspond to the presence of multiple forms in a subject. A classification region can be enriched by one or more probes. Furthermore, a classification region can be defined by a pair of primer binding sites. Furthermore, a classification region can be defined by a predetermined start genomic locus and a predetermined end genomic locus. A classification region can include from about 25 nucleotides to about 250 nucleotides, from about 50 nucleotides to about 200 nucleotides, or from about 75 nucleotides to about 150 nucleotides. For example, classification region can be a differentially methylated region. " Differentially methylated region " or " DMR " refers to the region of DNA that has detectably different degree of methylation in at least one cell or tissue type compared with the degree of methylation in the same region of DNA from at least one other cell or tissue type, or that has detectably different degree of methylation in at least one cell or tissue type obtained from a subject with disease or disorder compared with the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject.In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., hypermethylated region / hypermethylated target region) in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type that contributes to cfDNA in a healthy individual, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., hypomethylated region / hypomethylated target region) in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in a healthy individual, or from the same cell or tissue type from a healthy subject. In some embodiments, the classification region comprises a hypermethylated target region and / or a hypomethylated target region.

[0126] Communications Network: As used herein, a "communications network" refers to one or more portions of a network 114, 1034, which may be an ad-hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a plain old telephone service (POTS) network, a cellular network, a wireless network, a Wi-Fi network, another type of network, or a combination of two or more such networks. For example, the network 114, 1034 or a portion of a network may include a wireless or cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or other type of cellular or wireless coupling.In this example, the coupling may implement any of various types of data transport technologies, such as Single Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP), including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High Speed ​​Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) standards, other technologies defined by various standards-setting organizations, other long-range protocols, or other data transport technologies.

[0127] Confidence interval: As used herein, "confidence interval" refers to a set of values ​​defined within which there is a certain probability that the value of a given parameter falls within that range of values.

[0128] Control sample: As used herein, a "control sample" or "reference sample" refers to a sample obtained from an individual who does not have a known copy number variation.

[0129] Coverage: As used herein, "coverage" or "coverage metrics" refers to the number of nucleic acid molecules or sequencing reads that correspond to a particular genomic region of a reference sequence.

[0130] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides that have a hydrogen group at the 2' position of the sugar moiety. DNA can contain a chain of nucleotides containing four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides that have a hydroxyl group at the 2' position of the sugar moiety. RNA can contain a chain of nucleotides containing four types of nucleotides: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural or modified nucleotides. Certain pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to those of the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "sequence representation," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," "fragment sequence," "sequencing read," or "nucleic acid sequencing read" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule of nucleic acid, such as DNA or RNA (e.g., a whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment).It should be understood that the present teachings contemplate sequence information obtained using all available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.

[0131] Differentially methylated region: As used herein, "differentially methylated region" refers to a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, or that has a detectably different degree of methylation in at least one cell or tissue type obtained from a subject with a disease or disorder compared to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type, such as at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., other immune cell types and / or cell types that contribute to cfDNA in a healthy individual, or from the same cell or tissue type from a healthy subject.

[0132] Driver mutation: As used herein, "driver mutation" means a mutation that drives cancer progression.

[0133] Epigenetic target region: As used herein, "epigenetic target region" refers to a target region that may exhibit sequence-independent differences in different cell or tissue types (e.g., different types of immune cells) or in neoplastic cells (e.g., tumor cells and cancer cells) compared to normal cells, or that may exhibit sequence-independent differences (i.e., differences that do not result in a change in nucleotide sequence, e.g., differences in methylation, nucleosome distribution, or other epigenetic features) in DNA, e.g., cfDNA, from different cell types or from subjects with cancer compared to DNA, e.g., cfDNA, from healthy subjects, or in cfDNA originating from different cell or tissue types (e.g., immune, lung, colon, etc.) that do not normally contribute substantially to cfDNA compared to background cfDNA (e.g., cfDNA originating from hematopoietic cells). Examples of sequence-independent changes include, but are not limited to, methylation (increase or decrease), nucleosome distribution, cfDNA fragmentation pattern, CCCTC binding factor ("CTCF") binding, transcription start site (e.g., in relation to any one or more of the following: RNA polymerase component binding, regulatory protein binding, fragmentation characteristics, and nucleosome distribution), and changes in regulatory protein binding regions.Thus, epigenetic target region sets include, but are not limited to, hypermethylated target region sets, hypomethylated target region sets, and fragmentation variable target region sets, such as CTCF binding sites and transcription start sites.For the present purposes, epigenetic target region sets may also include loci that are susceptible to local amplification and / or gene fusion associated with neoplasia, tumor, or cancer, because the detection of copy number changes by sequencing or fusion sequences that map to more than one locus in a reference genome tends to be more similar to the detection of exemplary epigenetic changes discussed above than the detection of nucleotide substitutions, insertions, or deletions, for example, in that the detection of local amplification and / or gene fusion does not depend on the accuracy of base calls at one or a few individual positions and can therefore be detected at a relatively low sequencing depth. An epigenetic target region set is a set of epigenetic target regions.

[0134] Hypermethylated: As used herein, "hypermethylated" refers to an increased level or degree of methylation of a nucleic acid molecule(s) relative to other nucleic acid molecules within a population (e.g., a sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypermethylated DNA may include DNA molecules that contain at least one methylated cytosine, at least two methylated cytosines, at least three methylated cytosines, at least five methylated cytosines, or at least ten methylated cytosines.

[0135] Hypomethylated: As used herein, "hypomethylated" refers to a reduced level or degree of methylation of a nucleic acid molecule(s) relative to other nucleic acid molecules within a population (e.g., a sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA may include DNA molecules containing zero methylated cytosines, at most one methylated cytosine, at most two methylated cytosines, at most three methylated cytosines, at most four methylated cytosines, or at most five methylated cytosines.

[0136] Immunotherapy: As used herein, "immunotherapy" refers to treatment with one or more agents that stimulate the immune system to kill or at least inhibit the growth of cancer cells, preferably to reduce further cancer growth, shrink the size of cancer, and / or eliminate cancer. Some such agents bind to targets present in cancer cells, some bind to targets present in immune cells but not in cancer cells, and some bind to targets present in both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of immune system pathways that maintain self-tolerance and modulate the duration and magnitude of physiological immune responses in peripheral tissues, minimizing secondary tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Examples of agents include antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. Other examples of agents include pro-inflammatory cytokines such as IL-1β, IL-6, and TNF-α. Another example of an agent is T cells activated against tumors, for example, T cells activated by expressing a chimeric antigen that targets a tumor antigen recognized by the T cell.

[0137] Indel: As used herein, "indel" refers to a mutation involving the insertion or deletion of nucleotides in the genome of a subject.

[0138] Limit of Detection (LoD): As used herein, "limit of detection" refers to the smallest amount of a substance (e.g., nucleic acid) in a sample that can be measured by a given assay or analytical approach.

[0139] Machine-readable medium: As used herein, "machine-readable medium" refers to a component, device, or other tangible medium capable of temporarily or permanently storing instructions 502 and data, which may include, but are not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term "machine-readable medium" may be considered to include a single medium or multiple media (e.g., centralized or distributed databases, or associated caches and servers) capable of storing instructions 502. The term "machine-readable medium" shall also be considered to include any medium or combination of media capable of storing instructions 502 (e.g., code) for execution by a machine 500 such that the instructions 502, when executed by one or more processors 504 of the machine 500, cause the machine 500 to perform any one or more of the methodologies described herein. Thus, a "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network that includes multiple storage devices or apparatuses. Signals themselves are excluded from the term "machine-readable medium."

[0140] Maximum MAF: As used herein, "maximum MAF" or "max MAF" refers to the maximum MAF (variant allele fraction) of all somatic variants in a sample.

[0141] Methylation: As used herein, the term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to a cytosine at a CpG site (i.e., a cytosine followed by a guanine in the 5' to 3' direction of a nucleic acid sequence). In some embodiments, DNA methylation refers to the addition of a methyl group to an adenine, e.g., N 6 -methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the six-carbon ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine to create 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the six-carbon ring of cytosine). In some embodiments, 3C methylation includes the addition of a methyl group to the 3C position of cytosine to create 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA within a promoter region is methylated, gene transcription can be suppressed. DNA methylation is crucial for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption, e.g., suppression, in epigenetic regulation can cause diseases, such as cancer. Promoter methylation in DNA can indicate cancer.

[0142] Methylation-dependent nuclease: As used herein, the term "methylation-dependent nuclease" refers to a nuclease that preferentially cleaves methylated DNA compared to unmethylated DNA. For example, a methylation-dependent nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the methylation of at least one of the nucleic acid bases, e.g., cytosine, within the recognition sequence. In some embodiments, the nucleolytic activity of a methylation-dependent nuclease is at least 10-fold, 20-fold, 50-fold, or 100-fold higher for a methylated recognition site compared to an unmethylated control in a standard nucleolytic assay. Methylation-dependent nucleases include methylation-dependent restriction enzymes.

[0143] Methylation-dependent restriction enzyme: As used herein, "methylation-dependent restriction enzyme" or "MDRE" refers to a restriction enzyme that depends on DNA methylation (e.g., cytosine methylation), i.e., the presence or absence of methyl groups in nucleotide bases alters the rate at which the enzyme cleaves target DNA. In some embodiments, a methylation-dependent restriction enzyme does not cleave DNA if a particular nucleotide base is unmethylated in the recognition sequence. For example, MspJI is a methylation-dependent restriction enzyme with the recognition sequence "mCNNR(N9)," and does not cleave DNA if methylated cytosine (mC) is absent within the recognition sequence.

[0144] Methylation-sensitive nuclease: As used herein, the term "methylation-sensitive nuclease" refers to a nuclease that preferentially cleaves unmethylated DNA compared to methylated DNA. For example, a methylation-sensitive nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the lack of methylation of at least one of the nucleic acid bases, e.g., cytosine, within the recognition sequence. In some embodiments, the nucleolytic activity of a methylation-sensitive nuclease is at least 10-fold, 20-fold, 50-fold, or 100-fold higher for an unmethylated recognition sequence than for a methylated control in a standard nucleolytic assay. Methylation-sensitive nucleases include methylation-sensitive restriction enzymes.

[0145] Methylation-sensitive restriction enzyme: As used herein, "methylation-sensitive restriction enzyme" or "MSRE" refers to a restriction enzyme that is sensitive to the methylation state of DNA (e.g., cytosine methylation), i.e., the presence or absence of a methyl group at a nucleotide base alters the rate at which the enzyme cleaves a target DNA. In some embodiments, a methylation-sensitive restriction enzyme does not cleave DNA if a particular nucleotide base is methylated in the recognition sequence. For example, HpaII is a methylation-sensitive restriction enzyme with the recognition sequence "CCGG," and does not cleave DNA if the second cytosine in the recognition sequence is methylated.

[0146] Methylation rate: As used herein, "methylation rate" refers to the probability, likelihood, or percentage that a given base (e.g., a cytosine residue of a CpG) on a DNA molecule in a particular genomic region analyzed in a sample is methylated. In some embodiments, the methylation rate can be applied to a defined region containing one or more potentially methylated bases. In some embodiments, the methylation rate refers to the percentage of methylated CpG residues in a DNA molecule. In some embodiments, the methylation rate refers to the percentage of methylated CpG residues in molecules aligned to a particular genomic location or genomic region. Methylation rates can be measured by various methods, including, but not limited to, using bisulfite sequencing (any single-base resolution, such as TAPS, EM-SEQ, etc.) or partitioning (DNA molecule resolution). Methylation rates can be measured by various methods. One estimation can be made by counting how many DNA fragments end up in each methylation-dependent partition, or by counting the number of converted CpGs per fragment in the case of bisulfite sequencing or any other base-level resolution sequencing method. Furthermore, in the case of methylation-dependent partitioning, the calculation of the rate can be normalized using a set of predefined regions with known methylation status (i.e., positive and / or negative control regions) or spike-in synthetic DNA with known methylation status to derive a rate-parameterized partition distribution and estimate the rate using a maximum likelihood approach. In one or more examples, the methylation rate can be determined by determining the abundance of sequencing reads corresponding to a portion of a genomic region. The portion of a genomic region may include several genomic positions of the genomic region that are overlapped by at least a threshold number of sequencing reads.

[0147] Methylation state: As used herein, "methylation state" or "methylation state" can refer to the presence or absence of a methyl group on a DNA base (e.g., cytosine) at a particular genomic position in a nucleic acid molecule. It can also refer to the degree of methylation in a nucleic acid sequence (e.g., a hypermethylated, hypomethylated, intermediately methylated, or unmethylated nucleic acid molecule). Methylation state can also refer to the number of methylated nucleotides in a particular nucleic acid molecule.

[0148] Modified nucleotide-specific binding reagent: As used herein, refers to a binding reagent that is specific for or targets a modified nucleotide. For example, the modified nucleotide may be a methylated nucleotide, and thus the binding reagent may be specific for methylated nucleotides. Examples of binding reagents include, but are not limited to, the methyl-binding domain (MBD) of a methylation-binding protein ("MBP") or variants thereof, antibodies (and antibody variants, e.g., single-chain antibodies), aptamers, or combinations thereof. Thus, as disclosed throughout, the use of an MBD can be interchanged with any other modified nucleotide-specific binding reagent, provided that the modified nucleotide-specific binding reagent has the desired specificity and affinity for the particular modified base of interest in the selected implementation.

[0149] Variant allele fraction: As used herein, "variant allele fraction," "mutation dose," or "MAF" refers to the fraction of nucleic acid molecules that have an allelic alteration or mutation at a given genomic location in a given sample. The MAF is generally expressed as a fraction or percentage. For example, the MAF can be less than about 0.5, less than 0.1, less than 0.05, or less than 0.01 (i.e., less than about 50%, less than 10%, less than 5%, or less than 1%) of all somatic variants or alleles present at a given locus.

[0150] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence, including, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / abnormalities, insertions or deletions (indels), gene fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants.Mutation can be a germline mutation or a somatic mutation.In some cases, the reference sequence for comparison is the wild-type genome sequence of the species of the subject that provides the test sample, typically the human genome.

[0151] Variation caller: As used herein, "variation caller" refers to an algorithm (embodied in software or otherwise computer-implemented) used to identify variants in test sample data (e.g., sequence information obtained from a subject).

[0152] Mutation count: As used herein, "mutation count" or "mutational count" refers to the number of somatic mutations in the whole genome or exome or targeted region of a nucleic acid sample.

[0153] Negative control region: As used herein, "negative control region" refers to a genomic region that is expected to be unmethylated or hypomethylated in essentially all samples, regardless of whether the DNA is from cancer cells or normal cells.

[0154] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, presumably malignant, or malignant. A malignant tumor is called a cancer or cancerous tumor.

[0155] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, e.g., the ability to generate hundreds of thousands of relatively small sequencing reads at a time. Some examples of next-generation sequencing technologies include, but are not limited to, sequencing-by-synthesis, sequencing-by-ligation, and sequencing-by-hybridization.

[0156] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, less than about 50 nucleotides, or less than about 10 nucleotides in length) used to identify nucleic acids from different samples (e.g., representing a sample index) or to identify different nucleic acid molecules in the same sample that have been differently typed or processed (e.g., representing a molecular barcode). Nucleic acid tags comprise predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or various lengths, as desired. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, 5' or 3' single-stranded regions (e.g., overhangs), and / or one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample of origin, form, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, where the nucleic acids are subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or amplicons of different parent molecules in the same sample or subsample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.For non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their intrinsic sequence information (e.g., start and / or end positions where they map to a selected reference sequence, subsequences at one or both ends of the sequence, and / or sequence length) in combination with at least one molecular barcode. A sufficient number of different molecular barcodes are used so that the probability that any two molecules will have the same intrinsic sequence information (e.g., start and / or end positions, subsequences at one or both ends of the sequence, and / or length), as well as the same molecular barcode, is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% likelihood).

[0157] Partitioning: As used herein, "partitioning" refers to physically separating or fractionating a mixture of nucleic acid molecules in a sample based on characteristics of the nucleic acid molecules. Partitioning can be a physical partitioning of molecules. Partitioning can include separating nucleic acid molecules into groups or sets based on the level of epigenetic traits (e.g., related to methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is incorporated herein by reference in its entirety.

[0158] Distribution set: As used herein, "distribution set" or "distribution" refers to a set of nucleic acid molecules distributed into sets or groups based on the differential binding affinity of the nucleic acid molecules or proteins associated with the nucleic acid molecules to a binder. A distribution set may also be referred to as a subsample. A binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder may be a methyl-binding domain (MBD) protein. In some embodiments, a distribution set may include nucleic acid molecules belonging to a particular level or degree of epigenetic trait (e.g., methylation). For example, nucleic acid molecules may be distributed into three sets: one set of highly methylated nucleic acid molecules (first subsample, high distribution, high distribution set, or high methylation distribution set), a second set of hypomethylated nucleic acid molecules (second subsample, low distribution, low distribution set, or hypomethylation distribution set), and a third set of intermediately methylated nucleic acid molecules (third subsample, intermediate distribution set, intermediate methylation distribution set, residual distribution, or residual distribution set). In another example, the nucleic acid molecules may be distributed based on the number of methylated nucleotides, with one distribution set having nucleic acid molecules with 9 methylated nucleotides and another distribution set having unmethylated nucleic acid molecules (0 methylated nucleotides).

[0159] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," "polynucleotide molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleoside linkages. A polynucleotide can contain at least three nucleosides. Oligonucleotides often range in size from a small number of monomeric units, e.g., 3-4, to hundreds of monomeric units. Whenever a polynucleotide is represented by a sequence of letters, e.g., "ATGCCTG," it is understood that the nucleotides are in 5'→3' order from left to right, and that, in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine, unless otherwise specified. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide that includes the base, as is standard in the art.

[0160] Positive control region: As used herein, "positive control region" refers to a genomic region that is expected to be methylated or hypermethylated in essentially all samples, regardless of whether the DNA is from cancer cells or normal cells.

[0161] Probe: As used herein, "probe" refers to a polynucleotide that includes a function. The function may be a detectable label (fluorescence), a binding moiety (biotin), or a solid support (magnetically attractable particle or chip). A probe may include a single-stranded DNA / RNA polynucleotide or a double-stranded DNA polynucleotide (e.g., SureSelect® Probe, Agilent Technologies) that hybridizes to a target nucleic acid sequence. Sequence capture using a probe generally depends in part on the number of consecutive nucleotides in at least a portion of the target nucleic acid sequence that are complementary (or nearly complementary) to the sequence of the probe. In some examples, the probe may correspond to a driver mutation.

[0162] Processing: As used herein, the terms "processing," "calculating," and "comparing" can be used interchangeably. In certain applications, the term refers to determining differences, such as differences in number or sequence. For example, gene expression, copy number variation (CNV), indel, and / or single nucleotide variant (SNV) values ​​or sequences can be processed.

[0163] Processor: As used herein, "processor" refers to any circuit or virtual circuit (a physical circuit emulated by logic running on an actual processor) that manipulates data values ​​in accordance with control signals (e.g., "commands," "op codes," "machine code," etc.) to produce corresponding output signals that are applied to operate a machine. A processor may be, for example, a CPU, a RISC processor, a CISC processor, a GPU, a DSP, an ASIC, an RFIC, or any combination thereof. A processor may also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") capable of simultaneously executing instructions.

[0164] Promoter region: As used herein, "promoter region" refers to a DNA sequence recognized by the synthetic machinery of the cell, or introduced synthetic machinery, necessary to initiate the specific transcription of a gene.

[0165] As used herein, "quantitative measurement" refers to an absolute or relative measurement. A quantitative measurement can be, but is not limited to, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or relative quantity (e.g., high, medium, and low). A quantitative measurement can be a ratio of two quantitative measurements. A quantitative measurement can be a linear combination of quantitative measurements. A quantitative measurement can be a normalized measurement.

[0166] As used herein, " reference sequence " refers to the known sequence used for the purpose of comparison with experimentally determined sequences.For example, known sequence can be the whole genome, chromosome, or any segment thereof.Reference sequence can comprise at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000 or more nucleotides.Reference sequence can be aligned with a single continuous sequence of genome or chromosome, or can comprise discontinuous segments that are aligned with different regions of genome or chromosome.Examples of reference sequence include, for example, human genome reference sequence, such as hG19 and hG38.

[0167] As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.

[0168] Sensitivity: As used herein, "sensitivity" means the probability of detecting the presence of single-base variants, insertions, and deletions at a given MAF and coverage, and the probability of detecting the presence of copy number variants at a given tumor fraction and coverage.

[0169] As used herein, "sequencing" refers to any of several technologies used to determine the sequence (e.g., identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, and the like. Examples of sequencing techniques include high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperatures (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some implementations, sequencing can be performed by a genetic analyzer, such as a commercially available genetic analyzer from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.

[0170] Single nucleotide variant: As used herein, "single nucleotide variant" or "SNV" refers to a mutation or variation in a single base that occurs at a specific position in the genome.

[0171] Somatic mutation: As used herein, the terms "somatic mutation" or "somatic mutation" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.

[0172] Specific binding: As used herein, "specific binding," in the context of a probe or other oligonucleotide and a target sequence, means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a copy thereof to form a stable probe:target hybrid, while at the same time minimizing the formation of stable probe:non-target hybrids. Thus, the probe hybridizes to the target sequence or a copy thereof to a sufficiently greater extent than to non-target sequences, allowing for capture or detection of the target sequence. Suitable hybridization conditions are well known in the art and can be predicted based on sequence composition or determined using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, especially §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, which are incorporated herein by reference).

[0173] As used herein, "subject" refers to an animal, such as a mammalian species (e.g., a human), or an avian (e.g., a bird) species, or other organism, such as a plant. More specifically, a subject can be a vertebrate, such as a mammal, such as a mouse, a primate, a monkey, or a human. Animals include farm animals (e.g., beef cattle, dairy cattle, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or service animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended interchangeably with "subject."

[0174] For example, the subject can be the individual who has been diagnosed with cancer, who is going to receive cancer treatment, and / or who has received at least one cancer treatment.The subject can be in the remission stage of cancer.As another example, the subject can be the individual who has been diagnosed with autoimmune disease.As another example, the subject can be the female individual who is pregnant or planning to become pregnant, who has been diagnosed with or suspected to have disease, such as cancer, autoimmune disease.

[0175] Target region: As used herein, "target region" refers to a genomic locus that is targeted for identification and / or capture, e.g., by using a probe (e.g., by sequence complementarity). A "target region set" or "set of target regions" refers to a plurality of genomic loci that are targeted for identification and / or capture, e.g., by using a set of probes (e.g., by sequence complementarity).

[0176] Threshold: As used herein, "threshold" refers to a predetermined value used to characterize experimentally determined values ​​for the same parameter in different samples according to their relationship to the threshold.

[0177] Tumor fraction: As used herein, " tumor fraction " refers to the estimated fraction of nucleic acid molecules derived from tumor in a given sample.For example, the tumor fraction of sample can be the maximum MAF of sample or the pattern of sequencing coverage of sample or the length of the cfDNA fragments in sample or any other selected feature of sample.In some cases, the tumor fraction of sample is equal to the maximum MAF of sample.

[0178] Variant: As used herein, "variant" can be referred to as an allele. Variants are usually present at a frequency of 50% (0.5) or 100% (1), depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. However, somatic variants are acquired variants and usually have a frequency of <0.5. The major and minor alleles of a genetic locus refer to nucleic acids having a locus that is occupied by a nucleotide of a reference sequence, and each variant nucleotide is different from the reference sequence. Measurements at a locus can be obtained in the form of an allele fraction (AF), which measures the frequency at which an allele is observed in a sample. (Mode for Carrying Out the Invention)

[0179] Detailed Description Cancer is usually caused by the accumulation of mutations in the genes of an individual's cells, at least in part resulting in improperly regulated cell division. Such mutations can include single nucleotide variations (SNVs), gene fusions, insertions, transversions, translocations, and inversions. These mutations can also include copy number variations, which correspond to an increase or decrease in the copy number of genes in the tumor genome compared to the individual's non-cancerous cells. The degree of mutation present in cell-free nucleic acid and the amount of mutated cell-free nucleic acid in a sample can be used as biomarkers to determine tumor progression, predict patient outcome, and improve treatment selection. In various examples, the degree of mutation present in cell-free nucleic acid can be indicated by tumor cell copy number and tumor fraction for a given sample.

[0180] Additionally, cancer can be manifested by non-sequence alterations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation in CpG islands at the TSSs of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This increased methylation level can be associated with abnormal loss of transcriptional capacity of the associated gene and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression.

[0181] Therefore, DNA methylation profiling can be used to detect the abnormal methylation in the DNA of sample.The DNA is usually hypermethylated or hypomethylated in a given sample type (for example, the cfDNA derived from bloodstream), but for example, because the tissue contribution to this sample type is abnormally increased (for example, due to the increased DNA loss in or around neoplasia or cancer), and / or the degree of genomic methylation is changed during development or is disturbed by disease, for example, cancer or any cancer-related disease, the degree of abnormal methylation can correspond to a certain genomic region (" differentially methylated region " or " DMR "), which can show the degree of abnormal methylation that correlates with neoplasia or cancer.

[0182] In some methods for measuring DNA methylation, it may be difficult to accurately determine the amount of DNA methylation.The accuracy of determining DNA methylation may affect the accuracy of the estimated tumor fraction for a sample.Because tumor fraction can be used to determine whether a sample is derived from a subject with a tumor, the accuracy of determining the estimated tumor fraction may affect the diagnosis and / or treatment decision for an individual.

[0183] The methods and systems described herein are directed to accurately generating information indicative of multiple amounts of nucleic acid methylation using data indicative of the amount of nucleic acid binding to a methyl-binding domain (MBD). In various examples, the present application is directed to systems and processes for determining an estimate of the tumor fraction of a sample. In one or more examples, the amount of methylation of a nucleic acid can be determined based on the strength of binding of the nucleic acid to the methyl-binding domain (MBD). The nucleic acids can be distributed according to the strength of binding to the MBD. Furthermore, the number of cytosine-guanine (CG) regions can be determined for the nucleic acid. The amount of methylation of a classification region of a nucleic acid can be determined based on the distribution information associated with the nucleic acid and the number of cytosine-guanine regions of the nucleic acid. The classification region can have different amounts of methylation between tumor cells and non-tumor cells. The estimate of the tumor fraction of a sample can be determined according to the amount of methylation of the classification region.

[0184] In at least some implementations, the methods, systems, techniques, and architectures can implement models configured to have at least one parameter or weight that can be modified to more accurately fit the methylation data provided to the model. The methods, systems, techniques, and architectures are also directed to implementing several optimization procedures during model training to produce models that predict metrics indicative of the presence or absence of tumors more accurately than other systems, methods, techniques, and architectures. Furthermore, the methods, techniques, and processes used to generate the information used to produce the methylation data reduce the amount of noise present in the methylation data, thereby resulting in more accurate predictions of metrics indicative of the presence or absence of tumors than other methods, techniques, and processes.

[0185] 1 is a diagrammatic representation of an example environment 100 for identifying nucleic acids corresponding to classification regions of a reference sequence, where the classification regions have at least a threshold number of CpGs, according to one or more implementations. In one or more examples, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include bile duct cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.

[0186] The environment 100 may include a sample 102. The sample 102 may be derived from a biological fluid obtained from a subject. For example, the sample 102 may be derived from blood obtained from a subject. In one or more additional examples, the sample 102 may be derived from a tissue of the subject. In various examples, the sample 102 may be derived from multiple sources. For example, the sample 102 may be derived from one or more bodily fluids and / or tissues of the subject. In one or more instances, the subject may be a mammal. In one or more additional instances, the subject may be a human. In one or more further instances, the subject may be a non-human mammal.

[0187] The sample 102 may include several nucleic acids 104. Each nucleic acid 104 may include several regions having at least a threshold number of cytosine and guanine molecules. In one or more examples, each nucleic acid 104 may include a region having at least a threshold number of cytosine-guanine dinucleotides. In various examples, at least a portion of the cytosine-guanine pairs contained within a region may be sequentially positioned within the sequence of the nucleic acid 104. In one or more instances, a region of a nucleic acid having at least a threshold amount of cytosine-guanine pairs may be referred to herein as a "CG region" or "CpG region." In one or more examples, a CG region may include at least 200 CpG dinucleotides. In one or more instances, a CG region may contain between 200 and 5000 CpG dinucleotides, between 300 and 3000 CpG dinucleotides, between 200 and 2500 CpG dinucleotides, or between 500 and 1500 CpG dinucleotides. Furthermore, a CG region may have a GC percentage of at least 50% and an observed to predicted CpG ratio of at least 60%. The observed to predicted CpG ratio can be calculated as the number of cytosines multiplied by the number of guanines divided by the number of bases in a genomic region, where observed CpGs is the number of CpGs identified within a given genomic region. Predicted CpGs are: ((number of cytosines + number of guanines) / 2)2 / length of genomic region For example, CG regions can be determined using the techniques described by Gardiner-Garden M, Frommer M (1987). "CpG islands in vertebrate genomes". Journal of Molecular Biology. 196 (2): 261-282. and / or Saxonov S, Berg P, Brutlag DL (2006). "A genome-wide analysis of CpG dinucleotides in the human genome distinguishes two distinct classes of promoters". Proc Natl Acad Sci USA. 103 (5): 1412-1417.

[0188] 1, a portion of the sequence of an example nucleic acid 104 can include a first CG region 106, a second CG region 108, and a third CG region 110. Although the example of FIG. 1 illustrates a portion of the sequence of a nucleic acid 104 having three CG regions, the nucleic acids 104 included in the sample 102 can have a different number of CG regions. For example, an individual nucleic acid 104 included in the sample 102 can include at least 1 CG region, at least 5 CG regions, at least 10 CG regions, at least 25 CG regions, at least 50 CG regions, at least 100 CG regions, at least 250 CG regions, at least 500 CG regions, or at least 1000 CG regions.

[0189] An individual CG region may correspond to several molecules with one or more methylated cytosines. In the example of FIG. 1, CG region 106 may include molecules with methylated cytosine 112. In the example of FIG. 1, the molecules with methylated cytosine 112 are 5-methylcytosines. An individual CG region may also correspond to several molecules with unmethylated cytosines. For example, CG region 106 may include molecules with unmethylated cytosine 116. In various examples, at least a portion of the CG region of nucleic acid 104 may correspond to a classification region of a reference genome. The classification region may correspond to a genomic region of the reference genome that corresponds to a non-sequence difference consistent with one or more biological conditions, such as one or more types of cancer. In at least some examples, the non-sequence difference may include one or more mutations consistent with one or more biological conditions. In one or more examples, the classification region may correspond to a genomic region of a reference sequence for molecules derived from subjects with at least one form of cancer. In at least some instances, nucleic acid molecules having at least a threshold amount of methylated cytosines in at least one CG region (e.g., hypermethylated molecules) may be derived from a subject with cancer and may correspond to a classification. In one or more additional instances, nucleic acid molecules having less than a threshold amount of methylated cytosines in at least one CG region (e.g., hypomethylated molecules) may be derived from a subject with cancer and may correspond to a classification region.

[0190] In addition to the classification region, the CG region may include one or more positive control regions, such as positive control region 118. The positive control region 108 has at least a threshold number of methylated cytosine molecules within at least one CG region and may be mapped to nucleic acid molecules derived from subjects without cancer and from subjects with cancer. In various examples, the positive control region 106 may be hypermethylated in cells derived from subjects without cancer and also from subjects with cancer. The CG region may also include one or more negative control regions, such as negative control region 120. The negative control region 120 has less than a threshold number of methylated cytosine molecules within at least one CG region and may be mapped to nucleic acid molecules derived from subjects without cancer and also from subjects with cancer. In one or more examples, the negative control region 120 may be hypomethylated in subjects without cancer and also from subjects with cancer. In various examples, the positive control region and the negative control region can be used to perform normalization calculations. Normalization calculations can be performed to generate input data for one or more models implemented to determine tumor metrics for a given sample 102 .

[0191] A first molecular separation process 122 can be performed. The first molecular separation process 122 can separate the nucleic acids 104 contained in the sample 102 based on the amount of methylated cytosine in each individual nucleic acid 104. In one or more examples, the first molecular separation process can separate the nucleic acids 104 contained in the sample 102 based on the amount of methylated cytosine in a CG region of each individual nucleic acid 104. In various examples, the first molecular separation process 122 can separate the nucleic acids 104 into multiple groups, each group corresponding to a respective amount of methylated cytosine in the nucleic acids 104.

[0192] 1 , a first molecular separation process 122 can be performed in association with a first methylation threshold 124. Performing the first molecular separation process 122 in association with the first methylation threshold 124 can result in a first distribution 126 of nucleic acids. In one or more examples, the first methylation threshold 124 can indicate a first threshold number of molecules having methylated cytosines located within CG regions of the nucleic acids 104. The first molecular separation process 122 can identify some nucleic acids 104 having fewer molecules with methylated cytosines within the CG regions than the first methylation threshold 124. In various examples, the first methylation threshold 124 can correspond to a first methylation rate.

[0193] The first molecular separation process 122 can also be performed with respect to a second methylation threshold 128. The second methylation threshold 128 can indicate an amount of methylated cytosine in one or more genomic regions of the nucleic acids 104 that is greater than the amount of methylated cytosine in one or more regions corresponding to the first methylation threshold 124. The second methylation threshold 124 can indicate a number of molecules having methylated cytosine, depending on the number of nucleic acids. In one or more additional examples, the second methylation threshold 124 can correspond to a higher rate of methylation of the nucleic acids than the rate of methylation corresponding to the first methylation threshold 124. Performing the first molecular separation process 122 with respect to the second methylation threshold 128 can result in a second distribution 130 of the nucleic acids. In one or more examples, the first molecular separation process 122 can identify nucleic acids 104 having an amount of methylated cytosine greater than a first methylation threshold 124 and an amount of methylated cytosine less than a second methylation threshold 128, resulting in a second distribution 130 of nucleic acids.

[0194] Additionally, the first molecular separation process 122 can be performed with respect to a third methylation threshold 132. The third methylation threshold 132 can indicate an amount of methylated cytosine in one or more genomic regions of the nucleic acids 104 that is greater than the amount of methylated cytosine in one or more regions corresponding to the first methylation threshold 124 and greater than the amount of methylated cytosine in one or more regions corresponding to the second methylation threshold 128. The third methylation threshold 132 can indicate a number of molecules having methylated cytosine, depending on the number of nucleic acids. In one or more additional examples, the third methylation threshold 132 can correspond to a rate of methylated cytosine that is greater than the rate of methylation corresponding to the first methylation threshold 124 and greater than the rate of methylation corresponding to the second methylation threshold 128. Performing the first molecular separation process 122 with respect to the third methylation threshold 132 can result in a third distribution 134 of nucleic acids. In one or more examples, the first molecular separation process 122 can identify nucleic acids 104 that have a higher amount of methylated cytosine than the nucleic acids 104 contained in the second distribution of nucleic acids 128. Thus, the amount of methylated cytosine of the nucleic acids contained in the first distribution 122, the second distribution 126, and the third distribution 130 increases from the first distribution 122 to the second distribution 126 and increases from the second distribution 126 to the third distribution 130. In one or more instances, the first distribution of nucleic acids 126 can be referred to as a low-methylated distribution, the second distribution of nucleic acids 130 can be referred to as an intermediate distribution, and the third distribution of nucleic acids 134 can be referred to as a high-methylated distribution.

[0195] In one or more examples, the amount of methylated cytosine in a nucleic acid may correspond to the strength of binding to a methyl-binding domain (MBD). In these scenarios, the first partitioning 126, second partitioning 130, and third partitioning 134 may occur based on the different strengths of binding to the MBD for nucleotides with different amounts of methylated cytosine. In one or more examples, the first molecular separation process 122 may include a series of washes in which the nucleic acid 104 is contacted with solutions having different concentrations of sodium chloride (NaCl).

[0196] Partitioning of nucleic acids can be performed by contacting the nucleic acid with a modified nucleotide-specific binding reagent, such as the MBD of MBP. The modified nucleotide-specific binding reagent can bind to 5-methylcytosine (5mC). The modified nucleotide-specific binding reagent, such as the MBD, can be coupled to paramagnetic beads such as Dynabeads® M-280 streptavidin via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by increasing the NaCl concentration in a series of washes. Sequences eluted from the modified nucleotide-specific binding reagent are partitioned into two or more fractions (e.g., low, high) depending on which wash (e.g., NaCl concentration) the sequence eluted from. The resulting partition can contain one or more of the following nucleic acid forms: double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments.

[0197] Binding of nucleic acids using modified nucleotide-specific binding reagents can be a function of the number of methylated (or modified) sites per molecule, with molecules with more methylation eluted by increasing salt concentration. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into distinct populations based on the degree of methylation. In one or more implementations, the salt concentration ranges from about 100 nM to about 2500 mM NaCl. In various implementations, the process results in three partitions. Molecules are contacted with a solution containing molecules containing a methyl-binding domain at a first salt concentration, and the molecules may be bound to a capture moiety, such as streptavidin. At the first salt concentration, one population of molecules binds to the MBD and one population remains unbound. The unbound population can be separated as a "hypomethylated" population (low partition). For example, the first partition 126 can represent hypomethylated forms of DNA that remain unbound at low salt concentrations. In one or more instances, the NaCl concentration of the solution used to generate first partition 126 can be about 100 nM, about 120 nM, about 140 nM, about 160 nM, about 180 nM, about 200 nM, or about 250 nM. Second partition 130 can be referred to as the "residual partition" or "intermediate partition" and can represent intermediate methylated DNA eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM. In one or more additional examples, the NaCl concentration of the solution used to generate second partition 130 can be about 100 mM to about 500 mM, about 100 mM to about 1000 mM, about 100 mM to about 1500 mM, about 250 mM to about 1000 mM, about 250 mM to about 1500 mM, about 500 mM to about 1500 mM, about 250 mM to about 2000 mM, about 500 mM to about 2000 mM, or about 1000 mM to about 2000 mM. This can also be separated from the sample. Third partition 134 can represent highly methylated forms of DNA (high partition) and is eluted using a high salt concentration, e.g., at least about 2000 mM.In one or more further examples, the NaCl concentration of the solution used to generate the third distribution 134 can be between about 2000 mM and about 5000 mM, between about 2000 mM and about 4000 mM, between about 2000 mM and about 3500 mM, between about 2000 mM and about 3000 mM, or between about 2500 mM and about 4000 mM.

[0198] In various examples, the first distribution 126 may correspond to a first range of binding strengths of nucleic acids to MBDs and a first range of methylated CG regions, and the second distribution 130 may correspond to a second range of binding strengths of nucleic acids to MBDs and a second range of methylated CG regions. The first range of binding strengths may be less than the second range of binding strengths. In one or more scenarios, a first solution having a first NaCl concentration can separate a first group of nucleic acids having the first range of binding strengths from the MBDs, and a second solution having a second NaCl concentration higher than the first NaCl concentration can separate a second group of nucleic acids having the second range of binding strengths from the MBDs. Furthermore, the third distribution 134 may correspond to a third range of binding strengths and a third range of methylated CG regions. The third range of binding strengths may be greater than the first range of binding strengths and the second range of binding strengths. In one or more cases, a third group of nucleic acids having a third range of binding strengths can be separated from NaCl by a third solution having a third NaCl concentration, which can be higher than the first NaCl concentration and the second NaCl concentration.

[0199] In one or more instances, a plurality of nucleic acids from at least one of a subject's blood or tissue can be combined with a solution containing a quantity of MBD to produce a nucleic acid-MBD solution. A first wash of the nucleic acid-MBD solution can be performed using a first solution containing a first NaCl concentration to produce a first nucleic acid fraction and a first residual solution. The first nucleic acid fraction can include a first portion of the plurality of nucleic acids, and the first residual solution can include a second portion of the plurality of nucleic acids. In one or more instances, the first portion of the plurality of nucleic acids can have a first range of binding strength to the MBD that is less than a second range of binding strength to the MBD of the second portion of the plurality of nucleic acids.

[0200] Further, a second wash of the first residual solution can be performed with a second solution containing a second NaCl concentration higher than the first NaCl concentration to produce a second nucleic acid fraction and a second residual solution. The second nucleic acid fraction can include a first subset of a second portion of the plurality of nucleic acids, and the second residual solution can include a second subset of a second portion of the plurality of nucleic acids. The first subset of a second portion of the plurality of nucleic acids can have a third range of binding strengths to the MBD that is less than the fourth range of binding strengths to the MBD of the second subset of a second portion of the plurality of nucleic acids. Further, a third wash of the second residual solution can be performed with a third solution containing a third NaCl concentration higher than the second NaCl concentration to produce a third nucleic acid fraction containing a second subset of a second portion of the plurality of nucleic acids.

[0201] After the first wash, the second wash, and the third wash, a determination can be made that a first portion of the plurality of nucleic acids is associated with the first distribution 126. Molecular barcodes from the first set of molecular barcodes indicative of the first distribution 126 can be bound to the first portion of the plurality of nucleic acids. In this manner, a sequencing read corresponding to the first distribution 126 can be identified based on determining that the sequencing read includes the first molecular barcode. Further, a determination can be made that a first subset of a second portion of the plurality of nucleic acids is associated with an additional distribution of the plurality of distributions. In these circumstances, a second set of molecular barcodes different from the first set of molecular barcodes can be bound to the second portion of the plurality of nucleic acids, the second molecular barcodes indicative of the additional distribution. As a result, a sequencing read corresponding to the additional distribution can be identified based on determining that the sequencing read includes one or more molecular barcodes from the second set of molecular barcodes. Further, a determination can be made that a second subset of the second portion of the plurality of nucleic acids is associated with the second distribution 130. A third set of molecular barcodes, different from the first set of molecular barcodes and the second set of molecular barcodes, can then be attached to a second subset of a second portion of the plurality of nucleic acids, where the third set of molecular barcodes indicates a second distribution 130. In these cases, sequencing reads corresponding to the second distribution 130 can be identified based on determining that the sequencing reads include a third molecular barcode from among the third set of molecular barcodes.

[0202] In at least some examples, the first molecular separation process 122 may result in nucleic acids present in at least one of the first distribution 126, the second distribution 130, or the third distribution 134 having an amount of methylation that differs from the amount of methylation of other nucleic acids in the respective distribution. For example, the first distribution 126 may include some nucleic acids having an amount of methylation that corresponds to the amount of methylation of nucleic acids contained in at least one of the second distribution 130 or the third distribution 134. Furthermore, at least one of the second distribution 130 or the third distribution 134 may include nucleic acids having an amount of methylation that corresponds to the amount of methylation of nucleic acids contained in the first distribution 126. The presence of nucleic acids in at least one of the first distribution 126, the second distribution 130, or the third distribution 134 that do not correspond to the amount of methylation of at least a majority of the other nucleic acids contained in the respective distribution may cause data noise when performing computational operations on sequence reads produced from nucleic acids contained in the first distribution 126, the second distribution 130, and the third distribution 134. Data noise can result in inaccuracies with respect to calculations made based on sequence reads derived from nucleic acids contained in first distribution 126, second distribution 130, and third distribution 134.

[0203] A second molecular separation process 136 can be performed after the first molecular separation process 122 to reduce or eliminate data noise associated with nucleic acids present in at least one of the first distribution 126, the second distribution 130, or the third distribution 134 having an amount of methylation that does not match the amount of methylation of at least a majority of the other molecules contained in the respective distribution. The second molecular separation process 136 can be performed on the nucleic acids contained in the first distribution 126, the nucleic acids contained in the second distribution 130, and the nucleic acids contained in the third distribution 134. In one or more examples, the second molecular separation process 136 can include performing digestion of the nucleic acids contained in the first distribution 126 using a methylation-dependent restriction enzyme (MDRE), and the nucleic acids contained in the second distribution 130 and the third distribution 134 can be digested using a methylation-sensitive restriction enzyme (MSRE). Digestion of the nucleic acids contained in the first distribution 126 with an MDRE can result in the separation of nucleic acids contained in the first distribution having an amount of methylation corresponding to the second distribution 130 and the third distribution 134 from nucleic acids having an amount of methylation corresponding to the first distribution. Furthermore, digestion of the nucleic acids contained in the second distribution 130 and the third distribution 134 with an MSRE can result in the separation of nucleic acids having an amount of methylation corresponding to the first distribution 126 from the nucleic acids of the second distribution 130 and the third distribution 134. By removing nucleic acids having an amount of methylation corresponding to the second distribution 130 and the third distribution 134 from the first distribution 126 and by removing nucleic acids having an amount of methylation corresponding to the first distribution 126 from the second distribution 130 and the third distribution 134, an additional group 138 of nucleic acids can be produced. The additional group 138 of nucleic acids may include nucleic acids that correspond to the methylation amounts of the second distribution 130 and the third distribution 134, and have a minimal amount or no nucleic acids with an amount of methylation corresponding to the first distribution 126.For example, less than 50% of the nucleic acids included in the additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134, at least 50% of the nucleic acids included in the additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134, at least 60% of the nucleic acids included in the additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134, at least 70% of the nucleic acids included in the additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134, at least 90% of the nucleic acids included in the additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134, At least 95% of the nucleic acids included in group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134; at least 97% of the nucleic acids included in additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134; at least 99% of the nucleic acids included in additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134; at least 99.5% of the nucleic acids included in additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134; or at least 99.9% of the nucleic acids included in additional group 138 may have an amount of methylation corresponding to the second distribution 130 and the third distribution 134.

[0204] The architecture 100 may include a sequencing machine 140. In one or more examples, the sequencing machine 140 may be any of several sequencing machines capable of performing one or more sequencing operations that amplify nucleic acids present in the sample 104. In various examples, the sequencing machine 140 may perform next-generation sequencing operations. In one or more examples, the sample 104 may include a quantity of at least one bodily fluid extracted from a subject. In one or more additional examples, the sample 104 may include a tissue sample obtained from the subject.

[0205] In one or more examples, prior to sequencing, extracted polynucleotides can be partitioned into two or more partitions based on the binding strength of the polynucleotides to the MBD. Blunt-end ligation can be performed on the partitioned polynucleotides and adapters, and tags (e.g., molecular barcodes) can be added to the partitioned polynucleotides. Tagged polynucleotides in one or more partitions (e.g., high and / or medium partitions) can be treated with one or more methylation-sensitive restriction enzymes (MSREs). In some examples, low partitions can be treated with one or more methylation-dependent restriction enzymes (MDREs). After MSRE and / or MDRE treatment, molecules can be enriched by hybridizing the extracted polynucleotides with probes corresponding to target regions of a reference sequence. The enrichment process can identify thousands, hundreds of thousands, or even millions of polynucleotides corresponding to the on-target regions associated with the probes.

[0206] After and / or before the enrichment process, the molecules can be amplified according to one or more amplification processes. The one or more amplification processes can produce thousands, up to millions, of copies of individual nucleic acid molecules. In one or more instances, a portion of the non-enriched polynucleotides can be amplified, but in some cases not to the same extent as the enriched polynucleotides. The one or more amplification processes can produce amplification products that are subjected to one or more sequencing operations. After performing one or more sequencing operations on the sample 104, sequencing data 142 can be produced by the sequencing machine 140.

[0207] Sequencing data 142 may include an alphanumeric representation of the nucleic acids contained in the amplification product. For example, sequencing data 142 may include, for each nucleic acid in the amplification product, data corresponding to a character string representing each strand of nucleotides corresponding to the individual nucleic acid.

[0208] The sequencing data 142 may be stored in one or more data files. For example, the sequencing data 142 may be stored in a FASTQ file, which includes a text-based sequencing data file format that stores raw sequence data and quality scores. In one or more additional examples, the sequencing data 142 may be stored in a data file according to the binary base call (BCL) sequence file format. In one or more further examples, the sequencing data 142 may be stored in a BAM file. In one or more examples, the sequencing data 142 may include at least about 1 gigabyte (GB), at least about 2 GB, at least about 3 GB, at least about 4 GB, at least about 5 GB, at least about 8 GB, or at least about 10 GB. An individual sequence representation included in the sequencing data 106 may be referred to herein as a "read" or "sequencing read." In various examples, an individual first nucleic acid included in the pool 138 may correspond to multiple sequence representations included in the sequencing data 142 as a result of amplifying the individual first nucleic acid. In one or more additional examples, an individual second nucleic acid included in pool 138 may correspond to a single sequence representation included in sequencing data 142 as a result of the individual second nucleic acid not being amplified.

[0209] 2 is an example architecture 200 for analyzing sequencing data to determine one or more metrics indicative of the presence of a tumor in a subject, according to one or more implementations. The architecture 200 may include one or more sequencing machines 202 that perform one or more sequencing operations on a number of samples 204. The one or more samples 204 may be obtained from a subject 206. In one or more instances, a first portion of the subject 206 may be cancer-free; that is, no tumor is detected in the first portion of the subject 206. Additionally, a tumor may be present in a second portion of the subject 206.

[0210] One or more molecular separation processes 208 can be performed on the sample 204. The one or more separation processes 208 can correspond to separating nucleic acid molecules into several partitions based on the characteristics of the nucleic acid molecules. Examples of characteristics that can be used to partition nucleic acid molecules include multiple different nucleotide modifications, methylation levels, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins that bind to DNA. In one or more instances, a heterogeneous population of nucleic acid molecules can be partitioned into nucleic acid molecules with one or more epigenetic modifications and nucleic acid molecules without one or more epigenetic modifications. Examples of epigenetic modifications include, but are not limited to, the presence or absence of methylation; the level of methylation, hydroxymethylation, and the type of methylation (5' cytosine or 6 methyladenine).

[0211] Prior to the one or more molecular separation processes 208, nucleic acid molecules can be extracted from the sample 204. In one or more implementations, the nucleic acid molecules include cell-free nucleic acids (e.g., cell-free DNA). In various implementations, the sample 204 can be a sample selected from one or more of blood, plasma, serum, urine, feces, saliva samples, combinations thereof, and the like. In one or more additional examples, the sample 204 can include a sample selected from one or more of whole blood, blood fractions, tissue biopsies, pleural fluid, pericardial fluid, cerebrospinal fluid, and peritoneal fluid. In one or more instances, the cell-free nucleic acid molecules can be extracted from the sample 204, where the sample 204 is obtained from a subject 206 known to have cancer (e.g., a cancer patient) or suspected of having cancer.

[0212] Extraction of nucleic acid molecules from sample 204 may include implementing one or more cell lysis techniques to cleave membranes of cells contained in sample 204 and applying one or more proteases to degrade proteins contained in sample 204. Extraction of nucleic acid molecules from sample 204 may also include several washing and / or elution techniques to separate the nucleic acid molecules from other components contained in sample 204. In various examples, thousands, up to millions, or up to billions of nucleic acid molecules may be extracted from sample 204 and then subjected to one or more separation processes 208.

[0213] The nucleic acid molecules extracted from sample 204 may include molecules with various levels of methylation. Methylation may occur due to any one or more post-replication or post-transcriptional modifications. Post-replication modifications include, but are not limited to, modifications of the nucleotide cytosine, including 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine. One or more molecular separation processes 208 may separate the nucleic acid molecules extracted from sample 204 into several partitions, each corresponding to a different level of methylation. For example, molecular separation process 208 may produce a first partition of nucleic acid molecules having a first level of methylation, a second partition of nucleic acid molecules having a second level of methylation, and a third partition of nucleic acid molecules having a third level of methylation. In various examples, the second level of methylation may be higher than the first level of methylation, and the third level of methylation may be higher than both the first and second levels of methylation. In one or more instances, the one or more molecular separation processes 208 may include the first molecular separation process 122 and the second molecular separation process 136 of FIG.

[0214] The one or more molecular separation processes 208 can result in pools 210 that include a portion of the nucleic acid molecules extracted from the one or more samples 204 and are subjected to the one or more molecular separation processes 208. For example, pool 210 can include some nucleic acid molecules having a second level of methylation and some nucleic acid molecules having a third level of methylation. Thus, the nucleic acid molecules included in pool 210 can have at least a threshold amount of methylation. In one or more instances, the nucleic acid molecules included in pool 210 can have at least a threshold amount of methylation within CG regions of the nucleic acid molecules.

[0215] One or more sequencing operations can be performed by one or more sequencing machines 202 to produce sequencing data 212 corresponding to pool 210. Architecture 200 can include a computing system 214 that obtains sequencing data 212 from one or more sequencing machines 202 and analyzes the sequencing data 212. For example, computing system 214 can analyze sequencing data 212 to determine one or more metrics indicative of a possible tumor presence in a subject 206 that provided at least one sample 204. Computing system 214 can include one or more computing devices 216. One or more computing devices 216 can include at least one of one or more desktop computing devices, one or more mobile computing devices, or one or more server computing devices. In various examples, at least a portion of one or more computing devices 216 can be included in a remote computing environment, such as a cloud computing environment. In one or more examples, computing system 214 and sequencing machine 202 can be owned, operated, maintained, and / or controlled by a single organization. In one or more additional examples, the computing system 214 and the sequencing machine 202 may be owned, operated, maintained, and / or controlled by multiple organizations.

[0216] In operation 218, the computing system 214 can analyze the sequencing data 212. Analyzing the sequencing data 212 can include determining one or more first sequence representations 220 included in the sequencing data 212 that correspond to one or more classification regions of the reference sequence. The one or more classification regions can correspond to genomic regions of the reference sequence that map to nucleic acid molecules that have a certain amount of methylation in cfDNA obtained from a subject with cancer, relative to a certain amount of methylation in molecules that map to the same genomic region of the reference sequence in cfDNA obtained from a subject without a tumor. In at least some examples, the amount of methylation present in nucleic acid molecules that map to the classification region and are derived from a subject with cancer is less than the amount of methylation present in nucleic acid molecules that map to the classification region and are derived from a subject without cancer. In one or more additional examples, the amount of methylation present in nucleic acid molecules that map to the classification region and are derived from a subject with cancer is greater than the amount of methylation present in nucleic acid molecules that map to the classification region and are derived from a subject without cancer. The one or more classification regions can also include at least a threshold amount of cytosine-guanine content. In various examples, one or more classification regions may include a series of cytosine-guanine (CG) pairs (CpG sites) in the 5'→3' direction, for example, at least 3 CpG sites, at least 5 CpG sites, at least 8 CpG sites, at least 10 CpG sites, at least 12 CpG sites, at least 15 CpG sites, at least 18 CpG sites, or at least 20 CpG sites.

[0217] Additionally, the computing system 214 can analyze the sequencing data 212 to determine one or more second sequence representations 222 corresponding to one or more control regions of the reference sequence. The one or more control regions can include one or more positive control regions and / or one or more negative control regions. In various examples, the positive control region can include a genomic region of the reference sequence having molecules with at least a threshold amount of methylated cytosines and containing at least a threshold number of CpG sites. The positive control region can correspond to nucleic acid molecules having at least a threshold amount of methylation within one or more CpG regions in samples obtained from subjects with cancer and from subjects without tumors. In at least some examples, the threshold amount of methylation can correspond to at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or more CpGs being methylated in the nucleic acid molecule. In one or more embodiments, the positive control region may be hypermethylated in one or more CG regions and mapped to nucleic acid molecules derived from samples from both subjects with and without cancer. In one or more embodiments, the negative control region may include molecules with less than a threshold amount of methylated cytosines and a genomic region of the reference sequence with at least a threshold number of CpG sites. The negative control region may correspond to nucleic acid molecules with less than an additional threshold amount of methylation in one or more CG regions in samples from subjects with cancer and from subjects without tumors. In various examples, the additional threshold amount of methylation may correspond to 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, or 7 or less CpGs being methylated in the nucleic acid molecule. In one or more additional embodiments, the negative control region may be hypomethylated in one or more CG regions and mapped to nucleic acid molecules derived from samples from both subjects with and without cancer.

[0218] In one or more instances, the first sequence representation 220 can be determined by aligning the sequence representation included in the sequencing data 212 with one or more classification regions of the reference sequence. Additionally, the second sequence representation 222 can be determined by aligning the sequence representation included in the sequencing data 212 with one or more control regions of the reference sequence. In the alignment process, the first sequence representation 220 can be identified by determining the number of sequence representations included in the sequencing data 212 that correspond to one or more classification regions of the reference sequence. Additionally, in the alignment process, the second sequence representation 222 can be identified by determining the number of sequence representations that correspond to one or more control regions of the reference sequence.

[0219] In one or more instances, the alignment process can determine the amount of homology between each sequence representation included in sequence data 212 and a portion of a reference sequence. The amount of homology between a given sequence representation and a reference sequence can indicate the number of positions in the reference sequence that have the same nucleotide as the corresponding position in the given sequence representation. The computing system 214 can determine that a sequence representation is aligned with a portion of the reference sequence based on determining that the sequence representation and the portion of the reference sequence have at least a threshold amount of homology. In scenarios where a sequence representation has at least a threshold amount of homology to multiple portions of the reference sequence, the portion of the reference sequence that has the greatest amount of homology with the sequence representation can be determined to be aligned with the sequence representation.

[0220] The amount of homology between a given sequence representation and a portion of a reference sequence can be determined using the BLAST program (basic local alignment search tool) and PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), or by using the Gap program (Wisconsin Sequence Analysis Package, Genetics Computer Group, University Research Park, Madison Wis.) with default settings, which uses the Needleman and Wunsch algorithm (J. Mol. Biol. 48; 443-453 (1970)). The amount of homology between a sequence representation and a portion of a reference sequence can also be determined using the Burrows-Wheeler aligner (Li, H., & Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics, 25(14), 1754-1760).

[0221] In one or more examples, after aligning the sequence representations included in the sequencing data 212 with a reference sequence, the aligned sequence representations can be analyzed to identify one or more groups of sequence representations. For example, each aligned sequence representation can correspond to an individual sequencing read included in the sequencing data 212. In these scenarios, the aligned sequence representations can include multiple reads corresponding to a single nucleic acid molecule included in the sample pool 210. In one or more additional examples, the aligned sequence representations can correspond to individual nucleic acid molecules included in the pool 210. In these situations, the computing system can determine groups of reads included in the sequence data 212 that correspond to individual nucleic acid molecules included in the pool 210 based on a molecular barcode common to each group of sequencing reads. That is, each nucleic acid molecule included in the pool 210 can be encoded with a molecular barcode that uniquely identifies the individual nucleic acid molecule, and in at least some cases, the individual nucleic acid molecule can be represented by multiple sequencing reads included in the sequencing data 212. Thus, if multiple sequence representations corresponding to a single nucleic acid molecule included in the pool 210 are present in the sequencing data 212, the computing system 214 can group the multiple sequence representations together. In various examples, a group of sequence representations corresponding to a single nucleic acid molecule included in pool 210 may be referred to herein as a "family." Furthermore, the start and end positions relative to a reference sequence of aligned sequence representations having a common molecular barcode may be used to group the sequence representations corresponding to the individual nucleic acids included in pool 210. In one or more instances, an individual sequence representation representing a family of sequence representations corresponding to a single nucleic acid molecule included in pool 210 may be referred to herein as a "consensus sequence representation."

[0222] In operation 224, the computing system 214 can analyze the first and second sequence representations 222 to generate metrics corresponding to individual classification regions. In the example of FIG. 2, the computing system 214 can analyze the first and second sequence representations 220 and 222 to generate classification region metrics 226. The classification region metrics 226 can include a quantitative measure determined based on the number of first sequence representations 220 having at least a threshold amount of methylated cytosines. In one or more examples, the classification region metrics 226 can include a quantitative measure determined based on the number of sequencing reads corresponding to the number of first sequence representations 220 having at least a threshold amount of located methylated cytosines. In one or more additional examples, the classification region metrics 226 can include a quantitative measure determined based on the number of nucleic acid molecules corresponding to the number of first sequence representations 220. In various examples, classification region metric 226 may include a quantitative measure determined based on the number of first sequence representations 220 having at least a threshold amount of methylated cytosines and the number of second sequence representations 222 corresponding to control regions of the reference sequence. In one or more further examples, classification region metric 226 may include a quantitative measure related to the ratio of the number of first sequence representations 220 to the number of second sequence representations having at least a threshold amount of methylated cytosines. In at least some examples, the sequence representation of second sequence representation 222 used by computing system 214 to generate the quantitative measure included in classification region metric 226 may include a sequence representation corresponding to a positive control region of the reference sequence.

[0223] The classification region metrics 226 may also be determined by performing one or more normalization operations on the quantitative measures generated by the computing system 214 using at least one of the first sequence representation 220 and the second sequence representation 222. For example, the computing system 214 may perform a logarithmic calculation on the quantitative measures generated using at least one of the first sequence representation 220 or the second sequence representation 222. Additionally, the classification region metrics 226 may be determined by adding a pseudocount to the quantitative measures determined by the computing system 214 using at least one of the first sequence representation 220 or the second sequence representation 222. In one or more instances, the one or more normalization operations may include determining a quantitative measure corresponding to the ratio of the first sequence representation 220 to the number of second sequence representations 222 corresponding to a positive control region of the reference sequence for each classification region.

[0224] In one or more instances, the computing system 214 can determine the number of first sequence representations 220 that correspond to individual classification regions of the reference sequence and have at least a threshold amount of methylated cytosines located within the individual classification regions. In these scenarios, the computing system 214 can determine individual classification region metrics 226 for the individual classification regions. Additionally, the computing system 214 can determine the number of second sequence representations 222 that correspond to the positive control regions. In at least some instances, the computing system 214 can determine, for the individual classification regions, a ratio including a number of first sequence representations 220 that correspond to the individual classification regions and have molecules with at least a threshold amount of methylated cytosines within the classification region relative to the total number of second sequence representations 222 that correspond to the positive control regions of the reference sequence. In one or more instances, the computing system 214 can add a pseudocount value to the ratio to determine the classification region metrics 226 for the individual classification regions. The value of the pseudo count may be at least 1, at least 1.2, at least 1.4, at least 1.6, at least 1.8, or at least 2. Additionally, the computing system 214 may perform a base 1 logarithm operation on the combination of the ratio and pseudo count to determine classification region metrics 226 for individual classification regions. In at least some instances, the computing system 214 may calculate at least a portion of the classification region metrics according to the following equation:

number

[0225] In operation 228, the computing system 214 may execute the model to determine an indicator of cancer based on the classification region metrics 226. In the example of FIG. 2, the computing system 214 may execute the model using the classification region metrics 226 to generate model output 230. In one or more examples, the model output 230 may indicate a tumor detected 232 status or a tumor not detected 234 status for the sample 204 provided by the subject 206. In one or more additional examples, the computing system 214 may execute the model to determine an estimate of a tumor fraction 236 for the sample 204. In one or more further examples, the computing system 214 may execute the model to determine the probability of the presence of a tumor in the subject 206 who provided the sample 204.

[0226] In one or more examples, the model may include a classification model that implements one or more machine learning techniques. In one or more instances, the model may include a linear regression model. In various examples, the model may be executed to determine a probability 238 of the presence of a tumor in the subject 206 that provided the sample 204 based on the classification region metrics 226. In one or more instances, the computing system 214 may execute the model to determine weights for individual classification regions. The weights for individual classification regions may be different. For example, the computing system 214 may determine that a first weight of a first classification region metric 226 for a first classification region is different from a second weight of a second classification region metric 226 for a second classification region. In at least some instances, the computing system 214 may calculate the probability 238 of the presence of a tumor in the subject 206 that provided the sample 204 using the following equation:

number

[0227] In one or more additional examples, the computing system 214 can run a model to determine a maximum variant allele fraction (MAF). In various examples, the computing system 214 can run a model using the maximum MAF value to determine a tumor fraction 236 for the sample 204. In one or more instances, the computing system 214 can run a model using the classification region metrics 226 to determine a logit-transformed maximum MAF value, which can then be used by the computing system 214 to estimate the tumor fraction for the sample 204. In various examples, the computing system 214 can analyze the maximum MAF value to determine a probability 238 of the presence of cancer in the subject 206 who provided the sample 204. In various examples, a Huber regression (Huber, P. J. 1964. "Robust Estimation of a Location Parameter." Annals of Mathematical Statistics 35 (1): 73-101) can be performed to determine the maximum MAF value based on the classification region metrics 226.

[0228] In various examples, the model output 230 may also include a tumor tissue index 240. The tumor tissue index 240 may indicate one or more tissues of origin of the cancer cells that produced the genomic material detected in the sample 204. In one or more examples, the tumor tissue index 240 may correspond to one or more tissues of origin of the cancer cells that produced the genomic material detected in the sample 204. In these scenarios, the computing system 214 can generate multiple models with individual models corresponding to a given tissue type. The outputs from the individual models can be analyzed to determine additional metrics indicative of the tissues of origin of the cancer cells that produced the genomic material detected in the one or more samples. In at least some examples, the output for the individual models can indicate at least one of a tumor fraction 236 or a probability 238 that a tumor is present. The computing system 214 can analyze each model output to determine the model with at least one of the greatest tumor fraction or the greatest probability that cancer is present. The computing system 214 can then generate a tumor tissue index 240 corresponding to the model with the greatest tumor fraction and / or the greatest probability that cancer is present.

[0229] For example, samples 204 can be obtained from subjects 206 with different types of cancer. By way of example, a first sample can be obtained from a first group of subjects with a first class of cancer, and a second sample can be obtained from a second group of subjects with a second class of cancer. The sequencing data generated from the first sample can be analyzed by the computing system 214 to generate a first metric corresponding to a classification region for the first class of cancer, and the first metric can be used to generate a first model corresponding to the first class of cancer. Furthermore, the sequencing data generated from the second sample can be analyzed by the computing system 214 to generate a second metric corresponding to the classification region for the first class of cancer, and the second metric can be used to generate a second model corresponding to the second class of cancer. After training the model, the computing system 214 can analyze sequencing data obtained from one or more additional subjects not included in the training subjects to determine classification region metrics for the one or more additional subjects. The classification region metrics can then be analyzed using different tumor classification models to generate model outputs. The model outputs can be analyzed by a computing system to determine the model with the highest value for each model output and determine the tumor tissue classification corresponding to the model.

[0230] In one or more additional exemplary implementations, the model output 230 may also indicate the methylation state of one or more genomic regions of the reference sequence. For example, the computing system 214 may analyze the classification region metrics 226 to determine the methylation state of one or more promoter regions of the reference sequence. In various examples, the one or more promoter regions may include at least one promoter region associated with the presence of a tumor in a subject. In one or more instances, the classification region metrics 226 may indicate a number of sequence representations having at least a threshold amount of methylation for one or more promoter regions. In these scenarios, the computing system 214 may determine that the promoter region is methylated by determining that the number of sequence representations having molecules with at least a threshold amount of methylated cytosines in the promoter region is greater than a threshold number.

[0231] In yet another exemplary implementation, the computing system 214 can analyze the sequencing data 212 to determine a quantitative measure corresponding to the number of sequence representations corresponding to the promoter region and an additional quantitative measure corresponding to the number of additional sequence representations corresponding to the number of positive control regions. A normalized metric can be determined based on the quantitative measure and the additional quantitative measure. The normalized metric can be analyzed with respect to a threshold to determine the methylation status of the promoter region. In various examples, the threshold can be different for different promoter regions. The threshold can be determined using a cancer-free dataset, and the threshold is set to achieve a false positive rate of 5% or less, 4% or less, 3% or less, 2% or less, or 1% or less and a specificity of at least 80%, at least 90%, at least 95%, at least 98%, or at least 99%. In at least some examples, the computing system 214 can analyze the promoter region for a threshold where at least four sequence representations correspond to the promoter region, at least five sequence representations correspond to the promoter region, at least six sequence representations correspond to the promoter region, at least seven sequence representations correspond to the promoter region, at least eight sequence representations correspond to the promoter region, at least nine sequence representations correspond to the promoter region, at least ten sequence representations correspond to the promoter region, at least eleven sequence representations correspond to the promoter region, or at least twelve sequence representations correspond to the promoter region.

[0232] In one or more further examples, the computing system 214 can combine results from multiple models to determine a model output 230. For example, the computing system 214 can run a model on one or more epigenetic signals, such as methylation of a classification region, to determine one or more first tumor metrics. With respect to methylation, the computing system 214 can run both a classification model, such as a logistic regression model, that generates an indication that cancer is present in the subject providing the sample, and an additional model that predicts the tumor fraction for the sample. In one or more examples, the epigenetic signal can also correspond to the fragment length of a sequence representation generated from the sample. Furthermore, the computing system 214 can run one or more additional models on the genomic signal to generate further tumor metrics for the sample. In various examples, the genomic signal can correspond to the presence of one or more single nucleotide variants (SNVs) and / or the presence of insertions or deletions in one or more genomic regions of a reference sequence. In at least some examples, the computing system may include an integrated system that can combine tumor metrics and epigenetic signals generated by running several models on data corresponding to genomic signals to produce aggregate tumor metrics for a given sample. In some embodiments, the quantitative measure obtained from the model output 230 can be used to analyze with respect to a threshold value. In situations where the quantitative measure obtained from the model output is at least the threshold value, the computing system 214 can determine an indication that cancer is present in the subject. In situations where the quantitative measure obtained from the model output is less than the threshold value, the computing system 214 can determine an indication that cancer is not present or is not detected in the subject. In some embodiments, the threshold value used to determine the indication of whether cancer is present is calculated using a set of normal samples and set to a particular value that results in high specificity.

[0233] In various additional implementations, the computing system 214 may determine the methylation state of individual genomic regions. In one or more instances, the computing system 214 may determine the methylation state of one or more promoter regions. In one or more instances, the sequencing data 212 may be analyzed to determine sequence representations corresponding to one or more genomic regions. For example, the sequencing data 212 may be analyzed to determine the number of sequence representations corresponding to one or more promoter regions. In at least some instances, the computing system 214 may determine the number of sequence representations corresponding to individual promoter regions having at least a threshold amount of methylated cytosines.

[0234] For each genomic region and for each individual sample, the computing system 214 can determine the number of sequence representations corresponding to polynucleotide molecules having at least a threshold number of methylated cytosines in the genomic region. The computing system 214 can perform one or more normalization operations using the count of polynucleotide molecules or sequence reads that correspond to the genomic region and have at least a threshold number of methylated cytosines to generate normalized metrics. As an example, the computing system 214 can divide the count of polynucleotide molecules or reads that correspond to the genomic region and have at least a threshold number of methylated cytosines by the number of molecules or sequencing reads that correspond to a control region, e.g., a positive control region. In another case, the computing system 214 can perform a normalized metric by dividing the count of polynucleotide molecules or reads that correspond to the genomic region and have at least a threshold number of methylated cytosines by the number of molecules or sequencing reads in a control dataset (i.e., a control dataset composed of samples in which no tumor is detected) that correspond to the same genomic region and have at least the same threshold number of methylated cytosines.

[0235] The normalized metric can be analyzed with respect to a threshold value. The threshold value can correspond to a given genomic region, such as a given promoter region. In various examples, the threshold value can be different for different promoter regions. In these scenarios, a first promoter region can have a first threshold value, and a second promoter region can have a second threshold value. In situations where the normalized metric is at least the threshold value, the computing system 214 can determine that the genomic region has a first methylation state. In scenarios where the normalized metric is less than the threshold value, the computing system 214 can determine that the genomic region has a second methylation state. In one or more instances, the first methylation state can be labeled "methylated" and the second methylation state can be labeled "unmethylated."

[0236] The threshold value for a given genomic region can be determined based on training data obtained from samples of individuals who are not detected with cancer.In one or more examples, the sequence representation obtained from the training sample can be analyzed to determine a z-score for the number of polynucleotide molecules that correspond to the genomic region and have at least a threshold amount of methylated cytosine.In one or more examples, the threshold value for the promoter region used to determine the normalization metric for the promoter region can be obtained from the z-score calculated based on the training sample for the promoter region.

[0237] While the example in FIG. 2 describes generating a model to determine a number of indicators for the presence or absence of cancer in a given subject, in at least some additional examples, the sequencing data 212 can be analyzed by the computing system 214 to determine an indicator of the presence of cancer without training a specific model. In one or more examples, the computing system 214 can determine a tumor fraction value based on sequencing data 212 generated from one or more samples obtained from a single subject, where it is unknown whether the subject has cancer. In one or more examples, the computing system 214 can determine a change in the tumor fraction value based on sequencing data 212 generated from one or more samples obtained from a single subject at two or more time points. The change in the tumor fraction value can be used to monitor the subject's response to treatment. In one or more additional examples, a first sample can be obtained from the subject before or at the start of at least one administration of a cancer-related treatment or procedure, and one or more second samples can be obtained from the subject after at least one administration of a cancer-related treatment or procedure. In one or more instances, the one or more second samples may be obtained at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 8 weeks, or at least 10 weeks after administration of the treatment or procedure. In at least some instances, the first sample and the second sample may be derived from at least one of a bodily fluid obtained from the subject or a tissue obtained from the subject.

[0238] In one or more examples, one or more samples can be obtained from a given subject. Sequencing data 212 generated from one or more samples can be analyzed by a computing system to determine quantitative measures for several classification regions. In one or more examples, the quantitative measure can correspond to the amount of sequence representations having at least a threshold amount of overlap with one or more classification regions. In one or more additional examples, the quantitative measure can correspond to sequence representations having at least a threshold amount of methylated cytosines within CpG regions having at least a threshold amount of CG content. In various examples, the indication that cancer is present in a subject can include a tumor fraction. In one or more additional examples, the indication that cancer is present in a subject can include a variant allele fraction. In at least some examples, the quantitative measure can correspond to the number of sequencing reads corresponding to a given classification region in relation to the total number of sequencing reads across multiple positive control regions. In one or more further examples, the indication that cancer is present can be used to determine an output corresponding to the presence or absence of cancer in a given individual by analyzing one or more cancer presence indicators with respect to one or more thresholds. In one or more instances, the tumor fraction determined from one or more samples obtained from the subject can be analyzed with respect to one or more thresholds. If the tumor fraction is higher than the threshold level, the computing system 214 can determine that the probability that the subject has cancer is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%. Furthermore, in situations where multiple samples are obtained from the subject, a first quantitative measure generated from a first sample obtained from the subject can be analyzed with respect to a second quantitative measure generated from a second sample obtained from the subject. In at least some instances, the difference between the first and second quantitative measures can be analyzed to determine an indication of treatment response in the subject.

[0239] The quantitative scale used to determine the indicator that cancer exists in a subject can be determined by analyzing the quantitative scale of a subset of classification regions.In at least some examples, the subset of classification regions can be different for different subjects.In one or more embodiments, the quantitative scale values ​​for several classification regions can be analyzed against each other and ranked according to the magnitude of the value of the quantitative scale.In various examples, the classification regions for a given sample can be ranked in descending order from the one or more classification regions with the largest value of quantitative scale to the one or more classification regions with the smallest value of quantitative scale.

[0240] In various examples, after the quantitative scales of classification regions are ranked, quantitative scales corresponding to the group of classification regions can be removed before determining the indicator that the subject has cancer.For example, the group of classification regions that are not used to determine the indicator that the subject has cancer can include 1% of the classification regions with the highest quantitative scale value, 2% of the classification regions with the highest quantitative scale value, 3% of the classification regions with the highest quantitative scale value, 4% of the classification regions with the highest quantitative scale value, 5% of the classification regions with the highest quantitative scale value, or 6% of the classification regions with the highest quantitative scale value.In at least some examples, some classification regions with relatively high quantitative scale values ​​can be excluded from the group of classification regions used to determine the indicator that the subject has cancer, because at least in some cases, the classification regions corresponding to the quantitative scale values ​​at or near the top of the ranked list may have non-tumor origin and / or be related to sequencing artifacts.Therefore, by removing the quantitative scales corresponding to these classification regions from the analysis used to determine the indicator that the subject has cancer, the accuracy of the indicator that the subject has cancer can be increased.

[0241] In one or more examples, after determining the group of classification regions used to determine an indication that cancer is present in a subject, a subset of the classification regions of the group can then be determined by identifying at least 10 classification regions of the group, at least 25 classification regions of the group, at least 50 classification regions of the group, at least 75 classification regions of the group, at least 100 classification regions of the group, at least 150 classification regions of the group, at least 200 classification regions of the group, at least 250 classification regions of the group, at least 300 classification regions of the group, at least 350 classification regions of the group, at least 400 classification regions of the group, at least 450 classification regions of the group, or at least 500 classification regions of the group that have the maximum value of each quantitative measure.

[0242] In at least some examples, one or more statistical measures, such as at least one of the mean, median, or mode, can be applied to a subset of the quantitative measures of the classification regions of the group to generate an initial indicator of the presence of cancer in the subject. In various examples, the initial cancer indicator can be modified according to a scaling factor. A scaling factor can be applied to the initial indicator of the presence of cancer in the subject because, at least in some scenarios, the positive control regions may have different amounts of methylated CpGs. For example, at least a portion of the positive control regions may have fully methylated CpGs, while other positive control regions may not be fully methylated. Furthermore, in various situations, some classification regions may correspond to high values ​​of the indicator of the presence of cancer in the subject, such as 90% tumor fraction, 95% tumor fraction, 99% tumor fraction, or 100% tumor fraction, but the nucleic acid molecules corresponding to these classification regions may not be fully methylated. To account for these cases, a scaling factor can be applied to the initial indicator of the presence of cancer in the subject to provide a more accurate determination of the indicator. In one or more instances, the scaling factor can be determined by analyzing the determined indicator of the presence of cancer in a subject using one or more techniques described herein for additional data corresponding to the additional indicator of the presence of cancer in a subject, e.g., validation data, or other techniques that generate data orthogonal to the indicator of the presence of a tumor in a subject described herein.

[0243] In various examples, the classification region used to determine the quantitative measure may correspond to a classification region corresponding to one or more portions of a differentially methylated region. In one or more examples, the differentially methylated region may include a promoter region corresponding to one or more classifications of cancer. For example, the classification region may be determined by analyzing several sequencing representations across the differentially methylated region. In these scenarios, one or more portions of the differentially methylated region that overlap with at least a threshold number of sequencing representations may be included in the classification region. In one or more examples, the quantitative measure of one or more portions of the differentially methylated region may be determined based on the molecular number distribution of the differentially methylated region. For example, the quantitative measure may be determined based on the number of molecules within one or more peaks of the molecular distribution of the differentially methylated region. For example, in various examples, the distribution of molecules across the differentially methylated region may show one or more peaks with a higher amount of molecules overlapping with one or more subregions within the differentially methylated region. In various examples, one or more genomic regions corresponding to one or more subregions of the differentially methylated region that correspond to the greatest amount of sequence representation for a sample may be defined as the classification region. In at least some cases, the distribution of sequence representations may have a peak corresponding to a subregion of the differentially methylated region that has a higher number of sequence representations than other subregions of the differentially methylated region.In these scenarios, the subregion can be identified as a classification region.By determining the subregion of at least a portion of the differentially methylated region that is used to determine the indicator that a subject has cancer, the amount of computing resources and memory resources used to determine the indicator that a subject has cancer can be reduced.

[0244] For example, a classification region may include one or more portions of a differentially methylated region in which at least 50% of the sequencing representations are obtained from sample overlaps, at least 55% of the sequencing representations are obtained from sample overlaps, at least 60% of the sequencing representations are obtained from sample overlaps, at least 65% of the sequencing representations are obtained from sample overlaps, at least 70% of the sequencing representations are obtained from sample overlaps, at least 75% of the sequencing representations are obtained from sample overlaps, at least 80% of the sequencing representations are obtained from sample overlaps, at least 85% of the sequencing representations are obtained from sample overlaps, at least 90% of the sequencing representations are obtained from sample overlaps, at least 95% of the sequencing representations are obtained from sample overlaps, or at least 99% of the sequencing representations are obtained from sample overlaps. In one or more instances, one or more portions of a differentially methylated region comprising a classification region may be contiguous with respect to a reference sequence.

[0245] FIG. 3 is a diagrammatic representation of an example framework 300 for training a computational model 302 to determine one or more tumor metrics for a sample, according to one or more implementations. The framework 300 may include a computing system 214. The computing system 214 may execute the computational model 302 to generate one or more model outputs 304. In one or more examples, the computational model 302 may be a machine learning model. The model output 304 may include an indicator corresponding to the presence or absence of a tumor in the subject who provided the sample. In one or more instances, the model output 304 may include a tumor fraction. In one or more additional instances, the model output 304 may include a probability that the subject has cancer. In one or more further instances, the model output may include an indicator that the subject has cancer or that the subject does not have cancer. In yet other instances, the model output 304 may indicate the methylation state of one or more regions of a nucleic acid molecule. By way of example, the computing system 214 may execute the computational model 302 on a quantitative measure corresponding to a promoter region to determine the amount of methylation of the promoter region. In other instances, the model output 304 may include a tumor tissue index for the sample.

[0246] The framework 300 may also include a sequence representation 306. In one or more examples, the sequence representation 306 can be generated based on analyzing nucleic acid molecules derived from a sample provided by the subject. The sequence representation 306 can include genomic regions having a number of nucleotides corresponding to the number of regions of interest. For example, the sequence representation 306 can include a sequence of nucleotides corresponding to a first classification region 308. Further, the sequence representation 306 can include a sequence of nucleotides corresponding to a second classification region 310. Further, the sequence representation 306 can include a sequence of nucleotides corresponding to a third classification region 312. In various examples, the first classification region 308, the second classification region 310, and the third classification region 312 of the sequence representation 306 can have different amounts of methylated cytosines contained in the respective classification regions 308, 310, 312. In one or more additional examples, the sequence representation 306 can include a sequence of nucleotides corresponding to a positive control region 314 and a sequence of nucleotides corresponding to a negative control region 316.

[0247] The computational model 302 may include several components corresponding to individual classification regions. In one or more examples, the components of the computational model 302 may have respective values ​​corresponding to quantitative metrics for the respective classification regions. The quantitative metrics may indicate the number of sequence representations corresponding to the respective classification regions. In one or more examples, the computational model 302 may include several weights associated with each component of the computational model 302. For example, the computational model 302 may include a first model component 318 having a first weight 320. The first model component 318 may correspond to the first classification region 308. Furthermore, the computational model 302 may include a second model component 322 having a second weight 324. The second model component 322 may correspond to the second classification region 310. In various examples, at least one of the first weight 320, the second weight 324, or the third weight 328 may be different from at least another one of the first weight 320, the second weight 324, or the third weight 328.

[0248] In one or more instances, values ​​for the first model component 318, the second model component 322, and the third model component 326 can be determined for each sample. By way of example, for different samples, the computational model 302 can determine different values ​​for at least one of the first model component 318, the second model component 322, or the third model component 326. In various examples, the computing system 214 can determine a first quantitative measure for the first classification region 308 based on sequencing data for the sample. The computing system 214 can execute the computational model 302 to determine a value for the first model component 318 based on the first quantitative measure. Additionally, the computing system 214 can determine a second quantitative measure for the second classification region 310 based on the sequencing data for the sample. The computing system 214 can execute the computational model 302 to determine a value for the second model component 322 based on the second quantitative measure. Additionally, the computing system 214 can determine a third quantitative measure for the third classification region 312 based on the sequencing data for the sample. The computing system 214 can execute the computational model 302 to determine a value for the third model component 326 based on the third quantitative measure. The first quantitative measure, the second quantitative measure, and the third quantitative measure can be determined based on the number of sequence representations having at least a threshold amount of methylation in the CG regions corresponding to the first classification region 308, the second classification region 310, and the third classification region 312, respectively. In one or more additional instances, a value for the first weight 320, a value for the second weight 324, and a value for the third weight 328 can be determined for each sample. For example, for different samples, the computational model 302 can determine different values ​​for at least one of the first weight 320, the second weight 324, or the third weight 328.

[0249] In one or more examples, the computing system 214 may perform a training process to generate the computational model 302. In various examples, the training process may determine one or more features associated with classification domain metrics that may be used to determine the model output 304. Additionally, the training process may determine one or more parameters associated with the classification domain metrics that may be used to determine the model output 304. For example, the training process may be used to determine the model components and corresponding weights for the model components to be included in the computational model 302.

[0250] In the example of FIG. 3 , the training process can be performed using training data 330. The training data 330 can include information obtained for at least a first group of subjects 332 and information obtained for at least a second group of subjects 334. In one or more examples, the first group of subjects 332 can include subjects in which no tumor is detected, and the second group of subjects 334 can include subjects in which a tumor is detected. In various examples, the training data 330 can include features related to the amount of methylation of the classification regions of the reference sequence 306 for the first group of subjects 332 and the second group of subjects 334. For example, the training data 330 can indicate a quantitative measure corresponding to the number of sequence representations having at least a threshold level of methylation for the classification regions 308, 310, 312 for the first group of subjects 332 and the second group of subjects 334. The training data 330 can also include weights for the model components based on an analysis of the sequencing data for the first group of subjects 332 and the second group of subjects 334. In one or more instances, training data 330 may include values ​​for first weights 320, second weights 324, and third weights 328 based on classification region metrics determined from sequencing data obtained from samples provided by a first group of subjects 332 and a second group of subjects 334.

[0251] The training data 330 may also include information corresponding to additional characteristics of the first group of subjects 332 and the second group of subjects 334. By way of example, the training data 330 may include medical record information, medical history information, cancer treatment history information, demographic information, genomics information, one or more combinations thereof, and the like.

[0252] In one or more examples, the computing system 214 can train the computational model 302 to determine an indicator associated with one or more types of cancer present in an individual. Furthermore, in various examples, the computational model 302 can include multiple different models, and thus the computational model 302 is an ensemble model. In these situations, the computing system 214 can perform one or more training processes on the individual models of the ensemble model. In one or more instances, the computational model 302 can include several individual models, each corresponding to determining a model output for an individual genomic region, such as a gene or a particular group of genes. For example, the computational model 302 can include several individual models for generating a maximum MAF value for an individual gene or a particular group of genes. In these scenarios, the training of the computational model 302 can have greater constraints than other models used to determine an indicator that an individual has cancer due to the use of genomic information in the training process. As a result, in situations where genomic information is used to train the computational model 302, the accuracy of the output of the computational model 302 can be increased. Additionally, while training of the computational model 302 can incorporate genomic information to determine maximum MAF values, use of the computational model 302 to determine a cancer indicator for a test subject after training can be performed without the use of genomic information and can be based on input vectors corresponding to quantitative measures determined from sequencing data obtained from the test subject.

[0253] Computing system 214 may obtain first through Nth training data sets 336 through 338 to perform a training process to generate computational model 302. In one or more examples, first training data set 336 may include a first portion of training data 330 corresponding to a first group of subjects 332 and a second group of subjects 334 used to train computational model 302, and Nth training data set 338 may include a second portion of training data 330 corresponding to the first group of subjects 332 and the second group of subjects 334 as part of a validation process for computational model 302. In various examples, computational model 302 may be updated over time and undergo multiple training processes. In these scenarios, the first training data set 336 may include a portion of the training data 330 for the first group of subjects 332 and the second group of subjects 334 corresponding to a first time period, and the Nth training data set 338 may include a portion of the training data 330 for the first group of subjects 332 and the second group of subjects 334 corresponding to a second time period.

[0254] In one or more examples, one or more optimization operations can be performed by computing system 214 during a training process for computational model 302. In one or more instances, computing system 214 can identify, during a training process for computational model 302, one or more samples obtained from at least one of first group of subjects 332 or second group of subjects 334 that are outliers relative to samples obtained from other subjects in at least one of first group of subjects 332 or second group of subjects 334. By way of example, computing system 214 can determine that model outputs 304 generated for one or more subjects in at least one of first group of subjects 332 or second group of subjects 334 differ by at least a threshold amount from model outputs 304 generated for one or more additional subjects in at least one of first group of subjects 332 or second group of subjects 334. In one or more examples, the computing system 214 can identify at least one of the one or more first subjects 332 or one or more second subjects 334 having a model output 304 that differs by at least 1 standard deviation, at least 1.5 standard deviations, at least 2 standard deviations, at least 2.5 standard deviations, or at least 3 standard deviations from the mean model output 304 determined for at least one additional group of the first group of subjects 332 or the second group of subjects 334. In various examples, the computing system 214 can apply a penalty to information generated from samples corresponding to subjects that are outliers relative to information generated from samples corresponding to the additional subjects.

[0255] In one or more additional examples, the one or more optimization processes implemented by the computing system 214 in training the computational model 302 may correspond to the number of training cycles and / or the number of iterations for each training cycle performed during the training process.

[0256] In one or more examples, computing system 214 may perform at least 1,000 iterations of the training process to generate computational model 302, at least 3,000 iterations of the training process to generate computational model 302, at least 5,000 iterations of the training process to generate computational model 302, at least 8,000 iterations of the training process to generate computational model 302, at least 10,000 iterations of the training process to generate computational model 302, at least 12,000 iterations of the training process to generate computational model 302, or at least 15,000 iterations of the training process to generate computational model 302. In various examples, computing system 214 may terminate the training process before convergence of a loss function associated with computational model 302. In one or more examples, the number of iterations of the training process to produce computational model 302 may correspond to the number of iterations of the training process performed before the training process is stopped and before the loss function converges.

[0257] In one or more examples, a first stage of a training process implemented by computing system 214 to generate computational model 302 may include determining which samples in training data 330 contain somatic mutations indicative of one or more types of cancer versus samples in training data 330 that do not contain somatic mutations indicative of one or more types of cancer. Computing system 214 may then perform a training process for computational model 302 using samples in training data 330 that contain one or more somatic mutations indicative of one or more types of cancer, as well as using some samples obtained from subjects without detected tumors. In various examples, at least 100 iterations of the first stage of the training process may be performed.

[0258] Additionally, the training process performed by computing system 214 may include a second stage that includes predicting values ​​of tumor metrics for samples that do not contain somatic mutations for one or more types of cancer. The computing system 214 may perform at least 100 additional iterations of the second stage of the training process to generate the computational model 302. The second stage of the training process performed by computing system 214 to generate the computational model 302 may also include training the computational model 302 using a portion of the training data 330 corresponding to samples with somatic mutations indicative of one or more types of cancer, and using a portion of the training data 330 corresponding to predicted values ​​for samples that do not contain somatic mutations indicative of one or more types of cancer, and samples obtained from subjects with no detectable tumors. In various examples, the second stage of the training process performed by computing system 214 to generate the computational model 302 may be performed at least two additional times, at least three additional times, at least four additional times, at least five additional times, or at least six additional times. Once the first stage of the training process and the second stage of the training process are complete, the computing system 214 can perform a validation process on the computational model 302 using information obtained from different samples included in the training data 330.

[0259] In one or more instances, the computing system 214 may perform the training process for multiple computational models 302. In these scenarios, the individual computational models 302 trained by the computing system 214 may correspond to different tissue types that are the source of the genomic material obtained from the subject included in the training data 330. In one or more examples, the individual computational models 302 trained by the computing system 214 may correspond to different classifications of cancer, such as colorectal cancer, lung cancer, pancreatic cancer, bladder cancer, breast cancer, liver cancer, skin cancer, or one or more additional classifications of cancer. In situations where the computing system 214 trains multiple computational models 302 corresponding to different classifications of cancer, the outputs from the individual computational models 302 may be aggregated and analyzed by the computing system 214 to determine the tissue of origin for the subject.

[0260] In various examples, individual computational models 302 corresponding to a given tissue from which genomic material contained in a sample is derived may have different model components. For example, a first computational model generated by the computing system 214 corresponding to a first tissue type may have a first model component corresponding to a first set of classification regions. Furthermore, a second computational model generated by the computing system 214 corresponding to a second tissue type may have a second model component corresponding to a second set of classification regions having at least one classification region different from the first set of classification regions. Furthermore, the weights for individual components of computational models corresponding to different tissue types may be different. That is, in a situation where the first set of classification regions of the first computational model and the second set of classification regions of the second computational model have at least one classification region in common, the weights for the model component corresponding to at least one common classification region may be different for the first computational model and the second computational model.

[0261] Additionally, one or more additional normalization processes can be performed by the computing system when generating the computational model 302. For example, in at least some scenarios, molecules processed with the MBD may be distributed differently across different samples. In one or more examples, molecules may be distributed differently across different samples due to differences in the composition of the reagents used to process the molecules with the MBD. In one or more additional examples, molecules may be distributed differently across different samples due to differences in at least one of the equipment or process conditions used to process the molecules with the MBD.

[0262] For example, for one or more first samples, treatment with the MBD may result in first molecules having a region with a first CG content that are separated into a first partition, and second molecules having a region with a second CG content that are separated into a second partition. Furthermore, for one or more second samples, treatment with the MBD may result in third molecules having a third CG content different from the first CG content that are separated into the first partition, and fourth molecules having a region with a fourth CG content different from the second CG content that are separated into the second partition. In various examples, first molecules can be treated with the MBD and separated into a first partition, and second molecules can be treated with the MBD and separated into a second partition, across a first cutoff range of CG content. Additionally, a third molecule can be treated with the MBD over a second cutoff range of CG content different from the first cutoff range and separated into a first partition, and a fourth molecule can be treated with the MBD and separated into a second partition.

[0263] In one or more instances, a first cutoff range for CG content may include CpGs with 3 to 10 methylated cytosines, and a second cutoff range may include CpGs with 6 to 14 methylated cytosines. In one or more additional instances, a first cutoff range for CG content may include CpGs with 4 to 9 methylated cytosines, and a second cutoff range may include CpGs with 7 to 13 methylated cytosines. In one or more further instances, a first cutoff range for CG content may include CpGs with 5 to 8 methylated cytosines, and a second cutoff range may include 8 to 12 CpGs. In yet other instances, a first cutoff range for CG content may include 4 to 7 CpGs, and a second cutoff range may include 6 to 10 CpGs. In various examples, a first cutoff range of CG content and a second cutoff range of CG content can be used to determine the threshold amount of methylated cytosine used to determine at least one of training sequencing reads or test sequencing reads.In at least some examples, the threshold amount of methylated cytosine can include a cutoff number corresponding to the probability that, for example, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% of the individual molecules processed by MBD are separated into a given distribution.In one or more examples, the threshold amount of methylated cytosine can correspond to 5 methylated cytosines, 6 methylated cytosines, 7 methylated cytosines, 8 methylated cytosines, 9 methylated cytosines, 10 methylated cytosines, 11 methylated cytosines, 12 methylated cytosines, 13 methylated cytosines, or 14 methylated cytosines.

[0264] In one or more examples, the computing system 214 can generate a metric for each classification region based on a quantitative measure determined by analyzing a first number of sequencing reads to identify a first number of nucleic acid molecules having a first amount of CG content and by analyzing a second number of sequencing reads to identify a second number of nucleic acid molecules having a second amount of CG content. In at least some examples, the second number of nucleic acid molecules can be used to modify the metric determined using the first number of nucleic acid molecules to account for variations in the separation of molecules processed using MBD for different samples. In various examples, a first metric can be determined for each classification region by determining, for a given sample, a first quantitative measure corresponding to the number of molecules having a threshold amount of methylated cytosine and a first amount of cytosine-guanine content in one or more partitions (e.g., the second partition 130 and / or the third partition 134) corresponding to the individual classification region. In some embodiments, the first amount of CG content can be at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, or at least 30 CpGs in the nucleic acid molecule. In some embodiments, the first amount of CG content can be between 5-10, 5-15, 5-20, 5-30, 10-15, 10-20, 10-30, 15-20, 15-30, or 20-40 CpGs in the nucleic acid molecule.

[0265] The first metric can also be determined for a given sample by determining a second quantitative measure corresponding to the number of molecules having a threshold amount of methylated cytosine and a first amount of cytosine-guanine content in one or more distributions (e.g., second distribution 130 and / or third distribution 134) corresponding to a plurality of control regions (e.g., positive control regions). By way of example, for each classification region, the first metric can be determined using the first quantitative measure for the individual classification region and a second quantitative measure corresponding to a plurality of control regions.

[0266] The normalization process may also include determining a second metric for a given sample by determining one or more additional quantitative measures based on the number of molecules having at least a threshold amount of methylated cytosine in one or more distributions (e.g., second distribution 130 and / or third distribution 134) and a second amount of cytosine-guanine content corresponding to a plurality of control regions, where the second amount of cytosine-guanine content is less than the first amount of cytosine-guanine content. In some embodiments, the second amount of CG content may be between 5-10, 5-15, 10-15, 10-20, or 15-20 CpGs in the nucleic acid molecule. In one or more examples, the plurality of control regions may be positive control regions and / or negative control regions. In one or more examples, the second metric may be determined using the additional quantitative measure and the second quantitative measure. In at least some examples, the second metric can be determined by determining a ratio of the one or more additional quantitative measures to the second quantitative measure for a given sample. In one or more additional examples, the second metric can be determined by determining the logarithm, e.g., the base 10 logarithm, of the ratio of the one or more additional quantitative measures to the second quantitative measure for a given sample.

[0267] In one or more embodiments, the second metric for a given sample may include a combination of values, where each value corresponds to an additional quantitative measure based on the number of molecules having at least a threshold amount of methylated cytosine and a given number of CpGs for a plurality of control regions, and a second quantitative measure. For example, the first additional quantitative measure can be determined based on the first number of molecules having at least a threshold amount of methylated cytosine in a control region having a first number, for example, 6 CpGs, and the second additional quantitative measure can be determined based on the second number of molecules having at least a threshold amount of methylated cytosine in a control region having a second number, for example, 7 CpGs. In at least some examples, more additional quantitative measures can be determined based on the additional number of molecules having a threshold amount of methylated cytosine in a control region having an additional number of CpGs, for example, 8 CpGs, 9 CpGs, 10 CpGs, etc., up to an upper threshold of CpGs, for example, 12 CpGs, 13 CpGs, or 14 CpGs. Additional quantitative measures can be used to determine ratios determined relative to the second quantitative measure and summed to determine a second metric.

[0268] In various examples, for each classification region, a correlation coefficient can also be determined for different amounts of CpG, which can be used to determine a second metric.In one or more examples, each additional quantitative measure can be corrected by the correlation coefficient, and then the corrected each additional quantitative measure can be aggregated to determine a second metric.In one or more additional examples, the first metric and the second metric can be combined to determine a normalized metric corresponding to a given classification region.In one or more examples, the second metric can be subtracted from the first metric to determine a normalized metric.

[0269] In one or more additional examples, correlation coefficients for a given classification region for each of a plurality of different amounts of cytosine-guanine content can be determined, for example, a first correlation coefficient for 6 CpGs, a second correlation coefficient for 7 CpGs, a third correlation coefficient for 8 CpGs, and so on, up to a threshold amount of CG content. In at least some examples, the correlation coefficients can be determined by analyzing the training data using one or more linear regression techniques. For example, the training data 330 can be fitted to a linear regression model for each classification region to determine the correlation coefficients. In various examples, fitting at least a portion of the training data 330 to the linear regression model can be performed by aggregating additional quantitative measures for the given classification region across various CG contents, for example, 6 CpGs, 7 CpGs, up to a threshold number of CpGs, and determining the average value of the quantitative measures.

[0270] In one or more examples, normalized metrics can reduce variability in quantitative measures determined for individual samples, and in at least some examples, the reduced variability can result in increased accuracy of the model outputs 304 for at least some model outputs 304 determined without implementing an additional normalization process to determine the normalized metrics.

[0271] FIG. 4 is a flowchart of an example method 400 for determining tumor metrics in a subject based on the level of methylation of a classification region, according to one or more implementations. In operation 402, the method 400 may include obtaining training sequence data including training sequencing reads from a plurality of samples of a plurality of subjects. Each training sequencing read may include a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one of the plurality of samples. Each training sequencing read may have a threshold amount of molecules with methylated cytosines contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. In one or more instances, the plurality of samples may include cell-free nucleic acids. In one or more instances, the methylated cytosines may be determined using at least one of bisulfite conversion and sequencing, Tet-assisted bisulfite sequencing (TAB-Seq), differential enzymatic cleavage, treatment with MSRE and / or MDRE, or MBD partitioning. In one or more additional examples, methylated cytosines can be determined using one or more single molecule sequencing methods, such as nanopore DNA sequencing or those described in Eid, J., et al. (2009) Real-time DNA sequencing from single polymerase molecules. Science, 323 (5910), 133-138.

[0272] In one or more examples, the training process may include, by a computing system, obtaining test sequence data from an additional subject not included in the plurality of subjects. The test sequence data may include test sequencing reads from a sample of the additional subject. Each test sequencing read may include a nucleotide sequence corresponding to a fragment of a nucleic acid included in the additional sample. Furthermore, each test sequencing read may have at least a threshold amount of molecules with methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. Based on the additional sequence data, a model may be executed to determine an indication of the presence of cancer in the additional subject. The test sequencing reads may then be analyzed to determine a first quantitative measure derived from the test sequencing reads corresponding to each classification region of the plurality of classification regions. Furthermore, the test sequencing reads may be analyzed to determine a second quantitative measure derived from the test sequencing reads corresponding to each control region of the plurality of control regions. Metrics for each classification region may then be determined based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. An input vector including metrics for each classification region may then be generated. The input vector can be used in the model to determine an indication that additional subjects have cancer.

[0273] In the situation of training a model to determine an estimate of tumor fraction, the training sequence data can include a first portion of the training sequence data, and the second portion of the training sequence data includes additional training sequence data that are different from the training sequence data.In these scenarios, at least one of the first portion of the training sequence data or the second portion of the training sequence data can be analyzed to determine the individual frequencies of the multiple variants present in each sample of the multiple samples.Then, for each sample, the variant among the multiple variants with the highest frequency corresponding to the individual frequency with the maximum value among the individual frequencies derived from each sample can be determined.In one or more embodiments, the maximum variant allele frequency can be determined for each sample.In various examples, the individual measure of tumor fraction for each sample can then be determined based on the maximum individual frequency derived from each sample.

[0274] In at least some examples, the training process for the model can include one or more optimization operations.For example, the training process can include determining one or more additional weights for each sample included in the training data based on the cancer indicator for each sample being within a threshold confidence level.By determining that the cancer indicator for each sample is outside the threshold confidence level, a penalty can be applied to the individual sample during the training process.

[0275] The one or more training optimization operations may also include performing one or more first iterations of a training process on the model using one or more machine learning algorithms and a portion of the training data. Furthermore, first output data for the model may be generated based on the one or more first iterations of the training process. The first output data may correspond to one or more first additional indicators of the presence of cancer in a first individual subject of the plurality of subjects, where the first individual subject may correspond to the portion of the training data. Furthermore, the training process may include combining the first output data and the training data to generate additional training data, and performing one or more second iterations of the training process on the model using the portion of the additional training data. Then, second output data for the model may be generated based on the one or more second iterations of the training process. The second output data may indicate one or more second additional indicators of the presence of cancer in a second individual subject of the plurality of subjects, where the second individual subject corresponds to the portion of the additional training data. In one or more instances, weights for individual classification regions of the plurality of classification regions can be determined based on the first output data and the second output data.

[0276] Furthermore, the training process may include determining that there are some cancer indicators that have at least the threshold value for one or more samples included in the training data, as determined during one or more iterations of the training process.In these scenarios, the modification of one or more weights of the model is not modified or is modified by a minimum amount.Furthermore, the number of additional indicators that cancer is present that are less than the threshold value for one or more additional samples included in the training data, as determined during one or more iterations of the training process, may be determined.In these scenarios, the modification of one or more additional weights of the model may be determined, and the one or more additional weights are modified by more than a minimum amount.

[0277] Furthermore, process 400 may include, at operation 404, analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each individual classification region of the plurality of classification regions. In one or more examples, the first quantitative measure may be determined based on the number of training sequencing reads. In one or more additional examples, the first quantitative measure may be determined based on the number of polynucleotide molecules corresponding to the training sequencing reads. At least a portion of each individual classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having molecules with a threshold amount of methylated cytosine and having at least a threshold cytosine-guanine content in a subject in which cancer is detected. In various examples, the plurality of classification regions may correspond to a genomic region in which at least one mutation is present in a patient in which cancer is detected. Furthermore, the plurality of classification regions may correspond to a first plurality of classification regions for a first cancer type, and a model for a second cancer type may be generated based on a second plurality of classification regions different from the first plurality of classification regions.

[0278] At operation 406, process 400 may include analyzing the training sequencing reads to determine a second quantitative measure from the training sequencing reads corresponding to the plurality of control regions. In one or more examples, the second quantitative measure may be determined based on the number of training sequencing reads. In one or more additional examples, the second quantitative measure may be determined based on the number of polynucleotide molecules corresponding to the training sequencing reads. Each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content. Furthermore, each control region may have molecules with at least a threshold amount of methylated cytosine in subjects in which cancer is detected and in additional subjects in which cancer is not detected.

[0279] Furthermore, process 400 may include, at operation 408, determining a metric for each classification region of the plurality of classification regions based on the first quantitative measure for the each classification region and the second quantitative measure for the plurality of control regions. In one or more examples, the metric for each classification region is determined based on a scaling factor and an error correction factor. In one or more instances, the scaling factor may include a logarithmic function, and the error correction factor may include a pseudo count.

[0280] Process 400 may include generating, by a computing device, training data including metrics for each of the plurality of classification regions for the training sequence reads at operation 410. In an implementation in which the cancer indicator is tumor fraction, the training data may include each measure of tumor fraction for each of the plurality of samples, and the model may be run on each measure of tumor fraction for each of the plurality of samples.

[0281] Process 400 may also include, at operation 412, using the training data to implement one or more machine learning algorithms to generate a model for determining an indication that cancer is present in a subject based on the amount of molecules having methylated cytosines in at least a portion of the plurality of classification regions. The model may determine weights for individual classification regions of the plurality of classification regions, at least a portion of which weights for the individual classification regions may differ from one another. In various examples, the one or more machine learning algorithms may include one or more classification algorithms, and the indication that cancer is present corresponds to a probability that cancer is present in the additional subject. In one or more additional examples, the one or more machine learning algorithms include one or more regression algorithms, and the indication corresponds to an estimate of tumor fraction for the additional sample. In one or more instances, the detection limit of the model for determining the tumor fraction of a sample can be 0.01% or less assuming 95% sensitivity, 0.05% or less assuming 95% sensitivity, 0.1% or less assuming 95% sensitivity, 0.15% or less assuming 95% sensitivity, 0.2% or less assuming 95% sensitivity, 0.25% or less assuming 95% sensitivity, or 0.3% or less assuming 95% sensitivity.

[0282] In various examples, sequence reads provided to the model during or after the training process have at least a threshold amount of methylated cytosines within the classification region. Sequence reads that meet the methylation level can be produced, at least in part, using one or more molecular separation processes. The molecular separation process can include combining a plurality of nucleic acids from at least one of a subject's blood or tissue with a solution containing a certain amount of methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution. Multiple washes of the nucleic acid-MBD protein solution can then be performed with a salt solution to produce several nucleic acid fractions. Each nucleic acid fraction can have a threshold number of molecules with methylated cytosines within a region of the plurality of nucleic acids that has at least a threshold cytosine-guanine content. In one or more instances, one of the multiple washes can be performed with a solution having a sodium chloride (NaCl) concentration to produce one of several nucleic acid fractions with a range of binding strengths to the MBD protein.

[0283] In one or more examples, a first nucleic acid fraction associated with a first distribution of the plurality of distributions of nucleic acids can be determined. The first distribution corresponds to a first range of binding strength to the MBD protein. Further, a first molecular barcode can be bound to the nucleic acid of the first nucleic acid fraction. The first molecular barcode can be associated with the first distribution. Further, a second nucleic acid fraction associated with a second distribution of the plurality of distributions of nucleic acids can be determined. The second distribution may correspond to a second range of binding strength to the MBD protein that is different from the first range of binding strength to the MBD protein. A second molecular barcode can be bound to the nucleic acid of the second nucleic acid fraction. The second molecular barcode is associated with the second distribution.

[0284] In one or more additional examples, at least a portion of some nucleic acid fractions can be combined with a restriction enzyme that cleaves molecules with one or more unmethylated cytosines in a certain amount to produce at least a portion of the samples used to generate sequencing reads. In these scenarios, molecules with a threshold amount of methylated cytosines correspond to the lowest frequency of molecules with methylated cytosines in regions with at least a threshold cytosine-guanine content. In one or more further examples, at least a portion of some nucleic acid fractions can be combined with a restriction enzyme that cleaves molecules with methylated cytosines in a certain amount to produce at least a portion of the samples used to generate sequencing reads. In these situations, molecules with a threshold amount of methylated cytosines correspond to the highest frequency of molecules with methylated cytosines in regions with at least a threshold cytosine-guanine content. Exemplary Methods A. Determine the cancer indicators in the sample

[0285] In some embodiments, the methods disclosed herein include sequencing cfDNA from a sample and determining methylation levels for multiple target regions comprising DNA sequences, including differentially methylated and control regions. In some embodiments, the methods disclosed herein include capturing at least one set of epigenetic target regions from cfDNA or a subsample thereof, comprising contacting the cfDNA or a subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions; determining methylation levels for the target regions; and determining whether an indicator of cancer is present in a sample obtained from a subject. In any of these embodiments, the methylation level / status of a nucleic acid molecule can be determined using one or more of the following methods: partitioning the molecule using a binder that recognizes modified cytosines, MSRE / MDRE digestion, methylation-sensitive conversion such as bisulfite conversion, direct detection during sequencing, or any other suitable approach. Various approaches are described herein.

[0286] In some embodiments, the methods disclosed herein include partitioning a sample containing DNA by contacting the DNA with an agent that recognizes modified cytosines in the DNA, sequencing the DNA, and determining a quantitative measure of the nucleic acid within a plurality of regions.

[0287] In one or more embodiments, the method includes obtaining, by a computing system having one or more hardware processors and memory, training sequence data including training sequencing reads from a plurality of samples from a plurality of subjects, wherein each training sequencing read comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The method also includes analyzing, by the computing system, the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosine in the subject in which cancer is detected and having at least the threshold cytosine-guanine content. The method also includes analyzing the training sequencing reads by a computing system to determine second quantitative measures derived from the training sequencing reads corresponding to a plurality of control regions, each control region of the plurality of control regions having at least a threshold cytosine-guanine content and corresponding to additional genomic regions of the reference genome having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The method also includes determining, by the computing system, metrics for each classification region of the plurality of classification regions based on the first quantitative measures for each classification region and the second quantitative measures for the plurality of control regions. The method also includes generating, by the computing device, training data including metrics for each classification region of the plurality of classification regions for the training sequence reads from the sample of the training subject.The method also includes implementing, by the computing system, one or more machine learning algorithms using the training data to generate a model for determining an indication of the presence of cancer in a subject based on the amount of methylated cytosines in at least a portion of a plurality of classification regions, wherein the model includes weights for each classification region of the plurality of classification regions, at least a portion of the weights for the individual classification regions being different from one another.

[0288] In one or more embodiments, a computing system includes one or more hardware processors and one or more non-transitory computer-readable storage media containing computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including obtaining training sequence data including training sequencing reads from a plurality of samples of a plurality of subjects, wherein each training sequencing read includes a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosines in the subject in which cancer is detected and having at least the threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine second quantitative measures derived from the training sequencing reads corresponding to a plurality of control regions, each of the plurality of control regions corresponding to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in an additional subject in which cancer is not detected. The operations also include determining metrics for each of the plurality of classification regions based on the first quantitative measures for each of the classification regions and the second quantitative measures for the plurality of control regions. The operations also include generating training data including metrics for each of the plurality of classification regions for the training sequence reads from a sample of the training subject.The operations also include using the training data to implement one or more machine learning algorithms to generate a model for determining an indication of the presence of cancer in a subject based on the amount of methylated cytosines in at least a portion of the plurality of classification regions, wherein the model includes weights for each classification region of the plurality of classification regions, and at least a portion of the weights for the individual classification regions are different from one another.

[0289] In one or more embodiments, one or more computer-readable storage media include computer-readable instructions that, when executed by one or more processors of the computing system, cause the computing system to perform operations including obtaining training sequence data including training sequencing reads from a plurality of samples of a plurality of subjects, wherein each training sequencing read includes a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, and each training sequencing read corresponds to a molecule having a threshold amount of methylated cytosines included within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of the each classification region of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosines in the subject in which cancer is detected and having at least the threshold cytosine-guanine content. The operations also include analyzing the training sequencing reads to determine second quantitative measures derived from the training sequencing reads corresponding to a plurality of control regions, each of the plurality of control regions corresponding to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in an additional subject in which cancer is not detected. The operations also include determining metrics for each of the plurality of classification regions based on the first quantitative measures for each of the classification regions and the second quantitative measures for the plurality of control regions. The operations also include generating training data including metrics for each of the plurality of classification regions for the training sequence reads from a sample of the training subject.The operations also include using the training data to implement one or more machine learning algorithms to generate a model for determining an indication of the presence of cancer in a subject based on the amount of methylated cytosines in at least a portion of the plurality of classification regions, wherein the model includes weights for each classification region of the plurality of classification regions, and at least a portion of the weights for the individual classification regions are different from one another.

[0290] In one or more embodiments, the method includes obtaining, by a computing system having one or more hardware processors and memory, sequencing reads from a sample obtained from the subject, wherein each sequencing read comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in the sample and corresponds to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The method may also include determining, by the computing system, a first quantitative measure derived from the sequencing reads corresponding to each classification region of the plurality of classification regions, wherein at least a portion of each classification region of the plurality of classification regions corresponds to a genomic region of the reference genome having the threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in which cancer is detected. The method may also include analyzing, by the computing system, the sequencing reads to determine a second quantitative measure derived from the sequencing reads corresponding to a plurality of control regions, wherein each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least the threshold cytosine-guanine content and having at least the threshold amount of methylated cytosine in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The method may also include determining, by the computing system, a plurality of metrics, each of the plurality of metrics corresponding to a respective one of the plurality of classification regions, based on a first quantitative measure for the respective classification region and a second quantitative measure for the plurality of control regions. The method may also include determining, by the computing system, an indication that cancer is present in the subject based at least in part on the plurality of metrics.

[0291] In one or more embodiments, a computing system includes a processor and a memory storing instructions that, when executed by the processor, configure the computing system to obtain sequencing reads from a sample obtained from the subject, each sequencing read comprising a nucleotide sequence corresponding to a fragment of nucleic acid contained in the sample and corresponding to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The computing system can also determine a first quantitative measure derived from the sequencing reads corresponding to each classification region of the plurality of classification regions, at least a portion of each classification region of the plurality of classification regions corresponding to a genomic region of the reference genome having the threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in which cancer is detected. The computing system can also analyze the sequencing reads to determine a second quantitative measure derived from the sequencing reads corresponding to a plurality of control regions, each control region of the plurality of control regions corresponding to an additional genomic region of the reference genome having at least the threshold cytosine-guanine content and having at least the threshold amount of methylated cytosine in the subject in which cancer is detected and in additional subjects in which cancer is not detected. The computing system can also determine a plurality of metrics, each of which corresponds to a respective one of the plurality of classification regions, based on a first quantitative measure for each of the classification regions and a second quantitative measure for the plurality of control regions. The computing system can also determine an indication that cancer is present in the subject based at least in part on the plurality of metrics.

[0292] The computing system includes a processor and a memory storing instructions that, when executed by the processor, configure the system to obtain test sequence data from the subject, the test sequence data including test sequencing reads derived from a sample from the subject, each test sequencing read including a nucleotide sequence corresponding to a fragment of nucleic acid contained in an additional sample, each test sequencing read corresponding to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. The computing system can also analyze the test sequencing reads to determine a first quantitative measure derived from the test sequencing reads corresponding to each classification region of the plurality of classification regions, at least a portion of the each classification region of the plurality of classification regions corresponding to a genomic region of a reference genome having the threshold amount of methylated cytosine and having at least the threshold cytosine-guanine content in the subject in whom cancer is detected. The computing system can also analyze the test sequencing reads to determine a second quantitative measure derived from the test sequencing reads corresponding to each control region of the plurality of control regions, where each control region of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least a threshold cytosine-guanine content and having at least a threshold amount of methylated cytosine in the subject in which cancer is detected and in the additional subject in which cancer is not detected. The computing system can also determine metrics for each classification region based on the first quantitative measure for each classification region and the second quantitative measure for the plurality of control regions. The computing system can also generate an input vector including the metrics for each classification region. The computing system can also determine an indication that cancer is present in the subject by providing the input vector to a model implementing one or more machine learning techniques to generate an indication that cancer is present in the subject, the model including weights for each classification region of the plurality of classification regions, at least a portion of the weights for the individual classification regions being different from one another. B. Divide the sample into multiple aliquots

[0293] In some embodiments described herein, different forms of DNA (e.g., hypermethylated DNA and hypomethylated DNA) are physically partitioned based on one or more characteristics of the DNA. This approach can be used, for example, to determine whether a particular site or region is hypermethylated or hypomethylated. Partitioning can be performed, for example, before adapters are attached to DNA molecules in a sample, making it easy to include partitioning tags in the adapters. The partitioning tag can be used to identify which partition a molecule is found in. After partitioning (and, if applicable, adapter attachment), further steps such as amplification, target capture, and sequencing can be performed.

[0294] Methylation profiling can include determining the methylation pattern across different regions of genome.For example, after the further steps discussed above, including the distribution and sequencing of molecules based on the degree of methylation (for example, the relative number of methylated nucleic acid bases per molecule), the sequences of molecules in different distributions can be mapped to a reference genome.This can indicate the regions of genome that are more highly methylated or less highly methylated compared to other regions.In this way, genome regions can have different degrees of methylation compared to individual molecules.

[0295] By partitioning nucleic acid molecules in a sample, for example, rare nucleic acid molecules that are predominant in one partition of the sample can be enriched, thereby increasing rare signals. For example, genetic variations that are present in hypermethylated DNA but present less (or not at all) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple partitions of a sample, multidimensional analysis of single molecules can be performed, thus achieving higher sensitivity. Partitioning can include physically dividing nucleic acid molecules into partitions or subsamples based on the presence or absence of one or more methylated nucleic acid bases. Samples can be divided into partitions or subsamples based on features that indicate differential gene expression or disease state. During the analysis of nucleic acids, such as cell-free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA), and cell-free nucleic acid (cfNA), samples can be partitioned based on features or combinations thereof that result in differences in signals between normal and diseased states.

[0296] In some embodiments, hypermethylated and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit differential methylation signatures of particular immune cell types, such as rare immune cell types, tumor cells, or cell types that do not normally contribute to the DNA sample (e.g., cfDNA, etc.) being analyzed.

[0297] In some cases, the heterogeneous DNA in a sample is divided into two or more partitions (for example, at least three, four, five, six or seven partitions).In some embodiments, each partition is differentially tagged.The tagged partitions can then be pooled together for collective sample preparation and / or sequencing.The partition-tagging-pooling step can be carried out more than once, and each partition is carried out based on different characteristics (examples are provided herein) and tagged with a differential tag that distinguishes it from other partitions and partitioning means.In other cases, the differentially tagged partitions are sequenced separately.

[0298] In some embodiments, sequence reads are obtained from differentially tagged and pooled DNA and analyzed in silico. Tags can be used to distinguish between reads from different distributions. Analysis to detect genetic variants can be performed at the distribution level as well as at the total nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants, such as CNVs, SNVs, indels, and fusions, in the nucleic acids in each distribution. In some cases, in silico analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the position of nucleosomes in chromatin. Higher coverage can be correlated with higher nucleosome occupancy in a genomic region, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0299] In some embodiments, the partitioning is based on one or more characteristics, such as methylation. Where applicable, as part of the data analysis or partitioning, appropriate techniques can be used to sort molecules according to other characteristics, such as sequence length, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins binding to DNA. The resulting partitioning can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning based on cytosine modification (e.g., cytosine methylation) or methylation is typically performed, optionally combined with at least one additional partitioning step that can be based on any of the aforementioned DNA characteristics or forms. In some embodiments, a heterogeneous nucleic acid population is partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (for example, 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation), and the association and level of association with one or more proteins, such as histones.Alternatively or additionally, heterogeneous nucleic acid populations can be divided into nucleic acid molecules associated with nucleosomes and nucleic acid molecules that lack nucleosomes.Alternatively or additionally, heterogeneous nucleic acid populations can be divided into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA).Alternatively or additionally, heterogeneous nucleic acid populations can be divided based on nucleic acid length (for example, molecules that are up to 160 bp and molecules that have a length longer than 160 bp).

[0300] The agent used to partition the population of nucleic acids within a sample can be an affinity agent, such as an antibody with a desired specificity, a natural binding partner or variant thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected to have specificity for a given target, for example, by phage display. In some embodiments, the agent used for partitioning is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, e.g., a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is the product of a procedure that affects a first nucleobase in the DNA of the sample differently from a second nucleobase in the DNA. In some embodiments, the modified nucleobase can be a "converted nucleobase," i.e., one whose base-pairing specificity has been altered by a procedure. For example, in certain procedures, unmethylated or unmodified cytosine is converted to dihydrouracil, or more commonly, at least one modified or unmodified form of cytosine undergoes deamination, thereby generating uracil (considered a modified nucleobase in DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, for example, antibodies that recognize modified nucleobases, such as methylcytosine (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes modified cytosines other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Alternative partitioning agents include the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as MeCP2.

[0301] Additional non-limiting examples of partitioning agents are histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.

[0302] The binding of the partitioning agent to specific nucleic acids and the partitioning of nucleic acids into subsamples can be carried out to a certain extent, or can be carried out essentially in a binary manner.In some cases, nucleic acids containing a greater proportion of a particular modification bind to the agent to a greater extent than nucleic acids containing a smaller proportion of the modification.Similarly, partitioning can produce subsamples containing a greater proportion and a smaller proportion of nucleic acids containing a particular modification.Alternatively, partitioning can produce subsamples containing essentially all or none of the nucleic acids containing the modification.In all cases, various levels of modification can be sequentially eluted from the partitioning agent.

[0303] In some embodiments, partitioning can include both binary partitioning and partitioning based on the degree / level of modification. For example, methylated fragments can be partitioned by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, further partitioning can include eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, fragments with higher methylation levels are eluted.

[0304] In some cases, the final distribution is enriched for nucleic acids with different degrees of modification (over- or under-representation of the modification). Over- and under-representation can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is two, nucleic acids containing more than two 5-methylcytosine residues are over-represented in this modification, and nucleic acids with one or zero 5-methylcytosine residues are under-represented. The effect of affinity separation is to enrich nucleic acids with over-represented modifications in the binding phase and under-represented modifications in the non-binding phase (i.e., in solution). Nucleic acids in the binding phase can be eluted prior to subsequent processing.

[0305] When using MeDIP or the MethylMiner® Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be separated using sequential elution. For example, the low-methylation (unmethylated) distribution can be separated from the methylated distribution by contacting the nucleic acid population with MBD from the kit attached to magnetic beads. The beads are used to separate the methylated nucleic acids from the unmethylated nucleic acids. Subsequently, one or more sequential elution steps are performed to elute nucleic acids with different levels of methylation. For example, the first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, for example, at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After eluting such methylated nucleic acids, magnetic separation is again used to separate nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps can be repeated to create various partitions, such as a low-methylated partition (enriched for nucleic acids with no methylation), a methylated partition (enriched for nucleic acids with low levels of methylation), and a high-methylated partition (enriched for nucleic acids with high levels of methylation).

[0306] In some methods, nucleic acids bound to the agent used for affinity separation-based partitioning are subjected to a wash step. The wash step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched for nucleic acids that have modifications at a level close to the average or median (i.e., halfway between the nucleic acids that remain bound to the solid phase when the sample is first contacted with the agent and the nucleic acids that are not bound to the solid phase).

[0307] Affinity separation results in at least two, sometimes three or more, distributions of nucleic acids with different degrees of modification. The distributions are still separate, but at least one distribution's nucleic acid, and usually two or three (or more) distributions, is linked to a nucleic acid tag, usually provided as an adapter component, and the nucleic acids in different distributions receive different tags that distinguish one distribution's members from another. The tags linked to nucleic acid molecules of the same distribution can be the same or different from each other. However, if different from each other, the tags can have a common code portion so that the molecules to which they are attached are identified as belonging to a specific distribution.

[0308] For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.

[0309] In some embodiments, nucleic acid molecules may be fractionated into different partitions based on nucleic acid molecules that are bound to a particular protein or fragment thereof and those that are not bound to that particular protein or fragment thereof.

[0310] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific protein properties. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and serve as a basis for fractionation include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate nucleic acid molecules based on protein-bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein-bound regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow field separation (AF4).

[0311] In some embodiments, the sample is divided into multiple aliquots by contacting the nucleic acid with an antibody that recognizes a modified nucleic acid base in DNA, and the modified nucleic acid base can be a modified cytosine or the product of a procedure that affects a first nucleic acid base in DNA differently from a second nucleic acid base in the DNA of the sample. In some embodiments, the modified nucleic acid base is 5mC. In some embodiments, the modified nucleic acid base is 5caC. In some embodiments, the modified nucleic acid base is dihydrouracil (DHU). In some embodiments, the single-stranded DNA is divided using an antibody that recognizes a modified nucleic acid base in DNA.

[0312] In some embodiments, partitioning is performed by contacting the nucleic acid with the methyl-binding domain ("MBD") of a methyl-binding protein ("MBP"). In some such embodiments, the nucleic acid is contacted with the entire MBP. In some embodiments, the MBD binds to 5-methylcytosine (5mC), and the MBP comprises an MBD, referred to herein interchangeably as a methyl-binding protein or a methyl-binding domain protein. In some embodiments, the MBD binds to 5mC and 5hmC. In some embodiments, the MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 streptavidin, via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentration.

[0313] In some embodiments, the bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This can be performed instead of or in addition to the elution step using NaCl discussed above.

[0314] Examples of agents that recognize modified nucleobases contemplated herein include, but are not limited to, the following: (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine. (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 bind preferentially to 5-hydroxymethyl-cytosine over unmodified cytosine. (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind 5-formylcytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)). (d) An antibody specific for one or more methylated or modified nucleobases or their conversion products, e.g., 5mC, 5caC, or DHU.

[0315] Generally, elution is a function of the number of modifications (e.g., the number of methylation sites) per molecule, with higher salt concentrations resulting in more methylated molecules. A series of elution buffers with increasing NaCl concentrations can be used to separate DNA into distinct populations based on the degree of methylation. Salt concentrations can range from about 100 mM to about 2500 mM NaCl. In one embodiment, the process results in three partitions. The molecules are contacted with a solution containing molecules at a first salt concentration, including an agent that recognizes modified nucleobases, and the molecules may be bound to a capture moiety, e.g., streptavidin. At the first salt concentration, one population of molecules binds to the agent, while another population remains unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first partition enriches for hypomethylated forms of DNA that remain unbound at low salt concentrations, e.g., 100 mM or 160 mM. A second partition enriched for intermediately methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM. This partition is also separated from the sample. A third partition enriched for highly methylated forms of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0316] In some embodiments, methylated DNA is purified using a monoclonal antibody raised against 5-methylcytidine (5mC). To obtain single-stranded DNA fragments, the DNA is denatured, for example, at 95°C. The DNA bound to the antibody is immunoprecipitated using standard or magnetic bead-coupled protein G and washed after incubation with the anti-5mC antibody. The DNA can then be eluted. The partition may include unprecipitated DNA and one or more partitions eluted from the beads.

[0317] In some embodiments, sample DNA (e.g., between 5 and 200 ng) is mixed with a methyl-binding domain (MBD) buffer, conjugated to magnetic beads with MBD protein, and incubated overnight. Methylated DNA (hypermethylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated (hypomethylated DNA) or less methylated DNA (intermediate methylation) is washed from the beads with buffers containing increasing concentrations of salt. For example, one, two, or more fractions containing unmethylated, hypomethylated, and / or intermediate methylated DNA can be obtained from such washes. Finally, highly methylated DNA (hypermethylated DNA) is eluted from the MBD protein using a high-salt buffer. In some embodiments, these washes result in three partitions of DNA with increasing methylation levels: a hypomethylated partition, an intermediate methylated fraction, and a hypermethylated partition.

[0318] In some embodiments, the partitioning procedure may result in incomplete sorting of DNA molecules among the subsamples. For example, a small number of molecules in an unmethylated or hypomethylated subsample may be highly modified (e.g., hypermethylated), and / or a small number of molecules in a hypermethylated subsample may be unmodified or mostly unmodified (e.g., unmethylated or mostly unmethylated). Such molecules are considered to be non-specifically partitioned.

[0319] In some embodiments, non-specifically partitioned molecules are removed using a methylation-dependent nuclease, e.g., a methylation-dependent restriction enzyme (MDRE), which digests / cleaves DNA whose restriction enzyme (RE) recognition site contains a methylated nucleotide but does not cleave DNA whose restriction enzyme (RE) recognition site contains an unmethylated nucleotide. In some embodiments, non-specifically partitioned molecules are removed using a methylation-sensitive nuclease, e.g., a methylation-sensitive restriction enzyme (MSRE), which digests / cleaves DNA whose restriction enzyme (RE) recognition site contains an unmethylated nucleotide but does not cleave DNA whose restriction enzyme (RE) recognition site contains a methylated nucleotide. For example, in some embodiments, a low-methylation aliquot is contacted with a methylation-dependent nuclease, e.g., a methylation-dependent restriction enzyme, thereby degrading non-specifically partitioned DNA, e.g., methylated DNA, in the aliquot. Alternatively, or in addition, a high-methylation aliquot is contacted with a methylation-sensitive endonuclease, e.g., a methylation-sensitive restriction enzyme, thereby degrading non-specifically partitioned DNA in the aliquot.

[0320] The degradation of non-specifically distributed DNA in one or more divided aliquots can improve the performance of the method that relies on the accurate distribution of DNA based on cytosine modification.For example, such degradation can improve sensitivity and / or simplify downstream analysis.In some embodiments, DNA is divided based on modification such as methylation, and then non-specifically distributed DNA is removed using MDRE and / or MSRE as described herein, thereby improving efficiency and / or cost compared with the DNA analysis method that comprises a procedure that affects first nucleobase differently from second nucleobase, such as bisulfite sequencing or bisulfite conversion.

[0321] In some embodiments, one or more nucleases are used to degrade non-specifically distributed DNA molecules. In some embodiments, the aliquot is contacted with multiple nucleases. The aliquot can be contacted with the nucleases sequentially or simultaneously. Simultaneous use of nucleases can be advantageous to avoid unnecessary sample manipulation if the nucleases are active under similar conditions (e.g., buffer composition). Contacting the aliquot with more than one methylation-dependent restriction enzyme can more thoroughly degrade non-specifically distributed hypermethylated DNA. Contacting the aliquot with more than one methylation-sensitive restriction enzyme can more thoroughly degrade non-specifically distributed hypomethylated and / or unmethylated DNA.

[0322] In some embodiments, the methylation-dependent nuclease comprises one or more of MspJI, LpnPI, FspEI, or McrBC. In some embodiments, at least two methylation-dependent nucleases are used. In some embodiments, at least three methylation-dependent nucleases are used.

[0323] In some embodiments, the methylation-sensitive nucleases include one or more of AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Pspl406I, PvuI, SacII, SalI, SmaI, and SnaBI. In some embodiments, at least two methylation-sensitive nucleases are used. In some embodiments, at least three methylation-sensitive nucleases are used. In some embodiments, the methylation-sensitive nuclease includes BstUI and HpaII. In some embodiments, the two methylation-sensitive nucleases include HhaI and AccII. In some embodiments, the methylation-sensitive nucleases include BstUI, HpaII, and Hin6I.

[0324] In some embodiments, the DNA fraction is desalted and concentrated in preparation for the enzymatic steps of library preparation. C. Adapter Ligation

[0325] In some embodiments, adapters are added to DNA. This can be done in conjunction with an amplification procedure, for example, by providing adapters to the 5' portion of primers (when PCR is used, this can be referred to as library prep-PCR or LP-PCR). In some embodiments, adapters are added by other approaches, such as ligation. In some such methods, a first adapter is added to a nucleic acid by ligation to its 3' end before partitioning or capture, which can include ligation to single-stranded DNA. The adapter can be used as a priming site for second-strand synthesis, for example, using a universal primer and DNA polymerase. A second adapter can then be ligated to the 3' end of at least the second strand of the now double-stranded molecule. In some embodiments, the first adapter includes an affinity tag, such as biotin, and the nucleic acid ligated to the first adapter is attached to a solid support (e.g., beads), which can include a binding partner for the affinity tag, such as streptavidin. For further discussion of related procedures, see Gansauge et al., Nature Protocols 8: 737-748 (2013). Commercially available kits for preparing sequencing libraries compatible with single-stranded nucleic acids are available, such as the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, the nucleic acid is amplified.

[0326] Preferably, the adapters contain a sufficient number of different tags such that the number of tag combinations results in a low probability, e.g., 95, 99, or 99.9%, that two nucleic acids with the same start and end points will receive the same tag combination. Adapters, whether they have the same or different tags, may contain the same or different primer binding sites, although preferably the adapters contain the same primer binding sites.

[0327] In some embodiments, after binding of the adapters, the nucleic acids are subjected to amplification, which can be, for example, using universal primers that recognize primer binding sites in the adapters.

[0328] In some embodiments, after adapter binding, the DNA is partitioned, which involves contacting the DNA with an agent that preferentially binds to nucleic acids with epigenetic modifications. The nucleic acids are partitioned into at least two aliquots that differ in the degree to which the nucleic acids have the modification, following binding with the agent. For example, if the agent has affinity for nucleic acids with the modification, nucleic acids in which the modification is overrepresented (relative to the median representation in the population) will preferentially bind to the agent, while nucleic acids in which the modification is underrepresented will not bind to the agent or will be more easily eluted from the agent. The nucleic acids can then be amplified from primers that bind to primer binding sites in the adapters. Alternatively, partitioning can be performed before adapter binding, in which case the adapters can contain differential tags that contain components that identify which partitions contain molecules.

[0214] In some embodiments, the nucleic acids are ligated at both ends to Y-shaped adapters that contain primer binding sites and tags. The molecules are amplified. D. Tagging

[0329] "Tagging" a DNA molecule is the procedure of attaching or associating a tag to a DNA molecule. The tag can be a molecule, e.g., a nucleic acid, that contains information that indicates the characteristics of the molecule to which the tag is associated. For example, a molecule can have a sample tag (that distinguishes a molecule in one sample from molecules in a different sample) or a molecular tag / molecular barcode / barcode (that distinguishes different molecules from each other (in both unique and non-unique tagging scenarios)). For methods that involve a partitioning step, a partitioning tag (that distinguishes a molecule in one partition from molecules in a different partition) can be included. In some embodiments, the adapter added to the DNA molecule comprises a tag. In certain embodiments, the tag can comprise a barcode or a combination of barcodes. As used herein, the term "barcode" refers to a nucleic acid molecule having a specific nucleotide sequence, or the nucleotide sequence itself, depending on the context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or sequences with a certain Hamming distance, as desired for a particular purpose. Thus, for example, a molecular barcode can be composed of one barcode or a combination of two barcodes, each attached to a different end of the molecule. Additionally or alternatively, different sets of molecular barcodes, or molecular tags, can be used for different partitions and / or samples, such that the barcodes, through their individual sequences, serve as molecular tags and also function to identify corresponding partitions and / or samples based on the sets they are members of.

[0330] In some embodiments, two or more distributions, for example, each distribution, are differentially tagged. To correlate the tag(s) with a specific distribution, tags can be used to label individual polynucleotide population distributions. Alternatively, tags can be used in embodiments that do not use a distribution step. In some embodiments, a single tag can be used to label a specific distribution. In some embodiments, multiple different tags can be used to label a specific distribution. In embodiments that use multiple tags to label a specific distribution, the set of tags used to label one distribution can be easily distinguished from the set of tags used to label other distributions. In some embodiments, the tag may have additional functionality, for example, the tag may be used to index the origin of the sample or as a unique molecular identifier (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations, e.g., as in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011); Kou et al., PLoS ONE, 11: e0146638 (2016)), or as a non-unique molecular identifier, e.g., as described in U.S. Pat. No. 9,598,731. Similarly, in some embodiments, the tag may have additional functionality, for example, the tag may be used to index the origin of the sample or as a non-unique molecular identifier (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).

[0217] In some embodiments, the partitioning tagging comprises tagging molecules in each partition with a partitioning tag. After the distributions are recombined (e.g., to reduce the number of sequencing runs required and avoid unnecessary costs) and the molecules are sequenced, the distribution tag identifies the distribution of origin. In some embodiments, the distribution tag can serve as an identifier for the distribution of origin and the molecule, i.e., different distributions are tagged with different sets of molecular tags, e.g., composed of barcode pairs.In this way, one or more molecular barcodes attached to molecules are useful not only to indicate the distribution of origin, but also to identify molecules within the distribution. For example, a first set of 35 barcodes can be used to tag molecules in a first distribution, while a second set of 35 barcodes can be used to tag molecules in a second distribution.

[0331] In some embodiments, after partitioning and tagging with partition tags, the molecules can be pooled for single sequencing. In some embodiments, sample tags are added to the molecules, for example, in a step after partition tag addition and pooling. Sample tags can facilitate pooling materials generated from multiple samples for single sequencing.

[0332] Alternatively, in some embodiments, the distribution tag may be correlated with the sample and distribution. As a simple example, a first tag may indicate a first distribution of a first sample, a second tag may indicate a second distribution of the first sample, a third tag may indicate a first distribution of a second sample, and a fourth tag may indicate a second distribution of the second sample.

[0333] Tags may be attached to molecules that have already been distributed based on one or more characteristics, but the final tagged molecules in the library may no longer have those characteristics. For example, single-stranded DNA molecules may be distributed and tagged, but the final tagged molecules in the library are likely to be double-stranded. Similarly, DNA may be distributed based on different levels of methylation, but the tagged molecules derived from these molecules in the final library are likely to be unmethylated. Thus, tags attached to molecules in the library typically represent the characteristics of the "parent molecules" from which the final tagged molecules are derived, not necessarily the characteristics of the tagged molecules themselves.

[0334] For example, molecules in a first distribution are tagged and labeled using barcodes 1, 2, 3, 4, etc., molecules in a second distribution are tagged and labeled using barcodes A, B, C, D, etc., molecules in a third distribution are tagged and labeled using barcodes a, b, c, d, etc. Differ...

Claims

1. A computing system having one or more hardware processors and memory obtains training sequence data including training sequence reads derived from multiple samples of multiple subjects, wherein each training sequence read includes a nucleotide sequence corresponding to a nucleic acid fragment contained in one of the multiple samples, and each training sequence read corresponds to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least a threshold cytosine-guanine content. A step of analyzing the training sequencing reads using the computing system to determine a first quantitative measure derived from the training sequencing reads corresponding to individual classification regions of a plurality of classification regions, wherein at least a portion of the individual classification regions of the plurality of classification regions corresponds to a genomic region of a reference genome having the threshold amount of methylated cytosine and at least the threshold cytosine-guanine content in a subject in which cancer is detected. A step of analyzing the training sequencing reads using the computing system to determine a second quantitative measure derived from the training sequencing reads corresponding to a plurality of control regions, wherein each of the plurality of control regions corresponds to an additional genomic region of the reference genome having at least the threshold cytosine-guanine content and having at least the threshold amount of methylated cytosine in subjects in which cancer is detected and in additional subjects in which cancer is not detected. The computing system determines metrics for each of the multiple classification regions based on a first quantitative measure for each classification region and a second quantitative measure for the multiple control regions. The computing system generates training data including the metrics for each of the multiple classification regions for training sequence reads derived from the sample to be trained, The steps of the computing system to implement one or more machine learning algorithms using the training data to generate a model for determining an indicator of cancer presence in a subject based on the amount of methylated cytosine in at least a portion of the plurality of classification regions, wherein the model includes weights for each of the plurality of classification regions, and at least a portion of the weights for each of the individual classification regions are different from each other. A method that includes this.

2. The steps of obtaining test sequence data from an additional target not included in the plurality of targets using the computing system, wherein the test sequence data includes test sequencing reads derived from a sample of the additional target, each test sequencing read includes a nucleotide sequence corresponding to a nucleic acid fragment contained in the additional sample, and each test sequencing read corresponds to a molecule having a threshold amount of methylated cytosine contained within a region of the nucleotide sequence having at least the threshold cytosine-guanine content; A step of determining the indicators for the presence of cancer in the additional subjects using the model and the additional sequence data. The method according to claim 1, including the method described in claim 1.

3. The computing system performs the steps of analyzing the test sequencing reads and determining a first quantitative measure derived from the test sequencing reads corresponding to each of the multiple classification regions, The computing system performs the steps of analyzing the test sequencing reads to determine a second quantitative measure derived from the test sequencing reads corresponding to each of the multiple control regions, The computing system determines the metrics for each of the classification domains based on the first quantitative measure for each of the classification domains and the second quantitative measure for the plurality of control domains. The computing system generates an input vector including the metrics for each of the classification domains. Includes, The model uses the input vector to determine the indicator that cancer is present in the additional target. The method according to claim 2.

4. The aforementioned one or more machine learning algorithms include one or more regression algorithms, The aforementioned index corresponds to the estimated value of the tumor fraction of the additional sample. The method according to claim 2.

5. The training sequence determination read includes a first portion of the training sequence data, and the additional training sequence determination read includes a second portion of the training sequence data, and the additional training sequence determination read differs from the training sequence determination read. The method described above is The computing system performs the steps of analyzing at least one of the first portion of the training sequence data or the second portion of the training sequence data to determine the individual frequencies of the multiple variants present in each of the multiple samples, The computing system performs the steps of determining, for each sample, one variant among the plurality of variants that has the highest frequency corresponding to the individual frequency having the maximum value among the individual frequencies originating from each sample, The computing system performs the steps of determining individual scales of tumor fractions for individual samples based on the maximum value of the individual frequencies derived from each sample, including, The method according to claim 4.

6. The training data includes the individual scales of the tumor fraction for each of the plurality of samples, The model is generated based on the individual scales of the tumor fractions for each of the plurality of samples. The method according to claim 5.

7. (a) The metrics for each of the classification regions are determined based on a scaling factor and an error correction factor, and / or (b) The plurality of samples and the additional samples include cell-free nucleic acids, The method according to claim 1.

8. The computing system performs a training process using the training data to generate the model, wherein the training process is: The computing system determines one or more additional weights for individual samples included in the training data based on the cancer index for each individual sample within a threshold confidence level. The method according to claim 1, comprising the step of including claim 1.

9. The cancer index for each individual sample is outside the range of the threshold confidence level, and the method is The computing system includes the step of applying a penalty to the weight of each individual sample during the training process, The method according to claim 8.

10. The computing system performs one or more first iterations of the training process on the model using one or more machine learning algorithms, using a portion of the training data. The computing system generates first output data for the model based on one or more first iterations of the training process, wherein the first output data corresponds to one or more first additional indicators indicating the presence of cancer in a first individual subject among the plurality of subjects, and the first individual subject corresponds to the portion of the training data. Steps and The method according to claim 8, including the method described in claim 8.

11. The computing system performs the steps of combining the first output data and the training data to produce additional training data, The computing system performs one or more second iterations of the training process on the model using a portion of the additional training data. The computing system generates second output data for the model based on one or more second iterations of the training process, wherein the second output data indicates one or more second additional indicators of the presence of cancer in a second individual subject among the plurality of subjects, and the second individual subject corresponds to the portion of the additional training data. Steps and This includes, if necessary, determining the weights for each of the multiple classification regions based on the first and second output data. The method according to claim 10.

12. The computing system determines, during one or more iterations of the training process, that several indicators of cancer presence are thresholds for at least one or more samples included in the training data. The computing system determines whether any modifications to one or more weights of the model are made or made by a minimum amount. The method according to claim 8, including the method described in claim 8.

13. The computing system determines, during the one or more iterations of the training process, that the number of additional indicators in which cancer is present is less than the threshold for one or more additional samples included in the training data. The computing system determines that the modification to one or more additional weights of the model is greater than the minimum amount. The method according to claim 12, including the method described in claim 12.

14. The method according to claim 1, further comprising the step of generating a model output, for example, the model output indicating a tumor status of whether or not it is detected, an estimate of the tumor fraction, the probability of the tumor being present, or a tumor tissue index.

15. (i) a computer system including means for carrying out the method described in any one of claims 1 to 14; (ii) a computer program, which, when the program is executed by a computer, includes an instruction causing the computer to carry out the method described in any one of claims 1 to 14; or (iii) a computer-readable storage medium, which, when executed by a computer, includes an instruction causing the computer to carry out the method described in any one of claims 1 to 14.