Methods and systems for analyzing nucleic acid molecules

The method improves liquid biopsy assays by tagging and amplifying different nucleic acid forms in bodily fluids, addressing sensitivity and complexity issues, enabling effective detection of cancer-related genetic markers.

JP7756676B2Active Publication Date: 2025-10-20GUARDANT HEALTH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023060942
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-08-25
Filing Date
2023-04-04
Publication Date
2025-10-20
Estimated Expiration
2037-12-22

AI Technical Summary

Technical Problem

Existing liquid biopsy assays for cancer detection face challenges in sensitivity due to low amounts of nucleic acids in bodily fluids and heterogeneity in nucleic acid forms, leading to loss of data and complexity in analysis.

Method used

A method involving linking different forms of nucleic acids (e.g., double-stranded DNA, single-stranded DNA, single-stranded RNA) to distinct tags, amplifying and assaying the tagged nucleic acids to decode their forms, and optionally enriching or differentially tagging based on modifications like 5-methylcytosine binding, to enhance sensitivity and specificity.

Benefits of technology

Enhances the sensitivity and specificity of nucleic acid analysis in bodily fluids, allowing for better detection of somatic mutations, copy number variations, and other genetic markers associated with cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756676000004
    Figure 0007756676000004
  • Figure 0007756676000005
    Figure 0007756676000005
  • Figure 0007756676000006
    Figure 0007756676000006
Patent Text Reader

Abstract

To provide methods and systems for analyzing nucleic acid molecules.SOLUTION: The present disclosure provides methods for processing nucleic acid populations containing different forms (e.g., RNA and DNA, single-stranded or double-stranded) and / or extents of modification (e.g., cytosine methylation, association with proteins). These methods accommodate multiple forms and / or modifications of nucleic acid in a sample, such that sequence information can be obtained for multiple forms. The methods also preserve the identity of multiple forms or modified states through processing and analysis, such that analysis of sequence can be combined with epigenetic analysis.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Table 1] References to Related Patent Applications This application claims the benefit of the priority dates of U.S. Provisional Patent Application Nos. 62 / 438,240, filed December 22, 2016, 62 / 512,936, filed May 31, 2017, and 62 / 550,540, filed August 25, 2017, all of which are incorporated by reference herein in their entireties. [Background technology]

[0002] Cancer is a leading cause of disease worldwide. Tens of millions of people worldwide are diagnosed with cancer each year, and more than half ultimately die from it. In many countries, cancer ranks as the second most common cause of death after cardiovascular disease. Early detection is associated with improved outcomes for many cancers.

[0003] Cancer can be caused by the accumulation of genetic mutations within an individual's normal cells, at least in part resulting in improperly regulated cell division. Such mutations generally include copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (IDs), and epigenetic mutations include 5-methylation of cytosine (5-methylcytosine) and association of DNA with chromatin and transcription factors.

[0004] Cancer is often detected by tumor biopsy followed by analysis of cells, markers, or DNA extracted from the cells. However, more recently, it has been proposed that cancer can also be detected from cell-free nucleic acids in bodily fluids such as blood or urine. Such tests have the advantage of being non-invasive and can be performed without identifying suspected cancer cells in a biopsy. However, such tests are complicated by the fact that the amount of nucleic acid in bodily fluids is extremely low and the types of nucleic acid present are heterogeneous in form (e.g., RNA and DNA, single-stranded and double-stranded, various states of post-replicative modifications, and association with proteins such as histones).

[0005] It is desirable to increase the sensitivity of liquid biopsy assays while reducing the loss of circulating nucleic acids (original material) or data in the process. Summary of the Invention [Means for solving the problem]

[0006] overview The present disclosure provides methods, compositions, and systems for analyzing a nucleic acid population containing at least two forms of nucleic acid selected from double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, the method includes: (a) linking at least one of the forms of nucleic acid to at least one tag nucleic acid to distinguish the forms from each other; (b) amplifying the form of nucleic acid, at least one of which is linked to at least one nucleic acid tag, where the nucleic acid and the linked nucleic acid tag, if present, are amplified to produce amplified nucleic acids, of which at least one form is tagged; (c) assaying the sequence data of the amplified nucleic acids, at least some of which are tagged; and (d) decoding the tag nucleic acid molecules of the amplified nucleic acids to identify the forms of nucleic acids in the population and provide original templates of the amplified nucleic acids linked to the tag nucleic acid molecules whose sequence data was assayed.

[0007] In some embodiments, the method further comprises enriching at least one of the forms relative to one or more of the other forms. In some embodiments, at least 70% of the molecules of each form of nucleic acid in the population are amplified in step (b). In some embodiments, at least three forms of nucleic acid are present in the population, and at least two of the forms are linked to different tag nucleic acid forms that distinguish each of the three forms from one another. In some embodiments, each of the at least three forms of nucleic acid in the population is linked to a different tag. In some embodiments, each molecule of the same form is linked to a tag containing the same identifying information tag (e.g., a tag having or containing the same sequence). In some embodiments, molecules of the same form are linked to different types of tags. In some embodiments, step (a) comprises subjecting the population to reverse transcription using tagged primers, and the tagged primers are incorporated into cDNA made from the RNA in the population. In some embodiments, the reverse transcription is sequence-specific. In some embodiments, the reverse transcription is random. In some embodiments, the method further comprises degrading RNA that is duplexed with the cDNA. In some embodiments, the method further comprises separating single-stranded DNA from double-stranded DNA and ligating nucleic acid tags to the double-stranded DNA. In some embodiments, the single-stranded DNA is separated by hybridization with one or more capture probes. In some embodiments, the method further comprises differentially tagging the single-stranded DNA with single-stranded tags using a ligase that functions on single-stranded nucleic acids and differentially tagging the double-stranded DNA with double-stranded adapters using a ligase that functions on double-stranded nucleic acids. In some embodiments, the method further comprises pooling tagged nucleic acids containing different forms of nucleic acids prior to assaying. In some embodiments, the method further comprises analyzing the pool of DNA divided separately in individual assays. The assays can be identical, substantially similar, equivalent, or different.

[0008] In any of the above methods, the sequence data may indicate the presence of somatic or germline mutations or copy number variations or single nucleotide variations or indels or gene fusions.

[0009] The present disclosure further provides methods for analyzing a nucleic acid population containing nucleic acids with different degrees of modification. In some examples, the present disclosure provides methods for screening for characteristics associated with disease (e.g., 5' methylcytosine). The method includes contacting a population of nucleic acids with an agent (such as a methyl-binding domain or protein) that preferentially binds nucleic acids bearing the modification; separating nucleic acids in a first pool that are bound to the agent from nucleic acids in a second pool that are not bound to the agent, wherein the nucleic acids in the first pool are over-represented for the modification and the nucleic acids in the second pool are under-represented for the modification; linking the nucleic acids in the first pool and / or the second pool to one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool to produce a population of tagged nucleic acids; amplifying the tagged nucleic acids, wherein the nucleic acids and the linked tags are amplified; assaying sequence data of the amplified nucleic acids and the linked tags; and decoding the tags to determine whether the nucleic acids whose sequence data was assayed were amplified from templates in the first or second pool.

[0010] In some embodiments, the modification is binding of the nucleic acid to a protein. In some embodiments, the protein is a histone or a transcription factor. In some embodiments, the nucleic acid modification is a post-replication modification to a nucleotide. In some embodiments, the post-replication modification is 5-methylcytosine, and the extent of binding of the capture agent to the nucleic acid increases with the amount of 5-methylcytosine in the nucleic acid. In some embodiments, the post-replication modification is 5-hydroxymethylcytosine, and the extent of binding of the agent to the nucleic acid increases with the amount of 5-hydroxymethylcytosine in the nucleic acid. In some embodiments, the post-replication modification is 5-formylcytosine or 5-carboxylcytosine, and the extent of binding of the agent increases with the amount of 5-formylcytosine or 5-carboxylcytosine in the nucleic acid. In some embodiments, the post-replication modification is N 6 -methyladenine. In some embodiments, the method further comprises washing the nucleic acids bound to the agent and recovering the wash as a third pool containing nucleic acids with an intermediate degree of post-replication modification relative to the first and second pools. Some methods further comprise pooling the tagged nucleic acids from the first and second pools prior to the assaying step. In some embodiments, the agent comprises a methyl-binding domain or a methyl-CpG binding domain (MBD). The MBD can be a protein, antibody, or any other agent capable of specifically binding to the modification of interest. Preferably, the MBD further comprises magnetic beads, streptavidin, or other binding domains for performing an affinity separation step.

[0011] The present disclosure further provides a method for analyzing a population of nucleic acids, at least some of which contain one or more modified cytosine residues. The method includes linking a capture moiety, e.g., biotin, to nucleic acids in the population that serve as templates for amplification, performing an amplification reaction to produce amplification products from the templates, separating the capture moiety-linked templates from the amplification products, assaying the sequence data of the capture moiety-linked templates by bisulfite sequencing, and assaying the sequence data of the amplification products.

[0012] In some embodiments, the capture moiety comprises biotin. In some embodiments, the separating step is carried out by contacting the template with streptavidin beads. In some embodiments, the modified cytosine residue is 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, or 5-carboxylcytosine. In some embodiments, the capture moiety comprises biotin linked to a nucleic acid tag comprising one or more modified residues. In some embodiments, the capture moiety is linked to the nucleic acids in the population by a cleavable linkage. In some embodiments, the cleavable linkage is a photocleavable linkage. In some embodiments, the cleavable linkage comprises a uracil nucleotide.

[0013] The present disclosure further provides methods for analyzing a nucleic acid population that includes nucleic acids with different degrees of 5-methylcytosine. The method includes: (a) contacting a population of nucleic acids with an agent that preferentially binds 5-methylated nucleic acids; (b) separating nucleic acids in a first pool that are bound to the agent from nucleic acids in a second pool that are not bound to the agent, wherein the nucleic acids in the first pool are over-represented in 5-methylcytosine and the nucleic acids in the second pool are under-represented in 5-methylation; (c) linking the nucleic acids in the first pool and / or the second pool to one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool, wherein the nucleic acid tags linked to the nucleic acids in the first pool comprise a capture moiety (e.g., biotin); (d) amplifying the labeled nucleic acids, wherein the nucleic acids and the linked tags are amplified; (e) separating the amplified nucleic acids that have the capture moieties from the amplified nucleic acids that do not have the capture moieties; and (f) assaying sequence data of the separated amplified nucleic acids.

[0014] The present disclosure further provides a method for analyzing a nucleic acid population containing nucleic acids with different degrees of modification, comprising: contacting the nucleic acids in the population with adapters to produce a population of nucleic acids flanked by adapters containing primer binding sites; amplifying the adapter-flanked nucleic acids primed by the primer binding sites; contacting the amplified nucleic acids with an agent that preferentially binds nucleic acids with the modification; separating the nucleic acids of a first pool that are bound to the agent from the nucleic acids of a second pool that are not bound to the agent, wherein the nucleic acids of the first pool are over-represented in the modification and the nucleic acids in the second pool are under-represented in the modification; performing a second amplification step of the nucleic acids in the first and second pools; and assaying the sequence data of the amplified nucleic acids in the first and second pools. The amplification of each pool can occur separately in different reaction vessels. The use of pool-specific tags allows the amplicons to be subsequently pooled and then sequenced.

[0015] The present disclosure further provides a method of analyzing a population of nucleic acids, at least a portion of which nucleic acids comprise one or more modified cytosine residues, the method comprising: contacting the nucleic acid population with adapters comprising a primer binding site that comprises at least one modified cytosine to form adapter-flanked nucleic acids; amplifying the adapter-flanked nucleic acids that are primed from the primer binding site in the adapter flanking the nucleic acids; dividing the amplified nucleic acids into first and second aliquots; assaying sequence data for the nucleic acids of the first aliquot; contacting the nucleic acids of the second aliquot with bisulfite, which converts unmodified cytosines (C) to uracils (U); amplifying nucleic acids resulting from the bisulfite treatment that are primed from the primer binding site flanking the nucleic acids, wherein the U introduced by the bisulfite treatment is converted to T; assaying sequence data for the amplified nucleic acids from the second aliquot; and comparing the sequence data of the nucleic acids in the first and second aliquots to identify which nucleotides in the nucleic acid population were modified cytosines.

[0016] In any of the above methods, the nucleic acid population can be derived from a bodily fluid sample, such as blood, serum, or plasma. In some embodiments, the nucleic acid population is a cell-free nucleic acid population. In some embodiments, the bodily fluid sample is derived from a subject suspected of having cancer.

[0017] In one aspect, the present disclosure provides a method for analyzing a nucleic acid population comprising at least two forms of nucleic acid selected from double-stranded DNA, single-stranded DNA, and single-stranded RNA, each of the at least two forms comprising a plurality of molecules, the method comprising: linking at least one of the forms of nucleic acid to at least one tag nucleic acid to distinguish the forms from each other; amplifying the form of nucleic acid, at least one of which is linked to at least one nucleic acid tag, wherein the nucleic acid and the linked nucleic acid tag are amplified to produce amplified nucleic acids, of which at least one amplified form is tagged; and assaying the sequence data of the amplified nucleic acids, at least some of which are tagged, to obtain sufficient sequence information to decode the tag nucleic acid molecules of the amplified nucleic acids, thereby identifying the forms of nucleic acids in the population and providing the original template of the amplified nucleic acid, whose sequence data is linked to the tag nucleic acid molecule that is assayed. In one embodiment, the method further comprises decoding the tag nucleic acid molecules of the amplified nucleic acids to identify the forms of nucleic acids in the population and providing the original template of the amplified nucleic acid, whose sequence data is linked to the tag nucleic acid molecule that is assayed. In another embodiment, the method further comprises enriching at least one of the forms relative to one or more of the other forms. In another embodiment, at least 70% of the molecules of each form of nucleic acid in the population are amplified. In another embodiment, at least three forms of nucleic acid are present in the population, and at least two forms are linked to different tag nucleic acid forms that distinguish each of the three forms from one another. In another embodiment, each of the at least three forms of nucleic acid in the population is linked to a different tag. In another embodiment, each molecule of the same form is linked to a tag containing the same tag information. In another embodiment, molecules of the same form are linked to different types of tags. In another embodiment, the method further comprises subjecting the population to reverse transcription using tagged primers, and the tagged primers are incorporated into cDNA generated from the RNA in the population. In another embodiment, the reverse transcription is sequence-specific. In another embodiment, the reverse transcription is random.In another embodiment, the method further comprises degrading RNA duplexed with the cDNA. In another embodiment, the method further comprises separating single-stranded DNA from double-stranded DNA and ligating a nucleic acid tag to the double-stranded DNA. In another embodiment, the single-stranded DNA is separated by hybridization with one or more capture probes. In another embodiment, the method further comprises circularizing the single-stranded DNA using circligase and ligating a nucleic acid tag to the double-stranded DNA. In another embodiment, the method further comprises pooling tagged nucleic acids comprising different forms of nucleic acid prior to the assaying step. In another embodiment, the nucleic acid population is derived from a bodily fluid sample. In another embodiment, the bodily fluid sample is blood, serum, or plasma. In another embodiment, the nucleic acid population is a cell-free nucleic acid population. In another embodiment, the bodily fluid sample is derived from a subject suspected of having cancer. In another embodiment, the sequence data indicates the presence of somatic or germline mutations. In another embodiment, the sequence data indicates the presence of copy number variations. In another embodiment, the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions. In another embodiment, the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions.

[0018] In another aspect, provided herein is a method of analyzing a nucleic acid population containing nucleic acids with different degrees of modification, the method comprising: contacting the nucleic acid population with an agent that preferentially binds nucleic acids with the modification; separating nucleic acids in a first pool that are bound to the agent from nucleic acids in a second pool that are not bound to the agent, wherein the nucleic acids in the first pool are over-represented with the modification and the nucleic acids in the second pool are under-represented with the modification; linking the nucleic acids in the first pool and / or the second pool to one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool to produce a population of tagged nucleic acids; amplifying the labeled nucleic acids, wherein the nucleic acids and the linked tags are amplified; and assaying sequence data of the amplified nucleic acids and the linked tags, wherein the assaying provides sequence data for decoding the tags and reveals whether the nucleic acids whose sequence data was assayed were amplified from templates in the first or second pool. In one embodiment, the method includes decoding the tag to determine whether the nucleic acid for which sequence data was assayed was amplified from a template in the first or second pool. In another embodiment, the modification is binding of the nucleic acid to a protein. In another embodiment, the protein is a histone or a transcription factor. In another embodiment, the modification is a post-replication modification to a nucleotide. In another embodiment, the post-replication modification is 5-methyl-cytosine, and the extent of binding of the agent to the nucleic acid increases with the amount of 5-methyl-cytosine in the nucleic acid. In another embodiment, the post-replication modification is 5-hydroxymethyl-cytosine, and the extent of binding of the agent to the nucleic acid increases with the amount of 5-hydroxymethyl-cytosine in the nucleic acid. In another embodiment, the post-replication modification is 5-formyl-cytosine or 5-carboxyl-cytosine, and the extent of binding of the agent increases with the amount of 5-formyl-cytosine or 5-carboxyl-cytosine in the nucleic acid.In another embodiment, the method further comprises washing the nucleic acids bound to the agent and recovering the wash as a third pool containing nucleic acids with an intermediate degree of post-replication modification relative to the first and second pools. In another embodiment, the method comprises pooling the tagged nucleic acids from the first and second pools prior to the assaying step. In another embodiment, the agent is a 5-methyl binding domain magnetic bead. In another embodiment, the nucleic acid population is derived from a bodily fluid sample. In another embodiment, the bodily fluid sample is blood, serum, or plasma. In another embodiment, the nucleic acid population is a cell-free nucleic acid population. In another embodiment, the bodily fluid sample is derived from a subject suspected of having cancer. In another embodiment, the sequence data indicates the presence of somatic or germline mutations. In another embodiment, the sequence data indicates the presence of copy number variations. In another embodiment, the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions.

[0019] In another aspect, provided herein is a method for analyzing a nucleic acid population, at least some of the nucleic acids of which contain one or more modified cytosine residues, comprising the steps of: linking a capture moiety to nucleic acids in the population that serve as amplification templates; performing an amplification reaction to produce amplification products from the templates; separating the templates linked to the capture tags from the amplification products; assaying the sequence data of the templates linked to the capture tags by bisulfite sequencing; and assaying the sequence data of the amplification products. In one embodiment, the capture moiety comprises biotin. In another embodiment, the separating step is performed by contacting the templates with streptavidin beads. In another embodiment, the modified cytosine residue is 5-methyl-cytosine, 5-hydroxymethylcytosine, 5-formylcytosine, or 5-carboxylcytosine. In another embodiment, the capture moiety comprises biotin linked to a nucleic acid tag containing one or more modified residues. In another embodiment, the capture moiety is linked to the nucleic acids in the population by a cleavable linkage. In another embodiment, the cleavable linkage is a photocleavable linkage. In another embodiment, the cleavable linkage comprises a uracil nucleotide. In another embodiment, the nucleic acid population is derived from a bodily fluid sample. In another embodiment, the bodily fluid sample is blood, serum, or plasma. In another embodiment, the nucleic acid population is a cell-free nucleic acid population. In another embodiment, the bodily fluid sample is derived from a subject suspected of having cancer. In another embodiment, the sequence data indicates the presence of somatic or germline mutations. In another embodiment, the sequence data indicates the presence of copy number variations. In another embodiment, the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions.

[0020] In another aspect, provided herein is a method for analyzing a nucleic acid population comprising nucleic acids having different degrees of 5-methylation, the method comprising: contacting the nucleic acid population with an agent that preferentially binds 5-methylated nucleic acids; separating nucleic acids in a first pool that are bound to the agent from nucleic acids in a second pool that are not bound to the agent, wherein the nucleic acids in the first pool are over-represented in 5-methylation and the nucleic acids in the second pool are under-represented in 5-methylation; linking the nucleic acids in the first pool and / or the second pool to one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool, wherein the nucleic acid tags linked to the nucleic acids in the first pool comprise a capture moiety (e.g., biotin); amplifying the labeled nucleic acids, wherein the nucleic acids and the linked tags are amplified; separating the amplified nucleic acids with the capture moieties from the amplified nucleic acids that do not have the capture moieties; and assaying sequence data of the separated amplified nucleic acids.

[0021] In another aspect, provided herein is a method for analyzing a population of nucleic acids containing nucleic acids with different degrees of modification, the method comprising: contacting the nucleic acids in the population with adapters to produce a population of nucleic acids flanked by adapters containing primer binding sites; amplifying the adapter-flanked nucleic acids primed from the primer binding sites; contacting the amplified nucleic acids with an agent that preferentially binds nucleic acids with the modification; separating nucleic acids in a first pool that are bound to the agent from nucleic acids in a second pool that are not bound to the agent, wherein the nucleic acids in the first pool are over-represented with the modification and the nucleic acids in the second pool are under-represented with the modification; performing parallel amplification of the tagged nucleic acids in the first and second pools; and assaying sequence data of the amplified nucleic acids in the first and second pools. In another embodiment, the adapters are hairpin adapters.

[0022] In another aspect, provided herein is a method of analyzing a population of nucleic acids, at least a portion of which nucleic acids comprise one or more modified cytosine residues, the method comprising: contacting the nucleic acid population with adapters comprising primer binding sites that comprise modified cytosines to form adapter-flanked nucleic acids; amplifying the adapter-flanked nucleic acids that are primed from the primer binding sites in the adapters that are flanked by the nucleic acids; dividing the amplified nucleic acids into first and second aliquots; assaying sequence data for the nucleic acids in the first aliquot; contacting the nucleic acids in the second aliquot with bisulfite, which converts unmodified C to U; amplifying the nucleic acids resulting from bisulfite treatment that are primed from the primer binding sites that are flanked by the nucleic acids, wherein U introduced by bisulfite treatment is converted to T; and assaying the sequence data of the amplified nucleic acids from the second aliquot; comparing the sequence data of the nucleic acids in the first and second aliquots to result in sequence data that can be used to identify which nucleotides in the nucleic acid population were modified cytosines. In one embodiment, the method includes comparing sequence data of the nucleic acids in the first and second aliquots to identify which nucleotides in the nucleic acid population were modified cytosines. In another embodiment, the adaptor is a hairpin adaptor.

[0023] In another aspect, provided herein is a method comprising physically fractionating DNA molecules from a human sample to create two or more aliquots, applying differential molecular tags and NGS-enabling adapters to each of the two or more aliquots to create molecularly tagged aliquots, and assaying the molecularly tagged aliquots on an NGS instrument to generate sequence data for deconvoluting the sample into differentially resolved molecules. In one embodiment, the method further comprises analyzing the sequence data by deconvoluting the sample into differentially resolved molecules. In another embodiment, the DNA molecules are derived from extracted plasma. In another embodiment, the physically fractionating step comprises fractionating the molecules based on varying degrees of methylation. In another embodiment, the varying degrees of methylation include hypermethylated and hypomethylated. In another embodiment, the physically fractionating step comprises fractionating using methyl-binding domain protein ("MBD")-beads to stratify the varying degrees of methylation. In another embodiment, the differential molecular tags are different sets of molecular tags corresponding to the MBD-dividends. In another embodiment, the physical fractionation comprises separating DNA molecules using immunoprecipitation. In another embodiment, the method further comprises recombining two or more of the generated molecularly tagged fractions. In another embodiment, the method further comprises enriching the recombined molecularly tagged fractions or groups. In another embodiment, the one or more characteristics is methylation. In another embodiment, the fractionation comprises separating methylated nucleic acids from unmethylated nucleic acids using a protein containing a methyl-binding domain to generate groups of nucleic acid molecules containing various degrees of methylation. In another embodiment, one of the groups contains hypermethylated DNA. In another embodiment, at least one group is characterized by the degree of methylation. In another embodiment, the fractionation comprises isolating nucleic acids to which the protein is bound. In another embodiment, the isolating comprises immunoprecipitation.

[0024] In another aspect, provided herein is a method for molecular tag identification of MBD-bead fractionated libraries by NGS, comprising: physically fractionating an extracted DNA sample using a methyl-binding domain protein-bead purification kit while retaining all eluates for downstream processing; applying differential molecular tags and NGS-enabling adapter sequences to each fraction or group in parallel; recombining all molecularly tagged fractions or groups and subsequently amplifying them using adapter-specific DNA primer sequences; (d) enriching / hybridizing the recombined and amplified total library while targeting genomic regions of interest; reamplifying the enriched total DNA library while adding sample tags; and pooling the different samples and assaying them in multiplex on an NGS instrument, wherein the NGS sequence data generated by the instrument provides the sequences of the molecular tags that are used to identify unique molecules and the sequence data for deconvolution of the samples into differentially MBD-fractionated molecules. In one embodiment, the method includes analyzing the NGS data using molecular tags that are used to identify unique molecules, as well as deconvolving the sample into differentially segmented molecules. In another embodiment, the fractionation includes physical fractionation. In another embodiment, the population of nucleic acid molecules is divided based on one or more features selected from the group consisting of methylation status, glycosylation status, histone modification, length, and start / stop location. In another embodiment, the method further includes pooling the nucleic acid molecules. In another embodiment, the fractionation includes fractionation based on differences in mononucleosome profiles. In another embodiment, the fractionation can produce a different mononucleosome profile for at least one group of nucleic acid molecules compared to normal. In another embodiment, the method further includes fractionating at least one group of nucleic acid molecules based on different features. In another embodiment, the analyzing step includes comparing a first feature corresponding to a first group of nucleic acid molecules with a second feature corresponding to a second group of nucleic acid molecules at one or more loci.In another embodiment, the nucleic acid molecule is circulating tumor DNA. In another embodiment, the nucleic acid molecule is cell-free DNA ("cfDNA"). In another embodiment, the tag is used to distinguish different molecules in the same sample. In another embodiment, one or more features are cancer markers.

[0025] In another aspect, a method is provided herein that includes providing a population of nucleic acid molecules obtained from a subject's body sample; fractionating the population of nucleic acid molecules based on one or more features to create a plurality of groups of nucleic acid molecules; differentially tagging the nucleic acid molecules in the plurality of groups to distinguish the nucleic acid molecules in each of the plurality of groups from one another based on one or more features; and sequencing the plurality of groups of nucleic acid molecules to generate sequence read data containing sufficient data to generate relative information about nucleosome positioning, nucleosome modification, or binding DNA-protein interactions for each of the plurality of groups of nucleic acid molecules. In one embodiment, the method further includes analyzing the sequence read data to generate relative information about nucleosome positioning, nucleosome modification, or binding DNA-protein interactions for each of the plurality of groups of nucleic acid molecules. In another embodiment, the method further includes using a trained classifier to classify the subject based on one or more features. In another embodiment, the one or more features include quantitative features of the mapped read data. In another embodiment, the fractionation includes physical fractionation. In another embodiment, the method further includes pooling the nucleic acid molecules. In another embodiment, the fractionation comprises fractionating based on differences in mononucleosome profiles. In another embodiment, the fractionation can produce a different mononucleosome profile for at least one group of nucleic acid molecules compared to a normal. In another embodiment, the method further comprises fractionating at least one group of nucleic acid molecules based on different characteristics. In another embodiment, the analyzing step comprises comparing a first characteristic corresponding to a first group of nucleic acid molecules with a second characteristic corresponding to a second group of nucleic acid molecules at one or more loci. In another embodiment, the analyzing step comprises analyzing one characteristic of the one or more characteristics in the group relative to a normal sample at one or more loci.In another embodiment, the one or more features are selected from the group consisting of: a base call frequency at a base position on the reference sequence; the number of molecules mapping to a base or sequence on the reference sequence; the number of molecules with a start site mapping to a base position on the reference sequence; the number of molecules with a stop site mapping to a base position on the reference sequence; and the length of a molecule mapping to a locus on the reference sequence. In another embodiment, the method further comprises using a trained classifier to classify the subject based on the one or more features. In another embodiment, the trained classifier classifies the one or more features as associated with a tissue in the subject. In another embodiment, the trained classifier classifies the one or more features as associated with a type of cancer in the subject. In another embodiment, the one or more features indicate gene expression or a disease state. In another embodiment, the nucleic acid molecule is circulating tumor DNA. In another embodiment, the nucleic acid molecule is cell-free DNA ("cfDNA"). In another embodiment, the tag is used to distinguish different molecules in the same sample. In another embodiment, the one or more features are cancer markers.

[0026] In another aspect, provided herein is a method comprising: providing a population of nucleic acid molecules obtained from a subject's bodily sample; fractionating the population of nucleic acid molecules based on methylation status to create a plurality of groups of nucleic acid molecules; differentially tagging the nucleic acid molecules in the plurality of groups to distinguish the nucleic acid molecules in each of the plurality of groups from one another based on one or more features; sequencing the plurality of groups of nucleic acid molecules to generate sequence read data; and analyzing the sequence read data to detect one or more features in one of the plurality of groups of nucleic acid molecules, wherein the one or more features indicate nucleosome positioning, nucleosome modification, or DNA-protein interaction. In another embodiment, the method further comprises using a trained classifier to classify the subject based on the one or more features. In another embodiment, the one or more features comprise quantitative features of the mapped read data. In another embodiment, the fractionation comprises physical fractionation. In another embodiment, the method further comprises pooling the nucleic acid molecules. In another embodiment, the fractionation comprises fractionation based on differences in mononucleosome profiles. In another embodiment, the fractionation can produce a different mononucleosome profile for at least one group of nucleic acid molecules compared to a normal. In another embodiment, the method further comprises fractionating at least one group of nucleic acid molecules based on different features. In another embodiment, the analyzing step comprises comparing a first feature corresponding to a first group of nucleic acid molecules with a second feature corresponding to a second group of nucleic acid molecules at one or more loci. In another embodiment, the analyzing step comprises analyzing one of the one or more features in the group relative to a normal sample at one or more loci. In another embodiment, the one or more features are selected from the group consisting of: a base call frequency at a base position on the reference sequence; a number of molecules mapping to a base or sequence on the reference sequence; a number of molecules having a start site mapping to a base position on the reference sequence; a number of molecules having a stop site mapping to a base position on the reference sequence; and a length of a molecule mapping to a locus on the reference sequence.In another embodiment, the method further comprises using the trained classifier to classify the subject based on the one or more features. In another embodiment, the trained classifier classifies the one or more features as associated with a tissue in the subject. In another embodiment, the trained classifier classifies the one or more features as associated with a type of cancer in the subject. In another embodiment, the one or more features indicate gene expression or a disease state. In another embodiment, the nucleic acid molecule is circulating tumor DNA. In another embodiment, the nucleic acid molecule is cell-free DNA ("cfDNA"). In another embodiment, the tag is used to distinguish different molecules in the same sample. In another embodiment, the one or more features are cancer markers.

[0027] In another aspect, the present disclosure provides a method comprising: providing a collection of nucleic acid molecules obtained from a subject's body sample; fractionating the collection of nucleic acid molecules to produce a plurality of groups of nucleic acid molecules, comprising protein-bound cell-free nucleic acid; differentially tagging the nucleic acid molecules in the plurality of groups, so as to distinguish the nucleic acid molecules in each of the plurality of groups from each other based on one or more features; sequencing the plurality of groups of nucleic acid molecules to generate sequence reading data, wherein the sequence information obtained is sufficient to map the sequence reading data to one or more loci on a reference sequence, and analyze the sequence reading data to detect one or more features in one of the plurality of groups of nucleic acid molecules, wherein the one or more features represent nucleosome positioning, nucleosome modification or DNA-protein interaction.In one embodiment, the method further comprises: mapping the sequence reading data to one or more loci on a reference sequence, and analyzing the sequence reading data to detect one or more features in one of the plurality of groups of nucleic acid molecules, wherein the one or more features represent nucleosome positioning, nucleosome modification or DNA-protein interaction. In another embodiment, the method further comprises classifying the subjects based on one or more features using the trained classifier. In another embodiment, the one or more features comprise quantitative features of the mapped read data. In another embodiment, the fractionation comprises physical fractionation. In another embodiment, the population of nucleic acid molecules is divided based on one or more features selected from the group consisting of methylation state, glycosylation state, histone modification, length, and start / stop position. In another embodiment, the method further comprises pooling the nucleic acid molecules. In another embodiment, the one or more features are methylation. In another embodiment, the fractionation comprises separating methylated nucleic acids from unmethylated nucleic acids using a protein comprising a methyl-binding domain to create groups of nucleic acid molecules comprising various degrees of methylation. In another embodiment, one of the groups comprises hypermethylated DNA. In another embodiment, at least one group is characterized by the degree of methylation.In another embodiment, the fractionation comprises separating single-stranded and / or double-stranded DNA molecules. In another embodiment, double-stranded DNA molecules are separated using hairpin adapters. In another embodiment, the fractionation comprises isolating protein-bound nucleic acids. In another embodiment, the fractionation comprises fractionating based on differences in mononucleosome profiles. In another embodiment, the fractionation can produce a distinct mononucleosome profile for at least one group of nucleic acid molecules compared to a normal. In another embodiment, the isolating comprises immunoprecipitation. In another embodiment, the method further comprises fractionating at least one group of nucleic acid molecules based on distinct features. In another embodiment, the analyzing step comprises comparing a first feature corresponding to a first group of nucleic acid molecules with a second feature corresponding to a second group of nucleic acid molecules at one or more loci. In another embodiment, the analyzing step comprises analyzing one feature of the one or more features in the group relative to a normal sample at one or more loci. In another embodiment, the one or more features are selected from the group consisting of: a base call frequency at a base position on the reference sequence; the number of molecules mapping to a base or sequence on the reference sequence; the number of molecules with a start site mapping to a base position on the reference sequence; the number of molecules with a stop site mapping to a base position on the reference sequence; and the length of the molecule mapping to a locus on the reference sequence. In another embodiment, the method further comprises using a trained classifier to classify the subject based on the one or more features. In another embodiment, the trained classifier classifies the one or more features as associated with a tissue in the subject. In another embodiment, the trained classifier classifies the one or more features as associated with a type of cancer in the subject. In another embodiment, the one or more features indicate gene expression or a disease state. In another embodiment, the nucleic acid molecule is circulating tumor DNA. In another embodiment, the nucleic acid molecule is cell-free DNA ("cfDNA"). In another embodiment, the tags are used to distinguish different molecules in the same sample.

[0028] In another aspect, the present disclosure provides a method comprising: providing a collection of nucleic acid molecules obtained from a subject's body sample; dividing the collection of nucleic acid molecules according to one or more characteristics to create a plurality of groups of nucleic acid molecules; differentially tagging the nucleic acid molecules in the plurality of groups to distinguish the nucleic acid molecules in each of the plurality of groups according to one or more characteristics; and sequencing the plurality of groups of nucleic acid molecules to generate sequence read data, wherein the sequence information obtained is sufficient to map the sequence read data to one or more loci on a reference sequence, and analyze the sequence read data to detect one or more features in one of the plurality of groups of nucleic acid molecules, and the one or more features are not detectable in the pool of sequence read data from the plurality of groups.In one embodiment, the method further comprises: mapping the sequence read data to one or more loci on a reference sequence, and analyzing the sequence read data to detect one or more features in one of the plurality of groups of nucleic acid molecules, and the one or more features are not detectable in the pool of sequence read data from the plurality of groups.In another embodiment, the division comprises physical division.

[0029] In another aspect, provided herein is a method comprising: providing a population of nucleic acid molecules obtained from a subject's bodily sample; fractionating the population of nucleic acid molecules based on one or more characteristics to create a plurality of groups of nucleic acid molecules, wherein each nucleic acid molecule of the plurality of groups comprises a distinct identifier; pooling the plurality of groups of nucleic acid molecules; sequencing the pooled plurality of groups of nucleic acid molecules to generate a plurality of sets of sequence read data; and fractionating the sequence read data based on the identifier.

[0030] In another aspect, provided herein is a composition comprising a pool of nucleic acid molecules, including differentially tagged nucleic acid molecules, wherein the pool comprises a plurality of sets of nucleic acid molecules that are differentially tagged based on one or more characteristics selected from the group consisting of: methylation state, glycosylation state, histone modification, length, and start / stop position, and the pool is derived from a biological sample. In one embodiment, the plurality of sets is any of 2, 3, 4, 5, or more than 5.

[0031] In another aspect, the present disclosure provides a method comprising: dividing a population of nucleic acid molecules into a plurality of groups that contain nucleic acids with different characteristics; tagging the nucleic acids in each of the plurality of groups with a set of tags that distinguish the nucleic acids in each of the plurality of groups to produce a population of tagged nucleic acids, wherein each tagged nucleic acid comprises one or more tags; sequencing the population of tagged nucleic acids to generate sequence read data; using one or more tags to group the sequence read data of each group; and analyzing the sequence read data to detect a signal in at least one of the groups relative to a normal sample or a classifier.In one embodiment, the method further comprises normalizing the signal in at least one of the groups relative to another group or the whole genome sequence.

[0032] In another aspect, the present disclosure provides a method comprising the steps of: providing a population of cell-free DNA from a biological sample; fractionating the population of cell-free DNA based on features that are present at different levels in cell-free DNA derived from cancerous cells compared to non-cancerous cells, thereby creating subpopulations of cell-free DNA; amplifying at least one of the subpopulations of cell-free DNA; and sequencing at least one of the amplified subpopulations of cell-free DNA. In one embodiment, the features are the methylation level of cell-free DNA, the glycosylation level of cell-free DNA, the length of cell-free DNA fragments, or the presence of single-strand breaks in cell-free DNA.

[0033] In another aspect, provided herein is a method that includes providing a population of cell-free DNA from a biological sample; fractionating the population of cell-free DNA based on methylation levels of the cell-free DNA, thereby creating subsets of cell-free DNA; amplifying at least one of the subsets of cell-free DNA; and sequencing at least one of the amplified subsets of cell-free DNA.

[0034] In another aspect, provided herein is a method for determining the methylation status of cell-free DNA, the method comprising: providing a population of cell-free DNA from a biological sample; fractionating the population of cell-free DNA based on methylation levels of the cell-free DNA, thereby creating subsets of cell-free DNA; sequencing at least one subset of the cell-free DNA, thereby generating sequence reads; and assigning a methylation state to each cell-free DNA according to the subset from which the corresponding sequence reads originate.

[0035] In another aspect, provided herein is a method for classifying a subject, the method comprising: providing a population of cell-free DNA from a biological sample from a subject; fractionating the population of cell-free DNA based on the methylation level of the cell-free DNA, thereby creating subpopulations of cell-free DNA; sequencing the subpopulations of cell-free DNA, thereby generating sequence read data; and classifying the subject using a trained classifier according to which sequence read data occurs in which subpopulation. In another embodiment, the population of cell-free DNA is fractionated by one or more features that provide signal differences between healthy and diseased states. In another embodiment, the population of cell-free DNA is fractionated based on the methylation level of the cell-free DNA. In another embodiment, determining the fragmentation pattern of the cell-free DNA further comprises analyzing the number of sequence read data that are mapped to each base position in the reference genome. In another embodiment, the method further comprises determining the fragmentation pattern of the cell-free DNA in each subpopulation by analyzing the number of sequence read data that are mapped to each base position in the reference genome.

[0036] In another aspect, provided herein is a method for analyzing the fragmentation pattern of cell-free DNA, comprising: providing a population of cell-free DNA from a biological sample; fractionating the population of cell-free DNA to thereby create subpopulations of cell-free DNA; sequencing at least one subpopulation of cell-free DNA to thereby generate sequence read data; aligning the sequence read data to a reference genome; and determining the fragmentation pattern of cell-free DNA in each subpopulation by analyzing any number of the following: the length of each sequence read mapped to each base position in the reference genome; the number of sequence reads mapped to each base position in the reference genome as a function of the length of the sequence read data; the number of sequence reads starting at each base position in the reference genome; or the number of sequence reads ending at each base position in the reference genome. In another embodiment, the one or more characteristics comprise a chemical modification selected from the group consisting of methylation, hydroxymethylation, formylation, acetylation, and glycosylation.

[0037] In any of the methods described herein, the DNA:bead ratio is 1:100.

[0038] In any of the methods described herein, the DNA:bead ratio is 1:50.

[0039] In any of the methods described herein, the DNA:bead ratio is 1:20.

[0040] In one aspect, provided herein is the use of physical fractionation based on the degree of DNA methylation in the analysis of circulating tumor DNA (ctDNA) to determine gene expression or disease state.

[0041] In one aspect, provided herein is the use of features that result in a difference in signal between normal and diseased states to physically separate ctDNA during analysis of the ctDNA.

[0042] In one aspect, provided herein is the use of features that result in a difference in signal between normal and diseased states to physically separate ctDNA.

[0043] In one aspect, provided herein is the use of features that result in a difference in signal between normal and diseased states to physically separate ctDNA prior to sequencing and optional downstream analysis.

[0044] In one aspect, the present invention provides the use of features that cause signal differences between normal and diseased states to physically divide ctDNA for differential labeling / tagging.In one embodiment, the differential fragmentation pattern indicates gene expression or disease state.In another embodiment, the differential fragmentation pattern is characterized by one or more differences from normal, selected from the group consisting of: the length of each sequence read that maps to each base position in the reference genome; the number of sequence reads that map to each base position in the reference genome as a function of the length of the sequence read; the number of sequence reads that start at each base position in the reference genome; and the number of sequence reads that end at each base position in the reference genome.

[0045] In one aspect, the present disclosure provides the use of fractionation based on differential fragmentation patterns in analyzing ctDNA.In one embodiment, the differential fragmentation patterns indicate gene expression or disease state.In another embodiment, the differential fragmentation patterns are characterized by one or more differences from normal, selected from the group consisting of: the length of each sequence read data that is mapped to each base position in reference genome; the number of sequence read data that is mapped to each base position in reference genome as a function of the length of sequence read data; the number of sequence read data that starts at each base position in reference genome; and the number of sequence read data that ends at each base position in reference genome.

[0046] In one aspect, the present disclosure provides the use of differential fragmentation patterns for dividing ctDNA.In one embodiment, the differential fragmentation patterns indicate gene expression or disease state.In another embodiment, the differential fragmentation patterns are characterized by one or more differences from normal, selected from the group consisting of: the length of each sequence read data that is mapped to each base position in reference genome; the number of sequence read data that is mapped to each base position in reference genome as a function of the length of sequence read data; the number of sequence read data that starts at each base position in reference genome; and the number of sequence read data that ends at each base position in reference genome.

[0047] In one aspect, provided herein is the use of differential fragmentation pattern for dividing ctDNA before sequencing and optional downstream analysis.In one embodiment, differential fragmentation pattern indicates gene expression or disease state.In another embodiment, differential fragmentation pattern is characterized by one or more differences from normal, selected from the group consisting of: the length of each sequence read data that is mapped to each base position in reference genome; the number of sequence read data that is mapped to each base position in reference genome as a function of the length of sequence read data; the number of sequence read data that starts at each base position in reference genome; and the number of sequence read data that ends at each base position in reference genome.

[0048] In one aspect, provided herein is the use of differential fragmentation patterns to partition ctDNA for differential labeling / tagging.

[0049] In one aspect, provided herein is the use of differential molecular tagging of DNA molecules separated by molecular binding domain (MBD)-beads to stratify into different degrees of DNA methylation, which are then quantified by next generation sequencing (NGS).

[0050] In one aspect, the present disclosure provides a method for analyzing a nucleic acid population comprising at least two forms of nucleic acid selected from double-stranded DNA, single-stranded DNA, and single-stranded RNA, wherein each of the at least two forms comprises a plurality of molecules, and the method comprises: linking at least one of the nucleic acid forms to at least one tag nucleic acid to distinguish the forms from each other; amplifying the nucleic acid form, at least one of which is linked to at least one nucleic acid tag, wherein the nucleic acid and the linked nucleic acid tag are amplified to produce amplified nucleic acids, of which at least one amplified form is tagged; and sequencing the plurality of amplified nucleic acids linked to the tag, wherein the sequence data is sufficient to be deciphered to identify the form of nucleic acid in the population before linking to the at least one tag.In one embodiment, the molecular tag comprises one or more nucleic acid barcodes.In another embodiment, in a pool of tagged nucleic acid molecules, the combination of any two barcodes in a set has a different combined sequence from the combination of any two barcodes in any other set.

[0051] In another aspect, the present disclosure provides a pool of tagged nucleic acid molecules, wherein each nucleic acid molecule in the pool comprises a molecular tag selected from one of a plurality of tag sets, each tag set comprises a plurality of different tags, and the tag in any one set is distinct from the tag in any other set, and each tag set (i) represents the characteristic of the molecule to which it is attached or the characteristic of the parent molecule from which it originates, and (ii) contains information that, alone or in combination with the information from the molecule to which it is attached, uniquely distinguishes the molecule to which it is attached from other molecules that are tagged with tags from the same tag set.In one embodiment, the molecular tag comprises two nucleic acid barcodes attached to opposite ends of the molecule.In another embodiment, the barcodes are between 10 and 30 nucleotides in length.

[0052] In another aspect, a system is provided that includes a nucleic acid sequencer; a digital processing device that includes at least one processor, an operating system configured to execute executable instructions, and a memory; and a data link communicatively connecting the nucleic acid sequencer and the digital processing device, wherein the digital processing device is configured to receive from the nucleic acid sequencer via the data link sequence data of at least some of the tagged amplified nucleic acids, the sequence data being used to identify at least one of the forms of nucleic acids and to distinguish the forms from one another. The present disclosure provides a system further comprising executable instructions for creating an application comprising: a software module that is created by linking at least one tagged nucleic acid to at least one form of nucleic acid, at least one of which is linked to at least one nucleic acid tag, and amplifying the nucleic acid and the linked nucleic acid tag to produce amplified nucleic acids, at least one of which is tagged; and a software module that obtains sufficient sequence information to decode the tagged nucleic acid molecules of the amplified nucleic acids, thereby revealing the form of nucleic acids in the population and providing the original template of the amplified nucleic acid linked to the tag nucleic acid molecule whose sequence data is assayed. In one embodiment, the application further comprises a software module that decodes the tagged nucleic acid molecules of the amplified nucleic acids, revealing the form of nucleic acids in the population and providing the original template of the amplified nucleic acid linked to the tag nucleic acid molecule whose sequence data is assayed. In another embodiment, the application further comprises a software module that transmits the results of the assay via a communication network.

[0053] In another aspect, the present disclosure provides a system comprising: a next-generation sequencing (NGS) instrument; a digital processing device comprising at least one processor, an operating system configured to execute executable instructions, and a memory; and a data link communicatively connecting the NGS instrument and the digital processing device, wherein the digital processing device further comprises executable instructions for creating an application comprising: a software module for receiving sequence data from the NGS instrument through the data link, wherein the sequence data is generated by physically dividing DNA molecules from a human sample to produce two or more aliquots, applying differential molecular tags and NGS-enabling adapters to each of the two or more aliquots to produce molecularly tagged aliquots, and assaying the molecularly tagged aliquots with the NGS instrument; a software module for generating sequence data for deconvoluting the sample into differentially divided molecules; and a software module for analyzing the sequence data by deconvoluting the sample into differentially divided molecules.In one embodiment, the application further comprises a software module for transmitting the results of the assay via a communication network.

[0054] In another aspect, a system is provided that includes a next-generation sequencing (NGS) instrument; a digital processing device that includes at least one processor, an operating system configured to execute executable instructions, and a memory; and a data link communicatively connecting the NGS instrument and the digital processing device, wherein the digital processing device includes a software module configured to receive sequence data from the NGS instrument via the data link for an application for molecular tag identification of a library fractionated by MBD-beads, the software module comprising: a methyl-binding domain protein-bead purification kit; a methyl-binding domain protein-bead purification kit that performs parallel application of differential molecular tags and NGS-enabling adapter sequences to each fraction or group while retaining all eluate for downstream processing; and a data link communicatively connecting the NGS instrument and the digital processing device; wherein the digital processing device includes a software module configured to receive sequence data from the NGS instrument via the data link for application of differential molecular tags and NGS-enabling adapter sequences to each fraction or group while retaining all eluate for downstream processing. The present invention provides a system further comprising instructions executable by at least one processor to create an application comprising: a software module that generates NGS sequence data generated by the steps of: performing enrichment / hybridization of the recombined and amplified total library while targeting genomic regions of interest; adding sample tags and reamplifying the enriched total DNA library; pooling the different samples; and assaying them in multiplex on an NGS instrument, where the NGS sequence data generated by the instrument provides sequence data for deconvoluting the sample into differentially divided molecules and molecular tag sequences that are used to identify unique molecules; and a software module configured to perform analysis of the sequence data by using the molecular tags to identify unique molecules and deconvoluting the sample into differentially divided molecules. In one embodiment, the application further comprises a software module configured to transmit the results of the analysis via a communication network.

[0055] The summary provided above is an exemplary list of embodiments and is not intended to be an exhaustive list of embodiments.

[0056] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In certain embodiments, for example, the following items are provided: (Item 1) 1. A method for analyzing a population of nucleic acids comprising at least two forms of nucleic acid selected from double-stranded DNA, single-stranded DNA, and single-stranded RNA, each of the at least two forms comprising a plurality of molecules, the method comprising: (a) linking at least one of said forms of nucleic acid to at least one tag nucleic acid to distinguish said forms from one another; (b) amplifying said forms of nucleic acid, at least one of which is linked to at least one nucleic acid tag, wherein said nucleic acid and linked nucleic acid tag are amplified to produce amplified nucleic acids, of which those amplified from said at least one form are tagged; (c) assaying sequence data of the amplified nucleic acids tagged at least in part, wherein said assaying provides sufficient sequence information to decode the tag nucleic acid molecules of the amplified nucleic acids, revealing the morphology of nucleic acids in the population, and providing original templates of the amplified nucleic acids linked to the tag nucleic acid molecules whose sequence data was assayed; A method comprising: (Item 2) 2. The method of claim 1, further comprising decoding the tag nucleic acid molecules of the amplified nucleic acids to reveal the form of nucleic acids in the population and provide original templates of the amplified nucleic acids linked to the tag nucleic acid molecules whose sequence data was assayed. (Item 3) 3. The method of claim 1 or 2, further comprising enriching at least one of said forms relative to one or more of the other forms. (Item 4) 3. The method of claim 1 or 2, wherein at least 70% of the molecules of each form of nucleic acid in the population are amplified in step (b). (Item 5) 3. The method of claim 1 or 2, wherein at least three forms of nucleic acid are present in the population, and at least two of the forms are linked to different tag nucleic acid forms that distinguish each of the three forms from one another. (Item 6) 6. The method of claim 5, wherein each of the at least three forms of nucleic acid in the population is linked to a different tag. (Item 7) 3. The method of claim 1 or 2, wherein each molecule of the same form is linked to a tag that contains the same identifying tag. (Item 8) 3. The method according to item 1 or 2, wherein molecules of the same form are linked to tags of different types. (Item 9) 3. The method of claim 1 or 2, wherein step (a) comprises subjecting the population to reverse transcription using tagged primers, wherein the tagged primers are incorporated into cDNA made from RNA in the population. (Item 10) 10. The method of claim 9, wherein the reverse transcription is sequence-specific. (Item 11) 10. The method of claim 9, wherein the reverse transcription is random. (Item 12) 10. The method according to item 9, further comprising the step of degrading RNA that forms a duplex with the cDNA. (Item 13) 6. The method of claim 5, further comprising separating single-stranded DNA from double-stranded DNA and ligating a nucleic acid tag to the double-stranded DNA. (Item 14) 14. The method of claim 13, wherein the single-stranded DNA is separated by hybridization with one or more capture probes. (Item 15) 6. The method of claim 5, further comprising the steps of circularizing the single-stranded DNA using circligase and ligating a nucleic acid tag to the double-stranded DNA. (Item 16) 2. The method of claim 1, further comprising pooling tagged nucleic acids comprising different forms of nucleic acid prior to the assaying step. (Item 17) 17. The method of items 1 to 16, wherein the population of nucleic acids is derived from a body fluid sample. (Item 18) 18. The method of claim 17, wherein the body fluid sample is blood, serum, or plasma. (Item 19) 3. The method of claim 1 or 2, wherein the population of nucleic acids is a cell-free population of nucleic acids. (Item 20) 20. The method of claim 18, wherein the body fluid sample is derived from a subject suspected of having cancer. (Item 21) 21. The method of items 1 to 20, wherein the sequence data indicates the presence of a somatic or germline mutation. (Item 22) 22. The method of items 1 to 21, wherein the sequence data indicates the presence of copy number variation. (Item 23) 23. The method of items 1 to 22, wherein the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions. (Item 24) 1. A method for analyzing a population of nucleic acids comprising nucleic acids with different degrees of modification, comprising: contacting the population of nucleic acids with an agent that preferentially binds nucleic acids having the modification; separating nucleic acids of a first pool that are bound to the agent from nucleic acids of a second pool that are not bound to the agent, wherein the nucleic acids of the first pool are over-represented for the modification and the nucleic acids in the second pool are under-represented for the modification; linking the nucleic acids in the first pool and / or the second pool with one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool to produce a population of tagged nucleic acids; amplifying the labeled nucleic acid, wherein the nucleic acid and linked tag are amplified; assaying sequence data of the amplified nucleic acids and linked tags, said assaying providing sequence data for decoding said tags to determine whether said nucleic acids whose sequence data were assayed were amplified from templates in said first or said second pool; A method comprising: (Item 25) 25. The method of claim 24, comprising decoding the tag to determine whether the nucleic acid for which sequence data was assayed was amplified from a template in the first or second pool. (Item 26) 27. The method according to item 25 or 26, wherein the modification is binding of the nucleic acid to a protein. (Item 27) 27. The method of item 25 or 26, wherein the protein is a histone or a transcription factor. (Item 28) 27. The method of claim 25 or 26, wherein the modification is a post-replication modification to a nucleotide. (Item 29) 28. The method of claim 27, wherein the post-replicative modification is 5-methyl-cytosine and the extent of binding of the agent to the nucleic acid increases with the extent of 5-methyl-cytosine in the nucleic acid. (Item 30) 28. The method of claim 27, wherein the post-replicative modification is 5-hydroxymethyl-cytosine and the extent of binding of the agent to the nucleic acid increases with the extent of 5-hydroxymethyl-cytosine in the nucleic acid. (Item 31) 28. The method of claim 27, wherein the post-replicative modification is 5-formyl-cytosine or 5-carboxyl-cytosine, and the extent of binding of the agent increases with the extent of 5-formyl-cytosine or 5-carboxyl-cytosine in the nucleic acid. (Item 32) 27. The method of claim 25 or 26, further comprising washing nucleic acids bound to the agent and recovering the washes as a third pool containing nucleic acids with an intermediate degree of post-replication modification relative to the first and second pools. (Item 33) 27. The method of claim 25 or 26, comprising pooling tagged nucleic acids from the first and second pools prior to assaying. (Item 34) 27. The method of claim 25 or 26, wherein the agent is a 5-methyl binding domain magnetic bead. (Item 35) 35. The method of items 24 to 34, wherein the population of nucleic acids is derived from a body fluid sample. (Item 36) 36. The method of claim 35, wherein the body fluid sample is blood, serum, or plasma. (Item 37) 27. The method of claim 25 or 26, wherein the population of nucleic acids is a cell-free population of nucleic acids. (Item 38) 36. The method of claim 35, wherein the body fluid sample is derived from a subject suspected of having cancer. (Item 39) 39. The method of items 25 to 38, wherein the sequence data indicates the presence of a somatic or germline mutation. (Item 40) 40. The method of items 25 to 39, wherein the sequence data indicates the presence of copy number variation. (Item 41) 40. The method of any one of items 25 to 39, wherein the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions; iv) a nucleic acid, wherein the nucleic acid and the linked tag are amplified and the molecularly tagged aliquots are assayed using an NGS instrument; v) a software module for generating sequence data for decoding the tags; vi) a software module for decoding the tags and analyzing the sequence data to determine whether the nucleic acid for which sequence data was assayed was amplified from a template in the first or second pool. [Brief explanation of the drawings]

[0057] [Figure 1] FIG. 1 shows an exemplary schematic for cleaving RNA, single-stranded DNA, and double-stranded DNA.

[0058] [Figure 2] FIG. 2 shows further exemplary schematics for cleaving RNA, single-stranded DNA, and double-stranded DNA.

[0059] [Figure 3] FIG. 3 shows a schematic diagram for analyzing DNA containing different degrees of 5-methylcytosine representation.

[0060] [Figure 4] Figure 4 shows a schematic diagram for bisulfite sequencing of methylated DNA.

[0061] [Figure 5] FIG. 5 shows a further schematic for analyzing DNA containing different degrees of 5-methylcytosine representation.

[0062] [Figure 6] FIG. 6 shows a further schematic diagram for bisulfite sequencing of methylated DNA.

[0063] [Figure 7] Figure 7 shows an overview of differential tagging.

[0064] [Figure 8] Figure 8 shows an overview of the division method.

[0065] [Figure 9] Figure 9 shows an overview of the methodology.

[0066] [Figure 10] Figure 10 shows an example of using fragmentomic data analysis on fractionated nucleic acid molecules, where genomic position is shown on the X-axis, fragment length on the Y-axis, and coverage or copies on the Z-axis, with corresponding regions of elevated hypo- or hypermethylation indicated.

[0067] [Figure 11] FIG. 11 shows the methylation profiling of normal and lung cancer samples.

[0068] [Figure 12] Figures 12A, 12B, and 12C show methylation profiling using whole genome sequencing. Figure 12A shows the position along a 600 bp region in the transcription start site (TSS) on the X-axis and the frequency of hypermethylated sites along the Y-axis. Figure 12B shows the position along a 600 bp region in the transcription start site (TSS) on the X-axis and the frequency of hypomethylated sites along the Y-axis. Figure 12C shows percent hypermethylation on the X-axis and fragment length on the Y-axis.

[0069] [Figure 13]Figures 13A and 13B show methylation profiling of MOB3A and WDR88. Figure 13A shows the genomic location of the MOB3A gene on the X-axis, and the fragment lengths of nucleic acid molecules from different fractionation groups are shown in separate columns. The fractionation groups included hypermethylated, hypomethylated, hypermethylated mixed with hypomethylated (high + low), and a non-fractionated group (no MBD) for comparison.

[0070] [Figure 14] Figures 14A and 14B show the methylation profiling of the fractionated and unfractionated groups. Figure 14A shows a heatmap with coverage from the unfractionated group (no MBD) and from the combined fractionated samples on the X and Y axes, respectively.

[0071] [Figure 15] FIG. 15 shows the nucleosome organization for the fractionated and unfractionated samples.

[0072] [Figure 16] FIG. 16 shows the validation of the MBD signal.

[0073] [Figure 17] Figure 17 shows statistics for associating input genomic regions with the TSSs of all genes putatively regulated by the genomic region. The X-axis shows the distance to the TSS in kilobases (kb), and the Y-axis shows the region-gene association in percent (%). Above each bar in the graph, the absolute number of counted items is listed. The foreground genomic regions, represented by black bars, were selected from a superset of the background genomic regions, represented by white bars. The background genomic regions were repetitive elements utilized in functional roles selected from all repetitive elements in the genome.

[0074] [Figure 18]Figures 18A and 18B show methylation profiling of the AP3D1 gene. Figure 18A shows the genomic location of the AP3D1 gene on the X-axis, and the read coverage for nucleic acid molecules from different groups is shown in separate columns. The groups include fractionated groups such as hypermethylated and hypomethylated, and an unfractionated group (without MBD) for comparison. The TSS is shown as a vertical line in the center of the heatmap, and the arrow indicates the direction of transcription. Figure 18B shows percent hypermethylation on the X-axis and fragment length on the Y-axis. For example, in Figure 18B, the percent methylation in the unfractionated nucleic acid sample can be approximately 65%, as indicated by the red dotted line.

[0075] [Figure 19] Figures 19A and 19B show the methylation profiling of the DNMT1 gene. Figure 19A shows the genomic location of the DNMT1 gene on the X-axis, and the read coverage for nucleic acid molecules from different groups is shown in separate columns. The groups included fractionated groups such as hypermethylated and hypomethylated, and an unfractionated group (no MBD) for comparison. The TSS is shown as a vertical line in the center of the heatmap, and the arrow indicates the direction of transcription. Figure 19B shows percent hypermethylation on the X-axis and fragment length on the Y-axis.

[0076] [Figure 20] FIG. 20 shows a procedure for strand-based fractionation of nucleic acid molecules.

[0077] [Figure 21] Figure 21 shows the fractionation of nucleic acid molecules into ssDNA and dsDNA. The X axis shows two technical replicates of two samples with varying input DNA (200 ng and 500 ng). The Y axis shows the copy number of the on-target molecule using quantitative PCR amplification. The figure shows the quantitative determination of the target sequence in each group of fractionated cfDNA.

[0078] [Figure 22]Figure 22 shows PCR yields following fractionation of nucleic acid molecules into ssDNA and dsDNA. The X-axis shows the cfDNA input (200 ng and 500 ng) in two technical replicates, and the Y-axis shows the PCR yield in pmol.

[0079] [Figure 23] FIG. 23 shows methylation profiling of promoter regions using whole genome sequencing.

[0080] [Figure 24] Figure 24 provides three examples of strategies for tagging split or fractionated (MBD split) nucleic acid molecules using methyl-binding domain proteins.

[0081] [Figure 25] Figures 25A and 25B show a comparison between coverage for MBD and non-MBD samples in a targeted sequencing assay.

[0082] [Figure 26A] Figures 26A and 26B show coverage for genes in the panel using 15 ng of cfDNA input and two clinical samples (PowerpoolV1 and PowerpoolV2). [Figure 26B] Figures 26A and 26B show coverage for genes in the panel using 15 ng of cfDNA input and two clinical samples (PowerpoolV1 and PowerpoolV2).

[0083] [Figure 27A] Figures 27A and 27B show coverage for genes in the panel using 150 ng of cfDNA input and two clinical samples (PowerpoolV1 and PowerpoolV2). [Figure 27B]Figures 27A and 27B show coverage for genes in the panel using 150 ng of cfDNA input and two clinical samples (PowerpoolV1 and PowerpoolV2).

[0084] [Figure 28AB] Figure 28A, Figure 28B, and Figure 28C show the specificity and sensitivity of variant or mutation detection for the genes in the panel using 15 ng of cfDNA input. [Figure 28C] Figure 28A, Figure 28B, and Figure 28C show the specificity and sensitivity of variant or mutation detection for the genes in the panel using 15 ng of cfDNA input.

[0085] [Figure 29AB] Figure 29A, Figure 29B, and Figure 29C show the specificity and sensitivity of variant or mutation detection for the genes in the panel using 150 ng of cfDNA input. [Figure 29C] Figure 29A, Figure 29B, and Figure 29C show the specificity and sensitivity of variant or mutation detection for the genes in the panel using 150 ng of cfDNA input.

[0086] [Figure 30] FIG. 30 shows the correlation between the average methylation levels measured by whole genome bisulfite sequencing (WGBS) and MBD partitioning.

[0087] [Figure 31A] Figures 31A and 31B show the sensitivity (Figure 31A) and specificity (Figure 31B) of detecting methylated DNA using MBD splitting (Y-axis) and using a whole genome bisulfite sequencing assay (WGBS, X-axis). [Figure 31B]Figures 31A and 31B show the sensitivity (Figure 31A) and specificity (Figure 31B) of detecting methylated DNA using MBD splitting (Y-axis) and using a whole genome bisulfite sequencing assay (WGBS, X-axis).

[0088] [Figure 32] FIG. 32 shows an embodiment of a digital processing device.

[0089] [Figure 33] FIG. 33 shows an embodiment of an application provision system.

[0090] [Figure 34] FIG. 34 illustrates an embodiment of an application delivery system that uses a cloud-based architecture. DETAILED DESCRIPTION OF THE INVENTION

[0091] The terms "cell-free DNA" and "cell-free DNA population," as used herein, refer to DNA originally found in a cell or cells in a large, complex biological organism, e.g., a mammal, and released from the cell into a liquid fluid found in the organism, e.g., plasma, lymph, cerebrospinal fluid, urine, where the DNA can be obtained by obtaining a sample of the fluid without having to perform an in vitro cell lysis step.

[0092] General

[0093] The present disclosure provides a number of methods, reagents, compositions and systems for analyzing complex genomic materials while reducing or eliminating the loss of information (e.g., epigenetic or other types of structural) of molecular features originally present in the complex genomic materials. In some embodiments, molecular tags can be used to track and enumerate different forms of nucleic acids for the purpose of determining genetic modifications (e.g., SNVs, indels, gene fusions and copy number variations). In some embodiments, the methods described herein are used to detect, analyze or monitor conditions such as cancer in a subject or a fetal condition. In some embodiments, the subject is not pregnant.

[0094] The present disclosure provides a method for processing a nucleic acid population containing different forms. As used herein, different forms of nucleic acid have different characteristics. For example, but not limited to, RNA and DNA are different forms based on sugar identity. Single-stranded (ss) and double-stranded (ds) nucleic acids differ in terms of the number of strands. Nucleic acid molecules can differ based on epigenetic characteristics, such as 5-methylcytosine or association with proteins such as histones. Nucleic acids can have different nucleotide sequences, for example, specific genes or gene loci. Characteristics can also differ in degree.

[0095] For example, DNA molecules can differ in the degree of their epigenetic modification. The degree of modification can refer to the number of modification events a molecule has undergone, such as the number of methylation groups (degree of methylation) or other epigenetic changes. For example, methylated DNA can be hypomethylated or hypermethylated. Forms can be characterized by a combination of features, such as single-stranded unmethylated or double-stranded methylated. Fractionation of molecules based on one or a combination of features can be useful for multidimensional analysis of single molecules. These methods accommodate multiple forms and / or modifications of nucleic acids in a sample, allowing sequence information to be obtained for multiple forms. The methods also preserve the identity of the initial multiple forms or modification states throughout processing and analysis, allowing analysis of nucleobase sequences to be combined with epigenetic analysis. Some methods involve separating, tagging, and then pooling different forms or modification states, reducing the number of processing steps required to analyze multiple forms present in a sample. Analyzing multiple forms of nucleic acid in a sample provides somewhat more information because there are more molecules to analyze (which can be important when very little total amount of nucleic acid is available), because different forms or modification states can provide different information (e.g., mutations may only be present in RNA), and because different types of information (e.g., genetic and epigenetic) can be correlated, thereby providing greater precision, certainty, or leading to the discovery of novel correlations with medical conditions.

[0096] CpG dinucleotides are underrepresented in the normal human genome, and the majority of CpG dinucleotide sequences are transcriptionally inactive (e.g., DNA heterochromatin regions in pericentromeric portions of chromosomes and in repetitive elements) and methylated. However, many CpG islands, particularly around transcription start sites (TSSs), are protected from such methylation.

[0097] Cancer can be manifested by epigenetic alterations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation in CpG islands at the transcription start sites (TSSs) of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation can also be associated with abnormal loss of transcriptional capacity of the involved genes, occurring at least as frequently as point mutations and deletions as causes of altered gene expression. DNA methylation profiling can be used to detect regions of the genome with different degrees of methylation ("differentially methylated regions" or "DMRs") that are altered during development or disrupted by disease, such as cancer or any cancer-related disease. The genomes of cancer cells have imbalances in the above-mentioned DNA methylation patterns and therefore in the functional packaging of DNA. Therefore, abnormalities in chromatin organization, coupled with methylation changes, can contribute to enhanced cancer profiling when analyzed together. Combining MBD partitioning with fragmentome data, such as fragment-mapped start and stop positions (correlated with nucleosome positions), fragment lengths and associated nucleosome occupancy, can be used for chromatin structure analysis in hypermethylation studies aimed at improving biomarker detection rates.

[0098] Methylation profiling can involve determining the methylation pattern across different regions of genome.For example, after dividing molecules based on methylation level (for example, the relative number of methylation sites per molecule) and sequencing, the sequences of molecules in different divisions can be mapped to a reference genome.This can show the regions of genome that are more highly methylated or less highly methylated than other regions.In this way, genome regions, as opposed to individual molecules, can differ in their methylation level.

[0099] A feature of a nucleic acid molecule can be a modification, which can include various chemical or protein modifications (i.e., epigenetic modifications). Non-limiting examples of chemical modifications can include covalent DNA modifications, including, but not limited to, DNA methylation. In some embodiments, DNA methylation involves the addition of a methyl group to a cytosine at a CpG site (a cytosine followed by a guanine in a nucleic acid sequence). In some embodiments, DNA methylation involves the addition of an N 6 In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the six-carbon ring of cytosine). In some embodiments, 5-methylation involves the addition of a methyl group to the 5C position of cytosine to create 5-methylcytosine (m5c). In some embodiments, methylation involves derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the six-carbon ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine to create 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be suppressed. DNA methylation is important for normal development, and aberrant methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, for example, suppression, can cause diseases such as cancer. Promoter methylation in DNA can be indicative of cancer.

[0100] Protein modification includes binding with chromatin components, particularly histones containing modified forms, and binding with other proteins, such as proteins involved in replication or transcription.The present disclosure provides a method for processing and analyzing nucleic acids with different degrees of modification, so that the nature of their original modification correlates with nucleic acid tag, and can be deciphered by sequencing the tag when nucleic acid is analyzed.Then, the genetic variation of sample nucleic acid modification can be related to the degree of modification (epigenetic variation) of that nucleic acid in original sample.

[0101] As used herein, the terms "fractionating" and "dividing" refer to separating molecules based on different characteristics. Nucleic acid molecules in a sample can be fractionated based on one or more characteristics. Fractionation can include physically dividing nucleic acid molecules into subsets or groups based on the presence or absence of a genomic characteristic. Fractionation can include physically dividing nucleic acid molecules into groups based on the degree to which a genomic characteristic is present. A sample can be fractionated or divided into one or more groups of divisions based on differential gene expression or characteristics indicative of a disease state. A sample can be fractionated based on a characteristic or combination thereof that provides a difference in signal between normal and diseased states when analyzing nucleic acids, such as cell-free DNA ("cfDNA"), non-cfDNA, tumor DNA, circulating tumor DNA ("ctDNA"), and cell-free nucleic acid ("cfNA").

[0102] The present disclosure provides a method and system for efficiently analyzing nucleic acid molecules.The method can include: dividing nucleic acid molecules into different fractions based on one or more characteristics, and then sequencing (alone or together) and analyzing the nucleic acid molecules in each fraction.In some cases, the fractions of nucleic acid molecules are amplified before and / or after sequencing.The method can be used in various applications such as prognosis, diagnosis, and / or disease monitoring.

[0103] Nucleic acid molecules may be characterized by one or more characteristics. The characteristics of nucleic acid molecules may include strandedness, protein-binding regions, nucleic acid length, start / stop positions, and chemical or protein modifications. The stranded state of nucleic acid molecules may include single-stranded (e.g., ssDNA or RNA) or double-stranded molecules (e.g., dsDNA).

[0104] The genomic feature of the nucleic acid molecule can be a modification, which can include various chemical modifications. By way of non-limiting example, chemical modifications include DNA methylation (5mC), hydroxyl methylation (5hmC), formyl methylation (5fC), carboxyl methylation (5CaC), N 6 This may include covalent DNA modifications such as methyladenine or glycosylation. DNA methylation involves the addition of methyl groups to DNA (e.g., CpG) and can alter the expression of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be suppressed. DNA methylation is important for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, for example, suppression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.

[0105] By way of non-limiting example, benefits of methods that involve splitting single-stranded RNA and / or DNA and double-stranded DNA to characterize a sample include: 1. Further support for SNV, CNV and indel calls from ssDNA and RNA molecules in addition to dsDNA; 2. Easier identification (targeting) of gene fusions in RNA compared to DNA, since variable breakpoints in intronic DNA result in defined exon-exon junctions in RNA; 3. Identification or differential expression levels of messenger RNA (mRNA), microRNA (miRNA), and long non-coding RNA (lncRNA) can be characteristic of numerous disease states. Confirmation and further support of expression signatures observed in altered nucleosome positioning within circulating tumor DNA (ctDNA) populations relative to healthy cell-free DNA (cfDNA) from white blood cells may be important in the early detection of cancer. Furthermore, changes in leukocyte-derived cfDNA and cfRNA expression may also indicate an immune response to disease. 4. Evidence of unstable molecules. Capture of shorter circulating tumor DNA (ctDNA) - Cell-free DNA studies have shown that tumor DNA (ctDNA) can be significantly shorter in length than normal DNA. Some evidence indicates that these shorter sequences may be unstable and exist as ssDNA. They may also provide information about transcription factor binding changes in ctDNA compared to cfDNA, which may be important in early cancer detection. Similarly, cfDNA may also indicate disease response, as well as 5. Capture of damaged / degraded DNA that may be clinically relevant and contains single-stranded "gapped" regions. Analyzing multiple forms of nucleic acid in a sample can occur, for example, by differentially tagging and / or partitioning different forms of nucleic acid prior to sequencing.

[0106] II. Differentially tagging different nucleic acid forms in a sample The sample of nucleic acid, such as acellular nucleic acid in body fluid, often contains nucleic acid in multiple forms, including single-stranded and double-stranded DNA and single-stranded RNA.Because the total amount of nucleic acid in such sample may be small, and the nucleic acid of different forms with different characteristics and / or modifications can bring about different information about sample, the method herein provides for analyzing two, three or all of these forms.

[0107] The preparation and analysis of multiple forms is more efficient when at least some steps can be carried out in parallel.The information determined from such samples is most informative when the sequence information of specific nucleic acid after processing can be correlated with the original form of nucleic acid in sample.For example, when SNV is determined in specific nucleic acid after processing, it can be determined whether this nucleic acid originates from RNA, single-stranded DNA or double-stranded DNA in original sample.

[0108] Identification of different forms of nucleic acid in a sample can be achieved by differentially tagging the different forms of nucleic acid in the sample before the form is altered in a way that obscures its original form, such as by second-strand synthesis or amplification.Therefore, for nucleic acids containing multiple forms, at least one form is linked to a nucleic acid tag to distinguish it from one or more other forms present in the sample.For a sample containing three forms of nucleic acid, such as single-stranded DNA, single-stranded RNA, and double-stranded DNA, the three forms can be distinguished by differentially labeling at least two of the forms, or by differentially labeling all three.The tags linked to nucleic acid molecules of the same form can be identical or different from each other.However, when different from each other, in some embodiments, the tags can share part of their code to identify the attached molecule as a specific form.For example, nucleic acid molecules of a specific form can have codes such as form A1, A2, A3, A4, etc., and nucleic acid molecules of different forms can have codes such as B1, B2, B3, B4, etc. Such a coding system allows for differentiation both between forms and between molecules within a form. An exemplary strategy for differentially tagging nucleic acid molecules with different characteristics, e.g., degree of methylation as determined using methyl-binding domain proteins, is provided in Figure 24 (described below).

[0109] After the differential labeling of one, some or all forms of nucleic acid in sample with nucleic acid tag, the form can be amplified, so that the nucleic acid tag is amplified together with the form in original sample.Then, the amplified nucleic acid can be subjected to sequence analysis to read some or all of the sequences of the nucleic acid in original sample and that of the linked nucleic acid tag.Then, the sequence of tag can be deciphered to show the form of nucleic acid in original sample.Then, the sequence of different forms can be compared to determine whether genetic mutations are mainly or exclusively found in a certain form of nucleic acid, or whether they occur at approximately the same frequency independently of the original form.Some or all of the steps after the differential tagging of different forms, particularly amplification and sequencing, can be carried out using pooled different forms of nucleic acid.This method preferably results in the amplification and sequencing of at least 40, 50, 60, 70, 80, 90 or 95% of the nucleic acid molecules of two, three or more forms present in sample.

[0110] Double-stranded nucleic acids can be differentially labeled by ligating to at least partially double-stranded adapters. Usually, double-stranded nucleic acids are ligated to such adapters at both ends. Either or both of such adapters can contain nucleic acid tags. When two adapters, each with a tag, are linked to the respective ends of a nucleic acid, the tag combination can function as an identifier. Single-stranded DNA or RNA molecules do not ligate to a significant extent with the double-stranded end of the adapter, and therefore do not receive a nucleic acid tag. Double-stranded adapters can be fully double-stranded, as in the case of Y-type adapters or hairpin adapters, or can be partially double-stranded. Exemplary sequences of Y-type adapters are shown below. Universal Adapter

[0111] Universal Adapter SEQ ID NO:1: [ka]

[0112] Adapter Tag SEQ ID NO:2: [ka]

[0113] Truncated versions of these adapter sequences are described in Rohland et al., Genome Res., May 2012. ;Vol. 22 (No. 5): pp. 939-946.

[0114] Because Y-shaped adapters have single-stranded ends, they may need to be avoided (e.g., by separating single-stranded DNA using a probe that does not bind to the Y-shaped adapter) or protected if a subsequent step of separating single-stranded sample nucleic acid from other sample nucleic acid is to be performed.

[0115] RNA molecules can be differentially labeled with nucleic acid tags because they are the only form of molecule in a sample on which reverse transcriptase with an RNA-dependent DNA polymerase can act. The nucleic acid tag can be introduced as a 5' tag on a primer used to prime reverse transcription. Reverse transcription can be random or sequence-specific. After reverse transcription, the original RNA strand is degraded, and a second, complementary DNA strand can then be synthesized. The resulting double-stranded DNA can be optionally blunt-ended and ligated to an adapter in the same manner as double-stranded DNA molecules already present in the sample. Alternatively, the RNA / DNA hybrid molecule can be ligated directly to an adapter.

[0116] Single-stranded DNA molecules can be separated from double-stranded DNA molecules by treatment with an intramolecular ligase. In some embodiments, the intramolecular ligase is CircLigase™ ssDNA ligase, which differentially tags ssDNA with a 3' tag. Prior to treatment with the intramolecular ligase, the ssDNA is dephosphorylated at the 5' end to prevent circularization of the ssDNA. In one example, the ligase used to attach tags to single-stranded DNA is CircLigase™ ssDNA ligase. CircLigase™ ssDNA ligase is a thermostable ATP-dependent ligase. Second-strand synthesis can occur by several mechanisms, including ligating the single-stranded DNA at one end with an oligonucleotide (e.g., using T4 RNA ligase) to provide a primer binding site and hybridizing the single-stranded DNA with a complementary oligonucleotide that serves as a primer for extension based on the hybridized template sequence, or hybridizing with random oligonucleotides that also serve as primers for extension based on the hybridized template sequence. One method uses single-strand ligase to add an oligonucleotide with an extendable 3' end to the single-stranded DNA library member (Gansauge & Meyer, Nature Protocols, Vol. 8, p. 737 ( (See, e.g., 2013). The second DNA strand is filled in using an adapter as a primer binding site. A 5' DNA phosphorylation step and standard (dsDNA) ligation is then performed to add the adapter to the 5' end of the library molecules.

[0117] Alternatively, the method may include steps from the commercially available NEBDirect methodology, in which single-stranded DNA molecules are hybridized with sequence-specific primers for second-strand synthesis, followed by end-repair and ligation to adjacent adapters (see neb.com / nebnext-direct / nebnext-direct-for-target-enrichment). The second DNA strand is degraded and therefore not sequenced. Another method uses random primers with an adapter sequence at the 5' end and random bases at the 3' end. There are usually six random bases, but they can be between four and nine bases long. This approach is particularly suitable for low-input / single-cell amplification for RNA-seq or bisulfite-sequencing (Smallwood et al., Nat. Methods, 2014 Aug;11(8):817-820).

[0118] ssDNA can be selectively captured by nucleic acid (NA) probes by eliminating the standard denaturation step prior to hybridization. ssDNA-probe hybrids can be isolated from cell-free nucleic acid (cfNA) populations by conventional methods (e.g., biotinylated DNA / RNA probes captured by streptavidin-bead magnets). The probe sequences are target-specific and can be the same or different from panels using dsDNA workflows or subsets of those workflows (e.g., targeting exon-exon junctions, RNA fusions at "hotspot" DNA sequences). All single-stranded nucleic acids (ssNA) can be captured in a sequence-agnostic manner at this step by utilizing probes with "universal nucleotide bases" such as deoxyinosine, 3-nitropyrrole, and 5-nitroindole.

[0119] Figure 1 shows an exemplary scheme for separating forms of nucleic acids. The top of the diagram shows a sample containing double-stranded DNA, single-stranded DNA, and single-stranded RNA. RNA is reverse transcribed using a sequence-specific or random poly-T primer with a 5' RNA identification nucleic acid tag. After synthesis of the complementary DNA strand, the RNA template is degraded using RNase H or NaOH or ribosome depletion by selective hybridization. The sample is then treated with capture probes (which can be sequence-specific or sequence-agnostic) without sample denaturation. These probes hybridize with single-stranded molecules and remove them from the sample. In this example, double-stranded DNA molecules in the sample are then blunt-ended and ligated to adapters containing nucleic acid tags. In this example, the adapters are Y-shaped, and the double-stranded arm of the Y is ligated to the DNA molecule. Meanwhile, the separated single-stranded nucleic acids are treated using the DNA protocol discussed above, including tag attachment, or the NEBdirect protocol.

[0120] Figure 2 shows a further exemplary scheme that uses a simplified workflow, most notably by eliminating the 5' DNA phosphorylation step, starting with a sample containing double-stranded DNA, single-stranded DNA, and single-stranded RNA. The double-stranded DNA in the sample is first ligated to a hairpin adapter containing a nucleic acid tag. The sample is then 5' DNA dephosphorylated, and the RNA is then converted to cDNA and ligated to a different tag. The single-stranded DNA is then processed as in Figure 1. In some embodiments, the hairpin adapter can be cleaved into two strands prior to library amplification.

[0121] FIG. 7 illustrates one embodiment of differential tagging. In step 701, a population of nucleic acids is obtained. The nucleic acids can be circulating nucleic acids (cNAs), such as from a liquid biopsy sample (serum, plasma, or blood). In step 702, a first form of nucleic acid is differentially tagged to form a mixture (703) of a first tagged nucleic acid form and a second untagged nucleic acid form. Subsequently, in step 704, a second form of nucleic acid (or the remaining nucleic acid) is tagged with a different label. The method may include two or more different differential tagging steps (702) prior to step 704. After tagging two or more forms of nucleic acid in the population, in some embodiments, the different forms can be separated. If the different forms are separated, the differentially tagged nucleic acids can then be pooled together prior to sequencing or can be sequenced separately. Differential tagging of different forms of nucleic acid preferably occurs in one tube or reaction volume, and the entire tagged molecule is sequenced (without division).The read data obtained from sequencing can be used for analysis to be performed on the read data from different nucleic acid forms and collective nucleic acid samples.

[0122] In some embodiments, the first form of nucleic acid to be differentially tagged is dsDNA, and the differential tagging is performed by attaching dsDNA double-stranded adapters containing a first set of tags. The ssDNA (remaining nucleic acid) is then tagged with a different set of tags (a second set of tags).

[0123] In some embodiments, the first form of nucleic acid to be differentially tagged is DNA derived from an open chromatin region, and tagging is performed by contacting the population of nucleic acids with Tn5-mediated transposase activity.

[0124] In some embodiments, the first form of nucleic acid to be differentially tagged is a double-stranded nucleic acid, and the tagging is performed by attaching a hairpin adaptor to the double-stranded nucleic acid.

[0125] III. Partitioning of nucleic acids with different degrees of modification In certain embodiments described herein, prior to tagging and sequencing, a population of nucleic acids of different forms can be divided based on one or more characteristics of nucleic acids.By dividing a heterogeneous nucleic acid population, for example, rare nucleic acid molecules that are more prevalent in a certain fraction (or fraction) of the population can be enriched, thereby increasing rare signal.For example, by dividing RNA from DNA, it is possible to detect genetic variations that exist in RNA but are less (or not) present in DNA.Similarly, by dividing a sample into hypermethylated and hypomethylated nucleic acid molecules, it is possible to more easily detect genetic variations that exist in hypermethylated DNA but are less (or not) present in hypomethylated DNA.By analyzing multiple fractions of a sample, it is possible to perform multidimensional analysis of single molecules, and therefore achieve higher sensitivity.

[0126] In some examples, a heterogeneous nucleic acid sample is divided into two or more aliquots (e.g., at least three, four, five, six, or seven aliquots). In some embodiments, each aliquot is differentially tagged. The tagged aliquots are then pooled together for collective sample preparation and / or sequencing. The split-tagging-pooling step may be performed more than once, with each round of division being tagged using a differential tag and division means based on a different characteristic (examples are provided herein) that distinguishes it from the other aliquots.

[0127] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitions may contain one or more of the following nucleic acid forms: ribonucleic acid (RNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively, or in addition, a heterogeneous population of nucleic acids is partitioned into RNA and DNA. Alternatively, or in addition, a heterogeneous population of nucleic acids may be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or in addition, a heterogeneous population of nucleic acids may be partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation, the level of methylation, the type of methylation (5' cytosine), and the association and level of association with one or more proteins such as histones. Alternatively, or in addition, a heterogeneous population of nucleic acids may be divided based on nucleic acid length (e.g., molecules up to 160 bp and molecules having a length greater than 160 bp).

[0128] In some instances, each aliquot (representing a different nucleic acid form) is differentially labeled and the aliquots are pooled together prior to sequencing, while in other instances, the different forms are sequenced separately.

[0129] FIG. 8 illustrates one embodiment of the present disclosure. A population of distinct nucleic acids (801) is partitioned (802) into two or more distinct aliquots (803a, b). Each aliquot (803a, b) represents a distinct nucleic acid form. Each aliquot is tagged (804). The tagged nucleic acids are pooled together (807) and then sequenced (808). The reads are analyzed in silico. The tags are used to separate the reads from the distinct aliquots. Analysis to detect genetic variations can be performed at the aliquot level as well as at the total nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variations, e.g., CNVs, SNVs, indels, and fusions, in the nucleic acids in each aliquot. In some examples, in silico analysis can include determining chromatin structure. For example, the coverage or copy number of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage may correlate with higher nucleosome occupancy in a genomic region, and lower coverage may correlate with lower nucleosome occupancy or nucleosome depleted regions (NDRs).

[0130] The sample may contain nucleic acids that vary in modifications, including post-replication modifications to nucleotides and usually non-covalent association with one or more proteins.

[0131] In one embodiment, the nucleic acid population is obtained from serum, plasma, or blood samples from subjects suspected of having cancer or previously diagnosed with cancer. The nucleic acids include those with varying levels of methylation. Methylation can result from any one or more post-replication or transcriptional modifications. Post-replication modifications include modifications of the nucleotide cytosine, particularly 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.

[0132] Partitioning of nucleic acids is achieved by contacting the nucleic acid with the methylation binding domain ("MBD") of methylation binding protein ("MBP"). The MBD binds 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads such as Dynabeads® M-280 streptavidin via a biotin linker. Partitioning into fractions with different degrees of methylation can be achieved by eluting the fractions with increasing NaCl concentrations.

[0133] Generally, elution is a function of the number of methylation sites per molecule, with molecules having more methylated elution at increasing salt concentrations. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into distinct populations based on the degree of methylation. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three (3) aliquots. The molecules are contacted with a solution containing molecules containing a methyl-binding domain at a first salt concentration, which can be attached to a capture moiety such as streptavidin. At the first salt concentration, some of the molecules bind to the MBD, and some remain unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first aliquot, representing hypomethylated forms of DNA, is the one that remains unbound at a low salt concentration, e.g., 160 nM. A second aliquot, representing intermediately methylated DNA, is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM, and is also separated from the sample. A third aliquot, representing highly methylated forms of DNA, is eluted using a high salt concentration, e.g., at least about 2000 nM.

[0134] Each aliquot is differentially tagged. The tag can be a molecule, such as a nucleic acid, that contains information indicating the properties of the molecule to which the tag is associated. For example, a molecule can have a sample tag (which distinguishes molecules in one sample from those in a different sample), a aliquot tag (which distinguishes molecules in one aliquot from those in a different aliquot), or a molecular tag (which distinguishes different molecules from each other (in both unique and non-unique tagging scenarios)). In certain embodiments, the tag can include one or a combination of barcodes. As used herein, the term "barcode" refers to a nucleic acid molecule having a specific nucleotide sequence or the nucleotide sequence itself, depending on the context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have a degenerate sequence or a sequence with a certain Hamming distance, as desired for a particular purpose. Thus, for example, a sample index, aliquot index, or molecular index can be composed of one barcode or a combination of two barcodes, each attached to a different end of the molecule.

[0135] Tags can be used to label individual polynucleotide population aliquots, allowing the tag(s) to be correlated with a particular aliquot. In some embodiments, a single tag can be used to label a particular aliquot. In some embodiments, multiple different tags can be used to label a particular aliquot. In embodiments using multiple different tags to label a particular aliquot, the set of tags used to label one type of aliquot can be easily distinguished from the set of tags used to label other aliquots. In some embodiments, tags can have additional functions, for example, they can be used to index sample sources or as unique molecular identifiers (which can be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations). Similarly, in some embodiments, tags can have additional functions, for example, they can be used to index sample sources or as non-unique molecular identifiers (which can be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).

[0136] In one embodiment, tagging the aliquots comprises tagging the molecules in each aliquot with the equivalent of the sample tag. After recombining the aliquots and sequencing the molecules, the source aliquots are identified by the sample tag. In another embodiment, different aliquots are tagged with different sets of molecular tags, for example, consisting of paired barcodes. In this way, each molecular barcode represents the source aliquot and is also useful for distinguishing molecules within the aliquots. For example, a first set of 35 barcodes can be used to tag the molecules in the first aliquot, and a second set of 35 barcodes can be used to tag the molecules in the second aliquot.

[0137] Tags can be attached to molecules that have already been divided based on one or more characteristics, but the final tagged molecules in library may no longer have these characteristics.For example, single-stranded DNA molecules can be divided and tagged, but the final tagged molecules in library are likely to be double-stranded.Similarly, RNA can be divided, but in final library, the tagged molecules derived from these RNA molecules are likely to be DNA.Therefore, the tags attached to molecules in library usually represent the characteristics of the "parent molecule" from which the final tagged molecules are derived, and are not necessarily the characteristics of the tagged molecules themselves.

[0138] For example, use barcode 1, 2, 3, 4, etc. to tag and label the molecules in the first aliquot, use barcode A, B, C, D, etc. to tag and label the molecules in the second aliquot, and use barcode a, b, c, d, etc. to tag and label the molecules in the third aliquot. Differentially tagged aliquots can be pooled before sequencing. Differentially tagged aliquots can be sequenced separately, or can be sequenced together at the same time, for example, in the same flow cell of an Illumina sequencer.

[0139] After sequencing, the read data can be analyzed to detect genetic variations at the level of each division and at the level of the entire nucleic acid population. Tags are used to sort the read data from different divisions. Analysis can include in silico analysis to determine genetic variations and chromatin structure using sequence information, genome coordinate length and coverage or copy number. Higher coverage can be correlated with higher nucleosome occupancy in genomic regions, and lower coverage can be correlated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0140] In some embodiments, the nucleic acids in the original population can be DNA and / or RNA, single-stranded and / or double-stranded.Dividing based on single-stranded versus double-stranded state can be achieved, for example, by using labeled capture probes to divide ssDNA and using double-stranded adapters to divide dsDNA.Dividing based on RNA versus DNA composition includes, but is not limited to, using double-stranded adapters to divide dsDNA and using reverse transcription with or without capture probes to divide RNA.

[0141] The affinity agent can be an antibody with the desired specificity, a natural binding partner or variant thereof (Bock et al., Nat Biotech, 28:1106-1114 (2010); Song et al., Nat Biotech, 29:68-72 (2011)), or an artificial peptide selected, for example, by phage display, to have specificity for a given target.

[0142] Examples of capture moieties contemplated herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs). Examples of MBPs contemplated herein include, but are not limited to: (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine. (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine over unmodified cytosine. (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind 5-formyl-cytosine over unmodified cytosine (Iurlaro et al., Genome Biol., 14:R119 (2013)). (d) an antibody specific for one or more methylated nucleotide bases;

[0143] Similarly, partitioning different forms of nucleic acid can be performed using histone-binding proteins that can separate histone-bound nucleic acid from free or unbound nucleic acid. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.

[0144] For some affinity agents and modifications, binding to the agent can occur in an essentially all-or-nothing manner, depending on whether the nucleic acid has the modification, but there can be a degree of separation. In such cases, nucleic acids with over-represented modifications will bind to the agent to a greater extent than nucleic acids with under-represented modifications. Alternatively, nucleic acids with modifications can bind in an all-or-nothing manner. However, various levels of modifications can then be sequentially eluted from the bound agents.

[0145] For example, in some embodiments, partitioning can be binary or based on the degree / level of modification. For example, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, further partitioning can involve eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, fragments with higher methylation levels are eluted.

[0146] In some instances, the final fraction represents nucleic acids with different degrees of modification (over- or under-represented modification). Over- and under-representation can be defined by the number of modifications carried by a nucleic acid relative to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is two, nucleic acids containing more than two 5-methylcytosine residues are over-represented in this modification, and nucleic acids with one or zero 5-methylcytosine residues are under-represented. The effect of affinity separation is to enrich for nucleic acids that are over-represented in the modification in the bound phase and for nucleic acids that are under-represented in the modification in the non-bound phase (i.e., in solution). Nucleic acids in the bound phase may be eluted prior to further processing.

[0147] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), sequential elution can be used to resolve different levels of methylation. For example, the hypomethylated aliquot (no methylation) can be separated from the methylated aliquot by contacting the nucleic acid population with MBD from the kit attached to magnetic beads. The beads are used to separate the methylated nucleic acids from the unmethylated nucleic acids. One or more sequential elution steps are then performed to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After these methylated nucleic acids are eluted, magnetic separation is again used to separate the more highly methylated nucleic acids from those with lower levels of methylation. The elution and magnetic separation steps may be repeated to generate various aliquots, such as a hypomethylated aliquot (representing no methylation), a methylated aliquot (representing low levels of methylation), and a hypermethylated aliquot (representing high levels of methylation).

[0148] In some methods, nucleic acids bound to an agent used in affinity separation are subjected to a wash step, which washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids may be enriched for nucleic acids that have modifications near the average or median (i.e., halfway between nucleic acids that remain bound to the solid phase and nucleic acids that are not bound to the solid phase upon initial contact of the sample with the agent).

[0149] Affinity separation results in at least two, sometimes three or more aliquots of nucleic acids with different degrees of modification. While the aliquots are still separated, at least one, usually two or three (or more) aliquots of nucleic acid are linked to nucleic acid tags, usually provided as components of an adapter, and nucleic acids in different aliquots receive different tags that distinguish members of one aliquot from another. Tags linked to nucleic acid molecules of the same aliquot can be identical or different from each other. However, if different from each other, the tags can share part of their code in common to identify the attached molecules as a specific aliquot.

[0150] Figure 3 shows an exemplary scheme. A sample contains nucleic acids with different degrees of methylation, some of which also have genetic mutations. The sample is contacted with magnetic beads linked to an affinity reagent that preferentially binds 5-methylcytosine over cytosine. Affinity purification results in two aliquots of nucleic acids. The aliquot on the left represents nucleic acids that bind to the affinity reagent and is enriched for nucleic acids that over-represent 5-methylcytosine. The aliquot on the right represents nucleic acids that do not bind to the affinity reagent and is enriched for nucleic acids that lack or are under-represented in 5-methylcytosine. The two aliquots are then attached to Y-shaped adapters containing differential nucleic acid tags and amplified. The amplified nucleic acids are then assayed for sequence data, the sequence of the sample nucleic acid indicative of the genetic mutation and the sequence of the tag indicating the aliquot into which the sample nucleic acid was divided, thereby indicating the degree of modification.

[0151] Figure 24 provides an illustrative example of an MBD splitting and tagging approach. In workflow (1), one set of molecular tags (e.g., 35x35 tags) can be applied to all samples before splitting. After splitting, in this example, the molecules in each split are optionally amplified and then independently sequenced for hypermethylated and hypomethylated forms. In workflow (2), the molecules in a sample are split, for example, based on methylation characteristics. Each split is individually tagged, amplified, and sequenced. In workflow (3), the molecules in each of multiple samples are split, tagged with split-specific tags, pooled, and amplified. The molecules in each sample are then provided with a sample tag in order to deconvolute the sample from which they originate.

[0152] In some embodiments, nucleic acid molecules can be fractionated into different fractions based on nucleic acid molecules bound to specific proteins or fragments thereof and those not bound to specific proteins or fragments thereof. Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific protein properties. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that bind to DNA and can serve as the basis for fractionation include, but are not limited to, Protein A and Protein G. Any suitable method can be used to fractionate nucleic acid molecules based on protein-bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein-bound regions include, but are not limited to, SDS-PAGE, chromatin-immunoprecipitation (ChIP), heparin chromatography, and asymmetric field-flow fractionation (AF4).

[0153] IV. Determining 5-methylcytosine patterns in nucleic acids Bisulfite-based sequencing and its variants provide a means for determining the methylation pattern of nucleic acids. In some embodiments, determining the methylation pattern comprises distinguishing 5-methylcytosine (5mC) from unmethylated cytosine. In some embodiments, determining the methylation pattern comprises distinguishing N 6 In some embodiments, determining the methylation pattern includes distinguishing 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC) from unmethylated cytosine. Examples of bisulfite sequencing include, but are not limited to, oxidized bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reduced bisulfite sequencing (redBS-seq).

[0154] Oxidized bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC by first converting 5hmC to 5fC and then proceeding with bisulfite sequencing, as previously described. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish between 5mC and 5hmC. In TAB-seq, 5hmC is protected by glycosylation. The Tet enzyme is then used to convert 5mC to 5caC before proceeding with bisulfite sequencing, as previously described. Reduced bisulfite sequencing is used to distinguish 5fC from modified cytosines.

[0155] Generally, in bisulfite sequencing, a nucleic acid sample is divided into two aliquots, and one aliquot is treated with bisulfite.Bisulfite converts natural cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine or 5-carboxylcytosine) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) remain unconverted.Comparing the nucleic acid sequences of molecules from the two aliquots reveals which cytosines have been converted to uracil and which have not.As a result, modified and unmodified cytosines can be determined.First dividing the sample into two aliquots is disadvantageous for samples that contain only a small amount of nucleic acid and / or samples that are composed of heterogeneous cell / tissue origin, such as body fluids containing cell-free DNA.

[0156] The present disclosure provides methods that enable bisulfite sequencing and its variants. These methods work by linking nucleic acids in a population to a capture moiety, i.e., a label that can be captured or immobilized. Capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety attached to an analyte is captured by its binding partner, which is attached to an isolatable moiety, such as a magnetically attractable particle or a larger particle that can be sedimented by centrifugation. The capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties include biotin, which allows affinity separation by binding to streptavidin bound or linkable to a solid phase, or oligonucleotides, which allow affinity separation by binding to complementary oligonucleotides bound or linkable to a solid phase. After the capture moiety is ligated to the sample nucleic acid, the sample nucleic acid serves as a template for amplification. After amplification, the original template remains ligated to the capture moiety, but the amplicon is no longer ligated to the capture moiety.

[0157] The capture moiety can be linked to the sample nucleic acid as a component of an adapter, which can also provide an amplification and / or sequencing primer binding site. In some methods, the sample nucleic acid is linked to an adapter at both ends, with both adapters carrying a capture moiety. Preferably, any cytosine residues in the adapter are modified, such as with 5-methylcytosine, to protect them from the action of bisulfite. In some examples, the capture moiety is linked to a cleavable linkage (e.g., photocleavable desthiobiotin-TEG or USER™ enzyme, Chem. Commun.). (Camb)., 2015 February 21; Vol. 51 (No. 15): pp. 3266-3269 The capture moiety is linked to the original template by a cleavable uracil residue, in which case the capture moiety can be removed as needed.

[0158] The amplicons are denatured and contacted with an affinity reagent for the capture tag. The original template binds to the affinity reagent, but the amplified nucleic acid molecules do not. Thus, the original template can be separated from the amplified nucleic acid molecules.

[0159] After separation or division, each population of nucleic acids (i.e., the original template and the amplified product) can be subjected to bisulfite treatment, with the original template population receiving bisulfite treatment and the amplified product not. Alternatively, the amplified product can be subjected to bisulfite treatment and the original template population not. After such treatment, each population can be amplified (if the original template population converts uracil to thymine). The population can also be subjected to biotin probe hybridization for enrichment. Each population is then analyzed and the sequences are compared to determine which cytosines were 5-methylated (or 5-hydroxymethylated) in the original. Detection of T nucleotides (corresponding to unmethylated cytosines converted to uracil) in the template population and C nucleotides at the corresponding positions in the amplified population indicates unmodified C. The presence of C at the corresponding positions in the original template and the amplified population indicates modified C in the original sample.

[0160] In some embodiments, the method uses sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation of molecularly tagged DNA libraries (see Figure 4). This process is carried out by adaptor labeling (e.g., biotin), DNA-seq amplification of the entire library, parent molecule recovery (e.g., streptavidin bead pulldown), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosines at single-base resolution by sequential NGS preparative amplification of parent library molecules with and without bisulfite treatment. This can be achieved by modifying one of the two adaptor strands, the 5-methylated NGS adaptor (directional adaptor; Y-shaped / branched with 5-methylcytosine substitutions) used in BIS-seq, with a label (e.g., biotin). Sample DNA molecules are adaptor-ligated and amplified (e.g., by PCR). Because only the parent molecules have labeled adapter ends, they can be selectively recovered from their amplified progeny by label-specific capture methods (e.g., streptavidin magnetic beads). Because the parent molecules retain the 5-methylation mark, bisulfite conversion on the captured library provides single-base resolution 5-methylation status during BIS-seq while preserving the corresponding DNA-seq molecular information. In some embodiments, the bisulfite-treated library can be combined with the untreated library prior to enrichment / NGS by adding sample tag DNA sequences in a standard multiplexed NGS workflow. Similar to the BIS-seq workflow, bioinformatics analysis can be performed for genome alignment and 5-methylated base identification. In essence, this method provides the ability to selectively recover parental, ligated molecules that retain the 5-methylcytosine mark after library amplification, thereby enabling parallel processing of bisulfite-converted DNA. This overcomes the destructive nature of bisulfite treatment on the quality and sensitivity of the DNA-seq information extracted from the workflow.Using this method, the recovered ligated parent DNA molecules (with labeled adapters) can be amplified to form a complete DNA library and subjected to parallel treatments that induce epigenetic DNA modifications. While this disclosure discusses the use of BIS-seq methods to identify 5-methylated cytosine (5-methylcytosine), this is not a limitation. Variants of BIS-seq have been developed to identify hydroxymethylated cytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxylcytosine. These methodologies can be implemented using the sequential / parallel library preparation described herein.

[0161] Alternative methods for modified nucleic acid analysis

[0162] The present disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, histone-linked, and other modifications discussed above). In some such methods, a population of nucleic acids having modifications to different degrees (e.g., 0, 1, 2, 3, 4, 5, or more methyl groups per nucleic acid molecule) is contacted with adapters prior to fractionation of the population according to the degree of modification. The adapters are attached to either one or both ends of the nucleic acid molecules in the population. Preferably, the adapters contain a sufficient number of different tags to result in a low probability of tag combinations, e.g., 95, 99, or 99.9% of two nucleic acids with the same start and stop points will receive the same combination of tags. After attachment of the adapters, the nucleic acids are amplified from primers that bind to the primer binding sites in the adapters. The adapters, whether with the same or different tags, can contain identical or different primer binding sites, although it is preferred that the adapters contain identical primer binding sites. After amplification, the nucleic acids are preferably contacted with an agent that binds to nucleic acids having modifications (such as those agents previously described). Nucleic acid is separated into at least two divided fractions, and the degree that nucleic acid has modification from binding with active substance is different.For example, if active substance has affinity for the nucleic acid with modification, the nucleic acid that over-represents modification (compared to the median representation in the population) will preferentially bind with active substance, while the nucleic acid that under-represents modification will not bind or be more easily eluted from active substance.After separation, different divided fractions can then be subjected to further processing steps, usually in parallel but separately, including further amplification and sequence analysis.Then, the sequence data from different divided fractions can be compared.

[0163] An exemplary scheme for performing such separation is shown in Figure 5. Nucleic acids are ligated to both ends of a Y-shaped adapter containing a primer binding site and a tag. The molecule is amplified. The amplified molecule is then fractionated by contacting it with an antibody that preferentially binds 5-methylcytosine to obtain two aliquots. One aliquot contains the original molecule lacking methylation and an amplified copy that has lost methylation. The other aliquot contains the original DNA molecule with methylation. The two aliquots are then processed and sequenced separately, with further amplification of the methylated aliquot. The sequence data from the two aliquots can then be compared. In this example, tags are not used to distinguish between methylated and unmethylated DNA, but rather to distinguish between different molecules within these aliquots, allowing for the determination of whether reads with identical start and stop points are based on the same or different molecules.

[0164] The present disclosure provides further methods for analyzing a population of nucleic acids, at least some of which contain one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications previously described. In these methods, the population of nucleic acids is contacted with adapters containing one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably, all cytosine residues in such adapters are also modified, or all such cytosines in the primer binding region of the adapter are modified. Adapters are attached to both ends of nucleic acid molecules in the population. Preferably, the adapters contain a sufficient number of different tags to provide a low probability of tag combinations, e.g., 95, 99, or 99.9% of two nucleic acids with identical start and stop points will receive the same combination of tags. The primer binding sites in such adapters can be identical or different, but are preferably identical. After attachment of the adapters, the nucleic acids are amplified from primers that bind to the primer binding sites of the adapters. The amplified nucleic acids are divided into first and second aliquots. The first aliquot is assayed for sequence data, with or without further processing. The sequence data for the molecules in the first aliquot is determined in this manner, regardless of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are treated with bisulfite, which converts unmodified cytosines to uracil. The bisulfite-treated nucleic acids are then subjected to amplification primed by a primer for the original primer binding site of the adapter ligated to the nucleic acid. Because these nucleic acids retain cytosines in the primer binding site of the adapter, only the nucleic acid molecules originally ligated to the adapter (unlike the amplification products) are now amplifiable, but the amplification products have lost the methylation of these cytosine residues and have been converted to uracil during the bisulfite treatment. Therefore, only original molecules in the population that are at least partially methylated undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis. Comparison of the sequences determined from the first and second aliquots can indicate, among other things, which cytosines in the nucleic acid population have been methylated.

[0165] An exemplary scheme for this analysis is shown in Figure 6. Methylated DNA is ligated to Y-shaped adapters at both ends, containing primer binding sites and tags. The cytosines in the adapters are 5-methylated. Primer methylation serves to protect the primer binding sites in the subsequent bisulfite step. After adapter attachment, the DNA molecules are amplified. The amplification products are divided into two aliquots for sequencing with and without bisulfite treatment. The aliquot not subjected to bisulfite sequencing can be subjected to sequence analysis with or without further treatment. The other aliquot is treated with bisulfite, which converts unmethylated cytosines to uracil. Only the primer binding sites protected by cytosine methylation can support amplification when contacted with a primer specific for the original primer binding site. Therefore, only the original molecules are subjected to further amplification, not copies from the first amplification. The further amplified molecules are then subjected to sequence analysis. The sequences from the two aliquots can then be compared. As in Figure 5, the nucleic acid tags in the adapters are not used to distinguish between methylated and unmethylated DNA, but rather to distinguish nucleic acid molecules within the same partition.

[0166] V. General characteristics of the method 1. Sample The sample can be any biological sample isolated from a subject. The sample can be a bodily sample. Samples can include bodily tissues such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid (CSF), synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, fluid in intercellular spaces including gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a bodily fluid, particularly blood and its fractions, and urine. The sample may be in the form originally isolated from the subject, or may have been further processed to remove or add components such as cells, or to enrich one component relative to another. Thus, preferred bodily fluids for analysis are plasma or serum, which contain cell-free nucleic acids. The sample may be isolated or obtained from the subject and transported to a site for sample analysis. The sample may be stored and shipped at a desired temperature, for example, room temperature, 4°C, -20°C, and / or -80°C. The sample may be isolated or obtained from the subject at the location of sample analysis. The subject may be a human, mammal, animal, companion animal, service animal, or pet. The subject may have cancer. The subject may not have cancer or detectable symptoms of cancer. The subject may have been treated with one or more cancer therapies, for example, any one or more of chemotherapy, antibody, vaccine, or biologic. The subject may be in remission. The subject may or may not have been diagnosed with cancer or susceptible to any cancer-related genetic mutation / disorder.

[0167] The volume of plasma can vary depending on the desired read depth of the region being sequenced. Exemplary volumes are 0.4-40 ml, 5-20 ml, and 10-20 ml. For example, volumes can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. The volume of plasma sampled can be 5-20 ml.

[0168] A sample can contain various amounts of nucleic acid, including genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (104 ) haploid human genome equivalents, and in the case of cfDNA, approximately 200 billion (2 × 10 11 Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, or in the case of cfDNA, about 600 billion individual molecules.

[0169] A sample may contain nucleic acids from various sources, e.g., from cells and acellular matter of the same subject, or from cells and acellular matter of different subjects. A sample may contain nucleic acids that harbor mutations. For example, a sample may contain DNA that harbors germline mutations and / or somatic mutations. A germline mutation refers to a mutation present in a subject's germline DNA. A somatic mutation refers to a mutation that originates in a subject's somatic cells, e.g., cancer cells. A sample may contain DNA that harbors a cancer-associated mutation (e.g., a cancer-associated somatic mutation). A sample may contain epigenetic variants (i.e., chemical or protein modifications), where the epigenetic variants are associated with the presence of a genetic mutation, such as a cancer-associated mutation. In some embodiments, a sample contains epigenetic variants associated with the presence of a genetic mutation, and the sample does not contain a genetic mutation.

[0170] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can include obtaining between 1 femtogram (fg) and 200 ng.

[0171] Cell-free nucleic acids are nucleic acids that are not contained within or otherwise bound to cells, or in other words, nucleic acids that remain in a sample after the removal of intact cells. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as cell necrosis and apoptosis. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released into body fluids from cancer cells. Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.

[0172] Cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of the molecules, with a mode of about 168 nucleotides, and a second minor peak in the range between 240-440 nucleotides.

[0173] Cell-free nucleic acids can be isolated from bodily fluids by a fractionation or partitioning step in which the cell-free nucleic acids, as found in solution, are separated from intact cells and other non-soluble components of the bodily fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid can be lysed and the cell-free and cellular nucleic acids processed together. Generally, after the addition of buffer and a washing step, the nucleic acids can be precipitated with alcohol. Further purification steps, such as silica-based columns, can be used to remove contaminants or salts. C is run through the reaction to optimize certain aspects of the procedure, such as yield. o Non-specific bulk carrier nucleic acids such as t-1 DNA, DNA or proteins for bisulfite sequencing, hybridization and / or ligation may be added.

[0174] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA are converted to double-stranded form so that they can be included in subsequent processing and analysis steps.

[0175] 2. Ligation of DNA molecules to adapters Double-stranded DNA molecules in a sample and single-stranded RNA or DNA molecules converted into double-stranded DNA molecules can be ligated to an adapter at either or both ends. Typically, double-stranded molecules are blunt-ended by treatment with a polymerase that has a 5'-3' polymerase and a 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecules can be ligated to at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acid and adapter to facilitate ligation. Both blunt-end and sticky-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In sticky-end ligation, the nucleic acid molecule typically has an "A" overhang and the adapter typically has a "T" overhang.

[0176] 3. Amplification The sample nucleic acid flanked by the adapter can be amplified by PCR and other amplification methods. Amplification is usually primed by a primer that binds to the primer binding site in the adapter adjacent to the DNA molecule to be amplified. Amplification methods can involve cycles of denaturation, annealing, and extension due to thermal cycling, or can be isothermal, as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication.

[0177] Preferably, the method performs dsDNA "T / A ligation" using T-tailed and C-tailed adapters, resulting in at least 50, 60, 70, or 80% amplification of the double-stranded nucleic acid prior to ligation with the adapters. Preferably, the method increases the amount or number of molecules amplified by at least 10, 15, or 20% relative to a control method performed using T-tailed adapters alone.

[0178] 4. Tags Tags, including barcodes, can be incorporated into or otherwise attached to the adapters. Tags can be incorporated by ligation, overlap extension PCR, among other methods.

[0179] Molecular tagging strategies

[0180] Molecular tagging refers to tagging practices that allow distinguishing the molecules to which sequence read data originates. Tagging strategies can be divided into unique tagging and non-unique tagging strategies. In unique tagging, all or substantially all molecules in a sample have different tags, so that read data can be assigned to the original molecule based solely on tag information. Tags used in such methods are sometimes referred to as "unique tags." In non-unique tagging, different molecules in the same sample may have the same tag, so that other information, in addition to tag information, is used to assign sequence read data to the original molecule. Such information may include start and stop coordinates, coordinates to which the molecule is mapped, start or stop coordinates alone, etc. Tags used in such methods are sometimes referred to as "non-unique tags." Therefore, it is not necessary to uniquely tag every molecule in a sample. It is sufficient to uniquely tag molecules that fall within an identifiable class within a sample. Thus, molecules in different identifiable families can have the same tag without losing information about the identity of the tagged molecule.

[0181] In certain embodiments of non-unique tagging, the number of different tags used may be sufficient so that there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%) that all molecules in a particular group will have different tags. It should be noted that when barcodes are used as tags, and when barcodes are attached, for example, randomly, to both ends of molecules, a combination of barcodes may together constitute a tag. This number, in terms, is a function of the number of molecules going into the call. For example, a class may be all molecules that map to the same start-stop position on a reference genome. A class can be all molecules that map to a particular locus, e.g., a particular base, or across a particular region (e.g., up to 100 bases or a gene or exons of a gene). In certain embodiments, the number of distinct tags used to uniquely identify the number of molecules in a class, z, can be between any of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, ​​9*z, 10*z, 11*z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z, or 100*z (e.g., a lower limit) and any of 100,000*z, 10,000*z, 1000*z, or 100*z (e.g., an upper limit).

[0182] For example, in a sample of about 5 ng to 30 ng of cell-free DNA, approximately 3,000 molecules are expected to map to a particular nucleotide coordinate, with between about 3 and 10 molecules having any given start coordinate and sharing the same stop coordinate. Therefore, about 50 to about 50,000 different tags (e.g., between about 6 and 220 barcode combinations) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3,000 molecules mapped across nucleotide coordinates would require about 1 million to about 20 million different tags.

[0183] Generally, the assignment of unique or non-unique tag barcodes in the reaction follows the methods and systems described by U.S. Patent Applications Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731. Tags may be randomly or non-randomly linked to the sample nucleic acids.

[0184] In some embodiments, tagged nucleic acid is loaded into microwell plate and then sequenced.Microwell plate can have 96, 384 or 1536 microwells.In some cases, they are introduced into microwell with expected ratio of unique tags.For example, unique tags can be loaded so that each genome sample is loaded with about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or more than 1,000,000,000 unique tags. In some cases, unique tags may be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the average number of unique tags loaded per sample genome is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or or less than 1,000,000,000 or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags.

[0185] A preferred format uses 20-50 different tag barcodes ligated to both ends of a target nucleic acid. For example, 35 different tag barcodes ligated to both ends of a target molecule creates 35 x 35 permutations (which equals 1225 for 35 tag barcodes). This number of tags is sufficient so that different molecules with identical start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different tag combinations. Other barcode combinations include any number between 10 and 500, such as about 15 x 15, about 35 x 35, about 75 x 75, about 100 x 100, about 250 x 250, or about 500 x 500.

[0186] In some cases, the unique tag can be an oligonucleotide of predetermined or random or semi-random sequence. In other cases, multiple barcodes can be used, so that the barcodes are not necessarily unique to each other. In this example, the barcode can be ligated to each molecule, so that the combination of the barcode and the ligated sequence creates a unique sequence that can be tracked individually. As described herein, the detection of a non-unique barcode in combination with the sequence data of the beginning (start) and end (stop) of the sequence read data can allow a unique identity to be assigned to a specific molecule. The length or number of base pairs of each sequence read data can also be used to assign a unique identity to such a molecule. As described herein, a fragment derived from a single strand of nucleic acid that has been assigned a unique identity can thereby allow subsequent identification of fragments derived from the parent strand.

[0187] 5.Target enrichment In certain embodiments, nucleic acids in a sample can be subjected to target enrichment, in which molecules with target sequences are captured for subsequent analysis. Target enrichment can involve the use of a bait set containing oligonucleotide baits labeled with a capture moiety, such as biotin. Probes can have sequences selected to tile a panel of regions, such as genes. In some embodiments, the bait set can have a higher relative concentration for more specifically desired sequences of interest. Such a bait set is combined with the sample under conditions that allow hybridization of the target molecules with the bait. The captured molecules are then isolated using a capture moiety, for example, a bead-based streptavidin biotin capture moiety. Such methods are further described, for example, in U.S. Patent No. 15 / 426,668, filed February 7, 2017 (U.S. Patent No. 9,850,523, issued December 26, 2017).

[0188] 6. Sequencing The sample nucleic acid adjacent to adapter can be subjected to sequencing, with or without prior amplification.Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing (NGS), single molecule sequencing by synthesis (SMSS) (Helicos), massively parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing by PacBio, SOLiD, Ion Torrent or Nanopore platform. The sequencing reactions can be performed in a variety of sample processing units, which can be multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The sample processing unit can also include multiple sample chambers, allowing multiple runs to be processed simultaneously.

[0189] The sequencing reaction can be performed on one or more forms of nucleic acid, at least one of which is known to contain a marker for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. The sequencing reaction can provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% sequence coverage of the genome. In other cases, the sequence coverage of the genome can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%. Sequence coverage may be performed over at least 5, 10, 20, 70, 100, 200 or 500 different genes, or at most 5000, 2500, 1000, 500 or 100 different genes.

[0190] Simultaneous sequencing reaction can be carried out using multiplex sequencing.In some cases, cell-free nucleic acid can be sequenced using at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other cases, cell-free nucleic acid can be sequenced using less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.Sequencing reaction can be carried out sequentially or simultaneously.Subsequent data analysis can be carried out for all or part of sequencing reaction. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, data analysis may be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. Exemplary read depths are 1000 to 50,000 reads per locus (base).

[0191] 7.Analysis The present method can be used to diagnose the presence of a condition, particularly cancer, in a subject to characterize the condition (e.g., stage the cancer or determine the heterogeneity of the cancer), monitor the response to treatment of the condition, or for an effective prognostic risk of developing the condition or the subsequent course of the condition. The present disclosure can also be useful in determining the effectiveness of a particular treatment option. If treatment is successful, more cancer cells may die and shed DNA, so a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood. In other examples, this may not occur. In another example, perhaps a particular treatment option can be correlated with the genetic profile of the cancer over time. This correlation can be useful in selecting a therapy. Furthermore, if the cancer is observed to be in remission after treatment, the present method can be used to monitor residual disease or disease recurrence.

[0192] The types and number of cancers that can be detected may include blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid state tumors, heterogeneous tumors, homogeneous tumors, etc. The type and / or stage of cancer may be detected from genetic variations including mutations, rare mutations, insertions / deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, changes in chromosomal structure, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0193] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage classification. Genetic profile data can make it possible to characterize specific subtypes of cancer, which can be important in diagnosing or treating that specific subtype. This information can also provide the subject or practitioner with clues regarding the prognosis of a particular type of cancer, allowing either the subject or practitioner to tailor treatment options as the disease progresses. Some cancers may progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful in determining disease progression.

[0194] This analysis is also useful in determining the effectiveness of certain treatment options.If treatment is successful, more cancers may die and shed DNA, so successful treatment options may increase the amount of copy number variations or rare mutations detected in the subject's blood.In other cases, this may not happen.In another example, perhaps a certain treatment option may be correlated with the genetic profile of cancer over time.This correlation may be useful in selecting therapy.In addition, if cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.

[0195] The method can also be used to detect genetic mutations in conditions other than cancer. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection, and certain immune conditions can be monitored. In this example, copy number variation analysis can be performed over time to profile how a particular disease may progress. Copy number variation or even rare mutation detection can be used to determine how pathogen populations change during the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, whereby viruses can change life cycle states and / or mutate to more pathogenic forms during the course of infection. The method can also be used to determine or profile the host body's rejection activity as immune cells attempt to destroy transplanted tissue, monitoring the status of the transplanted tissue and altering the course of treatment or preventing rejection.

[0196] Furthermore, the methods of the present disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods may include, for example, creating a genetic profile of extracellular polynucleotides from a subject, the genetic profile including multiple data resulting from copy number variation and rare mutation analysis. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may result in a heterogeneous genomic population. In the example of cancer, several tumors are known to contain tumor cells at different stages of cancer. In other examples, the heterogeneity may include multiple foci. Again, in the example of cancer, there may be multiple tumor foci, perhaps one or more of which is the result of metastasis spreading from the primary site.

[0197] The methods may be used to generate or profile a fingerprint or data set that is the sum of genetic information from different cells in a heterogeneous disease, which may include copy number variation and mutation analysis, either alone or in combination.

[0198] This method can be used to diagnose, predict prognosis, monitor or observe cancer or other diseases.In some embodiments, the method herein does not include diagnosing, predicting prognosis or monitoring fetus, and therefore is not intended for non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, predict prognosis, monitor or observe cancer or other diseases in unborn subjects, whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0199] An exemplary method for molecular tag identification of MBD-bead partitioned libraries by NGS is as follows: 1. Physical fractionation of extracted DNA samples (e.g., plasma DNA extracted from human samples) using a methyl-binding domain protein-bead purification kit, while retaining all eluate from the process for downstream processing. 2. Applying differential molecular tags and NGS-enabling adapter sequences to each fraction in parallel, for example, hypermethylated, residually methylated ("wash"), and hypomethylated aliquots are ligated with NGS adapters bearing molecular tags. 3. Recombination of all molecularly tagged aliquots and subsequent amplification using adaptor-specific DNA primer sequences. 4. Enrichment / hybridization of the recombined and amplified total library targeting genomic regions of interest (e.g., cancer-specific genetic mutations and differentially methylated regions). 5. Re-amplify the enriched total DNA library while adding sample tags. Different samples are pooled and assayed in multiplex on an NGS instrument. 6. Bioinformatics analysis of NGS data using molecular tags used to identify unique molecules as well as deconvolution of samples into differentially segmented molecules. This analysis can yield information about relative 5-methylcytosine content for genomic regions consistent with standard gene sequencing / genetic variation detection.

[0200] VI. FORM OF CARRYING OUT THE DISCLOSURE The present disclosure provides methods that include dividing a cell-free nucleic acid (cfNA) population into aliquots that share one or more similar characteristics.

[0201] The disclosed methods can be implemented to separate single-stranded nucleic acids (ssNAs; ssDNA, RNA) and dsDNA, where dsDNA molecules are prepared through standard library preparation and ssNAs are prepared through an assisted library preparation workflow that preserves information about the original biomolecule type (i.e., RNA, ssDNA, dsDNA) while converting the ssNAs into a form that can be enriched, sequenced (e.g., NGS), and analyzed.

[0202] The approach to cfNA comprehensive library preparation can include (a) converting RNA to identifiable ssDNA and (b) partitioning ssDNA and dsDNA molecules for parallel NGS library preparation, (c) followed by (optional) target enrichment, and (d) NGS and downstream data analysis to identify molecular types using sequences (see Figure 1).

[0203] In some embodiments, dsDNA-specific NGS adapter ligation of the cfNA population can be performed prior to RNA molecule tagging, specific ligation, cDNA conversion, and NGS library preparation. A simultaneous sequencing method, as shown in Figure 2, in which dsDNA and then RNA are ligated sequentially without partitioning to create an NGS library, can be applied to cfNA samples.

[0204] In some embodiments, platform ligation uses Y-shaped or "branched" adapters that generate ligated ds-cfDNA molecules with ssDNA 5' and 3' ends. These ends can be accidentally ligated by RNA ligase (or Circligase™ II) during simultaneous sequencing or conventional ssDNA library preparation. By modifying the ends of the Y-shaped adapters to "hairpin" or "bubble" shapes, the ligated cf-dsDNA molecules no longer have ssDNA ends and are no longer substrates for subsequent ssNA ligation during simultaneous sequencing / conventional DNA library preparation. Therefore, by designing NGS adapters that do not contain free ssDNA ends, RNA and ssDNA library preparation can be performed in addition to dsDNA workflows without separating the molecule types.

[0205] The disclosed method can be performed on a cfNA population using gene-specific / random / poly-T DNA primers with molecular tagging tails and reverse transcriptase, followed by RNA removal by RNase H or NaOH hydrolysis to produce tagged ssDNA (cDNA) to replace each RNA molecule. Additional methodologies known to those skilled in the art, such as ribosomal RNA depletion by selective hybridization, can be used to remove undesired RNA sequences.

[0206] ssDNA can be selectively captured by NA probes by omitting the standard denaturation step prior to hybridization. ssDNA-probe hybrids can be isolated from cfNA populations by methods known in the art (e.g., biotinylated DNA / RNA probes, captured by streptavidin-bead magnets). The probe sequence is target-specific and can be the same as or different from a panel with a dsDNA workflow, a subset of that workflow (e.g., RNA-fusion at exon-exon junctions, targeting "hotspot" DNA sequences). Furthermore, all ssNAs can be captured in a sequence-agnostic manner in this step by utilizing probes with "universal nucleotide bases" such as deoxyinosine, 3-nitropyrrole, and 5-nitroindole.

[0207] In addition to genetic mutations such as SNVs, indels, gene fusions, and CNVs identified by DNA sequencing, epigenetic mutations (5-methylcytosine, histone methylation, nucleosome positioning, and micro- and long non-coding RNA expression) can lead to or contribute to disease progression, such as cancer. High-throughput measurement of epigenetic markers requires complex molecular biology techniques specifically developed for each type of epigenetic mark. Therefore, epigenetic sequencing projects are typically parallel to DNA (gene) sequencing and require large amounts of input. In other words, multi-analyte biomarker detection involves sample destruction.

[0208] Both genetic (DNA) sequencing and epigenetic sequencing of cell-free DNA have diagnostic value for noninvasive prenatal testing (NIPT) and cancer monitoring / detection. In both applications, the amount of genetic material is limited, making identification of rare molecular events paramount. Therefore, when epigenetic sequencing is performed using current methodologies, each type of marker requires a dedicated sample, reducing the sensitivity of detecting genetic variants.

[0209] While this disclosure provides methods for obtaining information about the epigenetic processes of DNA 5-methylcytosine, the "molecular tag-based partitioning" method outlined for 5-methylcytosine can be applied to other epigenetic mechanisms. Similarly, labeling and recovery of NGS-adapter-ligated parent DNA molecules, as outlined in this disclosure for 5-methylcytosine (5mC) identification, can be used to identify other epigenetic DNA modification markers (e.g., hydroxymethylation, formyl, and carboxyl; 5hmC, 5fC, and 5caC, respectively).

[0210] Bisulfite sequencing is the most widespread approach to 5-methylcytosine, allowing for single-base resolution. This method involves a chemical treatment (bisulfite) of all cytosine bases, converting them to uracil unless they are 5-methylated or 5-hydroxymethylated. Following bisulfite treatment, sequencing results in 5-methylated and 5-hydroxymethylated cytosine residues being detected as cytosine, while unmethylated cytosine, 5-formylmethylated cytosine, and 5-carboxymethylated cytosine are detected as thymine. Variations of bisulfite sequencing have been described, but they can further distinguish between 5mC, 5hmC, 5fC, and 5caC. The main weakness of this approach is the loss of the majority of genetic material. The harsh bisulfite treatment degrades less than 99% of the input DNA, thus reducing the molecular complexity of the sample and the achievable detection limit. Current molecular biology DNA amplification techniques (e.g., PCR, LAMP, RCA) are agnostic to the 5-methylation state of cytosines; therefore, 5-methylation marks are lost upon amplification, which is highly undesirable for liquid biopsy applications. Furthermore, using bisulfite-converted DNA libraries makes it more difficult to detect somatic variants (e.g., distinguishing C->T SNVs from unmethylated cytosines). Therefore, bisulfite-treated DNA is not used for genetic variant detection in liquid biopsy applications. Performing 5-methylcytosine analysis and genetic variant calling on DNA requires sample splitting, which reduces the input / detection sensitivity of each workflow and prevents the identification of both 5-methylcytosine information and genetic variants on a single molecule.

[0211] In certain embodiments, nucleic acids are partitioned based on methylation differences. "Hypermethylated" and "hypomethylated" forms of nucleic acids can be defined as molecules that are above and below a specific degree of methylation, respectively, as determined by the particular partitioning method used. For example, a partitioning method can select molecules with at least two, at least three, at least four, at least five, or at least six methylated nucleotides. The degree of methylation refers to the number of methylated nucleotides in a nucleic acid fragment. Identifying relatively "hypermethylated" DNA molecules in a DNA sample can be achieved by capturing molecules that bind to methyl-binding domain (MBD) proteins, or fragments or variants thereof. MBDs can also be referred to as methyl-CpG-binding domains. MBD proteins can be complexed with magnetic beads. In some embodiments, the proteins that bind to MBDs are MECP2, MBD1, MBD2, MBD3, MBD4, or fragments or variants thereof. Although 5-methylation sites are not directly indicated by this method (not bisulfite conversion), bioinformatic analysis of overlapping hypermethylated fragments can determine the specific site(s) of 5-methylcytosine. A major drawback of this method is that by sequencing only the hypermethylated tract, the vast majority (approximately 80-97% by mass) of the unmethylated human genome is not sequenced, thereby preventing / limiting the identification of genetic variants (e.g., SNVs, indels, and CNVs) because these variants are present in low-coverage regions or are not present at all in the hypermethylated tract.

[0212] The present disclosure provides methods for obtaining 5-methylcytosine data and sequencing data to detect rare genetic variants in the same low-input sample (e.g., liquid biopsy workflows). For example, approaches involving MBD fractionation and tagging are not destructive to the nucleic acids in the sample and preserve genomic complexity after amplification. Furthermore, fractionation-tagging approaches (e.g., MBD fractionation and tagging) can ensure the preservation of genomic complexity by recombining differentially divided nucleic acid molecules, enabling multi-analyte biomarker detection (genetic variants and epigenetic variants). In contrast, other approaches may be destructive to the nucleic acids in the sample. These other approaches may include bisulfite sequencing, methyl-sensitive restriction enzyme digestion, and MBD enrichment when only a fraction or group of nucleic acid molecules is analyzed (e.g., hypermethylated nucleic acid molecules). For example, bisulfite sequencing causes physical damage to nucleic acid molecules. Methyl-sensitive restriction enzyme digestion reduces genomic complexity by destroying the unmethylated fraction, leaving only methylated nucleic acids intact. MBD enrichment can also be used to isolate only a single fraction of nucleic acids in a sample, provided that only MBD-bound nucleic acid molecules are analyzed. This approach destroys information about the nucleic acid molecules present in the non-enriched fraction.

[0213] The methods provided herein for obtaining 5-methylcytosine data (or other methylation status data) can be performed in combination with the above-described methods for obtaining single-stranded and double-stranded nucleic acid information. In some embodiments, the methods described herein quantify the percentage of hypermethylated DNA by differentially tagging DNA molecules split by MBD-beads with varying degrees of methylation (see Figure 3). In this manner, all eluates from the MBD splitting protocol can be collected, and NGS libraries can be prepared using different sets of molecular tags corresponding to the MBD splits. The MBD splitting process therefore reduces the loss of material present in typical bisulfite treatments. Because ligated splits can be recombined before amplification / enrichment / NGS, the drawbacks to the DNA sequencing workflow are minimal. Because MBDs bind to double-stranded DNA (dsDNA), MBD splits preserve the double-stranded nature of the sample DNA, enabling sensitive DNA sequencing methods to tag double-stranded molecules.

[0214] In the MBD partitioned molecular tag NGS workflow, molecular tags can serve two purposes: identifying unique DNA molecules from a sample (by combining the tag with genomic start / end coordinates) and indicating the molecules' relative 5-methylcytosine levels. Using molecular tags, unique nucleic acid molecules can be identified and counted. This information can be used to calculate amplification imbalance. Molecular tags allow for the initial complexity of a sample to be discerned. Molecular tagging allows for the identification and counting of nucleic acid molecules in a sample even in the presence of uneven amplification. The above methodology describes physical partitioning by 5-methylcytosine content, application of differential molecular tags, and optional library recombination, enrichment, NGS, and bioinformatics deconvolution of the partitions, resulting in individual molecules being used for gene sequencing / variant detection simultaneously with DNA-seq. The methodology can be extended to characterize other epigenetic interactions by substituting different DNA- and protein-binding elements that preserve the double-stranded nature of DNA molecules in place of methylation-binding protein (MBD) division. For example, antibodies to histones, modified histones, and transcription factors used in various immunoprecipitation protocols can be substituted for MBD division to yield relative information about nucleosome positioning, nucleosome modifications, and transcription factor binding associated with every DNA molecule in a sample through the use of different sets of molecular tags.

[0215] Data analysis

[0216] A major challenge facing cancer methylation analysis in liquid biopsies is cell type heterogeneity. In addition to the inherent and well-documented heterogeneity of cancer, cell-free DNA in plasma represents a mixture of cell death types that are not primarily cancer-related. For example, cell death can occur in non-malignant organs, physiological hematopoietic lineages, etc. Adding to this complexity is the highly diverse nature of non-cancerous cells in the stromal component, such as vascular and lymphatic endothelial cells and immune cells such as pericytes, macrophages, leukocytes, and lymphocytes, stromal fibroblasts, myofibroblasts, myoepithelial cells, and even adipocytes, endocrine cells, neural cells, and other cells and tissue elements with different developmental origins. Therefore, in some embodiments, adjustments for variations in cell type composition are made when analyzing and interpreting findings from liquid biopsies.

[0217] The analysis pipeline has the following steps: a)-ome occupancy determination b) Locate the dyad and assign strictness c) Fitting Gaussian mixture models within individual genomic elements across the whole genome d) Deconvolving cell lineages at the gene level May include:

[0218] For example, cfDNA fragment starting enrichment profiles can be determined separately for samples from each aliquot. For example, aliquot samples may contain highly, hypo, or intermediately methylated DNA. Using the determined cfDNA fragment starting enrichment profiles, nucleosome occupancy within relevant regulatory elements, such as TSSs, enhancer regions, and distal intergenic elements, can be established. For each aliquot, occupancy peaks, e.g., dyads, can be determined and their stringency assigned. A canonical profile associated with the observed cellular state in a healthy plasma sample can be established by determining the cfDNA fragment starting enrichment profile and locating the dyads in a large, non-malignant control (e.g., a healthy individual or samples from multiple healthy individuals). For any sample, a Gaussian mixture model can be fitted using the canonical profile defined above to generate residual occupancies corresponding to the malignant (non-canonical) chromatin state observed in the aliquot sample, thereby determining the non-canonical cfDNA fragment peaks and profiles. Non-canonical cfDNA fragment peaks and profiles may be associated with the malignant chromatin state of cancer in each divided sample. Biological regulation by methylation can be mediated by a single CpG or by a group of CpGs in close proximity to each other. Therefore, regional analysis of DNA methylation provides a more comprehensive and systematic view of methylation data. Typically, methylation information is summarized over tiling windows or over a set of predefined regions (promoters, CpG islands, introns, etc.).

[0219] Nucleosome organization can be determined by two independent metrics: nucleosome occupancy and nucleosome positioning. Nucleosome occupancy can be understood as the probability that a nucleosome is present at a particular genomic region in a cell population. Nucleosome occupancy can be measured in sequencing-based experiments as coverage (the number of aligned sequencing reads mapped to a genomic region). Nucleosome positioning can be the probability that a nucleosome reference point (e.g., a dyad) is present at a particular genomic coordinate relative to the surrounding coordinates. As shown in Figure 9, good nucleosome positioning can be biologically interpreted as a nucleosome dyad occurring at the same genomic coordinate every time it is present. Poor positioning can be interpreted as a nucleosome dyad occupying a range of positions within the same general footprint of the entire nucleosome. In one example, samples from eight subjects with lung cancer were used to determine dyad centers. Nucleosome positioning and nucleosome occupancy were determined. For example, high occupancy and good positioning can be indicated when coverage is >0.5 quantile (Qu) and peak width is <0.5Qu. In some instances, the distance between dyad centers in fractionated samples (such as high / low methylation fractions) can be compared to unfractionated samples (no MBD). In some cases, dyad centers and adjacent chromatin structure can be determined by assigning dyad centers to all peaks, and occupancy for all peaks is greater than 5% across the genome. Occupancy coverage can be 15%, 20%, 25%, or 30%. Occupancy coverage can be assigned using machine learning approaches by determining peak position, width, length, center, and width resolution. This provides an empirical determination of chromatin structure for plasma DNA.

[0220] Increased sequence read coverage may correlate with greater nucleosome occupancy. Furthermore, nucleosome occupancy may be inversely related to nucleosome-depleted regions (NDRs). Increased nucleosome occupancy may indicate altered chromatin structure, such as more compact chromatin. Compact chromatin may indicate down-regulation of gene expression, which may disrupt normal cellular function. Disruption of normal cellular function can serve as a sign of diseases such as cancer.

[0221] Cell-free DNA contains signals from a heterogeneous population of cells (e.g., dying, malignant, non-malignant, etc.). A heterogeneous population of cells can have nucleic acids in multiple chromatin states. In some instances, the multiple chromatin states can include different states of nucleosome occupancy, such as well-positioned or dispersed ("fuzzy") nucleosomes. Well-positioned nucleosomes exhibit greater coverage of sequence read data, while fuzzy nucleosomes exhibit lower coverage. Based on the coverage of sequence read data, it is possible to determine the nucleosome occupancy across chromatin.

[0222] "Deconvolution" can refer to the process of resolving overlapping cell-free DNA fragment occupancy peaks, thus deriving information about "hidden peaks." Deconvolution of nucleosome occupancy peaks can be achieved by MBD fractionation. Splitting nucleic acids into hypermethylated and hypomethylated fractions can result in two distinct peaks, Peak 1 and Peak 2. However, if the nucleic acid is not fractionated, one continuous peak is obtained, and it may not be feasible to deconvolute malignancy-associated Peak 1 from non-malignant Peak 2.

[0223] A dyad can be a DNA region occupied by a nucleosome center. The dyad can be located in a split sample. In some cases, the nucleic acid is split into high and low methylation fractions. Positioning or localizing the dyad can be performed using a reference-free method or a reference-based method. A reference-free method can include combining both high and low methylation fractions in silico to determine the underlying dyad location, thereby determining a dyad map. In some cases, sequencing data from high and low methylation fractions is combined to determine nucleosome occupancy and compared between the fractions, e.g., combining the signal from all fractions to detect occupancy peaks and then comparing the positions of the peaks found in the high versus low fractions. A reference-based method can include independent analysis of the fractions. For example, nucleosome occupancy for high and low methylation fractions is determined. The nucleosome occupancy per aliquot from the first experiment can be used for the corresponding aliquots in subsequent experiments, where the same part 1 is performed independently on a large set of samples (standard WGS is considered sufficient, as aliquot-based information is not used, but rather information is combined to improve peak resolution), and the maps of occupancy peaks are saved as "references" for each, allowing single aliquots (or both) to be compared.

[0224] Fragmentome signature based on fragmentome data

[0225] Methods for examining fragmentome data are described, for example, in U.S. Patent Application Publication No. 2016 / 0201142 (Lo), International Publication No. 2016 / 015058 (Shendure), and PCT / US17 / 40986 ("Methods For Fragmentome Profiling Of Cell-Free Nucleic Acids"), filed July 6, 2017, all of which are incorporated herein by reference. Fragmentome data refers to sequence data obtained by analyzing nucleic acid fragments. For example, sequence data can include fragment length (in base pairs), genomic coordinates (e.g., start and stop positions on a reference genome), coverage (e.g., copy number), or sequence information (e.g., bases A, G, C, T). Fragmentome data refers to sequence information and associated occupancies of fragment starts and stops in cell-free DNA that correspond to the enrichment of protected content of cell-free DNA observed in blood or plasma.

[0226] For example, it may be possible to determine the number of cfDNA molecules in a sample whose center points map to specific nucleotide coordinates across the genome or a target portion thereof. In healthy individuals, this would typically result in a waveform graph in which peaks represent nucleosomal positions (e.g., where cellular DNA was not cleaved during conversion to cfDNA) and dips represent internucleosomal positions (e.g., where many molecules have been cleaved and therefore few molecules are centered there). The distance between peaks represents nucleosomal dyads. In malignant cells, nucleosomal positions may shift, for example, as a function of methylation. In this case, shifts in the positions of peaks and dips in the graph are expected. Such shifts can be more easily detected by splitting molecules based on different features and examining the fragment distribution for each split. Fragment data can then be further analyzed in one or more additional dimensions. For example, for any coordinate, the number of molecules mapping to that coordinate can be further identified based on fragment size. In graphs based on such data, the third "Z" dimension represents fragment size. Thus, for example, in a two-dimensional graph, the X-axis represents genomic coordinates and the Y-axis represents the number of molecules mapping to the coordinates. In a three-dimensional graph, the X-axis represents genomic coordinates, the Z-dimension represents fragment length, and the Y-axis represents the number of molecules of each size mapping to the coordinates. Such three-dimensional graphs can be represented as two-dimensional heat maps, where the X and Z-axes are displayed in two dimensions and the values ​​on the Y-axis are represented, for example, by color intensity (e.g., darker represents higher values) or color "hotness" (e.g., blue represents lower values ​​and red represents higher values). Mining such data allows for the determination of nucleosome positioning patterns characteristic of the condition being investigated, such as the presence or absence of cancer, type of cancer, extent of metastasis, etc.

[0227] A cohort of individuals may all share a shared characteristic. The shared characteristic may be selected from the group consisting of tumor type, inflammatory state, apoptotic state, necrotic state, tumor recurrence, and resistance to treatment. In some examples, the cohort includes individuals with a particular type of cancer (e.g., breast cancer, colorectal cancer, pancreatic cancer, prostate cancer, melanoma, lung cancer, or liver cancer). To obtain a nucleosome signature of a cancer, individuals with cancer provide a blood sample. Cell-free DNA is obtained from the blood sample. The cell-free DNA is sequenced (with or without selective enrichment of a set of regions from the genome). Sequence information in the form of sequence reads from the sequencing reaction is mapped to a human reference genome. In some embodiments, molecules are collapsed before or after the mapping operation to unique molecular reads.

[0228] Because the cell-free DNA fragments in a given sample represent a mixture of cells from which the cell-free DNA originated, differential nucleosome occupancy from each cell type may contribute to a mathematical model representing a given cell-free DNA sample. For example, fragment length distributions may arise due to differential nucleosome protection across different cell types or across tumor versus non-tumor cells. This method can be used to develop a set of clinically useful assessments based on single-parametric, multiparametric, and / or statistical analyses of sequence data.

[0229] Nucleic acid molecules in a sample may be fractionated based on one or more characteristics. Fractionation may involve physically dividing nucleic acid molecules into subsets or groups based on the presence or absence of a genomic feature. Fractionation may involve physically dividing nucleic acid molecules into groups based on the degree to which a genomic feature is present. Samples may be fractionated or divided into one or more groups based on features indicative of differential gene expression or disease state. Samples may be fractionated based on features that provide a difference in signal between normal and pathological states during the analysis of nucleic acids, such as cfDNA, non-cfDNA, tumor DNA, and circulating tumor DNA (ctDNA).

[0230] Fragmentome data may be used to infer genetic variants. Genetic variants include copy number variations (CNVs), insertions and deletions (indels), single nucleotide variations (SNVs), and / or gene fusions. Fragmentome data may also be used to infer epigenetic variants, such as variants indicative of cancer. One or more genetic variants may be determined in each fractionated or divided group and / or unfractionated nucleic acid. Fractionation or division can be performed based on at least one of various characteristics, including, but not limited to, the methylation status, size, length, and transcriptional binding of the nucleic acid. The genetic variants determined in the fractionated or divided groups may be compared with each other and / or with unfractionated nucleic acids, which may or may not have the same characteristics. The fractionated or divided nucleic acids can be recombined, and the fragmentome data can be compared with unfractionated nucleic acids and / or nucleic acids that do not have the same characteristics as the fractionated or divided nucleic acids to determine the presence of genetic variants.

[0231] The model may be used in a panel setting to selectively enrich regions (e.g., fragmentome profile-associated regions) to ensure a large number of reads spanning specific mutations, and to examine important chromatin-centric events such as transcription start sites (TSSs), promoter regions, splice sites, and intron regions.

[0232] In one example, differences in fragmentome profiles are found at or near intron-exon junctions (or borders). Identifying one or more somatic mutations can be correlated with one or more multiparametric or single-parametric models to reveal the genomic locations where cfDNA fragments are distributed. This correlation analysis can reveal one or more intron-exon junctions where fragmentome profile disruption is most pronounced.

[0233] As another example, hypermethylation in a sample can be observed in regions further away from the TSS. Enrichment of hypermethylated regions can be observed between 0 kb and 5 kb, 5 kb and 50 kb, and / or 50 kb and 500 kb from the TSS. Enrichment of hypermethylated regions can be observed between 5 kb and 50 kb from the TSS. Enrichment of hypermethylated regions can be observed less than 5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 50 kb, 100 kb, 200 kb, 300 kb, 400 kb, and / or 500 kb from the TSS. Enrichment of hypermethylated regions can be observed at distances greater than 5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 50 kb, 100 kb, 200 kb, 300 kb, 400 kb, and / or 500 kb from the TSS. The location and enrichment of hypermethylation can vary between DNA obtained from healthy or normal subjects (normal DNA) and diseased subjects. For example, DNA from subjects suspected of or with lung cancer (lung cancer DNA) may show enrichment for hypermethylation at the most distal distance from the canonical position within the TSS, with well-positioned nucleosomes in the hypermethylated fraction occupying the vicinity of promoter regions (Figure 17). For example, unfractionated nucleic acid (without MBD) from a lung cancer patient was used for sequencing. Based on the fragmentome data, such as genome location, the nucleosome dyad center was determined for the sequence read data. Furthermore, based on the fragmentome data, sequence read data with coverage less than or equal to 5% or coverage less than or equal to 95% were further analyzed. A gene annotation tool, such as Genomic Regions Enrichment of Annotations Tool (GREAT), was used to assign functionality to a set of genomic regions based on nearby genes. The distance between the sequence read data and its putative regulated genes were determined (Figure 17).Distance was divided into four separate bins: one from 0 to 5 kb, another from 5 kb to 50 kb, a third from 50 kb to 500 kb, and a final bin for all associations greater than 500 kb. Specifically, the bins were [0, 5 kb], [5 kb, 50 kb], [50 kb, 500 kb], and [500 kb, infinity]. On the graph, all associations exactly at 0 (i.e., on the TSS) were evenly distributed between the [-5 kb, 0] and [0, 5 kb] bins. Using this method, hypermethylation in a sample was observed in regions further distal from the TSS in both background genomic regions (e.g., all nucleosomes) and foreground genomic regions (e.g., methylated nucleosomes). For example, enrichment of hypermethylated regions was observed between the [5 kb, 50 kb] bins.

[0234] Fragmentome signatures can help determine nucleosome occupancy, nucleosome positioning, RNA polymerase II pausing, death-specific DNase hypersensitivity, and chromatin condensation during cell death. Such signatures can also provide insight into cellular debris clearance and transport. For example, cellular debris clearance can involve DNA fragmentation performed by caspase-activated DNase (CAD) in cells dying from apoptosis, but can also be performed by lysosomal DNase II after dying cells are phagocytosed, resulting in distinct cleavage maps.

[0235] Genome partition maps can be constructed by genome-wide identification of differential chromatin states in malignant versus non-malignant states associated with the aforementioned properties of chromatin through the assembly of meaningful windows into regions of interest, commonly referred to as genome partition maps.

[0236] Methylation status-based fractionation

[0237] Nucleic acid molecules in a sample can be fractionated based on the characteristics of 5-methylcytosine. DNA can be methylated at cytosines, such as at CpG dinucleotide regions. DNA methylation, in conjunction with histone complexes, can affect DNA packaging into chromatin and the epigenetic regulation of gene expression. Epigenetic alterations may play a crucial role in various diseases, including all steps of cancer progression, the initiation of primary or early-stage cancers, and recurrent or metastatic cancers. For example, hypermethylation of normally hypomethylated regions, such as transcription start sites (TSSs) of genes involved in normal growth, DNA repair, cell cycle regulation, and cell differentiation, may indicate cancer. Hypermethylation can alter gene expression by repressing transcription. In some cases, hypermethylation can reduce and / or suppress gene expression. For example, hypermethylation can reduce and / or suppress the expression of oncogene repressors. In some cases, hypermethylation can increase and / or promote gene expression. For example, hypermethylation of a suppressor may increase and / or promote gene expression of downstream responders, such as oncogenes that are normally repressed by the suppressor.

[0238] Based on DNA methylation status, nucleic acid molecules in a sample can be fractionated into different groups, allowing for the enrichment of nucleic acid molecules with similar methylation status using experimental procedures. For example, methyl-binding domain (MBD) proteins can be used to affinity purify nucleic acid molecules with similar methylation status, such as hypermethylated, hypomethylated, and residually methylated. In another example, antibodies specific to 5-methylcytosine can be used to immunoprecipitate nucleic acid molecules with similar levels of methylation. In another example, bisulfite-based methods can be used to selectively enrich highly methylated nucleic acid molecules. In yet another example, methylation-sensitive restriction enzymes can be used to selectively enrich highly methylated nucleic acid molecules.

[0239] Once fractionated using one of the features, the nucleic acid molecules in each group can be sequenced to generate sequence read data. The sequence read data can be mapped to a reference genome. The mapping can generate sequence information. The sequence information can be analyzed to determine genetic variations, including, for example, single-base variants, copy number variations, indels, or fusions. In examples where cell-free DNA is assayed, the methods disclosed herein can be used to generate fragmentome data, which may vary between groups of fractionated nucleic acid molecules. The fragmentome data can include genomic coordinates, size, coverage, or sequence information. The disclosure provides methods for integrating fragmentome data with sequence read data from each fraction. Such integration can be useful for accurate and rapid detection of biomarkers indicative of disease states.

[0240] Using the methods described herein, it is possible to enrich nucleic acid molecules in silico based on fragmentome data. For example, unfractionated nucleic acid molecules (without MBD) from lung cancer patients can be used for sequencing. In another example, fractionation can be achieved based on differences in mononucleosome or dinucleosome profile alone or in combination with other characteristics such as size and / or methylation status. A mononucleosome profile can refer to the coverage or total number of fragments of the approximate length (e.g., about 146 bp) required to wrap around a single nucleosome. A dinucleosome profile can refer to the coverage or total number of fragments of the approximate length (e.g., about 292 bp) required to wrap around a single nucleosome twice.

[0241] Data analysis

[0242] In certain embodiments, data from subjects of different classes, for example, cancer / cancer-free, cancer type 1 / cancer type 2, can be used to train a machine learning algorithm to classify samples as belonging to one of the classes. The term "machine learning algorithm," as used herein, refers to an algorithm that is executed by a computer and automates the construction of analytical models, for example, for clustering, classification, or pattern recognition. Machine learning algorithms can be supervised or unsupervised. Learning algorithms include, for example, artificial neural networks (e.g., backpropagation networks), discriminant analysis (e.g., Bayesian classifiers or Fisher analysis), support vector machines, decision trees (e.g., recursive partitioning processes such as CART-classification and regression trees), random forests, linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal component regression (PCR)), hierarchical clustering, and cluster analysis. The dataset from which a machine learning algorithm learns can be referred to as "training data."

[0243] The term "classifier," as used herein, refers to algorithmic computer code that receives test data as input and provides as output a classification of the input data as belonging to one of several classes.

[0244] The term "dataset," as used herein, refers to a collection of values ​​characterizing elements of a system. The system may be, for example, cfDNA derived from a biological sample. The elements of such a system may be gene loci. Examples of datasets (or "data sets") include values ​​indicating quantitative measures of features selected from: (i) DNA sequences mapping to gene loci, (ii) DNA sequences starting at gene loci, (iii) DNA sequences ending at gene loci; (iv) dinucleosome protection or mononucleosome protection of DNA sequences; (v) DNA sequences located in introns or exons of a reference genome; (vi) size distribution of DNA sequences having one or more features; (vii) length distribution of DNA sequences having one or more features, etc.

[0245] The term "value," as used herein, refers to an item in a dataset that can be anything that characterizes the feature to which the value refers, including, without limitation, a number, a word or phrase, a symbol (e.g., + or -), or a degree.

[0246] Digital Processing Device In some embodiments, the methods described herein utilize a digital processing device. In further embodiments, the digital processing device includes one or more hardware central processing units (CPUs) or general purpose graphics processing units (GPGPUs) that perform the functions of the device. In still further embodiments, the digital processing device further includes an operating system configured to execute the executable instructions. In some embodiments, the digital processing device is optionally connected to a computer network. In further embodiments, the digital processing device is optionally connected to the Internet to access the World Wide Web. In still further embodiments, the digital processing device is optionally connected to a cloud computing infrastructure. In other embodiments, the digital processing device is optionally connected to an intranet. In other embodiments, the digital processing device is optionally connected to a data storage device.

[0247] In accordance with the description herein, suitable digital processing devices include, by way of non-limiting example, server computers, desktop computers, laptop computers, notebook computers, handheld computers, internet appliances, mobile smartphones, and tablet computers.

[0248] In some embodiments, the digital processing device includes an operating system configured to execute executable instructions. An operating system is software, including programs and data, that manages the device's hardware and provides services for the execution of applications. Those skilled in the art will recognize that suitable server operating systems include, by way of non-limiting example, FreeBSD, OpenBSD, NetBSD®, Linux®, and Apple® Mac OS®. Those skilled in the art will recognize that suitable personal computer operating systems include, by way of non-limiting example, Microsoft® Windows®, Apple® Mac OS®, X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those skilled in the art will recognize that suitable mobile smartphone operating systems include, by way of non-limiting example, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry® OS, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.

[0249] In some embodiments, the device comprises a storage and / or memory device. A storage and / or memory device is one or more physical devices used to temporarily or permanently store data or programs. In some embodiments, the device is volatile memory and requires power to maintain the stored information. In some embodiments, the device is non-volatile memory and retains the stored information when power is not applied to the digital processing device. In further embodiments, the non-volatile memory comprises flash memory. In some embodiments, the non-volatile memory comprises dynamic random access memory (DRAM). In some embodiments, the non-volatile memory comprises ferroelectric random access memory (FRAM®). In some embodiments, the non-volatile memory comprises phase change random access memory (PRAM). In other embodiments, the device is a storage device, including, by way of non-limiting example, a CD-ROM, a DVD, a flash memory device, a magnetic disk drive, a magnetic tape drive, an optical disk drive, and cloud computing-based storage. In further embodiments, the storage and / or memory device is a combination of devices, such as those disclosed herein.

[0250] In some embodiments, the digital processing device includes a display that transmits visual information to a user. In some embodiments, the display is a liquid crystal display (LCD). In further embodiments, the display is a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, the display is an organic light emitting diode (OLED) display. In various further embodiments, the OLED display is a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display. In some embodiments, the display is a plasma display. In other embodiments, the display is a video projector. In still other embodiments, the display is a head-mounted display in communication with the digital processing device, such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting example, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headsets, and the like. In further embodiments, the display is a combination of devices, such as those disclosed herein.

[0251] In some embodiments, the digital processing device includes an input device that receives information from a user. In some embodiments, the input device is a keyboard. In some embodiments, the input device is a pointing device, including, by way of non-limiting example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some embodiments, the input device is a touchscreen or multi-touchscreen. In other embodiments, the input device is a microphone that captures voice or other sound input. In other embodiments, the input device is a video camera or other sensor that captures motion or visual input. In further embodiments, the input device is a Kinect, Leap Motion, or the like. In further embodiments, the input device is a combination of devices, such as those disclosed herein.

[0252] Referring to Figure 32, in certain embodiments, an exemplary digital processing device 101 is programmed or otherwise configured to analyze, assay, decode, and / or deconvolute sequence and / or tag data. In embodiments, the digital processing device 101 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 105, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The digital processing device 101 also includes memory or memory locations 110 (e.g., random access memory, read-only memory, flash memory), electronic storage 115 (e.g., hard disk), communication interface 120 (e.g., network adapter) for communicating with one or more other systems, and peripherals 125, such as cache, other memory, data storage, and / or electronic display adapters. The memory 110, storage 115, interface 120, and peripherals 125 communicate with the CPU 105 through a communication bus (a physical line), such as a motherboard. The storage 115 can be a data storage device (or data repository) for saving data. Digital processing device 101 may be operatively coupled to a computer network ("network") 130 with the aid of communication interface 120. Network 130 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. Network 130, in some cases, is a telecommunications and / or data network. Network 130 may include one or more computer servers, which may operate distributed computing such as cloud computing. Network 130, in some cases, may implement a peer-to-peer network with the aid of device 101, which may operate devices coupled to device 101 acting as clients or servers.

[0253] Continuing with reference to FIG. 32, CPU 105 may execute a sequence of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 110. The instructions may be directed to CPU 105, which may then be programmed or otherwise configured to perform the methods of the present disclosure. Examples of operations performed by CPU 105 may include fetching, decoding, executing, and writing back. CPU 105 may be part of a circuit, such as an integrated circuit. One or more other components of device 101 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0254] Continuing with reference to Figure 32, storage device 115 may store files such as drivers, libraries, and saved programs. Storage device 115 may store user data, such as user preferences and user programs. Digital processing device 101 may include one or more additional data storage devices that are external, in some cases, such as located on a remote server communicating over an intranet or the Internet.

[0255] 32, digital processing device 101 is capable of communicating with one or more remote computer systems over network 130. For example, device 101 is capable of communicating with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab, and Microsoft® Surface®), and smartphones (e.g., Apple® iPhone® or Android-enabled devices).

[0256] The methods described herein may be performed at least in part by machine (e.g., computer processor) executable code stored on an electronic storage location of the digital processing device 101, such as, for example, on memory 110 or electronic storage 115. The machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor 105. In some cases, the code may be retrieved from storage 115 and stored on memory 110 for easy access by the processor 105. In some situations, the electronic storage 115 may be eliminated, and the machine-executable instructions may be stored on memory 110.

[0257] Non-transitory computer-readable storage medium

[0258] In some embodiments, the methods disclosed herein utilize one or more non-transitory computer-readable storage media encoded with a program, the non-transitory computer-readable storage media including instructions executable by an operating system of a digital processing device, optionally connected to a network. In further embodiments, the computer-readable storage medium is a tangible component of the digital processing device. In still further embodiments, the computer-readable storage medium is optionally removable from the digital processing device. In some embodiments, the computer-readable storage medium includes, by way of non-limiting example, CD-ROMs, DVDs, flash memory devices, solid-state memory, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, and the like. In some cases, the programs and instructions are encoded on the medium permanently, substantially permanently, semi-permanently, or non-transitoryly.

[0259] Executable Instructions

[0260] In some embodiments, the methods disclosed herein utilize instructions executable by a digital processing device in the form of at least one computer program. For example, a computer program includes a sequence of instructions written to perform particular tasks that are executable by a CPU of a digital processing device. The computer-readable instructions may be implemented as program modules such as functions, objects, application program interfaces (APIs), data structures, and the like that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those skilled in the art will recognize that computer programs may be written in a variety of languages ​​and versions.

[0261] The functionality of the computer-readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises a single sequence of instructions. In some embodiments, a computer program comprises multiple sequences of instructions. In some embodiments, a computer program is provided from a single location. In other embodiments, a computer program is provided from multiple locations. In various embodiments, a computer program comprises one or more software modules. In various embodiments, a computer program comprises, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or a combination thereof.

[0262] Web Applications

[0263] In some embodiments, the computer program comprises a web application. In light of the disclosure provided herein, those skilled in the art will recognize that web applications, in various embodiments, utilize one or more software frameworks and one or more database systems. In some embodiments, the web application is created on a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, the web application utilizes one or more database systems, including, by way of non-limiting example, relational, non-relational, object-oriented, associative, and XML database systems. In further embodiments, suitable relational database systems include, by way of non-limiting example, Microsoft® SQL Server, mySQL™, and Oracle®. Those skilled in the art will also recognize that web applications, in various embodiments, are written in one or more versions of one or more languages. Web applications may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, the web application is written in part in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or eXtensible Markup Language (XML). In some embodiments, the web application is written in part in a presentation definition language such as Cascading Style Sheets (CSS).In some embodiments, the web application is written in part in a client-side scripting language such as Asynchronous Javascript and XML (AJAX), Flash Actionscript, Javascript, or Silverlight. In some embodiments, the web application is written in part in a server-side coding language such as Active Server Pages (ASP), ColdFusion, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tcl, Smalltalk, WebDNA, or Groovy. In some embodiments, the web application is written in part in a database query language such as Structured Query Language (SQL). In some embodiments, the web application integrates with an enterprise server product such as IBM Lotus Domino. In some embodiments, the web application includes a media player element. In various further embodiments, the media player element utilizes one or more of a number of suitable multimedia technologies, including, by way of non-limiting example, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.

[0264] 33 , in a particular embodiment, the application delivery system includes one or more databases 200 accessed by a relational database management system (RDBMS) 210. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBM DB2, IBM Informix, SAP Sybase, Teradata, and the like. In this embodiment, the application delivery system further includes one or more application servers 220 (such as a Java server, a .NET server, a PHP server, and the like) and one or more web servers 230 (such as Apache, IIS, GWS, and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 240 over a network such as the Internet, and the system provides a browser-based and / or mobile-native user interface.

[0265] Referring to FIG. 34, in certain embodiments, the application delivery system alternatively has a distributed cloud-based architecture 300, including elastic load-balanced, auto-scaling web server resources 310 and application server resources 320, as well as a synchronized replicated database 330.

[0266] Mobile Applications

[0267] In some embodiments, the computer program comprises a mobile application that is provided to the mobile digital processing device. In some embodiments, the mobile application is provided to the mobile digital processing device at the time of its manufacture. In other embodiments, the mobile application is provided to the mobile digital processing device via a computer network as described herein.

[0268] In view of the disclosure provided herein, mobile applications are created using techniques known to those skilled in the art using hardware, languages, and development environments known in the art. Those skilled in the art will recognize that mobile applications may be written in a number of languages. Suitable programming languages ​​include, by way of non-limiting example, C, C++, C#, Objective-C, Java™, Javascript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or a combination thereof.

[0269] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non-limiting example, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available free of charge, including, by way of non-limiting example, Lazarus, MobiFlex, MoSync, and PhoneGap. Additionally, mobile device manufacturers sell software developer kits, including, by way of non-limiting example, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.

[0270] Those skilled in the art will recognize that several commercial forums are available for the distribution of mobile applications, including, by way of non-limiting example, the Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.

[0271] Standalone Applications

[0272] In some embodiments, the computer program comprises a standalone application, which is a program that runs as an independent computer process rather than as an add-on to an existing process, e.g., as a plug-in. Those skilled in the art will recognize that standalone applications are often compiled. A compiler is a computer program or programs that convert source code written in a programming language into binary object code, such as assembly language or machine code. Suitable compiled programming languages ​​include, by way of non-limiting example, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB.NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, the computer program comprises one or more executable compilation applications.

[0273] Software Module

[0274] In some embodiments, the methods disclosed herein utilize software, server, and / or database modules. In light of the disclosure provided herein, software modules are created by techniques known to those of ordinary skill in the art using machines, software, and languages ​​known in the art. The software modules disclosed herein are implemented in numerous ways. In various embodiments, a software module comprises a file, a portion of code, a programming object, a programming structure, or a combination thereof. In further various embodiments, a software module comprises multiple files, multiple portions of code, multiple programming objects, multiple programming structures, or a combination thereof. In various embodiments, one or more software modules include, by way of non-limiting examples, a web application, a mobile application, and a standalone application. In some embodiments, a software module is in one computer program or application. In other embodiments, a software module is in more than one computer program or application. In some embodiments, a software module is provided on one machine. In other embodiments, a software module is provided on more than one machine. In further embodiments, a software module is provided on a cloud computing platform. In some embodiments, a software module is provided on one or more machines in a single location. In other embodiments, software modules are provided on one or more machines in more than one location.

[0275] Database

[0276] In some embodiments, the methods disclosed herein utilize one or more databases. In light of the disclosure provided herein, one of skill in the art will recognize that many databases are suitable for storing and retrieving patient, sequence, tag, code / decode, genetic variant, and disease information. In various embodiments, suitable databases include, by way of non-limiting example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, and Sybase. In some embodiments, the database is internet-based. In further embodiments, the database is web-based. In yet further embodiments, the database is cloud computing-based. In other embodiments, the database is based on one or more local computer storage devices.

[0277] In one aspect, provided herein is a system including a computer including a processor and computer memory, where the computer is in communication with a communications network, and the computer memory comprises code that, when executed by the processor, (1) receives sequence data from the communications network into the computer memory; (2) determines, using a method described herein, whether genetic variants in the sequence data represent germline mutations or somatic mutations; and (3) reports the determination to the communications network.

[0278] The communication network can be any available network that connects to the Internet. The communication network can be, for example, but not limited to, Broadband over High speed transmission networks are available including Powerlines (BPL), Cable Modem, Digital Subscriber Line (DSL), Fiber, Satellite, and Wireless.

[0279] In one aspect, provided herein is a system comprising: a local area network; one or more DNA sequencers connected to the local area network, the one or more DNA sequencers comprising computer memory configured to store DNA sequence data; and a bioinformatics computer connected to the local area network, the bioinformatics computer comprising the computer memory and a processor, wherein the computer further comprises code that, when executed, copies the DNA sequence data stored in the DNA sequencers, writes the copied data to the memory of the bioinformatics computer, and performs the steps described herein.

[0280] Numerous systems for carrying out the described methods are also provided herein. In some embodiments, the system includes a nucleic acid sequencer, including a next-generation DNA sequencer, in data communication with a digital processing device, where the software module(s) on the digital processing device receive data generated by the sequencer as the sequencer obtains DNA sequence information from the segmented and tagged DNA sequences segmented and tagged by the subject methods. The sequencer and digital processing device need not be located near each other and, in some embodiments, can be separated by a large physical distance, provided that appropriate data communication exists between the system components. The specific system embodiments described below are exemplary of the many types of systems provided by the present invention. The methods described herein, including data analysis steps, can be readily implemented through the systems disclosed herein, and those skilled in the art will understand that software module(s) on the digital processing device are used to analyze sequence data obtained by sequencing tagged nucleic acid populations generated by the subject methods.

[0281] An embodiment includes a system including a nucleic acid sequencer; a digital processing device including at least one processor, an operating system configured to execute executable instructions, and a memory; and a data link communicatively connecting the nucleic acid sequencer and the digital processing device, wherein the digital processing device is configured to receive an application for analyzing a population of nucleic acids including at least two forms of nucleic acid selected from double-stranded DNA, single-stranded DNA, and single-stranded RNA, each of the at least two forms comprising a plurality of molecules, and the application is configured to receive (i) sequence data of at least some tagged amplified nucleic acids from the nucleic acid sequencer via the data link, the sequence data identifying at least one of the forms of nucleic acid and identifying the forms from each other. The system further comprises executable instructions for creating an application comprising: (i) a software module created by linking at least one tagged nucleic acid to distinguish between nucleic acids and amplifying forms of the nucleic acid, at least one of which is linked to at least one nucleic acid tag, wherein the nucleic acid and the linked nucleic acid tag are amplified to produce amplified nucleic acids, at least one of which is tagged; and (ii) a software module for assaying the sequence data of the amplified nucleic acids by obtaining sufficient sequence information to decode the tagged nucleic acid molecules of the amplified nucleic acids to reveal the forms of the nucleic acids in the population and provide the original templates of the amplified nucleic acids linked to the tag nucleic acid molecules whose sequence data is assayed. In another embodiment of the system, the application further comprises a software module for decoding the tagged nucleic acid molecules of the amplified nucleic acids to reveal the forms of the nucleic acids in the population and provide the original templates of the amplified nucleic acids linked to the tag nucleic acid molecules whose sequence data is assayed. In another embodiment of the system, the application further comprises a software module for transmitting the results of the assay via a communication network. Another embodiment is a system comprising: a next-generation sequencing (NGS) instrument; a digital processing device comprising at least one processor, an operating system configured to execute executable instructions, and a memory; and a data link communicatively connecting the NGS instrument and the digital processing device, wherein the digital processing device further comprises executable instructions to create an application comprising: (i) a software module for receiving sequence data from the NGS instrument via the data link, the sequence data being created by physically partitioning DNA molecules from a human sample to create two or more aliquots, applying differential molecular tags and NGS-enabling adapters to each of the two or more aliquots to create molecularly tagged aliquots, and assaying the molecularly tagged aliquots on the NGS instrument; (ii) a software module for generating sequence data for deconvoluting the sample into differentially partitioned molecules; and (iii) a software module for analyzing the sequence data by deconvoluting the sample into differentially partitioned molecules. In yet another embodiment of the system, the system further comprises a software module that transmits the results of the assay over a communication network.

[0282] Another embodiment is a system comprising: a next generation sequencing (NGS) instrument; a digital processing device comprising at least one processor, an operating system configured to execute executable instructions, and a memory; and a data link communicatively connecting the NGS instrument and the digital processing device, wherein the digital processing device comprises a software module configured to receive sequence data from the NGS instrument via the data link, an application for molecular tag identification of a library fractionated by MBD-beads, the software module being configured to receive sequence data from the NGS instrument via the data link, the sequence data comprising: physically fractionating the extracted DNA sample using a methyl-binding domain protein-bead purification kit while retaining all eluate for downstream processing; performing parallel application of differential molecular tags and NGS-enabling adapter sequences to each fraction or group; and recombining all molecularly tagged fractions or groups to identify adapter-specific DNA primer sequences. The system further comprises instructions executable by at least one processor to create an application comprising: (i) a software module configured to: (i) generate NGS sequence data generated by the steps of: subsequently amplifying the combined and amplified total library using a molecular tag; performing enrichment / hybridization of the recombined and amplified total library while targeting genomic regions of interest; re-amplifying the enriched total DNA library while adding sample tags; pooling the different samples; and assaying them in multiplex on an NGS instrument, wherein the NGS sequence data generated by the instrument provides sequences of molecular tags that are used to identify unique molecules and sequence data for deconvolution of the samples into differentially partitioned molecules; and (ii) a software module configured to perform analysis of the sequence data by using the molecular tags to identify unique molecules and deconvoluting the samples into differentially partitioned molecules. Another embodiment is a system wherein the application further comprises a software module configured to transmit the results of the analysis over a communications network.

[0283] Another embodiment is a system including: (a) a next generation sequencing (NGS) instrument; (b) a digital processing device including at least one processor, an operating system configured to execute executable instructions, and a memory; and (c) a data link communicatively connecting the NGS instrument and the digital processing device, wherein the digital processing device includes: i) a software module for receiving sequence data from the NGS instrument via the data link, the sequence data comprising: contacting a population of nucleic acids with an agent that preferentially binds nucleic acids having a modification; separating nucleic acids of a first pool that are bound to the agent from nucleic acids of a second pool that are not bound to the agent, wherein the nucleic acids of the first pool are over-represented for the modification and the nucleic acids in the second pool are under-represented for the modification; and and / or a software module loaded with and created with labeled nucleic acids prepared by linking the nucleic acids in the second pool to one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool to produce a population of tagged nucleic acids, amplifying the labeled nucleic acids, whereby the nucleic acids and the linked tags are amplified, and assaying the molecularly tagged aliquots using an NGS instrument; ii) a software module for generating sequence data for decoding the tags; and iii) a software module for decoding the tags and analyzing the sequence data to determine whether the nucleic acids whose sequence data were assayed were amplified from templates in the first or second pool. Another embodiment is a system further comprising a software module for transmitting the results of the assay over a communications network. [Example]

[0284] VII. Working Examples Example 1 Experimental procedure for methyl-binding domain (MBD)-based fractionation

[0285] Sample collection

[0286] Samples such as blood, serum, or plasma from subjects with lung cancer (e.g., NSCLC) were selected from the Guardant Health repository that exhibited high circulating tumor DNA (ctDNA) content as determined by the GUARDANT360™ assay. Cell-free DNA (cfDNA) from healthy normal donors was extracted from blood-isolated plasma as previously described (Lanman et al., Analytical and clinical validation of a digital sequencing panel for quantitative, highly accurate evaluation of cell-free circulating tumor DNA, PLoS ONE 10(10):e0140712 (2015)).

[0287] cfDNA extraction

[0288] The samples were subjected to proteinase K digestion. DNA was precipitated with isopropanol. DNA was captured on a DNA purification column (e.g., QIAamp DNA Blood Mini Kit) and eluted in 100 μl of solution. DNA below 500 bp was purified by Ampure Selection was performed using SPRI magnetic bead capture (PEG / salt). The resulting products were suspended in 30 μl of HO. Size distribution was examined (major peak = 166 nucleotides; minor peak = 330 nucleotides) and quantified. Typically, 5 ng of extracted DNA contains approximately 1700 haploid genome equivalents ("HGE"). The general correlation between the amount of DNA and HGE was listed as follows: 3 pg of DNA = 1 HGE; 3 ng of DNA = 1K HGE; 3 ng of DNA = 1M HGE; 10 pg of DNA = 3HGE; 10 ng of DNA = 3K HGE; 10 ng of DNA = 3M HGE.

[0289] DNA fractionation

[0290] DNA was fractionated into multiple fractions. cfDNA (10–150 ng) was fractionated into highly, intermediately, and hypomethylated fractions using the MethylMiner™ affinity enrichment protocol (Thermo Fisher Scientific, catalog number ME10025), except that the reaction conditions were modified to use 300 mM NaCl incubation and wash buffers, and the 1 microgram DNA input protocol was scaled down to submicrogram amounts of DNA input.

[0291] Bead preparation

[0292] Washing Dynabeads® M-280 Streptavidin: Dynabeads® M-280 Streptavidin was washed using a wash buffer containing 300 mM NaCl prior to coupling with the MBD-Biotin protein. The Dynabeads® M-280 Streptavidin stock was resuspended to obtain a homogenous suspension. For each microgram of input DNA, 10 μl of beads was added to a 1.7 ml DNase-free microcentrifuge tube. The bead volume was brought to 100 μl with 1x Bind / Wash buffer. The tube was placed on a magnetic rack for 1 minute to allow all beads to collect on the inner wall of the tube before removing and discarding the liquid. The tube was removed from the magnetic rack and an equal volume (e.g., approximately 100-250 μl) of 1x Bind / Wash buffer was added to resuspend the beads. The resuspended beads were collected and washed one more time before subsequent coupling of MBD-biotin protein to the beads.

[0293] Dynabeads® M-280 Streptavidin was coupled to the MBD-Biotin protein: For each microgram of input DNA, 7 μl (3.5 μg) of MBD-Biotin protein was added to a 1.7 ml DNase-free microcentrifuge tube. The bead volume was brought to 100 μl with 1× Bind / Wash buffer (containing 300 mM NaCl). The MBD-Biotin protein was diluted and transferred to the tube of resuspended beads from the first bead wash. The bead-protein mixture was mixed for 1 hour at room temperature on a rotating plate mixer before continuing to wash the MBD-Beads.

[0294] Washing MBD-Beads: The MBD-Beads in the tube were collected by placing the tube on a magnetic rack for 1 minute. The liquid was removed and discarded. The beads were resuspended in 100-250 μl of 1x Bind / Wash buffer (containing 300 mM NaCl) and mixed on a rotary mixer for 5 minutes at room temperature. The beads were collected, washed, and resuspended as above two more times. The tube was then placed on a magnetic rack for 1 minute, and the liquid was carefully removed and discarded. The beads were resuspended in 100-250 μl of 1x Bind / Wash buffer (containing 300 mM NaCl) before methylated DNA capture.

[0295] Fragmented methylated DNA is captured onto MBD-beads, which are then incubated with the fragmented DNA. Generally, input DNA can range from 5 ng to 1 μg. Control reactions typically use 1 μg of K-562 DNA. To a clean 1.7 ml DNase-free microcentrifuge tube, 20 μl of 5x Wash / Bind buffer (containing 300 mM NaCl) was added. Fragmented sample DNA, e.g., 5 ng to 1 μg, was added to the tube, and the final volume was brought to 100 μl with DNase-free water. The DNA / buffer mixture was transferred to the tube containing the MBD-beads and mixed for 1 hour at room temperature on a rotary mixer. Alternatively, the mixture can be mixed overnight at 4°C.

[0296] Uncaptured DNA was collected from the bead solution; uncaptured / unmethylated DNA was collected from the DNA and MBD-bead mixture. The tube containing the DNA and MBD-bead mixture was placed on a magnetic rack for 1 minute to collect the beads, and the supernatant was removed and saved in a clean, DNase-free microcentrifuge tube. This saved supernatant is the uncaptured DNA supernatant and can be stored on ice. The beads were washed with 200 μl of 1x Bind / Wash buffer (containing 300 mM NaCl) for 3 minutes on a rotating mixer. The beads were collected as above, and the supernatant containing uncaptured / unmethylated / hypomethylated DNA was removed and saved as above and stored on ice. The beads were washed, mixed, and collected, and the supernatant was removed and saved once more to collect two wash fractions. Each wash fraction was stored on ice. The wash fractions can be pooled together and labeled appropriately.

[0297] Elution of captured DNA: The captured DNA was eluted using elution buffer containing 2000 mM NaCl. The beads were resuspended in 200 μl of elution buffer (2000 mM NaCl). The beads were incubated on a rotary mixer for 3 minutes, placed on a magnetic rack for 1 minute to collect all the beads, and the liquid containing the captured / hypermethylated DNA was removed and saved in a clean DNase-free microcentrifuge tube. The first pooled fraction of captured / methylated DNA was stored on ice. The beads were resuspended and incubated once more, and the liquid containing the captured / methylated DNA was removed and saved in a second clean tube. The first and second pools of captured / methylated DNA were pooled and saved on ice.

[0298] Preparation of methylated-fractionated DNA for analysis: The fractionated cfDNA, hypermethylated, intermediately methylated, and unmethylated DNA, was purified, for example, by SPRI bead cleanup (Ampure XP, Beckman Coulter), and then prepared for ligation (using the NEBNext® Ultra™ End Repair / dA-Tailing Module). Subsequently, the fractionated cfDNA was ligated with modified Y-shaped dsDNA adapters containing non-random molecular barcodes as described in Lanman et al. (2015). The hypermethylated, intermediately methylated, and hypomethylated cfDNA fractions were ligated with 11, 12, and 12 different non-random molecular barcode adapters, respectively. After ligation, the fractionated cfDNA molecules for each sample were purified again with SPRI beads (Ampure XP) and then recombined in a PCR reaction using universal oligos (NEBNext Ultra II™ Q5 master mix) for all adapter-ligated molecules to amplify all cfDNA molecules from a single sample. The amplified DNA libraries were again purified using SPRI beads (Ampure XP) in preparation for target enrichment or whole genome sequencing (WGS) using standard prep techniques.

[0299] Target capture and enrichment: DNA samples are sequenced using commercially available protocols, e.g., SureSelect for Illumina multiplex sequencing. XT Enrichment may also be performed using the Target Enrichment System.

[0300] Example 3 CDKN2A methylation profiling

[0301] DNA methylation profiling in conjunction with fragmentome data was used to capture differentially methylated regions (DMRs) in the CDKN2A gene. The CDKN2A gene is a tumor suppressor gene encoding the p16INK4A and p14ARF proteins, which are involved in cell cycle regulation. cfDNA samples were fractionated into hypomethylated and hypermethylated aliquots using MBD-affinity purification. Once fractionated, nucleic acid molecules within each aliquot were sequenced to generate sequence read data. The sequence read data, when mapped to the reference genome, provided fragmentome data, which were then combined with the sequence read data from each of the fractionated aliquots (Figure 10). The CDKN2A gene showed an overall increase in coverage in the hypomethylated aliquot compared to the hypermethylated aliquot.

[0302] Example 4 Methylation profiles of normal and lung cancer samples

[0303] As shown in Figure 11, the MBD splitting process was applied to four cfDNA samples from healthy donors (Norm13893, Norm13959, Norm13961, and Norm13962) and two cfDNA samples from lung cancer patients with a high percentage of ctDNA (LungA1345402 and LungA0516902) using varying input amounts (10–150 ng of cfDNA) and replicates (e.g., three replicates). Samples were hierarchically clustered by the percent of hypermethylated DNA across all targeted genomic loci within the panel. Percent hypermethylated DNA can be determined by dividing the number of hypermethylated cell-free DNA fragments by the total number of cell-free DNA fragments observed across all splits. The panel is a custom gene panel covering an approximately 30 kb genomic region. The panel also has higher sensitivity for detecting different cancers, such as lung cancer, colorectal cancer, etc. Samples from healthy donors clustered separately from samples from lung cancer patients. Individual lung cancer samples had distinct methylation profiles that further clustered separately (i.e., replicates of each lung cancer sample were precisely identified and grouped together). See, e.g., WO 2017 / 181146, October 19, 2017.

[0304] Example 5 Methylation profiling using whole genome sequencing

[0305] DNA methylation profiling was integrated with fragmentome data to determine aberrant fragmentation patterns and, therefore, altered chromatin structure in clinical samples (Figures 12A, 12B, and 12C). Nucleic acid molecules were derived from lung cancer patients. The nucleic acid molecules were fractionated into hypomethylated and hypermethylated aliquots using MBD-affinity purification. Once fractionated, the nucleic acid molecules in each aliquot were sequenced to generate sequence read data. The sequence read data, when mapped to the reference genome, provided fragmentome data. Fragmentome data, such as genomic location, fragment length, and coverage, were combined with the sequence read data from each aliquot. As shown in Figures 12A and 12B, a 600-bp region of the transcription start site (TSS) is on the X-axis, and the frequency or coverage is shown on the Y-axis. Figure 12C shows the percentage of hypermethylated fragments relative to total fragments on the X-axis and frequency on the Y-axis. For example, in Figure 12C, the proportion of hypermethylated fragments among all fragments is about 0.2 (ie, about 20%).

[0306] Example 6 Methylation profiling of MOB3A and WDR88

[0307] DNA methylation profiling was integrated with fragmentome data to determine differences in epigenetic regulation (Figures 13A and 13B). Nucleic acid molecules were fractionated into hypomethylated and hypermethylated aliquots using MBD-affinity purification. Once fractionated, nucleic acid molecules from each aliquot were sequenced to generate sequence read data. The sequence read data, when mapped to a reference genome, provided fragmentome data. Fragmentome data, such as genomic location and coverage, were combined with the sequence read data from each of the fractionated groups.

[0308] The MOB3A gene may have unknown biochemical functions and may be implicated in maintaining tumor growth and proliferation. The heat map in Figure 13A showed greater coverage of hypermethylation compared to hypomethylation near the start of the TSS in samples from healthy individuals. This example provided applications for combining fractionated groups with fragmentome data to detect markers at the TSS of genes that may be indicative of cancer. These data indicated that fractionated groups (or aliquots) provide better resolution for distinguishing methylation status across genomic regions, such as TSSs, both hypermethylated and hypomethylated. As described above, the coverage of fractionated groups indicated differences in methylation status across TSSs. This example provided applications for fractionating nucleic acid molecules to provide better resolution of methylation status across genes.

[0309] The WDR88 gene may be involved in cell cycle regulation, apoptosis, and autophagy. Heat maps showed greater coverage of hypermethylation compared to hypomethylation near the start of TSSs in samples from healthy individuals (Figure 13B). Furthermore, Figure 13B showed that fractionated groups, both hypermethylated and hypomethylated, provide better resolution for distinguishing methylation status across genomic regions, such as TSSs. As noted above, coverage of fractionated groups indicated differences in methylation status across TSSs. This example demonstrated the utility of fractionating nucleic acid molecules to provide better resolution of methylation status across genes.

[0310] Example 7 Methylation profiling of recombined fractionated and unfractionated samples

[0311] Figure 14A shows a heatmap with coverage from the unfractionated population (no MBD) and the recombined aliquots (all MBDs) after MBD affinity partitioning on the X and Y axes, respectively. The aliquots were recombined in silico after partitioning into high and low methylation aliquots to form "high + low" or "all MBDs." The heatmap shows a linear correlation between coverage for no MBDs and all MBDs. A linear correlation indicates similar coverage, potentially providing similar resolution of methylation status across genomic loci. The level of resolution afforded by no MBDs and / or all MBDs may not be sufficient to distinguish differences in methylation status across loci, indicating an unexpected advantage of partitioning based on MBD affinity.

[0312] Figure 14B shows an MVA plot heatmap using all MBDs. The X-axis shows the average fragments (recombined hyper- and hypomethylated fractions) in all MBDs as (a+b) / 2, where a = all MBDs and b = no MBDs.

[0313] Example 8 Nucleosome organization between recombined fractions (all MBD) and unfractionated samples

[0314] As shown in Figure 15, the difference in distance between nucleosome occupancy centers for all MBDs (high and low methylation fractions recombined in silico) across a genomic region and for the no MBD (unfractionated) sample is plotted on the X-axis. The difference in the distribution of distance between nucleosome occupancy centers for all MBDs across a genomic region and for the no MBD sample is plotted on the Y-axis, indicated by "density." The all MBD sample was prepared by recombining the high and low methylation fractions in silico. These results indicate that MBD fractionation does not affect nucleosome occupancy.

[0315] Example 9 Validation of MBD signals

[0316] MBD-split samples were used to distinguish nucleosome occupancy between healthy and cancer samples. In this example, blood samples were obtained from six lung cancer patients and three non-malignant healthy adults. Cell-free nucleic acids from the samples were extracted and split into hyper- and hypomethylated fractions using MBD-affinity purification. The nucleic acid samples were sequenced using whole-genome sequencing. The percent hypermethylated fragments per fraction and for all samples were determined. Figure 16 shows MBD signals in hyper- and hypomethylated fractions from lung cancer patients (rows 1 and 2 from the top) and healthy adults (rows 3 and 4). As shown in Figure 16, cell-free DNA fragments from lung cancer patients show enrichment of distal intragenic regions in hypermethylated fractions (LungSigHyper) compared to hypermethylated fractions from healthy individuals. Furthermore, the distribution of features in the top 5% percent hypermethylated (LungSigHyper) and hypomethylated (LungSigHypo) peaks shows a significant enrichment of hypomethylated peaks in all exons in addition to exon 1 (Figure 16, columns 1 and 2).

[0317] Example 10 Methylation profiling of the AP3D1 gene

[0318] The method described herein was used for the prognosis of lung cancer. In the experiment, samples containing nucleic acid molecules from lung cancer patients were fractionated into hypomethylated and hypermethylated fractions using MBD-affinity purification. As a control, one sample was not fractionated (no MBD). The samples were sequenced using whole genome sequencing.

[0319] The AP3D1 gene may encode the AP-3 complex subunit delta-1, which may be involved in organelle transport. Heat maps showed greater coverage of hypermethylated aliquots compared with hypomethylated aliquots and / or no MBD near the TSS (Figure 18A). Hypermethylated aliquots showed stronger and / or more localized coverage than the no MBD group. As shown in the heat maps, hypermethylated aliquots had more localized strong coverage near the TSS, while the no MBD group had similar coverage across the genomic region. The average percent hypermethylation was also determined, as indicated by the red line in Figure 18B. This example can be used to fractionate nucleic acid molecules to provide better resolution of methylation status across genes. These results indicate that the AP3D1 gene is particularly hypermethylated near the TSS (Figure 18A) and that the AP31 gene is hypermethylated (>60%, as shown in Figure 18B). Deregulation of the AP3D1 gene may be involved in causing lung cancer. Thus, this example may provide an application of this method in the prognosis of lung cancer by monitoring the methylation profile of an individual.

[0320] Example 11 Methylation profiling of the DNMT1 gene

[0321] In another example, we investigated methylation profiling of the DNMT1 gene. The DNMT1 gene encodes an enzyme that catalyzes the transfer of methyl groups to specific CpG dinucleotides in DNA. DNMT1 has been implicated in maintaining DNA methylation to ensure the fidelity of replication of inherited epigenetic patterns. Aberrant methylation patterns may be associated with cancer and developmental abnormalities.

[0322] Heat maps of hypermethylated, hypomethylated, and MBD-free regions are shown relative to the TSS (Figure 19A). Hypermethylated segments showed stronger and / or more localized coverage than the MBD-free group. Hypermethylated segments had stronger coverage localized near the TSS, while the MBD-free group had similar coverage across the gene. The average percent hypermethylation was also determined to be approximately 75%, as indicated by the red line in Figure 19B. These results indicate that the DNMT1 gene is particularly hypermethylated near the TSS (Figure 19A) and that the DNMT1 gene is hypermethylated (approximately 75%, as shown in Figure 19B). Aberrant methylation patterns, along with changes in chromatin structure, may lead to deregulation of DNMT1, which may be involved in causing lung cancer. Therefore, this example may provide applications for this method in the prognosis of lung cancer by monitoring an individual's methylation profile. This example may also provide applications for fractionating nucleic acid molecules to provide better resolution of methylation status across genes.

[0323] Example 12 Modified histone fraction

[0324] This example demonstrates partitioning using a modified histone approach. DNA is partitioned based on histone modifications. Briefly, agarose beads are blocked with BSA, and following washing, the beads are preincubated with antibodies against H3K9me3 and H4K20me3 (Millipore, Temecula, CA, USA) for 4 hours at 4°C. Subsequently, 200 μl of plasma is diluted in 800 μl of partition dilution buffer and then added to the pelleted agarose beads preincubated with the antibodies. Following overnight incubation at 4°C, the beads are washed with low-salt, high-salt, LiCl, and Tris / EDTA buffers. Finally, chromatin is eluted by incubating the beads at 65°C, and proteins are removed by treatment with proteinase K. The partitioned DNA is then purified using an appropriate purification kit and stored at -20°C.

[0325] Example 13 Fractionation based on protein binding domains

[0326] This example demonstrates a partitioning approach using protein-binding regions. DNA is partitioned based on differential binding to Protein A. Nucleic acid molecules in a sample can also be fractionated based on protein-binding regions. For example, nucleic acid molecules can be fractionated into distinct groups based on nucleic acid molecules that bind to a specific protein and those that do not. Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the proteins. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and serve as the basis for fractionation include, for example, Protein A or Protein G. Experimental procedures such as chromatin immunoprecipitation are used to fractionate nucleic acid molecules based on Protein A-binding regions.

[0327] Example 14 Fractionation based on hydroxymethylation

[0328] This example demonstrates partitioning using a modified histone approach. DNA is partitioned based on hydroxymethylation. Briefly, 5-hmC-modified bases are glycosylated in vitro. Specific glycosylation of 5-hmC is achieved by following the protocol of a highly active 5-hmC glycosyltransferase enzyme from Zymo Research (zymoresearch.com / epigenetics / dna-hydroxymethylation / 5-hmc-glucosyltransferase). J-binding protein-1 (JBP-1) is a high-affinity glycosyltransferase. It specifically binds to glycosylated DNA, allowing 5-hmC levels to be determined by JBP-1-based enrichment. Furthermore, glycosylation of 5-hmC alters DNA digestion by several restriction enzymes, and therefore the digestion pattern of 5-hmC-glycosylated DNA can be used to assess DNA hydroxymethylation status.

[0329] Example 15 Fractionation based on the state of the nucleic acid molecule strands

[0330] Nucleic acid molecules in a sample are fractionated based on strand state. For example, ssDNA and dsDNA are fractionated into two groups. These groups are subjected to sequencing assays individually or simultaneously. Nucleic acid samples containing both ssDNA and dsDNA are fractionated by not subjecting the sample to a denaturation step during fractionation. The denaturation step converts dsDNA to ssDNA, preventing the fractionation of nucleic acid molecules based on strand state.

[0331] Example 16 Molecular partitioning of ssDNA and dsDNA using a modified pre-amplification target capture protocol (NEBNext Direct)

[0332] The novel hybrid capture method applied a pre-amplification hybrid capture targeted sequencing protocol (e.g., NEBNext Direct HotSpot Cancer Panel) to cell-free DNA (cfDNA) samples without DNA denaturation to capture ssDNA molecules (Figure 18).

[0333] The unbound fraction containing dsDNA molecules was isolated, denatured to ssDNA, and applied to the capture protocol.

[0334] The pre-amplification hybrid capture sequencing protocol used was the NEBNext Direct HotSpot Cancer Panel, which contains baits for 190 common cancer targets from 50 genes, encompasses approximately 40 kb of sequence, and contains over 18,000 COSMIC features (NEBNext Direct HotSpot Cancer Panel; neb.com / products / e7000-nebnext-direct-cancer-hotspot-panel). Briefly, the NEBNext Direct target enrichment approach rapidly hybridizes DNA samples to biotinylated oligonucleotide baits, which define the 3' end of each target of interest. The bait-target hybrids were bound to streptavidin beads, and 3' off-target sequences were removed using an enzymatic reaction. Subsequent library preparation converted the targets into Illumina-compatible libraries containing molecular tags and sample barcodes. The kit allowed for the capture of all ssDNA and dsDNA molecules in a sample by denaturing the DNA sample prior to hybridization with the bait.

[0335] cfDNA samples containing ss- and ds-cfDNA were subjected to a target capture protocol that omitted the pre-dsDNA denaturation step. The captured ssDNA molecules were prepared for NGS using the NEBNext protocol (left panel of Figure 20 ), and the supernatant from the capture was subjected to a second target capture protocol with a standard pre-dsDNA denaturation step and subsequently prepared for NGS (right panel of Figure 20 ). cfDNA extracted from plasma was quantified using electrophoresis-based measurements. Sample volumes equivalent to 200 ng or 500 ng were applied to the NEBNext Direct HotSpot Cancer Panel assay, which omitted the DNA denaturation step so that only ssDNA molecules hybridized to the bait. The supernatant from the capture, containing dsDNA molecules and non-target ssDNA molecules, was retained and subjected to a second target capture (Figure 20 ). Both ssDNA and dsDNA libraries were prepared separately for NGS, each with a unique sample barcode tag identified in downstream bioinformatics analysis. Both ssDNA and dsDNA prep libraries were sequenced on an Illumina NextSeq 500 (2 × 75 paired ends), and the total number of on-target molecules (corresponding to the 40 kb bait) was computationally calculated (Figure 1).

[0336] Cell-free DNA (cfDNA) samples containing both single-stranded cell-free DNA (ss-cfDNA) and double-stranded cell-free DNA (ds-cfDNA) were fractionated into ss-cfDNA and ds-cfDNA groups, respectively, using the method described above (Figure 20). In two of the sequenced samples, the ssDNA libraries contained approximately 80% dsDNA (on-target molecules, the first with 200 ng and the second with 500 ng cfDNA input). The second 200 ng cfDNA failed to generate both ssDNA and dsDNA libraries, indicating an expected error in sample processing upstream of the ssDNA / dsDNA partitioning process, while the first 500 ng cfDNA input generated only a significant dsDNA library, suggesting that the relative amounts of ssDNA and dsDNA in cfDNA samples are variable. On-target molecules were computed as defined by the Picard package from the Broad Institute (Picard metrics; broadinstitute.github.io / picard / picard-metric-definitions.html). PCR yields for this experiment are shown in Figure 20. The relative yield, ssDNA PCR yield / dsDNA PCR yield, was determined to be between 20% and 75% in all four samples.

[0337] Example 17 Sensitive detection of somatic mutations retained using MBD-based methylation partitioning method

[0338] Sample collection and pooling

[0339] Samples were selected from the Guardant Health repository, which demonstrated high cfDNA yields. Clinical samples were prepared by mixing 96 samples in equal volumes. This served as a test material for assay sensitivity for mutation detection, as the pool contained mutations from the reference genome ranging from <0.02% to 100%. Two different clinical samples (Powerpool V1 and Powerpool V2) with unique component samples were prepared.

[0340] DNA division

[0341] Powerpool cfDNA was divided into fractions: cfDNA (15 or 150 ng) was divided into highly, intermediately, and hypomethylated fractions using the MethylMiner™ affinity enrichment protocol (Thermo Fisher Scientific, catalog number ME10025), except that the reaction conditions were modified to use 300 mM NaCl incubation and wash buffers, and the 1 microgram DNA input protocol was linearly scaled down to submicrogram amounts of DNA input.

[0342] Bead preparation

[0343] Washing Dynabeads® M-280 Streptavidin

[0344] Dynabeads® M-280 Streptavidin was washed using 1x Bind / Wash buffer (containing 160 mM NaCl) prior to coupling with MBD-Biotin protein. Briefly, Dynabeads® M-280 Streptavidin stock was resuspended to a homogenous suspension. For each microgram of input DNA, 10 μl of beads was added to a 1.7 ml DNase-free microcentrifuge tube. The bead volume was brought to 100 μl with 1x Bind / Wash buffer. The tube was placed on a magnetic rack for 1 minute to allow all beads to collect on the inner wall of the tube before removing and discarding the liquid. The tube was removed from the magnetic rack and an equal volume (e.g., approximately 100–250 μl) of 1x Bind / Wash buffer was added to resuspend the beads. The resuspended beads were collected and washed once more before subsequent coupling of MBD-Biotin protein to the beads.

[0345] Dynabeads® M-280 streptavidin is coupled to the MBD-biotin protein

[0346] For each microgram of input DNA, 7 μl (3.5 μg) of MBD-biotin protein was added to a 1.7 ml DNase-free microcentrifuge tube. The bead volume was brought to 100 μl with 1x Bind / Wash buffer (containing 300 mM NaCl). The MBD-biotin protein was diluted and transferred to the tube of resuspended beads from the first bead wash. The bead-protein mixture was mixed for 1 hour at room temperature on a rotating plate mixer before continuing to wash the MBD-beads.

[0347] Washing MBD-beads

[0348] The tube containing the MBD-beads was collected by placing the MBD-beads on a magnetic rack for 1 minute. The liquid was removed and discarded. The beads were resuspended in 100-250 μl of 1x Bind / Wash buffer (containing 160 mM NaCl) and mixed on a rotary mixer for 5 minutes at room temperature. The beads were collected, washed, and resuspended as above two more times. The tube was then placed on a magnetic rack for 1 minute, and the liquid was carefully removed and discarded. The beads were resuspended in 10 μl of 1x DNA capture buffer (containing 300 mM NaCl) for each μl of streptavidin beads used.

[0349] Capture of fragmented methylated DNA on MBD-beads

[0350] Incubate MBD-beads with fragmented DNA

[0351] Generally, input DNA can range from 5 ng to 1 μg. Control reactions typically used 1 μg of K-562 DNA. To a clean 1.7 ml DNase-free microcentrifuge or PCR tube, add fragmented sample DNA (e.g., 5 ng to 1 μg) along with an equal volume of 2x DNA capture buffer (containing 300 mM NaCl), bringing the final volume to 100 or 200 μl with 1x DNA capture buffer. The DNA / buffer mixture was transferred to the tube containing the MBD-beads and mixed for 1 hour at room temperature on a rotating mixer. Alternatively, the mixture can be mixed overnight at 4°C.

[0352] Collecting uncaptured DNA from the bead solution

[0353] Uncaptured / unmethylated DNA was collected from the DNA and MBD-bead mixture. Briefly, the tube containing the DNA and MBD-bead mixture was placed on a magnetic rack for 1 minute to collect all beads, and the supernatant was removed and saved in a clean, DNase-free microcentrifuge tube. This saved supernatant is the uncaptured DNA / unmethylated DNA fraction and can be stored on ice. The beads were washed with 200 μl of 1x DNA capture buffer (containing 300 mM NaCl) for 3 minutes on a rotary mixer. The beads were collected as above, and the supernatant containing uncaptured / unmethylated / hypomethylated DNA was removed and saved and stored on ice as above. The beads were washed, mixed, and collected, and the supernatant was removed and saved once more to collect two wash fractions. Each wash fraction was stored on ice. The wash fractions can be pooled together and labeled appropriately.

[0354] Eluting the captured DNA

[0355] The captured DNA was eluted using elution buffer containing 2000 mM NaCl. The beads were resuspended in 200 μl of elution buffer (2000 mM NaCl). The beads were incubated on a rotary mixer for 3 minutes, placed on a magnetic rack for 1 minute to collect all the beads, and the liquid containing the captured / hypermethylated DNA was removed and saved in a clean DNase-free microcentrifuge tube. The first saved fraction was kept on ice. The beads were resuspended and incubated once more, and the liquid containing the captured / methylated DNA was removed and saved in a second clean tube. The first and second collections of captured / methylated DNA were pooled and saved on ice. Alternatively, multiple elutions using increasing NaCl concentrations can be performed to further partition the DNA into fractions with increasing DNA methylation.

[0356] Preparation of methylated fractionated DNA for analysis

[0357] The methylated, intermediately methylated, and unmethylated cfDNA fragments were purified, e.g., by SPRI bead cleanup (Ampure XP, Beckman Coulter), and then prepared for ligation (using the NEBNext® Ultra™ End Repair / dA-Tailing Module). They were then ligated with modified Y-shaped dsDNA adapters containing non-random molecular barcodes as described in Lanman et al. (2015). The hypermethylated, intermediately methylated, and hypomethylated cfDNA fragments were ligated with 11, 12, and 12 different non-random molecular barcode adapters, respectively. After ligation, the ligated and split cfDNA molecules for each sample were purified again with SPRI beads (Ampure XP) and then recombined in a PCR reaction using universal oligos (NEBNext Ultra II™ Q5 master mix) for all adapter-ligated molecules, amplifying all cfDNA molecules from a single sample together. The amplified DNA library was purified again using SPRI beads (Ampure XP) in preparation for target enrichment by hybrid capture (Agilent SureSelect 30 kb panel; "panel").

[0358] Preparation of unsplit DNA for analysis

[0359] Powerpool cfDNA (10 or 150 ng) was prepared for ligation (using the NEBNext® Ultra™ End Repair / dA-Tailing Module) and then ligated with modified Y-shaped dsDNA adapters containing non-random molecular barcodes as described in Lanman et al. (2015). cfDNA was ligated with 35 different non-random molecular barcode adapters. The ligated cfDNA molecules for each sample were re-purified with SPRI beads (Ampure XP) and then subjected to PCR reactions using universal oligos (NEBNext Ultra II™ Q5 master mix) for all adapter-ligated molecules, amplifying all cfDNA molecules from a single sample together. The amplified DNA library was again purified using SPRI beads (Ampure XP) in preparation for target enrichment by hybrid capture (Agilent SureSelect 30 kb panel; "panel").

[0360] The present disclosure provides methods for processing nucleic acid populations containing different forms (e.g., RNA and DNA, single-stranded or double-stranded) and / or degrees of modification (e.g., cytosine methylation, protein association). These methods accommodate multiple forms and / or modifications of nucleic acids in a sample so that sequence information can be obtained for multiple forms. The methods also preserve the identity of multiple forms or modification states throughout processing and analysis so that sequence analysis can be combined with epigenetic analysis.

[0361] Data analysis

[0362] DNA libraries from different samples were pooled and sequenced on an Illumina HiSeq2500, 2x150 paired-end sequencing. Bioinformatics processing was performed using the standard GUARDANT360™ protocol described in Lanman et al., 2015, and elsewhere. For MBD-split samples, molecular barcodes were further used to identify MBD splits into which DNA was fractionated (hypermethylated, intermediately methylated, and hypomethylated). For each genomic locus targeted by the panel, hypermethylated, intermediately methylated, and hypomethylated aligned molecules were summed. % hypermethylation was defined as the proportion of total molecules spanning the locus that were hypermethylated at a given locus. For both MBD-split and unsplit DNA samples, the mutant allele fraction (MAF) from the reference genome was called in the target regions using proprietary Guardant Health variant calling software.

[0363] Example 18 Comparison between coverage for MBD and non-MBD samples in targeted sequencing assays

[0364] In this example, samples were processed as described in Example 17. Different clinical samples of cfDNA (Powerpool V1 and Powerpool V2) were assayed in triplicate in a targeted sequencing assay with and without MBD-splitting, "MBD" and "non-MBD," respectively. The unique molecules sequenced at each target genomic location for genes from the panel were compared in MBD and non-MBD with an assay input of 15 ng (FIG. 25A) and 150 ng (FIG. 25B) for Powerpool V1. The panel is a custom gene panel covering an approximately 30 kb genomic region. The panel also has high sensitivity for detecting different cancers, such as lung cancer, colorectal cancer, etc. Figures 25A and 25B show the high-efficiency recovery of molecules in the targeted sequencing assay retained in the use of MBD-splitting. The number of molecules from the targeted sequencing assay with inputs of 15 ng and 150 ng of powerpool V1 (a) and (b) increased with or without MBD-splitting (Y-axis). A linear correlation was observed between the number of MBD to non-MBD molecules or coverage, indicating that MBD partitioning does not bias the recovery of the assay.

[0365] The number of molecules or coverage for genes from the panel was compared between non-MBD and MBD samples. MBD and non-MBD samples were prepared using 15 ng input cfDNA extracted from two clinical samples (Figure 26A-PowerpoolV1 and Figure 26B-PowerpoolV2) or 150 ng input cfDNA extracted from two clinical samples (Figure 27A-PowerpoolV1; Figure 27B-PowerpoolV2). The X-axis of the left graph represents the number of molecules or coverage, the X-axis of the center graph represents the mutations confirmed in both paired-end read data (double-stranded overlap; DSO), and the X-axis of the left graph represents the number of molecules where both DNA strands are sequenced (double-stranded support; DS). The strong correlation between MBD and non-MBD samples for molecule count, DSO, and DS indicates that MBD samples are able to capture the majority of molecules compared to non-MBD (approximately 94% in Figure 26A, approximately 80-85% in Figures 26B and 27A, and approximately 90% in Figure 27B). There is no positional bias in molecular coverage, as well as other important variant calling metrics (DSO, DS), across panels.

[0366] Example 19 Sensitivity and specificity of variant detection in MBD and non-MBD samples

[0367] In this example, samples were processed as described in Example 17. To measure the impact on variant or mutation detection in terms of sensitivity and specificity, mutant allele fractions (MAFs) were compared between MBD (Y-axis) and non-MBD (X-axis) samples for genes in the panel using 15 ng input cfDNA. Different MAF ranges, e.g., 0-100% (Figure 28A), 0-5% (Figure 28B), and 0-0.5% (Figure 28C), were plotted on the X-axis. MAF values ​​were derived from triplicate samples of MBD and non-MBD origin. The MAFs determined for MBD samples were consistent with those determined for non-MBD samples. The MAFs between MBD and non-MBD showed a linear correlation for Powerpool V1 at 15 ng input (Figure 28A; 0-100%) and at the low detection limit (Figure 28B; 0-5%). The MAF between MBD and non-MBD was not well correlated below the detection limit (Figure 28C; 0-0.5% MAF). Similarly, MBD and non-MBD samples showed concordance in MAF at 150 ng cfDNA input from Powerpool V1 (Figures 29A and 29B), but there was no strong agreement in the 0-0.5% range (Figure 29C).

[0368] Example 20 Methylation profiling of promoter regions using whole genome sequencing

[0369] Molecularly partitioned samples can enhance the analysis of genomic structures, such as cell-free DNA fragment occupancy and detection in cancer. For example, transcription-associated hypermethylation events can be detected by considering cell-free DNA fragment occupancy when analyzing promoter regions of tumor suppressor genes, which are commonly targeted by cancers through methylation-driven gene silencing. Examining cell-free DNA fragment occupancy signals and hypermethylated fractions together in different MBD partitions can confirm the feasibility of MBD-driven discovery of transcription-associated hypermethylation events and gene silencing in cancer samples.

[0370] As an illustrative example, the publicly available gencode (v26lift3 7) The data can be used to generate percent hypermethylation (number of fragments in hypermethylated aliquots / total number of fragments in all MBD aliquots) in TSS regions of all gene-coding genes in available cohorts of non-malignant healthy adults. Cell-free DNA fragment occupancy signals can be collected across cohorts of non-malignant healthy adults. All TSSs may be binned based on the percent hypermethylated fraction observed in the MBD aliquot assay. Fragment occupancy of the non-MBD WSG cohort in each bin. The percentage of methylated nucleic acid fragments can be examined. Figure 23 shows the correlation between gene expression and methylation status. WGS occupancy versus percent MBD methylation in promoter profiles is shown. As seen in Figure 23, hypomethylated DNA (0-0.1% high) has low fragment occupancy coverage near the TSS, while hypermethylated DNA (10-50% high or >50% high) has high fragment occupancy coverage near the TSS and distinct NDRs. In some cases, the fragment occupancy coverage of hypomethylated DNA is used to normalize sequence depth and / or sequence mappability. The percentage of hypermethylated or hypomethylated nucleic acid fragments can be determined by dividing the number of hypermethylated or hypomethylated cell-free fragments by the total number of cell-free DNA fragments observed across all partitions.

[0371] Example 21 Comparison between methylation levels in MBD samples and whole genome bisulfite sequencing (WGBS) samples

[0372] To assess the methylation levels of fragments in various aliquots prepared using the MBD protocol, a well-characterized sample, NA12878 (catalog.coriell.org / 0 / Sections / Search / Sample_Detail.aspx?Ref=GM12878), was used. The sample was partitioned into highly, hypo, and intermediately methylated portions, followed by in silico recombination of the aliquots (MBD sample) as described in Example 1. The MBD sample was compared to a publicly available standard methylation dataset (basespace.illumina.com / datacentral (HiSeq 4000: TruSeq DNA Methylation (NA12878, 1x151))) using whole-genome bisulfite sequencing (WGBS). WGBS examines the methylation status of individual cytosines. Figure 31 shows the correlation of average methylation levels as measured by WGBS (X-axis) and MBD (Y-axis) in 160-bp windows. MBD methylation levels were computed by dividing the number of reads in the hypermethylated aliquots that fell within that partitioned window by the total number of reads in the hyper- and hypomethylated aliquots. WGBS methylation levels were computed by dividing the number of methylated bases in the window by the number of methylated and unmethylated bases. This experiment was performed over several different bead ratios, and the bead ratio affects the partitioning of methylated fragments. Fewer beads restrict the hypermethylated aliquots to highly methylated fragments (i.e., making the assay more specific for methylation), while more beads reduce the amount of methylation required to place a fragment in the hypermethylated aliquot (i.e., making the assay more sensitive to methylation). Empirically, an input DNA:bead ratio of 1:50 was found to correlate well between the partitioned fragments and their methylation levels. These results demonstrate that the MBD partitioning does accurately reflect the underlying methylation state of a sample.

[0373] In this analysis, we assessed the effect of the number of CG sites in a fragment on its segmentation. Publicly available fragments with a standard methylation dataset (NA12878; the same as in the previous analysis) showing very high or low methylation (whole-genome bisulfite sequencing methylation levels >90% or <10%, as calculated in the previous analysis) were selected for analysis. These fragments were stratified by the number of CG sites they contained. Highly methylated fragments with three or more CG sites ultimately fell into the high-methylation segment, demonstrating that the assay is sensitive to small amounts of methylation (Figure 31A). Conversely, fragments lacking methylation were primarily segmented into the low-methylation segment, regardless of the number of CG sites in the fragment, demonstrating the high specificity of the assay (Figure 31B).

[0374] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous modifications, changes, and substitutions will now occur to those skilled in the art without departing from the present disclosure. It is to be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the present disclosure. It is intended that the following claims define the scope of the disclosure, and that methods and structures within the scope of these claims and equivalents thereof be covered thereby.

[0375] Some embodiments of the present invention

[0376] Certain embodiments of the present invention, presented in claim form, are provided below. 1. A method for analyzing a population of nucleic acids comprising at least two forms of nucleic acid selected from double-stranded DNA, single-stranded DNA, and single-stranded RNA, each of the at least two forms comprising a plurality of molecules, the method comprising: (a) linking at least one of the forms of nucleic acid to at least one tag nucleic acid to distinguish the forms from one another; (b) amplifying forms of nucleic acid, at least one of which is linked to at least one nucleic acid tag, wherein the nucleic acid and the linked nucleic acid tag are amplified to produce amplified nucleic acids, of which at least one form amplified is tagged; (c) assaying sequence data of the amplified nucleic acids, wherein at least some of the amplified nucleic acids are tagged, the assaying providing sufficient sequence information to decode tag nucleic acid molecules of the amplified nucleic acids, revealing the form of the nucleic acids in the population, and providing original templates of the amplified nucleic acids whose sequence data are linked to the assayed tag nucleic acid molecules; A method comprising: 1A. The method of claim 1, further comprising decoding tag nucleic acid molecules of the amplified nucleic acids to reveal the form of the nucleic acids in the population and provide original templates of the amplified nucleic acids with sequence data linked to the tag nucleic acid molecules assayed. 2. The method of claim 1, further comprising enriching at least one of the forms relative to one or more of the other forms. 3. The method of claim 1, wherein at least 70% of the molecules of each form of nucleic acid in the population are amplified in step (b). 4. The method of claim 1, wherein at least three forms of the nucleic acid are present in the population, and at least two of the forms are linked to different tag nucleic acid forms that distinguish each of the three forms from one another. 5. The method of claim 4, wherein each of the at least three forms of nucleic acid in the population is linked to a different tag. 6. The method of claim 1, wherein each molecule of the same form is linked to a tag containing the same identifying information tag. 7. The method of claim 1, wherein molecules of the same type are linked to tags of different types. 8. The method of claim 1, wherein step (a) comprises subjecting the population to reverse transcription using tagged primers, wherein the tagged primers are incorporated into cDNA made from RNA in the population. 9. The method of claim 8, wherein the reverse transcription is sequence-specific. 10. The method of claim 8, wherein the reverse transcription is random. 11. The method of claim 8, further comprising a step of degrading RNA that forms a duplex with the cDNA. 12. The method of claim 4, further comprising the steps of separating single-stranded DNA from double-stranded DNA and ligating nucleic acid tags to the double-stranded DNA. 13. The method of claim 12, wherein the single-stranded DNA is separated by hybridization with one or more capture probes. 14. The method of claim 4, further comprising the steps of circularizing the single-stranded DNA with circligase and ligating a nucleic acid tag to the double-stranded DNA. 15. The method of claim 1, comprising pooling tagged nucleic acids comprising different forms of nucleic acid prior to the assaying step. 16. A method according to any preceding claim, wherein the nucleic acid population is derived from a body fluid sample. 17. The method of claim 16, wherein the body fluid sample is blood, serum or plasma. 18. The method of claim 1, wherein the population of nucleic acids is a cell-free population of nucleic acids. 19. The method of claim 17, wherein the body fluid sample is derived from a subject suspected of having cancer. 20. The method of claims 1 to 19, wherein the sequence data indicates the presence of a somatic or germline mutation. 21. The method of claims 1 to 20, wherein the sequence data indicates the presence of copy number variations. 22. The method of claims 1 to 21, wherein the sequence data indicates the presence of single nucleotide variations (SNVs), indels, or gene fusions. 23. A method for analyzing a population of nucleic acids comprising nucleic acids with different degrees of modification, comprising: contacting the population of nucleic acids with an agent that preferentially binds nucleic acids having the modification; separating nucleic acids of a first pool that are bound to the agent from nucleic acids of a second pool that are not bound to the agent, wherein the nucleic acids of the first pool are over-represented for the modification and the nucleic acids in the second pool are under-represented for the modification; linking the nucleic acids in the first pool and / or the second pool with one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool to produce a population of tagged nucleic acids; amplifying the labeled nucleic acid, wherein the nucleic acid and the linked tag are amplified; assaying sequence data of the amplified nucleic acids and linked tags, the assaying providing sequence data for decoding the tags to determine whether the nucleic acids whose sequence data were assayed were amplified from templates in the first or second pool; A method comprising: 23A. The method of claim 23, comprising decoding the tag to reveal whether the nucleic acid for which the sequence data was assayed was amplified from a template in the first or second pool. 24. The method of claim 23, wherein the modification is the binding of the nucleic acid to a protein. 25. The method of claim 23, wherein the protein is a histone or a transcription factor. 26. The method of claim 23, wherein the modification is a post-replication modification to a nucleotide. 27. The method of claim 26, wherein the post-replicative modification is 5-methyl-cytosine and the degree of binding of the agent to the nucleic acid increases with the degree of 5-methyl-cytosine in the nucleic acid. 28. The method of claim 26, wherein the post-replicative modification is 5-hydroxymethyl-cytosine and the extent of binding of the agent to the nucleic acid increases with the extent of 5-hydroxymethyl-cytosine in the nucleic acid. 29. The method of claim 26, wherein the post-replicative modification is 5-formyl-cytosine or 5-carboxyl-cytosine, and the degree of binding of the agent increases with the degree of 5-formyl-cytosine or 5-carboxyl-cytosine in the nucleic acid. 30. The method of claim 23, further comprising washing nucleic acids bound to the agent and recovering the wash as a third pool containing nucleic acids having an intermediate degree of post-replication modification relative to the first and second pools. 31. The method of claim 23, comprising pooling the tagged nucleic acids from the first and second pools prior to the assaying step. 32. The method of claim 23, wherein the agent is a 5-methyl binding domain magnetic bead. 33. A method according to any preceding claim, wherein the nucleic acid population is derived from a body fluid sample. 34. The method of claim 33, wherein the body fluid sample is blood, serum or plasma. 35. The method of claim 23, wherein the population of nucleic acids is a cell-free population of nucleic acids. 36. The method of claim 33, wherein the body fluid sample is derived from a subject suspected of having cancer. 37. The method of claims 23 to 36, wherein the sequence data indicates the presence of a somatic or germline mutation. 38. The method of claims 23 to 37, wherein the sequence data indicates the presence of copy number variations. 39. The method of any of 23 to 38, wherein the sequence data indicates the presence of a single nucleotide variation (SNV), an indel, or a gene fusion. 40. A method for analyzing a population of nucleic acids, at least some of the nucleic acids comprising one or more modified cytosine residues, comprising: Ligating a capture moiety to nucleic acids in the population that serve as templates for amplification; performing an amplification reaction to produce an amplification product from the template; Separating the capture tag-linked templates from the amplification products; Assaying the sequence data of the templates linked to the capture tags by bisulfite sequencing; Assaying the sequence data of the amplification products; A method comprising: 41. The method of claim 40, wherein the capture moiety comprises biotin. 42. The method of claim 41, wherein the separating step is carried out by contacting the template with streptavidin beads. 43. The method of claim 40, wherein the modified cytosine residue is 5-methyl-cytosine, 5-hydroxymethylcytosine, 5-formylcytosine or 5-carboxylcytosine. 44. The method of claim 40, wherein the capture moiety comprises biotin linked to a nucleic acid tag comprising one or more modified residues. 45. The method of claim 40, wherein the capture moieties are linked to the nucleic acids in the population by cleavable linkages. 46. ​​The method of claim 45, wherein the cleavable linkage is a photocleavable linkage. 47. The method of claim 45, wherein the cleavable linkage comprises a uracil nucleotide. 48. A method according to any preceding claim, wherein the nucleic acid population is derived from a body fluid sample. 49. The method of claim 48, wherein the body fluid sample is blood, serum, or plasma. 50. The method of claim 40, wherein the population of nucleic acids is a cell-free population of nucleic acids. 51. The method of claim 48, wherein the body fluid sample is derived from a subject suspected of having cancer. 52. A method according to any preceding claim, wherein the sequence data indicates the presence of a somatic or germline mutation. 53. A method according to any preceding claim, wherein the sequence data indicates the presence of copy number variation. 54. A method according to any preceding claim, wherein the sequence data indicates the presence of single nucleotide variations (SNVs), indels or gene fusions. 55. A method for analyzing a nucleic acid population containing nucleic acids with different degrees of 5-methylation, comprising: (a) contacting a population of nucleic acids with an agent that preferentially binds 5-methylated nucleic acids; (b) separating nucleic acids of a first pool that are bound to the agent from nucleic acids of a second pool that are not bound to the agent, wherein the nucleic acids of the first pool are over-represented for 5-methylation and the nucleic acids in the second pool are under-represented for 5-methylation; (c) linking the nucleic acids in the first pool and / or the second pool to one or more nucleic acid tags that distinguish the nucleic acids in the first pool and the second pool, wherein the nucleic acid tags linked to the nucleic acids in the first pool comprise a capture moiety (e.g., biotin); (d) amplifying the labeled nucleic acid, wherein the nucleic acid and the linked tag are amplified; (e) separating the amplified nucleic acids having the capture moiety from the amplified nucleic acids not having the capture moiety; (f) assaying the sequence data of the separated, amplified nucleic acids; A method comprising: 56. A method for analyzing a population of nucleic acids comprising nucleic acids with different degrees of modification, comprising: contacting the nucleic acids in the population with adaptors to produce a population of nucleic acids flanked by adaptors that contain primer binding sites; amplifying adapter-flanked nucleic acids primed from the primer binding sites; contacting the amplified nucleic acid with an agent that preferentially binds nucleic acid having the modification; separating nucleic acids of a first pool that are bound to the agent from nucleic acids of a second pool that are not bound to the agent, wherein the nucleic acids of the first pool are over-represented for the modification and the nucleic acids in the second pool are under-represented for the modification; performing parallel amplification of the tagged nucleic acids in the first and second pools; assaying the sequence data of the amplified nucleic acids in the first and second pools; A method comprising: 57. A method for analyzing a population of nucleic acids, at least some of the nucleic acids comprising one or more modified cytosine residues, comprising: contacting the population of nucleic acids with adaptors that include primer binding sites that include modified cytosines to form adaptor-flanked nucleic acids; amplifying adaptor-flanked nucleic acids primed from primer binding sites in the adaptors flanking the nucleic acids; dividing the amplified nucleic acid into first and second aliquots; assaying the sequence data for the first aliquot of nucleic acid; contacting a second aliquot of nuc...

Claims

1. (a) physically fractionating DNA molecules from a human sample to create two or more aliquots, wherein the physical fractionating comprises fractionating using methyl-binding domain protein (MBD) beads to stratify different degrees of methylation, including hypermethylated and hypomethylated; (b) applying differential molecular tags and NGS-enabling adaptors to each of the two or more aliquots to generate molecularly tagged aliquots, wherein the differential molecular tags are different sets of molecular tags corresponding to the MBD-aliquots; (c) assaying the molecularly tagged aliquots in an NGS instrument to generate sequence data for deconvoluting the sample into differentially resolved molecules; A method comprising:

2. 10. The method of claim 1, further comprising analyzing the sequence data by deconvolving the sample into differentially resolved molecules.

3. The method of claim 1 , wherein the DNA molecule is derived from a body fluid sample.

4. 2. The method of claim 1, wherein the differential molecular tags are different sets of molecular tags corresponding to MBD-splits.

5. 10. The method of claim 1, wherein the physical fractionation comprises separating DNA molecules using immunoprecipitation.

6. The method described in claim 1, further comprising a step of recombining the molecularly tagged divided fragments that have been prepared, the divided fragments being tagged with two or more types of molecular tags.

7. 10. The method of claim 1, further comprising amplifying the molecularly tagged aliquots.

8. The method of claim 3 , wherein the DNA molecule is cell-free DNA.

9. The method of claim 3 , wherein the body fluid sample is blood, serum, or plasma.

10. 4. The method of claim 3, wherein the body fluid sample is from a subject suspected of having cancer.

11. 10. The method of claim 1, further comprising enriching the recombined molecularly tagged aliquots for a target genomic region of interest.

12. 12. The method of claim 11, wherein the enrichment comprises using a bait set comprising oligonucleotide baits labeled with a capture moiety, such as biotin.

13. 13. The method of claim 12, wherein the bait set has a higher relative concentration for more specifically desired sequences of interest.

14. 10. The method of claim 1, further comprising classifying the object based on one or more characteristics using a trained classifier.

15. 9. The method of claim 8, wherein the cell-free DNA comprises circulating tumor DNA.

16. 10. The method of claim 1, wherein the DNA molecules from the human sample are sequenced to a data depth of 1,000 to 50,000 reads per locus.

17. 10. The method of claim 1, wherein the sequence data from the second aliquot is from at least 50,000 sequencing reactions.

18. 10. The method of claim 1, wherein the sequence data from the second aliquot has sequence coverage of at least 20 different genes.

Citation Information

Patent Citations

  • Processes and compositions for enriching maternal sample-derived fetal nucleic acids based on methylation, useful for non-invasive prenatal diagnosis.

    JP2013505019A

  • Processes and compositions for enriching maternal sample-derived fetal nucleic acids based on methylation, useful for non-invasive prenatal diagnosis.

    JP2015521862A

  • Varietal counting of nucleic acids for obtaining genomic copy number information

    WO2012054873A2

  • A method and kit for determining the tissue or cell origin of DNA

    WO2015159292A2

  • Method and system for determining cancer status

    WO2016115530A1