Methods for simultaneous mutation detection and methylation analysis

JP2024529674A5Pending Publication Date: 2025-08-19JOHNS HOPKINS UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024508357
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-12
Filing Date
2022-08-12
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Current nucleic acid sequencing technologies struggle to accurately detect rare mutations and epigenetic changes, such as cytosine methylation, due to high error rates and the need for separate aliquots for genetic and epigenetic analysis, which reduces sensitivity by 50%.

Method used

A method involving attaching adapter fragments with molecular barcodes to double-stranded DNA molecules, copying both strands, denaturing conditions, and separately recovering strands for sequencing to identify genetic and epigenetic characteristics, enabling simultaneous detection of mutations and methylation patterns from a single aliquot.

Benefits of technology

Enhances the accuracy and sensitivity of identifying rare mutations and epigenetic changes by improving library preparation and workflow, allowing for high-confidence differentiation between true mutations and artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0001_ABST
    Figure 00000000_0001_ABST
  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are methods for identifying the genetic, fragmentation, and epigenetic characteristics of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying both strands of the double-stranded DNA molecule. TIFF2024529674000004.tif119170
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 232,438, filed August 12, 2021. The disclosure of this prior application is considered part of the disclosure of this application and is incorporated herein in its entirety.

[0002] Technical Field The present disclosure relates to the field of nucleic acid analysis. In particular, the present disclosure relates to nucleic acid sequence analysis capable of detecting mutations and methylation of nucleic acid sequences.

[0003] Federally Sponsored Research and Development This invention was made with government support under grants GM136577, GM008752 and CA006973 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]

[0004] background Identification of rare mutations is useful not only in aspects of basic biology but also to improve clinical management of patients. Applications include infectious diseases, immune repertoire profiling, paleogenetics, forensics, aging, non-invasive prenatal testing, and cancer. Next-generation sequencing (NGS) technology is theoretically suitable for this application, and various NGS approaches exist for the detection of rare mutations. However, for conventional NGS approaches, the error rate of the sequencing itself is too high to confidently detect mutations, especially those present at low frequencies in the original sample.

[0005] The use of molecular barcodes to tag the original template molecules was devised to overcome various obstacles in the detection of rare mutations. With molecular barcoding, redundant sequencing of the PCR-generated products of each tagged molecule is performed, and sequencing errors are easily recognized. For example, if the same mutation falls within a given threshold of the products of the barcoded template molecule, the mutation is considered genuine. If the mutation of interest falls below a given threshold of the products, the mutation is considered an artifact. Two types of molecular barcodes have been described: exogenous and endogenous. Exogenous barcodes (also referred to as exogenous unique identifiers, or "UIDs") contain pre-specified or random nucleotides and are added during library preparation or PCR. Endogenous barcodes (also referred to as endogenous UIDs) are formed by sequences present in the template DNA being assayed, for example, fragments generated by random shearing of DNA, or fragments present in cell-free liquid biological samples. In some cases, the endogenous barcode is a sequence present at the 5' and / or 3' end of the fragment. Such barcodes have proven useful in tracing amplicons back to the original starting template, enabling molecular counting and improving identification of true mutations in clinically relevant samples.

[0006] Identification of rare epigenetic changes in DNA, such as those associated with cytosine methylation or hydroxymethylation, presents similar challenges as described above for mutations. Currently, experimental techniques that assess mutations in plasma (for example) with very high specificity require one aliquot of plasma. Experimental techniques that assess epigenetic changes require another aliquot of plasma. Because certain mutations and epigenetic changes in the plasma of early stage cancer patients are rare, it would be advantageous to assess as many molecules as possible for both genetic and epigenetic changes. Splitting the sample, half for genetic changes and half for epigenetic changes, reduces the sensitivity by 50% for both types of changes.

[0007] Therefore, there is a need to improve sequencing library preparation and workflows to enable accurate identification of mutations, e.g., rare mutations and epigenetic changes, from the same aliquot of DNA purified from clinically relevant samples, such as, but not limited to, plasma. Summary of the Invention

[0008] overview Provided herein is a method for identifying genetic and epigenetic characteristics of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule, the method comprising: (a) attaching an adapter fragment to each end of the double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) contacting the adapted double-stranded DNA molecule with a tagged primer. (c) subjecting the amplified products to denaturing conditions; (d) separately recovering the adapted Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson and Crick strands and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (f) generating a second population of analyte DNA fragments from the adapted Watson and Crick strands and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on the at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing reads according to the molecular barcode present on said at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family;(i) identifying genetic features of the tagged Watson and Crick strands in a first analyte DNA family; and (j) identifying epigenetic features of the matched Watson and Crick strands in a second analyte DNA family, thus identifying genetic and epigenetic features present on at least one strand of the double-stranded DNA molecule;

[0009] In some embodiments, the adapter fragment further comprises a sample barcode. In some embodiments, the molecular barcode comprises an intrinsic barcode, an exogenous barcode, or both.

[0010] In some embodiments, the copying step (b) comprises performing one, two, or three rounds of linear extension of the adapted double-stranded DNA molecule. In some embodiments, the tagged primer is a uracil-containing biotinylated primer, and the tagged Watson strand and Crick strand are generated from the uracil-containing biotinylated primer.

[0011] In some embodiments, the recovering step (d) comprises contacting the tagged Watson and Crick strands with streptavidin-functionalized beads, and wherein the tagged Watson and Crick strands bind to the streptavidin-functionalized beads. In some embodiments, the recovered, matched Watson and Crick strands that are not bound to the streptavidin-functionalized beads are treated with bisulfite to convert cytosine bases to uracil bases to generate a second population of analyte DNA fragments comprising a population of converted DNA molecules.

[0012] In some embodiments, the denaturing conditions comprise NaOH denaturation. In some embodiments, the denaturing conditions comprise thermal denaturation, chemical denaturation, or a combination thereof. In some embodiments, the generating steps (e) and (f) are performed under PCR conditions.

[0013] In some embodiments, the genetic feature is a mutation. In some embodiments, the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and combinations thereof. In some embodiments, the epigenetic feature is methylation. In some embodiments, the epigenetic feature is a methylation pattern. In some embodiments, the methylation pattern corresponds to a methylation pattern present in cells generated via clonal hematopoiesis of indeterminate origin. In some embodiments, the methylation pattern corresponds to a methylation pattern present in a tissue of origin. In some embodiments, the tissue of origin is anal, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus. In some embodiments, the epigenetic feature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation. In some embodiments, the methods identify the genetic and epigenetic characteristics of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying both strands of the double-stranded DNA molecule.

[0014] Also provided herein is a method for identifying a first feature and a second feature of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule, the method comprising: (a) attaching an adapter fragment to each end of the double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand; (c) subjecting the amplified products to denaturing conditions; (d) separately recovering the matched Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson and Crick strands and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (f) generating a second population of analyte DNA fragments from the matched Watson and Crick strands and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on the at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing read according to the molecular barcode present on the at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family;(i) identifying a first feature of the tagged Watson and Crick strands in a first analyte DNA family; and (j) identifying a second feature of the matched Watson and Crick strands in a second analyte DNA family, thus identifying the first and second features present on at least one strand of the double-stranded DNA molecule;

[0015] In some embodiments, the adapter fragment further comprises a sample barcode. In some embodiments, the molecular barcode comprises an intrinsic barcode, an exogenous barcode, or both.

[0016] In some embodiments, the copying step (b) comprises performing one, two, or three rounds of linear extension of the adapted double-stranded DNA molecule. In some embodiments, the tagged primer is a uracil-containing biotinylated primer, and the tagged Watson strand and Crick strand are generated from the uracil-containing biotinylated primer.

[0017] In some embodiments, the recovering step (d) comprises contacting the first single-stranded DNA fragment with streptavidin-functionalized beads, and wherein the first single-stranded DNA fragment binds to the streptavidin-functionalized beads. In some embodiments, the denaturing conditions comprise NaOH denaturation. In some embodiments, the denaturing conditions comprise thermal denaturation, chemical denaturation, or a combination thereof.

[0018] In some embodiments, the generating steps (e) and (f) are performed under PCR conditions. In some embodiments, the generating step utilizes whole genome PCR, whole genome bisulfite sequencing, or capture sequencing.

[0019] In some embodiments, the first characteristic is a genetic characteristic or an epigenetic characteristic. In some embodiments, the second characteristic is an epigenetic characteristic or an epigenetic characteristic. In some embodiments, the first characteristic and the second characteristic are both genetic characteristics. In some embodiments, the first characteristic and the second characteristic are both epigenetic characteristics.

[0020] In some embodiments, the genetic feature is a mutation. In some embodiments, the mutation is selected from the group consisting of insertion, deletion, substitution, deletion-insertion, duplication, inversion, frameshift, repeat expansion, translocation, and combinations thereof. In some embodiments, the identification of genetic feature comprises mutation analysis, aneuploidy analysis, or fragmentomics.

[0021] In some embodiments, the epigenetic signature is methylation. In some embodiments, the epigenetic signature is a methylation pattern. In some embodiments, the methylation pattern corresponds to the methylation pattern present in cells generated through clonal hematopoiesis of indeterminate origin. In some embodiments, the methylation pattern corresponds to the methylation pattern present in the tissue of origin. In some embodiments, the tissue of origin is anus, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus. In some embodiments, the epigenetic signature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation.

[0022] In some embodiments, the methods identify a first characteristic and a second characteristic of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying both strands of the double-stranded DNA molecule.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.Methods and materials similar or equivalent to those described herein can be used to practice the present invention, and suitable methods and materials are described below.All publications, patent applications, patents and other references mentioned herein are incorporated by reference in their entirety.In case of conflict, the present specification, including definitions, shall prevail.In addition, the materials, methods and examples are illustrative only and are not intended to be limiting.

[0024] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will become apparent from the description and drawings, and from the claims. [Brief description of the drawings]

[0025] [Figure 1] 1 shows an exemplary workflow for simultaneous mutation detection and methylation analysis. [Diagram 2] 1 shows duplex recovery according to the workflow described herein. [Diagram 3] 1 shows an exemplary workflow for simultaneous mutation detection and methylation analysis. [Figure 4] FIG. 1 shows an exemplary workflow for the simultaneous assessment of somatic mutations and methylation patterns. [Figure 5-1] FIG. 5 shows an exemplary workflow for mutation analysis, and for simultaneous mutation and methylation analysis. [Figure 5-2] FIG. 5 shows an exemplary workflow for mutation analysis, and for simultaneous mutation and methylation analysis. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0026] Detailed Description Identification of rare DNA mutations or rare epigenetic changes (e.g., cytosine methylation, hydroxymethylation) is useful not only in aspects of basic biology but also to improve the clinical management of patients. Currently, conventional techniques to evaluate mutations in plasma with very high specificity require one aliquot of plasma, while conventional techniques to evaluate epigenetic changes require another aliquot of plasma. Because certain mutations and epigenetic changes in the plasma of early cancer patients are rare, it would be advantageous to evaluate as many molecules as possible for both genetic and epigenetic changes. Splitting the sample, half for genetic changes and half for epigenetic changes, reduces the sensitivity by 50% for both types of changes.

[0027] Therefore, there is a need to improve sequencing library preparation and workflows to enable accurate identification of mutations, e.g., rare mutations and epigenetic alterations, from the same aliquot of DNA purified from clinically relevant samples.

[0028] Provided herein is a method for identifying genetic and epigenetic characteristics of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule, the method comprising: (a) attaching an adapter fragment to each end of the double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) contacting the adapted double-stranded DNA molecule with a tagged primer. (c) subjecting the amplified products to denaturing conditions; (d) separately recovering the adapted Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson and Crick strands and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (f) generating a second population of analyte DNA fragments from the adapted Watson and Crick strands and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on the at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing reads according to the molecular barcode present on said at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family;(i) identifying genetic features of the tagged Watson and Crick strands in a first analyte DNA family; and (j) identifying epigenetic features of the matched Watson and Crick strands in a second analyte DNA family, thus identifying genetic and epigenetic features present on at least one strand of the double-stranded DNA molecule;

[0029] Various non-limiting aspects of these methods are described herein and may be used in any combination, without limitation. Additional aspects of the various components of the methods for identifying the presence or absence of mutations and methylation are known in the art.

[0030] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0031] As used herein, "adaptor", "adapter" and "tag" are terms used interchangeably and refer to a species that can be attached to a polynucleotide sequence (e.g., in a process referred to as "tagging") using any one of a number of different techniques, including, but not limited to, ligation, hybridization, and tagmentation. In some embodiments, an adapter can also be a nucleic acid sequence that adds a function, such as a spacer sequence, a primer sequence / site, a barcode sequence, or a unique molecular identifier sequence.

[0032] As used herein, the term "barcode" refers to a label or identifier that conveys or is capable of conveying information (e.g., information about an analyte in a sample). The barcode may be part of the analyte or may be independent of the analyte. In some embodiments, the barcode may be attached to the analyte. In some embodiments, a particular barcode may be unique relative to other barcodes. In some embodiments, the barcode may have a variety of different formats. For example, the barcode may include non-random, pseudo-random and / or random nucleic acid and / or amino acid sequences, as well as synthetic nucleic acid and / or amino acid sequences. In some embodiments, the barcode may be attached to the analyte or another moiety or structure in a reversible or irreversible manner. In some embodiments, the barcode may be added, for example, to fragments of a deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample prior to or during sequencing of the sample. In some embodiments, the barcode may allow for identification and / or quantification of individual sequencing reads. In some embodiments, the barcode may refer to a unique identifier (UID), and the terms "barcode" and "UID" may be used interchangeably.

[0033] As used herein, the terms "nucleotide" and "nt" are used interchangeably herein to broadly refer to biological molecules, including nucleic acids. Nucleotides can have moieties that contain known purine and pyrimidine bases. Nucleotides can also have other heterocyclic bases that are modified. Such modifications include, for example, methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses, or other heterocycles. The terms "polynucleotide", "nucleic acid" and "oligonucleotide" can be used interchangeably and refer to polymers of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise a sequence that is not naturally occurring. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.

[0034] As used herein, "primer" refers broadly to a polynucleotide molecule that comprises a nucleotide sequence (e.g., an oligonucleotide) that generally has a free 3'-OH group that can hybridize with a template sequence (such as a target polynucleotide or a primer extension product) and promote polymerization of a polynucleotide that is complementary to the template. In some embodiments, the primer is a biotinylated primer.

[0035] Overview Provided herein are methods and materials useful for accurately identifying genetic and epigenetic features present in a nucleic acid sample.In some embodiments, the method includes identifying genetic and epigenetic features when they are present in at least one of the Watson and Crick strands of a double-stranded nucleic acid template.In some embodiments, the method includes identifying genetic and epigenetic features when they are present in both the Watson and Crick strands of a double-stranded nucleic acid template.In some embodiments, the double-stranded nucleic acid template can include a Watson strand and a Crick strand.In some embodiments, the double-stranded nucleic acid template can include a plus strand and a minus strand.In some embodiments, the double-stranded nucleic acid template can include a first strand and a second strand.As recognized in the art, Watson / Crick, plus / minus, and first / second refer to the two strands of a double-stranded nucleic acid molecule.Such methods are particularly useful for distinguishing true mutations from artifacts derived from, for example, DNA damage, PCR, and other sequencing artifacts, allowing identification of mutations with high confidence.

[0036] In some embodiments, a method for identifying genetic and epigenetic features of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule comprises: (a) attaching an adapter fragment to each end of the double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand; (c) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand; (d) subjecting the amplified products to denaturing conditions; (e) separately recovering the matched Watson and Crick strands and the tagged Watson and Crick strands; (f) generating a first population of analyte DNA fragments from the tagged Watson and Crick strands and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on the at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing read according to the molecular barcode present on the at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family;It can include (i) identifying the genetic features of the tagged Watson strand and Crick strand in the first analyte DNA family; and (j) identifying the epigenetic features of the matched Watson strand and Crick strand in the second analyte DNA family, thus identifying the genetic and epigenetic features present on at least one strand of the double-stranded DNA molecule.In some embodiments, the method includes identifying the genetic and epigenetic features present on both strands of the double-stranded DNA molecule (Figure 1).

[0037] In some cases, the methods and materials described herein can be used to achieve efficient double strand recovery. In some embodiments, the methods described herein can be used to recover amplification products from at least one of the Watson strand and the Crick strand of a double stranded nucleic acid template. For example, the methods described herein can be used to recover amplification products from both the Watson strand and the Crick strand of a double stranded nucleic acid template. In some cases, the methods described herein can be used to achieve at least 50% (e.g., about 50%, about 60%, about 70%, about 75%, about 80%, about 82%, about 85%, about 88%, about 90%, about 93%, about 95%, about 97%, about 99% or 100%) double strand recovery (Figure 2).

[0038] In some embodiments, a method for detecting one or more mutations present in at least one strand of a double-stranded nucleic acid can include generating a double-stranded sequencing library having a double-stranded molecular barcode at each end (e.g., 5'-end and 3'-end) of each nucleic acid in the library, generating a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand from the double-stranded sequencing library, and detecting the presence of one or more mutations present in at least one strand of a double-stranded nucleic acid in each single-stranded library. In some embodiments, a method for detecting one or more mutations present in both strands of a double-stranded nucleic acid can include generating a double-stranded sequencing library having a double-stranded molecular barcode at each end (e.g., 5'-end and 3'-end) of each nucleic acid in the library, generating a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand from the double-stranded sequencing library, and detecting the presence of one or more mutations present in both strands of a double-stranded nucleic acid in each single-stranded library. The presence of a first molecular barcode in the 3' duplex adapter and a second molecular barcode present in the 5' adapter can be used to distinguish amplification products derived from the Watson strand from amplification products derived from the Crick strand.

[0039] In some cases, the methods and materials described herein can be used to independently evaluate each strand of double-stranded nucleic acid.For example, when nucleic acid mutation is identified in the double-stranded nucleic acid strands that are independently evaluated as described herein, the materials and methods described herein can be used to determine which strand of double-stranded nucleic acid the nucleic acid mutation originates from.

[0040] Any suitable method can be used to generate a double-stranded sequencing library. As used herein, a double-stranded sequencing library can be a plurality of nucleic acid fragments that includes a double-stranded molecular barcode at one end (e.g., 5'-end and / or 3'-end) of each nucleic acid fragment in the library, allowing at least one strand of the double-stranded nucleic acid to be sequenced. In some embodiments, both strands of the double-stranded nucleic acid are sequenced. In some cases, a nucleic acid sample (e.g., a double-stranded DNA molecule) can be fragmented to generate nucleic acid fragments (e.g., analyte DNA fragments), and the generated nucleic acid fragments can be used to generate a double-stranded sequencing library. The nucleic acid fragments used to generate a double-stranded sequencing library can also be referred to as input nucleic acids herein. For example, when the nucleic acid fragments used to generate a double-stranded sequencing library are DNA fragments, the DNA fragments can also be referred to as input DNA herein. A double-stranded sequencing library can include any suitable number of nucleic acid fragments. In some cases, generating a double-stranded sequencing library can include fragmenting a nucleic acid template and ligating an adaptor to each end of each nucleic acid fragment in the library.

[0041] (1) A conformed double-stranded DNA molecule In some embodiments, the methods described herein can include the steps of: (a) attaching an adapter fragment to each end of a double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; and (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand.

[0042] analyte nucleic acid The nucleic acid analyzed by any of the various methods provided herein can include any type of nucleic acid (e.g., DNA, RNA, and DNA / RNA hybrids). Examples of nucleic acids that can be analyzed include, but are not limited to, genomic DNA and cell-free DNA (cfDNA) (e.g., circulating tumor DNA (ctDNA), or cell-free fetal DNA (cffDNA). In some embodiments, the nucleic acid analyzed can be a double-stranded DNA molecule. In some embodiments, the double-stranded DNA molecule can include a Watson strand, where the Watson strand is the first single strand of the double-stranded DNA molecule. In some embodiments, the double-stranded DNA molecule can include a Crick strand, where the Crick strand is the second single strand of the double-stranded DNA molecule.

[0043] In some embodiments, the double-stranded DNA molecules analyzed are nucleic acid fragments (e.g., DNA fragments). In some embodiments, the nucleic acid fragments are produced manually. In some embodiments, the fragments are produced by shearing (e.g., enzymatic shearing, shearing by chemical means, acoustic shearing, nebulization, centrifugal shearing, point-sink shearing, needle shearing, sonication, restriction endonucleases, non-specific nucleases (e.g., DNase I), or any combination thereof). In some embodiments, the nucleic acid fragments are naturally produced in the subject. For example, the nucleic acid fragments analyzed can be cfDNA (e.g., circulating tumor DNA (ctDNA) or cell-free fetal DNA (cffDNA).

[0044] In some embodiments, the nucleic acid fragments to be analyzed are from about 4 to about 1000 nucleotides (e.g., from about 10 to about 1000, from about 20 to about 1000, from about 30 to about 1000, from about 40 to about 1000, from about 50 to about 1000, from about 60 to about 1000, from about 70 to about 1000, from about 80 to about 1000, from about 90 to about 1000, from about 100 to about 1000, from about 250 to about 1000, from about 500 to about 1000, from about 750 to about 1000, from about 4 to about 750, from about 10 to about 750, from about 20 to about 750, from about 30 to about 750, from about 40 to about 750, from about 50 to about 750, from about 60 to about 750, about 70 to about 750, about 80 to about 750, about 90 to about 750, about 100 to about 750, about 250 to about 750, about 500 to about 750, about 4 to about 500, about 10 to about 500, about 20 to about 500, about 30 to about 500, about 40 to about 500, about 50 to about 500, about 60 to about 500, about 70 to about 500, about 80 to about 500, about 90 to about 500, about 100 to about 500, about 250 to about 500, about 4 to about 250, about 10 to about 250, about 20 to about 250, about 30 to about 250, about 40 to about 250, about 50 to about 250, about 60 to about 250, about 70 to about 250, 80 to about 250, about 90 to about 250, about 100 to about 250, about 4 to about 100, about 10 to about 100, about 20 to about 100, about 30 to about 100, about 40 to about 100, about 50 to about 100, about 60 to about 100, about 70 to about 100, about 80 to about 100, about 90 to about 100, about 4 to about 90, about 10 to about 90, about 20 to about 90, about 30 to about 90, about 40 to about 90, about 50 to about 90, about 60 to about 90, about 70 to about 90, about 80 to about 90, about 4 to about 80, about 10 to about 80, about 20 to about 80, about 30 to about 80, about 40 to about 80, about 50 to about 80, about 60 to about 100 The length of the ribosome is about 80, about 70 to about 80, about 4 to about 70, about 10 to about 70, about 20 to about 70, about 30 to about 70, about 40 to about 70, about 50 to about 70, about 60 to about 70, about 4 to about 60, about 10 to about 60, about 20 to about 60, about 30 to about 60, about 40 to about 60, about 50 to about 60, about 4 to about 50, about 10 to about 50, about 20 to about 50, about 30 to about 50, about 40 to about 50, about 4 to about 40, about 10 to about 40, about 20 to about 40, about 30 to about 40, about 4 to about 30, about 10 to about 30, about 20 to about 30, about 4 to about 20, about 10 to about 20, or about 4 to about 10).In some embodiments, the nucleic acid fragments analyzed can be less than 1000 (eg, less than 750, less than 500, less than 250, less than 100, less than 50, or less than 20) nucleotides in length.

[0045] In some embodiments, a sequence present in the nucleic acid being analyzed (e.g., one or both ends of the nucleic acid) is used as an intrinsic barcode. In some embodiments, the ends of a DNA fragment correspond to a unique sequence that can be used as an intrinsic barcode (e.g., a unique identifier) ​​of the fragment. One skilled in the art can determine the length of the intrinsic barcode required to uniquely identify a nucleic acid template, using factors such as, for example, the overall template length, the complexity of the nucleic acid template in the partition or starting nucleic acid sample, etc. In some embodiments, about 10 to about 500 nucleotides (e.g., about 25 to about 500, about 50 to about 500, about 100 to about 500, about 250 to about 500, about 10 to about 250, about 25 to about 250, about 50 to about 250, about 100 to about 250, about 10 to about 100, about 25 to about 100, about 50 to about 100, about 10 to about 50, about 25 to about 50, or about 10 to about 25 nucleotides) at the end of the nucleic acid template are used as an endogenous barcode. In some embodiments, both ends of the nucleic acid template are used as endogenous barcodes. In some embodiments, only one end of the nucleic acid template is used as an endogenous barcode.

[0046] In some embodiments, the nucleic acid to be analyzed is present in and / or can be obtained from a biological sample. The biological sample can be obtained from a subject. In some embodiments, the subject is a mammal. Examples of mammals from which nucleic acid can be obtained and used as nucleic acid template in the methods described herein include, but are not limited to, humans, non-human primates (e.g., monkeys), dogs, cats, sheep, rabbits, mice, hamsters and rats. In some embodiments, the subject is a human subject.

[0047] Biological samples include, but are not limited to, plasma, serum, blood, tissue, tumor samples, stool, sputum, saliva, urine, sweat, tears, ascites, bronchoalveolar lavage fluid, semen, archaeological specimens, and forensic samples. In some embodiments, the biological sample is a solid biological sample, such as a tumor sample. In some embodiments, the solid biological sample is processed. The solid biological sample can be processed by fixation in formalin solution and then embedding in paraffin (e.g., FFPE sample). Alternatively, processing can include freezing the sample before performing the probe-based assay. In some embodiments, the sample is neither fixed nor frozen. Non-fixed, non-frozen samples can be stored in a storage solution configured for preservation of nucleic acids, by way of example only.

[0048] In some embodiments, the biological sample is a liquid biological sample. Liquid biological samples include, but are not limited to, plasma, serum, blood, sputum, saliva, urine, sweat, tears, peritoneal fluid, bronchoalveolar lavage fluid, and semen. In some embodiments, the liquid biological sample is acellular or substantially acellular. In some embodiments, the biological sample is a plasma or serum sample. In some embodiments, the liquid biological sample is a whole blood sample. In some embodiments, the liquid biological sample comprises peripheral blood mononuclear cells.

[0049] In some embodiments, the nucleic acid to be analyzed is isolated and purified from biological samples. Nucleic acid can be isolated and purified from biological samples using any means known in the art. For example, biological samples can be treated to release nucleic acid from cells or to separate nucleic acid from unwanted components of biological samples (e.g., proteins, cell walls, other contaminants). Additionally or alternatively, nucleic acid can be extracted from biological samples using liquid extraction (e.g., Trizol, DNAzol) techniques. Nucleic acid can also be extracted using commercially available kits (e.g., Qiagen DNeasy kit, QIAamp kit, Qiagen Midi kit, QIAprep spin kit).

[0050] Nucleic acids can be concentrated by known methods, including, by way of example only, centrifugation. Nucleic acids can be bound to selective membranes (e.g., silica) for purification purposes. Nucleic acids can also be concentrated for fragments of desired length, for example, fragments less than 1000, 500, 400, 300, 200 or 100 base pairs in length. Such size-based concentration can be carried out, for example, using PEG-induced precipitation, electrophoretic gels or chromatographic materials (Huber et al. (1993) Nucleic Acids Res. 21:1061-6), gel filtration chromatography, TSK gel (Kato et al. (1984) J. Biochem, 95:83-86), the publications of which are incorporated herein by reference.

[0051] In some embodiments, the nucleic acid sample that contains the nucleic acid to be analyzed contains less than about 35 ng of nucleic acid. For example, a nucleic acid sample may contain about 1 ng to about 35 ng of nucleic acid (e.g., about 1 ng to about 30 ng, about 1 ng to about 25 ng, about 1 ng to about 20 ng, about 1 ng to about 15 ng, about 1 ng to about 10 ng, about 1 ng to about 5 ng, about 5 ng to about 35 ng, about 5 ng to about 30 ng, about 5 ng to about 25 ng). ng, about 5 ng to about 20 ng, about 5 ng to about 15 ng, about 5 ng to about 10 ng, about 10 ng to about 35 ng, about 10 ng to about 30 ng, about 10 ng to about 25 ng, about 10 ng to about 20 ng, about 10 ng to about 25 ng, about 10 ng to about 20 ng, about 10 ng to about 15 ng, about 15 ng to about 35 The nucleic acid sample may comprise about 15 ng, about 15 ng to about 30 ng, about 15 ng to about 25 ng, about 15 ng to about 20 ng, about 20 ng to about 35 ng, about 20 ng to about 30 ng, about 20 ng to about 25 ng, about 25 ng to about 35 ng, about 25 ng to about 30 ng, or about 30 ng to about 35 ng of nucleic acid. In some cases, the nucleic acid sample may comprise nucleic acid from a genome that includes nucleic acids of more than about several hundred nucleotides.

[0052] In some cases, the nucleic acid sample that comprises the nucleic acid to be analyzed can be essentially free of contamination. For example, when the nucleic acid sample comprises the cfDNA nucleic acid to be analyzed, the cfDNA can be essentially free of genomic DNA contamination. In some cases, the nucleic acid sample that comprises the cfDNA that is essentially free of genomic DNA contamination can contain minimal (or no) high molecular weight (e.g., greater than 1000 bp) DNA. In some cases, the method described herein can include a step of determining whether the nucleic acid sample is essentially free of contamination. Any suitable method can be used to determine whether the nucleic acid sample is essentially free of contamination. Examples of methods that can be used to determine whether the nucleic acid sample is essentially free of contamination include, for example, TapeStation system and Bioanalyzer. For example, when using TapeStation system and / or Bioanalyzer to determine whether the cfDNA sample is essentially free of genomic DNA contamination, a prominent peak at approximately 180 bp (e.g., corresponding to mononuclear DNA) can be used to indicate that the nucleic acid sample is essentially free of genomic DNA contamination.

[0053] In some cases, the nucleic acid fragments that can be used to generate a duplex sequencing library (e.g., before attaching a 3' duplex adaptor to the 3' end of the nucleic acid fragment) can be end-repaired. Any suitable method can be used to end-repair the nucleic acid template. For example, a blunt-end reaction (e.g., blunt-end ligation) and / or a dephosphorylation reaction can be used to end-repair the nucleic acid template. In some cases, the blunt-end reaction can include filling in the single-stranded region. In some cases, the blunt-end reaction can include degrading the single-stranded region. In some cases, the blunt-end reaction and a dephosphorylation reaction can be used to end-repair the nucleic acid template.

[0054] adapter As used herein, "adapter" and "adapter fragment" can refer to a species that can be attached to a polynucleotide sequence using any one of many different techniques, including, but not limited to, ligation, hybridization, and tagmentation. In some embodiments, the adapter fragment can also be a nucleic acid sequence that adds a function, such as a spacer sequence, a primer sequence / site, or a barcode sequence (e.g., a UID sequence).

[0055] In some embodiments, the method described herein comprises attaching an adaptor fragment to each end of a double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adaptor fragment comprising a molecular barcode, a primer sequence, and an adaptor sequence, and the molecular barcode of the adapted Watson strand is the reverse complement of the molecular barcode of the adapted Crick strand. In some embodiments, the primer sequence can be the reverse complement of the adaptor sequence. In some embodiments, the adaptor sequence can comprise a specific sequence to enable sequencing in generating a sequence library. In some embodiments, the adaptor sequence comprises a sequencing primer sequence (e.g., R1, R2).

[0056] In some embodiments, the adapter fragment comprises a double-stranded portion comprising a molecular barcode and a fork portion comprising (i) a single-stranded 3' adapter sequence and (ii) a single-stranded 5' adapter sequence. In some embodiments, the single-stranded 3' adapter sequence is not complementary to the single-stranded 5' adapter sequence. In some embodiments, the 3' adapter sequence comprises a second (e.g., R2) sequencing primer site and the 5' adapter sequence comprises a first (e.g., R1) sequencing primer site. It should be understood that the "R1" and "R2" sequencing primer sites are used by a sequencing system that produces paired-end reads, e.g., reads from opposite ends of a DNA fragment to be sequenced. In some embodiments, the R1 sequencing primer is used to produce a first population of reads from a first end of a DNA fragment and the R2 sequencing primer is used to produce a second population of reads from an opposite end of a DNA fragment. The first population is referred to herein as "R1" or "Read 1" reads. The second population is referred to herein as the "R2" or "Read 2" reads. The R1 and R2 reads can be aligned as a "read pair" or "mate pair" that corresponds to each strand of a double-stranded analyte DNA fragment.

[0057] Certain sequencing systems (e.g., Illumina) utilize what are referred to as "R1" and "R2" primers, and "R1" and "R2" reads. For the purposes of this application, it should be noted that the terms "R1" and "R2", and "read 1" and "read 2" are not limited to the manner in which they are referred to in relation to a particular sequencing platform. For example, when an Illumina sequencer is used, the "R2" primer and corresponding R2 read disclosed herein can refer to the Illumina "R2" primer and read, or the "R1" primer and corresponding R1 read disclosed herein can refer to other Illumina primers and reads. For clarity, in some embodiments where the "R2" primer provided herein is an Illumina "R1" primer that produces an "R1" read, the corresponding "R1" primer provided herein is an Illumina "R2" primer that produces an "R2" read. For clarity, in some embodiments where the "R2" primer provided herein is an Illumina "R2" primer that provides an "R2" read, the "R1" primer provided herein is an Illumina "R1" primer that provides an R1 read.

[0058] In some embodiments, the adapted double-stranded DNA molecule can be a double-stranded DNA molecule to which an adaptor is attached. In some embodiments, the adaptor fragment further comprises a sample barcode. In some embodiments, the sample barcode is different from the molecular barcode, where the sample barcode is unique to the sample from which the double-stranded DNA molecule was obtained. In some embodiments, the first double-stranded DNA molecule from the first sample can be contacted with a first adaptor fragment comprising a first sample barcode unique to the first sample. In some embodiments, the second double-stranded DNA molecule from the second sample can be contacted with a second adaptor fragment comprising a second sample barcode unique to the second sample. In some embodiments, the first adapted double-stranded DNA molecule and the second adapted double-stranded DNA molecule can be mixed in a population of adapted double-stranded DNA molecules used in any of the methods described herein. In some embodiments, the mixing of the first and second adapted double-stranded DNA molecules can be performed after the attaching step (a) and the copying step (b). In some embodiments, the mixing of the first and second adapted double-stranded DNA molecules can be performed after contacting the adapted double-stranded DNA molecules with the tagged primer. In some embodiments, the mixing of the first and second adapted double-stranded DNA molecules can be performed after subjecting the amplified products to denaturing conditions in step (c).

[0059] In some embodiments, the population of double-stranded DNA molecules can include a plurality of double-stranded DNA molecules that include the same sample barcode. In some embodiments, the population of double-stranded DNA molecules can include a plurality of double-stranded DNA molecules that include different sample barcodes.

[0060] Molecular barcoding As used herein, "molecular barcode" refers to a barcode that serves to identify individual nucleic acid fragments in an original sample prior to barcoding and amplification. In some embodiments, each individual nucleic acid fragment has a unique molecular barcode. In some embodiments, the barcode may be a randomly generated nucleotide sequence or an intentionally selected nucleotide run. In particular, when attaching molecular barcodes, the number of individual molecular barcodes in the reaction mixture will exceed the number of nucleic acid fragments.

[0061] In some embodiments, the molecular barcode is unique to each double-stranded DNA fragment in the nucleic acid sample. In some embodiments, the molecular barcode comprises an intrinsic barcode, an exogenous barcode, or both.

[0062] In some embodiments, the molecular barcode has a length of about 2 to about 4000 (e.g., about 2 to about 3500, about 2 to about 3000, about 2 to about 2500, about 2 to about 2000, about 2 to about 1500, about 2 to about 1000, about 2 to about 500, about 2 to about 100, about 2 to about 50, about 2 to about 20, about 2 to about 10, about 10 to about 4000, about 10 to about 3500, about 10 to about 3000, about 10 to about 2500, about 10 to about 2000, about 10 to about 1500, about 10 to about 1000, about 10 to about 500, about 10 to about 100, about 10 to about 50, about 10 to about 20, about 20 to about about 4000, about 20 to about 3500, about 20 to about 3000, about 20 to about 2500, about 20 to about 2000, about 20 to about 1500, about 20 to about 1000, about 20 to about 500, about 20 to about 100, about 20 to about 50, about 50 to about 4000, about 50 to about 3500, about 50 to about 3000, about 50 to about 2500, about 50 to about 2000, about 50 to about 1500, about 50 to about 1000, about 50 to about 500, about 50 to about 100, about 100 to about 4000, about 100 to about 3500, about 1 00 to about 3000, about 100 to about 2500, about 100 to about 2000, about 100 to about 1500, about 100 to about 1000, about 100 to about 500, about 500 to about 4000, about 500 to about 3500, about 500 to about 3000, about 500 to about 2500, about 500 to about 2000, about 500 to about 1500, about 500 to about 1000, about 1000 to about 4000, about 1000 to about 3500, about 1000 to about 3000, about 1000 to about 2500, about 1000 to about 2000, about 100 The molecular barcode has a length of 0 to about 1500, about 1500 to about 4000, about 1500 to about 3500, about 1500 to about 3000, about 1500 to about 2500, about 1500 to about 2000, about 2000 to about 4000, about 2000 to about 3500, about 2000 to about 3000, about 2000 to about 2500, about 2500 to about 4000, about 2500 to about 3500, about 2500 to about 3000, about 3000 to about 4000, about 3000 to about 3500, or about 3500 to about 4000 nucleotides. In some embodiments, the length of the molecular barcode is sufficient to uniquely barcode a molecule and the length / sequence of the molecular barcode does not interfere with downstream amplification steps.

[0063] In some embodiments, the molecular barcode sequence can be random. In some embodiments, the molecular barcode sequence can be a random N-mer. For example, if the molecular barcode sequence has a length of 6 nt, it can be a random hexamer. If the molecular barcode sequence has a length of 12 nt, it can be a random 12-mer.

[0064] In some embodiments, molecular barcodes can be created using random addition of nucleotides to form a sequence with a length that is used as an identifier. At each addition position, a selection from four deoxyribonucleotides can be used. Alternatively, a selection from three, two or one deoxyribonucleotides can be used. Thus, molecular barcodes can be completely random, somewhat random, or non-random at certain positions. In some embodiments, molecular barcodes are not random N-mers, but are selected from a predefined set of molecular barcode sequences. Exemplary molecular barcodes suitable for use in the methods disclosed herein are described in PCT / US2012 / 033207, which is incorporated herein by reference in its entirety.

[0065] Attachment of molecular barcodes to nucleic acid fragments can be performed by any means known in the art, including enzymatic, chemical, or biological means. In some embodiments, one means utilizes polymerase chain reaction. In some embodiments, another means utilizes ligase enzymes. For example, the ligase enzyme can be mammalian or bacterial. Other enzymes that can be used for attachment are other polymerase enzymes. Molecular barcodes can be added to one or both ends of a fragment, preferably both ends. In some embodiments, molecular barcodes can be included within a nucleic acid molecule that includes other regions for other intended functionality. For example, universal priming sites can be added to allow for subsequent amplification. In some embodiments, another addition site can be a region of complementarity to a specific region or gene in the nucleic acid fragment.

[0066] (2) Tagged double-stranded DNA molecule In some embodiments, the method described herein includes (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising (i) contacting the adapted double-stranded DNA molecule with a tagged primer, and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand. In some embodiments, the copying can comprise performing a single round of linear extension. In some embodiments, the copying can comprise performing 1, 2, or 3 rounds of linear extension. In some embodiments, the copying can comprise performing 1 or multiple rounds (e.g., 1, 2, 3, 4, or 5 rounds) of linear extension. In some embodiments, the tagged primer is a uracil-containing biotinylated primer, and the tagged Watson strand and the Crick strand are generated from the uracil-containing biotinylated primer. In some embodiments, tagged Watson and Crick strands can be selected using biotinylated-streptavidin affinity in any number of ways known in the art (e.g., streptavidin beads).

[0067] As used herein, the term "extension" can refer to a method in which two nucleic acid sequences are linked (e.g., hybridized) by overlapping of their respective terminal complementary nucleic acid sequences (i.e., e.g., 3' termini). Such ligation can be followed by nucleic acid extension (e.g., enzymatic extension) of one or both ends using the other nucleic acid sequence as a template for extension.

[0068] In some embodiments, nucleic acid extension generally involves the incorporation of one or more nucleic acids (e.g., A, G, C, T, U, nucleotide analogs, or derivatives thereof) into a nucleic acid sequence in a template-dependent manner, such that the successive nucleic acids are incorporated by an enzyme (such as a polymerase or reverse transcriptase), thereby generating a newly synthesized nucleic acid molecule. In some embodiments, enzymatic extension can be performed by an enzyme, including, but not limited to, a polymerase and / or a reverse transcriptase. For example, a primer that hybridizes to a complementary nucleic acid sequence can be used to synthesize a new nucleic acid molecule by using a complementary nucleic acid sequence as a template for nucleic acid synthesis.

[0069] In some embodiments, primers can be single-stranded nucleic acid sequences with 3'-ends that can be used as chemical substrates for nucleic acid polymerase in nucleic acid extension reactions. RNA primers are formed from RNA nucleotides and used in RNA synthesis, while DNA primers are formed from DNA nucleotides and used in DNA synthesis. Primers can also contain both RNA and DNA nucleotides (e.g., in random or designed patterns). In some embodiments, primers can also contain other natural or synthetic nucleotides described herein that can have additional functionality.

[0070] In some embodiments, the primer can include a tag, which is a molecule or molecular moiety that has high affinity or preference for associating or binding with another specific or particular molecule or moiety. In some embodiments, the association or binding with another specific or particular molecule or moiety can be by non-covalent interactions, such as hydrogen bonds, ionic forces, and van der Waals interactions. For example, the affinity group can be biotin, which has high affinity or preference for associating or binding with the protein avidin or streptavidin. Alternatively, the affinity group can refer to avidin or streptavidin, which has affinity for biotin. Other examples of affinity groups and the specific or particular molecules or moieties that they bind or associate with include, but are not limited to, antibodies or antibody fragments and their respective antigens, such as digoxigenin and anti-digoxigenin antibodies, lectins, carbohydrates (e.g., saccharides, monosaccharides, disaccharides, or polysaccharides), and receptors and receptor ligands. In some embodiments, the tagged primer is a biotinylated primer, and the tagged Watson strand and the Crick strand are generated from the biotinylated primer.In some embodiments, the tagged primer is a uracil-containing biotinylated primer, and the tagged Watson strand and the Crick strand are generated from the uracil-containing biotinylated primer.In some embodiments, the tagged Watson strand and the Crick strand can be selected by using biotinylated-streptavidin affinity in any number of ways known in the art (e.g., streptavidin beads).

[0071] (3) Denaturation of double-stranded DNA molecules In some embodiments, the method also includes (c) subjecting the amplified products to denaturing conditions. In some embodiments, the denaturing conditions include NaOH denaturation. In some embodiments, the denaturing conditions can include, but are not limited to, heat denaturation, chemical denaturation, or a combination thereof. In some embodiments, the double-stranded DNA molecules can be denatured by using heat. In some embodiments, the denaturation of the double-stranded DNA molecules can be achieved by chemical denaturation. In some embodiments, the chemical denaturation can include NaOH treatment. In some embodiments, the double-stranded DNA molecules can be denatured by using salt. In some embodiments, the double-stranded DNA molecules can be denatured by using salt and additional chemicals (e.g., isopropanol and ethanol).

[0072] (4) Recovery and purification of analyte DNA fragments In some embodiments, any of the methods described herein can include: (d) separately recovering the matched Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson and Crick strands, and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; and (f) generating a second population of analyte DNA fragments from the matched Watson and Crick strands, and generating a second sequencing read for at least one member of the second population of analyte DNA fragments.In some embodiments, the recovering step (d) includes contacting the tagged Watson and Crick strands with streptavidin-functionalized beads, and wherein the tagged Watson and Crick strands bind to the streptavidin-functionalized beads.

[0073] In some embodiments, the collected matched Watson strand and Crick strand that are not bound to streptavidin-functionalized beads are treated with bisulfite to convert cytosine bases to uracil bases to generate a second population of analyte DNA fragments that includes a population of converted DNA molecules.In some embodiments, bisulfite treatment can effectively convert the C bases in DNA molecules to U bases.In some embodiments, this conversion makes the two strands (e.g., Watson strand and Crick strand) distinguishable.In some embodiments, bisulfite conversion can be used to distinguish the methylated C bases that are not converted to T bases from the unmethylated C bases, thereby revealing epigenetic changes.

[0074] In some embodiments, tagged Watson and Crick chains can be separated by using any pair of affinity group and its specific or particular molecule or moiety that it binds or associates with. For example, affinity group can be biotin, which has high affinity or preference to associate or bind to protein avidin or streptavidin. Alternatively, affinity group can refer to avidin or streptavidin, which has affinity for biotin. In some embodiments, tagged Watson and Crick chains can be selected using biotinylated-streptavidin affinity in any number of ways known in the art (e.g., streptavidin beads). Other examples of affinity group and its specific or particular molecule or moiety that it binds or associates with include, but are not limited to, antibodies or antibody fragments and their respective antigens, such as digoxigenin and anti-digoxigenin antibodies, lectins, carbohydrates (e.g., saccharides, monosaccharides, disaccharides or polysaccharides), and receptors and receptor ligands.

[0075] In some embodiments, the recovering step can include using magnetic beads to separate the tagged Watson and Crick strands. In some embodiments, the magnetic beads can be covalently coated with streptavidin and can bind to the biotinylated tagged Watson and Crick strands. In some embodiments, the magnetic beads can be purified by using a magnet. In some embodiments, the magnetic beads can be recovered by centrifugation and size fractionated by filtration or flow sorting.

[0076] In some embodiments, the tagged Watson and Crick strands can be attached to a single bead, where the bead is stained with a fluorescent probe and counted using flow cytometry. Beads that represent specific variants can be optionally recovered by flow sorting and used for subsequent confirmation and experimentation. In some embodiments, the beads can be microspheres or microparticles. Particle size can vary between about 0.1-10 microns in diameter. Typically, beads are made of polymeric materials such as polystyrene, although non-polymeric materials such as silica can also be used. Other materials that can be used include styrene copolymers, methyl methacrylate, functionalized polystyrene, glass, silicon, and carboxylates. Optionally, the particles can be superparamagnetic, which facilitates their purification after being used in reactions. In some embodiments, the beads can be modified by covalent or non-covalent interactions with other materials to change gross surface properties, such as hydrophobicity or hydrophilicity, or to attach molecules that confer binding specificity. Such molecules can include, but are not limited to, antibodies, ligands, members of specific binding protein pairs, receptors, and nucleic acids. Specific binding protein pairs include avidin-biotin, streptavidin-biotin, and Factor VII-tissue factor.

[0077] In some embodiments, the tagged Watson and Crick strands can be separated by using treatment with USER (Uracil-Specific Excision Reagent) enzyme, which contains a mixture of uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII, which targets deoxyuridine bases embedded within the 5' ends of the strands.

[0078] Genetics – Sequencing As used herein, the term "genetic feature" refers to the genetic information and / or material that is replicated and passed on from parent cells to progeny cells at each cell division. In some embodiments, the genetic feature can be a mutation in a nucleic acid (e.g., a DNA molecule). In some embodiments, the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and a combination thereof. In some embodiments, the identification of the genetic feature can include mutation analysis, aneuploidy analysis, or fragmentomics. Exemplary methods for identifying genetic features suitable for use in the methods disclosed herein are described in PCT / US2021 / 017937, which is incorporated herein by reference in its entirety.

[0079] (a) Initial amplification of adaptor-attached templates. After adapter attachment, the adapted double-stranded DNA molecule can be amplified (e.g., PCR amplified) in an initial amplification reaction. Any suitable method can be used to amplify the adapted double-stranded DNA molecule. Exemplary methods that can be used to amplify the adapted double-stranded DNA molecule include, but are not limited to, whole genome PCR. In some embodiments, the adapted double-stranded DNA molecule is amplified by performing a single round of linear extension. In some embodiments, the adapted double-stranded DNA molecule is amplified by performing one, two, or three rounds of linear extension. In some embodiments, the adapted double-stranded DNA molecule is amplified by performing one or multiple rounds (e.g., one, two, three, four, or five rounds) of linear extension.

[0080] Any suitable primer pair can be used to amplify the adapted double-stranded DNA molecule. In some embodiments, a universal primer pair can be used. The primers can include, but are not limited to, about 12 nucleotides to about 30 nucleotides. In some embodiments, any suitable PCR conditions can be used in the initial amplification. The PCR amplification can include a denaturation phase, an annealing phase, and an extension phase. Each phase of the amplification cycle can include any suitable conditions. In some cases, the denaturation phase can include a temperature of about 90°C to about 105°C (e.g., about 94°C to about 98°C) and a time of about 1 second to about 5 minutes (e.g., about 10 seconds to about 1 minute). For example, the denaturation phase can include a temperature of about 98°C for about 10 seconds. In some cases, the annealing phase can include a temperature of about 50°C to about 72°C and a time of about 30 seconds to about 90 seconds. In some cases, the extension phase can include a temperature of about 55°C to about 80°C and a time of about 15 seconds per kb of amplicon generated to about 30 seconds per kb of amplicon generated. In some cases, the annealing and extension phases can be performed in a single cycle. For example, the annealing and extension phases can include a temperature of about 65° C. for about 75 seconds.

[0081] The PCR conditions used in the initial amplification can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can be performed for about 1 to about 50 (e.g., about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 5 to about 45, about 10 to about 45, about 15 to about 45, about 20 to about 45, about 25 to about 45, about 30 to about 45, about 35 to about 45, about 40 to about 45, about 1 to about 40, about 5 to about 40, about 10 to about 40, about 15 to about 40, about 20 to about 40, about 25 to about 40, about 30 to about 40, about 35 to about 40, about 1 to about 35, about 5 The PCR amplification may comprise about 1 to about 35, about 10 to about 35, about 15 to about 35, about 20 to about 35, about 25 to about 35, about 30 to about 35, about 1 to about 30, about 5 to about 30, about 10 to about 30, about 15 to about 30, about 20 to about 30, about 25 to about 30, about 1 to about 25, about 5 to about 25, about 10 to about 25, about 15 to about 25, about 20 to about 25, about 1 to about 20, about 5 to about 20, about 10 to about 20, about 15 to about 20, about 1 to about 15, about 5 to about 15, about 10 to about 15, about 1 to about 10, about 5 to about 10, or about 1 to about 5 cycles. In some embodiments, the PCR amplification comprises 11 cycles or less. In some embodiments, the PCR amplification comprises 7 cycles or less. In some embodiments, the PCR amplification comprises 5 cycles or less.

[0082] In some cases, when the PCR conditions include a heat-activated polymerase, the PCR amplification can also include an initialization step. For example, the PCR amplification can include an initialization step before performing a PCR amplification cycle. In some cases, the initialization step can include a temperature of about 94°C to about 98°C, and a time period of about 15 seconds to about 1 minute. For example, the initialization step can include a temperature of about 98°C for about 30 seconds.

[0083] In some cases, the PCR amplification can also include a holding step. For example, the PCR amplification can include a holding step after performing the PCR amplification cycles, and optionally after performing a final extension step. In some cases, the holding step can include a temperature of about 4° C. to about 15° C. for an unlimited amount of time.

[0084] In some cases, the double-stranded sequencing library (e.g., the amplified double-stranded sequencing library) generated as described herein can be purified.Any suitable method can be used to purify the double-stranded sequencing library.The exemplary method that can be used to purify the double-stranded sequencing library includes, but is not limited to, magnetic beads (e.g., solid-phase reversible immobilization (SPRI) magnetic beads).

[0085] (b) Optional ssDNA library preparation In some cases, the double-stranded sequencing library can be used to generate a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands. By generating a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands, non-specific amplification (e.g., from primers complementary to ligation sequences such as 3' double-stranded adapters or 5' adapters) can be minimized. Any suitable method can be used to generate a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands (e.g., from a double-stranded sequencing library made as described herein). In some cases, the library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands can be generated from the amplified double-stranded sequencing library by dividing the amplification product into at least two aliquots and subjecting each aliquot to PCR amplification, in which Watson strands are amplified from the first aliquot and Crick strands are amplified from the second aliquot. For example, a first aliquot of the amplified products from the double-stranded sequencing library can be subjected to PCR amplification with a primer pair, the first primer being biotinylated and the second primer being non-biotinylated, to generate a single-stranded library of Watson strand, and a second aliquot of the amplified products from the double-stranded sequencing library can be subjected to PCR amplification with a primer pair, the first primer being non-biotinylated and the second primer being biotinylated, to generate a single-stranded library of Crick strand.In some cases, a library of sequences derived from single-stranded Watson strand and a library of sequences derived from single-stranded Crick strand can be generated.

[0086] Any suitable method can be used to generate a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand from an amplified double-stranded sequencing library. For example, the amplification products from the amplified double-stranded sequencing library can be separated into a first PCR amplification and a second PCR amplification, in which only one of the two primers in the PCR primer pair is tagged. For example, the first PCR amplification can use a primer pair that includes a tagged primer (e.g., the first primer) and an untagged primer (e.g., the second primer), and the second PCR amplification can use a primer pair that includes an untagged primer (e.g., the first primer) and a tagged primer (e.g., the second primer). The primer tag can be any tag that allows the PCR amplification products generated from the tagged primer to be recovered. In some cases, the tagged primer can be a biotinylated primer, and the PCR amplification products generated from the biotinylated primer can be recovered using streptavidin. In some cases, the tagged primer can be a biotinylated primer containing uracil, and the PCR amplification products generated from the biotinylated primer containing uracil can be recovered using streptavidin.For example, the library of sequences from single-stranded Watson strand and the library of sequences from single-stranded Crick strand can be generated in PCR amplification using a primer pair that includes a biotinylated primer and a non-biotinylated primer.In some cases, the tagged primer can be a phosphorylated primer, and the PCR amplification products generated from the phosphorylated primer can be recovered using lambda nuclease.For example, the library of sequences from single-stranded Watson strand and the library of sequences from single-stranded Crick strand can be generated in PCR amplification using a primer pair that includes a phosphorylated primer and a non-phosphorylated primer.

[0087] Any suitable primer pair can be used to generate a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand (e.g., from a duplex sequencing library generated as described herein). The primers can include, without limitation, about 12 nucleotides to about 30 nucleotides. In some cases, the primer pair can include at least one primer that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification product generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to the nucleic acid fragments in the duplex sequencing library prior to amplification). Examples of primer pairs that can be used to generate a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand described herein include, without limitation, P5 primer and P7 primer.

[0088] Any suitable PCR conditions can be used to generate a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand (e.g., from a double-stranded sequencing library generated as described herein). PCR amplification can include a denaturation phase, an annealing phase, and an extension phase. Each phase of an amplification cycle can include any suitable conditions. In some cases, the denaturation phase can include a temperature of about 90°C to about 105°C, and a time of about 1 second to about 5 minutes. For example, the denaturation phase can include a temperature of about 98°C for about 10 seconds. In some cases, the annealing phase can include a temperature of about 50°C to about 72°C, and a time of about 30 seconds to about 90 seconds. In some cases, the extension phase can include a temperature of about 55°C to about 80°C, and a time of about 15 seconds per kb of amplicon generated to about 30 seconds per kb of amplicon generated. In some cases, the extension phase reflects the processivity of the polymerase used. In some cases, the annealing phase and the extension phase can be performed in a single cycle. For example, the annealing and extension phases can include a temperature of about 65° C. for about 75 seconds.

[0089] The PCR conditions used to generate the libraries of sequences derived from the single-stranded Watson strand and the libraries of sequences derived from the single-stranded Crick strand (e.g., from a double-stranded sequencing library generated as described herein) can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can be, but is not limited to, about 1 to about 50 cycles. (For example, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 5 to about 45, about 10 to about 45, about 15 to about 45, about 20 to about 45, about 25 to about 45, about 30 to about 45, about 35 to about 45, about 40 to about 45, about 1 to about 40, about 5 to about 40, about 10 to about 40, about 15 to about 40, about 20 to about 40, about 25 to about 40, about 30 to about 40, about 35 to about 40, about 1 to about 35, about 5 about 1 to about 35, about 10 to about 35, about 15 to about 35, about 20 to about 35, about 25 to about 35, about 30 to about 35, about 1 to about 30, about 5 to about 30, about 10 to about 30, about 15 to about 30, about 20 to about 30, about 25 to about 30, about 1 to about 25, about 5 to about 25, about 10 to about 25, about 15 to about 25, about 20 to about 25, about 1 to about 20, about 5 to about 20, about 10 to about 20, about 15 to about 20, about 1 to about 15, about 5 to about 15, about 10 to about 15, about 1 to about 10, about 5 to about 10, or about 1 to about 5) cycles. For example, the PCR amplification can include about 4 amplification cycles. In some embodiments, the PCR amplification can include about 8 amplification cycles. In some embodiments, the PCR amplification can include about 11 amplification cycles.

[0090] In some cases, when the PCR conditions include a heat-activated polymerase, the PCR amplification can also include an initialization step. For example, the PCR amplification can include an initialization step before performing a PCR amplification cycle. In some cases, the initialization step can include a temperature of about 94°C to about 98°C, and a time period of about 15 seconds to about 1 minute. For example, the initialization step can include a temperature of about 98°C for about 30 seconds.

[0091] In some cases, the PCR amplification can also include a holding step. For example, the PCR amplification can include a holding step after performing the PCR amplification cycles, and optionally after performing a final extension step. In some cases, the holding step can include a temperature of about 4° C. to about 15° C. for an unlimited amount of time.

[0092] Any suitable method can be used to separate double-stranded amplification product into single-stranded amplification product.In some cases, double-stranded amplification product can be denatured to separate double-stranded amplification product into two single-stranded amplification products.The examples of the method that can be used to separate double-stranded amplification product into single-stranded amplification product include, but are not limited to, heat denaturation, chemical (e.g., NaOH) denaturation, and salt denaturation.

[0093] After PCR amplification, tagged Watson strand and Crick strand can be collected. Any suitable method can be used to collect tagged Watson strand and Crick strand produced by using tagged primer. If tagged primer is biotinylated primer, biotinylated amplification product (e.g., produced from biotinylated primer) can be collected by using streptavidin (e.g., streptavidin functionalized beads). For example, when the amplified double-stranded sequencing library is further amplified in a first PCR amplification using a primer pair comprising a first biotinylated primer and a second non-biotinylated primer, and a second PCR amplification using a primer pair comprising a first non-biotinylated primer and a second biotinylated primer, the biotinylated amplification products generated from the first PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first set of streptavidin-functionalized beads), and the biotinylated amplification products generated from the second PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first second streptavidin-functionalized beads), and the double-stranded amplification products can be separated (e.g., denatured) into single-stranded amplification products. In some cases, recovering the biotinylated PCR amplification products can also include releasing the biotinylated PCR amplification products from the streptavidin (e.g., the streptavidin-functionalized beads).Separating the double-stranded amplification products generated by the first PCR amplification using a primer pair comprising a first biotinylated primer and a second non-biotinylated primer, and the second PCR amplification using a primer pair comprising a first non-biotinylated primer and a second biotinylated primer, can allow the single-stranded amplification products generated from the biotinylated primer to remain bound to the streptavidin-functionalized beads, while the single-stranded amplification products generated from the non-biotinylated primer can be denatured (e.g., denatured and degraded) from the streptavidin-functionalized beads, thereby generating a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences of the double-stranded sequencing library.

[0094] When the tagged primer is a phosphorylated primer, the phosphorylated amplification product (e.g., generated from the phosphorylated primer) can be recovered using an exonuclease (e.g., lambda exonuclease).For example, when the amplified double-stranded sequencing library is further amplified in a first PCR amplification using a primer pair comprising a first phosphorylated primer and a second non-phosphorylated primer, and in a second PCR amplification using a primer pair comprising a first non-phosphorylated primer and a second phosphorylated primer, the double-stranded amplification product can be separated into a single-stranded amplification product. Separating the double-stranded amplification products generated by the first PCR amplification using a primer pair comprising a first phosphorylated primer and a second unphosphorylated primer, and the second PCR amplification using a primer pair comprising a first unphosphorylated primer and a second phosphorylated primer, can enable the single-stranded amplification products generated from the unphosphorylated primer to be recovered, while the single-stranded amplification products generated from the phosphorylated primer can be degraded by lambda exonuclease, thereby generating a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand of the double-stranded sequencing library.

[0095] (c) Target enrichment In some embodiments of any one of the methods herein, the amplification product produced by initial amplification is enriched for one or more target polynucleotides.In some embodiments, before target enrichment, a single-stranded DNA library is prepared from the amplification product produced by initial amplification.An exemplary method for producing a single-stranded DNA library is described herein.

[0096] Any suitable method can be used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand). In some cases, a target region can be amplified from a library of amplification products by subjecting the library of amplification products to PCR amplification using a primer pair, which is a primer (e.g., a first primer) that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to the nucleic acid fragments in the double-stranded sequencing library prior to amplification), and a primer (e.g., a second primer) that can target (e.g., target and bind to) the target region (e.g., a region of interest).

[0097] In some cases, a target region can be amplified from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) in a single PCR amplification. For example, a target region can be amplified from a library of amplification products in a single PCR amplification using a primer pair that includes a first primer that can target an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in an amplification product generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in the double-stranded sequencing library prior to amplification), and a second primer that can target the target region.

[0098] In some cases, the target region can be amplified from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) in multiple PCR amplifications.Multiple PCR amplifications (e.g., an initial PCR amplification and a subsequent nested PCR amplification) can be used to increase the specificity of amplifying the target region. For example, a target region can be amplified from a library of amplification products in a series of PCR amplifications in which a first PCR amplification is performed using a primer pair comprising a first primer capable of targeting an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a second primer capable of targeting the target region, and the amplification products generated in the first PCR amplification are subjected to a subsequent nested PCR amplification using a primer pair comprising a first primer capable of targeting an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a second primer capable of targeting a nucleic acid sequence from the target region present in the amplification products generated in the first PCR amplification.

[0099] Any suitable primer pair can be used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand). The primers can include, without limitation, about 12 nucleotides to about 30 nucleotides. In some cases, the primer pair can include a primer (e.g., a first primer) that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in the double-stranded sequencing library prior to amplification), as well as a primer (e.g., a second primer) that can target (e.g., target and bind to) a target region (e.g., a region of interest). Examples of primers that can target adapter sequences that include molecular barcodes present in amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification) include, but are not limited to, i5 index primers and i7 index primers. Primers that can target a target region can include a sequence that is complementary to the target region. When the target region is a nucleic acid encoding TP53, examples of primers that can target a nucleic acid encoding TP53 include, but are not limited to, TP53_342_GSP1 and TP53_GSP2.

[0100] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) can include one or more molecular barcodes.

[0101] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) can contain one or more grafting sequences (e.g., grafting sequences for next-generation sequencing).

[0102] In some embodiments, target enrichment comprises: (a) selectively amplifying an amplification product of a Watson strand comprising a target polynucleotide sequence with a first set of Watson target selective primer pairs, the first set comprising: (i) a first Watson target selective primer comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence; and (ii) a second Watson target selective primer comprising a target selective sequence, thereby generating a target Watson amplification product; and (b) selectively amplifying an amplification product of a Click strand comprising the same target polynucleotide sequence with a first set of Click target selective primer pairs, the first set comprising: (i) a first Click target selective primer comprising a sequence complementary to the R1 sequencing primer site of the universal 5' adapter sequence; and (ii) a second Click target selective primer comprising the same target selective sequence as the second Watson target selective primer sequence, thereby generating a target Click amplification product.

[0103] In some embodiments, the method further comprises purifying the target Watson amplification products and the target click amplification products from non-target polynucleotides. In some embodiments, the purifying step comprises attaching the target Watson amplification products and the target click amplification products to a solid support. In some embodiments, the first Watson target selective primer and the first click target selective primer comprise a first member of an affinity binding pair, and the solid support comprises a second member of the affinity binding pair. In some embodiments, the first member is biotin and the second member is streptavidin. In some embodiments, the solid support comprises a bead, a well, a membrane, a tube, a column, a plate, sepharose, a magnetic bead or a chip. In some embodiments, the method comprises removing polynucleotides that are not attached to the solid support.

[0104] In some embodiments, the method includes the steps of: (a) further amplifying the target Watson amplification product with a second set of Watson target selective primers, the second set comprising: (i) a third Watson target selective primer comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence; and (ii) a fourth Watson target selective primer comprising, in the 5' to 3' direction, an R1 sequencing primer site and a target selective sequence selective for the same target polynucleotide, thereby generating a target Watson library member; and (b) a second set of Click target selective primers, the second set comprising: (i) a third Click target selective primer comprising a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence; and (ii) a fourth Watson target selective primer comprising, in the 5' to 3' direction, an R1 sequencing primer site and a target selective sequence selective for the same target polynucleotide, thereby generating a target Watson library member. and a fourth click target selective primer comprising, in the 5' to 3' direction, an R2 sequencing primer site and a target selective sequence selective for the same target polynucleotide as the fourth Watson target selective primer, thereby creating target click library members.

[0105] In some embodiments, the third Watson target selective primer and the third click target selective primer further comprise a sample barcode sequence. In some embodiments, the third Watson target selective primer further comprises a first grafting sequence that allows hybridization to the first grafting primer on the sequencer, and wherein the third click target selective primer further comprises a second grafting sequence that allows hybridization to the second grafting primer on the sequencer. In some embodiments, the fourth Watson target selective primer further comprises a second grafting sequence, and wherein the fourth click target selective primer further comprises a first grafting sequence. In some embodiments, the first grafting sequence is a P7 sequence, and wherein the second grafting sequence is a P5 sequence.

[0106] Any suitable PCR conditions can be used to generate the amplified target regions described herein (e.g., from a library of amplification products such as a generated double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand). Exemplary PCR conditions are described herein. PCR conditions used to generate the amplified target regions described herein (e.g., from a library of amplification products such as a generated double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) can include any suitable number of PCR amplification cycles. In some cases, PCR amplification can be performed for, but is not limited to, about 1 to about 50 cycles. (For example, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 5 to about 45, about 10 to about 45, about 15 to about 45, about 20 to about 45, about 25 to about 45, about 30 to about 45, about 35 to about 45, about 40 to about 45, about 1 to about 40, about 5 to about 40, about 10 to about 40, about 15 to about 40, about 20 to about 40, about 25 to about 40, about 30 to about 40, about 35 to about 40, about 1 to about 35, about 5 About 1 to about 35, about 10 to about 35, about 15 to about 35, about 20 to about 35, about 25 to about 35, about 30 to about 35, about 1 to about 30, about 5 to about 30, about 10 to about 30, about 15 to about 30, about 20 to about 30, about 25 to about 30, about 1 to about 25, about 5 to about 25, about 10 to about 25, about 15 to about 25, about 20 to about 25, about 1 to about 20, about 5 to about 20, about 10 to about 20, about 15 to about 20, about 1 to about 15, about 5 to about 15, about 10 to about 15, about 1 to about 10, about 5 to about 10, or about 1 to about 5) cycles. For example, when the PCR amplification of the amplified target region includes a single PCR amplification, the PCR amplification can include about 18 amplification cycles. For example, where the PCR amplification of the amplified target region comprises an initial PCR amplification and a subsequent nested PCR amplification, the initial PCR amplification can comprise about 18 amplification cycles and the subsequent nested PCR amplification can comprise about 10 amplification cycles.

[0107] (d) Exemplary Targets Any suitable target region (e.g., a region of interest) can be amplified from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) and evaluated for the presence or absence of one or more mutations. In some cases, the target region can be a region of a nucleic acid in which one or more mutations are associated with a disease or disorder. Examples of target regions that can be amplified and evaluated for the presence or absence of one or more mutations include, but are not limited to, a nucleic acid encoding a tumor protein p53 (TP53), a nucleic acid encoding breast cancer 1 (BRCA1), a nucleic acid encoding BRCA2, a nucleic acid encoding a phosphatase and tensin homolog (PTEN) polypeptide, a nucleic acid encoding an AKT1 polypeptide, a nucleic acid encoding an APC polypeptide, a nucleic acid encoding a CDKN2A polypeptide, a nucleic acid encoding an EGFR polypeptide, a nucleic acid encoding an FBXW7 polypeptide, a nucleic acid encoding a GNAS polypeptide, a nucleic acid encoding a KRAS polypeptide, a nucleic acid encoding a NRAS polypeptide, a nucleic acid encoding a PIK3CA polypeptide, a nucleic acid encoding a BRAF polypeptide, a nucleic acid encoding a CTNNB1 polypeptide, a nucleic acid encoding a FGFR2 polypeptide, a nucleic acid encoding a HRAS polypeptide, and a nucleic acid encoding a PPP2R1A polypeptide. In some cases, a target region that can be amplified and evaluated for the presence or absence of one or more mutations can be a nucleic acid encoding TP53.

[0108] Any suitable method can be used to evaluate the target region (e.g., the amplified target region) for the presence or absence of one or more mutations. In some cases, one or more sequencing methods can be used to evaluate the amplified target region for the presence or absence of one or more mutations.

[0109] (e) Sequence determination In some cases, one or more sequencing methods can be used to evaluate the amplified target region and determine whether mutations are present in both Watson and Crick strands.In some cases, sequencing reads can be used to evaluate the amplified target region for the presence or absence of one or more mutations, and sequencing reads can be used to determine whether mutations are present in both Watson and Crick strands.Examples of sequencing methods that can be used to evaluate the amplified target region for the presence or absence of one or more mutations described herein include, but are not limited to, single-read sequencing, paired-end sequencing, NGS, and deep sequencing.In some embodiments, single-read sequencing includes sequencing that generates sequence reads over the entire length of the template.In some embodiments, sequencing includes paired-end sequencing.In some embodiments, sequencing is performed on a massively parallel sequencer.In some embodiments, the massively parallel sequencer is configured to determine sequence reads from both ends of the template polynucleotide.

[0110] (f) Analysis of sequence reads In some embodiments, the sequence reads are mapped to a reference genome.

[0111] In some embodiments, sequence reads are assigned to a barcode (e.g., UID) family. A barcode family can include sequence reads from an amplified product derived from an original template, e.g., an original double-stranded DNA fragment from a nucleic acid sample.

[0112] In some embodiments, each member of a barcode family comprises the same exogenous barcode sequence. In some embodiments, each member of a barcode family further comprises the same endogenous barcode sequence. Endogenous barcodes are described herein.

[0113] In some embodiments, each member of a barcode family further comprises the same exogenous barcode sequence and the same endogenous barcode sequence. In some embodiments, the combination of the exogenous barcode sequence and the endogenous barcode sequence is unique to the barcode family. In some embodiments, the combination of the exogenous barcode sequence and the endogenous barcode sequence is not present in another barcode family represented in the nucleic acid sample.

[0114] The number of members of a barcode family can depend on the sequencing depth. In some embodiments, a barcode family has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, In some embodiments, the UID family comprises about 2-1000 members, about 2-500 members, about 2-100 members, about 2-500 members, about 2-100 members, about 2-50 members, or about 2-20 members.

[0115] In some embodiments, the sequence reads of each barcode family are assigned to a Watson subfamily and a Crick subfamily. In some embodiments, the sequence reads of each barcode family are assigned to a Watson subfamily and a Crick subfamily based on the orientation of the insert relative to the adapter sequence. In some embodiments, the orientation of the insert relative to the adapter sequence is resolved by how the sequence reads are aligned as a "read pair" or a "mate pair."

[0116] In some embodiments, the assignment of sequence reads to Watson and Crick subfamilies is based on the spatial relationship of the exogenous barcode sequence to the R1 and R2 read sequences. In some embodiments, members of the Watson subfamily are characterized by the exogenous barcode sequence being downstream of the R2 sequence and upstream of the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous barcode sequence being downstream of the R1 sequence and upstream of the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by the exogenous barcode sequence being more proximal to the R2 sequence and less proximal to the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous barcode sequence being more proximal to the R1 sequence and less proximal to the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by an exogenous barcode sequence immediately downstream or within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides of an R2 sequence. In some embodiments, members of the Crick subfamily are characterized by an exogenous barcode sequence immediately downstream or within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides of an R1 sequence.

[0117] In some embodiments, the barcode subfamily (e.g., Watson subfamily and / or Crick subfamily) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290, 310, 320, 330, 340, 350, 360, In some embodiments, the barcode subfamily (e.g., the Watson subfamily and / or the Crick subfamily) comprises about 2-500 members, about 2-100 members, about 2-50 members, about 2-20 members, or about 2-10 members.

[0118] In some embodiments, a nucleotide sequence is determined to accurately represent the Watson strand of an analyte DNA fragment, e.g., a double-stranded DNA fragment from a nucleic acid sample, if a threshold percentage (or a percentage above a threshold) of members of the Watson subfamily contain that sequence. In some embodiments, a nucleotide sequence is determined to accurately represent the Crick strand of an analyte DNA fragment, e.g., a double-stranded DNA fragment from a nucleic acid sample, if a threshold percentage (or a percentage above a threshold) of members of the Crick subfamily contain that sequence.

[0119] The threshold value can be determined by those skilled in the art based on, for example, the number of subfamily members, the specific purpose of the sequencing experiment, and the specific parameters of the sequencing experiment.In some embodiments, the threshold value is set at 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%.In certain embodiments, the threshold value is set at 50%.By way of example only, in an embodiment where the threshold value is set at 50%, a nucleotide sequence is determined to accurately represent an analyte DNA fragment, for example, a Watson strand or a Crick strand of a double-stranded DNA fragment from a nucleic acid sample, if at least 50% of the subfamily members comprise that sequence. As just one other example, in embodiments in which the threshold is set at 50%, a nucleotide sequence is determined to accurately represent an analyte DNA fragment, e.g., a Watson strand or a Crick strand of a double-stranded DNA fragment from a nucleic acid sample, if more than 50% of the subfamily members contain that sequence.

[0120] In some embodiments, a sequence that accurately represents the Watson strand of an analyte DNA fragment is determined to have a mutation. In some embodiments, a sequence that accurately represents the Watson strand of an analyte DNA fragment is determined to have a mutation if the sequence differs from a reference sequence that lacks the mutation.

[0121] In some embodiments, a sequence that accurately represents the Crick strand of an analyte DNA fragment is determined to have a mutation. In some embodiments, a sequence that accurately represents the Crick strand of an analyte DNA fragment is determined to have a mutation if the sequence differs from a reference sequence that lacks the mutation.

[0122] In some embodiments, an analyte DNA fragment is determined to have a mutation if a sequence that exactly represents the Watson strand and a sequence that exactly represents the Crick strand contain the same mutation.

[0123] In some cases, the location of the molecular barcode in the paired-end sequencing reads of the amplified target region can be used to distinguish which strand of the double-stranded nucleic acid template the amplified target region originates from. For example, if the first paired-end sequencing read of the amplified target region indicates that the molecular barcode was read last, the amplified target region can be identified as originating from the sense strand of the nucleic acid template, and if the first paired-end sequencing read of the amplified target region indicates that the molecular barcode was read first, the amplified target region can be identified as originating from the antisense strand of the nucleic acid template. For example, if the second paired-end sequencing read of the amplified target region indicates that the molecular barcode was read first, the amplified target region can be identified as originating from the antisense strand of the nucleic acid template, and if the second paired-end sequencing read of the amplified target region indicates that the molecular barcode was read last, the amplified target region can be identified as originating from the sense strand of the nucleic acid template. In some cases, paired-end sequencing can be used to distinguish amplification products originating from the Watson strand from amplification products originating from the Crick strand.

[0124] After sequencing of a target region (e.g., a target region amplified as described herein), the sequencing reads can be aligned with a reference genome and classified by the molecular barcode present in each sequencing read.In some cases, a sequencing read that contains the same molecular barcode and maps to both the Watson strand and the Crick strand of a double-stranded nucleic acid template (e.g., both the Watson strand and the Crick strand of a target region) can be identified as having duplex support.For example, if a sequencing read that indicates the presence of one or more mutations in a target region contains the same molecular barcode and maps to both the Watson strand and the Crick strand of a target region, the mutation can be identified as having duplex support.

[0125] Amplification of nucleic acid fragments comprising molecular barcodes can be performed according to known techniques to generate a family of barcoded fragments. In some embodiments, polymerase chain reaction (PCR) can be used. In some embodiments, inverse PCR can be used. In some embodiments, rolling circle amplification can be used. Amplification of the fragments is typically performed using a primer complementary to the priming site that is attached to the fragment simultaneously with the molecular barcode. In some embodiments, the priming site is distal to the molecular barcode such that the amplification includes the molecular barcode.

[0126] In some embodiments, the amplification forms a family of fragments, with each member of the family sharing the same molecular barcode. In some embodiments, the diversity of molecular barcodes present on the adapter fragments greatly exceeds the diversity of the fragments, and thus each family is derived from a single nucleic acid fragment molecule. In some embodiments, the primers used for amplification can be chemically modified to make them more resistant to exonucleases. In some embodiments, the family members are sequenced and compared to identify any differences within the family. In some embodiments, the sequencing is performed on a massively parallel sequencing platform, many of which are commercially available. If the sequencing platform requires sequences for "grafting", i.e., attachment to the sequencing instrument, such sequences may be added during the addition of the molecular barcodes or may be added separately. The grafting sequences may be part of the molecular barcoded primer, the universal primer, the gene target specific primer, the amplification primer used to create the family, the sample barcoded primer, or may be separate. Redundant sequencing refers to the sequencing of multiple members of a single family.

[0127] In some embodiments, a threshold for identifying mutations in a nucleic acid fragment can be set. If a "mutation" appears in all members of a family, it originates from the nucleic acid fragment. If a mutation appears in less than all members, it may be an artifact introduced during analysis (e.g., during the amplification step). The threshold for calling a mutation can be set at, for example, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98% or 100%. In some embodiments, the threshold for calling a mutation is 95%, so if 95% of family members sharing the same barcode contain the mutation, the mutation is considered to be real and not an artifact. The threshold will be set based on the number of family members to be sequenced, as well as the specific purpose and situation.

[0128] In some embodiments, one or more sequencing methods can be used to evaluate the amplified DNA molecule and determine whether mutations are present on both strands of double-stranded DNA molecules. In some embodiments, sequencing reads can be used to evaluate the amplified DNA molecule for the presence or absence of one or more mutations, and sequencing reads can be used to determine whether mutations are present on both strands of double-stranded DNA molecules. Examples of sequencing methods that can be used to evaluate the amplified DNA molecule for the presence or absence of one or more mutations described herein include, but are not limited to, single-read sequencing, paired-end sequencing, NGS, and deep sequencing. In some embodiments, single-read sequencing includes sequencing that generates sequence reads over the entire length of the template. In some embodiments, sequencing includes paired-end sequencing. In some embodiments, sequencing is performed on a massively parallel sequencer. In some embodiments, the massively parallel sequencer is configured to determine sequence reads from both ends of the template polynucleotide.

[0129] In some embodiments, the method described herein includes: (g) classifying the first sequencing read according to the molecular barcode present on at least one member of the first group of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing read according to the molecular barcode present on at least one member of the second group of analyte DNA fragments to generate a second analyte DNA family; (i) identifying the genetic features of the tagged Watson strand and Crick strand in the first analyte DNA family; and (j) identifying the epigenetic features of the matched Watson strand and Crick strand in the second analyte DNA family, thus identifying the genetic and epigenetic features present on at least one strand of the double-stranded DNA molecule. In some embodiments, the method includes identifying the genetic and epigenetic features present on both strands of the double-stranded DNA molecule.

[0130] Epigenetic Signatures – Methylation Analysis As used herein, the term "epigenetic signature" can refer to a heritable phenotypic change without a change in DNA sequence. In some embodiments, the epigenetic signature comprises a functionally relevant change to the genome without a change in nucleotide sequence. In some embodiments, the epigenetic signature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation. In some embodiments, the epigenetic signature is methylation. In some embodiments, the epigenetic signature is a differentially methylated region (DMR). In some embodiments, the epigenetic signature is a methylation pattern. In some embodiments, the methylation pattern corresponds to the methylation pattern present in cells generated via clonal hematopoiesis of indeterminate origin. In some embodiments, the methylation pattern corresponds to the methylation pattern present in the tissue of origin. In some embodiments, the tissue of origin is anus, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus (Cypris et al., Front. Genet. 10:785 (2019), Liu et al., Ann Oncol.31(6):745-759 (2020)).

[0131] In some embodiments, the methods described herein can be used to detect methylation at CpG dinucleotides in one or both strands of a double-stranded DNA molecule (e.g., both strands simultaneously). In some embodiments, a population of DNA molecules is treated with bisulfite to convert cytosine bases in the DNA molecules to uracil bases, forming a population of converted DNA molecules. In some embodiments, an excess of target-specific amplification primers attached to molecular barcodes are used to attach molecular barcodes to both strands of the population of converted DNA molecules, forming a population of amplified, barcoded, converted DNA molecules. In some embodiments, the amplified, barcoded, converted DNA molecules are amplified in an amplification reaction to form a family of amplified, barcoded, converted DNA molecules, where the amplified, barcoded, converted DNA molecules that share the same molecular barcode form a family of DNA molecules. In some embodiments, a plurality of members of the family are subjected to a sequencing reaction to obtain the nucleotide sequence of both strands of the plurality of members of the family. In some embodiments, the nucleotide sequences of multiple members of a family are compared and families in which more than 90% of the members contain a selected methylated C at a CpG dinucleotide are identified. In some embodiments, the nucleotide sequences of two complementary strands of the amplified, barcoded, converted DNA molecule are compared and a methylated C at a CpG dinucleotide is identified in the two complementary strands.

[0132] In some embodiments, incubation of DNA fragments with sodium bisulfite at elevated temperature and low pH deaminates cytosine to form 5,6-dihydrocytosine-6-sulfonate. Exemplary methods of sodium bisulfite treatment for use in the methods disclosed herein are described in PCT / US2018 / 022664, which is incorporated herein by reference in its entirety. Subsequent hydrolytic deamination at high pH removes the sulfonate to yield uracil. Many modifications of this basic reaction have been described and are primarily used to distinguish between cytosine and 5-methylcytosine (5-mC), the latter of which is not susceptible to conversion with bisulfite. In addition to converting C to U, bisulfite treatment can denature DNA and degrade it. This degradation is not limited to standard applications of bisulfite treatment, but is important for applications involving mutation detection in clinical samples that are already degraded prior to conversion. In some embodiments, sequencing of these products revealed that, on average, greater than 99.8% of C bases were converted to T bases on both strands (except for C bases at 5'-CpG sites, which may be resistant to conversion with bisulfite because they are methylated or hydroxymethylated).

[0133] (5) Identification of multiple features of double-stranded DNA molecules Also provided herein is a method for identifying a first feature and a second feature of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule, the method comprising: (a) attaching an adapter fragment to each end of the double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) contacting the adapted double-stranded DNA molecule with a tagged primer. (c) subjecting the amplified products to denaturing conditions; (d) separately recovering the adapted Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson and Crick strands and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (f) generating a second population of analyte DNA fragments from the adapted Watson and Crick strands and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on the at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing reads according to the molecular barcode present on said at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family;The method includes: (i) identifying a first feature of the tagged Watson strand and Crick strand in a first analyte DNA family; and (j) identifying a second feature of the matched Watson strand and Crick strand in a second analyte DNA family, thus identifying the first feature and the second feature present on at least one strand of the double-stranded DNA molecule. In some embodiments, the method includes identifying the first feature and the second feature present on both strands of the double-stranded DNA molecule;

[0134] In some embodiments, the first characteristic is a genetic characteristic. In some embodiments, the second characteristic is an epigenetic characteristic. In some embodiments, the first characteristic is a genetic characteristic or an epigenetic characteristic. In some embodiments, the second characteristic is an epigenetic characteristic or a genetic characteristic. In some embodiments, the first characteristic and the second characteristic are both genetic characteristics. In some embodiments, the first characteristic and the second characteristic are both epigenetic characteristics.

[0135] In some embodiments, the genetic feature is a mutation. In some embodiments, the mutation is selected from the group consisting of insertion, deletion, substitution, deletion-insertion, duplication, inversion, frameshift, repeat expansion, translocation, and combinations thereof. In some embodiments, the identification of genetic feature comprises mutation analysis, aneuploidy analysis, or fragmentomics.

[0136] In some embodiments, the epigenetic signature is methylation. In some embodiments, the epigenetic signature is a methylation pattern. In some embodiments, the methylation pattern corresponds to the methylation pattern present in cells generated through clonal hematopoiesis of indeterminate origin. In some embodiments, the methylation pattern corresponds to the methylation pattern present in the tissue of origin. In some embodiments, the tissue of origin is anus, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus. In some embodiments, the epigenetic signature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation.

[0137] In some embodiments, the first feature and the second feature are both epigenetic features, where the first feature is methylation and the second feature is hydroxymethylation. In some embodiments, the first feature is methylation and the second feature is acetylation. In some embodiments, the first feature is methylation and the second feature is a histone modification. In some embodiments, the first feature is methylation and the second feature is microRNA regulation. In some embodiments, the first feature is methylation and the second feature is phosphorylation. In some embodiments, the first feature is methylation and the second feature is ubiquitination. In some embodiments, the first feature is methylation and the second feature is sumoylation. In some embodiments, the first feature is hydroxymethylation and the second feature is methylation. In some embodiments, the first feature is hydroxymethylation and the second feature is acetylation. In some embodiments, the first feature is hydroxymethylation and the second feature is a histone modification. In some embodiments, the first characteristic is hydroxymethylation and the second characteristic is microRNA regulation. In some embodiments, the first characteristic is hydroxymethylation and the second characteristic is phosphorylation. In some embodiments, the first characteristic is hydroxymethylation and the second characteristic is ubiquitination. In some embodiments, the first characteristic is hydroxymethylation and the second characteristic is sumoylation. In some embodiments, the first characteristic is histone modification and the second characteristic is methylation. In some embodiments, the first characteristic is histone modification and the second characteristic is acetylation. In some embodiments, the first characteristic is histone modification and the second characteristic is hydroxymethylation. In some embodiments, the first characteristic is histone modification and the second characteristic is microRNA regulation. In some embodiments, the first characteristic is histone modification and the second characteristic is phosphorylation. In some embodiments, the first characteristic is histone modification and the second characteristic is ubiquitination.In some embodiments, the first characteristic is a histone modification and the second characteristic is sumoylation. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is methylation. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is acetylation. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is hydroxymethylation. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is histone modification. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is phosphorylation. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is ubiquitination. In some embodiments, the first characteristic is microRNA regulation and the second characteristic is sumoylation. In some embodiments, the first characteristic is acetylation and the second characteristic is methylation. In some embodiments, the first characteristic is acetylation and the second characteristic is microRNA regulation. In some embodiments, the first characteristic is acetylation and the second characteristic is hydroxymethylation. In some embodiments, the first characteristic is acetylation and the second characteristic is a histone modification. In some embodiments, the first characteristic is acetylation and the second characteristic is phosphorylation. In some embodiments, the first characteristic is acetylation and the second characteristic is ubiquitination. In some embodiments, the first characteristic is acetylation and the second characteristic is sumoylation. In some embodiments, the first characteristic is phosphorylation and the second characteristic is methylation. In some embodiments, the first characteristic is phosphorylation and the second characteristic is microRNA regulation. In some embodiments, the first characteristic is phosphorylation and the second characteristic is hydroxymethylation. In some embodiments, the first characteristic is phosphorylation and the second characteristic is a histone modification. In some embodiments, the first characteristic is phosphorylation and the second characteristic is acetylation. In some embodiments, the first characteristic is phosphorylation and the second characteristic is ubiquitination.In some embodiments, the first characteristic is phosphorylation and the second characteristic is sumoylation. In some embodiments, the first characteristic is ubiquitination and the second characteristic is methylation. In some embodiments, the first characteristic is ubiquitination and the second characteristic is microRNA regulation. In some embodiments, the first characteristic is ubiquitination and the second characteristic is hydroxymethylation. In some embodiments, the first characteristic is ubiquitination and the second characteristic is histone modification. In some embodiments, the first characteristic is ubiquitination and the second characteristic is acetylation. In some embodiments, the first characteristic is ubiquitination and the second characteristic is phosphorylation. In some embodiments, the first characteristic is ubiquitination and the second characteristic is sumoylation. In some embodiments, the first characteristic is sumoylation and the second characteristic is methylation. In some embodiments, the first characteristic is sumoylation and the second characteristic is microRNA regulation. In some embodiments, the first characteristic is sumoylation and the second characteristic is hydroxymethylation. In some embodiments, the first characteristic is sumoylation and the second characteristic is a histone modification. In some embodiments, the first characteristic is sumoylation and the second characteristic is acetylation. In some embodiments, the first characteristic is sumoylation and the second characteristic is phosphorylation. In some embodiments, the first characteristic is sumoylation and the second characteristic is ubiquitination.

[0138] Genetics – Sequencing In some embodiments, the first and / or second characteristic can be a genetic characteristic, where the term "genetic characteristic" refers to genetic information and / or material that is replicated and passed on from parent cells to progeny cells at each cell division. In some embodiments, the genetic characteristic can be a mutation in a nucleic acid (e.g., a DNA molecule). In some embodiments, the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and combinations thereof. In some embodiments, the identification of the genetic characteristic can include mutation analysis, aneuploidy analysis, or fragmentomics. Exemplary methods for identifying genetic characteristics suitable for use in the methods disclosed herein are described in PCT / US2021 / 017937, which is incorporated herein by reference in its entirety.

[0139] (a) Initial amplification of adaptor-attached templates. After adapter attachment, the adapted double-stranded DNA molecule can be amplified in an initial amplification reaction (e.g., PCR amplification). Any suitable method can be used to amplify the adapted double-stranded DNA molecule. Exemplary methods that can be used to amplify the adapted double-stranded DNA molecule include, but are not limited to, whole genome PCR.

[0140] Any suitable primer pair can be used to amplify the adapted double-stranded DNA molecule. In some embodiments, a universal primer pair can be used. The primers can include, but are not limited to, about 12 nucleotides to about 30 nucleotides. In some embodiments, any suitable PCR conditions can be used in the initial amplification. The PCR amplification can include a denaturation phase, an annealing phase, and an extension phase. Each phase of the amplification cycle can include any suitable conditions. In some cases, the denaturation phase can include a temperature of about 90°C to about 105°C (e.g., about 94°C to about 98°C) and a time of about 1 second to about 5 minutes (e.g., about 10 seconds to about 1 minute). For example, the denaturation phase can include a temperature of about 98°C for about 10 seconds. In some cases, the annealing phase can include a temperature of about 50°C to about 72°C and a time of about 30 seconds to about 90 seconds. In some cases, the extension phase can include a temperature of about 55°C to about 80°C and a time of about 15 seconds per kb of amplicon generated to about 30 seconds per kb of amplicon generated. In some cases, the annealing and extension phases can be performed in a single cycle. For example, the annealing and extension phases can include a temperature of about 65° C. for about 75 seconds.

[0141] The PCR conditions used in the initial amplification can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can be performed for about 1 to about 50 (e.g., about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 5 to about 45, about 10 to about 45, about 15 to about 45, about 20 to about 45, about 25 to about 45, about 30 to about 45, about 35 to about 45, about 40 to about 45, about 1 to about 40, about 5 to about 40, about 10 to about 40, about 15 to about 40, about 20 to about 40, about 25 to about 40, about 30 to about 40, about 35 to about 40, about 1 to about 35, about 5 The PCR amplification may comprise about 1 to about 35, about 10 to about 35, about 15 to about 35, about 20 to about 35, about 25 to about 35, about 30 to about 35, about 1 to about 30, about 5 to about 30, about 10 to about 30, about 15 to about 30, about 20 to about 30, about 25 to about 30, about 1 to about 25, about 5 to about 25, about 10 to about 25, about 15 to about 25, about 20 to about 25, about 1 to about 20, about 5 to about 20, about 10 to about 20, about 15 to about 20, about 1 to about 15, about 5 to about 15, about 10 to about 15, about 1 to about 10, about 5 to about 10, or about 1 to about 5 cycles. In some embodiments, the PCR amplification comprises 11 cycles or less. In some embodiments, the PCR amplification comprises 7 cycles or less. In some embodiments, the PCR amplification comprises 5 cycles or less.

[0142] In some cases, when the PCR conditions include a heat-activated polymerase, the PCR amplification can also include an initialization step. For example, the PCR amplification can include an initialization step before performing a PCR amplification cycle. In some cases, the initialization step can include a temperature of about 94°C to about 98°C, and a time period of about 15 seconds to about 1 minute. For example, the initialization step can include a temperature of about 98°C for about 30 seconds.

[0143] In some cases, the PCR amplification can also include a holding step. For example, the PCR amplification can include a holding step after performing the PCR amplification cycles, and optionally after performing a final extension step. In some cases, the holding step can include a temperature of about 4° C. to about 15° C. for an unlimited amount of time.

[0144] In some cases, the double-stranded sequencing library (e.g., the amplified double-stranded sequencing library) generated as described herein can be purified.Any suitable method can be used to purify the double-stranded sequencing library.The exemplary method that can be used to purify the double-stranded sequencing library includes, but is not limited to, magnetic beads (e.g., solid-phase reversible immobilization (SPRI) magnetic beads).

[0145] (b) Optional ssDNA library preparation In some cases, the double-stranded sequencing library can be used to generate a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands. By generating a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands, non-specific amplification (e.g., from primers complementary to ligation sequences such as 3' double-stranded adapters or 5' adapters) can be minimized. Any suitable method can be used to generate a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands (e.g., from a double-stranded sequencing library made as described herein). In some cases, the library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands can be generated from the amplified double-stranded sequencing library by dividing the amplification product into at least two aliquots and subjecting each aliquot to PCR amplification, in which Watson strands are amplified from the first aliquot and Crick strands are amplified from the second aliquot. For example, a first aliquot of the amplified products from the double-stranded sequencing library can be subjected to PCR amplification with a primer pair, the first primer being biotinylated and the second primer being non-biotinylated, to generate a single-stranded library of Watson strand, and a second aliquot of the amplified products from the double-stranded sequencing library can be subjected to PCR amplification with a primer pair, the first primer being non-biotinylated and the second primer being biotinylated, to generate a single-stranded library of Crick strand.In some cases, a library of sequences derived from single-stranded Watson strand and a library of sequences derived from single-stranded Crick strand can be generated.

[0146] Any suitable method can be used to generate a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand from an amplified double-stranded sequencing library. For example, the amplification products from the amplified double-stranded sequencing library can be separated into a first PCR amplification and a second PCR amplification, in which only one of the two primers in the PCR primer pair is tagged. For example, the first PCR amplification can use a primer pair that includes a tagged primer (e.g., the first primer) and an untagged primer (e.g., the second primer), and the second PCR amplification can use a primer pair that includes an untagged primer (e.g., the first primer) and a tagged primer (e.g., the second primer). The primer tag can be any tag that allows the PCR amplification products generated from the tagged primer to be recovered. In some cases, the tagged primer can be a biotinylated primer, and the PCR amplification products generated from the biotinylated primer can be recovered using streptavidin. In some cases, the tagged primer can be a biotinylated primer containing uracil, and the PCR amplification products generated from the biotinylated primer containing uracil can be recovered using streptavidin.For example, the library of sequences from single-stranded Watson strand and the library of sequences from single-stranded Crick strand can be generated in PCR amplification using a primer pair that includes a biotinylated primer and a non-biotinylated primer.In some cases, the tagged primer can be a phosphorylated primer, and the PCR amplification products generated from the phosphorylated primer can be recovered using lambda nuclease.For example, the library of sequences from single-stranded Watson strand and the library of sequences from single-stranded Crick strand can be generated in PCR amplification using a primer pair that includes a phosphorylated primer and a non-phosphorylated primer.

[0147] Any suitable primer pair can be used to generate a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand (e.g., from a duplex sequencing library generated as described herein). The primers can include, without limitation, about 12 nucleotides to about 30 nucleotides. In some cases, the primer pair can include at least one primer that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification product generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to the nucleic acid fragments in the duplex sequencing library prior to amplification). Examples of primer pairs that can be used to generate a library of sequences derived from a single-stranded Watson strand and a library of sequences derived from a single-stranded Crick strand described herein include, without limitation, P5 primer and P7 primer.

[0148] Any suitable PCR conditions can be used to generate a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand (e.g., from a double-stranded sequencing library generated as described herein). PCR amplification can include a denaturation phase, an annealing phase, and an extension phase. Each phase of an amplification cycle can include any suitable conditions. In some cases, the denaturation phase can include a temperature of about 90°C to about 105°C, and a time of about 1 second to about 5 minutes. For example, the denaturation phase can include a temperature of about 98°C for about 10 seconds. In some cases, the annealing phase can include a temperature of about 50°C to about 72°C, and a time of about 30 seconds to about 90 seconds. In some cases, the extension phase can include a temperature of about 55°C to about 80°C, and a time of about 15 seconds per kb of amplicon generated to about 30 seconds per kb of amplicon generated. In some cases, the extension phase reflects the processivity of the polymerase used. In some cases, the annealing phase and the extension phase can be performed in a single cycle. For example, the annealing and extension phases can include a temperature of about 65° C. for about 75 seconds.

[0149] The PCR conditions used to generate the libraries of sequences derived from the single-stranded Watson strand and the libraries of sequences derived from the single-stranded Crick strand (e.g., from a double-stranded sequencing library generated as described herein) can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can be, but is not limited to, about 1 to about 50 cycles. (For example, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 5 to about 45, about 10 to about 45, about 15 to about 45, about 20 to about 45, about 25 to about 45, about 30 to about 45, about 35 to about 45, about 40 to about 45, about 1 to about 40, about 5 to about 40, about 10 to about 40, about 15 to about 40, about 20 to about 40, about 25 to about 40, about 30 to about 40, about 35 to about 40, about 1 to about 35, about 5 About 1 to about 35, about 10 to about 35, about 15 to about 35, about 20 to about 35, about 25 to about 35, about 30 to about 35, about 1 to about 30, about 5 to about 30, about 10 to about 30, about 15 to about 30, about 20 to about 30, about 25 to about 30, about 1 to about 25, about 5 to about 25, about 10 to about 25, about 15 to about 25, about 20 to about 25, about 1 to about 20, about 5 to about 20, about 10 to about 20, about 15 to about 20, about 1 to about 15, about 5 to about 15, about 10 to about 15, about 1 to about 10, about 5 to about 10, or about 1 to about 5) cycles. For example, the PCR amplification can include about 4 amplification cycles. In some embodiments, the PCR amplification can include about 8 amplification cycles.

[0150] In some cases, when the PCR conditions include a heat-activated polymerase, the PCR amplification can also include an initialization step. For example, the PCR amplification can include an initialization step before performing a PCR amplification cycle. In some cases, the initialization step can include a temperature of about 94°C to about 98°C, and a time period of about 15 seconds to about 1 minute. For example, the initialization step can include a temperature of about 98°C for about 30 seconds.

[0151] In some cases, the PCR amplification can also include a holding step. For example, the PCR amplification can include a holding step after performing the PCR amplification cycles, and optionally after performing a final extension step. In some cases, the holding step can include a temperature of about 4° C. to about 15° C. for an unlimited amount of time.

[0152] Any suitable method can be used to separate double-stranded amplification product into single-stranded amplification product.In some cases, double-stranded amplification product can be denatured to separate double-stranded amplification product into two single-stranded amplification products.The examples of the method that can be used to separate double-stranded amplification product into single-stranded amplification product include, but are not limited to, heat denaturation, chemical (e.g., NaOH) denaturation, and salt denaturation.

[0153] After PCR amplification, tagged Watson strand and Crick strand can be collected. Any suitable method can be used to collect tagged Watson strand and Crick strand produced by using tagged primer. If tagged primer is biotinylated primer, biotinylated amplification product (e.g., produced from biotinylated primer) can be collected by using streptavidin (e.g., streptavidin functionalized beads). For example, when the amplified double-stranded sequencing library is further amplified in a first PCR amplification using a primer pair comprising a first biotinylated primer and a second non-biotinylated primer, and a second PCR amplification using a primer pair comprising a first non-biotinylated primer and a second biotinylated primer, the biotinylated amplification products generated from the first PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first set of streptavidin-functionalized beads), and the biotinylated amplification products generated from the second PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first set of streptavidin-functionalized beads), and the double-stranded amplification products can be separated (e.g., denatured) into single-stranded amplification products.In some cases, recovering the biotinylated PCR amplification products can also include releasing the biotinylated PCR amplification products from streptavidin (e.g., streptavidin-functionalized beads).Separating the double-stranded amplification products generated by the first PCR amplification using a primer pair comprising a first biotinylated primer and a second non-biotinylated primer, and the second PCR amplification using a primer pair comprising a first non-biotinylated primer and a second biotinylated primer, can allow the single-stranded amplification products generated from the biotinylated primer to remain bound to the streptavidin-functionalized beads, while the single-stranded amplification products generated from the non-biotinylated primer can be denatured (e.g., denatured and degraded) from the streptavidin-functionalized beads, thereby generating a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences of the double-stranded sequencing library.

[0154] When the tagged primer is a phosphorylated primer, the phosphorylated amplification product (e.g., generated from the phosphorylated primer) can be separated from the non-phosphorylated amplification product by using an exonuclease (e.g., lambda exonuclease).For example, when the amplified double-stranded sequencing library is further amplified in a first PCR amplification using a primer pair comprising a first phosphorylated primer and a second non-phosphorylated primer, and in a second PCR amplification using a primer pair comprising a first non-phosphorylated primer and a second phosphorylated primer, the double-stranded amplification product can be separated into a single-stranded amplification product. Separating the double-stranded amplification products generated by the first PCR amplification using a primer pair comprising a first phosphorylated primer and a second unphosphorylated primer, and the second PCR amplification using a primer pair comprising a first unphosphorylated primer and a second phosphorylated primer, can enable the single-stranded amplification products generated from the unphosphorylated primer to be recovered, while the single-stranded amplification products generated from the phosphorylated primer can be degraded by lambda exonuclease, thereby generating a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand of the double-stranded sequencing library.

[0155] (c) Target enrichment In some embodiments of any one of the methods herein, the amplification product produced by initial amplification is enriched for one or more target polynucleotides.In some embodiments, before target enrichment, a single-stranded DNA library is prepared from the amplification product produced by initial amplification.An exemplary method for producing a single-stranded DNA library is described herein.

[0156] Any suitable method can be used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand). In some cases, a target region can be amplified from a library of amplification products by subjecting the library of amplification products to PCR amplification using a primer pair, which is a primer (e.g., a first primer) that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to the nucleic acid fragments in the double-stranded sequencing library prior to amplification), and a primer (e.g., a second primer) that can target (e.g., target and bind to) the target region (e.g., a region of interest).

[0157] In some cases, a target region can be amplified from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) in a single PCR amplification. For example, a target region can be amplified from a library of amplification products in a single PCR amplification using a primer pair that includes a first primer that can target an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in an amplification product generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in the double-stranded sequencing library prior to amplification), and a second primer that can target the target region.

[0158] In some cases, the target region can be amplified from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) in multiple PCR amplifications.Multiple PCR amplifications (e.g., an initial PCR amplification and a subsequent nested PCR amplification) can be used to increase the specificity of amplifying the target region. For example, a target region can be amplified from a library of amplification products in a series of PCR amplifications in which a first PCR amplification is performed using a primer pair comprising a first primer capable of targeting an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a second primer capable of targeting the target region, and the amplification products generated in the first PCR amplification are subjected to a subsequent nested PCR amplification using a primer pair comprising a first primer capable of targeting an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a second primer capable of targeting a nucleic acid sequence from the target region present in the amplification products generated in the first PCR amplification.

[0159] Any suitable primer pair can be used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand). The primers can include, without limitation, about 12 nucleotides to about 30 nucleotides. In some cases, the primer pair can include a primer (e.g., a first primer) that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in the double-stranded sequencing library prior to amplification), as well as a primer (e.g., a second primer) that can target (e.g., target and bind to) a target region (e.g., a region of interest). Examples of primers that can target adapter sequences that include molecular barcodes present in amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification) include, but are not limited to, i5 index primers and i7 index primers. Primers that can target a target region can include a sequence that is complementary to the target region. When the target region is a nucleic acid encoding TP53, examples of primers that can target a nucleic acid encoding TP53 include, but are not limited to, TP53_342_GSP1 and TP53_GSP2.

[0160] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) can include one or more molecular barcodes.

[0161] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) can contain one or more grafting sequences (e.g., grafting sequences for next-generation sequencing).

[0162] In some embodiments, target enrichment comprises: (a) selectively amplifying an amplification product of a Watson strand comprising a target polynucleotide sequence with a first set of Watson target selective primer pairs, the first set comprising: (i) a first Watson target selective primer comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence; and (ii) a second Watson target selective primer comprising a target selective sequence, thereby generating a target Watson amplification product; and (b) selectively amplifying an amplification product of a Click strand comprising the same target polynucleotide sequence with a first set of Click target selective primer pairs, the first set comprising: (i) a first Click target selective primer comprising a sequence complementary to the R1 sequencing primer site of the universal 5' adapter sequence; and (ii) a second Click target selective primer comprising the same target selective sequence as the second Watson target selective primer sequence, thereby generating a target Click amplification product.

[0163] In some embodiments, the method further comprises purifying the target Watson amplification products and the target click amplification products from non-target polynucleotides. In some embodiments, the purifying step comprises attaching the target Watson amplification products and the target click amplification products to a solid support. In some embodiments, the first Watson target selective primer and the first click target selective primer comprise a first member of an affinity binding pair, and the solid support comprises a second member of the affinity binding pair. In some embodiments, the first member is biotin and the second member is streptavidin. In some embodiments, the solid support comprises a bead, a well, a membrane, a tube, a column, a plate, sepharose, a magnetic bead or a chip. In some embodiments, the method comprises removing polynucleotides that are not attached to the solid support.

[0164] In some embodiments, the method includes the steps of: (a) further amplifying the target Watson amplification product with a second set of Watson target selective primers, the second set comprising: (i) a third Watson target selective primer comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence; and (ii) a fourth Watson target selective primer comprising, in the 5' to 3' direction, an R1 sequencing primer site and a target selective sequence selective for the same target polynucleotide, thereby generating a target Watson library member; and (b) a second set of Click target selective primers, the second set comprising: (i) a third Click target selective primer comprising a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence; and (ii) a fourth Watson target selective primer comprising, in the 5' to 3' direction, an R1 sequencing primer site and a target selective sequence selective for the same target polynucleotide, thereby generating a target Watson library member. and a fourth click target selective primer comprising, in the 5' to 3' direction, an R2 sequencing primer site and a target selective sequence selective for the same target polynucleotide as the fourth Watson target selective primer, thereby creating target click library members.

[0165] In some embodiments, the third Watson target selective primer and the third click target selective primer further comprise a sample barcode sequence. In some embodiments, the third Watson target selective primer further comprises a first grafting sequence that allows hybridization to the first grafting primer on the sequencer, and wherein the third click target selective primer further comprises a second grafting sequence that allows hybridization to the second grafting primer on the sequencer. In some embodiments, the fourth Watson target selective primer further comprises a second grafting sequence, and wherein the fourth click target selective primer further comprises a first grafting sequence. In some embodiments, the first grafting sequence is a P7 sequence, and wherein the second grafting sequence is a P5 sequence.

[0166] Any suitable PCR conditions can be used to generate the amplified target regions described herein (e.g., from a library of amplification products such as a generated double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand). Exemplary PCR conditions are described herein. PCR conditions used to generate the amplified target regions described herein (e.g., from a library of amplification products such as a generated double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) can include any suitable number of PCR amplification cycles. In some cases, PCR amplification can be performed for, but is not limited to, about 1 to about 50 cycles. (For example, about 5 to about 50, about 10 to about 50, about 15 to about 50, about 20 to about 50, about 25 to about 50, about 30 to about 50, about 35 to about 50, about 40 to about 50, about 45 to about 50, about 1 to about 45, about 5 to about 45, about 10 to about 45, about 15 to about 45, about 20 to about 45, about 25 to about 45, about 30 to about 45, about 35 to about 45, about 40 to about 45, about 1 to about 40, about 5 to about 40, about 10 to about 40, about 15 to about 40, about 20 to about 40, about 25 to about 40, about 30 to about 40, about 35 to about 40, about 1 to about 35, about 5 About 1 to about 35, about 10 to about 35, about 15 to about 35, about 20 to about 35, about 25 to about 35, about 30 to about 35, about 1 to about 30, about 5 to about 30, about 10 to about 30, about 15 to about 30, about 20 to about 30, about 25 to about 30, about 1 to about 25, about 5 to about 25, about 10 to about 25, about 15 to about 25, about 20 to about 25, about 1 to about 20, about 5 to about 20, about 10 to about 20, about 15 to about 20, about 1 to about 15, about 5 to about 15, about 10 to about 15, about 1 to about 10, about 5 to about 10, or about 1 to about 5) cycles. For example, when the PCR amplification of the amplified target region includes a single PCR amplification, the PCR amplification can include about 18 amplification cycles. For example, where the PCR amplification of the amplified target region comprises an initial PCR amplification and a subsequent nested PCR amplification, the initial PCR amplification can comprise about 18 amplification cycles and the subsequent nested PCR amplification can comprise about 10 amplification cycles.

[0167] (d) Exemplary Targets Any suitable target region (e.g., a region of interest) can be amplified from a library of amplification products (e.g., a double-stranded sequencing library generated as described herein, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand) and evaluated for the presence or absence of one or more mutations. In some cases, the target region can be a region of a nucleic acid in which one or more mutations are associated with a disease or disorder. Examples of target regions that can be amplified and evaluated for the presence or absence of one or more mutations include, but are not limited to, a nucleic acid encoding a tumor protein p53 (TP53), a nucleic acid encoding breast cancer 1 (BRCA1), a nucleic acid encoding BRCA2, a nucleic acid encoding a phosphatase and tensin homolog (PTEN) polypeptide, a nucleic acid encoding an AKT1 polypeptide, a nucleic acid encoding an APC polypeptide, a nucleic acid encoding a CDKN2A polypeptide, a nucleic acid encoding an EGFR polypeptide, a nucleic acid encoding an FBXW7 polypeptide, a nucleic acid encoding a GNAS polypeptide, a nucleic acid encoding a KRAS polypeptide, a nucleic acid encoding a NRAS polypeptide, a nucleic acid encoding a PIK3CA polypeptide, a nucleic acid encoding a BRAF polypeptide, a nucleic acid encoding a CTNNB1 polypeptide, a nucleic acid encoding a FGFR2 polypeptide, a nucleic acid encoding a HRAS polypeptide, and a nucleic acid encoding a PPP2R1A polypeptide. In some cases, a target region that can be amplified and evaluated for the presence or absence of one or more mutations can be a nucleic acid encoding TP53.

[0168] Any suitable method can be used to evaluate the target region (e.g., the amplified target region) for the presence or absence of one or more mutations. In some cases, one or more sequencing methods can be used to evaluate the amplified target region for the presence or absence of one or more mutations.

[0169] (e) Sequence determination In some cases, one or more sequencing methods can be used to evaluate the amplified target region and determine whether mutations are present in both Watson and Crick strands. In some cases, sequencing reads can be used to evaluate the amplified target region for the presence or absence of one or more mutations, and sequencing reads can be used to determine whether mutations are present in both Watson and Crick strands. Examples of sequencing methods that can be used to evaluate the amplified target region for the presence or absence of one or more mutations described herein include, but are not limited to, single-read sequencing, paired-end sequencing, NGS, and deep sequencing. In some embodiments, single-read sequencing includes sequencing that generates sequence reads over the entire length of the template. In some embodiments, sequencing includes paired-end sequencing. In some embodiments, sequencing is performed on a massively parallel sequencer. In some embodiments, the massively parallel sequencer is configured to determine sequence reads from both ends of the template polynucleotide. In some embodiments, sequencing includes whole genome PCR, whole genome bisulfite sequencing, or capture sequencing.

[0170] (f) Analysis of sequence reads In some embodiments, the sequence reads are mapped to a reference genome.

[0171] In some embodiments, sequence reads are assigned to a barcode (e.g., UID) family. A barcode family can include sequence reads from an amplified product derived from an original template, e.g., an original double-stranded DNA fragment from a nucleic acid sample.

[0172] In some embodiments, each member of a barcode family comprises the same exogenous barcode sequence. In some embodiments, each member of a barcode family further comprises the same endogenous barcode sequence. Endogenous barcodes are described herein.

[0173] In some embodiments, each member of a barcode family further comprises the same exogenous barcode sequence and the same endogenous barcode sequence. In some embodiments, the combination of the exogenous barcode sequence and the endogenous barcode sequence is unique to the barcode family. In some embodiments, the combination of the exogenous barcode sequence and the endogenous barcode sequence is not present in another barcode family represented in the nucleic acid sample.

[0174] The number of members of a barcode family can depend on the sequencing depth. In some embodiments, a barcode family has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, In some embodiments, the UID family comprises about 2-1000 members, about 2-500 members, about 2-100 members, about 2-500 members, about 2-100 members, about 2-50 members, or about 2-20 members.

[0175] In some embodiments, the sequence reads of each barcode family are assigned to a Watson subfamily and a Crick subfamily. In some embodiments, the sequence reads of each barcode family are assigned to a Watson subfamily and a Crick subfamily based on the orientation of the insert relative to the adapter sequence. In some embodiments, the orientation of the insert relative to the adapter sequence is resolved by how the sequence reads are aligned as a "read pair" or a "mate pair".

[0176] In some embodiments, the assignment of sequence reads to Watson and Crick subfamilies is based on the spatial relationship of the exogenous barcode sequence to the R1 and R2 read sequences. In some embodiments, members of the Watson subfamily are characterized by the exogenous barcode sequence being downstream of the R2 sequence and upstream of the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous barcode sequence being downstream of the R1 sequence and upstream of the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by the exogenous barcode sequence being more proximal to the R2 sequence and less proximal to the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous barcode sequence being more proximal to the R1 sequence and less proximal to the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by an exogenous barcode sequence immediately downstream or within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides of an R2 sequence. In some embodiments, members of the Crick subfamily are characterized by an exogenous barcode sequence immediately downstream or within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides of an R1 sequence.

[0177] In some embodiments, the barcode subfamily (e.g., Watson subfamily and / or Crick subfamily) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290, 310, 320, 330, 340, 350, 360, In some embodiments, the barcode subfamily (e.g., the Watson subfamily and / or the Crick subfamily) comprises about 2-500 members, about 2-100 members, about 2-50 members, about 2-20 members, or about 2-10 members.

[0178] In some embodiments, a nucleotide sequence is determined to accurately represent the Watson strand of an analyte DNA fragment, e.g., a double-stranded DNA fragment from a nucleic acid sample, if a threshold percentage (or a percentage above a threshold) of members of the Watson subfamily contain that sequence. In some embodiments, a nucleotide sequence is determined to accurately represent the Crick strand of an analyte DNA fragment, e.g., a double-stranded DNA fragment from a nucleic acid sample, if a threshold percentage (or a percentage above a threshold) of members of the Crick subfamily contain that sequence.

[0179] The threshold value can be determined by those skilled in the art based on, for example, the number of subfamily members, the specific purpose of the sequencing experiment, and the specific parameters of the sequencing experiment.In some embodiments, the threshold value is set at 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%.In certain embodiments, the threshold value is set at 50%.By way of example only, in an embodiment where the threshold value is set at 50%, a nucleotide sequence is determined to accurately represent an analyte DNA fragment, for example, a Watson strand or a Crick strand of a double-stranded DNA fragment from a nucleic acid sample, if at least 50% of the subfamily members comprise that sequence. As just one other example, in embodiments in which the threshold is set at 50%, a nucleotide sequence is determined to accurately represent an analyte DNA fragment, e.g., a Watson strand or a Crick strand of a double-stranded DNA fragment from a nucleic acid sample, if more than 50% of the subfamily members contain that sequence.

[0180] In some embodiments, a sequence that accurately represents the Watson strand of an analyte DNA fragment is determined to have a mutation. In some embodiments, a sequence that accurately represents the Watson strand of an analyte DNA fragment is determined to have a mutation if the sequence differs from a reference sequence that lacks the mutation.

[0181] In some embodiments, a sequence that accurately represents the Crick strand of an analyte DNA fragment is determined to have a mutation. In some embodiments, a sequence that accurately represents the Crick strand of an analyte DNA fragment is determined to have a mutation if the sequence differs from a reference sequence that lacks the mutation.

[0182] In some embodiments, an analyte DNA fragment is determined to have a mutation if a sequence that exactly represents the Watson strand and a sequence that exactly represents the Crick strand contain the same mutation.

[0183] In some cases, the location of the molecular barcode in the paired-end sequencing reads of the amplified target region can be used to distinguish which strand of the double-stranded nucleic acid template the amplified target region originates from. For example, if the first paired-end sequencing read of the amplified target region indicates that the molecular barcode was read last, the amplified target region can be identified as originating from the sense strand of the nucleic acid template, and if the first paired-end sequencing read of the amplified target region indicates that the molecular barcode was read first, the amplified target region can be identified as originating from the antisense strand of the nucleic acid template. For example, if the second paired-end sequencing read of the amplified target region indicates that the molecular barcode was read first, the amplified target region can be identified as originating from the antisense strand of the nucleic acid template, and if the second paired-end sequencing read of the amplified target region indicates that the molecular barcode was read last, the amplified target region can be identified as originating from the sense strand of the nucleic acid template. In some cases, paired-end sequencing can be used to distinguish amplification products originating from the Watson strand from amplification products originating from the Crick strand.

[0184] After sequencing of a target region (e.g., a target region amplified as described herein), the sequencing reads can be aligned with a reference genome and classified by the molecular barcode present in each sequencing read.In some cases, the sequencing reads that contain the same molecular barcode and map to both the Watson strand and the Crick strand of a double-stranded nucleic acid template (e.g., both the Watson strand and the Crick strand of a target region) can be identified as having double-stranded support.For example, if the sequencing reads that indicate the presence of one or more mutations in a target region contain the same molecular barcode and map to both the Watson strand and the Crick strand of a target region, the mutation can be identified as having double-stranded support.

[0185] Amplification of nucleic acid fragments comprising molecular barcodes can be performed according to known techniques to generate a family of barcoded fragments. In some embodiments, polymerase chain reaction (PCR) can be used. In some embodiments, inverse PCR can be used. In some embodiments, rolling circle amplification can be used. Amplification of the fragments is typically performed using a primer complementary to the priming site that is attached to the fragment simultaneously with the molecular barcode. In some embodiments, the priming site is distal to the molecular barcode such that the amplification includes the molecular barcode.

[0186] In some embodiments, the amplification forms a family of fragments, with each member of the family sharing the same molecular barcode. In some embodiments, the diversity of molecular barcodes present on the adapter fragments greatly exceeds the diversity of the fragments, and thus each family is derived from a single nucleic acid fragment molecule. In some embodiments, the primers used for amplification can be chemically modified to make them more resistant to exonucleases. In some embodiments, the family members are sequenced and compared to identify any differences within the family. In some embodiments, the sequencing is performed on a massively parallel sequencing platform, many of which are commercially available. If the sequencing platform requires sequences for "grafting", i.e., attachment to the sequencing instrument, such sequences may be added during the addition of the molecular barcodes or may be added separately. The grafting sequences may be part of the molecular barcoded primer, the universal primer, the gene target specific primer, the amplification primer used to create the family, the sample barcoded primer, or may be separate. Redundant sequencing refers to the sequencing of multiple members of a single family.

[0187] In some embodiments, a threshold for identifying mutations in a nucleic acid fragment can be set. If a "mutation" appears in all members of a family, it originates from the nucleic acid fragment. If a mutation appears in less than all members, it may be an artifact introduced during analysis (e.g., during the amplification step). The threshold for calling a mutation can be set at, for example, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98% or 100%. In some embodiments, the threshold for calling a mutation is 95%, so if 95% of family members sharing the same barcode contain the mutation, the mutation is considered to be real and not an artifact. The threshold will be set based on the number of family members to be sequenced, as well as the specific purpose and situation.

[0188] In some embodiments, one or more sequencing methods can be used to evaluate the amplified DNA molecule and determine whether mutations are present on both strands of double-stranded DNA molecules. In some embodiments, sequencing reads can be used to evaluate the amplified DNA molecule for the presence or absence of one or more mutations, and sequencing reads can be used to determine whether mutations are present on both strands of double-stranded DNA molecules. Examples of sequencing methods that can be used to evaluate the amplified DNA molecule for the presence or absence of one or more mutations described herein include, but are not limited to, single-read sequencing, paired-end sequencing, NGS, and deep sequencing. In some embodiments, single-read sequencing includes sequencing that generates sequence reads over the entire length of the template. In some embodiments, sequencing includes paired-end sequencing. In some embodiments, sequencing is performed on a massively parallel sequencer. In some embodiments, the massively parallel sequencer is configured to determine sequence reads from both ends of the template polynucleotide.

[0189] In some embodiments, the method described herein includes: (g) classifying the first sequencing read according to the molecular barcode present on at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing read according to the molecular barcode present on at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family; (i) identifying the first feature of the tagged Watson strand and Crick strand in the first analyte DNA family; and (j) identifying the second feature of the matched Watson strand and Crick strand in the second analyte DNA family, thus identifying the first feature and the second feature present on at least one strand of the double-stranded DNA molecule. In some embodiments, the method includes identifying the first feature and the second feature present on both strands of the double-stranded DNA molecule.

[0190] Epigenetic Signatures – Methylation Analysis In some embodiments, the first and / or second feature can be an epigenetic feature, where the term "epigenetic feature" can refer to a heritable phenotypic change without a change in DNA sequence. In some embodiments, the epigenetic feature comprises a functionally relevant change to the genome without a change in nucleotide sequence. In some embodiments, the epigenetic feature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation. In some embodiments, the epigenetic feature is methylation. In some embodiments, the epigenetic feature is a methylation pattern. In some embodiments, the methylation pattern corresponds to the methylation pattern present in cells generated via clonal hematopoiesis of indeterminate origin. In some embodiments, the methylation pattern corresponds to the methylation pattern present in the tissue of origin. In some embodiments, the tissue of origin is anus, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus (Cypris et al., Front. Genet. 10:785 (2019), Liu et al., Ann Oncol.31(6):745-759 (2020)).

[0191] In some embodiments, the methods described herein can be used to detect methylation at CpG dinucleotides in one or both strands of a double-stranded DNA molecule (e.g., both strands simultaneously). In some embodiments, a population of DNA molecules is treated with bisulfite to convert cytosine bases in the DNA molecules to uracil bases, forming a population of converted DNA molecules. In some embodiments, an excess of target-specific amplification primers attached to molecular barcodes are used to attach molecular barcodes to both strands of the population of converted DNA molecules, forming a population of amplified, barcoded, converted DNA molecules. In some embodiments, the amplified, barcoded, converted DNA molecules are amplified in an amplification reaction to form a family of amplified, barcoded, converted DNA molecules, where the amplified, barcoded, converted DNA molecules that share the same molecular barcode form a family of DNA molecules. In some embodiments, a plurality of members of the family are subjected to a sequencing reaction to obtain the nucleotide sequence of both strands of the plurality of members of the family. In some embodiments, the nucleotide sequences of multiple members of a family are compared and families in which more than 90% of the members contain a selected methylated C at a CpG dinucleotide are identified. In some embodiments, the nucleotide sequences of two complementary strands of the amplified, barcoded, converted DNA molecule are compared and a methylated C at a CpG dinucleotide is identified in the two complementary strands.

[0192] In some embodiments, incubation of DNA fragments with sodium bisulfite at elevated temperature and low pH deaminates cytosine to form 5,6-dihydrocytosine-6-sulfonate. Exemplary methods of sodium bisulfite treatment for use in the methods disclosed herein are described in PCT / US2018 / 022664, which is incorporated herein by reference in its entirety. Subsequent hydrolytic deamination at high pH removes the sulfonate to yield uracil. Many modifications of this basic reaction have been described and are primarily used to distinguish between cytosine and 5-methylcytosine (5-mC), the latter of which is not susceptible to conversion with bisulfite. In addition to converting C to U, bisulfite treatment can denature DNA and degrade it. This degradation is not limited to standard applications of bisulfite treatment, but is important for applications involving mutation detection in clinical samples that are already degraded prior to conversion. In some embodiments, sequencing of these products revealed that, on average, greater than 99.8% of C bases were converted to T bases on both strands (except for C bases at 5'-CpG sites, which may be resistant to conversion with bisulfite because they are methylated or hydroxymethylated). EXAMPLES

[0193] The present disclosure is further described in the following examples, which do not limit the scope of the disclosure described in the claims.

[0194] Example 1 – Bisulfite treatment, library preparation, and sequencing The EZ DNA Methylation Kit (Zymo Research, Cat. No. D5001) was selected to bisulfite-convert and desulfonate DNA samples according to the manufacturer's recommended protocol. DNA was denatured in dilute M-Dilution buffer at 37°C for 15 min, then bisulfite-converted at 50°C in the dark for 16 h, followed by 10 min on ice. After one wash with M-Wash buffer, samples were desulfonated at room temperature for 15 min. Samples were washed twice with M-Wash buffer, then eluted in 15 μL of Elution Buffer and stored at -20°C. Next-generation sequencing libraries were prepared using the Accel-NGS Methyl-Seq DNA Library Kit (Swift Bioscience, Cat. No. 30024) with 9 PCR cycles in the indexing step. Each library was paired-end sequenced to 150 bp in a single lane on an Illumina HiSeq 4000 instrument. Illumina CASAVA Chastity filtered reads were used for subsequent analysis. FASTQ files from bisulfite sequencing are available from the European Genome-phenome Archive.

[0195] Example 2 –DNA sequencing data analysis Illumina adaptors and bases with a quality score below 25 were trimmed from the beginning and end of each read using Trimmomatic. To allow whole-genome alignment to hg19, 14 bp UIDs and 13 bp constant sequences were trimmed from the beginning of reads 1 and 2 using Trimmomatic v0.38. Each paired-end read was aligned to the bisulfite-converted hg19 genome using BSMAP, and the average methylation of each CpG was calculated using the methratio.py script in BSMAP.

[0196] Example 3 – Identification of methylation markers for plasma cfDNA tissue deconvolution The average contribution of 12 tissue types (liver, lung, colon, small intestine, pancreas, adrenal gland, esophagus, heart, brain, T cells, B cells, and neutrophils) to the total cfDNA pool was determined using 5,653 differentially methylated 500 bp regions. To identify methylation markers for plasma DNA tissue mapping, bisulfite sequencing data of 12 human tissues was analyzed. Whole-genome bisulfite sequencing data for liver, lung, esophagus, heart, pancreas, colon, small intestine, adrenal gland, brain, and T cells were retrieved from the Human Epigenome Atlas (www.genboree.org / epigenomeatlas / index.rhtml) at Baylor College of Medicine.

[0197] All CpG islands (CGIs) and CpG shores on autosomes were assessed for potential inclusion in the methylation marker set. To minimize potential variations in methylation levels associated with sex-related chromosome dosage differences in the source data, CGIs and CpG shores on sex chromosomes were not used. CGIs were downloaded from the University of California, Santa Cruz (UCSC) database (genome.ucsc.edu / , 27,048 CGIs vs. human genome) and CpG shores were defined as 2 kb flanking windows of CGIs. CGIs and CpG shores were then subdivided into non-overlapping 500 bp units, and each unit was considered as a potential methylation marker.

[0198] The methylation densities of all potential marker loci (i.e., the percentage of CpGs that are methylated within 500 bp units) were compared across the 12 tissue types. Using the methylation profiles of the 12 tissue types, two types of methylation markers were identified. Type I markers refer to any genomic loci that have a methylation density that is plus or minus 3 SD in a tissue compared to the mean level of the 12 tissue types. Type II markers are genomic loci that show highly variable methylation densities across the 12 tissue types. A locus is considered highly variable if (A) the methylation density in the most hypermethylated tissue is at least 20% higher than that in the most hypomethylated tissue; (B) the SD when the methylation density across the 13 tissue types is divided by the group mean methylation density (i.e., the coefficient of variation) is at least 0.25. To reduce the number of potentially redundant markers, only one marker is selected in one contiguous block of two CpG shores flanking one CGI.

[0199] Example 4 – Plasma cfDNA tissue deconvolution The mathematical relationship between the methylation density of various methylation markers in plasma and the corresponding methylation markers in various tissues is It can be represented as TIFF2024529674000002.tif11128, where TIFF2024529674000003.tif5128 represents the methylation density of methylation biomarker i in plasma; p k represents the proportional contribution of tissue k to plasma; and MD ik represents the methylation density of methylation biomarker i in tissue k. The goal of the deconvolution process is to determine, for each member of the tissue panel, the proportional contribution of tissue k to plasma, i.e., p kThe aim of the study was to determine the p The quadratic programming was used to solve the simultaneous equations. For each methylation marker in the combined list of type I and type II markers (5,653 markers in total), a matrix containing the tissue panel and their corresponding methylation density was compiled. The program used p k A range of values ​​was entered to determine the expected plasma DNA methylation density for each marker. k The range of values ​​is 100% for the sum of the contributions of all candidate tissues, i.e., liver, neutrophils, and lymphocytes, to plasma DNA, and 100% for all p k These three tissue types were chosen because each could be validated by one or more clinical scenarios: liver in liver transplants and HCC, and blood cells in bone marrow transplants and lymphoma cases. The program then selected the p-value that yielded the expected methylation density across markers that most closely resembled the data obtained from plasma DNA bisulfite sequencing. k A set of values ​​was identified.

[0200] The sum of the contributions from T and B cells was considered as the contribution from lymphocytes, and the sum of the contributions from leukocytes was considered as the contribution from lymphocytes and neutrophils.

[0201] Example 5 – Generation of sequencing libraries Libraries were prepared as described herein (Figure 3). Custom 3' and 5' adapters containing 5-methylcytosine rather than unmethylated cytosine were used during ligation, followed by one, two or three cycles of linear PCR with a single deoxyuridine-containing biotinylated primer targeting the 3' adapter. The amplified products were bound to streptavidin beads and then denatured to separate the biotinylated amplified strand ("A" pool) from the original, non-biotinylated strand ("M" pool), which still retained the methylation pattern. The supernatant after heat denaturation, containing the "M" pool, was bisulfite converted and amplified. After capture-based or PCR-based enrichment and / or whole genome bisulfite sequencing, the "M" pool was analyzed for methylation changes. For simultaneous preparation of the "A" pool, the strands bound to streptavidin beads were released after treatment with USER (Uracil-Specific Excision Reagent) enzyme, which consists of a mixture of uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII that targets deoxyuridine bases embedded within the 5' ends of the strands. The released strands are amplified and sequenced for analysis of somatic mutations (e.g., Cohen et al. Nat Biotechnol. (2021) 39(10):1220-1227, this publication is incorporated herein by reference) (Figures 4-5).

Claims

1. 1. A method for identifying the genetic and epigenetic characteristics of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule, comprising the steps of: (a) attaching an adapter fragment to each end of a double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is the reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand; (c) subjecting the amplified products to denaturing conditions; (d) separately recovering the adapted Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson strands and Crick strands, and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (f) generating a second population of analyte DNA fragments from the adapted Watson and Crick strands, and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on said at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing reads according to the molecular barcode present on said at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family; (i) identifying the genetic characteristics of the tagged Watson and Crick strands in a first analyte DNA family; and (j) identifying the matched Watson and Crick strand epigenetic signatures in a second analyte DNA family. and thus identifying genetic and epigenetic features present on at least one strand of said double-stranded DNA molecule.

2. 10. The method of claim 1, wherein the adapter fragment further comprises a sample barcode.

3. 3. The method of claim 1 or 2, wherein said molecular barcode comprises an endogenous barcode, an exogenous barcode, or both.

4. 3. The method of claim 1 or 2, wherein the copying step (b) comprises performing one, two, or three rounds of linear extension of the adapted double-stranded DNA molecule.

5. 3. The method of claim 1 or 2, wherein the tagged primer is a uracil-containing biotinylated primer, and the tagged Watson strand and Crick strand are generated from the uracil-containing biotinylated primer.

6. 6. The method of claim 5, wherein the recovering step (d) comprises contacting the tagged Watson and Crick strands with streptavidin-functionalized beads, and wherein the tagged Watson and Crick strands bind to the streptavidin-functionalized beads.

7. 7. The method of claim 6, wherein the recovered adapted Watson and Crick strands that are not bound to the streptavidin-functionalized beads are treated with bisulfite to convert cytosine bases to uracil bases to generate a second population of analyte DNA fragments comprising a population of converted DNA molecules.

8. 3. The method of claim 1 or 2, wherein the denaturing conditions comprise NaOH denaturation.

9. 3. The method of claim 1 or 2, wherein the denaturing conditions comprise thermal denaturation, chemical denaturation, or a combination thereof.

10. 3. The method of claim 1 or 2, wherein said generating steps (e) and (f) are carried out under PCR conditions.

11. 3. The method of claim 1 or 2, wherein the genetic characteristic is a mutation.

12. 12. The method of claim 11, wherein the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and combinations thereof.

13. 3. The method of claim 1 or 2, wherein the epigenetic feature is methylation.

14. 14. The method of claim 13, wherein the epigenetic signature is a methylation pattern.

15. 15. The method of claim 14, wherein the methylation pattern corresponds to a methylation pattern present in cells generated through clonal hematopoiesis of indeterminate origin.

16. 16. The method of claim 15, wherein the methylation pattern corresponds to the methylation pattern present in the tissue of origin.

17. 17. The method of claim 16, wherein the tissue of origin is anus, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus.

18. 3. The method of claim 1 or 2, wherein the epigenetic signature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation.

19. 3. The method of claim 1 or 2, wherein the genetic and epigenetic characteristics of a double-stranded DNA molecule in a population of double-stranded DNA molecules are identified by assaying both strands of the double-stranded DNA molecule.

20. 1. A method for identifying a first characteristic and a second characteristic of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying at least one strand of the double-stranded DNA molecule, the method comprising the steps of: (a) attaching an adapter fragment to each end of a double-stranded DNA molecule to generate an adapted double-stranded DNA molecule, the adapted double-stranded DNA molecule comprising an adapted Watson strand and an adapted Crick strand, the adapter fragment comprising a molecular barcode, a primer sequence, and an adapter sequence, and the molecular barcode of the adapted Watson strand is the reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, the copying comprising: (i) contacting the adapted double-stranded DNA molecule with a tagged primer; and (ii) performing a round of linear extension of the adapted double-stranded DNA molecule to generate a tagged Watson strand and a tagged Crick strand; (c) subjecting the amplified products to denaturing conditions; (d) separately recovering the adapted Watson and Crick strands and the tagged Watson and Crick strands; (e) generating a first population of analyte DNA fragments from the tagged Watson strands and Crick strands, and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (f) generating a second population of analyte DNA fragments from the adapted Watson and Crick strands, and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (g) classifying the first sequencing read according to the molecular barcode present on said at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (h) classifying the second sequencing reads according to the molecular barcode present on said at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family; (i) identifying a first characteristic of the tagged Watson strand and Crick strand in a first analyte DNA family; and (j) identifying a second feature of the matched Watson and Crick strands in a second analyte DNA family. thus identifying a first feature and a second feature present on at least one strand of the double-stranded DNA molecule.

21. 21. The method of claim 20, wherein the adapter fragment further comprises a sample barcode.

22. 22. The method of claim 20 or 21, wherein said molecular barcode comprises an endogenous barcode, an exogenous barcode, or both.

23. 22. The method of claim 20 or 21, wherein the copying step (b) comprises performing one, two, or three rounds of linear extension of the adapted double-stranded DNA molecule.

24. 22. The method of claim 20 or 21, wherein the tagged primer is a uracil-containing biotinylated primer, and the tagged Watson strand and Crick strand are generated from the uracil-containing biotinylated primer.

25. 25. The method of claim 24, wherein the recovering step (d) comprises contacting the first single-stranded DNA fragment with streptavidin-functionalized beads, and wherein the first single-stranded DNA fragment binds to the streptavidin-functionalized beads.

26. 22. The method of claim 20 or 21, wherein the denaturing conditions comprise NaOH denaturation.

27. 22. The method of claim 20 or 21, wherein the denaturing conditions comprise thermal denaturation, chemical denaturation, or a combination thereof.

28. 22. The method of claim 20 or 21, wherein said generating steps (e) and (f) are carried out under PCR conditions.

29. 22. The method of claim 20 or 21, wherein said generating step utilizes whole genome PCR, whole genome bisulfite sequencing, or capture sequencing.

30. 22. The method of claim 20 or 21, wherein the first characteristic is a genetic or epigenetic characteristic.

31. 22. The method of claim 20 or 21, wherein the second feature is an epigenetic feature or an epigenetic feature.

32. 22. The method of claim 20 or 21, wherein the first characteristic and the second characteristic are both genetic characteristics.

33. 22. The method of claim 20 or 21, wherein the first feature and the second feature are both epigenetic features.

34. 31. The method of claim 30, wherein the genetic characteristic is a mutation.

35. 35. The method of claim 34, wherein the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and combinations thereof.

36. 31. The method of claim 30, wherein identifying the genetic characteristic comprises mutation analysis, aneuploidy analysis, or fragmentomics.

37. 31. The method of claim 30, wherein the epigenetic signature is methylation.

38. 31. The method of claim 30, wherein the epigenetic signature is a methylation pattern.

39. 39. The method of claim 38, wherein the methylation pattern corresponds to a methylation pattern present in cells generated through clonal hematopoiesis of indeterminate origin.

40. 40. The method of claim 39, wherein the methylation pattern corresponds to the methylation pattern present in the tissue of origin.

41. 41. The method of claim 40, wherein the tissue of origin is anus, bladder / urothelium, breast, cervix, colon / rectum, head and neck, kidney, liver / bile duct, lung, lymphoid neoplasm, melanoma, myeloid neoplasm, ovary, pancreas / gallbladder, prostate, thyroid, upper GI, or uterus.

42. 31. The method of claim 30, wherein the epigenetic signature is hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation.

43. 22. The method of claim 20 or 21, wherein the first characteristic and the second characteristic of a double-stranded DNA molecule in a population of double-stranded DNA molecules are identified by assaying both strands of the double-stranded DNA molecule.