Methods and materials for evaluating nucleic acids

JP2026133491A5Pending Publication Date: 2026-08-26JOHNS HOPKINS UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026080245
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-14
Filing Date
2026-05-12
Publication Date
2026-08-26

AI Technical Summary

Technical Problem

Conventional next-generation sequencing (NGS) methods suffer from high sequencing error rates, making it difficult to reliably detect rare mutations, especially in limited DNA samples such as cell-free plasma DNA, and existing molecular barcoding techniques face challenges in converting template molecules into duplex molecules with the same barcode on each strand, leading to low duplex recovery and inefficiency in target region enrichment.

Method used

A method for barcoding both strands of a template identically and enriching each strand using PCR-based techniques without hybridization capture, involving the use of duplex molecular barcodes and strand-specific anchor PCR to generate single-stranded libraries, which are then sequenced to detect mutations present in both strands.

Benefits of technology

This approach allows for highly reliable identification of rare mutations with minimal DNA damage and artifacts, enabling accurate detection of mutations in both strands of a double-stranded nucleic acid, particularly in limited DNA samples, at an affordable cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026133491000001
    Figure 2026133491000001
  • Figure 2026133491000002
    Figure 2026133491000002
  • Figure 2026133491000003
    Figure 2026133491000003
Patent Text Reader

Abstract

This invention provides systems, kits, compositions, and methods for the preparation of sequencing libraries and sequencing workflows. [Solution] A method for detecting one or more mutations present in both strands of a double-stranded nucleic acid includes the steps of: generating a duplex sequencing library having a duplex molecular barcode at each end (e.g., the 5' end and the 3' end) of each nucleic acid in the library; generating a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences from the duplex sequencing library; and detecting the presence of one or more mutations present in both strands of the double-stranded nucleic acid in each single-stranded library.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Provisional Application No. 62 / 977,066, filed on 14 February 2020, which is incorporated herein by reference in its entirety.

[0002] Statement regarding federal government funding This invention was made with government support under grants CA062924, CA152753, and CA230691 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0003] Technical field This invention relates to the field of nucleotide sequencing. In particular, this invention relates to the preparation of sequencing libraries for mutation identification and sequencing workflows. [Background technology]

[0004] Background information Identifying rare mutations is useful not only in the field of basic biology but also in improving the clinical management of patients. Applications include infectious diseases, immunological profiling, paleogenetics, forensic medicine, aging, non-invasive prenatal testing, and cancer. Next-generation sequencing (NGS) technology is theoretically suitable for this application, and a variety of NGS methods exist for detecting rare mutations. However, conventional NGS methods suffer from a high sequencing error rate, making them unreliable for confidently detecting mutations, especially those present at low frequencies in the original sample.

[0005] The use of molecular barcodes to tag the original template molecule was planned to overcome various obstacles in rare mutations. Molecular barcoding allows for redundant sequencing of PCR-generated offspring of each tagged molecule, making sequencing errors easily recognizable (Kinde et al., Proc Natl Acad Sci USA 108:9530-9535 (2011) (Non-Patent Literature 1)). For example, if a given threshold of offspring of a barcoded template molecule contains the same mutation, the mutation is considered genuine ("supermutant"). If offspring below a given threshold contain the mutation of interest, the mutation is considered an artifact. Two types of molecular barcodes are described: exogenous and endogenous. Exogenous barcodes (also referred to herein as exogenous UIDs) contain pre-specified or random nucleotides and are added during library preparation or PCR. Endogenous barcodes (also referred to herein as endogenous UIDs) are formed by the sequences at the 5' and 3' ends of the template DNA fragment to be assayed (e.g., DNA present in a cell-free liquid biological sample or fragments generated by random shearing of the fragment). Such barcodes have proven useful for tracing amplicons back to their original start template, making molecules countable, and improving the identification of true mutations in clinically relevant samples.

[0006] To enable duplex sequencing, a fork-type adapter for paired-end sequencing has been developed, which allows each of the two strands of the original DNA duplex (Watson and Crick) to be identified by the 5'→3' orientation revealed during sequencing. Duplex sequencing reduces sequencing errors because if a mutation is accidentally generated during library preparation or sequencing, it is extremely unlikely that both strands of DNA will contain the same mutation.

[0007] However, there are many problems that limit the scope and clinical applicability of molecular barcoding. For example, it is difficult to convert most of the initial template molecules into duplex molecules with the same barcode on each strand (Schmitt et al., Proc Natl Acad Sci USA 109:14508-14513 (2012) (Non-Patent Literature 2); Schmitt et al., Nat Methods 12:423-425 (2015) (Non-Patent Literature 3); and Newman et al., Nat Biotechnol 34:547-555 (2016) (Non-Patent Literature 4)). This challenge is particularly problematic when the initial amount of DNA is limited (e.g., <33 ng), as is typically seen in cell-free plasma DNA used for liquid biopsies.

[0008] The preparation of target-oriented sequencing libraries typically involves adapter binding to a sequencing template, library amplification, and hybridization capture to enrich the library with respect to the target of interest. While hybridization capture is effective for enriching large regions of interest, it is not scalable for smaller target regions (Springer et al., Elife 7:doi:10.7554 / eLife.32143 (2018) (Non-Patent Literature 5)) and exhibits low duplex recovery (Wang et al., Proc Natl Acad Sci USA 112:9704-9709 (2015) (Non-Patent Literature 6); and Wang et al., Elife 5:doi:10.7554 / eLife.15175 (2016) (Non-Patent Literature 7)). Although these limitations can be partially overcome with sequential capture rounds, even with such improvements, duplex recovery rates are typically around 1%. While CRISPR-DS can achieve recovery of up to 15%, it cannot be applied to cell-free DNA. Capture-based methods are suboptimal when the target region is very small (e.g., one or a few locations in the genome of particular interest, such as those required for disease monitoring in plasma) or when the amount of available DNA is limited (e.g., <33 ng, often found in plasma).

[0009] Therefore, there is a need to improve sequencing library preparation and workflows to enable the accurate identification of mutations, such as rare mutations, in clinically relevant samples, including liquid biopsy samples. [Prior art documents] [Non-patent literature]

[0010] [Non-Patent Document 1] Kinde et al., Proc Natl Acad Sci USA 108:9530-9535 (2011) [Non-Patent Document 2] Schmitt et al., Proc Natl Acad Sci U S A 109:14508-14513 (2012) [Non-Patent Document 3] Schmitt et al., Nat Methods 12:423-425 (2015) [Non-Patent Document 4] Newman et al., Nat Biotechnol 34:547-555 (2016) [Non-Patent Document 5] Springer et al., Elife 7:doi:10.7554 / eLife.32143 (2018) [Non-Patent Document 6] Wang et al., Proc Natl Acad Sci U S A 112:9704-9709 (2015) [Non-Patent Document 7] Wang et al., Elife 5:doi:10.7554 / eLife.15175 (2016) [Summary of the Invention]

[0011] Summary This document provides methods and materials for addressing these challenges by providing a method for barcoding both strands of a template identically and by providing a method for PCR-based enrichment of each strand that does not require hybridization capture.

[0012] This document relates to methods and materials that can be used to detect the presence of one or more mutations present in both strands of a double-stranded nucleic acid (e.g., DNA). In some cases, a method for detecting one or more mutations present in both strands of a double-stranded nucleic acid may include the steps of: generating a duplex sequencing library having a duplex molecular barcode at each end (e.g., the 5' end and the 3' end) of each nucleic acid in the library; generating a library of single-stranded Watson-derived sequences and a library of single-stranded Crick-derived sequences from the duplex sequencing library; and detecting the presence of one or more mutations present in both strands of the double-stranded nucleic acid in each single-stranded library.

[0013] As demonstrated herein, single-stranded DNA libraries corresponding to the Watson strand of a double-stranded nucleic acid template and single-stranded DNA libraries corresponding to the Click strand of a double-stranded nucleic acid template can be generated from a sequencing library incorporating duplex molecular barcodes, and each single-stranded DNA library can be enriched with target regions using a strand-specific anchor PCR method, and the target regions can be sequenced to detect the presence or absence of one or more mutations within the target region of the nucleic acid. For example, the methods and materials described herein can be used to detect the presence of one or more mutations present on both strands of a double-stranded nucleic acid. S equence A scertainment F ree of Er rors Sv uencing SThis can be named a ystem (SaferSeqS), which can include, for example, in-situ generation of double-stranded molecular barcodes (see, e.g., Figure 22a), target enrichment by anchored PCR (see, e.g., Figure 22b), and library construction by in silico reconstruction of template molecules (see, e.g., Figure 22c). True mutations present in the original start template can be identified by finding changes in both strands of the same initial nucleic acid molecule.

[0014] The ability to detect one or more mutations present in both strands of a double-stranded nucleic acid (e.g., a true somatic mutation) provides a unique and unrealized opportunity to evaluate multiple mutations simultaneously, at an affordable cost, accurately and efficiently. Using the methods and materials described herein (e.g., the SaferSeqS method) to detect one or more mutations present in both strands of a double-stranded nucleic acid allows for highly reliable identification of rare mutations while minimizing the amount of DNA damage, the amount of PCR performed, and / or the number of DNA damage artifacts. It should be noted that the terms “Watson strand” and “Crick strand” are used simply to distinguish between the two strands of a double-stranded start nucleic acid sequence. Either strand may be referred to as “Watson” or “Crick,” in which case the other strand will be referred to by the other name.

[0015] In some cases, a) i) Multiple dephosphorylated and terminally blunted double-stranded DNA fragments, each comprising a Watson strand and a Crick strand; ii) Multiple adapters, each containing A) a barcode and B) a universal 3' adapter array in the 5'→3' direction; and iii) Ligauze The steps include forming a reaction mixture containing; b)i) The adapter is ligated to the 3' ends of the Watson and Crick chains, and ii) The adapter is not ligated to the 5' ends of either the Watson or Crick chains. The reaction mixture is incubated in a manner that generates a double-stranded ligation product. A method including the above is provided herein.

[0016] In a particular embodiment, each of the multiple adapters includes a unique barcode. In a further embodiment, each of the double-stranded ligation products includes a Watson strand having only one barcode and a Crick strand having only one barcode different from the barcode on the Watson strand. In a further embodiment, the method further includes the step of c) sequencing at least a portion of the double-stranded ligation product.

[0017] In a particular embodiment, a) a step of attaching a partial double-stranded 3' adapter (3'PDSA) to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, wherein the first strand of the 3'PDSA comprises a universal 3' adapter sequence including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site in the 5'-3' direction, and the second strand of the 3'PDSA comprises (i) a segment complementary to the first segment, and optionally (ii) a 3' blocking group in the 5'→3' direction; b) a step of annealing the 5' adapter to the annealing site, wherein the 5' adapter comprises (i) a segment not complementary to the universal 3' adapter sequence and an R1 sequencing primer site in the 5'→3' direction. A method is provided herein that includes the steps of: (ii) a universal 5' adapter sequence including a Mer site, and (ii) a sequence complementary to an annealing site for the 5' adapter; (ii) extending the 5' adapter over an exogenous UID sequence and a first segment to generate complements for the exogenous UID sequence and a first segment; and (iii) covalently ligating the 3' end of the complement for the first segment to the 5' ends of the Watson and Crick strands of a double-stranded DNA fragment to generate a plurality of double-stranded DNA fragments ligated with the adapter.

[0018] In some embodiments, a) a step of attaching a partial double-stranded 3' adapter (3'PDSA) to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, wherein the first strand of the 3'PDSA comprises a universal 3' adapter sequence in the 5'-3' direction, including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the 3'PDSA comprises (i) a segment complementary to the first segment, and optionally (ii) a 3' blocking group, in the 5'→3' direction; b) a step of annealing the 5' adapter to the annealing site, wherein the 5' adapter comprises (i) a segment not complementary to the universal 3' adapter sequence, and an R1 sequencing primer site, in the 5'→3' direction; A method is provided herein that includes the steps of: (ii) a universal 5' adapter sequence including a rimer region, and (ii) a sequence complementary to an annealing region for the 5' adapter; c) extending the 5' adapter over an exogenous UID sequence to generate a complement to the exogenous UID sequence; and d) covalently ligating the 3' end of the complement to the exogenous UID sequence to the 5' end of a segment complementary to the first segments of the Watson strand and Crick strand of a double-stranded DNA fragment to generate a plurality of double-stranded DNA fragments ligated with the adapter.

[0019] In some embodiments, a) a step of attaching a partial double-stranded 3' adapter (3'PDSA) to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, wherein the first strand of the 3'PDSA comprises a universal 3' adapter sequence including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site in the 5'-3' direction, and the second strand of the 3'PDSA comprises (i) a segment complementary to the first segment, and optionally (ii) a 3' blocking group in the 5'→3' direction; b) a step of annealing the 5' adapter to the annealing site, wherein the 5' adapter comprises (i) a unit that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site in the 5'→3' direction. A method is provided herein that includes the steps of: (ii) a versal 5' adapter sequence, and (ii) a sequence complementary to the annealing site for the 5' adapter; (ii) extending the 5' adapter over an exogenous UID sequence and a first segment of a 3' PDSA to generate a complement of the exogenous UID sequence and a complement of the first segment of the 3' PDSA; and (iii) covalently ligating the 3' end of the complement of the first segment of the 3' PDSA to the 5' ends of the Watson strand and Crick strand of a double-stranded DNA fragment to generate a plurality of double-stranded DNA fragments ligated with the adapter.

[0020] In some cases, a) A population of partial double-stranded 3' adapters (3'PDSAs) configured to ligate to the 3' ends of both the Watson and Crick strands of a population of double-stranded DNA fragments, wherein the first strand of the 3'PDSA comprises a universal 3' adapter sequence including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site in the 5'-3' direction, and the second strand of the 3'PDSA comprises (i) a segment complementary to the first segment, and (ii) a 3' blocking group in the 5'→3' direction; and b) A group of 5' adapters configured to anneal to an annealing site, wherein the 5' adapters include, in the 5'→3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence that is complementary to the annealing site for the 3' adapter. Systems, kits, and compositions including the above are provided herein.

[0021] In further embodiments, the system, kit, and composition include: c) a population of double-stranded DNA fragments from a biological sample, and / or c) reagents for degrading the second strand of a 3'PDSA to produce a single-stranded 3' adapter (3'SSA); and / or c) a first primer complementary to the universal 3' adapter sequence and a second primer complementary to the complement of the universal 5' adapter sequence; and / or c) a sequencing system; and / or c) a Watson anchor primer complementary to the universal 3' adapter sequence, and d) a click anchor primer complementary to the complement of the universal 5' adapter sequence; and / or c) (i) universal 3' adapter (i) a first set of pairs of second Watson target-selective primers, each comprising (i) one or more first Watson target-selective primers containing a sequence complementary to a portion of the ter sequence, and (ii) one or more second Watson target-selective primers, each comprising a sequence complementary to a portion of the universal 5' adapter sequence, and / or c) a first set of pairs of second click target-selective primers, each comprising (i) one or more click target-selective primers containing the same target-selective sequence as the second Watson target-selective primer sequence.

[0022] In some embodiments, the method further comprises the step of amplifying a plurality of double-stranded DNA fragments ligated with an adapter using a first primer complementary to a universal 3' adapter sequence and a second primer complementary to a complement of a universal 5' adapter sequence, thereby generating an amplicon, wherein the amplicon comprises a plurality of double-stranded Watson templates and a plurality of double-stranded Crick templates. In certain embodiments, the method further comprises the step of selectively amplifying a double-stranded Watson template using a first set of Watson target-selective primer pairs, thereby producing a targeted Watson amplification product, the first set of Watson target-selective primer pairs comprising (i) a first Watson target-selective primer comprising a sequence complementary to a portion of the universal 3' adapter sequence, and (ii) a second Watson target-selective primer comprising a target-selective sequence. In a further embodiment, the method further comprises the step of selectively amplifying a double-stranded click template using a first set of click-target-selective primer pairs, thereby producing a targeted click amplification product, the first set of click-target-selective primer pairs comprising (i) a first click-target-selective primer containing a sequence complementary to a complement of a portion of the universal 5' adapter sequence, and (ii) a second click-target-selective primer containing the same target-selective sequence as a second Watson-target-selective primer sequence. In a particular embodiment, a population of double-stranded DNA fragments is incubated with a mixture of uracil-DNA glycosylase and DNA glycosylase-lyase endonuclease VIII before ligating with any adapter.

[0023] In some embodiments, the polymerase employed (e.g., for extending the 5' adapter sequence) has 5'→3' exonuclease activity (e.g., capable of digesting the second chain of 3'PDSA). In other embodiments, the polymerase employed (e.g., for extending the 5' adapter sequence) does not have 5'→3' exonuclease activity.

[0024] In other embodiments, the method further includes a step of removing the second strand of the 3'PDSA to generate a single-stranded 3' adapter (3'SSA). In other embodiments, the step of removing the second strand is after step b), or before step b), or during step b). In some embodiments, the step of removing the second strand of the 3'PDSA includes contacting the 3' duplex adapter with uracil-DNA glycosylase (UDG) to degrade the second strand. In further embodiments, the step of removing the second strand is achieved by a polymerase having exonuclease activity, in which the polymerase extends the 5' adapter over the exogenous UID sequence and the first segment.

[0025] In further embodiments, the method further includes the step of determining sequence reads of one or more amplicons. In other embodiments, the method further includes the step of assigning sequence reads to UID families, wherein each member of the UID family contains the same exogenous UID sequence. In specific embodiments, the method further includes the step of assigning sequence reads of each UID family to Watson subfamilies and Crick subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences. In other embodiments, the method further includes the step of identifying a sequence that accurately represents the Watson strand of the DNA fragment under analysis, if at least 50% (e.g., 50…75…95%) of the Watson subfamily contain a certain nucleotide sequence. In other embodiments, the method further includes the step of identifying a sequence that accurately represents the Crick strand of the DNA fragment under analysis, if at least 50% (e.g., 50…75…90%) of the Crick subfamily contain a certain nucleotide sequence.

[0026] In some embodiments, the method further includes identifying a mutation as accurately representing the Watson chain if the nucleotide sequence accurately representing the Watson chain differs from a reference sequence lacking a certain mutation in that sequence. In further embodiments, the method further includes identifying a mutation as accurately representing the Click chain if the nucleotide sequence accurately representing the Click chain differs from a reference sequence lacking a certain mutation in that sequence. In other embodiments, the method further includes identifying a mutation in the DNA fragment under analysis if the mutation in the nucleotide sequence accurately representing the Watson chain and the mutation in the nucleotide sequence accurately representing the Click chain are the same mutation. In some embodiments, each member of the UID family further includes the same endogenous UID sequence, where the endogenous UID sequence includes the ends of double-stranded DNA fragments from a population. In other embodiments, the double-stranded DNA fragment population has blunt ends.

[0027] A method for detecting the presence or absence of mutations in a target region of a double-stranded DNA template obtained from a mammalian sample and for determining whether the mutations are present on both strands of the double-stranded DNA template, comprising: A) generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; B) amplifying the double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment to generate an amplified duplex sequencing library, wherein the amplification includes contacting the double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment with a universal primer pair under whole-genome PCR conditions; C) optionally generating a Watson strand single-stranded DNA library from the amplified duplex sequencing library; D) optionally generating a click strand single-stranded DNA library from the amplified duplex sequencing library; E) hybridizing a first primer and a 3' duplex adapter capable of hybridizing to a target region. A) A step of amplifying a target region from a Watson strand DNA library (e.g., a single-stranded DNA library) using a primer pair consisting of a second primer capable of reducing the target region; F) A step of amplifying a target region from a click strand DNA library (e.g., a single-stranded DNA library) using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 5' adapter; G) A step of sequencing the amplified target region from a Watson strand DNA library (e.g., a single-stranded DNA library) (e.g., a DNA library (e.g., a single-stranded DNA library)) to generate sequencing reads and to detect the presence or absence of mutations in the Watson strand of the target region; H) A step of sequencing the amplified target region from a click strand DNA library (e.g., a single-stranded DNA library) (e.g., a DNA library (e.g., a single-stranded DNA library)) to generate sequencing reads and to detect the presence or absence of mutations in the click strand of the target region;Furthermore, the present invention provides a method comprising the step of I) grouping sequencing reads by molecular barcodes present in each sequencing read in order to determine whether mutations are present on both strands of a double-stranded DNA template. In some embodiments, the step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment comprises i) ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 3' duplex adapter comprises a first oligonucleotide comprising a) a 5' phosphate, a first molecular barcode, and a 3' oligonucleotide, and b) a second oligonucleotide annealed thereto comprising a degradable 3' blocking group, and the 3' oligonucleotide and the second oligonucleotide ii) the rheotide sequences are complementary; ii) a degradable 3' blocking group is degraded; iii) a 5' adapter is ligated to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 5' duplex adapter contains an oligonucleotide containing a second molecular barcode, the second barcode being different from the first molecular barcode, and the 5' adapter is ligated to the double-stranded DNA fragment upstream of the first molecular barcode, leaving a single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter;and iv) filling the gap of single-stranded nucleic acid between the 5' end of a double-stranded DNA fragment and a 5' adapter in order to generate a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment. In some embodiments, the step of generating a Watson strand DNA library (e.g., a single-stranded DNA library) from an amplified duplex sequencing library includes: i) amplifying a first aliquot of the amplified duplex sequencing library using a primer pair consisting of a first primer and a second primer to generate a double-stranded amplification product having tagged Watson strands, wherein the first primer is capable of hybridizing to Watson strands and the first primer contains a tag; ii) denaturing the double-stranded amplification product having tagged Watson strands to generate single-stranded tagged Watson strands and single-stranded click strands; and iii) recovering the single-stranded tagged Watson strands to generate a Watson strand DNA library (e.g., a single-stranded DNA library) from the amplified duplex sequencing library.

[0028] In some embodiments, the double-stranded DNA template is obtained from a mammalian sample. The step of generating a click-strand DNA library (e.g., a single-stranded DNA library) from an amplified duplex sequencing library includes: i) amplifying a second aliquot of the amplified duplex sequencing library using a primer pair consisting of a first primer and a second primer to produce a double-stranded amplification product having a tagged click strand, wherein the first primer is capable of hybridizing to a click strand and the first primer contains a tag; ii) denaturing the double-stranded amplification product having a tagged click strand to produce a single-stranded tagged click strand and a single-stranded Watson strand; and iii) recovering the single-stranded tagged click strand to generate a click-strand DNA library (e.g., a single-stranded DNA library) from the amplified duplex sequencing library. In some embodiments, the mammal is a human.

[0029] In some embodiments, the method further includes the steps of: fragmenting double-stranded DNA to generate a double-stranded DNA fragment; dephosphorylating the 5' end of the double-stranded DNA fragment; and smoothing the ends of the double-stranded DNA fragment, prior to the step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; dephosphorylating the 5' end of the double-stranded DNA fragment; and smoothing the ends of the double-stranded DNA fragment. In some embodiments, ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template includes contacting the 3' duplex adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase. In some embodiments, the ligase is a T4 DNA ligase.

[0030] In some embodiments, degrading a degradable 3' blocking group involves contacting the 3' duplex adapter with uracil-DNA glycosylase (UDG). In some embodiments, ligating the 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template involves contacting the 5' adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase. In some embodiments, the ligase is Escherichia coli ligase.

[0031] In some embodiments, filling the single-stranded nucleic acid gap between the 5' end of a double-stranded DNA fragment and its 5' adapter involves bringing the 5' end of the double-stranded DNA fragment and its 5' adapter into contact in the presence of a polymerase and a dNTP. In some embodiments, the polymerase is a Taq polymerase.

[0032] In some embodiments, ligating a 5' adapter to each 5' end of a double-stranded DNA fragment and filling the gap between the 5' end of the double-stranded DNA fragment and the 5' adapter are performed simultaneously. In some embodiments, the step of amplifying a double-stranded DNA fragment, which has a duplex molecular barcode at each end of the double-stranded DNA fragment, to generate an amplified duplex sequencing library, includes contacting the double-stranded DNA fragment, which has a duplex molecular barcode at each end of the double-stranded DNA fragment, with a universal primer pair under PCR conditions. In some embodiments, amplification includes whole-genome PCR. In some embodiments, the tagged primers are biotinylated primers, in which case the biotinylated primers can generate biotinylated single-stranded Watson strands and biotinylated single-stranded Crick strands. In some embodiments, the denaturation step includes NaOH denaturation, thermal denaturation, or a combination of both.

[0033] In some embodiments, the recovery step includes contacting tagged Watson chains with streptavidin-functionalized beads and tagged click chains with streptavidin-functionalized beads. In some embodiments, the recovery step further includes denaturing untagged Watson chains and denaturing untagged Watson chains. In some embodiments, the recovery step further includes releasing biotinylated single-stranded Watson chains from streptavidin-functionalized beads and releasing biotinylated single-stranded click chains from streptavidin-functionalized beads. In some embodiments, the tagged primer is a phosphorylated primer, and the phosphorylated primer can generate phosphorylated single-stranded Watson chains and phosphorylated single-stranded click chains. In some embodiments, the denaturing step includes lambda exonuclease digestion.

[0034] In some embodiments, amplifying a target region from a Watson strand DNA library (e.g., a single-stranded DNA library) further includes a second amplification using a second primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; amplifying a target region from a Crick strand DNA library (e.g., a single-stranded DNA library) further includes a second amplification using a second primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 5' adapter. In some embodiments, the sequencing step includes paired-end sequencing.

[0035] A method for detecting the presence or absence of mutations in a target region of a double-stranded DNA template obtained from a mammalian sample and for determining whether the mutations are present on both strands of the double-stranded DNA template, comprising: A) generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; B) generating a Watson strand DNA library (e.g., a single-stranded DNA library) and a Crick strand DNA library (e.g., a single-stranded DNA library) from an amplified duplex sequencing library derived from the double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; C) amplifying the target region from a single-stranded Watson strand using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; D) hybridizing to the target region A method is also provided herein that includes the steps of: a) amplifying a target region from a single-stranded click strand using a primer pair consisting of a first primer capable of hybridizing to a 5' adapter and a second primer capable of hybridizing to a 5' adapter; E) sequencing the amplified target region from a Watson strand DNA library (e.g., a single-stranded DNA library) to generate sequencing reads and to detect the presence or absence of mutations in the Watson strand of the target region; F) sequencing the amplified target region from a click strand DNA library (e.g., a single-stranded DNA library) to generate sequencing reads and to detect the presence or absence of mutations in the click strand of the target region; and G) grouping the sequencing reads by molecular barcodes present in each sequencing read to determine whether the mutations are present on both strands of a double-stranded DNA template.

[0036] In some embodiments, the double-stranded DNA template is a genomic DNA sample, and the step of generating double-stranded DNA fragments, each having a duplex molecular barcode at each end of the double-stranded DNA fragment, is to i) ligate a 3' duplex adapter to each 3' end of the double-stranded DNA fragment obtained from the double-stranded DNA template, wherein the 3' duplex adapter comprises a first oligonucleotide comprising a) a 5' phosphate, a first molecular barcode, and a 3' oligonucleotide, and a second oligonucleotide annealed thereto comprising b) a degradable 3' blocking group, wherein the sequences of the 3' oligonucleotide and the second oligonucleotide are complementary; and ii) degrading the degradable 3' blocking group. iii) Ligate a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 5' duplex adapter comprises an oligonucleotide containing a second molecular barcode, the second barcode being different from the first molecular barcode, and the 5' adapter is ligated to the double-stranded DNA fragment upstream of the first molecular barcode, leaving a single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter; and iv) Fill the single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter in order to generate a double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment.

[0037] In some embodiments, the double-stranded DNA template is a cell-free DNA sample, and generating Watson strand DNA libraries (e.g., single-stranded DNA libraries) and Crick strand DNA libraries (e.g., single-stranded DNA libraries) from an amplified duplex sequencing library derived from the double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment is to i) amplify the double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment using a universal primer pair consisting of a first primer and a second primer to generate a double-stranded amplification product having biotinylated Watson strands, wherein the amplification is to amplify the double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment The method comprises: ii) contacting a lymer pair with whole-genome PCR conditions such that the first primer is capable of hybridizing to the Watson chain and the first primer is biotinylated; ii) contacting a double-stranded amplification product having biotinylated Watson chains with streptavidin-functionalized beads under conditions such that the biotinylated Watson chains bind to the streptavidin-functionalized beads; iii) denaturing the double-stranded amplification product having biotinylated Watson chains in order to retain the single-stranded biotinylated Watson chains bound to the streptavidin-functionalized beads and to release the single-stranded click chains; iv) collecting the single-stranded click chains; v) releasing the single-stranded biotinylated Watson chains from the streptavidin-functionalized beads; and vi) collecting the single-stranded biotinylated Watson chains.

[0038] In some embodiments, the double-stranded DNA template is obtained from a mammalian sample. In some embodiments, the mammal is a human.

[0039] In some embodiments, the method further includes fragmenting double-stranded DNA to generate a double-stranded DNA fragment; dephosphorylating the 5' end of the double-stranded DNA fragment; and blunting the ends of the double-stranded DNA fragment, prior to the step of generating the double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment.

[0040] In some embodiments, ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template includes contacting the 3' duplex adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase. In some embodiments, the ligase is T4 DNA ligase. In some embodiments, a degradable 3' blocking group includes contacting the 3' duplex adapter with uracil-DNA glycosylase (UDG). In some embodiments, ligating a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template includes contacting the 5' adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase. In some embodiments, the ligase is E. coli ligase.

[0041] In some embodiments, filling the single-stranded nucleic acid gap between the 5' ends of a double-stranded DNA fragment and the 5' adapter involves bringing the 5' ends of the double-stranded DNA fragment and the 5' adapter into contact in the presence of a polymerase and dNTPs. In some embodiments, the polymerase is Taq-B polymerase. In some embodiments, ligating the 5' adapters to each 5' end of the double-stranded DNA fragment and filling the gap between the 5' ends of the double-stranded DNA fragment and the 5' adapter are performed simultaneously.

[0042] In some embodiments, amplifying a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment comprises contacting the double-stranded DNA fragment containing the duplex molecular barcode at each end of the double-stranded DNA fragment with a primer pair under PCR conditions. In some embodiments, the amplification step comprises whole-genome PCR. In some embodiments, amplifying a target region from a Watson-stranded DNA library (e.g., a single-stranded DNA library) further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; amplifying a target region from a Crick-stranded DNA library (e.g., a single-stranded DNA library) further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 5' adapter. In some embodiments, the sequencing step comprises paired-end sequencing or single-end sequencing.

[0043] A method for detecting the presence or absence of mutations in a target region of a double-stranded DNA template obtained from a mammalian sample and for determining whether a mutation is present on both strands of the double-stranded DNA template, comprising: A) generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; B) amplifying the double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment using a universal primer pair, comprising contacting the double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment with the primer pair under whole-genome PCR conditions; C) using a primer pair consisting of a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 3' duplex adapter, the Watson strand of the amplified double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment. A method is also provided herein that includes the steps of: a) amplifying a target region; D) amplifying a target region from the click strand of an amplified double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 5' adapter; E) sequencing the amplified target region from the Watson strand to generate sequencing reads and to detect the presence or absence of a mutation in the Watson strand of the target region; F) sequencing the amplified target region from the click strand to generate sequencing reads and to detect the presence or absence of a mutation in the click strand of the target region; and G) grouping the sequencing reads by the molecular barcode present in each sequencing read to determine whether the mutation is present on both strands of the double-stranded DNA template.

[0044] In some embodiments, the double-stranded DNA template is a genomic DNA sample, and the step of generating double-stranded DNA fragments, each having a duplex molecular barcode at each end of the double-stranded DNA fragment, is to i) ligate a 3' duplex adapter to each 3' end of the double-stranded DNA fragment obtained from the double-stranded DNA template, wherein the 3' duplex adapter comprises a first oligonucleotide comprising a) a 5' phosphate, a first molecular barcode, and a 3' oligonucleotide, and a second oligonucleotide annealed thereto comprising b) a degradable 3' blocking group, wherein the sequences of the 3' oligonucleotide and the second oligonucleotide are complementary; and ii) degrade the degradable 3' blocking group. iii) Ligate a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 5' duplex adapter comprises an oligonucleotide containing a second molecular barcode, the second molecular barcode being different from the first molecular barcode, and the 5' adapter is ligated to the double-stranded DNA fragment upstream of the first molecular barcode, leaving a single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter; and iv) Fill the single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter in order to generate a double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment. In some embodiments, the double-stranded DNA template is a cell-free DNA sample. In some embodiments, the double-stranded DNA template is a genomic DNA sample. In some embodiments, the mammal is a human.

[0045] In some embodiments, the method further includes the steps of: fragmenting double-stranded DNA to generate a double-stranded DNA fragment; dephosphorylating the 5' end of the double-stranded DNA fragment; and smoothing the ends of the double-stranded DNA fragment, prior to the step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment.

[0046] In some embodiments, ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template includes contacting the 3' duplex adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase. In some embodiments, the ligase is T4 DNA ligase. In some embodiments, degrading a degradable 3' blocking group includes contacting the 3' duplex adapter with uracil-DNA glycosylase (UDG). In some embodiments, ligating a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template includes contacting the 5' adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase. In some embodiments, the ligase is E. coli ligase.

[0047] In some embodiments, filling the single-stranded nucleic acid gap between the 5' end of a double-stranded DNA fragment and its 5' adapter involves bringing the 5' end of the double-stranded DNA fragment and its 5' adapter into contact in the presence of DNA polymerase and dNTPs. In some embodiments, the DNA polymerase is Taq-B polymerase.

[0048] In some embodiments, ligating a 5' adapter to each 5' end of a double-stranded DNA fragment and filling the gap between the 5' ends of the double-stranded DNA fragment and the 5' adapter are performed simultaneously. In some embodiments, amplifying a double-stranded DNA fragment having a duplex molecular barcode at each end of the fragment includes contacting the double-stranded DNA fragment containing the duplex molecular barcode at each end of the fragment with a primer pair under PCR conditions. In some embodiments, amplification includes whole-genome PCR. In some embodiments, amplifying a target region from a Watson strand DNA library (e.g., a single-stranded DNA library) further includes a second amplification using a second primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; amplifying a target region from a Crick strand DNA library (e.g., a single-stranded DNA library) further includes a second amplification using a second primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 5' adapter. In some embodiments, the sequencing step includes paired-end sequencing.

[0049] a. A step of attaching a partial double-stranded 3' adapter to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, wherein the first strand of the partial double-stranded 3' adapter comprises a universal 3' adapter sequence in the 5'-3' direction including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the partial double-stranded 3' adapter comprises (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction, and optionally the second strand is digestible; b. A step of annealing a 5' adapter to a 3' adapter via an annealing site, wherein the 5' adapter comprises (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence complementary to the annealing site for the 5' adapter, in the 5'→3' direction; c. A step in which a nick translation-like reaction is performed to extend the 5' adapter over the exogenous UID sequence of the 3' adapter (e.g., using DNA polymerase) and to covalently ligate the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment (e.g., using ligase); d. The first amplification step to amplify the double-stranded DNA fragment ligated with the adapter to produce an amplicon; e. The step of determining sequence reads of one or more amplicons of one or more double-stranded DNA fragments ligated with an adapter; f. A step in which a sequence read is assigned to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; g. The step of assigning sequence reads of each UID family to the Watson subfamily and Crick subfamily based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; h. A step of identifying a sequence that accurately represents the Watson strand of the DNA fragment being analyzed, given that a threshold percentage of members of the Watson subfamily contain a certain nucleotide sequence; i. The step of identifying a sequence that accurately represents the click strand of the DNA fragment being analyzed, given that a threshold percentage of members of the click subfamily contain a certain nucleotide sequence; j. The step of identifying a mutation when a nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the Watson chain but lacks a certain mutation in that sequence; k. The step of identifying a mutation when a nucleotide sequence that accurately represents a click chain differs from a reference sequence that accurately represents a click chain but lacks a certain mutation in that sequence; and l. The step of identifying mutations in the DNA fragment under analysis when the mutations in the nucleotide sequence that accurately represents the Watson chain and the mutations in the nucleotide sequence that accurately represents the Crick chain are the same mutation. Methods including the above are also provided herein.

[0050] In some embodiments, each member of the UID family further includes the same endogenous UID sequence, in which the endogenous UID sequence includes the end of a double-stranded DNA fragment from the population. In some embodiments, the endogenous UID sequence including the end of a double-stranded DNA fragment includes at least 8, 10, or 15 bases. In some embodiments, the exogenous UID sequence is specific to each double-stranded DNA fragment. In some embodiments, the exogenous UID sequence is not specific to each double-stranded DNA fragment. In some embodiments, each member of the UID family includes the same endogenous UID sequence and the same exogenous UID sequence. In some embodiments, step (d) includes 11 or fewer PCR amplification cycles. In some embodiments, step (d) includes 7 or fewer PCR amplification cycles. In some embodiments, step (d) includes 5 or fewer PCR amplification cycles. In some embodiments, step (d) includes at least 1 PCR amplification cycle.

[0051] In some embodiments, the amplicon is enriched with one or more target polynucleotides before determining the sequence read. In some embodiments, enrichment is a. Selective amplification of an amplicon of a Watson chain containing a target polynucleotide sequence using a first set of Watson target-selective primer pairs, thereby producing a target Watson amplification product, wherein the first set of Watson target-selective primer pairs comprises (i) a first Watson target-selective primer containing a sequence complementary to a portion of a universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is optionally an R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a second Watson target-selective primer containing a target-selective sequence; and b. Selectively amplifying an amplicon of a click chain containing the same target polynucleotide sequence using a first set of click target-selective primer pairs, wherein the first set of click target-selective primer pairs includes (i) a first click target-selective primer containing a sequence complementary to a portion of the universal 5' adapter sequence, optionally the portion of which is the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) a second click target-selective primer containing the same target-selective sequence as the second Watson target-selective primer sequence. Includes.

[0052] In some embodiments, the method further comprises the step of purifying the target Watson amplification product and the target click amplification product from a non-target polynucleotide. In some embodiments, the method further comprising the purification step comprises binding the target Watson amplification product and the target click amplification product to a solid support. In some embodiments, the first Watson target-selective primer and the first click target-selective primer comprise a first member of an affinity-binding pair, and the solid support comprises a second member of an affinity-binding pair. In some embodiments, the first member is biotin and the second member is streptavidin. In some embodiments, the solid support comprises beads, wells, membranes, tubes, columns, plates, Sepharose, magnetic beads, or tips. In some embodiments, the method further comprises the step of removing polynucleotides not bound to the solid support.

[0053] In some embodiments, the method is a. A step of further amplifying a target Watson amplification product using a second set of Watson target-selective primers to construct a member of a target Watson library, wherein the second set of Watson target-selective primers comprises (i) a third Watson target-selective primer comprising a sequence complementary to a portion of a universal 3' adapter sequence, optionally wherein the portion of the universal 3' adapter sequence is the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a fourth Watson target-selective primer comprising, in the 5'→3' direction, an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide; b. A step of further amplifying the target click amplification product using a second set of click target-selective primers to construct a member of the target click library, wherein the second set of click target-selective primers includes (i) a third click target-selective primer comprising a sequence complementary to a portion of the universal 5' adapter sequence, optionally wherein the portion of the universal 5' adapter sequence is the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) a fourth click target-selective primer comprising, in the 5'→3' direction, an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide as the fourth Watson target-selective primer. It also includes.

[0054] In some embodiments, the third Watson target-selective primer and the third click target-selective primer further include a sample barcode sequence. In some embodiments, the third Watson target-selective primer further includes a first transplant sequence that enables hybridization with the first transplant primer in a sequencer, and the third click target-selective primer further includes a second transplant sequence that enables hybridization with the second transplant primer in a sequencer. In some embodiments, the fourth Watson target-selective primer further includes a second transplant sequence, and the fourth click target-selective primer further includes the first transplant sequence. In some embodiments, the first transplant sequence is a P7 sequence, and the second transplant sequence is a P5 sequence. In some embodiments, the members of the target Watson library and the target click library correspond to at least 50% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, the members of the target Watson library and the target click library correspond to at least 70% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 80% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 90% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 50% of the total DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 70% of the total DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 80% of the total DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 90% of the total DNA fragment population.

[0055] a. A step of attaching an adapter to a population of double-stranded DNA fragments in a DNA sample to be analyzed, wherein the adapter comprises a double-stranded portion containing an exogenous UID and a fork-shaped portion containing (i) a single-stranded 3' adapter sequence containing an R2 sequencing primer site and (ii) a single-stranded 5' adapter sequence containing an R1 sequencing primer site; b. The initial amplification step to amplify the adapter-ligated double-stranded DNA fragment to produce an amplicon; c. A step of selectively amplifying an amplicon of a Watson chain containing a target polynucleotide sequence using a first set of Watson target-selective primer pairs, thereby producing a target Watson amplification product, wherein the first set of Watson target-selective primer pairs comprises (i) a first Watson target-selective primer containing a sequence complementary to a portion of a universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is optionally the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a second Watson target-selective primer containing a target-selective sequence; d. A step of selectively amplifying an amplicon of a click chain containing the same target polynucleotide sequence using a first set of click target selective primer pairs, wherein the first set of click target selective primer pairs comprises (ii) a first click target selective primer containing a sequence complementary to a portion of the universal 5' adapter sequence, wherein the portion of the universal 5' adapter sequence is optionally the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) a second click target selective primer containing the same target selective sequence as the second click target selective primer sequence. e. The step of determining the sequence reads of the target Watson amplification product and the target click amplification product; f. A step in which a sequence read is assigned to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; g. The step of assigning sequence reads of each UID family to the Watson subfamily and Crick subfamily based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; h. The step of identifying a sequence that accurately represents the Watson strand of the DNA fragment being analyzed, given that a threshold percentage of members of the Watson family contain a certain nucleotide sequence; i. The step of identifying a sequence that accurately represents the click strand of the DNA fragment under analysis if a threshold percentage of members of the click family contain a certain nucleotide sequence; and j. The step of identifying mutations in the DNA fragment under analysis when both the nucleotide sequence accurately representing the Watson chain and the nucleotide sequence accurately representing the Crick chain contain the same mutations. Methods including the above are also provided herein.

[0056] In some embodiments, the method further includes the step of purifying a targeted Watson amplification product and a targeted click amplification product from a non-target polynucleotide. In some embodiments, the method further includes the step of binding the targeted Watson amplification product and the targeted click amplification product to a solid support. In some embodiments, the first Watson target-selective primer and the first click target-selective primer include a first member of an affinity-binding pair, and the solid support includes a second member of an affinity-binding pair. In some embodiments, the first member is biotin and the second member is streptavidin. In some embodiments, the solid support includes beads, wells, membranes, tubes, columns, plates, Sepharose, magnetic beads, or tips. In some embodiments, the method further includes the step of removing polynucleotides not bound to the solid support.

[0057] In some embodiments, the method is a. A step of further amplifying a target Watson amplification product using a second set of Watson target-selective primers to construct a member of a target Watson library, wherein the second set of Watson target-selective primers comprises (i) a third Watson target-selective primer containing a sequence complementary to the R2 sequencing primer site of a universal 3' adapter sequence, and (ii) a fourth Watson target-selective primer containing, in the 5'→3' direction, an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide; b. A step of further amplifying the target click amplification product using a second set of click target-selective primers to construct a member of the target click library, wherein the second set of click target-selective primers includes (i) a third click target-selective primer containing a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence, and (ii) a fourth click target-selective primer containing, in the 5'→3' direction, an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide of the fourth Watson target-selective primer. It also includes.

[0058] In some embodiments, the third Watson target-selective primer and the third click target-selective primer further include a sample barcode sequence. In some embodiments, the third Watson target-selective primer further includes a first transplant sequence that enables hybridization with the first transplant primer in a sequencer, and the third click target-selective primer further includes a second transplant sequence that enables hybridization with the second transplant primer in a sequencer. In some embodiments, the fourth Watson target-selective primer further includes a second transplant sequence, and the fourth click target-selective primer further includes the first transplant sequence. In some embodiments, the first transplant sequence is a P7 sequence, and the second transplant sequence is a P5 sequence. In some embodiments, the binding step includes binding the adapter having an A tail to a population of double-stranded DNA fragments. In some embodiments, the binding step includes binding the adapter having an A tail to both ends of the DNA fragments in the population.

[0059] In some embodiments, the coupling step is a. A partial double-stranded 3' adapter is attached to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments, wherein the first strand of the partial double-stranded 3' adapter comprises a universal 3' adapter sequence in the 5'-3' direction including (i) a first segment, (ii) optionally an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the partial double-stranded 3' adapter comprises (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction, and optionally the second strand is gradable; and b. Annealing a 5' adapter to a 3' adapter via an annealing site, wherein the 5' adapter includes, in the 5'→3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence complementary to the annealing site for the 5' adapter; and c. Perform a nick translation-like reaction to extend the 5' adapter across the 3' adapter (e.g., using DNA polymerase) and covalently ligate the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment (e.g., using ligase). Includes.

[0060] In some embodiments, the UID sequence includes an endogenous UID sequence containing the end of a double-stranded DNA fragment from a population. In some embodiments, the endogenous UID sequence containing the end of a double-stranded DNA fragment contains at least 8, 10, or 15 bases. In some embodiments, the exogenous UID sequence is specific to each double-stranded DNA fragment. In some embodiments, the exogenous UID sequence is not specific to each double-stranded DNA fragment. In some embodiments, each member of the UID family contains the same endogenous UID sequence and the same exogenous UID sequence.

[0061] In some embodiments, amplifying a double-stranded DNA fragment ligated with an adapter to produce an amplicon involves PCR amplification of 11 cycles or less. In some embodiments, amplifying a double-stranded DNA fragment ligated with an adapter to produce an amplicon involves PCR amplification of 7 cycles or less. In some embodiments, amplifying a double-stranded DNA fragment ligated with an adapter to produce an amplicon involves PCR amplification of 5 cycles or less. In some embodiments, amplifying a double-stranded DNA fragment ligated with an adapter to produce an amplicon involves at least 1 cycle of PCR amplification. In some embodiments, members of the target Watson library and the target Crick library represent at least 50% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 70% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 80% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 90% of the target polynucleotides in the double-stranded DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 50% of the total DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 70% of the total DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 80% of the total DNA fragment population. In some embodiments, members of the target Watson library and the target Crick library represent at least 90% of the total DNA fragment population.

[0062] In some embodiments, sequence read determination allows sequencing of both ends of a template molecule. In some embodiments, sequencing of both ends of a template molecule includes paired-end sequencing. In some embodiments, sequence read determination includes single-read sequencing over the length of the template to generate sequence reads. In some embodiments, sequence read determination includes sequencing by a large-scale parallel sequencer. In some embodiments, the large-scale parallel sequencer is configured to determine sequence reads from both ends of a template polynucleotide. In some embodiments, the double-stranded DNA fragment population includes one or more fragments with a length of approximately 50–600 nt. In some embodiments, the double-stranded DNA fragment population includes one or more fragments with a length of less than 2000 nt, less than 1000 nt, less than 500 nt, less than 400 nt, less than 300 nt, or less than 250 nt.

[0063] In some embodiments, the methods provided herein further include the step of preparing a single-stranded (ss) DNA library corresponding to the sense and antisense strands of the amplicon after the initial amplification and before selective amplification. In some embodiments, the preparation of the ss DNA library is a. An amplification reaction is carried out using two primers, thereby producing an amplified product containing a chain containing the first member of the affinity-binding pair and a chain not containing the first member of the affinity-binding pair, wherein only one of the two primers contains the first member of the affinity-binding pair; b. Contacting the amplification product with a solid support, wherein the solid support contains a second member of the affinity bond; c. Denature the amplification product in order to separate the chain containing the first member of the affinity binding pair from the chain not containing the first member of the affinity binding pair; and d. Purify the isolated chain containing the first member of the affinity binding pair and the isolated chain not containing the first member of the affinity binding pair. Includes.

[0064] In some embodiments, the first member of the affinity binding pair is biotin, and the second member of the affinity binding pair is streptavidin. In some embodiments, the preparation of the ss DNA library is a. Dividing an amplicon into two amplification reactions, thereby producing an amplified product containing a phosphorylated chain and a non-phosphorylated chain, wherein each amplification reaction utilizes a forward primer and a reverse primer, and only one of the two primers is phosphorylated; b. Contact the amplified product with an exonuclease that selectively digests the chain containing the 5' phosphate. Includes.

[0065] In some cases, a. In the first amplification reaction, the forward primer is phosphorylated, while the reverse primer is not; b. In the second amplification reaction, the reverse primer is phosphorylated, while the forward primer is not phosphorylated.

[0066] In some embodiments, the exonuclease is a lambda exonuclease. In some embodiments, phosphorylation occurs at the 5' site.

[0067] In some embodiments, the initial amplification is a. Amplification using a primer pair to produce an amplified product comprising a chain containing the first member of the affinity-binding pair and a chain not containing the first member of the affinity-binding pair, wherein only one of the two primers in the primer pair contains the first member of the affinity-binding pair; b. Contacting the amplification product with a solid support, wherein the solid support contains a second member of the affinity bond; c. Denature the amplification product in order to separate the chain containing the first member of the affinity binding pair from the chain not containing the first member of the affinity binding pair; and d. Purify the isolated chain containing the first member of the affinity binding pair and the isolated chain not containing the first member of the affinity binding pair. Includes.

[0068] In some embodiments, the first member of the affinity-binding pair is biotin, and the second member of the affinity-binding pair is streptavidin. In some embodiments, the sequence reads of the UID family are assigned to the Watson subfamily when the exogenous UID sequence is downstream of the R2 sequence and upstream of the R1 sequence. In some embodiments, the sequence reads of the UID family are assigned to the Crick subfamily when the exogenous UID sequence is downstream of the R1 sequence and upstream of the R2 sequence. In some embodiments, the sequence reads of the UID family are assigned to the Watson subfamily when the exogenous UID sequence is more closely related to the R2 sequence and less closely related to the R1 sequence. In some embodiments, the sequence reads of the UID family are assigned to the Crick subfamily when the exogenous UID sequence is more closely related to the R1 sequence and less closely related to the R2 sequence. In some embodiments, a sequence read of the UID family is assigned to the Watson subfamily if the exogenous UID sequence is immediately downstream of the R2 sequence or is within 1-300, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides. In some embodiments, a sequence read of the UID family is assigned to the Crick subfamily if the exogenous UID sequence is immediately downstream of the R1 sequence or is within 1-300, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides.

[0069] In some embodiments, the double-stranded DNA fragment population is derived from a biological sample. In some embodiments, the biological sample is obtained from a subject. In some embodiments, the subject is a human subject. In some embodiments, the biological sample is a liquid sample. In some embodiments, the liquid sample is selected from whole blood, plasma, serum, sputum, urine, sweat, tears, ascites, semen, and bronchoalveolar lavage fluid. In some embodiments, the liquid sample is a cell-free sample or a sample that is essentially cell-free. In some embodiments, the biological sample is a solid biological sample. In some embodiments, the solid biological sample is a tumor sample.

[0070] In some embodiments, the identified mutation is present at a frequency of less than 0.1% in the double-stranded DNA fragment population. In some embodiments, the identified mutation is present at a frequency of 0.1% to 0.00001% in the double-stranded DNA fragment population. In some embodiments, the identified mutation is present at a frequency of 0.1% to 0.01% in the double-stranded DNA fragment population. In some embodiments, the step of determining sequence reads includes determining sequence reads from both the Watson and Crick strands of at least 50% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from both the Watson and Crick strands of at least 70% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from both the Watson and Crick strands of at least 80% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from at least 90% of both Watson and Crick strands of the double-stranded DNA fragment containing the target polynucleotide in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from at least 50% of both Watson and Crick strands of the double-stranded DNA fragment in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from at least 70% of both Watson and Crick strands of the double-stranded DNA fragment in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from at least 80% of both Watson and Crick strands of the double-stranded DNA fragment in the DNA sample to be analyzed. In some embodiments, the step of determining sequence reads includes determining sequence reads from at least 90% of both Watson and Crick strands of the double-stranded DNA fragment in the DNA sample to be analyzed.

[0071] In some embodiments, the error rate associated with identifying one or more mutations in the DNA fragment to be analyzed by the method according to any one of the claims is reduced to at least 1 / 2, 1 / 4, 1 / 5, 1 / 10, 1 / 20, 1 / 30, 1 / 40, 1 / 50, 1 / 60, 1 / 70, 1 / 80, 1 / 90, or 1 / 100 compared to alternative methods for identifying mutations that do not require the detection of mutations in both the Watson and Crick strands of the DNA fragment to be analyzed. In some embodiments, the alternative method includes standard molecular barcoding or standard PCR-based molecular barcoding. In some embodiments, the alternative method is a. A step of attaching an adapter to a population of double-stranded DNA fragments in a DNA sample to be analyzed, wherein the adapter contains an intrinsic exogenous UID; b. The initial amplification step to amplify the adapter-ligated double-stranded DNA fragment to produce an amplicon; c. The step of determining sequence reads of one or more amplicons of one or more double-stranded DNA fragments ligated with an adapter; d. A step in which a sequence read is assigned to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; e. If a threshold percentage of members of the UID family contain a certain nucleotide sequence, the step of identifying a sequence that accurately represents the DNA fragment to be analyzed; and f. The step of identifying a mutation when the sequence identified as accurately representing the DNA fragment under analysis differs from a reference sequence lacking a certain mutation in the DNA fragment under analysis. Includes.

[0072] In some embodiments, the error rate associated with identifying one or more mutations in the DNA fragment to be analyzed by the method according to any one of the claims is 1 × 10⁻⁶ -2 Below, 1 x 10 -3 Below, 1 x 10 -4 Below, 1 x 10 -5 Below, 1 x 10 -6 Below, 5 x 10 -6 The following, or 1 × 10 -7The following applies:

[0073] Also provided herein is a computer-readable medium comprising computer-executable instructions for analyzing sequence read data from a nucleic acid sample, wherein the data is generated by the method described in any one of the claims. In some embodiments, the computer-readable medium is a. Assigning sequence reads to a UID family (where each member of the UID family contains the same exogenous UID sequence); b. For assigning sequence reads of each UID family to Watson subfamilies and Crick subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; c. To accurately represent the Watson strand of the DNA fragment being analyzed and to identify the sequence if a threshold percentage of members of the Watson subfamily contains a certain nucleotide sequence; d. To accurately represent the click strand of the DNA fragment under analysis and identify the sequence if a threshold percentage of members of the click subfamily contains a certain nucleotide sequence; e. In cases where a nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the Watson chain but lacks a certain mutation, for the purpose of identifying that mutation; f. When a nucleotide sequence that accurately represents a click chain differs from a reference sequence that accurately represents a click chain but lacks a certain mutation, for the purpose of identifying that mutation; g. When the mutation in the nucleotide sequence that accurately represents the Watson chain and the mutation in the nucleotide sequence that accurately represents the Crick chain are the same mutation, for identifying the mutation in the DNA fragment being analyzed. Includes executable instructions.

[0074] In some embodiments, the computer-readable medium includes executable instructions for assigning a UID family member to a Watson subfamily when the exogenous UID sequence is immediately downstream of the R2 sequencing primer binding site or within 1 to 300 nucleotides. In some embodiments, the computer-readable medium includes executable instructions for assigning a UID family member to a Click subfamily when the exogenous UID sequence is immediately downstream of the R1 sequencing primer binding site or within 1 to 300 nucleotides. In some embodiments, the computer-readable medium includes executable instructions for mapping sequence reads to a reference genome. In some embodiments, the reference genome is a human reference genome.

[0075] In some embodiments, the computer-readable medium further includes computer-executable instructions for generating a report of treatment options based on the presence, absence, or amount of mutations in a sample. In some embodiments, the computer-readable medium further includes computer-executable code that enables data transmission over a network.

[0076] a. A storage device configured to receive sequence data from a nucleic acid sample, wherein the data is generated by the method described in any one of the claims; b. A processor communicatively coupled to a storage device, comprising a computer-readable medium according to any one of the claims. Computer systems including the above are also provided herein.

[0077] In some embodiments, the computer system further includes a sequencing system configured to communicate data to a storage device. In some embodiments, the computer system further includes a user interface configured to communicate or display reports to a user. In some embodiments, the computer system further includes a digital processor configured to transmit the results of data analysis over a network.

[0078] a. A collection of double-stranded DNA fragments from biological samples; b. A group of 3' adapters according to any one of the claims; c. A group of 5' adapters according to any one of the claims; d. Reagents for performing a nic translation-like reaction (e.g., using DNA polymerase, adherent end-specific ligase, and uracil-DNA glycosylase); e. Reagents for enriching amplicons with respect to one or more target polynucleotides; and f. Sequencing system. Systems including the above are also provided herein.

[0079] In some embodiments, the system further includes the computer system described in any one of the claims.

[0080] a. (i) One or more first Watson target-selective primers comprising a sequence complementary to a portion of the universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is optionally the R2 sequencing primer site of the universal 3' adapter sequence; and (ii) One or more second Watson target-selective primers, each comprising a target-selective sequence. A first set of Watson target-selective primer pairs including [specific component]; b. (i) One or more click target-selective primers comprising a sequence complementary to a portion of the universal 5' adapter sequence, wherein the portion of the universal 5' adapter sequence is optionally the R1 sequencing primer site of the universal 5' adapter sequence; and (ii) One or more second click target-selective primers, each comprising the same target-selective sequence as the second Watson target-selective primer sequence. A first set of click-target-selective primer pairs including; c. (i) One or more third Watson target-selective primers comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more fourth Watson target-selective primers, each comprising an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide in the 5'→3' direction. A second set of Watson target-selective primer pairs including; and d. (i) One or more third click target-selective primers comprising a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more fourth click target-selective primers each comprising, in the 5'→3' direction, an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide. A second set of click-target selective primers, including A kit including this is also provided herein.

[0081] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the invention pertains. While the invention can be carried out using methods and materials similar or equivalent to those described herein, suitable methods and materials are listed below. All publications, patent applications, patents, and other references referenced herein are incorporated by reference in their entirety. In case of any conflict, this specification, including its definitions, shall prevail. Furthermore, materials, methods, and examples are illustrative and not intended to limit the invention.

[0082] Details of one or more embodiments of the present invention are shown in the accompanying drawings and the following specification. Other features, purposes, and advantages of the present invention will become apparent from the specification, drawings, and claims. [Brief explanation of the drawing]

[0083] [Figure 1]This figure includes a schematic diagram of an exemplary duplex-anchored PCR method. A duplex adapter with a molecular barcode is ligated to the ends of a blunt-ended nucleic acid fragment to generate a duplex sequencing library, which is then subjected to PCR to generate an amplified duplex sequencing library. The amplification product in the amplified duplex sequencing library is separated into two aliquots, each aliquot being subjected to PCR in which the Watson strand is amplified from the first aliquot and the Crick strand is amplified from the second aliquot. [Figure 2] This figure includes a schematic diagram of an exemplary second round of library amplification, in which the Watson strands amplified from the first aliquot in Figure 1 are subjected to PCR using a primer pair in which the first primer is biotinylated and the second primer is not biotinylated to generate a single-stranded DNA library. This library can be used to amplify and evaluate Watson strands. [Figure 3] The figure includes a schematic diagram of an exemplary second round of library amplification, in which the click strands amplified from the first aliquot in Figure 1 are subjected to PCR using a primer pair in which the first primer is not biotinylated and the second primer is biotinylated to generate a single-stranded DNA library. This library can be used to amplify and evaluate click strands. [Figure 4] This figure includes a schematic diagram illustrating an example of Watson amplification. [Figure 5] This figure includes a schematic diagram of an example click amplification. [Figure 6] This figure includes schematic diagrams of exemplary amplified Watson chains and exemplary amplified Crick chains. [Figure 7] This figure includes a schematic diagram of an exemplary nested Watson amplification. [Figure 8] This figure includes a schematic diagram of an exemplary nested click amplifier. [Figure 9]This figure includes a schematic diagram illustrating the exemplary removal of 5' phosphate. [Figure 10] This figure includes a schematic diagram of an exemplary filling in of the 3' end of an amplification fragment having a 5' overhang for generating a blunt-end amplification product. [Figure 11] The figure includes a schematic diagram of an exemplary 3' duplex adapter containing a 3SpC3 spacer, an exogenous UID sequence containing a molecular barcode, and a 3' oligonucleotide (dT) hybridized with a 3' blocking group that can be degraded by uracil-DNA glycosylase (UDG). [Figure 12] This figure includes a schematic diagram of an exemplary 3' adapter ligation using a 3' duplex adapter. The 5' phosphate of the 3' duplex adapter is ligated to the 3' end of the nucleic acid template. [Figure 13] This figure includes a schematic diagram of an exemplary 5' adapter ligation. In a single reaction, the blocking group of the 3' duplex adapter is degraded, and the contained 5' adapter is ligated to the 5' end of the nucleic acid template via a nick translation reaction. [Figure 14] This figure includes a schematic diagram of exemplary library PCR amplification. [Figure 15] This figure includes a schematic diagram illustrating an example of Watson amplification. [Figure 16] This figure includes a schematic diagram of an exemplary nested Watson amplification. [Figure 17] This figure includes a schematic diagram of an example click amplification. [Figure 18] This figure includes a schematic diagram of an exemplary nested click amplifier. [Figure 19] This figure includes a schematic diagram of the final amplification product generated by an exemplary duplex-anchored PCR. [Figure 20]This figure includes a schematic diagram showing how paired-end sequencing can be used to distinguish the Watson strand from the click strand of the input nucleic acid using the final amplification product generated by an exemplary duplex-anchored PCR. [Figure 21] This figure includes a schematic diagram showing how paired-end sequencing can be used to distinguish the Watson strand from the click strand of the input nucleic acid using the final amplification product generated by an exemplary duplex-anchored PCR. [Figure 22A]This figure includes a schematic diagram illustrating an exemplary SaferSeqS method. (a) Library preparation begins with end repair (Step 1), in which DNA template molecules are dephosphorylated and blunted. Next, a 3' adapter (narrow or broad parallel diagonal lines) containing a unique identifier (UID) sequence is ligated to the 3' fragment end (Step 2). The UID sequence is converted into a double-stranded barcode through extension and ligation of the 5' adapter (Step 3). Finally, duplicate PCR copies of each original template molecule are generated during library amplification (Step 4). (b) Target enrichment is achieved by strand-specific hemi-nested PCR. The amplified library is split into Watson and Crick-specific reactions (Step 5), which selectively amplify products derived from one of the DNA strands (Step 6). Further on-target specificity and incorporation of sample barcodes are achieved by a second nested PCR (Step 7). The final PCR product (Step 8) is subjected to paired-end sequencing (Step 9). The endogenous barcode corresponds to the end of the template fragment before library construction. (c) After sequencing, it is determined that the read originates from either the Watson or Crick strand. Since each strand of the original template molecule is tagged with the same exogenous barcode and has the same endogenous barcode, reads originating from each of the two strands of the same parental DNA duplex can be grouped together into a duplex family. The various parallel line and dotted patterns at the right end of the strands represent different barcodes. In the example shown, each duplex family has eight members: four representing the Watson strand and four representing the Crick strand. In the actual experiments described herein, each family will contain at least two members from the Watson strand and two from the Crick strand, the actual number depending on the sequencing depth. True mutations, represented by asterisks within the true mutation family, are present in both parental strands of the DNA duplex and are therefore found in both the Watson and Crick families. In contrast, PCR or sequencing errors represented by an asterisk within the family of PCR or sequencing errors are limited to a subset of reads from one of the two strands.Watson-specific artifacts (indicated by an asterisk in the damaged Watson-chain family) and Crick-specific artifacts (indicated by an asterisk in the damaged Crick-chain family) are found in all copies of either the Watson or Crick family, but not in both. [Figure 22B] Please refer to the explanation in Figure 22A. [Figure 22C] Please refer to the explanation in Figure 22A. [Figure 23] This figure includes graphs demonstrating the analytical performance of SaferSeqS. The figure shows the mutation allele frequency (MAF) determined by SaferSeqS, compared to the expected frequency when DNA from cancer containing known mutations is mixed with leukocyte DNA from healthy donors at various ratios ranging from 10% downwards to 0.001%. A 0% control sample was also assayed to determine the specificity for the mutation of interest. The solid line represents the fit of a linear regression model with intercept y fixed at zero (slope = 0.776, R² > 0.999, P = 3.95 × 10⁻¹⁵). [Figure 24] This figure shows high duplex recovery and efficient target enrichment using SaferSeqS. 33 ng of mixed cfDNA samples were assayed for one of three different mutations in TP53 (p.L264fs, p.P190L, or p.R342X). Three libraries were prepared for each cfDNA sample, each containing approximately 11 ng of cfDNA. (a) The median number of duplex families (i.e., both Watson and Crick strands containing the same endogenous and exogenous barcodes) was 89% (range: 65%–102%) of the number of intrinsic template molecules. (b) The median percentage of on-target reads was 80% (range: 72%–91%). The lower and upper hinges correspond to the 25th and 75th percentiles, and the whiskers extend up to 1.5 times the interquartile range. For clarity, individual points are superimposed with random scattering. [Figure 25-1]This figure includes a graph showing the detection of exemplary mutations in liquid biopsy samples. Analysis of 33 ng of cell-free plasma DNA from healthy individuals mixed with cell-free plasma DNA from cancer patients. Mixtures were prepared to produce high-frequency (approximately 0.5–1%) mutations, low-frequency (approximately 0.01–0.1%) mutations, or no mutations. The mixed TP53 p.R342X samples were assayed using (a) SafeSeqs or (b) SaferSeqS. Similarly, the mixed TP53 p.L264fs samples were assayed using (c) SafeSeqs and (d) SaferSeqS, and the mixed TP53 p.P190L samples were assayed using (e) SafeSeqs and (f) SaferSeqS. The mutation numbers represent each of the 153 distinct mutations observed with SafeSeqS (defined in Table 8). The single supercalimutant detected by SaferSeqS (Table 9) was located outside the genomic region assayed by SafeSeqS and is therefore not shown. [Figure 25-2] Please refer to the explanation in Figure 25-1. [Figure 26-1] This figure shows the errors in SaferSeqS compared to errors in strand-agnostic ligation-based molecular barcoding methods. Analysis of 33 ng of cell-free plasma DNA from healthy individuals mixed with cell-free plasma DNA from cancer patients. Mixtures were prepared to produce high-frequency (approximately 0.5–1%) mutations, low-frequency (approximately 0.01–0.1%) mutations, or no mutations. SaferSeqS was used to assay the mixed TP53 p.R342X samples, either (a) ignoring strand information in the analysis to mimic strand-agnostic ligation-based molecular barcoding methods, or (b) considering strand information during mutation calling. Similarly, (c) without considering strand information and (d) using SaferSeqS, the mixed TP53 p.L264fs samples were assayed. (e) without considering strand information and (f) using SaferSeqS, the mixed TP53 p.P190L samples were assayed. Mutation numbers are defined in Table 3 of the appendix. The asterisk indicates a mixed mutation. A single unexpected superkali variant detected by SaferSeqS is shown in (e). [Figure 26-2] Please refer to the explanation in Figure 26-1. [Figure 27-1] This figure shows the evaluation of plasma samples from cancer patients. Plasma cell-free DNA samples from five cancer patients, each carrying eight known mutations at frequencies of 0.01%–0.1%, were assayed using the previously described PCR-based molecular barcoding method ("SafeSeqS" rather than "SaferSeqS") and with SaferSeqS. Mutation numbers are defined in Table 11. Asterisks indicate expected mutations. A single unexpected superkali variant detected by SaferSeqS (Table 11) is not shown because it was outside the genomic region assayed by SafeSeqS. [Figure 27-2] Please refer to the explanation in Figure 27-1. [Figure 28] This figure shows the effect of PCR efficiency and cycle count on duplex recovery rate. The probability of recovering both strands of the original DNA duplex (y-axis) is plotted against the number of library amplification cycles (x-axis). Each panel in the figure represents the assumed PCR efficiency shown at the top of the panel. The ratio of library amplification products used for strand-specific PCR is shown. The number of library amplification cycles was varied from 1 to 11. The PCR efficiency was varied from 100% to 50% in 10% increments. The ratio of library amplification products used for each strand-specific PCR was varied from 50% to 1.4%. Probability modeling was performed as described in Example 2. [Figure 29] This figure includes a graph showing a multiplex panel for the detection of exemplary cancer driver gene mutations. The recovery and coverage of 36 well-amplified amplicons within the multiplex panel are shown. The horizontal axis indicates the downstream position of the 3' end of the second gene-specific primer (GSP2). The gradual decrease in coverage with increasing distance from the 3' primer end is a result of the fragmentation pattern of the input DNA. Details regarding the theoretical recovery rate of reads with specific amplicon lengths are described in Example 2. [Figure 30]Performance of 48 primer pairs used in a multiplex panel to assay regions of driver genes commonly mutated in cancer. On-target read ratio (i.e., the percentage of total reads mapped to the intended target) for each of the 48 SaferSeqS primer pairs used in strand-specific PCR. Primers were used at equimolar concentrations in each gene-specific PCR. [Figure 31] Performance of 62 primer pairs. The on-target read ratio (i.e., the percentage of total reads mapped to the intended target) for each of the 62 SaferSeqS primer pairs tested to date. 50 out of 62 pairs (81%) showed an on-target rate greater than 50%. The presented results reflect one attempt at primer design. [Figure 32] An exemplary computer system adapted to enable a user to analyze nucleic acid samples according to the methods described herein is shown. [Modes for carrying out the invention]

[0084] Detailed explanation When used herein and in the appended claims, it should be noted that the singular forms “a,” “an,” and “the” also refer to plural objects unless the context explicitly indicates otherwise.

[0085] The terms "nucleotide" and "nt" are used interchangeably herein to generally refer to biomolecules containing nucleic acids. Nucleotides may have moieties containing known purine and pyrimidine bases. Nucleotides may have other modified heterocyclic bases. Such modifications include, for example, methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses, or other heterocyclic bases. The terms "polynucleotide," "nucleic acid," and "oligonucleotide" can be used interchangeably. They may refer to polymeric forms of nucleotides of any length, such as deoxyribocleotides or ribonucleotides or their analogues. Polynucleotides may have any three-dimensional structure and may perform any known or unknown function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain non-natural sequences. Polynucleotides may contain modified nucleotides such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure, if present, may be made before or after the assembly of the polymer. Non-nucleotide components may be interposed in the sequence of nucleotides. Polynucleotides may be further modified after polymerization, such as by conjugation with labeled components.

[0086] A "primer" is generally a polynucleotide molecule (e.g., oligonucleotide) containing a nucleotide sequence that typically has a 3′-OH group, capable of hybridizing with a template sequence (e.g., a target polynucleotide or primer extension product) and promoting the polymerization of a polynucleotide complementary to the template.

[0087] As used herein, the term "mammal" includes both human and non-human mammals, including but not limited to humans, non-human primates, dogs, cats, mice, cows, horses, and pigs.

[0088] Overview This document relates to methods and materials useful for accurately identifying mutations present in a nucleic acid sample. In some aspects, the methods include identifying mutations when the mutations are present in both the Watson and Crick strands of a double-stranded nucleic acid template. Such methods are particularly useful for discriminating true mutations from, for example, DNA damage, artifacts resulting from PCR, and other sequencing artifacts, and enabling the identification of mutations with high confidence.

[0089] In some cases, the methods and materials described herein can detect one or more mutations with a low error rate. For example, using the methods and materials described herein, the presence or absence of a nucleic acid mutation in a nucleic acid template can be detected with an error rate of less than about 1% (e.g., less than about 0.1%, less than about 0.05%, or less than about 0.01%). In some cases, using the methods and materials described herein, the presence or absence of a nucleic acid mutation in a nucleic acid template can be detected with an error rate of about 0.001% to about 0.01%. In some cases, the error rate associated with the identification of one or more mutations in an analyzed DNA fragment by the methods described herein is 1×10 -2 Hereinafter, 1×10 -3 Hereinafter, 1×10 -4 Hereinafter, 1×10 -5 Hereinafter, 1×10 -6 Hereinafter, 5×10 -6 Hereinafter, or 1×10 -7The following applies: In some cases, the error rate associated with identifying one or more mutations in the DNA fragment under analysis by the methods described herein is reduced by at least 1 / 2, 1 / 4, 1 / 5, 1 / 10, 1 / 20, 1 / 30 times, 1 / 40, 1 / 50, 1 / 60, 1 / 70, 1 / 80, 1 / 90, or 1 / 100 times compared to alternative methods for identifying mutations that do not require the detection of mutations in both the Watson and Crick strands of the DNA fragment under analysis.

[0090] In some embodiments, the alternative method includes sequencing following standard molecular barcoding or standard PCR-based molecular barcoding. In certain embodiments, the alternative method includes (a) a step of ligating an adapter to a population of double-stranded DNA fragments in a DNA sample to be analyzed, wherein the adapter contains a unique exogenous UID; (b) a step of performing an initial amplification to amplify the double-stranded DNA fragments ligated with the adapter to produce an amplicon; (c) a step of determining sequence reads of one or more amplicons of one or more double-stranded DNA fragments ligated with the adapter; (d) a step of assigning the sequence reads to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; (e) a step of identifying a sequence that accurately represents the DNA fragment to be analyzed if a threshold percentage of the members of the UID family contains a certain nucleotide sequence; and (f) a step of identifying a mutation if the sequence identified as accurately representing the DNA fragment to be analyzed differs from a reference sequence lacking a certain mutation in the DNA fragment to be analyzed.

[0091] In some cases, efficient duplex recovery can be achieved using the methods and materials described herein. For example, PCR amplification products derived from both the Watson and Crick strands of a double-stranded nucleic acid template can be recovered using the methods described herein. In some cases, duplex recovery rates of at least 50% (e.g., about 50%, about 60%, about 70%, about 75%, about 80%, about 82%, about 85%, about 88%, about 90%, about 93%, about 95%, about 97%, about 99%, or 100%) can be achieved using the methods described herein.

[0092] In some cases, the methods and materials described herein can be used to detect variants with low allele frequencies. For example, the methods described herein can be used to detect variants with low allele frequencies of less than about 1% (e.g., less than about 0.1%, less than about 0.05%, or less than about 0.01%). In some cases, the methods described herein can be used to detect variants with low allele frequencies of about 0.001%.

[0093] In some cases, the methods described herein can be used to detect mutations present in a nucleic acid sample at a frequency of 0.1% or less. In some embodiments, the methods described herein can be used to detect mutations present in a nucleic acid sample at a frequency of 0.1% to 0.00001%. In some embodiments, the methods described herein can be used to detect mutations present in a nucleic acid sample at a frequency of 0.1% to 0.01%.

[0094] In some cases, the methods and materials described herein can be used to detect mutations with minimal background artifact mutations (or without background artifact mutations). In some cases, the methods described herein can be used to detect mutations with background artifact mutations of less than 0.01%. In some cases, the methods described herein can be used to detect mutations without background artifact mutations.

[0095] In some cases, a method for detecting one or more mutations present on both strands of a double-stranded nucleic acid may include the steps of: generating a duplex sequencing library having duplex molecular barcodes at each end (e.g., the 5' and 3' ends) of each nucleic acid in the library; generating a library of single-stranded Watson-derived sequences and a library of single-stranded Crick-derived sequences from the duplex sequencing library; and detecting the presence of one or more mutations present on both strands of the double-stranded nucleic acid in each single-stranded library. The presence of a first molecular barcode in the 3' duplex adapter and a second molecular barcode in the 5' adapter can be used to distinguish amplification products derived from the Watson strand from amplification products derived from the Crick strand.

[0096] In some cases, a method for identifying mutations involves (a) attaching a partial double-stranded 3' adapter to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, wherein the first strand of the partial double-stranded 3' adapter is 5'- (b) A step of annealing the 5' adapter to the 3' adapter via the annealing site, wherein the 5' adapter comprises a universal 3' adapter sequence in the 3' direction, including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the partially double-stranded 3' adapter comprises (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction, and optionally the second strand is digestible; (b) A step of annealing the 5' adapter to the 3' adapter via the annealing site, wherein the 5' adapter comprises a universal 5' adapter sequence in the 5'→3' direction, including (i) a sequence not complementary to the universal 3' adapter sequence and including an R1 sequencing primer site, and (ii) a sequence complementary to the annealing site for the 5' adapter; (c) An extension of the 5' adapter across the exogenous UID sequence of the 3' adapter (e.g., using DNA polymerase), thereby extending the extended 5' adapter (d) A nick translation-like reaction to covalently ligate (e.g., using a ligase) the Watson and Crick strands of a double-stranded DNA fragment; (e) A first amplification to amplify the double-stranded DNA fragment ligated with the adapter to produce an amplicon; (f) A step to determine the sequence reads of one or more amplicons of one or more double-stranded DNA fragments ligated with the adapter; (g) A step to assign the sequence reads to UID families, wherein each member of the UID family contains the same exogenous UID sequence; (h) A step to assign the sequence reads of each UID family to Watson subfamilies and Crick subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; (h) A step to identify a sequence that accurately represents the Watson strand of the DNA fragment under analysis if a threshold percentage of members of the Watson subfamily contain a certain nucleotide sequence.(i) a step of identifying a nucleotide sequence that accurately represents the click chain of the DNA fragment under analysis if a threshold percentage of members of the click subfamily contains a certain nucleotide sequence; (j) a step of identifying a mutation if the nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the Watson chain and lacks a certain mutation in that sequence; (k) a step of identifying a mutation if the nucleotide sequence that accurately represents the click chain differs from a reference sequence that accurately represents the click chain and lacks a certain mutation in that sequence; and (l) a step of identifying a mutation in the DNA fragment under analysis if the mutation in the nucleotide sequence that accurately represents the Watson chain and the mutation in the nucleotide sequence that accurately represents the click chain are the same mutation.

[0097] In some cases, a method for identifying mutations may include: (a) a step of attaching an adapter to a population of double-stranded DNA fragments, wherein the adapter comprises a double-stranded portion containing an exogenous UID and a fork-shaped portion containing (i) a single-stranded 3' adapter sequence containing an R2 sequencing primer site and (ii) a single-stranded 5' adapter sequence containing an R1 sequencing primer site; (b) a step of performing an initial amplification to amplify the double-stranded DNA fragments ligated with the adapter to produce an amplicon; and (c) a Watson target-selective primer. - A step of selectively amplifying the amplicon of a Watson chain containing a target polynucleotide sequence using a first set of pairs of Watson target-selective primers, wherein the first set of pairs of Watson target-selective primers comprises (i) a first Watson target-selective primer containing a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a second Watson target-selective primer containing a target-selective sequence; (d) a step of using the first set of pairs of click target-selective primers, A step of selectively amplifying the amplicon of a click strand containing a target polynucleotide sequence, thereby producing a target click amplification product, wherein a first set of click target-selective primer pairs comprises (ii) a first click target-selective primer containing a sequence complementary to the R1 sequencing primer site of a universal 5' adapter sequence, and (ii) a second click target-selective primer containing the same target-selective sequence as the second click target-selective primer sequence; (e) determining the sequence reads of the target Watson amplification product and the target click amplification product; (f) assigning the sequence reads to UID families, wherein each member of the UID family contains the same exogenous UID sequence; (g) assigning the sequence reads of each UID family to Watson subfamilies and click subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; (h) identifying a sequence that accurately represents the Watson strand of the DNA fragment under analysis if a threshold percentage of members of the Watson family contains a certain nucleotide sequence;(i) a step of identifying a nucleotide sequence that accurately represents the click strand of the DNA fragment under analysis, if a threshold percentage of members of the click family contains that nucleotide sequence; and (j) a step of identifying a mutation in the DNA fragment under analysis, if both the nucleotide sequence that accurately represents the Watson strand and the nucleotide sequence that accurately represents the click strand contain the same mutation.

[0098] In some cases, the methods and materials described herein can be used to independently evaluate each strand of a double-stranded nucleic acid. For example, if a nucleic acid mutation is identified in an independently evaluated strand of a double-stranded nucleic acid as described herein, the materials and methods described herein can be used to determine which strand of the double-stranded nucleic acid the nucleic acid mutation originates from.

[0099] A duplex sequencing library can be generated using any suitable method. The duplex sequencing library used herein comprises multiple nucleic acid fragments, each containing a duplex molecular barcode at one end (e.g., the 5' end and / or 3' end) of each nucleic acid fragment in the library, enabling sequencing of both strands of a double-stranded nucleic acid. In some cases, a nucleic acid sample can be fragmented to generate nucleic acid fragments, and the generated nucleic acid fragments can be used to generate a duplex sequencing library. The nucleic acid fragments used to generate the duplex sequencing library may also be referred to herein as input nucleic acids. For example, if the nucleic acid fragments used to generate the duplex sequencing library are DNA fragments, the DNA fragments may also be referred to herein as input DNA. A duplex sequencing library can contain any suitable number of nucleic acid fragments. In some cases, generating a duplex sequencing library may involve fragmenting a nucleic acid template and ligating adapters to each end of each nucleic acid fragment in the library.

[0100] Nucleic acid sample to be analyzed The nucleic acid template in the nucleic acid sample to be analyzed may include any type of nucleic acid (e.g., DNA, RNA, and DNA / RNA hybrids). In some cases, the nucleic acid template may be a double-stranded DNA template. Examples of nucleic acids that can be used as templates for the methods described herein include, but are not limited to, genomic DNA and circulating free DNA (cfDNA; e.g., circulating tumor DNA (ctDNA), and cell-free fetal DNA (cffDNA)).

[0101] In some embodiments, the nucleic acid template in the nucleic acid sample is a nucleic acid fragment, such as a DNA fragment. In some embodiments, the ends of the DNA fragment represent a unique sequence that can be used as an endogenous unique identifier for the fragment. In some embodiments, the fragment is produced manually. In some embodiments, the fragment is produced by shearing, e.g., enzymatic shearing, chemical shearing, acoustic shearing, spraying, centrifugal shearing, point-sink shearing, needle shearing, sonication, restriction endonuclease, nonspecific nuclease (e.g., DNase I), or other methods. In some embodiments, the fragment is not produced manually. In some embodiments, the fragment is derived from a cfDNA sample.

[0102] In some embodiments, nucleic acid fragments in a nucleic acid sample have length. The length may be about 4 to 1000 nucleotides. The length may be about 60 to 300 nucleotides. The length may be about 60 to 200 nucleotides. The length may be about 140 to 170 nucleotides. The length may be less than 500, less than 400, less than 300, less than 250 nt, or less than 200 nt.

[0103] In some embodiments, the ends of the nucleic acid template are used as the endogenous UID. Those skilled in the art can determine the length of the endogenous UID required to uniquely identify the nucleic acid template using factors such as, for example, the total length of the template, the complexity of the nucleic acid template in a fragmented nucleic acid sample or a starting nucleic acid sample, and others. In some embodiments, the last 10 to 500 nucleotides of the nucleic acid template are used as the endogenous UID. In some embodiments, the last 15 to 100 nucleotides of the nucleic acid template are used as the endogenous UID. In some embodiments, the last 15 to 40 nucleotides of the nucleic acid template are used as the endogenous UID. In some embodiments, at least 10 nucleotides of the last nucleic acid template are used as the endogenous UID. In some embodiments, at least 15 nucleotides of the last nucleic acid template are used as the endogenous UID. In some embodiments, only one end of the nucleic acid template is used as the endogenous UID.

[0104] In some embodiments, the nucleic acid template comprises one or more target polynucleotides. The terms “target polynucleotide,” “target region,” “nucleic acid template of interest,” “desired locus,” “desired template,” or “target” are used interchangeably herein to refer to the polynucleotide of interest under investigation. In certain embodiments, the target polynucleotide comprises one or more sequences of interest under investigation. The target polynucleotide may include, for example, a genomic sequence. The target polynucleotide may include target sequences whose presence, quantity, and / or nucleotide sequence, or changes therein, are desirable to determine.

[0105] The target polynucleotide can be a region of a disease-related gene. In some embodiments, the gene is a target that could lead to the development of a new drug. As used herein, the term “target that could lead to the development of a new drug” generally refers to a gene or cellular pathway that is modulated by disease treatment. The disease can be cancer. Therefore, the gene can be a known cancer-related gene.

[0106] In some embodiments, the input nucleic acid, also referred to herein as the nucleic acid sample, is obtained from a biological sample. The biological sample may be obtained from a subject. In some embodiments, the subject is a mammal. Examples of mammals from which nucleic acids can be obtained and used as nucleic acid templates in the methods described herein include, but are not limited to, humans, non-human primates (e.g., monkeys), dogs, cats, sheep, rabbits, mice, hamsters, and rats. In some embodiments, the subject is a human subject. In some embodiments, the subject is a plant.

[0107] Biological specimens include, but are not limited to, plasma, serum, blood, tissue, tumor specimens, stool, sputum, saliva, urine, sweat, tears, ascites, bronchoalveolar lavage fluid, semen, archaeological specimens, and forensic specimens. In certain embodiments, biological specimens are solid biological specimens, such as tumor specimens. In some embodiments, solid biological specimens are processed. Solid biological specimens may be processed by paraffin embedding (e.g., FFPE specimens) following fixation in formalin solution. Alternatively, processing may include freezing the specimen before performing probe-based assays. In some embodiments, specimens are neither fixed nor frozen. Unfixed, unfrozen specimens may, as merely an example, be stored in preservation solutions configured for nucleic acid preservation.

[0108] In some embodiments, the biological sample is a liquid biological sample. A liquid biological sample includes, but is not limited to, plasma, serum, blood, sputum, saliva, urine, sweat, tears, ascites, bronchoalveolar lavage fluid, and semen. In some embodiments, the liquid biological sample is cell-free or substantially cell-free. In certain embodiments, the biological sample is a plasma or serum sample. In some embodiments, the liquid biological sample is a whole blood sample. In some embodiments, the liquid biological sample includes peripheral blood mononuclear cells.

[0109] In some embodiments, nucleic acid samples are isolated and purified from biological samples. Nucleic acids can be isolated and purified from biological samples using any means known in the art. For example, biological samples may be treated to release nucleic acids from cells or to separate nucleic acids from unwanted components of the biological sample (e.g., proteins, cell walls, other contaminants). For example, nucleic acids can be extracted from biological samples using liquid extraction techniques (e.g., Trizole, DNAzol). Nucleic acids can also be extracted using commercially available kits (e.g., Qiagen DNeasy kit, QIAamp kit, Qiagen Midi kit, QIAprep spin kit).

[0110] In some embodiments, the biological sample contains a small amount of nucleic acid. In some embodiments, the biological sample contains less than about 500 nanograms (ng) of nucleic acid. For example, the biological sample contains about 30 ng to about 40 ng of nucleic acid.

[0111] Nucleic acids can be concentrated by known methods, including centrifugation, as merely an example. Nucleic acids can be bound to a selective membrane (e.g., silica) for purification. Nucleic acids can also be enriched with fragments of desired length, e.g., fragments less than 1000, 500, 400, 300, 200, or 100 base pairs in length. Such size-based enrichment can be carried out, for example, using PEG-induced precipitation, electrophoretic gels or chromatographic materials (Huber et al. (1993) Nucleic Acids Res. 21:1061-6), gel filtration chromatography, or TSK gels (Kato et al. (1984) J. Biochem, 95:83-86), the publications of which are incorporated herein by reference.

[0112] Polynucleotides extracted from biological samples can be selectively precipitated or concentrated using any method known in the art.

[0113] In some embodiments, the nucleic acid sample contains less than approximately 35 ng of nucleic acid. For example, the nucleic acid sample may contain approximately 1 ng to approximately 35 ng of nucleic acid (e.g., approximately 1 ng to approximately 30 ng, approximately 1 ng to approximately 25 ng, approximately 1 ng to approximately 20 ng, approximately 1 ng to approximately 15 ng, approximately 1 ng to approximately 10 ng, approximately 1 ng to approximately 5 ng, approximately 5 ng to approximately 35 ng, approximately 10 ng to approximately 35 ng, approximately 15 ng to approximately 35 ng, approximately 20 ng to approximately 35 ng, approximately 30 ng to approximately 35 ng, approximately 5 ng to approximately 30 ng, approximately 10 ng to approximately 25 ng, approximately 15 ng to approximately 20 ng, approximately 15 ng to approximately 20 ng, approximately 20 ng to approximately 25 ng, or approximately 25 ng to approximately 30 ng of nucleic acid). In some cases, nucleic acid samples may contain nucleic acids from a genome that are larger than approximately several hundred nucleotides.

[0114] In some cases, nucleic acid samples may be inherently free of contamination. For example, if a nucleic acid sample is a cfDNA template, the cfDNA may be inherently free of genomic DNA contamination. In some cases, a cfDNA sample that is inherently free of genomic DNA contamination may contain (or not contain) a small amount of high molecular weight (e.g., >1000 bp) DNA. In some cases, the methods described herein may include determining whether a nucleic acid sample is inherently free of contamination. Any suitable method can be used to determine whether a nucleic acid sample is inherently free of contamination. Examples of methods that can be used to determine whether a nucleic acid sample is inherently free of contamination include, for example, the TapeStation system and the Bioanalyzer. For example, when using the TapeStation system and / or Bioanalyzer to determine whether a cfDNA sample is inherently free of genomic DNA contamination, a prominent peak at approximately 180 bp (e.g., corresponding to mononucleosomal DNA) can be used to indicate that the nucleic acid sample is inherently free of genomic DNA contamination.

[0115] In some cases, nucleic acid fragments that can be used to generate duplex sequencing libraries (e.g., before 3' duplex adapters are attached to the 3' end of the nucleic acid fragment) can be repaired at the ends. Nucleic acid templates can be repaired at the ends using any suitable method. For example, nucleic acid templates can be repaired at the ends using blunting reactions (e.g., blunt end ligation) and / or dephosphorylation reactions. In some cases, blunting may include filling in single-stranded regions. In some cases, blunting may include decomposing single-stranded regions. In some cases, nucleic acid templates can be repaired at the ends using blunting and dephosphorylation reactions as shown in Figure 9 and / or Figure 10.

[0116] adapter In some embodiments, the method includes the step of attaching an adapter to a population of double-stranded DNA fragments to produce a population of double-stranded DNA fragments bound to the adapter.

[0117] In some embodiments, the adapter comprises a double-stranded portion containing an exogenous UID and a fork-shaped portion containing (i) a single-stranded 3' adapter sequence and (ii) a single-stranded 5' adapter sequence. In some embodiments, the single-stranded 3' adapter sequence is not complementary to the single-stranded 5' adapter sequence. In some embodiments, the 3' adapter sequence includes a second (e.g., R2) sequencing primer site, and the 5' adapter sequence includes a first (e.g., R1) sequencing primer site. It should be understood that the "R1" and "R2" sequencing primer sites are used by a sequencing system to produce paired-end reads, e.g., reads from the reverse end of a DNA fragment to be sequenced. In some embodiments, the R1 sequencing primer is used to produce a first population of reads from the first end of the DNA fragment, and the R2 sequencing primer is used to produce a second population of reads from the reverse end of the DNA fragment. The first population is referred to herein as "R1" or "read 1" reads. The second group is referred to herein as “R2” or “Read 2” reads. The R1 and R2 reads can be aligned as “read pairs” or “mate pairs” corresponding to each strand of the double-stranded DNA fragment to be analyzed.

[0118] Certain sequencing systems, such as Illumina, utilize what are referred to as “R1” and “R2” primers, as well as “R1” and “R2” reeds. It should be noted that the terms “R1” and “R2,” and “Reed 1” and “Reed 2,” are not limited for the purposes of this application to how they refer to a particular sequencing platform. For example, if an Illumina sequencer is used, the “R2” primer and corresponding R2 reed disclosed herein may refer to either Illumina “R2” primers and reeds or Illumina “R1” primers and reeds, insofar as the “R1” primer and corresponding R1 reed disclosed herein refer to other Illumina primers and reeds. For clarity, in some embodiments, the “R2” primer provided herein is an Illumina “R1” primer producing an “R1” reed, and the corresponding “R1” primer provided herein is an Illumina “R2” primer producing an “R2” reed. For clarity, in some embodiments, the “R2” primer provided herein is an Illumina “R2” primer providing an “R2” lead, while the “R1” primer provided herein is an Illumina “R1” primer providing an R1 lead.

[0119] In some embodiments, the exogenous UID is specific to each double-stranded DNA fragment in the nucleic acid sample. In some embodiments, the exogenous UID is not specific to each double-stranded DNA fragment.

[0120] In some embodiments, the exogenous UID has length. The length can be approximately 2 to 4000 nt. The length can be approximately 6 to 100 nt. The length can be approximately 8 to 50 nt. The length can be approximately 10 to 20 nt. The length can be approximately 12 to 14 nt. In some embodiments, the length of the exogenous UID is sufficient to barcode the molecule uniquely, and the length / sequence of the exogenous UID does not interfere with downstream amplification steps.

[0121] In some embodiments, the exogenous UID sequence is not present in the nucleic acid template. In some embodiments, the exogenous UID sequence is not present in the desired template having the desired locus. Such unique sequences can be randomly generated, for example, by a computer-readable medium and selected by performing a BLAST against a known nucleotide database such as EMBL, GenBank, or DDBJ. In some embodiments, the exogenous UID sequence is present in the nucleic acid template. In such cases, the position of the exogenous UID sequence in the sequence read is used to identify the exogenous UID sequence from the sequence in the nucleic acid template.

[0122] In some embodiments, the exogenous UID sequence is random. In some embodiments, the exogenous UID sequence is a random N-mer. For example, if the exogenous UID sequence has a length of 6nt, it may be a random hexamer. If the exogenous UID sequence has a length of 12nt, it may be a random dodecamer.

[0123] Exogenous UIDs can be produced by randomly adding nucleotides to form a sequence of a length to be used as an identifier. A selection of one of four deoxyribonucleotides may be used at each addition site. Alternatively, a selection of three, two, or one deoxyribonucleotides may be used. Therefore, a UID can be completely random, partially random, or non-random at any given position.

[0124] In some embodiments, the exogenous UID is selected from a predetermined set of exogenous UID sequences, rather than being a random N-mer.

[0125] An exemplary exogenous UID suitable for use in the methods disclosed herein is described in PCT / US2012 / 033207, which is incorporated herein by reference in its entirety.

[0126] The fork-type adapter described herein can be attached to a double-stranded DNA fragment by any means known in the art.

[0127] In some embodiments, the fork-type adapter is (a) A partial double-stranded 3' adapter is attached to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments, wherein the first strand of the partial double-stranded 3' adapter comprises a universal 3' adapter sequence in the 5'-3' direction including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the partial double-stranded 3' adapter comprises (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction, and optionally the second strand is gradable; (b) Annealing the 5' adapter to the 3' adapter via an annealing site, wherein the 5' adapter includes, in the 5'→3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence complementary to the annealing site for the 5' adapter; and (c) Perform a Nick translation-like reaction to extend the 5' adapter across the exogenous UID sequence of the 3' adapter (e.g., using DNA polymerase) and covalently ligate the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment (e.g., using ligase). It is bound to the double-stranded DNA fragment by this process.

[0128] In some embodiments, the fork-type adapter is attached to a double-stranded DNA fragment by (a) attaching the 3' duplex adapter to both the 3' ends of the Watson and Crick strands of the double-stranded DNA fragment population. The 3' duplex adapter, also referred to herein as a partially double-stranded 3' adapter, is an oligonucleotide complex containing a molecular barcode, which may have a first oligonucleotide (also referred herein as the "first strand") annealed (hybridized) to a second oligonucleotide (also referred herein as the "second strand") such that a portion of the 3' duplex adapter (e.g., a first portion) is double-stranded and a portion of the 3' duplex adapter (e.g., a second portion) is single-stranded. In some cases, the first oligonucleotide of the 3'-duplex adapter described herein includes a first segment containing a nucleotide complementary to the nucleotide present in the second oligonucleotide of the 3'-duplex adapter (for example, as a result, the first and second oligonucleotides of the 3'-duplex adapter can be annealed in complementary regions). An exemplary structure of the 3'-duplex adapter may be as shown in Figure 11.

[0129] The first oligonucleotide of the 3' duplex adapter described herein may be an oligonucleotide comprising a 5' phosphate and a molecular barcode. The first oligonucleotide of the 3' duplex adapter described herein may contain any appropriate number of nucleotides. Any appropriate molecular barcode may be included in the first oligonucleotide of the 3' duplex adapter described herein. In some cases, the molecular barcode may contain a random sequence. In some cases, the molecular barcode may contain a constant sequence. Examples of molecular barcodes that may be included in the first oligonucleotide of the 3' duplex adapter described herein include, but are not limited to, IDT 8, IDT 10, ILMN 8, and ILMN 10 available from Integrated DNA Technologies. Any appropriate type of molecular barcode may be used. In some cases, the molecular barcode may contain an exogenous UID sequence. Exogenous UIDs are described herein. Examples of oligonucleotides that may be included in the first oligonucleotide of the 3' duplex adapter described herein, comprising a 5' phosphate and a molecular barcode, are: This includes, but is not limited to, TIFF2026133491000001.tif11157 (the asterisk indicates a phosphorothioate bond; SEQ ID NO: 1), and the sequence contains, TIFF2026133491000002.tif4128 is a molecular barcode, and the number of nucleotides in a molecular barcode can range from 0 to approximately 25.

[0130] In some embodiments, the first oligonucleotide of the 3' duplex adapter includes an annealing site for the 5' adapter.

[0131] In some embodiments, the first oligonucleotide of the 3' duplex adapter comprises a universal 3' adapter sequence. In some embodiments, the universal 3' adapter sequence comprises an R2 sequencing primer site.

[0132] In some cases, the first oligonucleotide of the 3' duplex adapter described herein may also include one or more features for preventing or reducing elongation during PCR. Features capable of preventing or reducing elongation during PCR can be of any kind (e.g., chemical modifications). Examples of features capable of preventing or reducing elongation during PCR and that may be included in the first oligonucleotide of the 3' duplex adapter described herein include, but are not limited to, 3SpC3 and 3Phos. Features capable of preventing or reducing elongation during PCR can be incorporated into the first oligonucleotide of the 3' duplex adapter described herein at any suitable position within the oligonucleotide. In some cases, molecules capable of preventing or reducing elongation during PCR can be incorporated inside the oligonucleotide. In other cases, molecules capable of preventing or reducing elongation during PCR can be incorporated at the end of the oligonucleotide (e.g., the 5' end).

[0133] In certain embodiments, the first oligonucleotide of the 3' duplex adapter comprises a 5' phosphate, a first segment containing nucleotides complementary to the nucleotides present in the second oligonucleotide of the 3' duplex adapter, an exogenous UID sequence, an annealing site to the 5' adapter, and a universal 3' adapter sequence.

[0134] The second oligonucleotide of the 3' duplex adapter described herein may be an oligonucleotide containing a blocked 3' group (for example, to reduce or eliminate dimerization of the two adapters). The second oligonucleotide of the 3' duplex adapter described herein may contain any appropriate number of nucleotides. In some embodiments, the second oligonucleotide of the 3' duplex adapter is complementary to the first segment of the first oligonucleotide of the 3' duplex adapter. Exemplary oligonucleotides that contain a blocked 3' group and may be included in the second oligonucleotide of the 3' duplex adapter described herein are: This includes, but is not limited to, TIFF2026133491000003.tif4128.

[0135] The second oligonucleotide of the 3' duplex adapter described herein may be degradable. The second oligonucleotide of the 3' duplex adapter described herein can be degraded using any suitable method. For example, the second oligonucleotide of the 3' duplex adapter described herein can be degraded using UDG.

[0136] In some cases, the 3' duplex adapter described herein is an array Sequence annealed to a second oligonucleotide containing TIFF2026133491000004.tif4128 It may contain a first oligonucleotide including TIFF2026133491000005.tif12157.

[0137] In some cases, the 3' duplex adapters described herein may include commercially available adapters. Exemplary commercially available adapters that can be used as (or used to generate) the 3' duplex adapters described herein include, but are not limited to, the adapter in the Accel-NGS 2S DNA Library Kit (Swift Biosciences, cat. # 21024). In some cases, the 3' duplex adapters described herein may be as described in Example 1.

[0138] The 3' adapter can be attached (e.g., covalently) to the 3' end of a double-stranded DNA fragment using any suitable method. In some embodiments, the 3' adapter is attached by ligation. In some embodiments, ligation involves the use of a ligase. Examples of ligases that can be used to attach the 3' adapter to the 3' end of each nucleic acid fragment include, but are not limited to, T4 DNA ligase, E. coli ligase (e.g., Enzyme Y3), CircLigase I, CircLigase II, Taq-ligase, T3 ligase, T7 ligase, and 9N ligase.

[0139] After the 3' duplex adapter is attached (e.g., covalently) to the 3' end of each nucleic acid fragment, the second oligonucleotide of the 3' duplex adapter described herein can be degraded, and the 5' adapter can be attached (e.g., covalently) to the 5' end of each nucleic acid fragment. In some embodiments, the 5' adapter sequence is not complementary to the first oligonucleotide of the 3' adapter. In some embodiments, the 5' adapter sequence includes a sequence complementary to the R1 sequencing primer site and the annealing site of the 3' adapter in the 5'→3' direction.

[0140] In some embodiments, coupling the 5' adapter involves annealing the 5' adapter to the 3' adapter via an annealing portion.

[0141] A 5' adapter can be annealed to a nucleic acid fragment upstream of the molecular barcode on the 3' duplex adapter such that a gap (e.g., a single-stranded nucleic acid fragment) containing a portion of the 3' duplex adapter (e.g., a molecular barcode) is present in the nucleic acid fragment. The gap containing a portion of the 3' duplex adapter can be filled (e.g., to produce a double-stranded nucleic acid fragment). The single-stranded gap can be filled using any suitable method. Examples of methods that can be used to fill a single-stranded gap on a nucleic acid fragment include, but are not limited to, polymerases such as DNA polymerase (e.g., Taq polymerase such as Taq-B polymerase) and nick translation reactions (e.g., including both ligases such as E. coli ligase and polymerases such as DNA polymerase). If filling a single-stranded gap on a nucleic acid fragment involves providing a polymerase, the method may also include providing deoxyribonucleotide triphosphates (dNTPs; e.g., dATP, dGTP, dCTP, and dTTP). In some cases, attaching a 5' adapter to the 5' end of each nucleic acid fragment and filling the single-strand gap can be done simultaneously (for example, in a single reaction tube).

[0142] In some cases, alternative methods can be used to bind adapters to templates. For example, nucleic acid fragments can be treated with single-stranded nucleases (e.g., to digest overhangs), and then ligation can be used to prepare a duplex sequencing library. For instance, a single nucleotide can be added to the 3' end of each nucleic acid fragment, and an adapter containing a complementary base at the 5' end (e.g., containing a molecular barcode) can be ligated to each nucleic acid fragment to prepare a duplex sequencing library of adapter-bound templates.

[0143] The first amplification of an adapter-coupled template Following adapter binding, the adapter-bound template can be amplified in the first amplification reaction (e.g., by PCR amplification). The adapter-bound template can be amplified using any suitable method. Exemplary methods that can be used to amplify the adapter-bound template include, but are not limited to, whole-genome PCR.

[0144] Adapter-bound templates can be amplified using any suitable primer pair. In some cases, a universal primer pair can be used. Primers may contain, but are not limited to, about 12 to about 30 nucleotides. Examples of primer pairs that can be used to amplify the adapter-bound templates described herein include, but are not limited to, those described in Example 1 and / or Example 2.

[0145] Any suitable PCR conditions can be used for the initial amplification. PCR amplification can include a denaturation phase, an annealing phase, and an extension phase. Each phase of the amplification cycle can include any suitable conditions. In some cases, the denaturation phase can include a temperature of about 90°C to about 105°C (e.g., about 94°C to about 98°C) and a time of about 1 second to about 5 minutes (e.g., about 10 seconds to about 1 minute). For example, the denaturation phase can include a temperature of about 98°C for about 10 seconds. In some cases, the annealing phase can include a temperature of about 50°C to about 72°C and a time of about 30 seconds to about 90 seconds. In some cases, the extension phase can include a temperature of about 55°C to about 80°C and a time of about 15 seconds to about 30 seconds per kb of amplicon that will be produced. In some cases, the annealing phase and extension phase can be performed in a single cycle. For example, the annealing phase and extension phase can include a temperature of about 65°C for about 75 seconds.

[0146] The PCR conditions used for the initial amplification can include any appropriate number of PCR amplification cycles. In some cases, PCR amplification can include approximately 1 to approximately 50 cycles. In some embodiments, PCR amplification includes 11 cycles or less. In some embodiments, PCR amplification includes 7 cycles or less. In some embodiments, PCR amplification includes 5 cycles or less.

[0147] In some cases, PCR amplification can also include a reprogramming step if the PCR conditions include a thermoactivated polymerase. For example, PCR amplification can include a reprogramming step before performing the PCR amplification cycle. In some cases, the reprogramming step can include a temperature of approximately 94°C to 98°C and a time of approximately 15 seconds to 1 minute. For example, the reprogramming step can include a temperature of approximately 98°C for approximately 30 seconds.

[0148] In some cases, PCR amplification may also include a retention step. For example, PCR amplification may include a retention step after the PCR amplification cycle and optionally after any final extension step. In some cases, the retention step may include an indeterminate time at a temperature of approximately 4°C to 15°C.

[0149] In some cases, the duplex sequencing library produced as described herein (e.g., an amplified duplex sequencing library) can be purified. The duplex sequencing library can be purified using any suitable method. Exemplary methods that can be used to purify the duplex sequencing library include, but are not limited to, magnetic beads (e.g., solid-phase reversible immobilization (SPRI) magnetic beads).

[0150] Preparation of any ssDNA library In some cases, a duplex sequencing library can be used to generate libraries of single-stranded Watson strand-derived sequences and single-stranded click strand-derived sequences. Generating libraries of single-stranded Watson strand-derived sequences and single-stranded click strand-derived sequences can minimize non-specific amplification (e.g., from primers complementary to ligated sequences, such as 3' duplex adapters or 5' adapters). Libraries of single-stranded Watson strand-derived sequences and single-stranded click strand-derived sequences can be generated using any suitable method (e.g., from a duplex sequencing library generated as described herein). In some cases, libraries of single-stranded Watson strand-derived sequences and single-stranded click strand-derived sequences can be generated from an amplified duplex sequencing library by splitting the amplified product into at least two aliquots and subjecting each aliquot to PCR amplification, in which case the Watson strand is amplified from the first aliquot and the click strand is amplified from the second aliquot. For example, a single-stranded Watson chain library can be generated by subjecting a first aliquot of the amplification product from an amplified duplex sequencing library to PCR amplification using a primer pair in which the first primer is biotinylated and the second primer is not biotinylated, and a single-stranded Crick chain library can be generated by subjecting a second aliquot of the amplification product from an amplified duplex sequencing library to PCR amplification using a primer pair in which the first primer is not biotinylated and the second primer is biotinylated. In some cases, libraries of single-stranded Watson chain-derived sequences and libraries of single-stranded Crick chain-derived sequences can be generated as shown in Figures 2 and 3.

[0151] Using any suitable method, libraries of single-stranded Watson strand-derived sequences and libraries of single-stranded Crick strand-derived sequences can be generated from an amplified duplex sequencing library. For example, the amplification product from the amplified duplex sequencing library can be separated into a first PCR amplification and a second PCR amplification, where only one of the two primers in the PCR primer pair is tagged. For example, the first PCR amplification can use a primer pair containing a tagged primer (e.g., the first primer) and an ungagged primer (e.g., the second primer), and the second PCR amplification can use a primer pair containing an ungagged primer (e.g., the first primer) and a tagged primer (e.g., the second primer). The primer tag can be any tag that allows for the recovery of the PCR amplification product generated from the tagged primer. In some cases, the tagged primer can be a biotinylated primer, and the PCR amplification product generated from the biotinylated primer can be recovered using streptavidin. For example, libraries of single-stranded Watson strand-derived sequences and libraries of single-stranded click strand-derived sequences can be generated in PCR amplification using primer pairs containing biotinylated and non-biotinylated primers. In some cases, the tagged primers can be phosphorylated primers, and the PCR amplification product generated from the phosphorylated primers can be recovered using lambda nuclease. For example, libraries of single-stranded Watson strand-derived sequences and libraries of single-stranded click strand-derived sequences can be generated in PCR amplification using primer pairs containing phosphorylated and non-phosphorylated primers.

[0152] Using any suitable primer pair, libraries of single-stranded Watson-derived sequences and libraries of single-stranded click-derived sequences (e.g., from duplex sequencing libraries generated as described herein) can be produced. Primers may contain, but are not limited to, about 12 to about 30 nucleotides. In some cases, a primer pair may contain at least one primer that can target (e.g., bind to, as a target) an adapter sequence (e.g., an adapter sequence containing a molecular barcode) present in the amplified product generated as described herein (e.g., by ligating a 3' duplex adapter containing a first molecular barcode and a 5' adapter containing a second barcode to a nucleic acid fragment in a duplex sequencing library before amplification). Examples of primer pairs that can be used to produce libraries of single-stranded Watson-derived sequences and libraries of single-stranded click-derived sequences described herein include, but are not limited to, P5 primers and P7 primers.

[0153] Libraries of single-stranded Watson strand-derived sequences and single-stranded Crick strand-derived sequences can be generated (for example, from duplex sequencing libraries generated as described herein) using any suitable PCR conditions. PCR amplification can include a denaturation phase, an annealing phase, and an extension phase. Each phase of the amplification cycle can include any suitable conditions. In some cases, the denaturation phase can include a temperature of approximately 90°C to approximately 105°C and a time of approximately 1 second to approximately 5 minutes. For example, the denaturation phase can include a temperature of approximately 98°C and a time of approximately 10 seconds. In some cases, the annealing phase can include a temperature of approximately 50°C to approximately 72°C and a time of approximately 30 seconds to approximately 90 seconds. In some cases, the extension phase can include a temperature of approximately 55°C to approximately 80°C and a time of approximately 15 seconds to approximately 30 seconds per kb of amplicon to be generated. In some cases, the extension phase reflects the processing capacity of the polymerase used. In some cases, the annealing and extension phases can be performed in a single cycle. For example, the annealing and extension phases can include a temperature of approximately 65°C for approximately 75 seconds.

[0154] The PCR conditions used to generate libraries of single-stranded Watson strand-derived sequences and single-stranded Crick strand-derived sequences (for example, from duplex sequencing libraries generated as described herein) may include any appropriate number of PCR amplification cycles. In some cases, PCR amplification may include, but is not limited to, about 1 to about 50 cycles. For example, PCR amplification may include about 4 amplification cycles.

[0155] In some cases, if the PCR conditions include a thermoactivated polymerase, PCR amplification may also include a reprogramming step. For example, PCR amplification may include a reprogramming step before performing the PCR amplification cycle. In some cases, the reprogramming step may include a temperature of approximately 94°C to 98°C and a time of approximately 15 seconds to 1 minute. For example, the reprogramming step may include a temperature of approximately 98°C for approximately 30 seconds.

[0156] In some cases, PCR amplification may also include a retention step. For example, PCR amplification may include a retention step after the PCR amplification cycle, and optionally after any final extension step. In some cases, the retention step may include an indeterminate time at a temperature of approximately 4°C to 15°C.

[0157] Double-stranded amplification products can be separated into single-stranded amplification products using any suitable method. In some cases, double-stranded amplification products can be denatured to separate them into two single-stranded amplification products. Examples of methods that can be used to separate double-stranded amplification products into single-stranded amplification products include, but are not limited to, thermal denaturation, chemical (e.g., NaOH) denaturation, and salt denaturation.

[0158] After PCR amplification, the tagged PCR amplification product can be recovered. The tagged PCR amplification product generated using the tagged primer can be recovered using any suitable method. If the tagged primer is a biotinylated primer, the biotinylated amplification product (e.g., generated from the biotinylated primer) can be recovered using streptavidin (e.g., streptavidin-functionalized beads). For example, if an amplified duplex sequencing library is further amplified in a first PCR amplification using a primer pair containing a first biotinylated primer and a second non-biotinylated primer, and a second PCR amplification using a primer pair containing a first non-biotinylated primer and a second biotinylated primer, the biotinylated amplification product generated from the first PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first set of streptavidin-functionalized beads), and the biotinylated amplification product generated from the second PCR amplification can be bound to streptavidin-functionalized beads (e.g., a second set of streptavidin-functionalized beads), and the double-stranded amplification product can be separated (e.g., denatured) into single-stranded amplification products. In some cases, recovering the biotinylated PCR amplification product may also involve releasing the biotinylated PCR amplification product from streptavidin (e.g., streptavidin-functionalized beads). Separating the double-stranded amplification products generated by a first PCR amplification using a primer pair containing a first biotinylated primer and a second non-biotinylated primer, and a second PCR amplification using a primer pair containing a first non-biotinylated primer and a second biotinylated primer, allows the single-stranded amplification product generated from the biotinylated primer to remain bound to the streptavidin-functionalized beads, while the single-stranded amplification product generated from the non-biotinylated primer can be denatured (e.g., denatured and degraded) from the streptavidin-functionalized beads, thereby generating libraries of single-stranded Watson strand-derived sequences and libraries of single-stranded Crick strand-derived sequences in a duplex sequencing library.

[0159] If the tagged primer is a phosphorylated primer, the phosphorylated amplification product (e.g., one generated from the phosphorylated primer) can be recovered using an exonuclease (e.g., lambda exonuclease). For example, if the amplified duplex sequencing library is further amplified in a first PCR amplification using a primer pair containing a first phosphorylated primer and a second non-phosphorylated primer, and then in a second PCR amplification using a primer pair containing a first non-phosphorylated primer and a second phosphorylated primer, the double-stranded amplification product can be separated into single-stranded amplification products. Separating the double-stranded amplification products generated by a first PCR amplification using a primer pair containing a first phosphorylated primer and a second non-phosphorylated primer, and a second PCR amplification using a primer pair containing a first non-phosphorylated primer and a second phosphorylated primer, allows for the recovery of single-stranded amplification products generated from the non-phosphorylated primers, while degrading the single-stranded amplification products generated from the phosphorylated primers with lambda exonuclease allows for the generation of libraries of single-stranded Watson strand-derived sequences and libraries of single-stranded Crick strand-derived sequences in a duplex sequencing library.

[0160] Enrichment of targets In some embodiments of any one of the methods described herein, the amplicon produced by the initial amplification is enriched with one or more target polynucleotides. In some embodiments, a single-stranded DNA library is prepared from the amplicon produced by the initial amplification prior to target enrichment. Exemplary methods for producing a single-stranded DNA library are described herein.

[0161] A target region can be amplified from a library of amplification products (e.g., a duplex sequencing library, a single-stranded Watson-derived sequence library, or a single-stranded Click-derived sequence library) using any suitable method. In some cases, a target region can be amplified from a library of amplification products by subjecting the library to PCR amplification using a primer pair, wherein a primer (e.g., a first primer) can target (e.g., bind to, as a target) an adapter sequence (e.g., an adapter sequence containing a molecular barcode) present in the amplification product generated as described herein (e.g., by ligating a 3' duplex adapter containing a first molecular barcode and a 5' adapter containing a second barcode to the nucleic acid fragment in the duplex sequencing library before amplification), and a primer (e.g., a second primer) can target (e.g., bind to, as a target) the target region (e.g., the region of interest). In some cases, libraries of single-stranded Watson-derived sequences and libraries of single-stranded Click-derived sequences can be generated as shown in Figures 4 and 5. In some cases, libraries of single-stranded Watson strand-derived sequences and libraries of single-stranded Crick strand-derived sequences can be generated as described in Example 2.

[0162] In some cases, the target region can be amplified from a library of amplification products (e.g., a duplex sequencing library generated as described herein, a library of single-stranded Watson-derived sequences, or a library of single-stranded Click-derived sequences) in a single PCR amplification. For example, the target region can be amplified from a library of amplification products in a single PCR amplification using a primer pair comprising a first primer that can target an adapter sequence present in the amplification product generated as described herein (e.g., an adapter sequence containing a molecular barcode) (e.g., by ligating a 3' duplex adapter containing a first molecular barcode and a 5' adapter containing a second molecular barcode to the nucleic acid fragment in the duplex sequencing library before amplification) and a second primer that can target the target region. For example, the target region can be amplified from a library of amplification products in a single PCR amplification as shown in Figures 4, 5, 15, and 17.

[0163] In some cases, the target region can be amplified from a library of amplification products in multiple PCR amplifications (e.g., a duplex sequencing library generated as described herein, a library of single-stranded Watson strand-derived sequences, or a library of single-stranded Crick strand-derived sequences). Multiple PCR amplifications (e.g., a first PCR amplification and subsequent nested PCR amplifications) can be used to increase the specificity of the amplification of the target region. For example, in a series of PCR amplifications in which a first PCR amplification is performed, the target region is amplified from a library of amplification products, and the amplification product produced in the first PCR amplification can be supplied to a subsequent nested PCR amplification using a primer pair comprising a first primer that can target an adapter sequence (e.g., an adapter sequence containing a molecular barcode) present in the amplification product produced as described herein (e.g., by ligating a 3' duplex adapter containing a first molecular barcode and a 5' adapter containing a second barcode to a nucleic acid fragment in a duplex sequencing library before amplification), and a second primer that can target a nucleic acid sequence from the target region present in the amplification product produced in the first PCR amplification. For example, the target region can be amplified from a library of amplification products in a series of PCR amplifications shown in Figures 7, 8, 16, and 18.

[0164] Using any suitable primer pair, a target region can be amplified from a library of amplification products (e.g., a duplex sequencing library generated as described herein, a library of single-stranded Watson-derived sequences, or a library of single-stranded Crick-derived sequences). Primers may contain, but are not limited to, about 12 to about 30 nucleotides. In some cases, a primer pair may include a primer (e.g., a first primer) that can target (e.g., bind as a target) an adapter sequence (e.g., an adapter sequence containing a molecular barcode) present in the amplification product generated as described herein (e.g., by ligating a 3' duplex adapter containing a first molecular barcode and a 5' adapter containing a second barcode to the nucleic acid fragment in the duplex sequencing library before amplification) and a primer (e.g., a second primer) that can target (e.g., bind as a target) a target region (e.g., a region of interest). Examples of primers that can target adapter sequences containing molecular barcodes present in the amplification product generated as described herein (for example, by ligating a 3' duplex adapter containing a first molecular barcode and a 5' adapter containing a second barcode to a nucleic acid fragment in a duplex sequencing library before amplification) include, but are not limited to, i5 index primers and i7 index primers. Primers that can target a target region may include sequences complementary to the target region. If the target region is a nucleic acid encoding TP53, examples of primers that can target the nucleic acid encoding TP53 include, but are not limited to, TP53_342_GSP1 and TP53_GSP2. In some cases, when the target region is a nucleic acid encoding TP53, primers that target the nucleic acid encoding TP53 may be as described in Example 2.

[0165] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a duplex sequencing library generated as described herein, a library of single-stranded Watson-derived sequences, or a library of single-stranded Click-derived sequences) may contain one or more molecular barcodes.

[0166] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a duplex sequencing library generated as described herein, a library of single-stranded Watson strand-derived sequences, or a library of single-stranded Crick strand-derived sequences) may contain one or more transplant sequences (e.g., transplant sequences for next-generation sequencing).

[0167] In this context, target enrichment includes (a) selectively amplifying the amplicon of a Watson chain containing a target polynucleotide sequence using a first set of Watson target-selective primer pairs, thereby producing a targeted Watson amplification product, wherein the first set of Watson target-selective primer pairs includes (i) a first Watson target-selective primer containing a sequence complementary to the R2 sequencing primer site of a universal 3' adapter sequence, and (ii) a second Watson target-selective primer containing a target-selective sequence; and (b) selectively amplifying the amplicon of a Crick chain containing the same target polynucleotide sequence as the first set of Crick target-selective primer pairs, thereby producing a targeted Crick amplification product, wherein the first set of Crick target-selective primer pairs includes (i) a first Crick target-selective primer containing a sequence complementary to the R1 sequencing primer site of a universal 5' adapter sequence, and (ii) a second Crick target-selective primer containing the same target-selective sequence as the second Watson target-selective primer sequence.

[0168] In some embodiments, the method further includes a step of purifying the targeted Watson amplification product and the targeted click amplification product from the non-target polynucleotide. In some embodiments, the purification step includes binding the targeted Watson amplification product and the targeted click amplification product to a solid support. In some embodiments, the first Watson target-selective primer and the first click target-selective primer include a first member of the affinity binding pair, and the solid support includes a second member of the affinity binding pair. In some embodiments, the first member is biotin and the second member is streptavidin. In some embodiments, the solid support includes beads, wells, membranes, tubes, columns, plates, Sepharose, magnetic beads, or tips. In some embodiments, the method includes a step of removing polynucleotides not bound to the solid support.

[0169] In some embodiments, the method comprises the steps of (a) further amplifying a target Watson amplification product using a second set of Watson target-selective primers to construct a member of a target Watson library, wherein the second set of Watson target-selective primers comprises (i) a third Watson target-selective primer comprising a sequence complementary to the R2 sequencing primer site of a universal 3' adapter sequence, and (ii) a fourth Watson target-selective primer comprising, in the 5'→3' direction, an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide; (b) A step of further amplifying a target click amplification product using a second set of click target-selective primers to produce members of a target click library, wherein the second set of click target-selective primers includes (i) a third click target-selective primer containing a sequence complementary to the R1 sequencing primer site of a universal 3' adapter sequence, and (ii) a fourth click target-selective primer containing, in the 5'→3' direction, an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide of a fourth Watson target-selective primer.

[0170] In some embodiments, the third Watson target-selective primer and the third click target-selective primer further include a sample barcode sequence. In some embodiments, the third Watson target-selective primer further includes a first transplant sequence that enables hybridization with the first transplant primer in the sequencer, and the third click target-selective primer further includes a second transplant sequence that enables hybridization with the second transplant primer in the sequencer. In some embodiments, the fourth Watson target-selective primer further includes a second transplant sequence, and the fourth click target-selective primer further includes the first transplant sequence. In some embodiments, the first transplant sequence is a P7 sequence, and the second transplant sequence is a P5 sequence.

[0171] The amplified target regions described herein can be generated using any suitable PCR conditions (e.g., from a library of amplification products, such as a generated duplex sequencing library, a library of single-stranded Watson-derived sequences, or a library of single-stranded click-derived sequences). Exemplary PCR conditions are described herein. The PCR conditions used to generate the amplified target regions described herein (e.g., from a library of amplification products, such as a generated duplex sequencing library, a library of single-stranded Watson-derived sequences, or a library of single-stranded click-derived sequences) may include any suitable number of PCR amplification cycles. In some cases, PCR amplification may include about 1 to about 50 cycles, but is not limited thereto. For example, if the PCR amplification of the amplified target region includes a single PCR amplification, the PCR amplification may include about 18 amplification cycles. For example, if the PCR amplification of the amplified target region includes a first PCR amplification and subsequent nested PCR amplifications, the first PCR amplification may include about 18 amplification cycles, and the subsequent nested PCR amplification may include about 10 amplification cycles.

[0172] Exemplary target Any suitable target region (e.g., region of interest) can be amplified from a library of amplification products (e.g., a duplex sequencing library generated as described herein, a library of single-stranded Watson-derived sequences, or a library of single-stranded Crick-derived sequences) and evaluated for the presence or absence of one or more mutations. In some cases, the target region may be a region of nucleic acid in which one or more mutations are associated with a disease or disorder. Examples of target regions that can be amplified and evaluated for the presence or absence of one or more mutations include, but are not limited to, nucleic acids encoding tumor protein p53 (TP53), nucleic acids encoding breast cancer 1 (BRCA1), nucleic acids encoding BRCA2, nucleic acids encoding phosphatase and tensin homolog (PTEN) polypeptides, nucleic acids encoding AKT1 polypeptide, nucleic acids encoding APC polypeptide, nucleic acids encoding CDKN2A polypeptide, nucleic acids encoding EGFR polypeptide, nucleic acids encoding FBXW7 polypeptide, nucleic acids encoding GNAS polypeptide, nucleic acids encoding KRAS polypeptide, nucleic acids encoding NRAS polypeptide, nucleic acids encoding PIK3CA polypeptide, nucleic acids encoding BRAF polypeptide, nucleic acids encoding CTNNB1 polypeptide, nucleic acids encoding FGFR2 polypeptide, nucleic acids encoding HRAS polypeptide, and nucleic acids encoding PPP2R1A polypeptide. In some cases, the target region that can be amplified and evaluated for the presence or absence of one or more mutations may be the nucleic acid encoding TP53. For example, the nucleic acid encoding TP53 can be amplified and evaluated as described in Example 2.

[0173] The target region (e.g., the amplified target region) can be evaluated for the presence or absence of one or more mutations using any appropriate method. In some cases, the amplified target region can be evaluated for the presence or absence of one or more mutations using one or more sequencing methods.

[0174] Determining the sequence In some cases, one or more sequencing methods can be used to evaluate the amplified target region to determine whether the mutation is present in both the Watson and Crick strands. In some cases, sequencing reads can be used to evaluate the amplified target region for the presence or absence of one or more mutations, and these can be used to determine whether the mutation is present in both the Watson and Crick strands. Examples of sequencing methods that can be used to evaluate the amplified target region for the presence or absence of one or more mutations described herein include, but are not limited to, single-read sequencing, paired-end sequencing, NGS, and deep sequencing. In some embodiments, single-read sequencing involves sequencing over the entire length of the template to generate sequence reads. In some embodiments, sequencing includes paired-end sequencing. In some embodiments, sequencing is performed using a large-scale parallel sequencer. In some embodiments, the large-scale parallel sequencer is configured to determine sequence reads from both ends of the template polynucleotide.

[0175] Analysis of sequence reads In some embodiments, sequence reads are mapped to a reference genome.

[0176] In some embodiments, sequence reads are assigned to a UID family. The UID family may include sequence reads from amplicons originating from the original template, for example, the original double-stranded DNA fragment from a nucleic acid sample.

[0177] In some embodiments, each member of the UID family contains the same exogenous UID sequence. In some embodiments, each member of the UID family further contains the same endogenous UID sequence. Endogenous UIDs are described herein.

[0178] In some embodiments, each member of the UID family further comprises the same exogenous UID sequence and the same endogenous UID sequence. In some embodiments, the combination of the exogenous and endogenous UID sequences is unique to the UID family. In some embodiments, the combination of the exogenous and endogenous UID sequences does not exist in another UID family representing a nucleic acid sample.

[0179] The number of members in a UID family can depend on the sequencing depth. In some embodiments, the UID family may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 2 The UID family includes 00, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, or 1000 members. In some embodiments, the UID family includes approximately 2 to 1000 members, approximately 2 to 500 members, approximately 2 to 100 members, approximately 2 to 50 members, or approximately 2 to 20 members.

[0180] In some embodiments, sequence reads of individual UID families are assigned to Watson subfamilies and Crick subfamilies. In some embodiments, sequence reads of individual UID families are assigned to Watson subfamilies and Crick subfamilies based on the orientation of the insertion portion relative to the adapter sequence. In some embodiments, the orientation of the insertion portion relative to the adapter sequence is separated by how the sequence reads are aligned as “read pairs” or “mate pairs”.

[0181] In some embodiments, the assignment of sequence reads to the Watson and Crick subfamilies is based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences. In some embodiments, members of the Watson subfamily are characterized by the exogenous UID sequence being downstream of the R2 sequence and upstream of the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous UID sequence being downstream of the R1 sequence and upstream of the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by the exogenous UID sequence being more closely adjacent to the R2 sequence and less closely adjacent to the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous UID sequence being more closely adjacent to the R1 sequence and less closely adjacent to the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by their exogenous UID sequence being immediately downstream of the R2 sequence or being within 1–70, 1–60, 1–50, 1–40, 1–30, 1–20, 1–10, or 1–5 nucleotides. In some embodiments, members of the Crick subfamily are characterized by their exogenous UID sequence being immediately downstream of the R1 sequence or being within 1–70, 1–60, 1–50, 1–40, 1–30, 1–20, 1–10, or 1–5 nucleotides.

[0182] In some embodiments, the UID subfamily (e.g., the Watson subfamily and / or the Crick subfamily) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, Includes 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 members. In some embodiments, a UID subfamily (e.g., Watson subfamily and / or Crick subfamily) includes approximately 2 to 500 members, approximately 2 to 100 members, approximately 2 to 50 members, approximately 2 to 20 members, or approximately 2 to 10 members.

[0183] In some embodiments, if a threshold rate (or rate above threshold) of a member of the Watson subfamily contains a nucleotide sequence, the sequence is determined to accurately represent the Watson strand of the DNA fragment under analysis, e.g., a double-stranded DNA fragment from a nucleic acid sample. In some embodiments, if a threshold rate (or rate above threshold) of a member of the Crick subfamily contains a nucleotide sequence, the sequence is determined to accurately represent the Crick strand of the DNA fragment under analysis, e.g., a double-stranded DNA fragment from a nucleic acid sample.

[0184] Those skilled in the art can determine the threshold based, for example, on the number of subfamily members, the specific purpose of the sequencing experiment, and specific parameters of the sequencing experiment. In some embodiments, the threshold is set to 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In certain embodiments, the threshold is set to 50%. Simply as an example, in embodiments where the threshold is set to 50%, a nucleotide sequence is determined to accurately represent the Watson or Crick strand of a double-stranded DNA fragment from a DNA fragment to be analyzed, e.g., a nucleic acid sample, if at least 50% of the subfamily members contain a certain nucleotide sequence. As just another example, in a configuration where the threshold is set to 50%, if more than 50% of the subfamily members contain a certain nucleotide sequence, then that sequence is determined to accurately represent the Watson or Crick strand of a double-stranded DNA fragment from the DNA fragment being analyzed, e.g., a nucleic acid sample.

[0185] In some embodiments, the sequence that accurately represents the Watson strand of the DNA fragment under analysis is determined to have a mutation. In some embodiments, the sequence that accurately represents the Watson strand of the DNA fragment under analysis is determined to have a mutation if the sequence differs from a reference sequence lacking a certain mutation.

[0186] In some embodiments, the sequence that accurately represents the click strand of the DNA fragment under analysis is determined to have a mutation. In some embodiments, the sequence that accurately represents the click strand of the DNA fragment under analysis is determined to have a mutation if the sequence differs from a reference sequence that lacks a certain mutation.

[0187] In some embodiments, the DNA fragment to be analyzed is determined to have a mutation if the sequence that accurately represents the Watson strand and the sequence that accurately represents the Crick strand contain the same mutation.

[0188] In some cases, the location of the molecular barcode within the paired-end sequencing reads of the amplified target region can be used to identify which strand of the double-stranded nucleic acid template the amplified target region originated from. For example, if the first paired-end sequencing read of the amplified target region indicates that the molecular barcode was read last, the amplified target region can be identified as originating from the sense strand of the nucleic acid template; if the first paired-end sequencing read of the amplified target region indicates that the molecular barcode was read first, the amplified target region can be identified as originating from the antisense strand of the nucleic acid template. For example, if the second paired-end sequencing read of the amplified target region indicates that the molecular barcode was read first, the amplified target region can be identified as originating from the antisense strand of the nucleic acid template; if the second paired-end sequencing read of the amplified target region indicates that the molecular barcode was read last, the amplified target region can be identified as originating from the sense strand of the nucleic acid template. In some cases, as shown in Figures 20 and 21, paired-end sequencing can be used to distinguish between amplification products originating from the Watson strand and amplification products originating from the Crick strand.

[0189] Following sequencing of a target region (e.g., an amplified target region as described herein), the sequencing reads can be aligned with a reference genome and grouped by the molecular barcode present in each read. In some cases, sequencing reads containing the same molecular barcode and mapping to both the Watson and Crick strands of a double-stranded nucleic acid template (e.g., both the Watson and Crick strands of the target region) can be identified as having duplex support. For example, if a sequencing read showing the presence of one or more mutations within a target region contains the same molecular barcode and maps to both the Watson and Crick strands of the target region, the mutations can be identified as having duplex support.

[0190] kit Kits are also provided herein. A kit may include a set of primer pairs for the amplification of one or more target polynucleotides.

[0191] In some embodiments, the kit is (a)(i) One or more first Watson target-selective primers comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more second Watson target-selective primers, each containing a target-selective sequence. A first set of Watson target-selective primer pairs, including; (b)(i) One or more click-target selective primers containing a sequence complementary to the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) One or more second click target-selective primers, each containing the same target-selective sequence as the second Watson target-selective primer sequence. A first set of click-target-selective primer pairs, including [specific component]; (c)(i) One or more third Watson target-selective primers comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more fourth Watson target-selective primers, each comprising an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide, in the 5'→3' direction. A second set of Watson target-selective primer pairs, including; and (d)(i) One or more third click-target selective primers comprising a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more fourth click target-selective primers, each comprising an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide in the 5'→3' direction. A second set of click-target selective primers, including Includes.

[0192] The kit may include a set of primer pairs for multiplex amplification of multiple target polynucleotides.

[0193] Computer-readable media Also provided herein is a computer-readable medium containing computer-executable instructions configured to perform any of the methods described herein. The computer-readable medium may contain computer-executable instructions for analyzing sequence data from a nucleic acid sample, wherein the data is generated by the method described in any one of the claims.

[0194] Computer-readable media can perform methods for semi-automatic or automated sequence data analysis.

[0195] In some embodiments, the computer-readable medium (a) assigns sequence reads to UID families such that each member of the UID family contains the same exogenous UID sequence; (b) assigns sequence reads of each UID family to Watson subfamilies and Crick subfamilies; (c) identifies a sequence that accurately represents the Watson strand of the DNA fragment to be analyzed if a threshold percentage of members of the Watson subfamily contains a certain nucleotide sequence; and (d) identifies a sequence that accurately represents the D of the DNA fragment to be analyzed if a threshold percentage of members of the Crick subfamily contains a certain nucleotide sequence. (e) to identify a sequence that accurately represents the click chain of an NA fragment; (f) to identify a mutation when the nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the Watson chain but lacks a certain mutation in that sequence; and (g) to identify a mutation in the DNA fragment under analysis when the mutation in the nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the click chain but lacks a certain mutation in that sequence; and (g) to identify a mutation in the DNA fragment under analysis when the mutation in the nucleotide sequence that accurately represents the Watson chain and the mutation in the nucleotide sequence that accurately represents the click chain are the same mutation.

[0196] In some embodiments, the computer-readable medium includes executable code for assigning UID family members to Watson subfamilies or Click subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences. In some embodiments, the computer-executable code assigns members of the UID family to the Watson subfamilies when the exogenous UID sequence is downstream of the R2 sequence and upstream of the R1 sequence. In some embodiments, the computer-executable code assigns members of the UID family to the Click subfamilies when the exogenous UID sequence is downstream of the R1 sequence and upstream of the R2 sequence. In some embodiments, the computer-executable code assigns members of the UID family to the Watson subfamilies when the exogenous UID sequence is more closely located to the R2 sequence and less closely located to the R1 sequence. In some embodiments, the computer-executable code assigns members of the UID family to the Click subfamilies when the exogenous UID sequence is more closely located to the R1 sequence and less closely located to the R2 sequence. In some embodiments, computer executable code assigns a member of the UID family to a Watson subfamily if the exogenous UID sequence is immediately downstream of the R2 sequence or is within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides. In some embodiments, computer executable code assigns a member of the UID family to a Click subfamily if the exogenous UID sequence is immediately downstream of the R1 sequence or is within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides.

[0197] In some embodiments, the computer-readable medium includes executable code for mapping sequence reads to a reference genome. In some embodiments, the reference genome is a human reference genome.

[0198] In some embodiments, the computer-readable medium includes executable code for generating a report on disease status, prognosis, or theranosis based on the presence, absence, or amount of mutations in the sample. In some embodiments, the disease is cancer.

[0199] In some embodiments, the computer-readable medium includes executable code for generating a report of treatment options based on the presence, absence, or amount of mutations in a sample.

[0200] In some embodiments, the computer-readable medium includes executable code for data transmission over a network.

[0201] Computer system Computer systems are also provided herein. In some embodiments, the computer system includes a memory device configured to receive and store sequence data from a nucleic acid sample, wherein the data is generated by the methods described herein; and a processor communicatively coupled to the memory device, the processor including a computer-readable medium disclosed herein.

[0202] Figure 32 shows an exemplary computer system 900 adapted to enable a user to analyze nucleic acid samples according to any of the methods described herein. The system 900 includes a central computer server 901 programmed to perform the exemplary methods described herein. The server 901 includes a central processing unit (CPU, also called a "processor") 905, which can be a single-core processor, a multi-core processor, or multiple processors for parallel processing. The server 901 also includes memory 910 (e.g., random-access memory, read-only memory, flash memory); electronic storage units 915 (e.g., hard disks); a communication interface 920 (e.g., a network adapter) for communicating with one or more other systems, e.g., a sequencing system; and peripheral devices 925, which may include a cache, other memory, data storage, and / or an electronic display adapter. The memory 910, storage units 915, interface 920, and peripheral devices 925 communicate with the processor 905, e.g., a motherboard, via a communication bus (solid line). The storage units 915 can be storage units for storing data. Server 901 is functionally connected to a computer network ("Network") 930 via a communication interface 920. Network 930 can be the Internet, an intranet and / or extranet, an intranet and / or extranet communicating with the Internet, a telecommunications network, or a data network. Network 930 may, in some cases, operate as a peer-to-peer network, allowing devices connected to Server 901 to act as clients or servers.

[0203] The storage unit 915 can store files such as sequence data, barcode sequence data, or any aspect of data related to the present invention. The data storage unit 915 may be connected to data relating to the positions of cells in a virtual grid.

[0204] The server can communicate with one or more remote computer systems via network 930. The one or more remote computer systems can be, for example, a personal computer, laptop, tablet, phone, smartphone, or personal digital assistant.

[0205] In some situations, system 900 includes a single server 901. In other situations, the system includes multiple servers that communicate with each other via an intranet, extranet, and / or the Internet.

[0206] Server 901 can be adapted to store array data, data regarding nucleic acid samples, data regarding biological samples, data regarding a subject, and / or other potentially relevant information. Such information can be stored in storage unit 915 or in server 901, and such data can be transmitted via the network.

[0207] The methods described herein can be executed by a machine (e.g., a computer processor) computer-readable medium (or software) stored in an electronic storage location of server 901, such as, for example, memory 910 or electronic storage unit 915. During use, processor 905 can execute the code.

[0208] In some cases, the code can be retrieved from storage unit 915 and stored in memory 910 for easy access by processor 905. In some situations, electronic storage unit 915 can be eliminated and the machine-executable instructions are stored in memory 910. Alternatively, the code can be executed on a second computer system 940.

[0209] Aspects of the systems and methods provided herein, such as Server 901, can be embodied in programming. Various aspects of the technology can be conceived as “products” or “manufactured articles” in the form of related data, typically carried or embodied in some kind of machine (or processor) executable code and / or machine-readable medium (e.g., computer-readable medium). Machine-executable code can be stored in electronic storage units, such as memory (e.g., read-only memory, random-access memory, and flash memory) or hard disks. The “storage” type of medium can include any or all of a computer, processor or other tangible memory, or their related modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transient storage for software programming at any time. All or part of the software may sometimes be communicated over the Internet or various other telecommunication networks. Such communication may enable, for example, the loading of software from one computer or processor to another, or from a management server or host computer to an application server computer platform. Therefore, other types of media capable of carrying software elements include those used over wired and optical terrestrial communication networks, as well as over various air links, such as optical, radio, and electromagnetic waves, across physical interfaces between local devices. Physical elements carrying such waves, such as wired or wireless links, optical links, or others, may also be considered software-carrying media. Unless not limited to non-transient, terms such as tangible “storage” media, computer or machine-readable media as used herein refer to any medium involved in providing instructions to a processor for execution.

[0210] Thus, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media can include, for example, optical disks or magnetic disks such as any computer or other storage device, which can be used to execute a system. Tangible transmission media can include coaxial cables, copper wire, and optical fibers (including the wires that make up a bus within a computer system). Carrier wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during high-frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media can include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs, DVD-ROMs, any other optical media, punch cards, paper tapes, any other physical storage media with patterns of holes, RAMs, ROMs, PROMs, and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, carrier wave transmission data or instructions, cables, or links that transmit such carrier waves, or any other media that a computer may read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0211] The results of the analysis can be presented to the user with the aid of a user interface, such as a graphical user interface.

[0212] The present invention is further described in the following examples, which do not limit the scope of the invention described in the claims.

Example

[0213] Example 1: Duplex Anchor PCR Materials and Methods Preparation of a duplex anchor PCR library This protocol allows for the preparation of a duplex library using the Swift Accel-NGS 2S PCR-Free Library Kit (Cat. # 20024 and 20096) along with specific cleavage adapters and primers. In some cases, full-length P5 and P7 transplant sequences can be added to the library by a separate PCR for sequencing using an Illumina instrument.

[0214] This protocol is for PCR tubes, but it can be scaled up for PCR plates.

[0215] material: 1. Swift Accel-NGS 2S PCR-Free Library Kit (Cat. # 20024 and 20096) 2. 3' Swift N14 Adapter 1 v3A a. Ordered PAGE-purified, 1 μmol synthesis-scale lyophilized product from IDT (TIFF2026133491000006.tif12156b). c. The / 3SpC3 / group can be replaced with / 3Phos / , eliminating the need for phosphorothioate bonding, and the oligonucleotide can be purified by HPLC. 3. 3' Swift Adapter 2 v3'dT a. TIFF2026133491000007.tif5128b. / 33dT / is an uncataloged modification of IDT for 3'-deoxy T. c. Order PAGE-purified, 1 μmol synthesis-scale lyophilized products from IDT. 4. 5' Swift Adapter a. Ordered PAGE-purified, 1 μmol synthesis-scale lyophilized product from IDT (TIFF2026133491000008.tif5139b). c. The / 5SpC3 / and phosphorothioate binding are unnecessary; the oligonucleotides should be purified by HPLC. d. The truB2 reagent from the 2S Dual Indexing Kit (Cat. No. 28096) can be used as a substitute. 5. NEB Ultra II Q5 Master Mix (Cat. No. M0544L) 6. Shortened P5 primer a. TIFF2026133491000009.tif4128b. No modification is necessary; desalting from IDT, 100 μM in IDTE. 7. Shortened P7 primer a. TIFF2026133491000010.tif4128b. Modification is unnecessary, desalting from IDT, 100 μM in IDTE. 8. SPRIselect beads (Beckman Coulter, Cat. No. B23317 / B23318 / B23319) 9. 80% EtOH (approximately 2 mL per sample) 10. PCR tube strips (e.g., GeneMate VWR Cat. No. 490003-710) 11. Magnetic rack (e.g., Permagen MSRLV08) 12. USER Enzyme (NEB Cat. No. M5505L) - This is a mixture of uracil-DNA glycosylase and DNA glycosylase-lyase endonuclease VIII.

[0216] Prepare a custom adapter (this can be done once per large batch): 1. If not using Swift's truB2 reagent, resuspend the 5' Swift Adapter in Low EDTA TE (included in the Swift 2S kit) to a concentration of 42 μM. 2. Resuspend the 3' Swift N14 Adapter 1 v3A to 100 μM in Low EDTA TE (included in the Swift 2S kit). Then store it at -20 °C for later use. 3. Resuspend the 3' Swift Adapter 2 v3'dT to 100 μM in Low EDTA TE (included in the Swift 2S kit). Then store it at -20 °C for later use. 4. Anneal the 3' Swift N14 Adapter 1 v3A to the 3' Swift Adapter 2 v3'dT by mixing 100 μl of each oligo at room temperature. Label the tube with 3' Swift N^{14} v3'dT Duplex Adapter, 50 μM. The final concentration of the 3' duplex adapter is 50 μM. Incubate at room temperature for at least 5 minutes before use. Then store it at -20 °C for later use.

[0217] Technical Note: Remove the enzyme tube from -20 °C storage and place it on ice for about 10 minutes to allow the enzyme to reach 4 °C before pipetting. Pipetting the enzyme at -20 °C may result in insufficient enzyme reagent.

[0218] After thawing the reagents at 4 °C, vortex the reagents (except the enzyme) briefly to ensure they are well mixed. Centrifuge all tubes in a microcentrifuge to collect the contents and then uncap them.

[0219] Combine all reagent master mixes on ice and adjust the volume appropriately using a 5% excess volume to account for pipetting losses.

[0220] Reagents should be added to the master mix in the specific order described throughout the protocol.

[0221] Reagents can be prepared in advance (e.g., to ensure that the magnetic beads do not dry out during the size selection step).

[0222] Step 1: Repair the template 1. Transfer 11 ng of cfDNA sample to a 0.2 mL PCR tube and adjust the sample volume to a final volume of 37 μl using Low EDTA TE as needed. 2. Add 3 μl of USER Enzyme to each sample. 3. Mix by vortexing and gently centrifuge to collect all liquid at the bottom of the tube. 4. Place the sample in the thermocycler, turn off the lid heating, and run the program at 37°C for 15 minutes.

[0223] Stage 2: Endodontic Repair 1 1. Gently centrifuge the sample to precipitate it, and collect any condensed material. 2. Add 20 μl of pre-mixed Repair I master mix (see Table 1) to each sample containing 40 μl of DNA sample.

[0224] (Table 1) Endoplasty I Master Mix TIFF2026133491000011.tif301283. Mix by vortexing, gently centrifuge and settle, place in a thermocycler, and run the Repair I thermocycler program in the following order. a. 37℃, 5 minutes, turn on the lid (set the lid to 75℃) b. Heat at 65°C for 2 minutes, then turn on the lid (set the lid to 75°C). c. 37℃, 5 minutes, turn on the lid (set the lid to 75℃) 4. After the thermocycling program is complete, gently centrifuge the tube to collect the aggregates. 5. Complete the Repair I reaction by adding 120 μl (2.0×) of SPRIselect beads. Mix by vortexing. Gently centrifuge to collect the beads and incubate at room temperature for 5 minutes. 6. Collect the beads by placing the sample on a magnetic rack for 5 minutes. 7. Carefully remove the supernatant without disturbing the pellets, and discard it. 8. Add 180 μl of freshly prepared 80% ethanol solution to the sample, leaving the sample in the magnetic rack. Take care not to disturb the pellet. Incubate for 30 seconds, then carefully remove the ethanol solution using a P20 pipette. 9. Repeat the above steps for a second wash using an 80% ethanol solution. 10. Remove any remaining ethanol solution with a P20 pipette and allow the beads to dry for approximately 30 seconds. Take care not to over-dry the beads, and immediately proceed to step 1 of end repair 2.

[0225] Stage 3: Endodontic Repair 2 1. Add 50 μl of pre-mixed Repair II master mix (see Table 2) to each sample and mix until homogeneous by vortexing.

[0226] (Table 2) Endodontic Repair II Master Mix TIFF2026133491000012.tif411282. Place the sample in the thermocycler, turn off the lid heating, and run the program at 20°C for 20 minutes. 3. After the thermocycling program is complete, gently centrifuge the tube to collect the aggregates. 4. Complete the Repair 2 reaction by adding 90 μl (1.8×) of PEG / NaCl solution. Mix by vortexing. Gently centrifuge to collect the beads and incubate at room temperature for 5 minutes. 5. Collect the beads by placing the sample on a magnetic rack for 5 minutes. 6. Carefully remove the supernatant without disturbing the pellets, and discard it. 7. Add 180 μl of freshly prepared 80% ethanol solution to the sample, leaving the sample in the magnetic rack. Take care not to disturb the pellet. Incubate for 30 seconds, then carefully remove the ethanol solution using a P20 pipette. 8. Repeat the above steps for a second wash using an 80% ethanol solution. 9. Remove any remaining ethanol solution with a P20 pipette and allow the beads to dry for approximately 30 seconds. Take care not to over-dry the beads, and immediately proceed to step 1 of ligation 1.

[0227] Stage 4: Ligation 1 1. Add 30 μl of pre-mixed Ligation I Master Mix (see Table 3) to each sample and mix until homogeneous by vortexing.

[0228] (Table 3) Ligation I Master Mix TIFF2026133491000013.tif351282. Place the sample in the thermocycler, turn off the lid heating, and run the program at 25°C for 15 minutes. 3. After the thermocycling program is complete, gently centrifuge the tube to collect the aggregates. 4. Complete ligation reaction 1 by adding 36 μl (1.2×) of PEG / NaCl solution. Mix by vortexing. Gently centrifuge to collect the beads and incubate at room temperature for 5 minutes. 5. Collect the beads by placing the sample on a magnetic rack for 5 minutes. 6. Carefully remove the supernatant without disturbing the pellets, and discard it. 7. Add 180 μl of freshly prepared 80% ethanol solution to the sample, leaving the sample in the magnetic rack. Take care not to disturb the pellet. Incubate for 30 seconds, then carefully remove the ethanol solution using a P20 pipette. 8. Repeat the above steps for a second wash using an 80% ethanol solution. 9. Remove any remaining ethanol solution with a P20 pipette and allow the beads to dry for approximately 30 seconds. Take care not to over-dry the beads, and immediately proceed to step 1 of ligation 2.

[0229] Stage 5: Ligation 2 1. Add 50 μl of pre-mixed Ligation II Master Mix (see Table 4) to each sample and mix until homogeneous by vortexing.

[0230] (Table 4) Ligation II Master Mix TIFF2026133491000014.tif641282. Place the sample in the thermocycler, turn off the lid heating, and run the program at 40°C for 10 minutes. 3. After the thermocycling program is complete, gently centrifuge the tube to collect the aggregates. 4. Complete ligation reaction 1 by adding 52.5 μl (1.05×) of PEG / NaCl solution. Mix by vortexing. Gently centrifuge to collect the beads and incubate at room temperature for 5 minutes. 5. Collect the beads by placing the sample on a magnetic rack for 5 minutes. 6. Carefully remove the supernatant without disturbing the pellets, and discard it. 7. Add 180 μl of freshly prepared 80% ethanol solution to the sample, leaving the sample in the magnetic rack. Take care not to disturb the pellet. Incubate for 30 seconds, then carefully remove the ethanol solution using a P20 pipette. 8. Repeat the above steps for a second wash using an 80% ethanol solution. 9. Remove any remaining ethanol solution with a P20 pipette and allow the beads to dry for approximately 30 seconds. Take care not to over-dry the beads, and immediately resuspend them in 24 μl of Low EDTA TE. Mix by vortexing and incubate for 2 minutes. 10. Gently centrifuge to collect the beads, then collect the beads on a magnetic rack for 2 minutes.

[0231] Step 6: Amplification of the PCR library 1. Add 26 μl of the pre-mixed PCR library amplification master mix (see Table 5) to a clean tube for each sample.

[0232] (Table 5) PCR Library Amplification Master Mix TIFF2026133491000015.tif301282. Carefully transfer the supernatant containing the library after final ligation to the PCR library amplification master mix. 3. If there is any remaining post-ligation library, transfer it using a P20 pipette. Take care to transfer as much supernatant as possible. 4. Mix by vortexing, gently centrifuge and precipitate, place in a thermocycle, and run the PCR library amplification thermocycle program in the following order.

[0233] (Table 6) Example PCR library amplification thermocycling program TIFF2026133491000016.tif351285. Complete the PCR library amplification reaction by adding 90 μl (1.8×) of SPRIselect beads. Mix by vortexing. Gently centrifuge to collect the beads and incubate at room temperature for 5 minutes. 6. Collect the beads by placing the sample on a magnetic rack for 5 minutes. 7. Carefully remove the supernatant without disturbing the pellets, and discard it. 8. Add 180 μl of freshly prepared 80% ethanol solution to the sample, leaving the sample in the magnetic rack. Take care not to disturb the pellet. Incubate for 30 seconds, then carefully remove the ethanol solution using a P20 pipette. 9. Repeat the above steps for a second wash using an 80% ethanol solution. 10. Remove any remaining ethanol solution with a P20 pipette and allow the beads to dry for approximately 30 seconds. Take care not to over-dry the beads, and immediately resuspend them in 47 μl of Low EDTA TE. Mix by vortexing and incubate for 2 minutes. 11. Gently centrifuge to collect the beads, then collect the beads on a magnetic rack for 2 minutes. 12. Carefully transfer the supernatant containing the final PCR amplification library into a clean tube, taking care not to transfer any beads. 13. Analyze 1 μl of the amplified library using TapeStation. A prominent peak should be located at approximately 300 bp (180 bp + 60 bp + 59 bp), corresponding to the mononucleosomal DNA ligated with the adapter. 14. Store the library at -20°C.

[0234] Accurate and efficient detection of rare mutations using duplex-anchored PCR. A sequencing library incorporating duplex molecular barcodes was generated by sequentially ligating two adapter molecules to double-stranded input DNA. First, the input DNA was repaired at the ends via blunting and dephosphorylation reactions (Figures 9 and 10). After end repair, a 3' adapter (3' oligo #1) containing a 5' phosphate, annealed to a short oligonucleotide (3' oligo #2) with a blocked 3' group, was ligated to each 3' end of the input DNA (Figure 12). Since one of the oligonucleotides contains a 3' blocking group, only the oligonucleotide containing the 5' phosphate (3' oligo #1) covalently bound to the input DNA at its 3' end. The bound 3' oligonucleotides also contained molecular barcodes that uniquely labeled each strand (Figure 11). Next, the 3' oligo containing the 3' blocking group was degraded, and the 5' adapter oligo was ligated to each 5' end via a nick translation-like reaction. Specifically, the 5' adapter oligo anneals, leaving a gap just upstream of the molecular barcode on 3' adapter oligo #1. This gap is then filled and sealed during a nick translation-like reaction, thereby generating in-situ duplex molecular barcodes at each end of the DNA fragment (Figure 13). The resulting ligation product was purified and amplified via the initial whole-genome PCR (Figure 14).

[0235] After the initial whole-genome PCR, the product can optionally be purified to generate single-stranded (ss)NA libraries corresponding to the sense and antisense strands (Figures 2 and 3).

[0236] The amplified DNA library was enriched for the desired target using a strand-specific anchor PCR technique. This PCR enrichment utilized a single primer targeting the desired region of interest and a second primer targeting the ligated adapter sequence (Figures 4, 5, 15, 17). To increase the specificity of target enrichment, a second nested PCR can be performed using a single primer targeting the desired region of interest and a second primer targeting the ligated adapter sequence (Figures 7, 8, 16, 18). To further increase the specificity of target enrichment, the second nested PCR can be used to incorporate not only the sample barcode but also essential transplant sequences required for next-generation sequencing. The resulting library was then quantified, normalized, and sequenced.

[0237] Following sequencing, reads were aligned to the genome and grouped by molecular barcoding. Fragments containing reads with the same molecular barcode mapped to both the sense and antisense strands of the target were designed to have "duplex support." Mutations were scored only if they were present on both strands (Figures 20 and 21).

[0238] Example 2: Targeted DNA sequencing of Watson and Crick strands of DNA The identification and quantification of rare nucleic acid sequences are crucial for many areas of biology and clinical medicine. This embodiment describes a method (named SaferSeqS) that addresses this challenge by (i) efficiently introducing identical molecular barcodes into the Watson and Crick strands of a template molecule, and (ii) enriching the genomic region of interest using a novel strand-specific PCR assay. This can be applied to evaluate mutations within a single amplicon or multiple amplicons simultaneously, and can evaluate limited amounts of DNA, such as those found in plasma, reducing the error rate of existing PCR-based molecular barcoding methods by at least two orders of magnitude.

[0239] result To address the inefficiencies and introduction errors typically associated with library construction, we designed a strategy involving sequential ligation of adapter sequences to the 3' and 5' ends of DNA fragments, as well as in-situ generation of double-stranded molecular barcodes (Figure 22a). In-situ generation of molecular barcodes represents a significant innovation in novel library preparation methods. The enzyme used for in-situ generation of double-stranded molecular barcodes uniquely affixed barcodes to each DNA fragment, avoiding the need for enzymatic preparation of duplex adapters (Figure 22a, steps 2 and 3). The adapters contained a stretch of 14 random nucleotides as the exogenous molecular barcode (unique identifier sequence [UID]). The adapter-ligated fragments were subjected to a limited number of PCR cycles to create duplicate copies of the two original DNA strands (Figure 22a, step 4). For clarity, in this exemplary embodiment, the UCSC reference sequence (available from genome.ucsc.edu / ) is arbitrarily defined as the "Watson" strand, and its reverse complement is defined as the "Crick" strand.

[0240] Another innovation in this protocol is the use of a hemi-nested PCR-based method for enrichment. While hemi-nested PCR has been used previously for targeted enrichment (see, e.g., Zheng et al., 2014, Nat Med 20:1479-1484), applying it to duplex sequencing required significant modifications. Specifically, two separate PCRs were performed, one on the Watson strand and the other on the Click strand. Both PCRs employed the same gene-specific primers, but each employed different anchor primers. The pairs of PCRs from each strand could be distinguished by the orientation of the insertion site relative to the exogenous UID (Figure 22b).

[0241] Following sequencing, reads corresponding to each strand of the original DNA duplex were grouped into the Watson family and the Crick family. Members of each family had the same endogenous barcode representing the sequence at one end of the initial template fragment and the same exogenous UID introduced in situ during library construction. Mutations present in >80% of the Watson strand family were called "Watson supervariants." Mutations present in >80% of the Crick strand family were called "Crick supervariants." Mutations present in >80% of both the Watson family and the Crick family ("duplex family") with the same UID were referred to as "supercali variants," or "supercali-fragileexpialidocious variants," as used herein (Figure 22c).

[0242] As an initial demonstration of SaferSeqS, mixed experiments were performed in which DNA with known mutations was spiked into DNA from normal leukocytes at variation ratios of 10% to 0%. These mixtures were predicted to result in 15,400, 150, 15, 15, 8, or 0 supercaliforms per assay. The on-target read rate (i.e., reads consisting of the intended amplicon) was 88%, which was much higher than the rate achievable with hybrid capture-based methods (see, e.g., Samorodnitsky et al., 2015 Hum Mutat 36:903-914). Furthermore, a strong correlation between expected and observed allele frequencies was demonstrated over five orders of magnitude (Figure 23, Pearson's r>0.999, p=2.02×10⁻⁶). -12 ). No variants corresponding to the pre-specified mixed variant were observed in the DNA of normal individuals. This indicates very high specificity for the variant of interest. Specificity was determined not only for the matched bases but also for any base within the amplicon. Of the total 37,747,670 bases matched across all DNA samples, only 6 supercali variants were observed, which corresponds to 1.59 × 10⁻⁶. -7 This corresponds to the mutation frequency of the supercali variant / bp (Table 7).

[0243] (Table 7) Mutations identified in analytical sensitivity and specificity validation experiments * Coordinates refer to the Human Reference Genome hg19 release (Genome Reference Consortium GRCh37, Feb 2009). TIFF2026133491000017.tif23886

[0244] Next, we attempted to determine whether SaferSeqS could be applied to clinical samples with limited DNA quantities. For example, as little as 33 ng of DNA is often present in 10 mL of cell-free plasma DNA samples used for liquid biopsies. The vast majority of DNA template molecules in these samples are wild-type, and out of 10,000 wild-type templates present in samples from patients with low tumor burden, only one or two contain mutant templates. To detect these extremely few mutant templates with high sensitivity, the assay should efficiently recover the initiation molecules.

[0245] To evaluate SaferSeqS under these challenging conditions, cell-free plasma DNA from cancer patients was mixed with cell-free plasma DNA from normal individuals to mimic the mutation frequencies typically observed in clinical samples. In these experiments, 33 ng of each sample was assayed for one of three distinct mutations in TP53. The median on-target read percentage across 27 experimental conditions (3 TP53 amplicons × 3 samples × 3 aliquots / sample) was 80% (range: 72%–91%) (Figure 24a). The median number of duplex families (i.e., both Watson and Crick strands containing the same endogenous and exogenous barcodes) was 89% (range: 65%–102%) of the original template molecule number (Figure 24b). Furthermore, the supercaliform variant of interest was identified at the expected frequency in a total of six mixed samples (Figures 25b, d, e, Table 9). Using the previously described molecular barcoding method ("SafeSeqS" rather than "SaferSeqS"), mutations were identified in these same samples at this expected frequency (Figure 25a, b, c, Table 8). The advantage of SaferSeqS was its specificity. There were a total of 1,406 supermutations, corresponding to the 153 distinct mutations observed using the previously described method, which had an average error rate of 9.39 × 10⁻⁶. -6These reflected supermutations / bp (Figure 25a, b, c, Table 8). The vast majority of these mutations were likely polymerase errors that occurred during the initial barcoding cycle on only one of the two strands. Similarly, when considering only Watson supermutations or Crick supermutations (i.e., those where the mutation is observed on only one of the two strands, Figure 22c) and not superkali mutants, the ratio was 6.56 × 10⁶. -6 An error rate of supermutant / bp was observed (Figure 26, Table 9). In contrast, only one superkali mutant was detected among 4,947,725 base pairs matched using SaferSeqS. This corresponds to an overall mutation rate of 2.02 × 10⁻⁶. -7 This corresponds to (Table 9). These differences in specificity between SaferSeqS and previously described molecular barcoding methods (i.e., methods employing direct PCR or adapter ligation to incorporate molecular barcodes before sequencing) were highly significant (P<3.5×10⁻⁶). -10 (A two-tailed Z-test on proportions is performed to compare SaferSeqS with each of the other methods.)

[0246] (Table 8) Comparison of mutations identified by SafeSeqS and SaferSeqS (see Appendix A)

[0247] (Table 9) Comparison of mutations identified by chain agnostic molecular barcoding and SaferSeqS (see Appendix B)

[0248] To further demonstrate the clinical applicability of SaferSeqS, five cancer patients with minimal tumor burden were evaluated. In each case, mutations in the primary tumor (not plasma) were identified as described separately (Tie et al., Sci Transl Med 8:346ra392 (2016)). Plasma from these patients was divided into two equal aliquots. One aliquot was evaluated using the molecular barcoding method described separately (Kinde et al., Proc Natl Acad Sci USA 108: 9530-9535 (2011)), and the other was evaluated with SaferSeqS. In both cases, primers were designed, resulting in small amplicons targeting the mutations of interest. Evaluation using the aforementioned barcoding method revealed that the plasma samples contained all eight mutations originally identified in the primary tumor. The frequency of these mutations in the plasma varied from 0.01% to 0.1% (Figure 27, Table 10). In addition to the eight known mutations, the previously described method identified 334 distinct mutations present at a frequency of up to 0.013%, none of which were found in the primary tumors of these patients. These 334 mutations comprise 10,347 supermutations, totaling 1.23 × 10⁶ -5 This reflected the average error rate of supermutation / bp (Figure 27, Table 10). Using SaferSeqS, the eight mutations found in the primary tumors were detected in all five patients at similar frequencies to those found using previously described methods (Figure 27, Table 10). However, out of 8,707,755 matched bases, only one additional superkali mutation (rather than 334 mutations) was identified using SaferSeqS. This was 1.15 × 10⁻¹⁶. -7 This corresponds to the average error rate (Table 10). This >100-fold improvement in specificity compared to previously described molecular barcoding methods was highly significant (P<2.2 × 10⁻¹⁰). -16 (Two-tailed Z-test on proportions).

[0249] (Table 10) Mutations identified by SafeSeqS and SaferSeqS in plasma samples obtained from cancer patients (see Appendix C)

[0250] Next, we investigated whether SaferSeqS, which can be useful for various sequencing applications, can assay multiple targets simultaneously. SaferSeqS allows for two types of multiplexing: one in which multiple targets are assayed in separate PCR reactions, and the other in which multiple targets are assayed in the same PCR reaction. Since overlapping Watson and Crick strand-derived copies are generated during library amplification, the library can be split into multiple PCR reactions without negatively impacting sample recovery. For example, assuming a PCR efficiency of 70%, if the DNA library is amplified in 11 PCR cycles, up to 22 targets can be assayed separately with a loss of <10% in recovery (Figure 28). In practice, we assayed 100% or 4.4% of the library. Whether using 100% or 4.4% of the library, the on-target rate was similar, with 82% and 92% of reads properly mapped to the intended region. The number of recovered duplex families was similar, with 7,825 and 6,769 recovered using 100% and 4.4% library splits, respectively.

[0251] While the above multiplexing methods are useful for assaying a limited number of targets simultaneously, applications evaluating many genomic regions may involve multiplexing into a small number of PCR reactions. To evaluate the multiplexing ability of SaferSeqS in this regard, we designed 48 primers for matching regions of driver genes commonly mutated in cancer (Table 11). These primers were combined in two reactions, one targeting 25 regions and the other targeting 23 regions. Each of the 48 primer pairs specifically amplified their intended target (Figure 30), and 36 were considered successful in that the number of duplex families was at least 50% of the number identified in singleplex reactions. Of these 36, the median on-target rate for Watson-derived reads was 95% (range: 39%–97%), and the median on-target rate for Crick-derived reads was 95% (range: 39%–98%). Most importantly, the targets showed relatively uniform recovery rates of input molecules, with a coefficient of variation of only 17% (Figure 29). The lengths of the sequenced amplicons (median 77 bp, interquartile range: 71–83 bp) were also similar across all amplicons, consistent with the initial size of cell-free plasma DNA, which is approximately 167 bp ± 10.4 bp (Figure 29).

[0252] (Table 11) Multiplex panel composition and GSP primer sequence * Coordinates refer to the human reference genome hg19 release (Genome Reference Consortium GRCh37, Feb 2009). The asterisk in the primer sequence indicates a phosphorothioate bond. TIFF2026133491000018.tif236148TIFF2026133491000019.tif236162TIFF2026133491000020.tif236165TIFF2026133491000021.tif23635

[0253] Multiple amplicons can be evaluated using two exemplary methods. The first involves parallel amplicon-specific PCR in different wells. For liquid biopsies to monitor disease recurrence, where only a few driver gene mutations are typically observed, this strategy can be easily applied without concerns regarding cross-hybridization between primers or other challenges commonly encountered in multiplex PCR reactions. For other applications of liquid biopsies, such as screening, where the mutation of interest is unknown, evaluation of more amplicons, e.g., multiple primer pair combinations in each PCR well, is useful. This embodiment demonstrates that at least 18 amplicons can be effectively analyzed in a single well using SaferSeqS, and that a hemi-nested PCR strategy without duplex sequencing is capable of simultaneously amplifying up to 313 amplicons.

[0254] By enabling the efficient detection and quantification of rare genetic mutations, SaferSeqS not only facilitates the development of highly sensitive and specific DNA-based molecular diagnostics, but also helps answer diverse and important fundamental chemical questions.

[0255] method Plasma and peripheral blood DNA samples DNA was purified from 10 mL of plasma using the cfPure MAX cell-free DNA extraction kit (BioChain, cat. # K5011625MA) as specified by the manufacturer. DNA was purified from peripheral WBCs using the QIAsymphony DSP DNA Midi Kit (Qiagen, cat. # 937255) as specified by the manufacturer. The purified DNA from all samples was quantified as described separately (see, for example, Douville et al., 2019 bioRxiv, 660258).

[0256] Library preparation We developed a custom library preparation workflow that efficiently recovers input DNA fragments and simultaneously incorporates double-stranded molecular barcodes. Briefly, we prepared a duplex sequencing library using cell-free DNA or peripheral WBC DNA with the Accel-NGS 2S DNA Library Kit (Swift Biosciences, cat. # 21024) with the following modifications: 1) DNA was pretreated with 3 units of USER enzyme (New England BioLabs, cat. # M5505L) at 37°C for 15 minutes to excavate uracil bases; 2) The SPRI bead / PEG NaCl ratios used after each reaction were 2.0×, 1.8×, 1.2×, and 1.05× for end repair 1, end repair 2, ligation 1, and ligation 2, respectively; 3) a custom 50 μM 3' adapter (Table 12) was used instead of reagent Y2; 4) a custom 42 μM 5' adapter (Table 12) was used instead of reagent B2. Next, the library was PCR-amplified in 50 μL of reactant using primers targeting the ligated adapter (Table 12). The reaction conditions were as follows: 1× NEBNext Ultra II Q5 master mix (New England BioLabs, cat. # M0544L), 2 μM universal forward primer, and 2 μM universal reverse primer (Table 12). The library was amplified by 5, 7, or 11 cycles of PCR, depending on how many experiments were planned according to the following protocol: cycles of 30 seconds at 98°C, 10 seconds at 98°C, 75 seconds at 65°C, and retention at 4°C. When 5 or 7 cycles were used, the library was amplified in a single 50 μL reactant.When 11 cycles were used, the library was divided into eight aliquots, each amplified in eight 50 μL reaction mixtures supplemented with an additional 0.5 units of Q5® Hot Start High-Fidelity DNA polymerase (New England BioLabs, cat. # M0493L), 1 μL of 10 mM dNTPs (New England BioLabs, cat. # N0447L), and 0.4 μL of 25 mM MgCl2 solution (New England BioLabs, cat. # B9021S). The products were purified using 1.8 × SPRI beads (Beckman Coulter cat. # B23317) and eluted in EB buffer (Qiagen).

[0257] (Table 12) Oligonucleotides for library construction, chain-specific PCR assay, and sequencing The asterisks within the primer sequence indicate phosphorothioate bonds, and other custom modifications are shown in the "Notes" column. TIFF2026133491000022.tif237168TIFF2026133491000023.tif23793

[0258] Building a Library To address the inefficiencies associated with library construction, we designed a strategy involving the sequential ligation of adapter sequences to the 3' and 5' ends of DNA fragments, as well as in-situ generation of double-stranded molecular barcodes (Figure 22a). After dephosphorylation and repair of the DNA ends (Figure 22a, step 1), the adapter was ligated to the 3' end of the DNA fragment (Figure 22a, step 2). The adapter was a partially double-stranded DNA fragment with terminal modifications that selectively ligated to the 3' DNA end, preventing the formation of adapter dimers. Specifically, this adapter consisted of one oligonucleotide containing a 5' phosphate terminal modification (Table 12, 3' N14 adapter oligo #1), which was hybridized with another oligonucleotide containing a 3' blocking group and deoxyuridine as a substitute for deoxythymidine (Table 12, 3' N14 adapter oligo #2). This design allowed for the use of a high concentration of adapter in the ligation reaction, which facilitated efficient binding to the 3' end without the risk of significant dimerization or concatemerization. Furthermore, the adapter contained a stretch of 14 random nucleotides on one of the two oligonucleotides, which damaged one strand of the duplex UID. Following the ligation of the 3' adapter, a second adapter (Table 12, 5' adapter) was ligated to the 5' DNA fragment end via a nic translation-like reaction consisting of DNA polymerase, a fixed end-specific ligase, and uracil-DNA glycosylase (Figure 22a, step 3). The coordinated action of these enzymes synthesized the complementary strand of the UID, degraded the blocking portion of the 3' adapter, and ligated the extended adapter to the 5' DNA fragment end. In-situ generation of double-stranded molecular barcodes avoided the need to enzymatically prepare duplex adapters, which have been shown to barcode each DNA fragment uniquely and adversely affect the recovery of the input DNA. Finally, the adapter-ligated fragments were subjected to a limited number of PCR cycles to create duplicate copies of the two original DNA strands (UID "family") (Figure 22a, step 4).

[0259] The effect of the number of amplification cycles and efficiency of the library. SaferSeqS parameters can be optimized by adjusting the number of PCR cycles and replication efficiency during library amplification. SaferSeqS can involve splitting duplicated Watson and Crick strand-derived copies for target enrichment into specific strand-specific PCRs. In a preferred embodiment, the required number of copies should be generated to ensure a high probability of duplex recovery. For example, assuming 100% efficiency, after one PCR cycle, each template DNA duplex is converted into two double-stranded copies (one representing each strand), with only a 25% probability that these two copies are properly distributed, resulting in one Watson strand-derived copy being split into a Watson-specific PCR and one Crick strand-derived copy into a Crick-specific PCR. Increasing the number of PCR cycles or increasing amplification efficiency generates more duplicated copies, which in turn increase the probability of recovering the original DNA duplex.

[0260] A probabilistic model was developed to estimate the number of PCR cycles and amplification efficiency required for efficient duplex recovery. This model consisted of three steps: 1) simulating the number of PCR offspring generated during library amplification; 2) randomly splitting these PCR copies into Watson and Crick strand-specific reactions; and 3) determining the duplex recovery rate, which is the ratio of the original DNA duplex having at least one Watson strand-derived copy split into the Watson strand-specific reaction and at least one Crick strand-derived copy split into the Crick strand-specific reaction.

[0261] The PCR copy numbers of the original template strands generated during each library amplification cycle follow a binomial distribution. For the first PCR cycle, the strand-specific copy number was initialized to 1. It should be noted that the first library amplification cycle initializes these numbers to 1 (not 2) because it simply functions to denature the two original template strands and convert them into physically distinct double-stranded forms. During the subsequent i-th PCR cycle, n i Each of the PCR copies replicates with probability p (i.e., amplification efficiency), and n i +Binom(n i The sum of n is equal to p. i+1 n PCR copies can be generated. This process was repeated to simulate the number of offspring generated after i PCR cycles. Formally, the total number of generated PCR copies can be expressed as follows: TIFF2026133491000024.tif15128

[0262] After library amplification, the original DNA duplexes are amplified individually, and the n of the Watson strand is as described above. i,W n copy and click chain i,C A copy was generated. i,W copy and n i,C Each copy is randomly split into Watson-specific PCR and Click-specific PCR reactions with a probability q equal to the proportion of the library used for each reaction. If the library is split into a single Watson-specific PCR and a single Click-specific PCR, q is equal to 50%. If the library is split into two Watson-specific PCRs and a Click-specific PCR, q is equal to 25%. The number of PCR copies split into appropriate strand-specific PCRs (N for the kth Watson-specific or Click-specific PCR, respectively) k,W or N k,C ) for Watson and Clickcopy respectively, n i,W or n i,CThis is obtained from a binomial distribution with n "trials" and a success rate q. Therefore, the probability of splitting at least one Watson-derived PCR copy into the k-th Watson-specific PCR reaction is: The filename is TIFF2026133491000025.tif6128.

[0263] Similarly, the probability of splitting at least one click-derived PCR copy into the k-th click-specific PCR reaction is: The filename is TIFF2026133491000026.tif6128.

[0264] Both strands of the original DNA duplex are N k,W and N k,C Recovery is possible only if the value is greater than 0. Since the PCR offspring splitting is independent, therefore the probability of duplex recovery is, It is predicted to be TIFF2026133491000027.tif6128.

[0265] The inventors varied the PCR efficiency from 100% to 50%, the number of library amplification cycles from 1 to 11, and the proportion of library used for each reaction from 50% to 1.4%. For each condition, the inventors performed 10,000 simulations of the above process and report the average duplex recovery rate in Figure 28.

[0266] Fragment size and recovery rate by anchored hemi-nested PCR Anchor hemi-nested PCR theoretically demonstrates a higher recovery rate of template molecules than conventional amplicon PCR. In conventional amplicon PCR, the template molecule must contain both forward and reverse primer binding sites, as well as intervening sequences that define the amplicon. In contrast, in anchor hemi-nested PCR, the template molecule only needs to have a combination of two gene-specific primer binding sites for recovery. The combined footprint of the nested gene-specific primers used in SaferSeqS is approximately 30 bp, while the amplicon lengths employed by SafeSeqS to profile cfDNA are typically 70–80 bp. Formally, assuming uniformly random start / end coordinates of fragments, the probability of recovering a template molecule of length L is: The formula is TIFF2026133491000028.tif7128 [wherein r is the amplicon length in the case of conventional PCR or the length of the combined footprint of the gene-specific primers in the case of anchored hemi-nested PCR]. Therefore, for a cell-free DNA fragment of approximately 167 bp in size, anchored hemi-nested PCR can theoretically recover about 25% more of the original template fragment than conventional amplicon PCR. Furthermore, unlike conventional amplicon PCR, which produces a predefined product, its size is indicated by the positions of the forward and reverse primers. Anchored hemi-nested produces fragments of varying lengths, with only one of the fragment ends indicated by the position of the gene-specific primer. Assuming a template molecule of length L with uniformly random start / end coordinates, the fragment lengths observed after anchored hemi-nested PCR are: TIFF2026133491000029.tif7128[r is likely the composite footprint length of the gene-specific primers].

[0267] Exemplary aspects of the SaferSeqS bioinformatics pipeline In an exemplary embodiment of the SaferSeqS bioinformatics pipeline, Watson and Crick reads for each sample were merged into a single BAM file and sorted by read name using SAMtools to facilitate mate pair extraction. A custom Python script was then used for subsequent duplex family reconstruction and identification of Watson supermutants, Crick supermutants, and superkali mutants.

[0268] First, we grouped the reads into UID families, noting which reads originated from the Watson strand and which from the Crick strand by examining the values ​​of the reads' bitwise flags (i.e., the FLAG field). Reads containing bitwise flag values ​​of 99 and 147 originated from the Watson strand, while reads containing bitwise flag values ​​of 83 and 163 originated from the Crick strand. Reads with any other bitwise flag values ​​were excluded from subsequent analysis. Bitwise flags are numerical values ​​assigned to read pairs during mapping. Their values ​​indicate how read mates align with each other in the genome. For example, if a read was mapped to a reference strand and its mate was mapped to a reverse (complementary) strand, this read pair originated from the Watson strand. Similarly, if a read was mapped to a reverse (complementary) strand and its mate was mapped to a reference strand, this read pair originated from the Crick strand.

[0269] Secondly, two additional quality control criteria were imposed during UID family grouping to improve the determination of endogenous molecular barcodes (i.e., fragment end coordinates): 1) reads with soft clipping at the 5' or 3' of the fragment end were excluded, and 2) reads had to contain the expected constant tag sequence (GCCGTCGTTTTAT; SEQ ID NO: 117) immediately following an exogenous UID with one or fewer mismatches.

[0270] Thirdly, since the number of possible exogenous UID sequences in this embodiment far exceeds the number of start template molecules, "barcode collisions" where two molecules share the same exogenous UID sequence but have different endogenous UIDs should be extremely rare. Specifically, the expected number of barcode collisions can be calculated from the classical "burst-day problem," The formula is TIFF2026133491000030.tif12128 [wherein n is equal to the number of template molecules and N is equal to the number of possible barcodes]. For 14 bp exogenous UID sequences (containing a total of 268, 435, and 456 possible sequences) and 10,000 genomic equivalents, the expected number of collisions is 0.37, or 0.0037% of the input. Therefore, for this embodiment, it was required that each exogenous UID sequence could be associated with only one endogenous UID. If an exogenous UID was associated with more than one endogenous UID, the largest family was preserved and all others were discarded.

[0271] For other experimental design parameters, non-specific exogenous UIDs may be used for assignment to UID families, and it should be noted that non-specific exogenous UIDs can be used in combination with endogenous UIDs.

[0272] Finally, since the exogenous barcodes themselves are susceptible to errors in PCR and sequencing, we corrected the errors in the UID sequences and regrouped the UID families using the UMI-tools network proximity method.

[0273] After collecting reads within the UID family, Watson super variants, Crick super variants, and superkali variants were called as described separately herein. To exclude common polymorphisms, all variants in the Genome Aggregation Database (gnomeAD) present at allele frequencies higher than 0.1% were excluded. Reads containing superkali variants were subjected to a manual final check to remove possible alignment artifacts.

[0274] Estimated nonclonal somatic cell mutation rate The DNA used in this study was obtained from a set of individuals with an average age of 30 years. As a result, the expected frequency of non-clonal somatic single base substitutions in these samples is 426 / diproid genome, or approximately 7 × 10⁶. -8 This is the mutation / bp. In this study, the inventors evaluated a total of 42,695,395 base pairs from DNA derived from healthy control subjects using SaferSeqS. Of these 42,695,395 base pairs, 12 × 10⁻¹⁶ -8 Five single-nucleotide substitution superkali variants corresponding to the mutation frequency were detected. To determine whether the observed superkali variant frequencies followed previous estimates of non-clonal somatic mutation rates in healthy blood cells, the following exact one-sided binomial p-values ​​were also calculated: TIFF2026133491000031.tif14142

[0275] Therefore, there is no statistically significant difference between the number of supercali mutations observed and the predicted number of age-related nonclonal somatic mutations originating from healthy hematopoietic stem cells.

[0276] Anchor-Hemi-Nested PCR Targeted enrichment of the region of interest was achieved using a key modification of anchor hemi-nested PCR required for duplex sequencing. During the development of this custom stand-specific assay, various reaction conditions, including cycle count, primer concentration, and polymerase formulation, were optimized. The final optimized protocol was as follows: The first round of PCR was performed in 50 μL of reactant under the following conditions: 1× NEBNext Ultra II Q5 master mix (New England BioLabs, cat. # M0544L) for Watson chain amplification, 2 μM GSP1 primer, and 2 μM P7 short-chain anchor primer. The GSP1 primer was specific to each amplicon, and the P7 short-chain anchor primer was used as the anchor primer for the Watson chain of all amplicons (Tables 11 and 12). The click chain was amplified in the same manner, in a separate well, except that the P5 short-chain primer anchor primer was substituted for the P7 short-chain primer. Note that the GSP1 primer used for Watson chain amplification was identical to the GSP1 primer used for the click chain. The only difference between Watson chain PCR and Click chain PCR was the anchor primer. Both reactants (Watson chain and Click chain) were amplified for 19 cycles according to the thermal cycling protocol described above.

[0277] For the Watson chain, a second round of PCR was performed in 50 μl of reaction using the same reaction conditions as those used for the first round of PCR. The differences were: (i) template: 1% of the product from the first anchor Watson chain PCR was used as the template (instead of the library used as the template for the first PCR), and (ii) primers: the gene-specific primer GSP2 was substituted for the GSP1 gene-specific primer, and the anchor P5 indexing primer was substituted for the P7 short-chain anchoring primer. A second round of PCR was performed for the Crick chain in the same manner, with the exceptions that (i) template: the first Crick chain PCR was used as the template, and (ii) primers: the anchor P7 indexing primer was substituted for the anchor P5 indexing primer. Both reactions (Watson chain and Crick chain) were amplified for 17 cycles according to the thermal cycling protocol described above. The sequences of the primers used for the second round of PCR are listed in Table 12. The products of the second round of PCR were pooled and purified using 1.8 × SPRI beads before sequencing.

[0278] For experiments in which multiple targets were simultaneously amplified in a single reaction, the PCR conditions were the same as those described above, with the exception that (i) each gene-specific primer was included at a final concentration of 0.25 μM, and (ii) the anchor primer was included at a final concentration of 0.25 μM / target (for example, if 25 targets were simultaneously amplified, the final concentration would be 6.25 μM).

[0279] Sequencing Library concentrations were determined as described by the manufacturer using the KAPA Library Quantification Kit (KAPA Biosystems, cat. # KK4824). Sequencing was performed using 2×75 paired-end reads with 8-nucleotide dual indexing using an Illumina MiSeq instrument. A dual-indexed PhiX control library (SeqMatic cat. # TM-502-ND) was spiked in at 25% of the total template to ensure nucleotide diversity throughout the cycle. Custom read 1, index, and read 2 sequencing primers (Table 12) were combined with standard Illumina sequencing primers at a final concentration of 1 μM.

[0280] Mutation calling and SaferSeqS analysis pipeline We performed analysis of SafeSeqS data using a custom Python script as described separately (see, for example, Kinde et al., 2011 Proc Natl Acad Sci USA 108:9530-9535). Sequencing reads underwent initial processing by extracting the first 14 nucleotides as the UID sequence and masking the adapter sequence using Picard's IlluminaBasecallsToSam (broadinstitute.github.io / picard). Reads were then mapped to the hg19 reference genome using BWA-MEM (version 0.7.17) and sorted by UID sequence using SAMtools. UID families were scored if they consisted of two or more reads and >90% of the reads were mapped to the reference genome using the expected primer sequences. "Supermutations" were identified as mutations present in >95% of mapped reads and having an average Phred score higher than 25.

[0281] A custom analysis pipeline was developed for SaferSeqS analysis. Briefly, reads were demultiplexed, and an index sequence was used to identify the strand from which the read originated. For clarity and brevity, reads originating from the Watson strand are referred to as "Watson reads," and reads originating from the Crick strand are referred to as "Crick reads." For Watson reads, the first 14 bases of read 1 were extracted as the UID sequence. Since the orientation of the insertion region is reversed in the Crick strand, for Crick reads, the first 14 bases of read 2 were extracted as the UID sequence. Adapter sequences were masked using Picard's IlluminaBasecallsToSam (broadinstitute.github.io / picard), and the template-specific portions of the resulting reads were mapped to the hg19 reference genome using BWA-MEM (version 0.7.17). Following alignment, the mapped Watson and Crick reads were merged and sorted using SAMtools.

[0282] Python scripts were used for subsequent reconstruction of the duplex family and for the identification of Watson super variants, Crick super variants, and superkali variants. After correcting errors in the PCR and sequencing of molecular barcode sequences (see, e.g., Smith et al., 2017 Genome Res 27:491-499), Watson and Crick reads belonging to the same duplex family were grouped together to reconstruct the original template molecule sequence. To exclude artifacts resulting from the end repair step of library construction, fewer than 10 bases from the 3' adapter sequence were not considered for variant analysis. Watson super variants and Crick super variants were identified as variants present in >80% of Watson or Crick reads of the duplex family, respectively. Superkali variants were defined as variants present in >80% of both the Watson and Crick families with the same UID.

[0283] statistical analysis Continuous variables were reported as medians and ranges, while categorical variables were reported as integers and percentages. All statistical tests were performed using R's stats package (version 3.5.1).

[0284] These results demonstrate that SaferSeqS can detect rare mutations with extremely high specificity. The technology is highly scalable, cost-effective, and easily automated for high throughput. SaferSeqS achieves up to 5-75 times improvement in input recovery compared to existing duplex sequencing technologies, can be applied to limited amounts of starting material, and provides a >50-fold improvement in error correction compared to standard PCR-based technologies employing molecular barcodes (Figure 23, Table 8). This also shows a >50-fold improvement in error correction compared to optimal ligation-based technologies employing only Watson or Crick supermutants rather than superkali variants (Figure 26, Table 9). Both reductions are useful for detecting mutations presenting as single units or at very low copy numbers, for example, in cancer screening and minimal residual disease settings. Finally, because it incorporates duplex sequencing, SaferSeqS is considerably more sensitive than digital droplet PCR for single amplicon analysis and, unlike digital droplet PCR, can be highly multiplexed.

[0285] Other embodiments The present invention has been described in detail, but the foregoing description is intended to be illustrative rather than limiting the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

[0286] Inclusion by reference All references, published patents, and patent applications cited within the body of this Specified Publication are incorporated herein by reference in their entirety for all purposes. Addendum A TIFF2026133491000032.tif217170TIFF2026133491000033.tif217170TIFF2026133491000034.tif217170 Addendum B TIFF2026133491000035.tif206170 Addendum C TIFF2026133491000036.tif219170TIFF2026133491000037.tif209170TIFF2026133491000038.tif209170TIFF2026133491000039.tif219170

Claims

1. a) The step of attaching a partial double-stranded 3' adapter (3'PDSA) to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, The first strand of the 3'PDSA includes a universal 3' adapter sequence in the 5'-3' direction, comprising (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site. The second chain of 3'PDSA is a step comprising (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction. b) A step of annealing the 5' adapter to the annealing site, wherein the 5' adapter includes, in the 5'→3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence that is complementary to the annealing site for the 5' adapter; c) extending a 5' adapter over an exogenous UID sequence and the first segment, thereby generating a complement of the exogenous UID sequence and a complement of the first segment, and d) The step of covalently ligating the 3' end of the complement of the first segment to the 5' ends of the Watson strand and Crick strand of the double-stranded DNA fragment, thereby generating a plurality of double-stranded DNA fragments ligated with adapters. A method that includes this.

2. The step of generating an amplicon involves amplifying the multiple double-stranded DNA fragments ligated with the adapter using a first primer complementary to the universal 3' adapter sequence and a second primer complementary to the complement of the universal 5' adapter sequence. The method according to claim 1, further comprising, wherein the amplicon comprises a plurality of double-stranded Watson templates and a plurality of double-stranded Crick templates.

3. A step of selectively amplifying the double-stranded Watson template using a first set of Watson target-selective primer pairs to produce a targeted Watson amplification product, wherein the first set of Watson target-selective primer pairs comprises (i) a first Watson target-selective primer containing a sequence complementary to a portion of the universal 3' adapter sequence, and (ii) a second Watson target-selective primer containing a target-selective sequence. The method according to claim 2, further comprising:

4. A step of selectively amplifying the double-stranded click template using a first set of click target-selective primer pairs to produce a targeted click amplification product, wherein the first set of click target-selective primer pairs includes (i) a first click target-selective primer containing a sequence complementary to a complement of a portion of the universal 5' adapter sequence, and (ii) a second click target-selective primer containing the same target-selective sequence as the second Watson target-selective primer sequence. The method according to claim 3, further comprising:

5. The method according to claim 1, further comprising the step of removing the second strand of the 3'PDSA to generate a single-stranded 3' adapter (3'SSA).

6. The method according to claim 5, wherein the step of removing the second chain occurs after step b), or before step b), or during step b).

7. The method according to claim 5, wherein the second strand comprises one or more deoxyuridines, and the step of removing the second strand of the 3'PDSA comprises contacting a 3' duplex adapter with uracil-DNA glycosylase (UDG) to degrade the second strand.

8. The method according to claim 5, wherein the step of removing the second chain is achieved by a polymerase having exonuclease activity, the polymerase extending a 5' adapter over the exogenous UID sequence and the first segment.

9. The method according to claim 2, further comprising the step of determining sequence reads of one or more amplicons.

10. The step of assigning sequence reads to a UID family, wherein each member of the UID family contains the same exogenous UID sequence. The method according to claim 9, further comprising:

11. The method according to claim 10, further comprising the step of assigning sequence reads of each UID family to Watson subfamilies and Crick subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences.

12. The method according to claim 11, further comprising the step of identifying a sequence that accurately represents the Watson strand of the DNA fragment under analysis, if at least 50% of the Watson subfamily contains a certain nucleotide sequence.

13. The method according to claim 12, further comprising the step of identifying a sequence that accurately represents the click strand of the DNA fragment to be analyzed, if at least 50% of the click subfamily contains a certain nucleotide sequence.

14. The method according to claim 12, further comprising the step of identifying a mutation that accurately represents the Watson chain when the nucleotide sequence that accurately represents the Watson chain differs from a reference sequence lacking a certain mutation in the sequence.

15. The method according to claim 14, further comprising the step of identifying a mutation when a nucleotide sequence that accurately represents a click chain differs from a reference sequence lacking a certain mutation in the sequence.

16. The method according to claim 15, further comprising the step of identifying a mutation in the DNA fragment to be analyzed when the mutation in the nucleotide sequence that accurately represents the Watson chain and the mutation in the nucleotide sequence that accurately represents the Crick chain are the same mutation.

17. The method according to claim 10, wherein each member of the UID family further comprises the same endogenous UID sequence, the endogenous UID sequence comprising the end of a double-stranded DNA fragment from a population.

18. The method according to claim 1, wherein the group of double-stranded DNA fragments has blunt ends.

19. a) A group of partial double-stranded 3' adapters (3'PDSAs) configured to ligate to the 3' ends of both the Watson strand and the Crick strand of a group of double-stranded DNA fragments, The first strand of the 3'PDSA includes a universal 3' adapter sequence in the 5'-3' direction, comprising (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site. The second chain of 3'PDSA comprises (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction; and b) A group of 5' adapters configured to anneal to the annealing site, wherein the 5' adapters include, in the 5'→3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence that is complementary to the annealing site for the 3' adapter. A system that includes this.

20. c) The population of double-stranded DNA fragments from biological samples The system according to claim 19, further comprising:

21. The system according to claim 20, wherein the group of double-stranded DNA fragments has blunt ends.

22. c) Reagents for decomposing the second strand of the 3'PDSA to produce a single-stranded 3' adapter (3'SSA) The system according to claim 19, further comprising:

23. c) A first primer complementary to the universal 3' adapter sequence, and a second primer complementary to the complement of the universal 5' adapter sequence. The system according to claim 19, further comprising:

24. c) Watson anchor primers complementary to the universal 3' adapter array, and d) Click anchor primers complementary to the complement of the universal 5' adapter array The system according to claim 19, further comprising:

25. c) (i) one or more first Watson target-selective primers comprising a sequence complementary to a portion of the universal 3' adapter sequence, and (ii) one or more second Watson target-selective primers, each comprising a target-selective sequence. A first set of Watson target-selective primer pairs including, d) (i) one or more click target-selective primers containing a sequence complementary to a complement of a portion of the universal 5' adapter sequence, and (ii) one or more second click target-selective primers, each containing the same target-selective sequence as the second Watson target-selective primer sequence. A first set of click-target-selective primer pairs including The system according to claim 19, further comprising:

26. a) i) Multiple dephosphorylated and terminally blunted double-stranded DNA fragments, each comprising a Watson strand and a Crick strand; ii) Multiple adapters, each containing A) a barcode and B) a universal 3' adapter array in the 5'→3' direction; and iii) Ligauze The steps include: forming a reaction mixture containing; b) i) The adapter is ligated to the 3' ends of the Watson and Crick chains, and ii) The adapter is not ligated to the 5' ends of either the Watson or Crick chains. The reaction mixture is incubated in such a manner to generate a double-stranded ligation product. A method that includes this.

27. The method according to claim 26, wherein each of the plurality of adapters includes a unique barcode.

28. The method according to claim 27, wherein each of the double-stranded ligation products includes a Watson chain having only one barcode and a Crick chain having only one barcode different from the barcode on the Watson chain.

29. A method for detecting the presence or absence of mutations in a target region of a double-stranded DNA template obtained from a mammalian sample, and for determining whether the mutations are present on both strands of the double-stranded DNA template, A) A step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; B) A step of amplifying a double-stranded DNA fragment, each containing a duplex molecular barcode at each end, in order to generate an amplified duplex sequencing library, wherein the amplification includes contacting the double-stranded DNA fragment, each containing a duplex molecular barcode at each end, with a universal primer pair under whole-genome PCR conditions; C) Optionally, a step of generating a Watson strand single-stranded DNA library from the amplified duplex sequencing library; D) Optionally, a step of generating a click-strand single-strand DNA library from the amplified duplex sequencing library; E) Amplifying the target region from a Watson strand DNA library using a primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; F) Amplifying the target region from a click strand DNA library using a primer pair comprising a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to the 5' adapter; G) Sequencing the amplified target region from a Watson strand DNA library in order to generate sequencing reads and to detect the presence or absence of mutations in the Watson strand of the target region; H) Sequencing the amplified target region from a click strand DNA library in order to generate sequencing reads and to detect the presence or absence of mutations in the click strand of the target region; I) The step of grouping sequencing reads by molecular barcodes present in each sequencing read in order to determine whether mutations are present on both strands of the double-stranded DNA template. A method that includes this.

30. The step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment is: i) Ligate a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 3' duplex adapter is a) A first oligonucleotide comprising a 5' phosphate, a first molecular barcode, and a 3' oligonucleotide, and annealed thereto b) A second oligonucleotide containing a degradable 3' blocking group and It contains such that the 3' oligonucleotide and the second oligonucleotide sequence are complementary; ii) Decomposing decomposable 3' blocking groups; iii) Ligate a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 5' duplex adapter comprises an oligonucleotide containing a second molecular barcode, the second molecular barcode being different from the first molecular barcode, and the 5' adapter is ligated to the double-stranded DNA fragment upstream of the first molecular barcode, leaving a single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter; and iv) Filling the single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter in order to generate a double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment. The method according to claim 29, including the method described in claim 29.

31. The step of generating a Watson strand DNA library from an amplified duplex sequencing library is i) Amplifying a first aliquot of an amplified duplex sequencing library using a primer pair consisting of a first primer and a second primer to produce a double-stranded amplification product having a tagged Watson strand, wherein the first primer is capable of hybridizing to a Watson strand and the first primer contains a tag; ii) Denature a double-stranded amplification product having a tagged Watson strand in order to generate a single-stranded tagged Watson strand and a single-stranded click strand; and iii) Recover single-stranded tagged Watson strands from the amplified duplex sequencing library to generate a Watson strand DNA library. The method according to claim 29, including the method described in claim 29.

32. The process involves obtaining a double-stranded DNA template from a mammalian sample and generating a click-strand DNA library from the amplified duplex sequencing library. i) Amplifying a second aliquot of an amplified duplex sequencing library using a primer pair comprising a first primer and a second primer to produce a double-stranded amplification product having a tagged click strand, wherein the first primer is capable of hybridizing to the click strand and the first primer contains a tag; ii) Denature a double-stranded amplification product having a tagged click strand in order to generate a single-stranded tagged click strand and a single-stranded Watson strand; and iii) Recover single-stranded tagged click strands from the amplified duplex sequencing library to generate a click strand DNA library. The method according to any one of claims 29 to 31, including the method described in any one of claims 29 to 31.

33. The method according to any one of claims 29 to 32, wherein the mammal is a human.

34. Prior to the step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment, The step of fragmenting double-stranded DNA in order to generate double-stranded DNA fragments; The step of dephosphorylating the 5' end of a double-stranded DNA fragment; and The step of smoothing the ends of a double-stranded DNA fragment. The method according to any one of claims 29 to 33, further comprising:

35. The method according to any one of claims 29 to 34, wherein ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template comprises contacting the 3' duplex adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase.

36. The method according to claim 35, wherein the ligase is a T4 DNA ligase.

37. The method according to any one of claims 29 to 36, wherein degrading a degradable 3' blocking group comprises contacting a 3' duplex adapter with uracil-DNA glycosylase (UDG).

38. The method according to any one of claims 29 to 37, wherein ligating a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template comprises contacting the 5' adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase.

39. The method according to claim 38, wherein the ligase is Escherichia coli ligase.

40. The method according to any one of claims 29 to 39, wherein filling the single-stranded nucleic acid gap between the 5' end of a double-stranded DNA fragment and its 5' adapter comprises bringing the 5' end of the double-stranded DNA fragment and its 5' adapter into contact in the presence of polymerase and dNTPs.

41. The method according to claim 40, wherein the polymerase is Taq polymerase.

42. The method according to any one of claims 29 to 31, wherein ligating 5' adapters to each 5' end of a double-stranded DNA fragment and filling the gap between the 5' ends of the double-stranded DNA fragments and the 5' adapters are performed simultaneously.

43. The method according to any one of claims 29 to 42, wherein the step of amplifying a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment to generate an amplified duplex sequencing library comprises contacting the double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment with a universal primer pair under PCR conditions.

44. The method according to claim 43, wherein amplification includes whole-genome PCR.

45. The method according to any one of claims 29 to 44, wherein the tagged primer is a biotinylated primer, and the biotinylated primer can generate biotinylated single-stranded Watson chains and biotinylated single-stranded Crick chains.

46. The method according to claim 45, wherein the denaturation step includes NaOH denaturation, thermal denaturation, or a combination of both.

47. The method according to claim 45 or claim 46, wherein the recovery step includes contacting tagged Watson chains with streptavidin-functionalized beads and contacting tagged click chains with streptavidin-functionalized beads.

48. The method according to claim 47, wherein the recovery step further comprises denaturing untagged Watson chains and denaturing untagged Watson chains.

49. The method according to claim 47 or claim 48, wherein the recovery step further comprises releasing biotinylated single-stranded Watson chains from streptavidin-functionalized beads and releasing biotinylated single-stranded click chains from streptavidin-functionalized beads.

50. The method according to any one of claims 29 to 44, wherein the tagged primer is a phosphorylated primer, and the phosphorylated primer can generate a phosphorylated single-stranded Watson chain and a phosphorylated single-stranded Crick chain.

51. The method according to claim 50, wherein the denaturing step includes lambda exonuclease digestion.

52. The method according to any one of claims 29 to 51, wherein the step of amplifying a target region from a Watson strand DNA library further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 3' duplex adapter; and the step of amplifying a target region from a Click strand DNA library further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 5' adapter.

53. The method according to any one of claims 29 to 52, wherein the sequencing step includes paired-end sequencing.

54. A method for detecting the presence or absence of mutations in a target region of a double-stranded DNA template obtained from a mammalian sample, and for determining whether the mutations are present on both strands of the double-stranded DNA template, A) A step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; B) The step of generating Watson strand DNA libraries and Crick strand DNA libraries from an amplified duplex sequencing library derived from a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; C) A step of amplifying the target region from a single-stranded Watson chain using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; D) A step of amplifying the target region from a single-stranded click strand using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 5' adapter; E) Sequencing the amplified target region from a Watson strand DNA library in order to generate sequencing reads and to detect the presence or absence of mutations in the Watson strand of the target region; F) Sequencing the amplified target region from a click strand DNA library in order to generate sequencing reads and to detect the presence or absence of mutations in the click strand of the target region; G) The step of grouping sequencing reads by molecular barcode present in each read in order to determine whether the mutation is present in both strands of the double-stranded DNA template. A method that includes this.

55. The step of generating a double-stranded DNA fragment in which the double-stranded DNA template is a genomic DNA sample and each end of the double-stranded DNA fragment has a duplex molecular barcode is: i) Ligate a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 3' duplex adapter is a) A first oligonucleotide comprising a 5' phosphate, a first molecular barcode, and a 3' oligonucleotide, and annealed thereto b) A second oligonucleotide containing a degradable 3' blocking group and It contains such that the 3' oligonucleotide and the second oligonucleotide sequence are complementary; ii) Decomposing decomposable 3' blocking groups; iii) Ligate a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 5' duplex adapter comprises an oligonucleotide containing a second molecular barcode, the second molecular barcode being different from the first molecular barcode, and the 5' adapter is ligated to the double-stranded DNA fragment upstream of the first molecular barcode, leaving a single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter; and iv) Filling the single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter in order to generate a double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment. The method according to claim 54, including the method described in claim 54.

56. The step of generating a Watson strand DNA library and a Crick strand DNA library from an amplified duplex sequencing library derived from a double-stranded DNA template, where the double-stranded DNA template is a cell-free DNA sample and each end of the double-stranded DNA fragment has a duplex molecular barcode, is as follows: i) Amplifying a double-stranded DNA fragment having a duplex molecular barcode at each end using a universal primer pair consisting of a first primer and a second primer to produce a double-stranded amplification product having biotinylated Watson strands, wherein the amplification comprises contacting the double-stranded DNA fragment containing the duplex molecular barcode at each end with the primer pair under whole-genome PCR conditions, wherein the first primer is capable of hybridizing to the Watson strand and the first primer is biotinylated; ii) Contacting a double-strand amplification product having biotinylated Watson chains with streptavidin-functionalized beads under conditions in which the biotinylated Watson chains bind to the streptavidin-functionalized beads; iii) Denaturing the double-stranded amplification product containing the biotinylated Watson chain in order to retain the single-stranded biotinylated Watson chain bound to the streptavidin-functionalized beads and to release the single-stranded click chain; iv) Collecting single-stranded click strands; v) Releasing single-stranded biotinylated Watson chains from streptavidin-functionalized beads; and vi) Collect single-stranded biotinylated Watson chains. The method according to claim 54, including the method described in claim 54.

57. The method according to any one of claims 54 to 56, wherein the double-stranded DNA template is obtained from a mammalian sample.

58. The method according to any one of claims 54 to 57, wherein the mammal is a human.

59. Prior to the step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment, The step of fragmenting double-stranded DNA in order to generate double-stranded DNA fragments; The step of dephosphorylating the 5' end of a double-stranded DNA fragment; and The step of smoothing the ends of a double-stranded DNA fragment. The method according to any one of claims 54 to 58, further comprising:

60. The method according to any one of claims 54 to 59, wherein ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template includes contacting the 3' duplex adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase.

61. The method according to claim 60, wherein the ligase is a T4 DNA ligase.

62. The method according to any one of claims 54 to 61, wherein degrading a degradable 3' blocking group comprises contacting a 3' duplex adapter with uracil-DNA glycosylase (UDG).

63. The method according to any one of claims 54 to 62, wherein ligating a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template comprises contacting the 5' adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase.

64. The method according to claim 63, wherein the ligase is Escherichia coli ligase.

65. The method according to any one of claims 54 to 64, wherein filling the single-stranded nucleic acid gap between the 5' end of a double-stranded DNA fragment and its 5' adapter comprises bringing the 5' end of the double-stranded DNA fragment and its 5' adapter into contact in the presence of polymerase and dNTPs.

66. The method according to claim 65, wherein the polymerase is Taq-B polymerase.

67. The method according to any one of claims 54 to 66, wherein ligating a 5' adapter to each 5' end of a double-stranded DNA fragment and filling the gap between the 5' ends of the double-stranded DNA fragment and the 5' adapter are performed simultaneously.

68. The method according to any one of claims 54 to 67, wherein amplification of a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment comprises contacting the double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment with a primer pair under PCR conditions.

69. The method according to claim 68, wherein amplification includes whole-genome PCR.

70. The method according to any one of claims 54 to 69, wherein the step of amplifying a target region from a Watson strand DNA library further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 3' duplex adapter; and the step of amplifying a target region from a Crick strand DNA library further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 5' adapter.

71. The method according to any one of claims 54 to 70, wherein the sequencing step includes paired-end sequencing or single-end sequencing.

72. A method for detecting the presence or absence of mutations in a target region of a double-stranded DNA template obtained from a mammalian sample, and for determining whether the mutations are present on both strands of the double-stranded DNA template, A) A step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment; B) A step of amplifying a double-stranded DNA fragment, each having a duplex molecular barcode at each end of the double-stranded DNA fragment, using a universal primer pair, the step comprising contacting the double-stranded DNA fragment, each containing a duplex molecular barcode at each end of the double-stranded DNA fragment, with the primer pair under whole-genome PCR conditions; C) A step of amplifying the target region from the Watson strand of an amplified double-stranded DNA fragment, each having a duplex molecular barcode at each end of the double-stranded DNA fragment, using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to a 3' duplex adapter; D) A step of amplifying the target region from the click strand of an amplified double-stranded DNA fragment, each having a duplex molecular barcode at each end of the double-stranded DNA fragment, using a primer pair consisting of a first primer capable of hybridizing to the target region and a second primer capable of hybridizing to the 5' adapter; E) Sequencing the amplified target region from the Watson strand to generate sequencing reads and to detect the presence or absence of mutations in the Watson strand of the target region; F) Sequencing the target region amplified from the click strand in order to generate sequencing reads and to detect the presence or absence of mutations in the click strand of the target region; G) The step of grouping sequencing reads by molecular barcode present in each read in order to determine whether the mutation is present in both strands of the double-stranded DNA template. A method that includes this.

73. The step of generating a double-stranded DNA fragment in which the double-stranded DNA template is a genomic DNA sample and each end of the double-stranded DNA fragment has a duplex molecular barcode is: i) Ligate a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 3' duplex adapter is a) A first oligonucleotide comprising a 5' phosphate, a first molecular barcode, and a 3' oligonucleotide, and annealed thereto b) A second oligonucleotide containing a degradable 3' blocking group and It contains such that the 3' oligonucleotide and the second oligonucleotide sequence are complementary; ii) Decomposing decomposable 3' blocking groups; iii) Ligate a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template, wherein the 5' duplex adapter comprises an oligonucleotide containing a second molecular barcode, the second molecular barcode being different from the first molecular barcode, and the 5' adapter is ligated to the double-stranded DNA fragment upstream of the first molecular barcode, leaving a single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter; and iv) Filling the single-stranded nucleic acid gap between the 5' end of the double-stranded DNA fragment and the 5' adapter in order to generate a double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment. The method according to claim 72, including the method described in claim 72.

74. The method according to claim 73, wherein the double-stranded DNA template is a cell-free DNA sample.

75. The method according to any one of claims 72 to 74, wherein the double-stranded DNA template is a genomic DNA sample.

76. The method according to any one of claims 72 to 75, wherein the mammal is a human.

77. Prior to the step of generating a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment, The step of fragmenting double-stranded DNA in order to generate double-stranded DNA fragments; The step of dephosphorylating the 5' end of a double-stranded DNA fragment; and The step of smoothing the ends of a double-stranded DNA fragment. The method according to any one of claims 72 to 76, further comprising:

78. The method according to any one of claims 72 to 77, wherein ligating a 3' duplex adapter to each 3' end of a double-stranded DNA fragment obtained from a double-stranded DNA template comprises contacting the 3' duplex adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase.

79. The method according to claim 50, wherein the ligase is a T4 DNA ligase.

80. The method according to any one of claims 72 to 79, wherein degrading a degradable 3' blocking group comprises contacting a 3' duplex adapter with uracil-DNA glycosylase (UDG).

81. The method according to any one of claims 72 to 80, wherein ligating a 5' adapter to each dephosphorylated 5' end of a double-stranded DNA fragment obtained from a double-stranded DNA template comprises contacting the 5' adapter with the double-stranded DNA fragment obtained from the double-stranded DNA template in the presence of a ligase.

82. The method according to claim 81, wherein the ligase is Escherichia coli ligase.

83. The method according to any one of claims 72 to 82, wherein filling the single-stranded nucleic acid gap between the 5' end of a double-stranded DNA fragment and its 5' adapter comprises bringing the 5' end of the double-stranded DNA fragment and its 5' adapter into contact in the presence of DNA polymerase and dNTPs.

84. The method according to claim 83, wherein the DNA polymerase is Taq-B polymerase.

85. The method according to any one of claims 72 to 84, wherein ligating 5' adapters to each 5' end of a double-stranded DNA fragment and filling the gap between the 5' ends of the double-stranded DNA fragment and the 5' adapters are performed simultaneously.

86. The method according to any one of claims 72 to 85, wherein amplification of a double-stranded DNA fragment having a duplex molecular barcode at each end of the double-stranded DNA fragment comprises contacting the double-stranded DNA fragment containing a duplex molecular barcode at each end of the double-stranded DNA fragment with a primer pair under PCR conditions.

87. The method according to claim 86, wherein amplification includes whole-genome PCR.

88. The method according to any one of claims 72 to 87, wherein the step of amplifying a target region from a Watson strand DNA library further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 3' duplex adapter; and the step of amplifying a target region from a Crick strand DNA library further comprises a second amplification using a second primer pair comprising a first primer capable of hybridizing to a target region and a second primer capable of hybridizing to a 5' adapter.

89. The method according to any one of claims 72 to 88, wherein the sequencing step includes paired-end sequencing.

90. a. A step of attaching a partial double-stranded 3' adapter to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments in the DNA sample to be analyzed, wherein the first strand of the partial double-stranded 3' adapter comprises a universal 3' adapter sequence in the 5'-3' direction including (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the partial double-stranded 3' adapter comprises (i) a segment complementary to the first segment and (ii) a 3' blocking group in the 5'→3' direction, and optionally the second strand is digestible; b. A step of annealing a 5' adapter to a 3' adapter via an annealing site, wherein the 5' adapter comprises (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence complementary to the annealing site for the 5' adapter, in the 5'→3' direction; c. A step in which a nick translation-like reaction is performed to extend the 5' adapter over the exogenous UID sequence of the 3' adapter and covalently ligate the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment; d. The first amplification step to amplify the adapter-ligated double-stranded DNA fragment to produce an amplicon; e. The step of determining sequence reads of one or more amplicons of one or more double-stranded DNA fragments ligated with an adapter; f. A step in which a sequence read is assigned to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; g. The step of assigning sequence reads of each UID family to the Watson subfamily and Crick subfamily based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; h. The step of identifying a sequence that accurately represents the Watson strand of the DNA fragment being analyzed, given that a threshold percentage of members of the Watson subfamily contain a certain nucleotide sequence; i. The step of identifying a sequence that accurately represents the click strand of the DNA fragment under analysis, given that a threshold percentage of members of the click subfamily contain a certain nucleotide sequence; j. The step of identifying a mutation when a nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the Watson chain but lacks a certain mutation in that sequence; k. The step of identifying a mutation when a nucleotide sequence that accurately represents a click chain differs from a reference sequence that accurately represents a click chain but lacks a certain mutation in that sequence; and l. The step of identifying the mutation in the DNA fragment being analyzed when the mutation in the nucleotide sequence that accurately represents the Watson chain and the mutation in the nucleotide sequence that accurately represents the Crick chain are the same mutation. A method that includes this.

91. The method according to claim 90, wherein each member of the UID family further comprises the same endogenous UID sequence, the endogenous UID sequence comprising the end of a double-stranded DNA fragment from a population.

92. The method according to claim 91, wherein the endogenous UID sequence including the ends of the double-stranded DNA fragment comprises at least 8, 10, or 15 bases.

93. The method according to any one of claims 90 to 92, wherein the exogenous UID sequence is unique to each double-stranded DNA fragment.

94. The method according to any one of claims 90 to 92, wherein the exogenous UID sequence is not specific to each double-stranded DNA fragment.

95. The method according to any one of claims 91 to 94, wherein each member of the UID family includes the same endogenous UID sequence and the same exogenous UID sequence.

96. The method according to any one of the claims, wherein step (d) includes PCR amplification of 11 cycles or less.

97. The method according to claim 96, wherein step (d) includes PCR amplification of 7 cycles or less.

98. The method according to claim 97, wherein step (d) includes PCR amplification of 5 cycles or less.

99. The method according to any one of the claims, wherein step (d) comprises at least one cycle of PCR amplification.

100. The method according to any one of the claims, wherein the amplicon is enriched with one or more target polynucleotides before determining the sequence read.

101. To become wealthy, a. Selective amplification of a Watson chain amplicon containing a target polynucleotide sequence using a first set of Watson target-selective primer pairs, thereby producing a target Watson amplification product, wherein the first set of Watson target-selective primer pairs is (i) A first Watson target-selective primer comprising a sequence complementary to a portion of the universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) A second Watson target-selective primer containing a target-selective sequence Including; and b. Selectively amplifying an amplicon of a click chain containing the same target polynucleotide sequence using a first set of click target-selective primer pairs, thereby producing a target click amplification product, wherein the first set of click target-selective primer pairs is (i) A first click target selective primer comprising a sequence complementary to a portion of the universal 5' adapter sequence, wherein the portion of the universal 5' adapter sequence is the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) A second Crick target-selective primer containing the same target-selective sequence as the second Watson target-selective primer sequence. including The method according to claim 100, including the method described in claim 100.

102. The method according to claim 101, comprising the step of purifying a targeted Watson amplification product and a targeted click amplification product from a non-targeted polynucleotide.

103. The method according to claim 102, wherein the purification step includes binding the target Watson amplification product and the target Crick amplification product to a solid support.

104. The method according to claim 103, wherein the first Watson target-selective primer and the first click target-selective primer comprise a first member of an affinity-binding pair, and the solid support comprises a second member of an affinity-binding pair.

105. The method according to claim 104, wherein the first member is biotin and the second member is streptavidin.

106. The method according to any one of claims 102 to 105, wherein the solid support comprises beads, wells, membranes, tubes, columns, plates, Sepharose, magnetic beads, or chips.

107. The method according to any one of claims 102 to 106, comprising the step of removing polynucleotides that are not bound to a solid support.

108. a. A step in which the target Watson amplification product is further amplified using a second set of Watson target-selective primers, thereby producing members of the target Watson library, wherein the second set of Watson target-selective primers is (i) A third Watson target-selective primer comprising a sequence complementary to a portion of the universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) A fourth Watson target-selective primer comprising an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide in the 5'→3' direction. Stages including; b. A step in which the target click amplification product is further amplified using a second set of click target-selective primers, thereby producing members of the target click library, wherein the second set of click target-selective primers is (i) A third click target selective primer comprising a sequence complementary to a portion of the universal 5' adapter sequence, wherein the portion of the universal 5' adapter sequence is the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) A fourth click target-selective primer comprising, in the 5'→3' direction, an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide of the fourth Watson target-selective primer. stages The method according to any one of claims 101 to 107, including the method described in any one of claims 101 to 107.

109. The method according to claim 108, wherein the third Watson target-selective primer and the third click target-selective primer further comprise a sample barcode sequence.

110. The method according to claim 108 or 109, wherein a third Watson target-selective primer further comprises a first transplant sequence that enables hybridization with a first transplant primer in a sequencer, and a third click target-selective primer further comprises a second transplant sequence that enables hybridization with a second transplant primer in a sequencer.

111. The method according to any one of claims 108 to 110, wherein a fourth Watson target-selective primer further comprises a second transplant sequence, and a fourth click target-selective primer further comprises a first transplant sequence.

112. The method according to claim 110 or 111, wherein the first transplant sequence is a P7 sequence and the second transplant sequence is a P5 sequence.

113. The method according to any one of claims 101 to 112, wherein the members of the target Watson library and the members of the target Crick library correspond to at least 50% of the target polynucleotides in the double-stranded DNA fragment population.

114. The method according to claim 113, wherein the members of the target Watson library and the members of the target Crick library correspond to at least 70% of the target polynucleotides in the double-stranded DNA fragment population.

115. The method according to claim 114, wherein the members of the target Watson library and the members of the target Crick library correspond to at least 80% of the target polynucleotides in the double-stranded DNA fragment population.

116. The method according to claim 115, wherein members of the target Watson library and members of the target Crick library correspond to at least 90% of the target polynucleotides in the double-stranded DNA fragment population.

117. The method according to any one of claims 101 to 112, wherein the members of the target Watson library and the members of the target Crick library represent at least 50% of the total DNA fragment population.

118. The method according to claim 117, wherein the members of the target Watson library and the members of the target Crick library represent at least 70% of the total DNA fragment population.

119. The method according to claim 118, wherein the members of the target Watson library and the members of the target Crick library represent at least 80% of the total DNA fragment population.

120. The method according to claim 119, wherein the members of the target Watson library and the members of the target Crick library represent at least 90% of the total DNA fragment population.

121. a. A step of attaching an adapter to a population of double-stranded DNA fragments in a DNA sample to be analyzed, wherein the adapter comprises a double-stranded portion containing an exogenous UID and a fork-shaped portion containing (i) a single-stranded 3' adapter sequence containing an R2 sequencing primer site and (ii) a single-stranded 5' adapter sequence containing an R1 sequencing primer site; b. The initial amplification step to amplify the double-stranded DNA fragment ligated with the adapter to produce an amplicon; c. A step of selectively amplifying the amplicon of a Watson chain containing a target polynucleotide sequence using a first set of Watson target-selective primer pairs, thereby producing a target Watson amplification product, wherein the first set of Watson target-selective primer pairs is (i) A first Watson target-selective primer comprising a sequence complementary to a portion of the universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) A second Watson target-selective primer containing a target-selective sequence stages d. A step of selectively amplifying an amplicon of a click chain containing the same target polynucleotide sequence using a first set of click target-selective primer pairs, thereby producing a target click amplification product, wherein the first set of click target-selective primer pairs is A first click-target selective primer comprising a sequence complementary to a portion of a universal 5' adapter sequence, wherein the portion of the universal 5' adapter sequence is the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) A second click target-selective primer containing the same target-selective sequence as the second click target-selective primer sequence. Stages including; e. The step of determining the sequence reads of the target Watson amplification product and the target click amplification product; f. A step in which a sequence read is assigned to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; g. The step of assigning sequence reads of each UID family to the Watson subfamily and Crick subfamily based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; h. The step of identifying a sequence that accurately represents the Watson strand of the DNA fragment being analyzed, given that a threshold percentage of members of the Watson family contain a certain nucleotide sequence; i. The step of identifying a sequence that accurately represents the click strand of the DNA fragment under analysis if a threshold percentage of members of the click family contain a certain nucleotide sequence; and j. The step of identifying mutations in the DNA fragment under analysis when both the nucleotide sequence accurately representing the Watson chain and the nucleotide sequence accurately representing the Crick chain contain the same mutations. A method that includes this.

122. The method according to claim 121, further comprising the step of purifying the target Watson amplification product and the target click amplification product from a non-target polynucleotide.

123. The method according to claim 122, wherein the purification step includes binding the target Watson amplification product and the target Crick amplification product to a solid support.

124. The method according to claim 123, wherein the first Watson target-selective primer and the first click target-selective primer comprise a first member of an affinity-binding pair, and the solid support comprises a second member of an affinity-binding pair.

125. The method according to claim 124, wherein the first member is biotin and the second member is streptavidin.

126. The method according to any one of claims 122 to 125, wherein the solid support comprises beads, wells, membranes, tubes, columns, plates, Sepharose, magnetic beads, or chips.

127. The method according to any one of claims 122 to 126, comprising the step of removing polynucleotides that are not bound to a solid support.

128. a. A step in which the target Watson amplification product is further amplified using a second set of Watson target-selective primers, thereby producing members of the target Watson library, wherein the second set of Watson target-selective primers is (i) A third Watson target-selective primer containing a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) A fourth Watson target-selective primer comprising an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide in the 5'→3' direction. Stages including; b. A step in which the target click amplification product is further amplified using a second set of click target-selective primers, thereby producing members of the target click library, wherein the second set of click target-selective primers is (i) A third click-target selective primer containing a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence, and (ii) A fourth click target-selective primer comprising, in the 5'→3' direction, an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide of the fourth Watson target-selective primer. stages The method according to any one of claims 121 to 127, including the method described in any one of claims 121 to 127.

129. The method according to claim 128, wherein the third Watson target-selective primer and the third click target-selective primer further comprise a sample barcode sequence.

130. The method according to claim 128 or 129, wherein a third Watson target-selective primer further comprises a first transplant sequence that enables hybridization with a first transplant primer in a sequencer, and the third click target-selective primer further comprises a second transplant sequence that enables hybridization with a second transplant primer in a sequencer.

131. The method according to any one of claims 128 to 130, wherein a fourth Watson target-selective primer further comprises a second transplant sequence, and a fourth click target-selective primer further comprises a first transplant sequence.

132. The method according to claim 130 or 131, wherein the first transplant sequence is a P7 sequence and the second transplant sequence is a P5 sequence.

133. The method according to any one of claims 121 to 131, wherein the binding step includes binding an adapter having an A-tail to a population of double-stranded DNA fragments.

134. The method according to claim 133, wherein the binding step includes binding an adapter having an A-tail to both ends of a DNA fragment in a population.

135. The step of combining them, a. A partial double-stranded 3' adapter is attached to the 3' ends of both the Watson strand and the Crick strand of a population of double-stranded DNA fragments, wherein the first strand of the partial double-stranded 3' adapter comprises a universal 3' adapter sequence in the 5'-3' direction including (i) a first segment, (ii) optionally an exogenous UID sequence, (iii) an annealing site for the 5' adapter, and (iv) an R2 sequencing primer site, and the second strand of the partial double-stranded 3' adapter comprises in the 5'→3' direction including (i) a segment complementary to the first segment and (ii) a 3' blocking group, and optionally the second strand is gradable; and b. Annealing a 5' adapter to a 3' adapter via an annealing site, wherein the 5' adapter includes, in the 5'→3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence complementary to the annealing site for the 5' adapter; and c. Perform a nick translation-like reaction to extend the 5' adapter across the 3' adapter and covalently ligate the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment. The method according to any one of claims 121 to 131, including the method described in any one of claims 121 to 131.

136. The method according to any one of claims 121 to 135, wherein the UID sequence includes an endogenous UID sequence comprising the ends of double-stranded DNA fragments from a population.

137. The method according to claim 136, wherein the endogenous UID sequence including the ends of the double-stranded DNA fragment comprises at least 8, 10, or 15 bases.

138. The method according to any one of claims 121 to 136, wherein the exogenous UID sequence is unique to each double-stranded DNA fragment.

139. The method according to any one of claims 121 to 136, wherein the exogenous UID sequence is not specific to each double-stranded DNA fragment.

140. The method according to any one of claims 136 to 139, wherein each member of the UID family includes the same endogenous UID sequence and the same exogenous UID sequence.

141. The method according to any one of claims 121 to 140, wherein the amplification of a double-stranded DNA fragment ligated with an adapter to produce an amplicon comprises 11 or fewer PCR amplification cycles.

142. The method according to claim 141, wherein the amplification of a double-stranded DNA fragment ligated with an adapter to produce an amplicon comprises PCR amplification of 7 cycles or less.

143. The method according to claim 142, wherein the amplification of a double-stranded DNA fragment ligated with an adapter to produce an amplicon comprises PCR amplification of 5 cycles or less.

144. The method according to any one of the claims, wherein the amplification of a double-stranded DNA fragment ligated with an adapter to produce an amplicon comprises at least one cycle of PCR amplification.

145. The method according to any one of claims 121 to 143, wherein the members of the target Watson library and the members of the target Crick library correspond to at least 50% of the target polynucleotides in the double-stranded DNA fragment population.

146. The method according to claim 145, wherein members of the target Watson library and members of the target Crick library correspond to at least 70% of the target polynucleotides in the double-stranded DNA fragment population.

147. The method according to claim 146, wherein members of the target Watson library and members of the target Crick library correspond to at least 80% of the target polynucleotides in the double-stranded DNA fragment population.

148. The method according to claim 147, wherein members of the target Watson library and members of the target Crick library correspond to at least 90% of the target polynucleotides in the double-stranded DNA fragment population.

149. The method according to any one of claims 121 to 143, wherein the members of the target Watson library and the members of the target Crick library represent at least 50% of the total DNA fragment population.

150. The method according to claim 149, wherein the members of the target Watson library and the members of the target Crick library represent at least 70% of the total DNA fragment population.

151. The method according to claim 150, wherein the members of the target Watson library and the members of the target Crick library represent at least 80% of the total DNA fragment population.

152. The method according to claim 151, wherein the members of the target Watson library and the members of the target Crick library represent at least 90% of the total DNA fragment population.

153. The method according to any one of the claims, wherein the determination of sequence reads enables the determination of sequences at both ends of a template molecule.

154. The method according to claim 153, wherein the determination of both ends of the template molecule includes paired-end sequencing.

155. The method according to any one of the claims, wherein the determination of the sequence reads includes single-read sequencing over the length of a template for generating the sequence reads.

156. The method according to any one of the claims, wherein the determination of sequence reads includes sequencing by a large-scale parallel sequencer.

157. The method according to claim 156, wherein a large-scale parallel sequencer is configured to determine sequence reads from both ends of a template polynucleotide.

158. The method according to any one of the claims, wherein the double-stranded DNA fragment population comprises one or more fragments having a length of approximately 50 to 600 nt.

159. The method according to any one of the claims, wherein the double-stranded DNA fragment population comprises one or more fragments having a length of less than 2000 nt, less than 1000 nt, less than 500 nt, less than 400 nt, less than 300 nt, or less than 250 nt.

160. The method according to any one of claims 101 to 159, further comprising the step of preparing single-stranded (ss) DNA libraries corresponding to the sense and antisense strands of the amplicon after the initial amplification and before selective amplification.

161. Preparation of ssDNA libraries a. An amplification reaction is carried out using two primers, thereby producing an amplified product containing a chain containing the first member of the affinity-binding pair and a chain not containing the first member of the affinity-binding pair, wherein only one of the two primers contains the first member of the affinity-binding pair; b. Contacting the amplification product with a solid support, wherein the solid support contains a second member of the affinity bond; c. Denature the amplification product in order to separate the chain containing the first member of the affinity bond from the chain not containing the first member of the affinity bond; and d. Purify the isolated chain containing the first member of the affinity binding pair and the isolated chain not containing the first member of the affinity binding pair. The method according to claim 160, including the method described in claim 160.

162. The method according to claim 161, wherein the first member of the affinity binding pair is biotin and the second member of the affinity binding pair is streptavidin.

163. Preparation of ssDNA libraries a. Dividing an amplicon into two amplification reactions, thereby producing an amplified product containing a phosphorylated chain and a non-phosphorylated chain, wherein each amplification reaction utilizes a forward primer and a reverse primer, and only one of the two primers is phosphorylated; b. Contact the amplified product with an exonuclease that selectively digests the chain containing the 5' phosphate. The method according to claim 160, including the method described in claim 160.

164. a. In the first amplification reaction, the forward primer is phosphorylated and the reverse primer is not phosphorylated; b. In the second amplification reaction, the reverse primer is phosphorylated and the forward primer is not phosphorylated. The method according to claim 163.

165. The method according to claim 163, wherein the exonuclease is lambda exonuclease.

166. The method according to any one of claims 163 to 165, wherein phosphorylation occurs at the 5' site.

167. The initial amplification was a. Amplifying using a primer pair to produce an amplified product comprising a chain containing the first member of the affinity-binding pair and a chain not containing the first member of the affinity-binding pair, wherein only one of the two primers in the primer pair contains the first member of the affinity-binding pair; b. Contacting the amplification product with a solid support, wherein the solid support contains a second member of the affinity bond; c. Denature the amplification product in order to separate the chain containing the first member of the affinity bond from the chain not containing the first member of the affinity bond; and d. Purify the isolated chain containing the first member of the affinity binding pair and the isolated chain not containing the first member of the affinity binding pair. The method according to any one of claims 90 to 153, including the method described in any one of claims 90 to 153.

168. The method according to claim 167, wherein the first member of the affinity binding pair is biotin and the second member of the affinity binding pair is streptavidin.

169. The method according to any one of the claims, wherein a sequence read of the UID family is assigned to the Watson subfamily when the exogenous UID sequence is downstream of the R2 sequence and upstream of the R1 sequence.

170. The method according to any one of the claims, wherein sequence reads of a UID family are assigned to a click subfamily when the exogenous UID sequence is downstream of the R1 sequence and upstream of the R2 sequence.

171. The method according to any one of the claims, wherein a sequence read of the UID family is assigned to the Watson subfamily when the exogenous UID sequence is more closely adjacent to the R2 sequence and less closely adjacent to the R1 sequence.

172. The method according to any one of the claims, wherein a sequence read of the UID family is assigned to a click subfamily when the exogenous UID sequence is more closely adjacent to the R1 sequence and less closely adjacent to the R2 sequence.

173. The method according to any one of the claims, wherein a sequence read of the UID family is assigned to the Watson subfamily if the exogenous UID sequence is immediately downstream of the R2 sequence or is within 1 to 300, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, or 1 to 5 nucleotides.

174. The method according to any one of the claims, wherein a sequence read of the UID family is assigned to a click subfamily if the exogenous UID sequence is immediately downstream of the R1 sequence or is within 1 to 300, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, or 1 to 5 nucleotides.

175. The method according to any one of the claims, wherein the double-stranded DNA fragment population is derived from a biological sample.

176. The method according to claim 175, wherein a biological sample is obtained from the subject.

177. The method according to claim 176, wherein the subject is a human subject.

178. The method according to any one of claims 175 to 177, wherein the biological sample is a liquid sample.

179. The method according to claim 178, wherein the liquid sample is selected from whole blood, plasma, serum, sputum, urine, sweat, tears, ascites, semen, and bronchoalveolar lavage fluid.

180. The method according to claim 178, wherein the liquid sample is a cell-free sample or an essentially cell-free sample.

181. The method according to any one of claims 175 to 177, wherein the biological sample is a solid biological sample.

182. The method according to claim 181, wherein the solid biological sample is a tumor sample.

183. The method according to any one of the claims, wherein the identified mutation is present at a frequency of 0.1% or less in a population of double-stranded DNA fragments.

184. The method according to claim 183, wherein the identified mutation is present at a frequency of 0.1% to 0.00001% in a population of double-stranded DNA fragments.

185. The method according to claim 183, wherein the identified mutation is present at a frequency of 0.1% to 0.01% in the double-stranded DNA fragment population.

186. The method according to any one of the claims, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 50% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed.

187. The method according to claim 186, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 70% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed.

188. The method according to claim 187, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 80% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed.

189. The method according to claim 188, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 90% of the double-stranded DNA fragments containing the target polynucleotide in the DNA sample to be analyzed.

190. The method according to any one of the claims, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 50% of the double-stranded DNA fragments in the DNA sample to be analyzed.

191. The method according to any one of the claims, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 70% of the double-stranded DNA fragments in the DNA sample to be analyzed.

192. The method according to any one of the claims, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 80% of the double-stranded DNA fragments in the DNA sample to be analyzed.

193. The method according to any one of the claims, wherein the step of determining sequence reads includes determining sequence reads from both Watson strands and Crick strands of at least 90% of the double-stranded DNA fragments in the DNA sample to be analyzed.

194. The method according to any one of the claims, wherein the error rate associated with identifying one or more mutations in the DNA fragment to be analyzed by the method according to any one of the claims is reduced to at least 1 / 2, 1 / 4, 1 / 5, 1 / 10, 1 / 20, 1 / 30, 1 / 40, 1 / 50, 1 / 60, 1 / 70, 1 / 80, 1 / 90, or 1 / 100 compared to alternative methods for identifying mutations that do not require the detection of mutations in both the Watson strand and the Crick strand of the DNA fragment to be analyzed.

195. The method according to claim 194, wherein the alternative method includes standard molecular barcoding or standard PCR-based molecular barcoding.

196. Alternative methods are a. A step of attaching an adapter to a population of double-stranded DNA fragments in a DNA sample to be analyzed, wherein the adapter contains an intrinsic exogenous UID; b. The initial amplification step to amplify the double-stranded DNA fragment ligated with the adapter to produce an amplicon; c. The step of determining sequence reads of one or more amplicons of one or more double-stranded DNA fragments ligated with an adapter; d. A step in which a sequence read is assigned to a UID family, wherein each member of the UID family contains the same exogenous UID sequence; e. The step of identifying a nucleotide sequence that accurately represents the DNA fragment to be analyzed, given that a threshold percentage of members of the UID family contain a certain nucleotide sequence; and f. The step of identifying a mutation when the sequence identified as accurately representing the DNA fragment under analysis differs from a reference sequence lacking a certain mutation in the DNA fragment under analysis. The method according to claim 195, including the method described in claim 195.

197. The error rate associated with identifying one or more mutations in the DNA fragment to be analyzed by the method according to any one of the above claims is 1 × 10⁻⁶ -2 Below, 1 x 10 -3 Below, 1 x 10 -4 Below, 1 x 10 -5 Below, 1 x 10 -6 Below, 5 x 10 -6 The following, or 1 × 10 -7 The method according to any one of the above claims, which is as follows:

198. A computer-readable medium comprising computer-executable instructions for analyzing sequence read data from a nucleic acid sample, wherein the data is generated by the method described in any one of the claims.

199. a. Assigning sequence reads to a UID family (where each member of the UID family contains the same exogenous UID sequence); b. For assigning sequence reads of each UID family to Watson subfamilies and Crick subfamilies based on the spatial relationship between the exogenous UID sequence and the R1 and R2 read sequences; c. To accurately represent the Watson strand of the DNA fragment being analyzed and to identify the sequence if a threshold percentage of members of the Watson subfamily contains a certain nucleotide sequence; d. To accurately represent the click strand of the DNA fragment under analysis and identify the sequence if a threshold percentage of members of the click subfamily contains a certain nucleotide sequence; e. For identifying a mutation when a nucleotide sequence that accurately represents the Watson chain differs from a reference sequence that accurately represents the Watson chain but lacks a certain mutation in that sequence; f. In cases where a nucleotide sequence that accurately represents a click chain differs from a reference sequence that accurately represents a click chain but lacks a certain mutation, for the purpose of identifying that mutation; g. When the mutation in the nucleotide sequence that accurately represents the Watson chain and the mutation in the nucleotide sequence that accurately represents the Crick chain are the same mutation, to identify the mutation in the DNA fragment being analyzed. A computer-readable medium according to claim 198, including executable instructions.

200. A computer-readable medium according to claim 199, comprising assigning a UID family member to the Watson subfamily when the exogenous UID sequence is immediately downstream of the R2 sequencing primer binding site or within 1 to 300 nucleotides.

201. A computer-readable medium according to any one of the claims, comprising assigning a UID family member to a click subfamily when the exogenous UID sequence is immediately downstream of the R1 sequencing primer binding site or within 1 to 300 nucleotides.

202. A computer-readable medium according to any one of the claims, comprising mapping sequence reads to a reference genome.

203. The computer-readable medium according to claim 202, wherein the reference genome is a human reference genome.

204. A computer-readable medium according to any one of the claims, further comprising computer-executable instructions for generating a report of treatment options based on the presence, absence, or amount of mutations in a sample.

205. A computer-readable medium according to any one of the claims, further comprising computer-executable code that enables data transmission over a network.

206. a. A storage device configured to receive sequence data from a nucleic acid sample, wherein the data is generated by the method described in any one of the claims; b. A processor communicatively coupled to the storage device, comprising a computer-readable medium according to any of the claims. A computer system including a computer system.

207. The computer system according to claim 206, further comprising a sequencing system configured to communicate data to a storage device.

208. The computer system according to any one of the claims, further comprising a user interface configured to communicate or to display reports to a user.

209. The computer system according to any one of the claims, further comprising a digital processor configured to transmit the results of data analysis over a network.

210. a. A collection of double-stranded DNA fragments from biological samples; b. A group of 3' adapters according to any one of the claims above; c. A group of 5' adapters according to any one of the claims; d. Reagents for performing a nick translation-like reaction; e. Reagents for enriching amplicons with respect to one or more target polynucleotides; and f. Sequencing system A system that includes this.

211. The system according to claim 210, further comprising the computer system according to any one of the preceding claims.

212. a. (i) One or more first Watson target-selective primers comprising a sequence complementary to a portion of the universal 3' adapter sequence, wherein the portion of the universal 3' adapter sequence is optionally the R2 sequencing primer site of the universal 3' adapter sequence; and (ii) One or more second Watson target-selective primers, each comprising a target-selective sequence. A first set of Watson target-selective primer pairs including; b. (i) One or more click target-selective primers comprising a sequence complementary to a portion of the universal 5' adapter sequence, wherein the portion of the universal 5' adapter sequence is optionally the R1 sequencing primer site of the universal 5' adapter sequence; and (ii) One or more second click target-selective primers, each comprising the same target-selective sequence as the second Watson target-selective primer sequence. A first set of click-target-selective primer pairs including; c. (i) One or more third Watson target-selective primers comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more fourth Watson target-selective primers, each comprising, in the 5'→3' direction, an R1 sequencing primer site and a target-selective sequence selective to the same target polynucleotide. A second set of Watson target-selective primer pairs including; and d. (i) One or more third click-target selective primers containing a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence, and (ii) One or more fourth click target-selective primers, each comprising an R2 sequencing primer site and a target-selective sequence selective to the same target polynucleotide, in the 5'→3' direction. A second set of click-target selective primers, including A kit that includes this.