Methods and compositions for DNA library preparation and analysis
By using linear end adaptors to connect with the target DNA template to form a circular construct, a DNA template of double length is generated, solving the problems of inconsistent target sequence replication and size control in existing technologies, and realizing efficient and accurate nucleic acid sequencing and bioinformatics analysis.
Patent Information
- Application Number
- CN202480022554.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-31
- Filing Date
- 2024-03-21
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to consistently and reliably replicate target sequences and control the size of the replicates in nucleic acid sequencing, while maintaining both strands of the target sequence within the replicate, thus limiting the efficiency and accuracy of bioinformatics analysis.
A circular construct is formed by ligating a linear end adaptor (EA) to the target DNA template. A DNA polymerase-mediated extension reaction generates a DNA template of double length, ensuring the continuous ligation of the parent and daughter strands. The template retains a unique molecular identifier (UMI) and sequence index (SID) for subsequent analysis.
It achieves accurate replication and size control of the target DNA template, improves the accuracy and efficiency of sequencing, and can simultaneously preserve genetic and epigenetic information, supporting efficient bioinformatics analysis.
Smart Images

Figure CN120898003A_ABST
Abstract
Description
[0001] Statement of the Sequence Listing
[0002] The Sequence Listing associated with this application is provided in xml format in place of a paper copy, and is hereby incorporated by reference into the specification. The name of the text file containing the Sequence Listing is P37440-WO_Sequence_Listing.xml. The xml file is 5,000 bytes and was created on March 7, 2024. TECHNICAL FIELD
[0003] The present invention relates generally to methods and compositions for preparing DNA libraries, and more particularly to methods and compositions for replicating target DNA templates and analyzing the genetic and / or epigenetic information of the replicated target DNA templates. BACKGROUND
[0004] Nucleic acid sequencing is a key technology in biology and medicine. While conventional polymerase chain reaction (PCR) technology is very useful and effective, the heating and cooling cycles required for PCR limit its practicality. For example, unzipping hybridized DNA during a heating cycle can degrade the target sample. Such heating / cooling cycles are also incompatible with studies of living systems. Due to the limitations of conventional PCT technology, many studies have focused on identifying isothermal methods of nucleic acid amplification.
[0005] One conventional isothermal method for producing multiple copies of a target nucleic acid includes rolling circle amplification (RCA), in which a small circular oligonucleotide provides a template for polymerase attachment and unidirectional replication. RCA produces a long single ss-DNA product, which consists of many sequentially linked (concatemeric) copies of the complement of the target DNA molecule. This method circularizes the target DNA and initiates polymerase extension with a primer. After replication around the circularized DNA, the primer is displaced, and the polymerase continues with additional rounds of replication of the target DNA, producing multiple copies, until a termination event occurs. This results in a long single-stranded DNA chain, with several copies of the target. The single-stranded DNA chain can then be read and analyzed.
[0006] However, single-read accuracy of single DNA molecule sequencing often has limited accuracy. Some techniques to improve accuracy are i) re-reading the molecule, ii) reading its complement, or iii) reading multiple copies of the DNA molecule (e.g., as employed by RCA). For example, a molecule can be read multiple times by circularizing the target DNA (including both complementary strands) and taking multiple measurements of it as it cycles around the sensing location. Other systems "strip" the complementary strand as one strand is read, and then capture the complement for a small fraction of time so that it can be read immediately thereafter. DNA-based universal molecular identifiers (UMIs) and sample identifiers (SIDs) are then spliced into the individual molecule prior to PCR amplification, such that a measured subset of the family of resulting amplicon copies can be attributed to a single parent molecule from a particular sample. Reading multiple copies within a family improves the accuracy of the sequence known for that molecule.
[0007] While the above methods are useful, what is needed are methods and compositions that consistently, reliably, and easily replicate target sequences, such as for library preparation while controlling the size of the replicates. For example, what is needed are methods and compositions that can replicate target sequences, thereby beneficially controlling the size of a template library. What is also needed are methods that replicate both strands of a target sequence while retaining the parent strands in the replicates, thereby facilitating bioinformatic analysis of the target sequence. SUMMARY
[0008] In certain example aspects, a linear end adapter (EA) for replicating a linear target DNA template is provided. The EA includes, for example, a first polynucleotide strand hybridized to a second polynucleotide strand, thereby forming a polynucleotide duplex. The polynucleotide duplex includes, for example, a first end and a second end. The EA further includes a first nick site within the first polynucleotide strand of the polynucleotide duplex and a second nick site within the second polynucleotide strand of the polynucleotide duplex. A spacer region separates the first and second nick sites from one another, thereby linearly offsetting the first and second nick sites, i.e., there is a linear offset between the first and second nick sites. Moreover, each end of the EA can be configured for ligation to the two ends of the target DNA template. For example, one or both of the nick sites facilitates polymerase binding and extension.
[0009] In certain example aspects, the linear end adapter includes a first Y-branch element sequence attached to the 5' end flanking the first nick site and / or a second Y-branch element sequence attached to the 5' end flanking the second nick site. For example, the Y-branch element can encode a primer binding sequence or other beneficial sequence.
[0010] In certain example aspects, the first polynucleotide strand and / or the second polynucleotide strand of the EA comprises a unique molecular identifier (UMI) sequence. For example, the UMI can be located within the spacer region. In certain example aspects, the first polynucleotide strand of the EA comprises a first sequence index (SID) and / or the second polynucleotide strand of the EA comprises a second SID.
[0011] In certain example aspects, a method of preparing a double-length DNA template from a target DNA template is provided. The method comprises, for example, performing a ligation reaction between the target DNA template and an end adapter as described herein to form a circular construct. For example, the target DNA template comprises a first target DNA template end and a second target DNA template end. Accordingly, the ligation reaction (i) ligates a first end of the end adapter to the first target DNA template end and (ii) ligates a second end of the end adapter to the second target DNA template end. This forms the circular construct. Thereafter, a DNA polymerase-mediated extension reaction is performed on the circular construct. For example, the circular construct is contacted with a plurality of strand displacement polymerases to initiate the extension reaction. The extension reaction forms the double-length DNA template, which comprises, for example, a first copy of the target DNA template and a second copy of the target DNA template.
[0012] In certain example aspects, the first copy of the target DNA template and the second copy of the target DNA template of the double-length DNA template are contiguously linked to each other by a DNA bridging region. For example, the bridging region is derived from the end adapter. For example, the bridging region is double-stranded.
[0013] In certain example aspects, each polynucleotide strand of the double-length DNA template comprises a 5’ to 3’ parent strand of the target DNA template and a 5’ to 3’ child strand copy of the parent strand of the target DNA template. In certain example aspects, the parent strand of the target DNA template and the child strand copy of the target DNA template can be contiguously linked to each other by a 5’ to 3’ strand of the DNA bridging region.
[0014] In certain example aspects, a strand of the bridging region comprises a unique molecular identifier (UMI) or a sequence index (SID). For example, the double-length DNA template comprises a first end and a second end, wherein the first end and / or the second end comprises a SID.
[0015] In certain example aspects, such as when the linear end adaptor comprises a first Y-branch element sequence and a second Y-branch element sequence, the DNA polymerase-mediated extension reaction positions the first Y-branch element sequence and the second Y-branch sequence at the 5' end of each parent strand of the double-length DNA template. Further, the polymerase-mediated extension reaction of the DNA circular construct synthesizes a first progeny Y-branch element sequence and a second progeny Y-branch element sequence, where the first progeny Y-branch element sequence is complementary to the first Y-branch element sequence and the second progeny Y-branch element sequence is complementary to the Y-branch element sequence. For example, the Y-branch element can encode a primer binding site for a subsequent PCR reaction.
[0016] In certain example aspects, these methods can be repeated in succession. For example, successive repetition of the method can yield a quadruple-length DNA template or a multiple-length DNA template. In such example aspects, the multiple-length DNA template comprises multiple copies of the target DNA template.
[0017] In certain example aspects, a method of identifying epigenetic information associated with a target nucleic acid sequence is provided. The method comprises, for example, ligating a linear target DNA template to both ends of a linear end adaptor as described herein, thereby forming a circular DNA construct. The circular DNA construct is then subjected to a DNA polymerase-mediated bidirectional extension reaction in the presence of a plurality of protected cytosine nucleotides. A double-length DNA template is then formed, which comprises, for example, the protected cytosine nucleotides in newly synthesized strands. The double-length DNA template is then denatured and subjected to a bisulfite conversion reaction, which forms a bisulfite-converted double-length DNA template strand of the double-length DNA template. A polymerase chain reaction (PCR) amplification reaction is then performed using the bisulfite-converted double-length DNA template strand, followed by a sequencing reaction of the PCR-amplified / bisulfite-converted double-length DNA template strand. Based on the sequencing of the PCR-amplified / bisulfite-converted double-length DNA template strand, epigenetic information associated with the target nucleic acid is identified. That is, bioinformatics analysis can be used to identify the epigenetic information.
[0018] In certain example aspects, each polynucleotide strand of the double-length DNA template of the method of identifying epigenetic information comprises a parent template strand from the target DNA template and a progeny copy strand of the parent template strand. For example, the parent template strand is continuously linked to the progeny copy strand of the parent template strand through a single-stranded bridging region, where the single-stranded bridging region is derived from the end adaptor. Further, during the DNA polymerase-mediated bidirectional extension reaction, the protected cytosine nucleotides are incorporated into the progeny copy strand of the parent template strand.
[0019] In certain example embodiments, sequencing of the PCR-amplified bisulfite-converted double-length DNA template strand provides the polynucleotide sequence of the parent template strand and the sequence of the progeny copy strand. Then, the step of identifying epigenetic information associated with the target nucleic acid comprises an intra-strand comparison of the polynucleotide sequence of the parent template strand to the polynucleotide sequence of the progeny copy strand. For example, a sequence difference position between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the progeny copy strand identifies an unprotected cytosine residue position in the parent template strand. The unprotected cytosine residue position in the parent template strand, for example, corresponds to an unprotected cytosine residue position in the target nucleic acid sequence.
[0020] In certain example aspects, the double-length DNA template of the method of identifying epigenetic information comprises a first copy and a second copy of a target DNA template. For example, the first copy and the second copy of the target DNA template can be linked together by a double-stranded bridging region, wherein the bridging region is derived from an end adapter. Further, each copy of the target DNA template within the double-length DNA template comprises a parent template strand and a progeny strand that is complementary to and hybridized to the parent template strand. During a DNA polymerase-mediated bidirectional extension reaction, for example, a protected cytosine nucleotide is incorporated into the hybridized complementary progeny strand.
[0021] In such example aspects, when the PCR-amplified bisulfite-converted double-length DNA template is sequenced, an inter-strand comparison of the polynucleotide sequence of the parent template strand to the polynucleotide sequence of the hybridized complementary progeny strand can be used to identify epigenetic information associated with the target nucleic acid. For example, a nucleotide mismatch position between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the hybridized complementary progeny identifies an unprotected cytosine residue position in the parent template strand, wherein the unprotected cytosine residue position in the parent template strand corresponds to an unprotected cytosine residue position in the target nucleic acid sequence.
[0022] In certain example aspects, the protected cytosine nucleotide comprises a methylated cytosine residue. In certain example aspects, the unprotected cytosine nucleotide is an unmethylated cytosine residue. In certain example aspects, the double-length DNA template of the method of identifying epigenetic information comprises a unique molecular identifier (UMI) and / or one or more sequencing index (SID).
[0023] In certain example aspects, a double-length DNA template is provided that is formed by the methods and compositions described herein. For example, the double-length DNA template includes a first copy and a second copy of a target DNA template, wherein the first copy and the second copy of the target DNA template are contiguously linked to each other by a double-stranded bridging region. Further, each polynucleotide strand of the double-length DNA template includes a parent template strand from the target DNA template and a child strand copy of the parent template strand. The parent template strand is contiguously linked to the child copy strand of the parent template strand, for example, by a strand of the bridging region. Additionally, each copy of the target DNA template within the double-length DNA template includes a parent template strand and a child strand that is complementary to and hybridized to the parent template strand.
[0024] In certain example aspects, the double-length DNA template includes a first end and a second end, wherein either end includes a sequence encoding a primer binding site. In certain example aspects, the bridging region - or a strand thereof - includes a unique molecular identifier (UMI) and / or a sequencing index (SID).
[0025] These and other aspects, objects, features, and advantages of the example embodiments will become apparent to those of ordinary skill in the art upon consideration of the following detailed description of illustrated example embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0026] FIG. 1A is an illustration of a linear end adapter for synthesizing a double-length DNA template, according to certain example embodiments.
[0027] FIG. 1B is a schematic depicting circularization of a target DNA template using an EA, according to certain example embodiments.
[0028] FIG. 1C is a schematic depicting initiation of polymerase attachment and bidirectional extension of a circular construct, according to certain example embodiments.
[0029] FIG. 1D is a schematic depicting continued polymerase extension of a circular construct and formation of a double-length DNA template, according to certain example embodiments.
[0030] FIG. 2A is an illustration showing a Y-branch end adapter 200 (YBEA), according to certain example embodiments.
[0031] FIG. 2B is a schematic depicting circularization of a target DNA template using a YBEA 200, according to certain example embodiments.
[0032] FIG. 2C is a schematic depicting initiation of polymerase attachment and bidirectional extension of a circular construct including a YBEA 200, according to certain example embodiments.
[0033] FIG. 2D is a schematic diagram depicting the continuous polymerase extension and formation of a double-length DNA template of a circularized construct using YBEA 200, according to certain example embodiments.
[0034] FIG. 2E is a graphical illustration showing the denatured (single-stranded) form of the double-length DNA template of FIG. 2D (lower panel), in which the original Y-branch element provides a predetermined oligonucleotide primer binding sequence upon replication.
[0035] FIG. 3A is a graphical illustration showing a Y-branch end adapter including a UMI (“YB-UMI-EA”), according to certain example embodiments.
[0036] FIG. 3B is a schematic diagram depicting circularization of a target DNA template using YB-UMI-EA 300, according to certain example embodiments.
[0037] FIG. 3C is a zoomed-in view of a portion of the target DNA template of FIG. 3B, showing example nucleic acid sequences, according to certain example embodiments.
[0038] FIG. 3D is a schematic diagram depicting the initiation of polymerase attachment and bidirectional extension of a circularized construct including YB-UMI-EA 300, according to certain example embodiments.
[0039] FIG. 3E is a schematic diagram depicting the continuous polymerase extension and formation of a double-length DNA template of a circularized target DNA template using YB-UMI-EA 300 example embodiments, according to certain example embodiments.
[0040] FIG. 3F is a schematic diagram showing example bisulfite conversion of a double-length DNA template and its PCR amplification product via use of a Y-branch end adapter with a UMI of FIG. 3A (i.e., YB-UMI-EA), according to certain example embodiments.
[0041] FIG. 3G is a schematic diagram showing the in-strand and inter-strand bioinformatics analysis of a portion of a double-length DNA template to determine epigenetic information associated with the original target DNA template, according to certain example embodiments.
[0042] FIG. 4A is a graphical illustration showing a Y-branch end adapter including two SID sequences and a UMI (i.e., YB-UMI / SID-EA), according to certain example embodiments.
[0043] Figure 4B is an illustration of a double-length DNA template generated using the YB-UMI / SID-EA 400 of Figure 4A, according to certain exemplary embodiments.
[0044] Figure 5A is an illustration of a modified Y-branched end adapter according to Figure 2A according to certain exemplary embodiments, but it has been modified to be adapted for single polymerase attachment and unidirectional extension only.
[0045] Figure 5B is a schematic diagram depicting the initiation of unidirectional extension of polymerase attachment and cyclic constructs according to certain exemplary embodiments.
[0046] Figure 5C is a schematic diagram depicting the continuous polymerase extension and asymmetric template formation of a cyclic construct using modified YBEA 500 according to certain exemplary embodiments.
[0047] Figure 6 is a schematic diagram depicting the formation of a quadruple-length DNA template from a double-length DNA template according to certain exemplary embodiments. Detailed Implementation
[0048] Overview
[0049] This article discloses methods and compositions for preparing DNA libraries containing replicated target nucleic acid sequences. For example, a double-length DNA template is formed by extending a target DNA template that includes or encodes a target nucleic acid sequence by adding a single copy of the target DNA template to an original target DNA template. That is, a double-length DNA template is "double-length" because it comprises two copies of the original target DNA template (and therefore, if a target sequence is present, two copies are included). Typically, the method includes, for example, the steps of circularizing the target DNA template and subsequently replicating it to form two copies of the target DNA template, each copy residing within the double-length DNA template.
[0050] Advantageously, each strand of the double-length DNA template includes a parental polynucleotide sequence sequentially linked to a newly synthesized daughter copy of the parental polynucleotide sequence. Furthermore, each copy of the target DNA template within the double-length DNA template includes a parental strand that hybridizes with a complementary daughter DNA strand. In some instances, predetermined sequences such as primer sequences, unique molecular identifiers (UMIs), and sample indexes (SIDs) may also be included in the double-length DNA template. And due to the association between parental and daughter polynucleotide sequences within the double-length DNA template, sequencing of the double-length DNA template can beneficially reveal genetic and epigenetic information associated with the target nucleic acid sequence.
[0051] To facilitate the preparation of double-length DNA templates, linear end adaptors (EAs) are provided in some instances, comprising hybridized polynucleotide chains to form a polynucleotide duplex, such as a DNA molecule. For example, the ends of the EA are each attached to the opposite end of the target DNA template to form a circular construct. The EA includes juxtaposed cleavage sites—one on each polynucleotide chain—separated by spacers. Because each cleavage site is located within the polynucleotide chain of the EA duplex, each cleavage site is flanked by both the 5' and 3' ends. Thus, in some instances, the EA provides an exposed 3' end for polymerase binding and extension along each strand of the EA.
[0052] For example, when a circular construct including EA comes into contact with DNA polymerase, the two juxtaposed 3' ends can be extended in opposite directions by the polymerase, while the opposing strands of the target DNA template are displaced. The complete extension of the two free 3' ends provided by EA produces a double-length DNA template, wherein each copy of the target DNA template within the double-length DNA template comprises an original (parental) DNA strand and a newly synthesized and complementary daughter strand. Each copy of the target DNA template is separated by EA, which forms a bridge between the two template copies. In this way, the bridges of the double-length DNA template originate from EA. Furthermore, each polynucleotide strand of the double-length DNA template comprises a parental polynucleotide sequence from the target DNA template and a new daughter copy of the parental polynucleotide sequence, the parental sequence and the daughter copy being continuous and covalently linked to each other and having the same sequence.
[0053] In some instances, single-stranded (ss) branched sequence elements (or Y-branched elements) can be added to the 5' end of each nick site of the EA to form one or more Y-branched adaptors within a double-length DNA template. Y-branched elements may include, for example, a polynucleotide sequence encoding a primer binding site. For instance, a Y-branched element may include a single-stranded polynucleotide sequence (e.g., ssDNA) whose complement encodes a primer binding site as described herein. The primer binding site can be used, for example, in a subsequent PCR reaction to efficiently and accurately amplify the double-length DNA template (and thus the original target DNA template).
[0054] In some instances, because the methods disclosed herein advantageously provide a double-length DNA template in which a parental polynucleotide sequence is covalently and sequentially linked to a daughter polynucleotide chain copy, both epigenetic (parental chain) and genetic (daughter chain) information are preserved in the double-length DNA template. That is, because both polynucleotide chains of the double-length DNA template compositions provided herein comprise a parental polynucleotide sequence from the target DNA template and a daughter copy of that parental sequence, chain-specific analysis and comparison can be used to identify parental chain methylation, thereby identifying epigenetic information associated with the parental chain, and thus identifying epigenetic information present in the target sequence. Furthermore, such genetic and epigenetic information can be advantageously obtained in a single read by sequencing the double-length DNA template.
[0055] In some instances, the methods provided herein can be used to generate double-length DNA templates that include a unique molecular identifier (UMI). For example, the UMI can be included in the spacer region of the end-adaptors provided herein, i.e., in the region between the juxtaposed nick sites of the end-adaptors. In such instances, Y-branching elements may also be included to allow for subsequent PCR amplification. For example, by including a UMI in the double-length DNA template, the double-length DNA template can be used for a variety of bioinformatics applications. For example, the sequence information from each strand of the double-length DNA template can be bioinformatically paired to advantageously confirm the accuracy of sequence reads. In some instances, such UMIs can also facilitate strand differentiation in the genetic and epigenetic analyses described herein.
[0056] In some instances, the methods provided herein can be advantageously used to generate double-length DNA template compositions including one or more sample indices (SIDs). Typically, the use of such SIDs is very useful in applications such as DNA multiplexing (i.e., processing multiple different samples simultaneously). For example, different SIDs can be included adjacent to the Y-branch sequence elements described herein. Subsequently, double-length DNA molecules with different SIDs can be processed simultaneously, allowing for sample differentiation after sequencing. Furthermore, because multiple copies of a SID can appear in a single repeat PCR product strand, the SID can be determined with high accuracy in bioinformatics, thereby reducing or eliminating the need for additional error correction. In such exemplary embodiments, the SID can also be used as a marker within a given strand, thereby allowing for additional analysis.
[0057] In some instances, methods and compositions for generating double-length DNA templates can be applied sequentially to multiply the number of parental target DNA templates on a single molecule in each iteration, such as to generate quadruple-length or multi-fold-length DNA templates. This can be advantageously used in sequencing applications, for example, to generate additional template reads in a single pass, thereby achieving higher read accuracy and confidence. In other exemplary instances, the target DNA template can be extended asymmetrically, generating asymmetric DNA templates. For example, nick sites on end-adaptors can be blocked, allowing extension from a single nick site.
[0058] Because double-length DNA templates can be limited to single-copy extension products (i.e., dual templates forming the original parental target DNA template), the methods and compositions provided herein also advantageously maintain library length consistency and read efficiency. The methods and compositions provided herein also improve sequencing accuracy while balancing other important characteristics of sequencing systems, such as throughput, efficiency, and read length. These and other examples and benefits will become apparent to those skilled in the art given the further detailed description provided herein.
[0059] Terminology and Naming
[0060] The invention will now be described in detail by reference only using the following definitions and examples. All patents and publications cited herein, including all sequences disclosed therein, are expressly incorporated herein by reference in their entirety.
[0061] Unless otherwise defined herein, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Such commonly used techniques and methods are described, for example, in Green and Sambrook, *Molecular Cloning: A Laboratory Manual* (4th Edition), Volumes 1–3, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, 2012 (hereinafter “Sambrook”) and *Current Protocols in Molecular Biology*, edited by FM Ausubel et al., originally published as a book in 1987 by Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., with regular supplements since 2011, and now available online as a journal, *Current Protocols in Molecular Biology*, Volumes 00–130, (1987–2020), published by Wiley & Sons, Inc. in the Wiley Online Library. Each of these documents provides a general dictionary of the many terms used in this invention for those skilled in the art.
[0062] While any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of the invention, preferred methods and materials are described. It should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. For the purposes of interpreting this disclosure, the following terminology will be applied in the description, and where appropriate, terms used in the singular will also include the plural, and vice versa.
[0063] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various implementations of exemplary embodiments. At the same time, those skilled in the art will fully understand that certain steps or actions used to describe the method can also be interchanged or adjusted in terms of order. Therefore, the various orders in the specification and drawings are merely for the purpose of clearly describing a particular embodiment and are not necessary orders, unless otherwise stated that a certain order must be followed or such an order is necessary from the context (e.g., polymerase must be added to the reaction mixture for polymerase-mediated replication to occur).
[0064] Unless otherwise stated, nucleic acids are written from left to right with a 5' to 3' orientation; amino acid sequences are written from left to right with an orientation from amino to carboxyl.
[0065] The headings provided herein are not intended to limit the various aspects or embodiments of the invention, which can be obtained by referring to the entire specification. Therefore, the terms defined below are defined more fully by referring to the entire specification.
[0066] As used herein, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include the plural objects referred to.
[0067] A range may be expressed herein as from “about” or “approximately” a specific value, and / or to “about” or “approximately” another specific value. When expressing such a range, the other side includes from a specific value in that range and / or to another specific value in that range. It should also be understood that each endpoint of a range is significant relative to and independent of the other endpoint. Similarly, when a value is expressed as an approximation using the antecedent “about”, it should be understood that the specific value forms the other side.
[0068] In some exemplary embodiments, the term “about” or “approximately” is understood to mean within the normal tolerance range in the field, such as within 2 standard deviations of the average. “About” or “approximately” can be understood to mean within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless the context clearly indicates otherwise, all numerical values provided herein may be modified by the term “about”. Furthermore, terms such as “example,” “exemplary,” or “illustrated” as used herein are not intended to indicate preference, but rather to explain that the aspects discussed herein are merely one instance of the presented aspects.
[0069] The term "amplification" refers to the process of preparing additional copies of the target nucleic acid. Amplification can involve more than one cycle, such as multiple cycles of exponential amplification. Amplification can also involve only one cycle (preparing a single copy of the target nucleic acid). This copy may contain additional sequences, such as those present in the primers used for amplification. Amplification can also produce copies of only one strand (linear amplification) or preferentially produce copies of one strand (asymmetric PCR).
[0070] As used herein, "polymerase" refers to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Typically, the enzyme begins synthesis at the 3' end of a primer annealed to the polynucleotide template sequence and proceeds toward the 5' end of the template strand. "DNA polymerase" catalyzes the polymerization of deoxynucleotides using complementary template DNA strands and primers, for example, by sequentially adding nucleotides to a free 3'-hydroxyl group. The template strand determines the sequence of the added nucleotides through Watson-Crick base pairing.
[0071] Typically, any DNA polymerase suitable for use with rolling circle amplification can be used for the replication reaction. In some embodiments, a suitable DNA polymerase will have strand displacement activity. The term strand displacement describes the ability to displace downstream DNA encountered during DNA synthesis. Several DNA polymerases with varying degrees of strand displacement activity are known in the art and are commercially available. In some exemplary embodiments, the polymerase is a φ29 polymerase, a bst polymerase, etc. In a preferred embodiment, the strand displacement polymerase is a φ29 polymerase.
[0072] In some embodiments, the DNA polymerase is a high-fidelity DNA polymerase. The fidelity of a DNA polymerase is the result of accurate replication of the desired template. Specifically, this involves multiple steps, including the ability to read the template strand, select the appropriate nucleoside triphosphate, and insert the correct nucleotide at the 3' primer end, thus maintaining Watson-Crick base pairing. In addition to effectively distinguishing between correct and incorrect nucleotide incorporation, some DNA polymerases possess 3'→5' exonuclease activity. This activity, known as "proofreading," is used to remove incorrectly incorporated mononucleotides and replace them with the correct nucleotides.
[0073] In some embodiments, suitable high-fidelity DNA polymerases for practicing the present invention include KAPA HiFi DNA polymerase, commercially available from Roche Diagnostics Corp., Q5® high-fidelity DNA polymerase, commercially available from New England Biolabs, Inc., and engineered Pfu DNA polymerases, such as Pfu-X, commercially available from Jena Biosciences.
[0074] As used herein, the terms "ligate / ligating / ligation," etc., generally refer to the process of covalently linking two or more molecules together, such as covalently linking two or more nucleic acid molecules to each other. A similar term, "ligateable," refers to having the ability to link. As those skilled in the art will understand, linking involves a condensation reaction that forms a covalent bond between the ends of a first nucleic acid molecule and the ends of a second nucleic acid molecule.
[0075] In some exemplary embodiments, ligation may include forming a covalent bond between the 5' phosphate group of one nucleic acid and the 3' hydroxyl group of a second nucleic acid, thereby forming a ligated nucleic acid molecule. Typically, for the purposes of this disclosure, a target DNA template sequence may be ligated to an end adaptor to generate a circularized construct. Ligation includes the ligation of two DNA molecules, each having a protruding end (i.e., a "sticky" end), where one strand is longer than the other (typically at least a few nucleotides longer), such that the longer strand has unpaired bases. Ligation also includes the ligation of DNA molecules in which each molecule has an equal strand length (i.e., a "flat" end, without protrusions).
[0076] In some exemplary embodiments, ligation can be achieved using an asymmetric 5' thymine nucleotide overhang on the target DNA template and a 5' adenine nucleotide overhang on the end adaptor. For example, the target DNA template and end adaptor can be combined at equimolar or near-equimolar concentrations for ligation. In some exemplary embodiments, the concentrations of the end adaptor and target DNA template can be optimized through trial and error to favor circularization ligation rather than tandem ligation; for example, the molar ratio of target DNA template to adaptor can be 1:1, 1:5, 1:10, 1:25, or 1:50. In some exemplary embodiments, improved circularization can be achieved when the target DNA template and / or end adaptor include sufficient flexibility to bend and align for a sufficient amount of time and frequency. It has been shown that ds-DNA >200 base pairs will link together to form “microloops,” and those with linear ds-DNA oligomers containing nick sites will be more prone to circularization (see, for example, “Small DNA Circles as Probes of DNA Topology,” Bates, AD et al., Biochem. Soc. Trans. (2013) 41, 565-570, which is incorporated herein by reference in its entirety). Target sequencing libraries are typically in this size range. In some exemplary embodiments, the target DNA template is 200 to 500 base pairs in length.
[0077] In some exemplary embodiments, the circularization of the target DNA template can be facilitated by reducing the concentration of the target DNA template and / or end adaptors to favor circularization rather than polymerization. In other embodiments, circularization can be facilitated by a “protein scaffold” strategy that uses one or more DNA-binding proteins to increase the local concentration of intramolecular linkable ends to shift the equilibrium toward circularization and physically bend the DNA to overcome the energy challenge of forming small loops. In some exemplary embodiments, suitable DNA-binding proteins for protein scaffolds include histones, Abf2p, DSP1, histone-like proteins AU, and CAP. In some exemplary embodiments, the circularized linker constructs can be enriched by treatment with one or more exonucleases because the circularized constructs lack free ends that would trigger exonuclease-mediated DNA degradation. Some exemplary exonucleases include ExoVIII, ExoIII, and T5 exonucleases.
[0078] As used herein, the terms “target,” “target sequence,” or “target nucleic acid sequence” are used interchangeably and refer to any target nucleic acid molecule that has been processed (e.g., for the purpose of generating a double-length DNA template as described herein). Target nucleic acid sequences may include genomic DNA, subgenomic DNA, chromosomal DNA (e.g., from a separated chromosome or a portion of a chromosome, such as one or more genes or loci from a chromosome), mitochondrial DNA, chloroplast DNA, DNA derived from plasmids or other episomes (or recombinant DNA contained therein), or double-stranded cDNA prepared by reverse transcription of RNA, or RNA that can subsequently be converted into cDNA by any method recognized in the art, or composed of such sequences. Furthermore, target nucleic acid sequences, such as target DNA or RNA, may be derived from any in vivo or in vitro source, including from one or more cells, tissues, organs, body fluids, or organisms (whether alive or dead), or from any biological or environmental source (e.g., water, air, soil).
[0079] The terms "DNA," "double-stranded DNA," or "dsDNA" generally refer to complementary deoxyribonucleic acid polynucleotide chains that hybridize to form a double helix. The two polynucleotide chains are held together by hydrogen bonds between complementary nucleotide base pairs (i.e., Watson-Crick). Each nucleotide in DNA consists of a sugar molecule, a phosphate group, and one of four nitrogenous bases: adenine (A), cytosine (C), guanine (G), or thymine (T). Perfect complementarity is not required to maintain the double helix. Double-stranded DNA can be found in the nucleus of eukaryotic cells, as well as in the cytoplasm and plasmids of prokaryotic cells. It is also used in various molecular biology techniques, such as PCR (polymerase chain reaction), DNA sequencing, and genetic engineering.
[0080] A DNA strand, or single-stranded DNA (ssDNA), refers to one of the polynucleotide chains in a DNA molecule; it may also be called ssDNA. For example, a daughter polynucleotide chain is a new strand of a DNA duplex produced by replicating a DNA molecule. For instance, a polymerase-mediated replication reaction uses a template DNA strand to produce a complementary strand as the daughter strand. In some exemplary embodiments, the DNA is cDNA that has been converted from a target RNA sequence or otherwise derived.
[0081] As used herein, the terms “target DNA template” and “DNA template” are used interchangeably and refer to a DNA molecule that encodes or includes a target nucleic acid sequence as genetic and / or epigenetic information. For example, one strand of the DNA molecule may include or encode a target sequence, while another hybrid and opposite strand of the DNA molecule is complementary to the strand that includes or encodes the target sequence. In some embodiments, the target DNA template may be a natural DNA target fragment (e.g., a genomic or cell-free DNA target fragment), or it may be a cDNA copy of a natural DNA or RNA target fragment. The target DNA templates disclosed herein are molecules that are replicated (e.g., duplicated) and / or subjected to DNA sequencing. Furthermore, when a subsequent DNA molecule is formed that includes, for example, a polynucleotide chain of the target DNA template, that chain may be referred to as the “original” or “parental” chain of the target DNA template, indicating that the chain was originally part of the target DNA template. For example, the target template may be prepared according to any method known in the art.
[0082] The term "primer" refers to a single-stranded oligonucleotide that hybridizes to a target nucleic acid sequence ("primer binding site") and is capable of acting as a starting point for synthesis along the complementary strand of the nucleic acid under conditions suitable for such synthesis. In other words, a "primer" serves as a substrate on which nucleotides can be polymerized by a polymerase. In various embodiments, the primer has a free 3'-OH group that can be extended by a nucleic acid polymerase. For template-dependent polymerases, typically at least the 3' portion of the primer oligonucleotide is complementary to a portion of the template nucleic acid. The primer oligonucleotide "binds" (or "complexes," "anneals," or "hybridizes") the primer oligonucleotide to the template via hydrogen bonding and other molecular forces to obtain a primer / template complex for initiating synthesis by a DNA polymerase and is extended during DNA synthesis by adding a covalently bound base complementary to the template (attached to the 3' end of the template) (i.e., "primer extension").
[0083] As used herein, a unique molecular identifier (UMI) is a sequence of nucleotides inserted into or identified within a DNA molecule that can be used to distinguish individual DNA molecules from one another. Due to their complementary nature within the DNA molecule, UMIs present in or inserted into a DNA molecule can also be used to identify individual strands of the DNA molecule, as the polarity (orientation) of the UMI sequence can be identified and distinguished between two complementary DNA strands. See, for example, Kivioja, Nature Methods 9, 72-74 (2012). UMIs can be sequenced along with the DNA molecule they are associated with to determine whether a read sequence belongs to one source DNA molecule or another. The term “UMI” is used herein to refer to both the sequence information of the polynucleotide and the physical polynucleotide itself. UMI sequences can be random, pseudo-random, partially random, or non-random nucleotide sequences inserted into or otherwise incorporated into, for example, end-adaptors as described herein.
[0084] The term "sample index" is a nucleotide sequence appended to a target polynucleotide, where the sequence identifies the origin of the target polynucleotide (i.e., the sample from which the target polynucleotide originated). Therefore, the sample index (or SID) is also referred to as a "sample identifier sequence," "index sequence identifier," "multiplexing identifier," or "MID." In practice, each sample includes a different sample index sequence (e.g., one sequence is appended to each sample, where different samples are appended to different sequences), and the samples are pooled. After sequencing the pooled samples, the sample identifier sequence can be used to identify the origin of the sequence. Conventionally, the sample identifier sequence can be added to the 5' end or the 3' end of the polynucleotide. In some cases, some of the sample identifier sequence may be at the 5' end of the polynucleotide, and the rest may be at the 3' end. When the elements of the sample identifier have sequences at each end, the 3' and 5' sample identifier sequences together identify the sample. In some instances, the sample identifier sequence is simply a subset of the bases appended to the target oligonucleotide. Furthermore, as described herein, end connectors can be used to include SIDs in a sample.
[0085] As used herein, the term "polymerase chain reaction" (or "PCR") refers to a method for increasing the concentration of a segment of a target polynucleotide in a mixture of genomic DNA without cloning or purification. See generally U.S. Patent Nos. 4,683,195 and 4,683,202 (description of PCR procedures). Methods for amplifying target polynucleotides typically consist of repeated cycles of denaturation, primer annealing, and extension using DNA polymerase. Because the amplified segment of the desired target polynucleotide becomes the dominant nucleic acid sequence in the mixture (in terms of concentration), they are referred to as "PCR-amplified." In modifications of the methods described above, target nucleic acid molecules can be amplified using multiple different primer pairs (in some cases, one or more primer pairs for each target nucleic acid molecule) to form a multiplex PCR reaction.
[0086] As used herein, the term "terminal adaptor" generally refers to a polynucleotide duplex, such as a DNA molecule, that can be added to (i.e., ligated) to a target DNA template. Terminal adaptors can be 5 to 100 bases in length and can provide, include, or encode amplification primer binding sites, sequencing primer binding sites, molecular identifiers, and / or sample identifier sequences, as described herein. Terminal adaptors can be added to the 5' and 3' ends of the target DNA template via ligation. When added to the target DNA template, for example, the terminal adaptor forms a circularized structure ("circularized DNA construct" or "circular construct") in which both ends of the target molecule bind to the ends of the terminal adaptor.
[0087] Double-length DNA template
[0088] Turning now to the accompanying drawings, where similar numbers throughout the drawings indicate similar (but not necessarily identical) elements, exemplary embodiments are described in detail. Furthermore, while some of the drawings provided herein illustrate the ligation, circularization, and replication of a single target DNA template, it should be understood that multiple target DNA templates are typically ligated, circularized, and replicated in a single library preparation reaction, such as when multiple reaction components (e.g., multiple target DNA templates, end-adaptors, polymerases, etc.) are combined. Multiple replicates can then be used for any number of different applications, such as sequencing or other analyses.
[0089] In some exemplary embodiments, a method for preparing a DNA library is provided, the method comprising synthesizing a double-length DNA template from a target nucleic acid via a linear end adaptor (EA). This is illustrated in Figures 1A through 1D, which collectively illustrate the characteristics of an example EA according to some exemplary embodiments and show how the example EA can be used to synthesize a double-length DNA template.
[0090] Referring to Figure 1A, a linear end adapter (EA) for synthesizing a double-length DNA template is illustrated according to certain exemplary embodiments. As shown, EA 100 is a duplex polynucleotide molecule, such as a DNA molecule, comprising hybridized oligonucleotide chains, namely a first polynucleotide chain 100a (shown in circles) and a hybridized second polynucleotide chain 100b (shown in rectangles). As used herein, in conjunction with the structure of the EA of the present invention, the term "polynucleotide chain" refers to one or more oligonucleotides having the same 5' to 3' polarity, which hybridizes with a portion of one or more complementary oligonucleotides to form the EA structure 100. In the embodiment shown in Figure 1A, polynucleotide chains 100a and 100b each comprise two oligonucleotide portions (described further below) separated by cleavage sites 101a and 101b, respectively. In this regard, reference to "100a" refers to the entire 5'→3' chain, wherein cleavage site 101a is within chain 100a. Similarly, references to "100b" refer to the entire 5'→3' strand that hybridizes with strand 100a, where the cleavage site 101b is within strand 100b. In some exemplary embodiments, EA is the entire length of 50 to 100 nucleotides, such as 75 to 80 nucleotides. In some exemplary embodiments, the length of the oligonucleotide used to generate EA is selected to ensure efficient and specific hybridization to form a stable EA structure, as discussed further herein.
[0091] As also shown in the example EA of Figure 1A, within EA are a first nick site 101a and a second nick site 101b. That is, in some exemplary embodiments, EA includes an internal first nick site 101a and a second nick site 101b, one of each of the first polynucleotide chain 100a and the second polynucleotide chain 100b of EA 100. A nick site includes, for example, any break or gap in a strand of a DNA molecule, causing the chain to be discontinuous. In some exemplary embodiments, the nick site is a break or disruption of the phosphodiester backbone, while in other exemplary embodiments, the nick site is a gap in one or more nucleotides in the DNA chain. It is noteworthy that each nick site 101a and 101b is associated with and located flanking a free 5' end and a free 3' end. Referring to Figure 1A, the length and location of the depicted nick sites 101a and 101b are not intended to be limiting, but are shown for illustrative purposes only.
[0092] In some exemplary embodiments, EA can facilitate polymerase-mediated chain elongation reactions by exposing the 3' ends of nick sites 101a and 101b. That is, the polymerase can use the exposed 3' ends to elongate the 3'-associated chain in conventional polymerization and chain displacement reactions, as described herein. Preferably, nick sites 101a and 101b are separated by a spacer subregion 102, allowing EA to accommodate the attachment of two polymerases for bidirectional elongation, as described herein. As shown, for example, the spacer subregion linearly offsets the first nick site 101a from the second nick site 101b. Therefore, nick sites 101a and 101b can be spaced sufficiently far apart, such as separated by the spacer subregion 102, so that the binding of one polymerase does not spatially impede and / or displace the binding of the second polymerase. EA 100 also includes ends 103 and 104 located on the flanks of each nick site, each end 103 and 104 being compatible with the ends of the target DNA template for effective ligation. In other words, the ends of EA can be ligated to the target DNA template.
[0093] Any means known in the art can be used to form or otherwise create an EA. For example, as shown in Figure 1A, in one embodiment, an EA is formed by hybridization of four incompletely contiguous synthetic oligonucleotides, leaving spacer regions (also referred to herein as “gaps” or “nicks”) during hybridization. Any other suitable method for generating nicks, gaps, or other sites for polymerase binding and initiation of DNA synthesis can be used. For example, an EA can be generated from a continuous chain of oligonucleotides designed to include recognition sites for one or more nicking enzymes (i.e., nicking endonucleases). Nicking enzymes are known in the art and hydrolyze (cut) only one strand of the DNA duplex to produce a “nick” rather than a cut DNA molecule. Treatment of an EA with a nicking enzyme generates a free 3' end, which provides a polymerase initiation site in each strand.
[0094] Referring to Figure 1B, a schematic diagram depicting the circularization of a target DNA template using EA 100 according to certain exemplary embodiments is provided. The target DNA template includes, for example, a target sequence encoding. Before ligating the adaptor, the ends of the target DNA template can be prepared for ligation. For example, by end repair and the generation of blunt ends with 5' phosphate groups. The DNA template can be blunted by many methods known to those skilled in the art. In a particular method, the ends of fragmented DNA are “polished” with T4 DNA polymerase and Klenow polymerase, a procedure well known to those skilled in the art, and then phosphorylated with a polynucleotide kinase. A single 'A' deoxynucleotide is then added to both 3' ends of the DNA molecule using Taq polymerase or Klenow exo minus polymerase, producing a 3' overhang complementary to a single 3' 'T' overhang on the double-stranded end of the adaptor.
[0095] As shown in Figure 1B, double-stranded EA 100 binds to target DNA template 107, which has a first end 105 and a second end 106. Also as shown, the target DNA template comprises complementary polynucleotide chains, namely, a first template strand 107a (dashed line) and a second template strand 107b (solid line), both referred to herein as the “parental strand” or “parental line”. That is, strands 107a and 107b of the target DNA template 107 correspond to the original strand of the target DNA template, which includes or encodes the target sequence as described herein. In some embodiments, the parental strand will include epigenetic information, such as methylated cytosine residues.
[0096] At step 1a, for example, EA 100 is ligated to either end of the target DNA template 107, which includes parental polynucleotide chains 107a and 107b. For example, end 103 of EA 100 is ligated to end 106 of the template (Figure 1B). Alternatively, at step 1a, and although not shown for simplicity, the other end (i.e., 104) of EA 100 is ligated to end 105 of the target DNA template 107. In either case, one end of EA 100 is ligated to an end of the target DNA template.
[0097] In step 1b of Figure 1B, the remaining free end of EA 100 is attached to the remaining free end of the target DNA template 107 to form a circular construct 109. For example, if the end 103 of EA 100 is attached to the template end 106 in step 1a, then the EA end 104 of EA 100 is attached to the template end 105 in step 1b, thereby forming a circular construct 109 of the original (parental) template 107. Alternatively, if the end 104 of EA 100 is attached to the template end 105 in step 1a, then the EA end 103 is attached to the template end 106 in step 1b, thereby forming a circular construct 109. In either case, at step 1b of Figure 1B, the two ends 103 and 104 of EA 100 connect to each end 105 and 106 of template 107, thereby forming a DNA bridge 108 between the ends of template 107. That is, the entire EA 100 forms a DNA bridge 108 between the two ends 105 and 106 of the parental template 107. This forms a circular construct 109 comprising EA 100 (as bridge 108) and complementary parental template strands 107a and 107b of the parental template 107.
[0098] In this manner, EA 100 (of Figure 1A) operates as a bridging precursor for the bridging region 108 of the circular construct 109. As shown, corresponding first nick site 101a and second nick site 101b are retained in the circular construct 109 (as part of the bridging region 108), and thus, in some exemplary embodiments, provide two corresponding 3' ends available for polymerase attachment and bidirectional extension, as described herein.
[0099] Continuing with the above examples, Figure 1C is a schematic diagram depicting the initiation of polymerase attachment and bidirectional extension of a circular construct 109 according to certain exemplary embodiments. As shown, once the circular construct 109 is formed (at step 1b, Figure 1B), DNA polymerases (shown as DNA polymerases 110a and 110b) are added to initiate the replication reaction. For example, in one embodiment, DNA polymerase 110a is attached to a nick site 101a of EA 100. In this respect, the nick site 101a and its available 3' end serve as the primer end for the attachment and extension initiation of DNA polymerase 101a. Similarly, DNA polymerase 110b is attached to a nick site 101b of EA 100, wherein the 3' end nick site 101b serves as the primer end for the attachment and extension of DNA polymerase 110b. As shown in the figure (with opposite arrows), polymerases 110a and 110b are positioned for bidirectional extension of the cyclic construct 109 in opposite directions (Figure 1C, top).
[0100] At step 1c in Figure 1C, the first polymerase 110a and the second polymerase 110b extend the circular construct 109 bidirectionally in opposite directions (Figure 1C, bottom, see arrows). For example, polymerase 110a extends the 3' end of the nick site 101a while also displacing the 5' end of the nick site 101a (and its associated parental template strand 107a). That is, as polymerase 110a proceeds, it uses the parental strand 107b as a template to extend the 3' end of the nick site 101a to synthesize a new daughter strand 107a' that is sequence-complementary to (and therefore shares the same sequence as) the parental strand 107a. The new daughter strand 107a' also includes a replicated ssDNA daughter strand bridge portion 108a as part of the DNA bridging region 108.
[0101] Similarly, polymerase 110b extends the 3' end of nick site 101b while simultaneously using parental strand 107a as a template to replace the 5' end of nick site 101b (and its associated parental template strand 107b) (Figure 1C, bottom). In other words, as polymerase 110b proceeds in step 1c, it uses parental strand 107a as a template to extend the 3' end of nick site 101b to synthesize a new daughter strand 107b' that is sequence-complementary to parental strand 107a (and therefore shares the same sequence as parental strand 107b). The new daughter strand 107b' also includes a replicated ssDNA daughter strand bridge portion 108b as part of bridging region 108.
[0102] Figure 1D is a schematic diagram depicting the continuous polymerase extension and double-length DNA template formation of the circular construct 109 according to certain exemplary embodiments. As shown (top), polymerases 110a and 110b continue to the end of the parental template 107. For example, polymerase 110a continues along the parental template strand 107b to the 5' end of the parental template strand 107b, completing the synthesis of the new daughter strand 107a'. Similarly, polymerase 110b continues along the parental template strand 107a to the 5' end of the parental template strand 107a, completing the synthesis of the new daughter strand 107b'. At step 1d, once polymerases 110a and 110b have completed the synthesis of daughter strands 107a' and 107b', respectively, polymerases 110a and 110b dissociate from the circular construct 109 to form a double-length DNA template 111 (as shown in Figure 1D, below).
[0103] As shown in Figure 1D (below), the double-length DNA dual template 111 comprises two copies of the original target DNA template 107, namely, a first copy 111a and a second copy 111b, each located on either side of a bridging region 108, which includes a spacer region 102. Notably, each template copy comprises both a parental polynucleotide strand (shown in black) and a newly synthesized daughter polynucleotide strand (shown in gray). For example, template copy 111a comprises the original (parental) template strand 107b and the newly synthesized daughter strand 107a'. On the other side of the bridging region 108, template copy 111b comprises the original (parental) template strand 107a and the newly synthesized daughter strand 107b'. Furthermore, the double-length DNA 111 template includes a first end 112a and a second end 112b. For example, the first end 112a includes the portion of the EA 100 chain 100b associated with the 5' end of the nick site 101b EA 100 (the hollow black rectangle at end 112a) and a copy thereof (the hollow gray circle at end 112a). Similarly, the second end 112b of the double-length DNA template 111 includes the portion of the EA 100 chain 100a associated with the 5' end of the nick site 101a EA 100 (the hollow black circle at end 112b) and a copy thereof (the hollow gray rectangle at end 112b).
[0104] It is noteworthy that the two strands of the double-length DNA template also include a parental strand (black) linked to a newly synthesized daughter strand (gray) of the parental strand. For example, parental strand 107a is covalently and continuously linked to the newly synthesized daughter strand 107a' in the 5'→3' direction via the strand of bridging region 108 (i.e., the strand of bridging region 108 including strand portions 100a and 108a). Furthermore, the nucleotide sequence of parental strand 107a matches the nucleotide sequence of the new daughter strand 107a' due to polymerase-mediated elongation of the circular construct 109 as described herein. In other words, daughter strand 107a' is a sequence copy (i.e., a daughter copy) of the parental strand 107a of the target DNA template.
[0105] Similarly, on the complementary strand of the double-length DNA template, the parental strand 107b is also covalently and continuously linked to the new daughter strand copy 107b' in the 5'→3' direction via the strand of DNA bridging region 108 (i.e., the strand including strand portions 100b and 108b of bridge 108). And similarly—and again due to polymerase-mediated elongation of the circular construct as described herein—the nucleotide sequence of the parental strand 107b matches the nucleotide sequence of the new daughter strand 107b'. In this way, each strand of the double-length DNA includes both a parental polynucleotide sequence and a daughter polynucleotide sequence copy on each of the two target DNA copies 111a and 111b, in addition to the parental template strand and its complementary daughter strand (Fig. 1D, bottom).
[0106] Double-length DNA template with Y-branched adaptor
[0107] In some exemplary embodiments, the design of the end-adaptor (EA) shown in Figure 1A can be modified to impart additional features to the resulting double-length DNA template. These include features, for example, facilitating subsequent PCR amplification and / or DNA sequencing. This is illustrated in Figures 2A through 2E, which collectively illustrate modified EAs according to certain exemplary embodiments and depict how modified EAs can be used to synthesize double-length DNA templates including primer-binding sequences.
[0108] Referring to Figure 2A, an illustration is provided of a Y-branched adapter 200 (YBEA) having hybrid strands 200a (circle) and 200b (rectangle) according to certain exemplary embodiments. As shown, YBEA 200 has the general double-stranded polynucleotide EA structure shown in Figure 1A, except that the first Y-branching element 213a and the second Y-branching element 213b are respectively connected to the 5' ends of the first nick site 201a and the second nick site 201b. For example, each Y-branching element 213a and 213b may include a predetermined oligonucleotide sequence, the design of which can be customized to achieve a specific purpose, such as, but not limited to, PCR amplification or DNA sequencing of double-length DNA templates. Also similar to the EA in Figure 1A, the reference to "200a" refers to the entire 5'→3' strand of YBEA 200, where the nick site 201a is within strand 200a. Similarly, the reference to "200b" refers to the entire 5'→3' strand that crosses with strand 200a, where the nick site 201b is within strand 200b.
[0109] In some exemplary embodiments, each Y-branch element 213a and 213b sequence may include a predetermined oligonucleotide sequence that provides a complementary or hybridizable primer binding site for use, for example, PCR amplification. That is, each Y-branch element 213a and 213b may, for example, include 10 to 30 nucleotides, such as 15 to 25 nucleotides or 18 to 22 nucleotides, whose complementary sequence includes the primer binding site sequence. In some exemplary embodiments, Y-branch elements 213a and 213b include the same sequence, while in other exemplary embodiments, Y-branch elements 213a and 213b include different sequences. In some exemplary embodiments, Y-branch elements 213a and 213b have the same length, while in other exemplary embodiments, Y-branch elements 213a and 213b may have different lengths.
[0110] YBEA 200 also includes ends 203 and 204 located flanking each nick site 201a and 201b, each end 203 and 204 being compatible with the end of the target DNA template for effective ligation. That is, the ends are ligable to the target DNA template. Preferably, the nick sites 201a and 201b are separated by a spacer region 202, allowing EA to accommodate the attachment of two polymerases for bidirectional extension, as described herein. That is, the nick sites 201a and 201b are spaced sufficiently far apart that the binding of one polymerase does not spatially impede and / or displace the binding of the second polymerase. This configuration is illustrated, for example, in Figure 2A, where the spacer region 202 linearly offsets the first nick site 201a from the second nick site 201b.
[0111] Figure 2B is a schematic diagram depicting the circularization of a target DNA template using YBEA 200 according to certain exemplary embodiments. Referring to Figure 2B, YBEA 200 and its corresponding first Y branching element 213a and second Y branching element 213b associated with ends 203 and 204, respectively, are combined with the target DNA template 207 to form a circular construct 209 (similar to the formation of the circular construct 109 in Figure 1B). That is, YBEA 200 is combined with the target DNA template 207, which has a first end 205 and a second end 206, and includes complementary strands 207a and 207b (Figure 2B).
[0112] At step 2a, for example, YBEA 200 is ligated to either end of the target DNA template 207, which includes parental polynucleotide chains 207a and 207b. For example, end 203 of YBEA 200 is ligated to end 206 of the template. Alternatively, at step 2a, and although not shown for simplicity, the other end (i.e., 204) of YBEA 200 is ligated to end 205 of the target DNA template 207.
[0113] In step 2b of Figure 2B, the unconnected (free) end of YBEA 200 is connected to the remaining free end of the target DNA template 207 to form a circular construct 209. For example, if the end 203 of YBEA 200 is connected to the template end 206 in step 2a, then the YBEA end 204 of YBEA 200 is connected to the template end 205 in step 2b, thereby forming a circular construct 209 of the original (parental) target DNA template 207. Alternatively, if the end 204 of YBEA 200 is connected to the template end 205 in step 2a, then the YBEA end 203 is connected to the template end 206 in step 2b, thereby forming a circular construct 209.
[0114] In either case, at step 2b of Figure 2B, the two ends 203 and 204 of YBEA 200 connect to each end 205 and 206 of template 207, thereby forming a YBEA bridging region 208 between the ends of template 207. That is, the entire YBEA 200 forms a bridging region 208 between the two ends 205 and 206 of the parental target DNA template 207. This forms a circular construct 209 comprising YBEA 200 (as bridging region 208) and complementary parental template strands 207a and 207b of the parental target DNA template 207. In this way, YBEA 200 (Figure 2A) operates as a bridging precursor for bridging region 208. Furthermore, the corresponding first nick site 201a and second nick site 201b are retained in the circular construct 209 (as part of the bridging region 208), and thus provide two corresponding free 3' ends for polymerase attachment and bidirectional extension, as described herein. Y-branching elements 213a and 213b are also present in the circular construct 209.
[0115] Continuing with the above examples, Figure 2C is a schematic diagram depicting the initiation of polymerase attachment and bidirectional extension of a circular construct including YBEA200 according to certain exemplary embodiments. As shown in Figure 2C (and similar to Figure 1C), when the first polymerase 210a and the second polymerase 210b are combined with the circular construct 209, they bind to nick sites 201a and 201b, respectively, in opposite directions (as indicated by the arrows in Figure 2C, upper view). Polymerases 210a and 210b also replace the 5' ends of the parental strands 207b and 207a, respectively (including their associated and respective 213b and 213a Y-branching elements).
[0116] At step 2c (Figure 2C, below), polymerases 210a and 210b extend the circular construct 209 bidirectionally in opposite directions (see arrows). That is, as polymerase 210a proceeds, it uses the parental strand 207b as a template to extend the 3' end of the nick site 201a to synthesize a new daughter strand 207a' that is sequence-complementary to (and therefore shares the same sequence as) the parental strand 207b. The new daughter strand 207a' also includes a replicated ssDNA daughter strand bridge portion 208a as part of the bridging region 208. Furthermore, the Y-branching element 213a is retained at the 5' end of the replaced template strand 207a.
[0117] Similarly, polymerase 210b extends the 3' end of cleavage site 201b while simultaneously using parental strand 207a as a template to replace the 5' end of cleavage site 201b (and its associated parental template strand 207b) (Figure 2C, bottom). Thus, as polymerase 210b proceeds in step 2c (bottom), it uses parental strand 207a as a template to extend the 3' end of cleavage site 201b to synthesize a new daughter strand 207b' that is sequence-complementary to (and therefore shares with) the same sequence of parental strand 207b. As shown, the Y-branch element 213b also remains attached to the replaced parental strand at the 5' end. The new daughter strand 207b' also includes a replicated ssDNA daughter strand bridge portion 208b as part of the bridging region 208.
[0118] Continuing with the exemplary embodiments described above, Figure 2D is a schematic diagram depicting the continuous polymerase extension and double-length DNA template formation of a circular construct using YBEA200 according to certain exemplary embodiments. As shown in Figure 2D (top), polymerase 210a continues along the parental template strand 207b to the 5' end of the parental template strand 207b, completing the synthesis of a new daughter strand 207a'. Similarly, as shown in Figure 2D (top), polymerase 210b continues along the parental template strand 207a to the 5' end of the parental template strand 207a, completing the synthesis of a new daughter strand 207b'. Notably, the new daughter strand 207a' includes a Y-branching element daughter strand 213b', which is complementary to the Y-branching element 213b attached to the parental strand 207b. Similarly, the new offspring chain 207b' includes a Y-branching element offspring chain 213a', which is complementary to the Y-branching element 213a attached to the parent chain 207a.
[0119] At step 2d in Figure 2D, once the first polymerase 210a and the second polymerase 210b have completed the synthesis of daughter strands 207a' and 207b', respectively, polymerases 210a and 210b dissociate from the circular construct 209 to form a double-length DNA template 211 (Figure 2D, bottom). As shown in Figure 2D (bottom), the double-length DNA template 211 comprises two copies of the target DNA template 207, namely, the first copy 211a and the second copy 211b, each located on the flanking side of the bridging region 208 (the bridge originates from YBEA 200 and includes the bridged daughter strand portions 208a and 208b).
[0120] As shown in the figure, each template copy 211a and 211b includes both a parental polynucleotide strand (shown in black) and a newly synthesized daughter polynucleotide strand (shown in gray). For example, template copy 211a includes the original (parental) template strand 207b and the newly synthesized daughter strand 207a'. On the other side of the bridging region 208, template copy 211b includes the original (parental) template strand 207a and the newly synthesized daughter strand 207b'. Furthermore, the double-length DNA 211 template includes a first end 212a and a second end 212b. For example, the first end 212a includes the portion of the YBEA 200 strand 200b associated with the 5' end of the nick site 201b YBEA 200 (a hollow black rectangle at end 212a) and its copy (a hollow gray circle at end 212a). Similarly, the second end 212b of the double-length DNA template 211 includes the portion of YBEA200 strand 200a associated with the 5' end of the nick site 201a YBEA 200 (the hollow black circle at 212b) and its copy (the hollow gray rectangle at end 212b).
[0121] As also shown in the figure, template copy 211a also includes a Y branching element 213b and its complementary sequence in the Y branching element progeny strand 213b' at its end 212a, while template copy 211b includes a Y branching element 213a and its complementary sequence in the Y branching element progeny strand 213a' at its end 212b (Figure 2D, bottom). That is, the Y branching elements 213a and 213b of YBEA 200 are present at the ends of the double-length DNA template 211, as derived from YBEA 200 as described herein. In this way, YBEA 200 and the parental target DNA template 207 are combined via steps 2a-2d of Figures 2B to 2D to produce a double-length DNA template 211 (Figure 2D, bottom), which includes predetermined oligonucleotide sequences at each end.
[0122] Furthermore, similar to the example double-length DNA template 111 shown in Figure 1D (below), the two strands of the double-length DNA template 211 in Figure 2D also include a parental polynucleotide strand (black) covalently linked to a newly synthesized daughter copy (grey) of the target DNA template. For example, parental strand 207a is covalently and continuously linked to the newly synthesized daughter strand 207a' in the 5'→3' direction via the strand of bridging region 208 (i.e., the strand of bridging region 208 including strand portions 200a and 208a). Moreover, due to polymerase-mediated elongation of the circular construct as described herein, the nucleotide sequence of parental strand 207a matches the nucleotide sequence of the new daughter strand 207a'. That is, daughter strand 207a' is a sequence copy of parental strand 207a of the target DNA template.
[0123] Similarly, parental strand 207b is covalently and continuously linked to the new daughter strand 207b' in the 5'→3' direction via the strand of bridging region 208 (i.e., the strand including strand portions 200b and 208b of bridging region 208). Likewise, the nucleotide sequence of parental strand 207b matches the nucleotide sequence of the new daughter strand copy 207b'. That is, daughter strand 207b' is a sequence copy (i.e., a daughter copy) of the parental strand 207b of the target DNA template. In this way, each strand of the double-length DNA includes both a copy of the parental polynucleotide sequence and a copy of the daughter polynucleotide sequence on each of the two target DNA copies 211a and 211b (Figure 2D, bottom).
[0124] As described herein, in some exemplary embodiments, Y branch elements 213a and 213b can be used to facilitate amplification, such as PCR amplification. Therefore, Figure 2E is an illustration of a denatured (single-stranded) form of the double-length DNA template 212 of Figure 2D (below) according to some exemplary embodiments, wherein the original Y branch elements 213a and 213b provide predetermined oligonucleotide primer-binding sequences during replication. For example, when Y branch element 213a replicates to 213a' (Figure 2D), the progeny Y branch element 213a' includes a primer-binding site at or within Y branch element 213a'. Similarly, when Y branch element 213b replicates to 213b' (Figure 2D), the progeny Y branch element 213b' includes a primer-binding site at or within Y branch element 213b'. As shown in Figure 2E, after the PCR denaturation step, primer 214 binds to the primer binding site of Y-branch element 213a', while primer 215 binds to the primer binding site of Y-branch element 213b'. And as with conventional PCR protocols and procedures, primers 214 and 215 provide the 3' end for polymerase extension. In this way, the YBEA 200 example can be used to replicate the template strand and then provide primer binding sites for conventional PCR-based template amplification.
[0125] In some exemplary embodiments, primers 214 and 215 have the same sequence and therefore bind the same sequence within their respective primer binding sites on Y branch elements 213a' and 213b'. That is, primers 214 and 215 are identical. Alternatively, in some exemplary embodiments, primers 214 and 215 have different sequences and therefore bind to different sequences within their respective primer binding sites on Y branch elements 213a' and 213b'. Thus, YBEA and its associated Y branch elements provide a unique ability to customize the replication of the target DNA template strand for downstream applications such as PCR amplification.
[0126] Epigenetic analysis using double-length DNA templates
[0127] As described herein, in some exemplary embodiments, the target DNA template includes a natural target sequence. Therefore, in some exemplary embodiments, the target DNA template may retain epigenetic information about the target sequence, such as the methylation pattern of the target sequence. And because the parental polynucleotide strand of the target DNA template is retained in a double-length DNA template as described herein—that is, each strand of the double-length DNA template includes a parental polynucleotide strand from the target DNA template (referred to in the context of a “parental copy” of the target sequence)—the double-length DNA template also retains epigenetic information from the target sequence.
[0128] Furthermore, progeny copies of the target sequence can be synthesized while preserving the genetic information of the target sequence, as further described herein. Therefore, the presence of both parental and progeny copies of the target sequence on the same strand of a double-length DNA template is particularly advantageous for "intra-strand" comparisons to identify epigenetic information. And because each parental copy of the target DNA template in the double-length DNA template also hybridizes with a complementary progeny sequence, this arrangement also allows for "inter-strand" comparisons to identify epigenetic information in some exemplary embodiments. This dual approach of comparing parental and progeny sequences advantageously increases the accuracy and confidence of the epigenetic information detected in the target sequence. These and other exemplary embodiments are shown and described with reference to Figures 3A to 3G.
[0129] To facilitate intra- and / or inter-strand comparisons, in some exemplary embodiments, the end-adaptors provided herein, such as YBEA of FIG. 2A, may be further modified to provide features capable of bioinformatically grouping sequence reads. For example, end-adaptors, such as YBEA of FIG. 2A, may be modified to include a unique molecular identifier (UMI). In some exemplary embodiments, the UMI may be included within a spacer region 202 of the YBEA, in which case the UMI sequences of the two strands of the double-length DNA template will have reverse complementary sequences. Such modified end-adaptors—and their use in methods for performing epigenetic analysis of the target DNA template strand and thus the target sequence—are shown in FIGS. 3A through 3F.
[0130] Referring to Figure 3A, an illustration is provided showing a Y-branched adapter (“YB-UMI-EA”) including UMI according to certain exemplary embodiments. As shown, YB-UMI-EA 300 has a general double-stranded polynucleotide YBEA structure as shown in Figure 2A. This includes, for example, hybridized strands 300a and 300b, where “300a” refers to the entire 5’→3’ strand of YBEA 300 (with cleavage site 301a within strand 300a), and “300b” refers to the entire 5’→3’ strand hybridized with strand 300a (with cleavage site 301b within strand 300b). Furthermore, single-stranded Y-branching elements 313a and 313b are attached to the 5’ ends of the first cleavage site 301a and the second cleavage site 301b of YB-UMI-EA, respectively. For example, Y-branch elements 313a and 313b may include predetermined oligonucleotide sequences whose design can be tailored to achieve specific purposes, such as, for example, PCR amplification, as described above with respect to Figures 2A to 2E.
[0131] For example, and as shown in FIG3A, each Y branch element 313a and 313b may include a predetermined oligonucleotide sequence that provides a complementary or hybridizable primer binding site for PCR amplification of the new daughter strand. That is, each Y branch element 313a and 313b may, for example, include 10 to 30 nucleotides, such as 15 to 25 nucleotides or 18 to 22 nucleotides, whose complementary sequence includes the primer binding site sequence. In some exemplary embodiments, Y branch elements 313a and 313b include the same sequence, as shown in FIG3A (at 313a and 313b). Alternatively, Y branch elements 313a and 313b include different sequences. In some exemplary embodiments, Y branch elements 313a and 313b have the same length, while in other exemplary embodiments, Y branch elements 313a and 313b may have different lengths.
[0132] YB-UMI-EA 300 also includes ends 303 and 304 located flanking each nick site, each end 303 and 304 being compatible with the end of the target DNA template for efficient ligation. Preferably, nick sites 301a and 301b are separated by a double-stranded spacer region 302, allowing EA to accommodate the attachment of two polymerases for bidirectional extension, as described herein. That is, nick sites 301a and 301b are spaced sufficiently far apart that the binding of one polymerase does not spatially impede and / or displace the binding of the second polymerase. This is illustrated in Figure 3A, where the spacer region 302 linearly offsets the first nick site 301a from the second nick site 301b.
[0133] As also shown in the figure, for example, the UMI sequence 316 is located within the spacer region 302 of YB-UMI-EA 300. A UMI, also known as a molecular barcode or random barcode, comprises a short, random, and / or predetermined nucleotide sequence incorporated into an oligonucleotide sequence. Typically, a UMI is 5 to 20 nucleotides in length, such as 8 to 16 nucleotides. Of course, this length can vary depending on the application. For example, a UMI can have a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. More conventionally, UMIs consist of sequences of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.
[0134] As shown in Figure 3A, because YB-UMI-EA 300 is double-stranded, the UMI sequence 316 has complementary UMI strand sequences 316a and 316b, where strand 316a exhibits 5'→3' polarity, while strand 316b has complementary 3'→5' polarity. Furthermore, because UMI strand sequences 316a and 316b are complementary sequences, the sequences of different strands in the resulting double-length DNA template can be bioinformatically paired to enable the comparison of sequencing reads from different strands as described herein.
[0135] Referring to Figure 3B, a schematic diagram illustrating the circularization of a target DNA template using YB-UMI-EA 300 according to certain exemplary embodiments is provided. As shown, YB-UMI-EA 300, along with its corresponding first single-stranded Y branching element 313a and second single-stranded Y branching element 313b, and double-stranded UMI sequence 316, is combined with a double-stranded target DNA template 307 to form a circular construct 309. That is, YB-UMI-EA 300 is combined with the target DNA template 307, which has a first end 305 and a second end 306 and includes complementary strands 307a and 307b. In step 3a, for example, YB-UMI-EA 300 is ligated to either end of the target DNA template 307, which includes polynucleotide chains 307a and 307b. For example, end 303 of YB-UMI-EA 300 is ligated to end 306 of the template. Alternatively, at step 3a, and although not shown for simplicity, the other end (i.e., 304) of YB-UMI-EA 300 is ligated to end 305 of the target DNA template 307.
[0136] At step 3b of Figure 3B, the unconnected (free) end of YB-UMI-EA 300 is connected to the other free end of the target DNA template 307 to form a circular construct 309 (or a circular construct). For example, if end 303 of YB-UMI-EA 300 is connected to template end 306 at step 3a, then end 304 of YB-UMI-EA 300 is connected to template end 305 at step 3b, thereby forming a circular construct 309 of the original (parental) target DNA template 307. Alternatively, if end 304 of YB-UMI-EA 300 is connected to template end 305 at step 3a, then end 303 is connected to template end 306 at step 3b, thereby forming a circular construct 309.
[0137] In either case, at step 3b of Figure 3B, the two ends 303 and 304 of YB-UMI-EA 300 connect to each end 305 and 306 of template 307, thereby forming a YB-UMI-EA bridging region 308 between the ends of template 307. That is, the entire YB-UMI-EA 300 forms a bridging region 308 between the two ends 305 and 306 of the target (parental) template 307. This forms a ring-shaped construct 309 (or a ring-shaped construct) comprising YB-UMI-EA 300 (as bridging region 308) and complementary parental template chains 307a and 307b of the parental template 307. In this way, YB-UMI-EA 300 (Figure 3A) operates as a bridging precursor for bridging region 308. Furthermore, the corresponding first nick site 301a and second nick site 301b are retained in the circular construct 309 (as part of the bridging region 308), and thus provide two corresponding 3' ends for polymerase attachment and bidirectional extension, as described herein. Single-stranded Y-branching elements 313a and 313b are also present in the circular construct 309 along with the UMI sequence 316.
[0138] Referring to Figure 3C, an enlarged view is provided of a portion of a target DNA template 307 according to certain exemplary embodiments, illustrating an example nucleic acid sequence. As shown, a portion of the target DNA template strand 307a includes endogenously methylated cytosine residues (i.e., 5-methylcytosine or "5mc") at nucleotide positions 3, 8, and 10 of the example sequence and an unmethylated cytosine residue at position 5 (arrow) (when reading strand 307a from left to right, i.e., in the 5'→3' direction of strand 307a). However, the cytosine residue at position 5 (see the asterisk in strand 307a) is unprotected (i.e., unmethylated). Furthermore, the target DNA template strand 307b, complementary to the target DNA template strand 307a, includes an endogenously methylated cytosine residue at position 9 (when reading from left to right, i.e., in the 3'→5' direction of strand 307b). The cytosine residue at position 6 (see asterisked chain 307b) is unprotected (i.e., unmethylated). In this example, the 5mC residue represents the epigenetic methylation pattern of the parental target DNA template.
[0139] Continuing with the exemplary embodiments of YB-UMI-EA 300, Figure 3D is a schematic diagram depicting the initiation of polymerase attachment and bidirectional extension of a circular construct including YB-UMI-EA 300 according to certain exemplary embodiments. As shown in Figure 3D (and similar to Figure 2C), when the first polymerase 310a and the second polymerase 310b are combined with the DNA circular construct 309, they bind to cleavage sites 301a and 301b, respectively, in opposite directions (as indicated by the arrows in the upper part of Figure 3D). Polymerases 310a and 310b also replace the 5' ends of the parental strands 307a and 307b, respectively (including their associated and respective 313a and 313b Y-branching elements). The UMI strand sequences 316a and 316b of UMI 316 remain unchanged.
[0140] At step 3c (Figure 3D, bottom), polymerases 310a and 310b extend the circular construct 309 bidirectionally in opposite directions (see arrows). That is, as polymerase 310a proceeds, it uses the parental strand 307b as a template to extend the 3' end of the nick site 301a to synthesize a new daughter strand 307a' that is sequence-complementary to (and therefore shares the same sequence as) the parental strand 307b. The new daughter strand 307a' also includes a replicated ssDNA daughter strand bridge portion 308a as part of the bridging region 308. Furthermore, the Y-branch element 313a remains at the 5' end of the replaced template strand 307a, while the UMI sequence 316 and its strand sequences 316a and 316b remain unchanged.
[0141] Similarly, polymerase 310b extends the 3' end of nick site 301b while simultaneously using parental strand 307a as a template to replace the 5' end of nick site 301b (and its associated parental template strand 307b) (Figure 3D, bottom). Thus, as polymerase 310b proceeds in step 3c (bottom), it uses parental strand 307a as a template to extend the 3' end of nick site 301b to synthesize a new daughter strand 307b' that is sequence-complementary to (and therefore shares with) the same sequence as parental strand 307b. As shown, the Y-branch element 313b is also attached at the 5' end to the replaced parental strand 307b. The new daughter strand 307b' also includes a replicated ssDNA daughter strand bridge portion 308b as part of the bridging region 308. Similarly, UMI sequence 316 and its chain sequences 316a and 316b remain unchanged because they are not replicated during chain extension.
[0142] Referring to Figure 3E, a schematic diagram is provided depicting the continuous polymerase extension of a circularized target DNA template and the formation of a double-length DNA template using an exemplary embodiment of the YB-UMI-EA 300 according to certain exemplary embodiments. As shown in Figure 3E (top), polymerase 310a continues along the parental template strand 307b to the 5' end of the parental template strand 307b, completing the synthesis of a new daughter strand 307a'. Similarly, as shown in Figure 3E (top), polymerase 310b continues along the parental template strand 307a to the 5' end of the parental template strand 307a, completing the synthesis of a new daughter strand 307b'. Notably, the new daughter strand 307a' includes a Y-branching element daughter strand 313b', which is complementary to the Y-branching element 313b attached to the parental strand 307b. Similarly, the new offspring chain 307b' includes a Y-branching element offspring chain 313a', which is complementary to the Y-branching element 313a attached to the parent chain 307a. The UMI sequence 316 and its chain sequences 316a and 316b remain unchanged.
[0143] At step 3d in Figure 3E, once the first polymerase 310a and the second polymerase 310b have completed the synthesis of daughter strands 307a' and 307b', respectively, polymerases 310a and 310b dissociate from the circular construct 309, forming a double-length DNA template 311 (Figure 3E, bottom). As shown in Figure 3E (bottom), the double-length DNA template 311 comprises two copies of the original (parental) target DNA template 307, namely, the first copy 311a and the second copy 311b, each located flanking the bridging region 308. The bridging region 308, derived from YB-UMI-EA 300, includes the bridged daughter strand portions 308a and 308b, as well as the unchanged UMI 316 and its complementary strand sequences 316a and 316b.
[0144] As shown in the figures, each template copy 311a and 311b includes both a parental polynucleotide strand and a newly synthesized daughter polynucleotide strand. For example, template copy 311a includes the original (parental) template strand 307b of the target DNA template 307 and the newly synthesized daughter strand 307a' (Figures 3B and 3G). On the other side of the bridging region 308, template copy 311b includes the original (parental) template strand 307a of the target DNA template 307 and the newly synthesized daughter strand 307b' (Figures 3B and 3G). Furthermore, the double-length DNA template 311 includes a first end 312a and a second end 312b. For example, the first end 312a includes the portion of YB-UMI-EA 300 chain 300b associated with the 5' end of the nick site 301b YB-UMI-EA 300 (the hollow black rectangle at end 312a) and its copy (the hollow gray circle at end 312a). Similarly, the second end 312b of the double-length DNA template 311 includes the portion of YB-UMI-EA 300 chain 300a associated with the 5' end of the nick site 301a YB-UMI-EA 300 (the hollow black circle at 312b) and its copy (the gap gray rectangle at end 312b).
[0145] As also shown in the figure, template copy 311a further includes a Y-branching element 313b and its complementary sequence in the Y-branching element progeny strand 313b' at the first end 312a, while template copy 311b includes a Y-branching element 313a and its complementary sequence in the Y-branching element progeny strand 313a' at the second end 312b (Figure 3E, bottom). In this way, YB-UMI-EA 300 and target DNA template strand 307 are combined via steps 3a-3d of Figures 3B to 3D to produce a double-length DNA template 311 (Figure 3E, bottom), which includes a predetermined oligonucleotide sequence at each end, as well as UMI 316 (and its strand sequences 316a and 316b).
[0146] Similar to the double-length DNA templates 111 and 211 in Figures 1D and 2D, the double-length DNA template 311 shown in Figure 3D also includes a parental copy (black) linked to a newly synthesized daughter copy (grey) of the target DNA template. For example, parental copy 307a is covalently and continuously linked to the newly synthesized daughter copy 307a' in the 5'→3' direction via the strand of bridging region 308 (i.e., the strand including strand portions 300a and 308a of bridging region 308). Furthermore, the nucleotide sequence of parental copy 307a matches the nucleotide sequence of the new daughter copy 307a' due to polymerase-mediated elongation of the circular construct as described herein.
[0147] Similarly, parental copy 307b is covalently and continuously linked to the new daughter copy 307b' in the 5'→3' direction via the strand of bridging region 308 (i.e., the strand including strand portions 300b and 308b of bridging region 308). Likewise, the nucleotide sequence of parental copy 307b matches the nucleotide sequence of the new daughter copy 307b'. In this way, each strand of the double-length DNA comprises both parental template DNA and daughter copy DNA, except for the parental template DNA that hybridizes with the complementary daughter DNA in each of the two target DNA copies (Figure 3E, bottom).
[0148] In some exemplary embodiments, the double-length DNA template of Figure 3E (below) can be advantageously used to identify epigenetic information associated with the parental target DNA template 307. For example, during the extension steps of polymerases 310a and 310b described in Figures 3E to 3F, protected nucleotides (such as methylated cytosine nucleotide residues) can be used for daughter strand extension and synthesis. This results in the incorporation of protected nucleotides (such as methylated cytosine nucleotide residues) into the newly synthesized daughter strands 307a' and 307b'. In such embodiments, the daughter strands—having protected cytosine residues—retain the genetic information of the target DNA template during, for example, a bisulfite treatment process that converts native cytosine to uracil. Subsequently, following the bisulfite conversion reaction and DNA sequencing, as further described herein, bioinformatics analysis of the sequence information can be performed to identify the methylated cytosine residues in the original (parental) target DNA template.
[0149] In some exemplary embodiments, the identification of methylated cytosine residues in the original (parental) target DNA template 307 provides epigenetic information associated with the original (parental) target DNA template 307. This is illustrated, for example, in Figure 3F, which provides a schematic diagram of an example bisulfite conversion of a double-length DNA template 311 and its subsequent PCR amplification product via a Y-branched adaptor with UMI (i.e., YB-UMI-EA 300) as shown in Figure 3A.
[0150] Referring to Figure 3F (and prior to step 3e), an example truncated portion of the double-length DNA 311 of Figure 3E is shown, which includes two copies of the example template sequence portion shown in Figure 3C. That is, Figure 3F Only a portion of the sequence of the original target DNA template 311 is shown (for illustrative purposes only), and the portion shown includes two copies (311a and 311b) of the example sequence according to Figure 3C. As shown, each strand includes a parent copy (black) and a daughter copy (grey) in its intrastrand arrangement, for example, copies generated by the formation of a double-length DNA template as described herein.
[0151] As also shown in Figure 3F, each template copy portion 311a and 311b includes 10 example nucleotide pairs, corresponding to the nucleotide pairs in Figure 3C. Due to the formation of a double-length DNA template, the nucleotide sequences are mirrored on each side of the double-length DNA template. For example, the sequence of the parental strand 307a (of template copy 311b) corresponds to the same sequence on the daughter strand 307a' (of template copy 311a), where both sequences 307a and 307a' are associated with the UMI strand sequence 316a. For example, the polynucleotide sequence associated with the UMI strand 316a from left to right (i.e., 5'→3') is the example parental copy (black) TACACGACGC (SEQ ID NO:1), while the daughter copy (gray) is the same polynucleotide sequence, i.e., TACACGACGC. Therefore, the example 5'→3' sequence associated with UMI 316a is TACACGACGC--UMI--TACACGACGC.
[0152] Similarly, considering the complementary base pairing of the strands, the sequence of the parental strand 307b (of template copy 311a) corresponds to the same sequence on the daughter strand 307b' (of template copy 311b), but both sequences 307b and 307b' are associated with the UMI strand sequence 316b. That is, the sequence associated with the UMI strand 316b from left to right (i.e., 3'→5') is, for example, the daughter strand sequence (grey) ATGTGCTGCG (SEQ ID NO:2), while the parental sequence (black) is also ATGTGCTGCG. In other words, the example 5'→3' sequence associated with UMI 316b is ATGTGCTGCG--UMI--ATGTGCTGCG. Therefore, each UMI strand sequence 316a and 316b of UMI 316 is associated with a portion of the parental strand (black) and the new daughter strand (grey) (Figure 3F).
[0153] As illustrated in this example epigenetic assessment of the target DNA template strand, prior to step 3e, the parental strand 307a of template copy 311b includes endogenously methylated (protected) cytosine residues at positions 3, 8, and 10, and an unmethylated cytosine residue at position 5 (arrow) (from left to right, i.e., 5'→3', and also shown in Figure 3C). Furthermore, the parental strand 307b of template copy 311a includes endogenously methylated (protected) cytosine residues at position 9, and an unmethylated residue at position 6 (also shown in Figure 3C, when reading from 3'→5'). However, each daughter strand 307a' and 307b' includes only methylated cytosine residues, a result of polymerase elongation, which provides only methylated cytosine residues during the elongation reaction. That is, neither daughter strand 307a' nor 307b' contains unmethylated cytosine residues. Therefore, the protected progeny strands 307a' or 307b' (gray) retain the genetic information of the target DNA template, while the natural parental target DNA template strands 307a and 307b (black) retain the epigenetic information of the target DNA template (and thus the target sequence).
[0154] At step 3e of Figure 3F, a double-length DNA template is subjected to bisulfite conversion using conventional methods. For example, bisulfite conversion is a method that uses bisulfite to determine DNA methylation patterns, such as the methylation of the target DNA template. DNA methylation, for example, is an endogenous biochemical process involving the addition of methyl groups to cytosine or adenine DNA nucleotides. For example, DNA methylation stably alters gene expression in cells as they divide from embryonic stem cells and differentiate into specific tissues. In bisulfite conversion (also known as bisulfite sequencing), the target nucleic acid is first treated with a bisulfite reagent that specifically converts unmethylated cytosine residues to uracil residues (i.e., C→U conversion), while having no effect on methylated cytosine residues (i.e., methylated cytosine residues are "protected" from C→U conversion). Subsequently, PCR reactions using natural adenine (A), cytosine (C), guanine (G), and thymine (T) nucleotides were performed, with the converted uracil residues replaced by thymine residues (i.e., U→T substitution). In this way, unmethylated (i.e., "unprotected") cytosine residues were converted to thymine via the intermediate uracil (i.e., C→U→T).
[0155] Therefore, as shown in step 3e of Figure 3F, the denatured double-length DNA template strand undergoes a bisulfite conversion reaction, resulting in the conversion of the unmethylated cytosine residue at position 5 of the parental strand 307a to a uracil residue (as shown in bold and underline, SEQ ID NO:3), i.e., 5C→5U conversion. Similarly, the unmethylated cytosine residue at position 6 of the parental strand 307b is converted to uracil (as shown in bold and underline, SEQ ID NO:4), i.e., 6C→6U conversion. However, the bisulfite reaction does not affect any methylated (protected) cytosine residues in the parental strands 307a and 307b', i.e., these cytosine residues remain cytosine residues. This includes the methylated cytosine residues in the daughter strands 307a' and 307b', which are included in the daughter strands 307a' and 307b' via polymerase extension using methylated cytosine nucleotides, as described herein.
[0156] Following the bisulfite conversion reaction in step 3e, at step 3f in Figure 3F, the bisulfite-converted strand of the denatured double-length DNA template 311 product is subjected to PCR amplification and sequencing using conventional methods. For example, PCR primers targeting Y-branching elements 313a' and 313b' can be used to amplify the denatured double-length DNA template 311 strand, as described herein. As shown at step 3f, and as can be determined via conventional DNA sequencing, the PCR product (shown in a denatured state for illustrative purposes) produces distinct strands, each associated with either the UMI strand sequence 316a or 316b. In the PCR product, uracil residues resulting from the bisulfite conversion of unmethylated (unprotected) cytosine residues are replaced by thymine. For example, during the PCR reaction, the uracil residue at position 5 of parental strand 307a is replaced by a thymine (T) residue (see arrow), i.e., 5U→5T substitution. Furthermore, the 5U→5T substitution of parental strand 307a is associated with UMI strand sequence 316a. Similarly, during the PCR reaction, the uracil residue at position 6 of parental strand 307b is replaced by a thymine (T) residue (see arrow), i.e., 6U→6T conversion. Furthermore, the 6U→6T substitution of parental strand 307b is associated with UMI strand sequence 316a. Therefore, in this example, each strand of UMI (316a and 316b) is localized using strand-specific nucleotide conversions (C→U→T) of the original (parental) DNA template.
[0157] At step 3g, following the PCR reaction in step 3f, the PCR product is sequenced. The resulting sequencing reads are used to identify the methylation pattern of the original parental copy of the DNA target sequence by intrastrand comparison of the parental and daughter sequences. That is, daughter strand copies with protected cytosine residues are resistant to bisulfite transformation and therefore retain the genetic sequence of the parental template. Therefore, where the original parental strand sequence includes each position of native (unmethylated) cytosine, the entire strand sequence read will indicate the difference between the parental and daughter sequences; conversely, where the parental strand sequence includes each position of methylated cytosine, the entire strand sequence read will show the consistency between the parental and daughter sequences.
[0158] Alternatively, comparisons of complementary parental-derived and daughter strand sequences (i.e., inter-strand comparisons) can also be used to identify and / or confirm parental sequence methylation patterns. That is, comparisons of parental-derived and daughter strand sequences from different strands of a double-length DNA template (achieved through bioinformatics grouping of UMI read sequences) will reveal mismatches between paired bases at the native cytosine position in the parental sequence, while the methylated cytosine position will show normal complementarity with the daughter sequence. Such intra-strand and inter-strand comparisons and analyses are illustrated in Figure 3G, where any of these methods, independently or in combination, are used to assess epigenetic information associated with the original target template sequence.
[0159] Referring to Figure 3G (before step 3h), a schematic diagram is provided illustrating intra-strand and inter-strand comparisons of a portion of a double-length DNA template according to certain exemplary embodiments to determine epigenetic information associated with the original target DNA template. In this example schematic diagram, the same exemplary sequence continues from Figure 3F, and the strands are shown in a double-length DNA dual-template configuration for illustrative purposes only. As shown, for intra-strand analysis, for example—when reading the sequence of the strand associated with UMI sequence 316a in the 5'→3' direction (i.e., from the 5' end of strand fragment 307a to the 3' end of strand fragment 307a' from left to right)—the intra-strand TC difference is identified at the fifth nucleotide position (see the arrow associated with UMI sequence 316a). That is, the sequence of strand fragment 307a (black) includes thymine (T) residues, while the sequence of strand fragment 307a' (gray) includes cytosine (C) residues. Importantly, because the daughter strands were synthesized using protected cytosine analogs, intra-strand TC differences identified strand fragment 307a' as a daughter copy of the target DNA template and fragment 307a as a parent copy of the target DNA template.
[0160] Similarly, in the strand associated with UMI sequence 316b, when reading the sequence of the strand associated with UMI strand 316b in a 3'→5' direction (i.e., from left to right from the 3' end of strand fragment 307b' to strand fragment 307b), the intra-strand TC difference is identified at the sixth nucleotide position (see the arrow associated with UMI sequence 316b). That is, the sequence of strand fragment 307b (black) includes thymine residues, while the sequence of strand fragment 307b' (gray) includes cytosine residues. And as with the TC difference associated with UMI sequence 316a discussed above, the presence of the cytosine residue at the sixth position in strand fragment 307b' identifies this strand as the daughter strand (gray), where strand fragment 307b (black) is the parental-derived strand. Furthermore, as described more fully below, the presence of a substituted thymine residue at the sixth position of chain 307b indicates that the thymine nucleotide is an unprotected cytosine residue in the original target sequence.
[0161] Alternatively or concurrently, prior to step 3h, in some exemplary embodiments, analysis of interstrand mismatches may be used to identify, assess, and / or confirm epigenetic information associated with the original target sequence. As shown in the figure, for example, interstrand alignment of the sequence of the example parental strand 307a (black) with the sequence of the daughter strand 307b' (gray) reveals a TG mismatch at position 5 of the 307a / 307b' alignment sequence. Based on the presence of this mismatch, it can also be determined that the sequence of strand 307a corresponds to the parental target sequence. This is because only unmethylated (unprotected) cytosine residues undergo C→U→T bisulfite / PCR conversion, and because daughter strand extension with methylated (protected) cytosine residues only incorporates the protected cytosine residues into the daughter strand. Therefore, during bisulfite / PCR conversion, only the unprotected cytosine residues in the parental strand are converted to thymine residues (i.e., not those in the daughter strand). Once fragment 307a is identified as a parental copy, when reading from left to right (i.e., 5'→3'), this parental copy can be identified as associated with the 5' end of UMI sequence 316a, with the offspring fragment 307a' located downstream of the 3' end of UMI 316a (as shown in the figure).
[0162] Similarly, alignment of the sequence of the example parental strand fragment 307b (black) with the sequence of the example daughter strand fragment 307a' (gray) reveals a TG mismatch at position 6 of the 307b / 307a' alignment sequence. Therefore, the presence of thymine residues in the TG mismatch identifies strand fragment 307b as the parental-derived strand, while strand fragment 307a' is the complementary daughter strand. Thus, once strand 307b is identified as the parental-derived copy, when reading from left to right (i.e., 3'→5'), this parental-derived copy can be identified as associated with the 3' end of UMI strand fragment 316a, with daughter strand fragment 307b' located upstream of the 5' end of UMI 316b, as illustrated. In some exemplary embodiments, this inter-strand and intra-strand analysis can be used to identify and confirm methylation patterns across multiple sequence reads due to UMI-based read grouping. This is particularly beneficial, for example, when large regions of the target sequence—such as those preserved in the target DNA template—include methylated cytosine residues.
[0163] At step 3h in Figure 3G, based on the intra-strand and / or inter-strand analyses described herein, protected (methylated) cytosine residues associated with the original (parental) target DNA template 307 can be identified. This, in turn, provides epigenetic information about the original (parental) DNA template 307. For example, as described above, C→U→T bisulfite / PCR transformation occurs only in the presence of unprotected (native) cytosine residues. Therefore, the presence of any cytosine residues in the strand identified as corresponding to the original (parental) target DNA template strand (i.e., strand fragments 307a and 307b in the above examples, shown in black) can be identified as previously protected (methylated) cytosine residues. This is shown, for example, after step 3h, where arrows indicate the identification of protected cytosine residues in the original parental (target) template (e.g., template 307). As shown for chain segment 307a, for example, previously protected cytosine residues are present at positions 3, 8, and 10 (see arrows, reading the chain from left to right). Similarly, the cytosine residue at position 9 of chain segment 307b can also be identified as previously protected (see arrows, reading the chain from left to right).
[0164] Finally, at step 3i of Figure 3G, in some exemplary embodiments, strand fragments (e.g., 307a and 307b) of the parental strand identified as corresponding to the original (parental) target DNA template 311—and their associated methylation patterns—can be aligned to reveal the epigenetic patterns associated with the original (parental) target DNA template 307. That is, epigenetic information associated with the original (parental) target DNA template 307 can be obtained by using the methods described in Figures 3A to 3G. As shown, for example, the aligned sequences of strand fragments 307a and 307b show methylation at positions 3, 8, and 10 of strand fragment 307a because these cytosine residues present in the strand corresponding to the parental template strand fragment 307a are necessarily protected (methylated) in the original (parental) template strand 307a (and therefore not converted via bisulfite conversion). Furthermore, TC differences and / or TG mismatches (which identify the presence of substituted thymine residues (due to bisulfite conversion)) can be used to specify that the C residue at position 5 replaces the thymine residue in fragment 307a (see the asterisk at the cytosine residue at position 5 of fragment 307a). Similarly, via this same or similar principle, fragment 307b shows methylation at position 9 (read from left to right) with unprotected cytosine at position 6 (asterisk) (when reading the sequence of fragment 307b from left to right, i.e., 3'→5'). And it is noteworthy that this identified epigenetic methylation pattern corresponds to the example methylation pattern provided as an example in Figure 3C (see Figure 3C inset).
[0165] Therefore, by incorporating methylated cytosine nucleotides during polymerase extension of the circularized target DNA template, and then subjecting the double-length DNA template to bisulfite / PCR transformation, epigenetic information associated with the original target DNA template can be readily obtained from intra- and inter-strand parental / daughter sequences.
[0166] In light of the disclosure herein, epigenetic detection methods may be incorporated into the methods of this invention. For example, enzymatic conversion of a modified target base or any other biochemical or chemical reaction that specifically converts a modified nucleobase or target nucleobase relative to a native base (or alternatively, conversion of an unmodified target nucleobase, as discussed herein in conjunction with the bisulfite conversion of native cytosine to uracil). Certain example methods for enzymatic conversion of modified target bases are disclosed, for example, in the applicant’s co-pending U.S. Provisional Patent Applications Nos. 63 / 380439 and 63 / 147959, which are incorporated herein by reference in their entirety.
[0167] Double-length DNA templates for PCR multiplexing
[0168] In some exemplary embodiments, the end-adaptors described herein may be additionally or alternatively modified to include one or more sequence indices (SIDs). That is, end-adaptors (such as those in Figures 1A, 2A, and / or 3A) may be modified to include one or more specific nucleotide sequences that, when analyzing multiple target DNA templates / target sequences, identify, for example, the original source of the target DNA template (and thus the target sequence). Such SIDs are particularly useful in applications such as DNA multiplexing (i.e., processing multiple different samples simultaneously, such as via PCR). Therefore, SIDs are also referred to as sample identifiers.
[0169] In some exemplary embodiments, the same or different SIDs may be included in a position adjacent to the Y-branching sequence element described herein, such as in a sequence continuous at the 3' end of the Y-branching sequence element described herein. Alternatively or additionally, one or more SIDs may be included on the same strand as the Y-branching sequence element, wherein an inserted non-SID nucleotide or a sequence of nucleotides separates the SID from the Y-branching element. In any case, each SID may be unique for the target sequence, wherein the complementary sequence of the SID is present in the opposing (complementary) strand of the end adaptor. Subsequently, double-length DNA molecules with different SIDs can be processed in a single PCR reaction; for example, the SID allows for differentiation of different DNA samples after sequencing. Furthermore, because multiple copies of the SID will appear in a single repeating PCR product strand, the SID can be determined with high accuracy in bioinformatics. This, in turn, reduces or eliminates the need for additional error correction. In such exemplary embodiments, the SID can also be used as a marker in a given strand, thereby allowing for additional analysis. Furthermore, such embodiments including SID may also include UMI, such as those described in Figures 3A to 3E.
[0170] Referring to Figure 4A, an illustration is provided of a Y-branched adapter comprising two SID sequences and (optionally) a UMI, according to certain exemplary embodiments. As shown, YB-UMI / SID-EA 400 has a general polynucleotide double-stranded YBEA structure as shown in Figure 3A, including a UMI 416 having UMI strands 416a and 416b. This includes, for example, Y-branching elements 413a and 413b attached to the 5' ends of a first nick site 401a and a second nick site 401b of YB-UMI / SID-EA. For example, Y-branching elements 413a and 413b may comprise predetermined oligonucleotide sequences whose design can be tailored to achieve specific purposes, such as PCR amplification of double-length DNA templates, as described above for... Figure 2A As described in Figures 2E and 3A to 3E.
[0171] As shown in Figure 4A, each Y-branch element 413a and 413b may include a predetermined oligonucleotide sequence that provides a complementary primer binding site for PCR amplification. That is, each Y-branch element 413a and 413b may, for example, include 10 to 30 nucleotides, such as 15 to 25 nucleotides or 18 to 22 nucleotides, whose complementary sequence includes the primer binding site sequence. In some exemplary embodiments, Y-branch elements 413a and 413b include the same sequence, as shown in Figure 4A (at 413a and 413b). Alternatively, Y-branch elements 413a and 413b include different sequences. In some exemplary embodiments, Y-branch elements 413a and 413b have the same length, while in other exemplary embodiments, Y-branch elements 413a and 413b may have different lengths.
[0172] YB-UMI / SID-EA 400 also includes ends 403 and 404 located flanking each nick site, each end 403 and 404 being compatible with the end of the target DNA template for efficient ligation. Preferably, nick sites 401a and 401b are separated by a double-stranded spacer region 402, allowing EA to accommodate the attachment of two polymerases for bidirectional extension, as described herein. That is, nick sites 401a and 401b are spaced sufficiently far apart that the binding of one polymerase does not spatially impede and / or displace the binding of the second polymerase. This is illustrated, for example, in Figure 4A, where the spacer region 402 linearly offsets the first nick site 401a from the second nick site 401b.
[0173] As also shown in the figure, for example, the UMI sequence 416 is located within the spacer region 402 of YB-UMI / SID-EA 400. Typically, UMIs are 5 to 20 nucleotides in length, such as 8 to 16 nucleotides. Of course, this length can vary depending on the application. For example, UMIs can have a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. More conventionally, UMIs comprise sequences of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides. As shown in Figure 4A, because YB-UMI / SID-EA 400 is double-stranded, the UMI sequence 416 has complementary UMI strand sequences 416a and 416b, where strand 416a exhibits a 5'→3' polarity, while strand 416b has a complementary 3'→5' polarity.
[0174] In addition to UMI 416, which is shown in YB-UMI / SID-EA 400 but may optionally be included, YB-UMI / SID-EA 400 also includes diagonally positioned SID 417a (gray circle with black cross lines) and SID 418a (direction, shaded box), which are respectively shown as being continuously connected to Y branching elements 413a and 416b (solid black circles). That is, in the example shown in Figure 4A, the sequences of each SID 417a and 418a appear in tandem with the 5'→3' polynucleotide sequences of Y branching elements 413a and 413b, respectively. Also as shown, each of SID 417a and 418a has a corresponding complementary strand, namely, SID complementary strands 417b (box with diagonal lines) and 418b (circle with gray fill).
[0175] Typically, a SID comprises a short, random, and / or predetermined nucleotide sequence that can be incorporated into a polynucleotide sequence. Typically, a SID is 5 to 20 nucleotides in length, such as 8 to 16 nucleotides. Of course, this length can vary depending on the application. For example, a SID can have a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. More conventionally, SIDs consist of sequences of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.
[0176] Although YB-UMI / SID-EA 400 shows SIDs 417a and 418a that are adjacent to and continuously connected to Y branch elements 413a and 413b, respectively, it should be understood that one or more SIDs may be located in YB-UMI / SID-EA 400 at any location that facilitates sample differentiation. For example, one or more SIDs may be located adjacent to UMI chain 416a or 416b, such as on the 5' side of UMI chain 416a or the 5' side of UMI chain 416b. In other exemplary embodiments, SIDs may be included within and / or as part of UMI 416. Additionally or alternatively, one or more SIDs may be located at either end of YB-UMI / SID-EA 400. For example, SID 417a may be located on the 3' end portion of end 404, while SID 418a may be located on the 3' end portion of end 403. Therefore, the SIDs described herein may be located at or within multiple different locations of the YB-UMI / SID-EA 400, as long as the SIDs allow for sample differentiation as described herein.
[0177] Referring to Figure 4B, a double-length DNA template generated using the YB-UMI / SID-EA400 of Figure 4A is shown according to certain exemplary embodiments. That is, the YB-UMI / SID-EA 400 is attached to both ends of the target DNA template to form a circularized end-adaptor / target DNA template, as described in Figures 1B, 2B, and 3B. In other words, Figure 1B , 2BThe same or similar steps described in 3B can be used to form a circular construct including YB-UMI / SID-EA 400, which forms a bridge covalently linking the two ends of the target DNA template. Thereafter, using the same or similar steps as described for Figures 1C-1D, 2C-2D, and 3D-3E, a first polymerase and a second polymerase can be used to bind and extend the 3' ends of the cleavage sites 401a and 401b of YB-UMI / SID-EA 400, such as in opposite directions, to form daughter strand copies as described herein. Once the polymerases have completed their extension, as in, for example, step 1d (Figure 1D), step 2d (Figure 3D), etc. Figure 2D ) and step 3d ( Figure 3E As described in section ), the polymerase dissociates from the circular construct, producing a double-length DNA template as shown in Figure 4B.
[0178] As shown in Figure 4B, the double-length DNA template 411 comprises two copies of the original (parental) DNA template 407, namely, a first copy 411a and a second copy 411b, each located flanking the bridging region 408. The bridging region 408, derived from YB-UMI / SID-EA 400, comprises a daughter strand SID sequence 417a' and a complementary sequence 417b, wherein the daughter strand (gray) is formed via polymerase-mediated extension at the 3' end of the nick site 401a. The bridging region 408 also includes a daughter strand SID sequence 418a' and a complementary sequence 418b, wherein the daughter strand (gray) is formed via polymerase-mediated extension at the 3' end of the nick site 401b. UMI 416 comprises complementary UMI strand sequences 416a and 416b.
[0179] As shown in the figure, each DNA template copy 411a and 411b includes a parental polynucleotide strand (black) and a newly synthesized daughter polynucleotide strand (gray). For example, template copy 411a includes the original (parental) template strand 407b and the newly synthesized daughter strand 407a'. On the other side of the bridging region 408, template copy 411b includes the original (parental) template strand 407a (dashed line and black) and the newly synthesized daughter strand 307b' (gray). Furthermore, the double-length DNA template 411 includes a first end 412a and a second end 412b. For example, the first end 412a includes the portion of YB-UMI / SID-EA 400 strand 400b associated with the 5' end of the nick site 401b YB-UMI / SID-EA 400 (a hollow black rectangle at end 412a) and its copy (a hollow gray circle at end 412a). Similarly, the second end 412b of the double-length DNA template 411 includes the portion of YB-UMI / SID-EA 400 chain 400a associated with the 5' end of the nick site 401a YB-UMI / SID-EA400 (the hollow black circle at 412b) and its copy (the hollow gray rectangle at end 412b).
[0180] As also shown in the figure, template copy 411a also includes, at end 412a, a Y-branching element 413b and its complementary sequence in the Y-branching element progeny strand 413b', as well as SID 418a and its complementary progeny SID copy 418b. Similarly, at the other end of the double-length DNA template (i.e., end 412b), Y-branching element 413a and its complementary sequence in the Y-branching element progeny strand 413a' are shown, as well as SID 417a and its complementary progeny SID copy 417b.
[0181] In this way, the combination of YB-UMI / SID-EA 400 with the target DNA template strand produces a double-length DNA template 411, which includes predetermined oligonucleotide sequences at each end (i.e., Y branching elements 413a and 413b and their corresponding complementary copies 413a' and 413b'), SIDs and their complementary copies at each end (i.e., SIDs 418a and 417a and their respective complementary copies 418b' and 417b'), and UMI 416 (and its strand sequences 416a and 416b). And although Figure 4B shows the positions of the SID sequences 418a, 418b', 417a, and 417b' adjacent to the Y branching elements 413b, 413b', 413a, and 413a', it should be understood that the SIDs can be located at different positions within the double-length DNA template. For example, if there is no Y-branching element at the end adapter that includes SID, the resulting double-length DNA template may include SID at the first end 412a and the second end 412b without having a Y-branching element.
[0182] By encoding this strand-specific information in each strand of the double-length DNA template 411, the individual strands of the double-length DNA template 411 can be easily identified and distinguished in multiplex PCR reactions. In fact, because multiple copies of the SID will appear in a single replicated PCR product strand, the SID of YB-UMI / SID-EA 400 and the resulting double-length DNA template 411 can be determined with high accuracy in bioinformatics, thereby reducing or eliminating the need for additional error correction. The SID can also be used as a marker within a given strand, allowing for additional analysis.
[0183] Asymmetric DNA template extension and formation
[0184] In some exemplary embodiments, methods for copying and preparing asymmetric DNA templates are provided. That is, the methods and compositions provided herein can be used to generate asymmetric DNA templates in which only one strand of the target DNA template is copied. Therefore, an asymmetric DNA template is an asymmetric DNA template because only one strand of the parental template is copied. Such asymmetric DNA templates can be used in sequence preparation workflows that require a single-stranded DNA molecule as a target template, such as the “amplification sequencing” method developed by the inventors (see, for example, U.S. Patent Application Publication No. 20220042075, which is incorporated herein by reference in its entirety).
[0185] Referring to Figure 5A, an illustration is provided of a modified Y-branched adaptor according to Figure 2A according to certain exemplary embodiments, but which has been modified to accommodate only a single polymerase attachment and unidirectional extension. In other words, the modified Y-branched adaptor (or "modified YBEA")—when ligated to a target DNA template—accompanies only unidirectional extension of the target DNA template. For example, modified YBEA 500 includes hybrid strands 500a (circle) and 500b (rectangle) to form a polynucleotide duplex. Modified YBEA 500 also has the general EA structure shown in Figure 2A, wherein a first Y-branching element 513a and a second Y-branching element 513b are added to the 5' ends of a first nick site 501a and a second nick site 501b, respectively. Furthermore, the modified YBEA 500 includes ends 503 and 504 located flanking each nick site 501a and 501b, each end 503 and 504 being compatible with the ends of the target DNA template as described herein. In other words, the ends are ligable to the target DNA template.
[0186] However, unlike YBEA 200 in Figure 2A, the 3' end of one of the nick sites—which provides a polymerase extension site as described herein—is blocked or otherwise modified to prevent polymerase binding and / or extension in the modified YBEA. It is noteworthy that any method known in the art for modifying the 3' end to prevent polymerase binding and / or extension can be used. For example, the 3' end may be phosphorylated. As shown in the example modified YBEA 500 in Figure 5A, the 3' end associated with nick site 501a includes a phosphorylated 3' end, thereby preventing polymerase extension of the 3' nick site 501a, as further described herein. And although not shown for simplicity, in some exemplary embodiments, the modified YBEA 500 may include a UMI and / or one or more SIDs as described herein.
[0187] In some exemplary embodiments, the modified YBEA—having a single extendable nick site—can be combined with a target DNA template to form a circular construct. That is, the modified YBEA can be ligated to both ends of the target DNA template, as described in Figures 2B and 3B. After ligation, a bridge is formed between the two ends of the target DNA template, wherein the modified YBEA 500 serves as the bridge therebetween. Once the modified YBEA forms a bridge connecting the ends of the target DNA template, a circular construct is formed, as described in Figures 2B (steps 2a and 2b) and 3B (steps 3a and 3b).
[0188] Referring to Figure 5B, a schematic diagram depicting the initiation of polymerase attachment and unidirectional extension of a circular construct according to certain exemplary embodiments is provided. As shown, the circular construct 509 includes parental strands (black) 507a (dashed line) and 507b (solid line) of the target DNA template, and Y-branching elements 513a and 513b. Furthermore, the nick site 501a includes a phosphorylated 3' end, thereby preventing polymerase extension at the 3' end of the nick site 501a. However, when polymerase 510b is combined with the DNA circular construct 509, it binds to the nick site 501b—in this example where the 3' modification is absent—to proceed in the opposite direction to the nick site 501a (as indicated by the arrow). Polymerase 510b also substitutes the 5' end of the parental template strand 507b, including the substitution of its associated Y-branching element 513b. However, in the absence of polymerase binding / extension at the nick site 501a, the parental strand 507a remains bound to its complementary strand 507b at the nick site; that is, the 5' end of strand 507a does not undergo the substitution as described in the bidirectional target DNA template extension.
[0189] At step 5a in Figure 5B, polymerase 510b continues to unidirectionally extend the circular construct 509 (see arrow). That is, polymerase 510b uses the parental strand 507a as a template to extend the 3' end of the nick site 501b, while also replacing the 5' end of the nick site 501b (and its associated parental template strand 507b) (Figure 5B, bottom). Thus, when polymerase 510b proceeds in step 5a (bottom), it uses the parental strand 507a as a template to extend the 3' end of the nick site 501b to synthesize a new daughter strand 507b' (grey) that is sequence-complementary to (and therefore shares the same sequence as) the parental strand 507b. As shown, the Y-branch element 513b also remains attached at the 5' end to the replaced parental strand 507b. The new daughter strand 507b' also includes a replicated ssDNA daughter strand bridge portion 508b as part of bridge 508. However, the nick site 501a is retained and does not extend due to its phosphorylated (and therefore blocked) 3' end.
[0190] Continuing with the exemplary embodiments described above, Figure 5C is a schematic diagram depicting the continuous polymerase extension and asymmetric template formation of a circular construct using modified YBEA 500 according to certain exemplary embodiments. As shown, polymerase 510b continues along the parental template strand 507a to the 5' end of the parental template strand 507a, completing the synthesis of a new daughter strand 507b'. Notably, the new daughter strand 507b' includes a Y-branching element daughter strand 513a', which is complementary to the Y-branching element 513a attached to the parental strand 507a. As shown in Figure 5C, in this exemplary embodiment, the parental template strand 507b is completely replaced, but the strand is not replicated due to the blocking 3' end associated with 501a (see Figure 5B).
[0191] At step 5b in Figure 5C, once polymerase 510b completes the synthesis of daughter strand 507b', it dissociates the DNA complex to form an asymmetric DNA template 511 (see figure below). As shown in Figure 5D (see figure below), the asymmetric DNA template 511 includes parental template strands 507a and 507b, each located on the flank of bridge 508 (the bridge is derived from modified YBEA 500 and includes the bridge daughter strand portion 508b). The symmetric DNA template 511 also includes a first end 512a and a second end 512b. However, for unidirectional extension via polymerase 510b, only a single daughter strand, strand 507b', exists in the asymmetric DNA template 511. As shown in the figure, the asymmetric portion of the asymmetric DNA template 511 includes strand 507b', wherein the Y-branching element 513b is located at the 5' end of the 507b' strand (at the end 512a). Furthermore, a bridge portion 508 with a newly synthesized daughter bridge portion 508b continuously and covalently links the parental template strand 507b to the newly synthesized daughter strand 507b'.
[0192] As also shown in the figure, template copy 511b includes daughter strand 507b', which is a complement to parental template strand 507a. Parental strand 507a also includes a Y-branching element 513a at its 5' end, while the new daughter strand 507b includes a daughter Y-branching element 513a' at its 3' end. In this way, the modified YBEA 500 is combined with the parental target DNA template to form a circular construct, followed by polymerase-mediated extension as described in steps 5a and 5b of Figures 5B and 5C to form an asymmetric DNA template.
[0193] Multiple length template extension
[0194] In some exemplary embodiments, the methods and compositions described herein can be repeated any number of times—starting from a first double-length DNA template—to form multiple-length DNA templates. For example, after forming a double-length DNA template according to the methods and compositions described herein, the two ends of the double-length DNA template can be connected to second end adapters (EAs)—second EAs, for example, having the features of the EA of Figure 1A. This forms a circular construct comprising the double-length DNA template and the connected EAs. Thereafter, as described herein, the circular construct can be bidirectionally replicated to form a quadruple-length DNA template, or a “double-double” template, i.e., a DNA molecule comprising a replicated copy of the original double-length DNA template. For example, a quadruple-length DNA template comprises two parental target DNA template strands derived from the original double-length DNA template, and their complementary daughter strands, as described herein. However, the quadruple-length DNA template also comprises repetitions of these strands and thus comprises four copies of the target DNA template. The formation of such a quadruple-length DNA template is illustrated, for example, in Figure 6.
[0195] As shown in Figure 6, the target DNA template in this example is a double-length DNA template, comprising two copies of the original target DNA template as described herein, and is therefore referred to in this example as the target double-length DNA template 607. For example, the two copies of the original DNA template comprise hybridized polynucleotide strands 607a and 607b' (first copy) and hybridized polynucleotide strands 607b and 607a' (second copy). Also, as described herein, both copies comprise parental DNA (strands 607a and 607b, shown in black) and its complementary copy strands (strands 607b' and 607a', shown in gray) from the original target DNA template. The two copies are separated by a first double-stranded bridging region 608a (i.e., the original bridge), which originates from the initial (or first) end adapter for forming the target double-length DNA template 607, as described herein. For example, the first bridge 608a comprises strands 620 and 630. The target double-length DNA template 607 also includes a first template end 605 and a second template end 606, both of which can be linked to the second EA 600.
[0196] A second end adaptor (EA) 600 is also shown, having, for example, the structure of EA 100 as shown in FIG. 1A. For example, the second EA 600 includes a first nick site 601a and a second nick site 601b, both of which are adaptable to polymerase binding and extension (e.g., bidirectional extension as described herein). EA 600 also includes a first EA end 605 and a second EA end 606, respectively. As described herein, for example, both EA ends 605 and 606 can be ligated to a target double-length DNA template 607.
[0197] At step 6a, for example, the second EA 600 is attached to either end of the target double-length DNA template 607. For example, end 603 of the second EA 600 is attached to end 606 of the template (Figure 6, step 6a). Alternatively, at step 6a, and although not shown for simplicity, the other end (i.e., end 604) of the second EA 600 is attached to end 605 of the target double-length DNA template 607. In either case, one end of the second EA 600 is attached to the end of the target double-length DNA template.
[0198] At step 6b in Figure 6, the remaining free ends of the second EA 600 are connected to the remaining free ends of the target double-length DNA template 607 to form a circular construct 609. That is, in step 6b of Figure 6B, the two ends 603 and 604 of the second EA 600 are connected to each end 605 and 606 of the target double-length DNA template 607, thereby forming a second DNA bridge 608b between the ends of the target double-length template 607. In other words, the entire second EA 600 forms a second DNA bridge 608b between the two ends 605 and 606 of the target double-length DNA template 607. This forms a circular construct 609 comprising the first EA (as bridge 608a) and the second bridge 608b (formed by the second EA 600). In this way, the second EA 600 operates as a bridging precursor for the second bridging region 608b of the circular construct 609. The corresponding first nick site 601a and second nick site 601b are retained in the loop construct 609 (as part of the second bridging region 608b), and thus, in some exemplary embodiments, two corresponding 3' ends are provided for polymerase attachment and bidirectional extension, as described herein.
[0199] In steps 6c to 6d, the circular construct 609 is replicated, as described with respect to steps 1c to 1d of Figures 1C and 1D (these steps are combined in Figure 6 for simplicity). As shown, completion of steps 6c to 6d in Figure 6 produces a DNA template 611 of four times the length. For example, in step 6c, the circular construct 609 is brought into contact with polymerases (e.g., a first polymerase and a second polymerase (not shown)) bound to nick sites 601a and 601b of the circular construct 609. Thereafter, the polymerases extend the circular construct 609 bidirectionally. For example, one of the polymerases extends the 3' end of nick site 601a while also displacing the 5' end of nick site 601a. Similarly, the other polymerase extends the 3' end of nick site 601b while also displacing the 5' end of nick site 601b. At step 6d, once the polymerases have completed the extension reaction of their circular construct 609, they dissociate from the circular construct 609 to form a DNA template 611 that is four times longer.
[0200] As shown in the figure, the quadruple-length DNA template 611 comprises four copies of the target sequence. For example, copy 1 comprises the original parental target template DNA strand 607a—as carried from the original parental target DNA template to the double-length DNA template—and its complementary, newly synthesized non-parental strand 607c. For example, copy 2 comprises the non-parental strand 607a' from the double-length DNA template, and the newly synthesized non-parental strand 607d. As shown in the figure, copies 1 and 2 are separated by the strand segment 620 of the first bridge 608a (black and gray circles) and its newly synthesized complementary portion (gray rectangle).
[0201] Similarly, copy 3 includes a non-parental strand 607b' from the double-length DNA template, and a newly synthesized non-parental strand 607e. As shown, copies 2 and 3 are separated by a second bridging region 608b, which includes the portion forming EA 600 (black) and its newly synthesized portion (gray). Furthermore, copy 4 includes the original parental target template DNA strand 607b—as carried from the original parental target DNA template to the double-length DNA template—and its complementary, newly synthesized non-parental strand 607f. As shown, copies 3 and 4 are separated by the strand segment 630 of the first bridge 608a (hollow and gray box) and its newly synthesized complementary portion (gray shaded circle).
[0202] It is noteworthy that, in the example shown in Figure 6, in the 5'→3' direction, the parental polynucleotide strand 607a is continuously linked to the non-parental strand copies 607a', 607e, and 607f via a bridging region sequence. Parental strand 607a also shares the same 5'→3' polynucleotide sequence as the non-parental strand copies 607a', 607e, and 607f. Similarly, in the 5'→3' direction, the parental polynucleotide strand 607b—via a bridging region sequence—is continuously linked to the non-parental strand copies 607b', 607d, and 607c. Parental strand 607b also shares the same 5'→3' polynucleotide sequence as the non-parental strand copies 607b', 607d, and 607c.
[0203] Although Figure 6 illustrates the formation of a quadruple-length DNA template using, for example, EA as shown in Figure 1A, it should be understood that the method of Figure 6 can be repeated iteratively, with the number of target DNA template copies doubling each time. For example, the initial double-length DNA template includes two copies of the target DNA template, as described herein, while additional repetitions—in the example method shown in Figure 6—generate four copies of the target DNA template, i.e., a quadruple-length DNA template 611. Subsequent additional iterations produce initial target DNA templates of 8, 16, 32, 64, etc.
[0204] Furthermore, it should be understood that any end-adaptor described herein and its associated methods of use can be used to form a DNA template of four times its length. Alternatively, any end-adaptor described herein and its associated methods can be used to form a DNA template of multiple lengths when repeating the replication described in Figure 6. For example, this includes using different EAs in different iterations when forming a DNA template of multiple lengths.
[0205] For example, the EA of Figure 1A can be used to form an initial double-length DNA template, and the EA of Figure 1A can also be used in a second iteration to form a quadruple-length DNA template 611. Subsequent iterations can use the EAs of Figure 2A (EA200), Figure 3A, and / or Figure 4A (EA 300) to form an 8-copy multi-length DNA template. Therefore, in some exemplary embodiments, the quadruple-length DNA template or multi-length DNA template may include a Y-branched EA to facilitate subsequent PCR amplification and / or UMI and SID to facilitate bioinformatics analysis (including genetic and epigenetic analyses as described herein). Indeed, such quadruple-length and multi-length DNA templates are particularly useful in the validation of sequencing reactions and their associated data. Additionally or alternatively, in some exemplary embodiments, one or more of the EAs may include protected nick sites as described herein (e.g., Figures 5A and 5B) to form asymmetric double-length or multi-length DNA templates.
Claims
1. A linear end adaptor for replicating a target DNA template, the linear end adaptor comprising: A first polynucleotide chain hybridizes with a second polynucleotide chain to form a polynucleotide duplex, the polynucleotide duplex comprising a first end and a second end; A first nick site and a second nick site, wherein the first nick site is located within the first polynucleotide chain of the polynucleotide duplex, and wherein the second nick site is located within the second polynucleotide chain of the polynucleotide duplex; and, A spacer subregion separates the first incision site and the second incision site from each other, thereby causing the first incision site and the second incision site to be linearly offset.
2. The linear end-connector of claim 1, wherein the first cleavage site of the first polynucleotide chain comprises a discontinuous break in the sequence of the first polynucleotide chain, and / or wherein the second cleavage site of the second polynucleotide chain comprises a discontinuous break in the sequence of the second polynucleotide chain.
3. The linear end adapter according to claim 1 or 2, wherein each end is configured to connect to both ends of the target DNA template.
4. The linear terminator of claim 3, wherein the first and / or second ends of the polynucleotide duplex comprise connectable blunt ends.
5. The linear terminator of claim 3, wherein the first and / or second ends of the polynucleotide duplex comprise connectable nucleic acid overhangs.
6. The linear end-connector according to any one of claims 1 to 5, wherein each nick site is configured for a polymerase-mediated extension reaction.
7. The linear end-connector according to any one of claims 1 to 6, wherein the linear offset between the first nick site and the second nick site corresponds to the distance for accommodating the polymerase binding to the first nick site and the second nick site.
8. The linear end connector according to any one of claims 1 to 7, wherein the first cut site and / or the second cut site comprises a 3' end and a 5' end.
9. The linear end connector of claim 8, wherein the linear end connector further comprises a first Y-branch element sequence attached to the 5' end located on the flank of the first incision site and / or a second Y-branch element sequence attached to the 5' end located on the flank of the second incision site.
10. The linear end-connector of claim 9, wherein the sequence of the first Y-branch element and / or the second Y-branch element encodes the primer binding sequence.
11. The linear end-connector according to claim 9 or 10, wherein the sequence of the first Y-branch element and / or the second Y-branch element is approximately 5 to 25 nucleotides in length.
12. The linear end-connector according to any one of claims 1 to 11, wherein the first polynucleotide chain and / or the second polynucleotide chain comprises a unique molecular identifier (UMI) sequence.
13. The linear connector of claim 12, wherein the UMI is located within the spacer region.
14. The linear end-adaptor according to any one of claims 1 to 13, wherein the first polynucleotide chain comprises a first sequence index (SID) and / or wherein the second polynucleotide chain comprises a second SID.
15. The linear end-adaptor according to any one of claims 9 to 13, wherein the first polynucleotide chain comprises a first sequence index (SID), and wherein the second polynucleotide chain comprises a second SID, wherein the sequence of the first SID is adjacent to the sequence of the first Y-branching element, and wherein the sequence of the second SID is adjacent to the sequence of the second Y-branching element.
16. The linear terminator according to any one of claims 1 to 15, wherein the linear terminator is about 50 to 100 nucleotides in length.
17. The linear end-connector according to any one of claims 1 to 16, wherein the spacer region is approximately 10 to 50 nucleotides in length.
18. The linear end-connector according to any one of claims 1 to 17, wherein the first nick site and / or the second nick site has a length corresponding to about 0 to 10 nucleotides.
19. The linear end-connector according to any one of claims 1 to 5, wherein only one of the nick sites is configured for a polymerase-mediated extension reaction.
20. The linear end-connector of claim 19, wherein the first nick site or the second nick site comprises a 3'-blocking group that prevents polymerase-mediated elongation reaction.
21. The linear end-connector according to claim 20, wherein the 3'-blocking group is a phosphate group.
22. The linear end connector according to any one of claims 19 to 21, wherein the first cut site and / or the second cut site comprises a 5' end, and wherein the 5' end of the first cut site comprises a first Y-branching element sequence and / or wherein the 5' end of the second cut site comprises a second Y-branching element sequence.
23. The linear end connector of claim 22, wherein the first Y-branch element and / or the second Y-branch element encodes a primer binding sequence.
24. The linear end connector according to any one of claims 19 to 23, wherein the spacer subregion comprises a UMI sequence.
25. A method for replicating a target DNA template, the method comprising: A ligation reaction is performed between the target DNA template and the linear end-adaptor according to any one of claims 1 to 24 to form a circular construct. The target DNA template includes a first target DNA template end and a second target DNA template end, and The ligation reaction (i) connects the first end of the end-joint to the end of the first target DNA template, and (ii) connects the second end of the end-joint to the end of the second target DNA template, thereby forming the circular construct; as well as, The DNA polymerase-mediated extension reaction of the circular construct is performed to replicate the target DNA template.
26. The method of claim 25, wherein performing the DNA polymerase-mediated extension reaction comprises contacting the circular construct with a plurality of strand displacement polymerases.
27. The method of claim 26, wherein the strand substitution polymerase is selected from the group consisting of: KAPAHiFi DNA polymerase, Q5® high-fidelity DNA polymerase, and Pfu DNA polymerase, such as Pfu-X.
28. The method of claim 26, wherein the chain displacement polymerase is phi 29 polymerase.
29. The method according to any one of claims 25 to 28, wherein the polymerase-mediated extension reaction comprises the extension of the 3' end of the first nick site or the 3' end of the second nick site of the end adapter.
30. The method of claim 29, wherein (i) polymerase-mediated extension of the 3' end of the first nick site of the end-adaptor or polymerase-mediated extension of the 3' end of the second nick site of the end-adaptor forms an asymmetric DNA template, or (ii) polymerase-mediated extension of both the 3' end of the first nick site and the 3' end of the second nick site of the end-adaptor forms a double-length DNA template.
31. The method of claim 30, wherein the strand of the asymmetric DNA template or the double-length DNA template contains a unique molecular identifier (UMI).
32. The method of claim 31, wherein the UMI is located within the strand of the bridging region of the asymmetric DNA template or the double-length DNA template.
33. The method according to any one of claims 30 to 32, wherein the strand of the asymmetric DNA template or the double-length DNA template comprises a sequence index (SID).
34. The method of claim 33, wherein the SID is located within the strand of the bridging region of the asymmetric DNA template or the double-length DNA template and / or at the end of the asymmetric DNA template or the double-length DNA template.
35. The method according to any one of claims 25 to 33, wherein the asymmetric DNA template or the double-length DNA template comprises a Y-branched adaptor at its end.
36. The method of claim 35, wherein the Y-branch end-connector encodes a primer binding site.
37. A method for preparing a double-length DNA template from a target DNA template, the method comprising: A ligation reaction is performed between the target DNA template and the end-adaptor according to any one of claims 1 to 18 to form a circular construct. The target DNA template includes a first target DNA template end and a second target DNA template end, and The ligation reaction (i) connects the first end of the end-adaptor to the end of the first target DNA template, and (ii) connects the second end of the end-adaptor to the end of the second target DNA template; and, The circular construct undergoes a DNA polymerase-mediated extension reaction to form a double-length DNA template containing a first and a second copy of the target DNA template.
38. The method of claim 37, wherein performing the DNA polymerase-mediated extension comprises contacting the circular construct with a plurality of strand displacement polymerases.
39. The method of claim 38, wherein the strand substitution polymerase is selected from the group consisting of: KAPAHiFi DNA polymerase, Q5® high-fidelity DNA polymerase, and Pfu DNA polymerase, such as Pfu-X.
40. The method of claim 38, wherein the chain displacement polymerase is phi 29 polymerase.
41. The method according to any one of claims 37 to 40, wherein the polymerase-mediated extension reaction comprises the extension of the 3' end of the first nick site and the 3' end of the second nick site of the end adapter.
42. The method according to any one of claims 37 to 41, wherein the polymerase-mediated extension is bidirectional.
43. The method according to any one of claims 37 to 41, wherein the first copy of the target DNA template and the second copy of the target DNA template are continuously connected to each other through a DNA bridging region.
44. The method of claim 43, wherein the bridging region originates from the end connector.
45. The method of claim 43 or 44, wherein each polynucleotide strand of the double-length DNA template comprises a 5' to 3' parental strand of the target DNA template and a 5' to 3' daughter strand copy of the parental strand of the target DNA template.
46. The method of claim 45, wherein copies of the parental strand and the daughter strand of the target DNA template are continuously linked to each other through the 5' to 3' strands of the DNA bridging region.
47. The method of claim 46, wherein the chain in the bridging region comprises a unique molecular identifier (UMI).
48. The method of claim 46 or 47, wherein the chain of the bridging region comprises a sequence index (SID).
49. The method of any one of claims 45 to 48, wherein the double-length DNA template comprises a first end and a second end, and wherein the first end and / or the second end comprises a SID.
50. The method according to any one of claims 37 to 49, wherein the linear end adapter comprises a first Y branching element sequence and a second Y branching element sequence, and wherein the DNA polymerase-mediated extension reaction is performed to position the first Y branching element sequence and the second Y branching sequence at the 5' end of each parent strand of the double-length DNA template.
51. The method of claim 50, wherein the polymerase-mediated elongation reaction of the DNA circular construct synthesizes a first progeny Y-branching element sequence and a second progeny Y-branching element sequence, wherein the first progeny Y-branching element sequence is complementary to the first Y-branching element sequence, and wherein the second progeny Y-branching element sequence is complementary to the Y-branching element sequence.
52. The method of claim 51, wherein the first daughter Y-branching element sequence and the second daughter Y-branching element sequence are located at the 3' end of each daughter strand copy of the double-length DNA template.
53. The method according to claim 50 or 52, wherein the Y-branch element encodes a primer binding site.
54. The method according to any one of claims 37 to 52, wherein the method is repeated continuously to form a DNA template of four times the length or multiple times the length.
55. A double-length DNA template formed by the method according to any one of claims 37 to 53.
56. A method for identifying epigenetic information associated with a target nucleic acid sequence, comprising: (a) Connecting a linear target DNA template to both ends of a linear end-connector according to any one of claims 1 to 18 to form a circular DNA construct; (b) A DNA polymerase-mediated bidirectional extension reaction of the circular DNA construct is performed in the presence of multiple protected cytosine nucleotides to form a double-length DNA template containing the protected cytosine nucleotides. (c) Denature the double-length DNA template; (d) subjecting a denatured double-length DNA template to a bisulfite conversion reaction to form a double-length DNA template strand that has undergone bisulfite conversion of the double-length DNA template; (e) Polymerase chain reaction (PCR) amplification of the double-length DNA template strand converted by the bisulfite; (f) Sequencing of the double-length DNA template strand converted from bisulfite by PCR amplification; (g) Based on the sequencing of the double-length DNA template strand amplified by PCR and converted by bisulfite, identify epigenetic information associated with the target nucleic acid.
57. The method of claim 56, wherein each polynucleotide strand of the double-length DNA template in step (b) comprises a parental template strand from the target DNA template and a daughter copy strand of the parental template strand.
58. The method of claim 57, wherein the parental template chain is continuously connected to the offspring copy chain of the parental template chain via a single-chain bridging region.
59. The method of claim 58, wherein the single-chain bridging region is derived from the end connector.
60. The method according to claim 57 or 58, wherein the protected cytosine nucleotide is incorporated into the daughter copy strand of the parent template strand during the DNA polymerase-mediated bidirectional elongation reaction in step (b).
61. The method according to any one of claims 57 to 60, wherein the sequencing of the PCR-amplified bisulfite-converted double-length DNA template strand in step (f) provides the polynucleotide sequences of the parental template strand and the daughter copy strand, and The identification of the epigenetic information associated with the target nucleic acid includes an intra-strand comparison of the polynucleotide sequence of the parental template strand with the polynucleotide sequence of the daughter copy strand.
62. The method of claim 61, wherein the sequence difference between the polynucleotide sequence of the parental template strand and the polynucleotide sequence of the daughter copy strand identifies the position of an unprotected cytosine residue in the parental template strand.
63. The method of claim 62, wherein the unprotected cytosine residue position in the parental template strand corresponds to the unprotected cytosine residue position in the target nucleic acid sequence.
64. The method according to any one of claims 61 to 63, wherein the position of a cytosine residue in the sequence of the parental template strand indicates the corresponding position of a protected cytosine in the target nucleic acid sequence.
65. The method of claim 56, wherein the double-length DNA template in step (b) comprises a first copy and a second copy of the target DNA template.
66. The method of claim 65, wherein the first copy and the second copy of the target DNA template are linked together by a double-stranded bridging region.
67. The method of claim 66, wherein the double-chain bridging region originates from the end connector.
68. The method according to any one of claims 65 to 67, wherein each copy of the target DNA template within the double-length DNA template comprises a parental template strand and a daughter strand that is complementary to and hybridizes with the parental template strand.
69. The method of claim 68, wherein the protected cytosine nucleotide is incorporated into the hybridized complementary daughter strand during the DNA polymerase-mediated bidirectional elongation reaction in step (b).
70. The method according to claim 68 or 69, wherein sequencing of the PCR-amplified, bisulfite-converted, double-length DNA template strand of step (f) provides a polynucleotide sequence of the parental template strand and its hybridized complementary daughter strand, and The identification of the epigenetic information associated with the target nucleic acid includes an inter-strand comparison of the polynucleotide sequence of the parent template strand with the polynucleotide sequence of the hybridized complementary offspring.
71. The method of claim 70, wherein the nucleotide mismatch position between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the hybridized complementary offspring identifies the position of an unprotected cytosine residue in the parent template strand.
72. The method of claim 71, wherein the unprotected cytosine residue position in the parental template strand corresponds to the unprotected cytosine residue position in the target nucleic acid sequence.
73. The method according to any one of claims 56 to 72, wherein the protected cytosine nucleotide comprises methylated cytosine residues.
74. The method according to any one of claims 56 to 72, wherein the unprotected cytosine nucleotide is an unmethylated cytosine residue.
75. The method according to any one of claims 56 to 74, wherein the double-length DNA template in step (b) contains a unique molecular identifier (UMI).
76. The method of claim 75, wherein the UMI is located in the single-chain bridging region of claim 59 or the double-chain bridging region of claim 67.
77. The method according to any one of claims 56 to 76, wherein the double-length DNA template in step (b) comprises a sequencing index (SID).
78. A double-length DNA template comprising a first copy and a second copy of a target DNA template, wherein the first copy and the second copy of the target DNA template are continuously connected to each other by a double-stranded bridging region.
79. The double-length DNA template of claim 78, wherein each polynucleotide strand of the double-length DNA template comprises a parental template strand from the target DNA template and a daughter copy strand of the parental template strand.
80. The double-length DNA template according to claim 78 or 79, wherein the parental template strand is continuously connected to the daughter copy strand of the parental template strand through the bridging region.
81. The double-length DNA template of claim 78, wherein each copy of the target DNA template within the double-length DNA template comprises a parental template strand and a daughter strand that is complementary to and hybridizes with the parental template strand.
82. The double-length DNA template according to any one of claims 78 to 81, wherein the double-length DNA template comprises a first end and a second end, wherein either end comprises a sequence encoding a primer binding site.
83. The double-length DNA template according to any one of claims 78 to 82, wherein the bridging region or the strand thereof includes a unique molecular identifier (UMI).
84. The double-length DNA template according to any one of claims 78 to 83, wherein the bridging region or the strand thereof comprises a sequencing index (SID).
Citation Information
Patent Citations
Methods, compositions, and devices for solid-state syntehsis of expandable polymers fo ruse in single molecule sequencings
US20220042075A1
Process for amplifying, detecting, and / or-cloning nucleic acid sequences
US4683195A
Process for amplifying nucleic acid sequences
US4683202A