Methods and compositions for the preparation and analysis of DNA libraries
The use of a linear end adapter with nick sites and unique molecular identifiers enhances DNA replication accuracy and sequencing efficiency by forming double-length templates, addressing the limitations of conventional PCR and isothermal methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2026-03-19
AI Technical Summary
Conventional PCR techniques for nucleic acid sequencing are limited by heating and cooling cycles, which can degrade the target sample and are not compatible with studying living systems, and existing isothermal methods like rolling circle amplification lack accuracy in single-read DNA sequencing.
A method using a linear end adapter with nick sites and spacer regions to replicate a target DNA template, forming a double-length DNA template with parent and daughter strands, incorporating unique molecular identifiers and sequence indices, enabling accurate and controlled replication and analysis of both DNA strands.
The method provides reliable and efficient replication of DNA templates, maintaining library length uniformity and improving sequencing accuracy by allowing simultaneous analysis of genetic and epigenetic information in a single read.
Smart Images

Figure 2026509614000001_ABST
Abstract
Description
Technical Field
[0001] [Description of the Sequence Listing] The sequence listing related to this application is provided in xml format instead of a paper copy and is incorporated herein by reference. The name of the text file containing the sequence listing is P37440-WO_Sequence_Listing.xml. The xml file is 5,000 bytes and was created on March 7, 2024.
[0002] [Field of the Invention] The present invention generally relates to methods and compositions for preparing DNA libraries, and more particularly to methods and compositions for replicating a target DNA template and analyzing the replicated target DNA template for genetic and / or epigenetic information.
Background Art
[0003] Nucleic acid sequencing is an important technique for biology and medicine. Although conventional polymerase chain reaction (PCR) techniques are very useful and effective, the required heating and cooling cycles of PCR limit their usefulness. For example, melting the hybridized DNA during the heating cycle can cause degradation of the target sample. Also, such heating / cooling cycles are not well compatible with the study of living systems. Due to the limitations of conventional PCT techniques, many studies have focused on finding an isothermal approach to nucleic acid amplification.
[0004] One conventional isothermal method for generating multiple copies of a target nucleic acid involves rolling circle amplification (RCA), in which a small circular oligonucleotide provides a template for polymerase binding and unidirectional replication. RCA generates a long single ss-DNA product consisting of many tandemly linked copies of the complementary strand of the target DNA molecule. This method involves circularizing the target DNA and initiating polymerase elongation using a primer. After replicating around the circularized DNA, the primer is replaced, and the polymerase proceeds through multiple rounds of the target DNA, generating multiple copies until a termination event occurs. This results in a long single-stranded DNA strand containing multiple copies of the target. The single-stranded DNA strand can then be read and analyzed.
[0005] However, the accuracy of a single read in single-DNA molecular sequencing has often been limited. Several techniques to improve accuracy include i) rereading the molecule, ii) reading its complementary strand, or iii) reading multiple copies of the DNA molecule (e.g., as used in RCA). For example, a molecule can be read multiple times by circularizing the target DNA (including both complementary strands) and examining it multiple times as it loops around the sensing site. Other systems "exfoliate" the complementary strand when reading one strand and then briefly capture that complementary strand for immediate reading. DNA-based universal molecular identifiers (UMIs) and sample identifiers (SIDs) are then spliced into the individual molecules before PCR amplification, and as a result, the examined subset of the resulting amplicon copy family can be attributed to a single parent molecule derived from a specific sample. Reading multiple copies within a family improves the accuracy of determining the sequence of that molecule.
[0006] While the above methods are useful, what is needed is a method and composition for consistently, reliably, and easily replicating a target sequence for library preparation and other purposes, while controlling the size of the replicas. For example, what is needed is a method and composition for replicating a target sequence, thereby enabling beneficial control over the size of a template library. There is also a need for a method for replicating both strands of a target sequence while preserving the parent strand in the replicas, thereby facilitating bioinformatics analysis of the target sequence. [Overview of the project]
[0007] In certain exemplary embodiments, a linear end adapter (EA) is provided for replicating a linear target DNA template. The EA includes, for example, a first polynucleotide chain that hybridizes to a second polynucleotide chain, thereby forming a polynucleotide duplex. The polynucleotide duplex includes, for example, a first terminal end and a second terminal end. The EA also includes a first nick site and a second nick site, the first nick site located within the first polynucleotide chain of the polynucleotide duplex, and the second nick site located within the second polynucleotide chain of the polynucleotide duplex. A spacer region separates the first and second nick sites from each other, thereby linearly offsetting the first nick site from the second nick site, i.e., there is a linear offset between the first and second nick sites. Furthermore, each terminal end of the EA can be configured to ligate both ends of the target DNA template. One or both of the nick sites facilitate, for example, polymerase binding and elongation.
[0008] In certain exemplary embodiments, the linear end adapter includes a first Y-branched element sequence bound to the 5' end adjacent to a first nick site and / or a second Y-branched element sequence bound to the 5' end adjacent to a second nick site. The Y-branched elements may encode, for example, a primer-binding sequence or other beneficial sequences.
[0009] In certain exemplary embodiments, the first polynucleotide chain and / or the second polynucleotide chain of the EA include a unique molecular identifier (UMI) sequence. For example, the UMI may be located within a spacer region. In certain exemplary embodiments, the first polynucleotide chain of the EA includes a first sequence index (SID), and / or the second polynucleotide chain of the EA includes a second SID.
[0010] In certain exemplary embodiments, a method is provided for preparing a double-length DNA template a from a target DNA template. This method includes, for example, performing a ligation reaction between the target DNA template and the terminal adapter described herein to form a circular construct. For example, the target DNA template includes a first target DNA template terminal end and a second target DNA template terminal end. Thus, the ligation reaction (i) ligates the first terminal end of the terminal adapter to the first target DNA template terminal end, and (ii) ligates the second terminal end of the terminal adapter to the second target DNA template terminal end. This forms a circular construct. Subsequently, a DNA polymerase-mediated extension reaction is performed on the circular construct. For example, the circular construct is brought into contact with a plurality of strand-displacing polymerases to initiate the extension reaction. The extension reaction forms a double-length DNA template including, for example, a first copy and a second copy of the target DNA template.
[0011] In certain exemplary embodiments, a first copy of a target DNA template (of a double-length DNA template) and a second copy of the target DNA template are joined adjacent to each other at a DNA crosslinking region. The crosslinking region originates, for example, from a terminal adapter. The crosslinking region is, for example, double-stranded.
[0012] In certain exemplary embodiments, each polynucleotide strand of a double-length DNA template includes a 5' to 3' parent strand of the target DNA template and a 5' to 3' daughter strand copy of the parent strand of the target DNA template. In certain exemplary embodiments, the parent strand and the daughter strand copy of the target DNA template may be adjacent to each other via a 5' to 3' strand of DNA crosslinking region.
[0013] In certain exemplary embodiments, the crosslinked region chains include a unique molecular identifier (UMI) or a sequence index (SID). For example, a double-length DNA template includes a first and a second terminal end, where the first and / or second terminal ends include an SID.
[0014] In certain exemplary embodiments, for example, if the linear end adapter includes a first Y-branching element sequence and a second Y-branching element sequence, the DNA polymerase-mediated extension reaction places the first Y-branching element sequence and the second Y-branching element sequence at the 5' ends of each parent strand of the double-length DNA template. Furthermore, the polymerase-mediated extension reaction of the DNA circular construct synthesizes a first daughter Y-branching element sequence and a second daughter Y-branching element sequence, where the first daughter Y-branching element sequence is complementary to the first Y-branching element sequence and the second daughter Y-branching element sequence is complementary to the Y-branching element sequence. The Y-branching elements can, for example, encode primer binding sites for subsequent PCR reactions.
[0015] In certain exemplary embodiments, this method can be repeated sequentially. For example, by repeating this method sequentially, a tetraplic or multiplicative DNA template can be produced. In such exemplary embodiments, the multiplicative DNA template includes multiple copies of the target DNA template.
[0016] In certain exemplary embodiments, a method is provided for identifying epigenetic information associated with a target nucleic acid sequence. This method includes, for example, ligating linear target DNA templates to both ends of a linear end adapter described herein to form a circular DNA construct. A DNA polymerase-mediated bidirectional elongation reaction is then performed on the circular DNA construct in the presence of multiple protected cytosine nucleotides. A diploid DNA template containing protected cytosine nucleotides in a newly synthesized strand is then formed. The diploid DNA template is then denatured and subjected to a bisulfite conversion reaction to form a bisulfite-converted diploid DNA template strand of the diploid DNA template. A polymerase chain reaction (PCR) amplification reaction is then performed using the bisulfite-converted diploid DNA template strand, followed by sequencing of the PCR-amplified / bisulfite-converted diploid DNA template strand. Based on the sequencing of the PCR-amplified / bisulfite-converted diploid DNA template strand, epigenetic information associated with the target nucleic acid is identified. In other words, epigenetic information can be identified using bioinformatics.
[0017] In certain exemplary embodiments, each polynucleotide strand of a double-length DNA template for a method of identifying epigenetic information includes a parent template strand derived from a target DNA template and a daughter copy strand of the parent template strand. The parent template strand is adjoiningly linked to the daughter copy strand of the parent template strand, for example, via a single-stranded crosslinking region (which originates from a terminal adapter). Furthermore, during a DNA polymerase-mediated bidirectional elongation reaction, protected cytosine nucleotides are incorporated into the daughter copy strand of the parent template strand.
[0018] In certain exemplary embodiments, sequencing of a PCR-amplified bisulfite-converted double-length DNA template strand provides the polynucleotide sequence of the parent template strand and the sequence of the daughter copy strand. The subsequent step of identifying epigenetic information related to the target nucleic acid includes an intra-stratum comparison between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the daughter copy strand. For example, the location of the sequence mismatch between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the daughter copy strand identifies the location of an unprotected cytosine residue in the parent template strand. For example, the location of an unprotected cytosine residue in the parent template strand corresponds to the location of an unprotected cytosine residue in the target nucleic acid sequence.
[0019] In certain exemplary embodiments, a doubling DNA template for a method of identifying epigenetic information comprises a first copy and a second copy of a target DNA template. The first and second copies of the target DNA template can be linked together, for example, via a double-stranded crosslinking region, which originates from a terminal adapter. Furthermore, each copy of the target DNA template within the doubling DNA template comprises a parent template strand and a daughter strand that is complementary to and hybridizes with the parent template strand. During a DNA polymerase-mediated bidirectional elongation reaction, for example, protected cytosine nucleotides are incorporated into the hybridized complementary daughter strand.
[0020] In such exemplary embodiments, when a PCR-amplified bisulfite-converted double-length DNA template is sequenced, inter-strand comparisons between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the hybridized complementary daughter strand can be used to identify epigenetic information related to the target nucleic acid. For example, nucleotide mismatch sites between the polynucleotide sequence of the parent template strand and the hybridized complementary daughter strand identify unprotected cytosine residue locations in the parent template strand, and these unprotected cytosine residue locations in the parent template strand correspond to unprotected cytosine residue locations in the target nucleic acid sequence.
[0021] In certain exemplary embodiments, protected cytosine nucleotides include methylated cytosine residues. In certain exemplary embodiments, unprotected cytosine nucleotides are unmethylated cytosine residues. In certain exemplary embodiments, a double-length DNA template for a method of identifying epigenetic information includes a unique molecular identifier (UMI) and / or one or more sequencing indices (SIDs).
[0022] In certain exemplary embodiments, a diploid DNA template formed by the methods and compositions described herein is provided. For example, the diploid DNA template comprises a first copy and a second copy of a target DNA template, the first and second copies of the target DNA template being adjacent to each other via a double-stranded crosslinking region. Furthermore, each polynucleotide chain of the diploid DNA template comprises a parent template strand and a daughter strand copy of the parent template strand derived from the target DNA template. The parent template strand is adjacent to the daughter copy strand of the parent template strand, for example, via a crosslinking region. In addition, each copy of the target DNA template within the diploid DNA template comprises a parent template strand and a daughter strand that is complementary to and hybridizes with the parent template strand.
[0023] In certain exemplary embodiments, the double-length DNA template includes a first and a second terminal end, either of which includes a sequence encoding a primer binding site. In certain exemplary embodiments, the crosslinking region (or its strand) includes a unique molecular identifier (UMI) and / or a sequencing index (SID).
[0024] These and other aspects, purposes, features and advantages of the exemplary embodiments described will become apparent to those skilled in the art when considering the following detailed description of the exemplary embodiments. [Brief explanation of the drawing]
[0025] [Figure 1A] Figure 1A shows a linear end adapter for synthesizing a double-length DNA template, according to a specific exemplary embodiment. [Figure 1B]Figure 1B is a schematic diagram showing the circularization of a target DNA template using EA according to a particular exemplary embodiment. [Figure 1C] Figure 1C is a schematic diagram showing the binding of polymerase and the initiation of bidirectional elongation of a circular construct according to a particular exemplary embodiment. [Figure 1D] Figure 1D is a schematic diagram showing the continuous polymerase elongation of a circular construct and the formation of a double-length DNA template according to a particular exemplary embodiment. [Figure 2A] Figure 2A is a diagram showing a Y-branched end adapter 200 (YBEA) according to a particular exemplary embodiment. [Figure 2B] Figure 2B is a schematic diagram showing the circularization of a target DNA template using YBEA200 according to a particular exemplary embodiment. [Figure 2C] Figure 2C is a schematic diagram showing the binding of polymerase and the initiation of bidirectional elongation of a circular construct containing YBEA200 according to a particular exemplary embodiment. [Figure 2D] Figure 2D is a schematic diagram showing the continuous polymerase elongation of a circular construct and the formation of a target DNA template using YBEA200 according to a particular exemplary embodiment. [Figure 2E] Figure 2E is a diagram showing the double-length DNA template (lower panel) of Figure 2D in denatured (single-stranded) form that provides a predetermined oligonucleotide primer binding sequence when the original Y-branched element is replicated. [Figure 3A] Figure 3A is a diagram showing a Y-branched end adapter containing a UMI ("YB-UMI-EA") according to a particular exemplary embodiment. [Figure 3B-C] Figure 3B is a schematic diagram showing the circularization of a target DNA template using YB-UMI-EA300 according to a particular exemplary embodiment. Figure 3C is an enlarged view of a portion of the target DNA template of Figure 3B showing an exemplary nucleic acid sequence according to a particular exemplary embodiment. [Figure 3D] Figure 3D is a schematic diagram showing the binding of polymerase and the initiation of bidirectional elongation of a circular construct containing YB-UMI-EA300 according to a particular exemplary embodiment. [Figure 3E] Figure 3E is a schematic diagram showing the continuous polymerase elongation of a circularized target DNA template and the formation of a double-length DNA template using an exemplary embodiment of YB-UMI-EA300 according to a particular exemplary embodiment. [Figure 3F] Figure 3F is a schematic diagram showing an exemplary bisulfite conversion of a double-length DNA template and its PCR amplification product using a Y-branched end adapter having the UMI (i.e., YB-UMI-EA) of Figure 3A, according to a particular exemplary embodiment. [Figure 3G] Figure 3G is a schematic diagram showing both intra- and inter-strand bioinformatics analysis of a portion of a double-length DNA template to confirm epigenetic information associated with the original target DNA template, according to a specific exemplary embodiment. [Figure 4A] Figure 4A shows a Y-branched terminal adapter (i.e., YB-UMI / SID-EA) containing two SID sequences and one UMI, according to a particular exemplary embodiment. [Figure 4B] Figure 4B shows a double-length DNA template resulting from the use of YB-UMI / SID-EA400 in Figure 4A, according to a specific exemplary embodiment. [Figure 5A] Figure 5A shows a modified Y-branched terminal adapter according to Figure 2A, according to a particular exemplary embodiment, however, it is modified to accept binding of a single polymerase and unidirectional extension only. [Figure 5B] Figure 5B is a schematic diagram showing the binding of polymerase and the initiation of unidirectional extension of the cyclic construct according to a particular exemplary embodiment. [Figure 5C] Figure 5C is a schematic diagram illustrating the continuous polymerase elongation of a cyclic construct and the formation of an asymmetric template using modified YBEA500 according to a specific exemplary embodiment. [Figure 6] Figure 6 is a schematic diagram illustrating the formation of a quadruple-length DNA template from a double-length DNA template according to a specific exemplary embodiment. [Modes for carrying out the invention]
[0026] overview Methods and compositions for preparing DNA libraries for replicating target nucleic acid sequences are disclosed herein. For example, a target DNA template containing or encoding a target nucleic acid sequence is extended by adding a single copy of the target DNA template to the original target DNA template, thereby forming a double-length DNA template. That is, the double-length DNA template is "double-length" in that it contains two copies of the original target DNA template (i.e., two copies in the case of a target sequence). Generally, this method includes, for example, the step of circularizing the target DNA template and then replicating it to form two copies of the target DNA template, each copy located within the double-length DNA template.
[0027] Beneficially, each strand of the diploid DNA template contains a parent polynucleotide sequence ligated adjacent to a newly synthesized daughter copy of the parent polynucleotide sequence. Furthermore, each copy of the target DNA template within the diploid DNA template contains a parent strand hybridized to a complementary daughter DNA strand. In certain examples, predetermined sequences, such as primer sequences, unique molecular identifiers (UMIs), and sample indices (SIDs), can also be included in the diploid DNA template. Moreover, the association between the parent and daughter polynucleotide sequences within the diploid DNA template allows sequencing of the diploid DNA template to beneficially reveal genetic and epigenetic information related to the target nucleic acid sequence.
[0028] To facilitate the preparation of a double-length DNA template, in certain examples, a linear end adapter (EA) is provided that comprises hybridized polynucleotide strands and thus forms a polynucleotide double helix such as a DNA molecule. For example, the ends of the EA are ligated to the opposing ends of a target DNA template, respectively, to form a circular construct. The EA contains juxtaposed nick sites (one on each polynucleotide strand) separated by a spacer region. Since each nick site is located on a polynucleotide strand of the EA double helix, each nick site is adjacent to the 5' and 3' ends. Thus, in certain examples, the EA provides an exposed 3' end for polymerase binding and extension on each strand of the EA.
[0029] For example, when a circular construct containing an EA is brought into contact with DNA polymerase, the two juxtaposed 3' ends can be extended in opposite directions by the polymerase, while the opposing strand of the target DNA template is replaced. The complete extension of both free 3' ends provided by the EA results in a diploid DNA template, where each copy of the target DNA template contains one original (parental) DNA strand and one newly synthesized complementary daughter strand. Each copy of the target DNA template is separated via the EA, which forms a crosslink between the two template copies. Thus, the crosslinks in the diploid DNA template originate from the EA. Furthermore, each polynucleotide strand of the diploid DNA template contains a parental polynucleotide sequence derived from the target DNA template and a new daughter copy of the parental polynucleotide sequence, where the parental sequence and daughter copy are adjacent to each other, covalently linked, and have the same sequence.
[0030] In certain cases, single-stranded (ss) branched sequence elements (or Y-branched elements) can be added to the 5' end of each nick site in the EA to form one or more Y-branched end adapters within the double-length DNA template. The Y-branched elements may include, for example, a polynucleotide sequence encoding a primer binding site. For example, the Y-branched elements may include a single-stranded polynucleotide sequence (e.g., ssDNA) whose complementary strand encodes a primer binding site as described herein. The primer binding sites can be used, for example, in a subsequent PCR reaction to efficiently and accurately amplify the double-length DNA template (therefore amplifying the original target DNA template).
[0031] In certain cases, the methods disclosed herein advantageously provide a diploid DNA template in which a parent polynucleotide sequence is covalently bonded to and adjacent to a daughter polynucleotide strand copy, so that both epigenetic information (parent strand) and genetic information (daughter strand) are stored in the diploid DNA template. That is, since both polynucleotide strands of the diploid DNA template composition provided herein contain a parent polynucleotide sequence derived from the target DNA template and a daughter copy of that parent sequence, parent strand methylation can be identified using strand-specific analysis and comparison, thereby identifying epigenetic information associated with the parent strand and, consequently, present in the target sequence. Furthermore, such genetic and epigenetic information can be usefully obtained in a single read by sequencing the diploid DNA template.
[0032] In certain cases, a diploid DNA template containing a unique molecular identifier (UMI) can be prepared using the method provided herein. The UMI can be included, for example, in the spacer region of the terminal adapter provided herein, i.e., in the region between the juxtaposed nick sites of the terminal adapter. In such cases, a Y-branch element can also be included to enable subsequent PCR amplification. By including a UMI in the diploid DNA template, the diploid DNA template can be used for a variety of bioinformatics applications. For example, sequence information from each strand of the diploid DNA template can be bioinformatically paired to advantageously verify the accuracy of sequence reads. Such UMIs can also, in certain cases, help distinguish strands for genetic and epigenetic analyses described herein.
[0033] In certain examples, the methods provided herein can be usefully used to prepare double-stranded DNA template compositions containing one or more sample indices (SIDs). Conventionally, the use of such SIDs is very useful in applications such as DNA multiplexing, i.e., processing multiple different samples simultaneously. For example, different SIDs can be included adjacent to the Y-branched sequence elements described herein. Double-stranded DNA molecules with different SIDs can then be processed simultaneously, and the SIDs allow for the differentiation of samples after sequencing. Furthermore, since multiple copies of the SID can appear on a single replicated PCR product strand, the SIDs can be determined with high bioinformatics accuracy, thereby reducing or eliminating the need for additional error correction. In such exemplary embodiments, the SIDs can also be used as landmarks on a given strand, allowing for further analysis.
[0034] In certain cases, methods and compositions for producing a double-length DNA template can be applied sequentially to increase the number of parent target DNA templates on a single molecule with each iteration, for example, to produce a quadruple-length or multi-length DNA template. This can be useful in sequencing applications, for example, to generate additional template reads in a single pass, thereby achieving higher read accuracy and reliability. In yet another example, an asymmetric DNA template can be obtained by asymmetrically extending the target DNA template. For example, a nick site in the terminal adapter can be blocked, thereby allowing extension from a single nick site.
[0035] Since the double-length DNA template may be limited to a single-copy extension product (i.e., forming a double template of the original parent target DNA template), the methods and compositions provided herein also beneficially maintain library length uniformity and read efficiency. The methods and compositions provided herein also improve sequencing accuracy while balancing other important properties of the sequencing system, such as throughput, efficiency, and read length. These and other examples and advantages will become apparent to those skilled in the art upon consideration of the further detailed description provided herein.
[0036] Terms and technical terms The present invention will be described in detail by reference only, using the following definitions and examples. All patents and publications referred to herein, including all sequences disclosed in such patents and publications, are expressly incorporated in their entirety by reference.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in the field to which this invention pertains. Such common techniques and methods are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition), Volumes 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, 2012 (hereinafter "Sambrook"), and in Current Protocols in Molecular Biology, edited by FMAusubel et al., which was first published in book form in 1987 by Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., regularly supplemented until 2011, and is now published by Wiley & Sons, Inc. and available online in journal form as Current Protocols in Molecular Biology, Volumes 00-130 (1987-2020) in the Wiley Online Library. Each of these documents provides a general dictionary of many of the terms used in this invention.
[0038] Any methods and materials similar to or equivalent to those described herein may be used in carrying out or testing the present invention, but preferred methods and materials are described herein. It should be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit them. For the purposes of interpreting this disclosure, the following definitions of terms apply, and where appropriate, terms used in the singular form also include the plural form, and vice versa.
[0039] Furthermore, the features, operations, or characteristics described herein can be combined in any suitable manner to form various implementations of the exemplary embodiments. On the other hand, those skilled in the art will fully understand that specific steps or operations for illustrating the method can also be interchanged or modified in terms of their order. Accordingly, the various orders in this specification and drawings are merely for the purpose of illustrating specific embodiments and are not essential orders unless it is specifically stated that a particular order must be followed or that such an order is necessary from the context (for example, polymerase must be added to the reaction mixture for polymerase-mediated replication to occur).
[0040] Unless otherwise specified, nucleic acids are written from left to right, in the 5' to 3' direction. Amino acid sequences are written from left to right, in the amino to carboxy direction.
[0041] The headings provided herein are not limitations on the various aspects or embodiments of the invention that can be obtained by referring to the entire specification. Accordingly, the terms defined below are further defined by referring to the entire specification.
[0042] As used herein, the singular forms "a," "an," and "the" refer to multiple objects unless the context explicitly indicates otherwise.
[0043] A range may be expressed herein as a range from a certain value "about" or "approximately" to another specific value "about" or "approximately". Where such a range is expressed, the alternative aspect includes that certain value and / or that other specific value. It will be further understood that the endpoint of each range is significant both in relation to the other endpoint and independently of the other endpoint. Similarly, it will be understood that, by the use of the antecedent "about," when a value is expressed as an approximation, the specific value forms an alternative aspect.
[0044] In certain exemplary embodiments, the terms “about” or “approximately” are understood to mean, for example, within two standard deviations of the mean, as is common in the art. “About” or “approximately” can be understood to mean within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise evident from the context, all numerical values provided herein can be modified by the term “about.” Furthermore, terms used herein such as “example,” “exemplary,” or “exemplified” are not intended to indicate priority, but rather to indicate that the embodiments described later are merely examples of the embodiments proposed.
[0045] The term "amplification" refers to the process of creating additional copies of a target nucleic acid. Amplification can have two or more cycles, for example, multiple cycles in exponential amplification. Amplification may have only one cycle (creating a single copy of the target nucleic acid). This copy may contain additional sequences, for example, sequences present in the primers used for amplification. Amplification can also produce a single-strand copy (linear amplification) or a preferential single-strand copy (asymmetric PCR).
[0046] As used herein, “polymerase” refers to an enzyme that catalyzes the polymerization (i.e., polymerase activity) of nucleotides. Generally, this enzyme initiates synthesis at the 3' end of a primer annealed to a polynucleotide template sequence and proceeds toward the 5' end of the template strand. “DNA polymerase” catalyzes the polymerization of deoxynucleotides by sequentially adding nucleotides to, for example, a free 3'-hydroxyl group, using a complementary template DNA strand and primer. The template strand determines the sequence of the added nucleotides by Watson-Crick base pairing.
[0047] In general, any DNA polymerase suitable for use in rolling circle amplification reactions, for example, can be used in the replication reaction. In certain embodiments, the suitable DNA polymerase will have strand displacement activity. The term strand displacement refers to the ability to replace downstream DNA encountered during DNA synthesis. Several DNA polymerases with varying degrees of strand displacement activity are known and commercially available in the art. In certain exemplary embodiments, the polymerase is phi29 polymerase, bst polymerase, etc. In preferred embodiments, the strand displacement polymerase is phi29 polymerase.
[0048] In some embodiments, the DNA polymerase is a high-fidelity DNA polymerase. The fidelity of the DNA polymerase is the result of accurate replication of the desired template. Specifically, this involves multiple steps, including the ability to read the template strand, select the appropriate nucleoside triphosphate, and insert the correct nucleotide into the 3' primer end, so that Watson-Crick base pairings are maintained. In addition to the effective identification of correct and incorrect nucleotide incorporation, some DNA polymerases have 3'→5' exonuclease activity. This activity, known as "proofreading," is used to remove an incorrectly incorporated mononucleotide and then replace it with the correct nucleotide.
[0049] In certain embodiments, suitable high-fidelity DNA polymerases for carrying out the present invention include KAPA HiFi DNA Polymerase, commercially available from Roche Diagnostics Corp., Q5® High-Fidelity DNA Polymerase, commercially available from New England Biolabs, Inc., and engineered Pfu DNA polymerases such as Pfu-X, commercially available from Jena Biosciences.
[0050] As used herein, terms such as “ligate,” “the act of ligating,” and “ligation” generally refer to a method of covalently bonding two or more molecules to one another, for example, two or more nucleic acid molecules to one another. Similarly, the term “ligatable” refers to having the ability to ligate. As those skilled in the art will understand, ligation involves a condensation reaction that forms a covalent bond between the ends of a first nucleic acid molecule and the ends of a second nucleic acid molecule.
[0051] In certain exemplary embodiments, ligation may involve forming a covalent bond between the 5' phosphate group of one nucleic acid and the 3' hydroxyl group of a second nucleic acid, thereby forming a ligated nucleic acid molecule. Generally, for the purposes of this disclosure, a target DNA template sequence can be ligated to terminal adapters to produce a circularized construct. Ligation involves the linking of two DNA molecules, each having overhanging ends (i.e., “sticky” ends), i.e., one strand being longer than the other (typically by at least a few nucleotides), so that the longer strand has unpaired bases. Ligation also involves the linking of DNA molecules, each having strands of equal length (i.e., “blunt ends” without overhangs).
[0052] In certain exemplary embodiments, ligation can be achieved using an asymmetric 5' thymine nucleotide overhang on the target DNA template and a 5' adenine nucleotide overhang on the terminal adapter. For example, ligation can be performed by combining the target DNA template and the terminal adapter at equimolar or nearly equimolar concentrations. In certain exemplary embodiments, the concentrations of the terminal adapter and the target DNA template can be optimized by trial and error to favor circular ligation over ligation, for example, the molar ratio of the target DNA template to the adapter may be 1:1, 1:5, 1:10, 1:25, or 1:50. In certain exemplary embodiments, improved circularization can be achieved if the target DNA template and / or terminal adapter have sufficient flexibility to bend and align for a sufficient amount of time and frequency. It is known that ds-DNA longer than 200 base pairs ligates to form "minicircles," and that minicircles containing linear ds-DNA oligos with nick sites can be further easily circularized (see, for example, “Small DNA Circles as Probes of DNA Topology”, Bates, AD et al., Biochem. Soc. Trans. (2013) 41, 565-570, the entire work of which is incorporated herein by reference). The target sequencing library is often in this size range. In certain exemplary embodiments, the target DNA template is 200–500 base pairs long.
[0053] In certain exemplary embodiments, cyclization of the target DNA template can be promoted by reducing the concentration of the target DNA template and / or terminal adapter, thereby promoting cyclization rather than concatemerization. In other embodiments, cyclization can be promoted by a “protein scaffold” strategy that overcomes the energetic challenge of physically bending the DNA to form a small circle by increasing the local concentration of intramolecularly ligable ends using one or more DNA-binding proteins to shift the equilibrium toward cyclization. In certain exemplary embodiments, suitable DNA-binding proteins for protein scaffold formation include histones, Abf2p, DSP1, histone-like protein AU, and CAP. In certain exemplary embodiments, since the cyclized construct does not present free ends to initiate exonuclease-mediated DNA degradation, the cyclized ligation construct can be enriched by treatment with one or more exonucleases. Specific exemplary exonucleases include ExoVIII, ExoIII, and T5 exonucleases.
[0054] As used herein, the terms “target,” “target sequence,” or “target nucleic acid sequence” are used interchangeably to refer to any target nucleic acid molecule subjected to a process for generating the double-length DNA template described herein. A target nucleic acid sequence may include or consist of genomic DNA, subgenomic DNA, chromosomal DNA (e.g., from an isolated chromosome or portion of a chromosome, e.g., from one or more genes or loci of chromosome origin), mitochondrial DNA, chloroplast DNA, plasmid or other episome-derived DNA (or recombinant DNA contained therein), or double-stranded cDNA produced by reverse transcription of RNA, or RNA that can subsequently be converted to cDNA by any method known in the art. Furthermore, a target nucleic acid sequence, such as target DNA or RNA, may be obtained from any in vivo or in vitro source, such as one or more living or dead cells, tissues, organs, bodily fluids, or organisms, or from any biological or environmental source (e.g., water, air, soil).
[0055] The terms "DNA," "double-stranded DNA," or "dsDNA" generally refer to complementary deoxyribonucleic acid polynucleotide chains that hybridize to form a double helix. The two polynucleotide chains are linked by hydrogen bonds between complementary nucleotide base pairs (i.e., Watson-Crick). Each nucleotide in DNA consists of a sugar molecule, a phosphate group, and one of four nitrogen bases: adenine (A), cytosine (C), guanine (G), or thymine (T). The chains do not need to be perfectly complementary to maintain the double helix. Double-stranded DNA can be found in the nucleus of eukaryotic cells, as well as in the cytoplasm and plasmids of prokaryotic cells. It can also be used in various molecular biology techniques such as PCR (polymerase chain reaction), DNA sequencing, and genetic engineering.
[0056] A DNA strand, or single-stranded DNA (ssDNA), refers to one of the polynucleotide strands of a DNA molecule and may also be called ssDNA. A daughter polynucleotide strand is, for example, a new strand of a DNA double helix produced from the replication of a DNA molecule. For example, polymerase-mediated replication reactions use a template DNA strand to produce a complementary strand, which is the daughter strand. In certain exemplary embodiments, the DNA is cDNA that has been converted to a target RNA sequence, or otherwise derived from a target RNA sequence.
[0057] As used herein, the terms “target DNA template” and “DNA template” are interchangeable and refer to a DNA molecule that codes for or contains the genetic and / or epigenetic information of a target nucleic acid sequence. For example, one strand contains or codes for the target sequence, and the other hybridized strand of the DNA molecule is complementary to the strand containing or coding for the target sequence. In certain embodiments, the target DNA template may be a native DNA target fragment (e.g., a genome or cell-free DNA target fragment) or a cDNA copy of a native DNA or RNA target fragment. The target DNA templates disclosed herein are molecules that are replicated (e.g., copied) and / or subjected to DNA sequencing. Furthermore, if a subsequent DNA molecule is formed that contains, for example, a polynucleotide strand of the target DNA template, that strand may be referred to as the “original” or “parent” strand of the target DNA template to indicate that it was originally part of the target DNA template. The target template can be prepared, for example, by any means known in the art.
[0058] The term "primer" refers to a single-stranded oligonucleotide that hybridizes with a target nucleic acid sequence ("primer binding site") and can act as a starting point for synthesis along the complementary strand of the nucleic acid under conditions suitable for such synthesis. In other words, a "primer" functions as a substrate that allows polymerase to polymerize nucleotides. In various embodiments, primers have a free 3'-OH group that can be extended by nucleic acid polymerase. In the case of template-dependent polymerases, typically at least the 3' portion of the primer oligonucleotide is complementary to a portion of the template nucleic acid that it "binds" (or "complexes", "anneals", or "hybridizes") to the template by hydrogen bonds and other molecular forces, giving a primer / template complex to initiate synthesis by DNA polymerase, which is then extended during DNA synthesis by the addition of complementary covalent bases to the template that are bound to their 3' portions (i.e., "primer extension").
[0059] As used herein, a unique molecular identifier (UMI) is a sequence of nucleotides inserted into or identified within a DNA molecule, which can be used to distinguish individual DNA molecules from one another. Due to their complementary nature within DNA molecules, UMIs present in or inserted into DNA molecules can also be used to identify individual strands of DNA molecules, insofar as the polarity (direction) of the UMI sequence can be identified and distinguished between two complementary DNA strands. See, for example, Kivioja, Nature Methods 9, 72-74 (2012). UMIs can be sequenced together with the DNA molecules to which they are bound to determine whether a read sequence belongs to one source DNA molecule or another. The term "UMI" is used herein to refer to both the sequence information of a polynucleotide and the physical polynucleotide itself. UMI sequences may be random, pseudo-random, partially random, or non-random nucleotide sequences, for example, inserted into or otherwise incorporated within the terminal adapters described herein.
[0060] The term "sample index" refers to a sequence of nucleotides attached to a target polynucleotide, which identifies the source of the target polynucleotide (i.e., the sample from which the target polynucleotide originates). Therefore, a sample index (or SID) is also called a "sample identifier sequence," "index sequence identifier," "multiple identifier," or "MID." In use, each sample contains a different sample index sequence (e.g., one sequence is attached to each sample, and different samples are attached to different sequences), and the samples are pooled. After the pooled samples are sequenced, the sample identifier sequences can be used to identify the source of the sequences. Traditionally, a sample identifier sequence may be attached to the 5' end or the 3' end of a polynucleotide. In specific cases, part of the sample identifier sequence may be at the 5' end of the polynucleotide, and the remainder may be at the 3' end. If the elements of the sample identifier have sequences at each end, the 3' and 5' sample identifier sequences together identify the sample. In specific examples, the sample identifier sequence is only a subset of the bases attached to the target oligonucleotide. Furthermore, as described herein, terminal adapters may be used to include SIDs in the sample.
[0061] As used herein, the term “polymerase chain reaction” (or “PCR”) refers to a method of increasing the concentration of a desired polynucleotide segment in a mixture of genomic DNA without cloning or purification. See U.S. Patents 4,683,195 and 4,683,202 in whole (which describe the PCR process). The process for amplifying a desired polynucleotide generally consists of repeated cycles of denaturation, primer-annealing, and extension using the DNA polymerase enzyme. The amplified segments of the desired polynucleotide are said to be “PCR-amplified” because they become the dominant nucleic acid sequence (in terms of concentration) in the mixture. In modifications of the methods described above, the target nucleic acid molecule may be PCR-amplified using multiple different primer pairs (possibly one or more primer pairs for each target nucleic acid molecule of interest) to form a multiplex PCR reaction.
[0062] As used herein, the term “terminal adapter” generally refers to a polynucleotide double helix, such as a DNA molecule, that can be attached (i.e., ligated) to a target DNA template. Terminal adapters may be 5 to 100 nucleotides long and may provide, include, or encode amplification primer binding sites, sequencing primer binding sites, molecular identifiers, and / or sample identifier sequences, as described herein. Terminal adapters can be attached to both the 5' and 3' ends of a target DNA template by ligation. For example, when a terminal adapter is attached to a target DNA template, it forms a cyclic structure ("cyclic DNA construct" or "cyclic construct") in which both ends of the target molecule are bound to the ends of the terminal adapter.
[0063] Double-length DNA template Referring here to the drawings, similar reference numerals throughout the drawings indicate similar (but not necessarily identical) elements, and exemplary embodiments are described in detail. Furthermore, while certain figures provided herein show target DNA template ligation, circularization, and replication of a single target DNA template, it should be understood that, in general, multiple target DNA templates can be ligated, circularized, and replicated in a single library preparation reaction, for example, when multiple reaction components (e.g., multiple target DNA templates, terminal adapters, polymerases, etc.) are combined. Multiple replicas can then be used for a variety of different applications, such as sequencing or other analyses.
[0064] In certain exemplary embodiments, a method for preparing a DNA library is provided, which includes synthesizing a double-length DNA template from a target nucleic acid by using a linear end adapter (EA). This is shown in Figures 1A to 1D, which summarize the features of an exemplary EA according to a particular exemplary embodiment and illustrate a method for synthesizing a double-length DNA template using the exemplary EA.
[0065] Referring to Figure 1A, a linear end adapter (EA) for synthesizing a double-length DNA template is shown according to a particular exemplary embodiment. As shown, EA100 is a DNA molecule comprising a double-stranded polynucleotide molecule, for example, a hybridized oligonucleotide chain, i.e., a first polynucleotide chain 100a (shown as a circle) and a hybridized second polynucleotide chain 100b (shown as a rectangle). As used herein, in relation to the structure of the EA of the present invention, the term “polynucleotide chain” refers to one or more oligonucleotides having the same 5’ to 3’ polarity that hybridize with a portion of one or more complementary oligonucleotides to form the EA structure 100. In the embodiment shown in Figure 1A, polynucleotide chains 100a and 100b each contain two oligonucleotide portions separated by nick sites 101a and 101b, respectively (as will be further described below). In this regard, the reference to “100a” refers to the entire 5’→3’ chain, with nick site 101a located within chain 100a. Similarly, the reference to "100b" refers to the entire 5'→3' chain hybridized to chain 100a, with the nick site 101b located within chain 100b. In certain exemplary embodiments, the total length of the EA is 50 to 100 nucleotides long, for example, 75 to 80 nucleotides long. In certain exemplary embodiments, the length of the oligonucleotide used to generate the EA is selected to ensure efficient and specific hybridization to form a stable EA structure, as will be further discussed herein.
[0066] As also shown in the exemplary EA of Figure 1A, the EA contains a first nick site 101a and a second nick site 101b. That is, in a particular exemplary embodiment, the EA contains one internal first and second nick site 101a and 101b in each of the first and second polynucleotide chains 100a and 100b of EA100. A nick site contains, for example, any break or gap in one strand of the DNA molecule so that the chain is not continuous. In a particular exemplary embodiment, the nick site is a break or discontinuity in the phosphodiester backbone, but in other exemplary embodiments, the nick site is a gap of one or more nucleotides in the DNA chain. In particular, each nick site 101a and 101b is associated with and adjacent to the 5' free end and the 3' free end. Referring to Figure 1A, the illustrated lengths and locations of the nick sites 101a and 101b are shown for illustrative purposes only and are not intended to be limiting.
[0067] In certain exemplary embodiments, by exposing the 3' ends of nick sites 101a and 101b, the EA can facilitate polymerase-mediated chain elongation reactions. That is, the polymerase can use the exposed 3' ends to elongate the 3' associated chain in conventional polymerization and chain substitution reactions, as described herein. Preferably, as described herein, the nick sites 101a and 101b are separated by a spacer region 102 so that the EA can accept binding of two polymerases for bidirectional elongation. As shown, for example, the spacer region linearly offsets the first nick site 101a from the second nick site 101b. Thus, the nick sites 101a and 101b can be separated by the spacer region 102 so that the binding of one polymerase does not sterically interfere with and / or substitute for the binding of the second polymerase. EA100 also includes terminal ends 103 and 104 adjacent to each nick site, and each terminal 103 and 104 is suitable for efficient ligation to the ends of the target DNA template. In other words, the ends of EA are ligable to the target DNA template.
[0068] An EA can be formed or otherwise prepared using any means known in the art. For example, as shown in Figure 1A, in one embodiment, an EA is formed by hybridization of four synthetic oligonucleotides that are not completely contiguous, thereby leaving spacer regions (also referred to herein as “gaps” or “nicks”) during hybridization. Any other suitable method can be used to generate nicks, gaps or other sites for polymerase binding and initiation of DNA synthesis. For example, an EA can be produced from a contiguous oligonucleotide chain designed to contain recognition sites for one or more appropriately positioned nicking enzymes (i.e., nicking endonucleases). Nicking enzymes are known in the art and hydrolyze (cleave) only one strand of a DNA double helix to produce a DNA molecule that is “nicked” rather than cleaved. Treatment of an EA with a nicking enzyme generates a free 3' end on each strand that provides a polymerase initiation site.
[0069] Referring to Figure 1B, a schematic diagram is shown illustrating the circularization of a target DNA template using EA100 according to a particular exemplary embodiment. The target DNA template contains or encodes, for example, a target sequence. Prior to ligation of the adapter, the ends of the target DNA template can be prepared for ligation, for example, by end repair and the creation of blunt ends with a 5' phosphate group. The DNA template can be blunt-ended by several methods known to those skilled in the art. In a particular method, the ends of the fragmented DNA are "polished" with T4 DNA polymerase and Krenow polymerase, procedures well known to those skilled in the art, and then phosphorylated with a polynucleotide kinase enzyme. Then, using Taq polymerase or Krenow exominus polymerase enzyme, a single "A" deoxynucleotide is added to both 3' ends of the DNA molecule to produce a single nucleotide 3' overhang complementary to the single nucleotide 3' "T" overhang at the double-stranded ends of the adapter.
[0070] As shown in Figure 1B, the double-stranded EA100 is combined with a target DNA template 107 having a first terminal end 105 and a second terminal end 106. As shown herein, the target DNA template includes complementary polynucleotide strands, namely a first template strand 107a (dashed line) and a second template strand 107b (solid line), both of which are referred herein as “parent strands” or “parental strands.” That is, strands 107a and 107b of the target DNA template 107 correspond to the original strands of the target DNA template, and the target DNA template contains or encodes the target sequence described herein. In some embodiments, the parent strands contain epigenetic information, such as methylated cytosine residues.
[0071] In step 1a, for example, EA100 is ligated to either end of the target DNA template 107, which includes parent polynucleotide strands 107a and 107b. For example, the terminal end 103 of EA100 is ligated to the template end 106 (Figure 1B). Alternatively, in step 1a, although not shown for simplification, the other terminal end of EA100 (i.e., 104) is ligated to the terminal end 105 of the target DNA template 107. In either case, one end of EA100 is ligated to the end of the target DNA template.
[0072] In step 1b of Figure 1B, the remaining free end of EA100 is ligated to the remaining free end of the target DNA template 107 to form a circular construct 109. For example, if the terminal end 103 of EA100 is ligated to the template end 106 in step 1a, then in step 1b, the EA terminal end 104 of EA100 is ligated to the template terminal end 105, thereby forming the circular construct 109 of the original (parent) template 107. Alternatively, if the terminal end 104 of EA100 is ligated to the template end 105 in step 1a, then in step 1b, the EA terminal end 103 is ligated to the template terminal end 106, thereby forming the circular construct 109. In either case, in step 1b of Figure 1B, the two terminal ends 103 and 104 of EA100 are linked to the respective ends 105 and 106 of the template 107, thereby forming a DNA crosslink 108 between the ends of the template 107. That is, the entirety of EA100 forms a DNA crosslink 108 between the two ends 105 and 106 of the parent template 107. This forms a ring structure 109 containing EA100 (as a crosslink 108) together with the complementary parent template strands 107a and 107b of the parent template 107.
[0073] In this way, EA100 (in Figure 1A) acts as a crosslinking precursor for the crosslinking region 108 of the cyclic construct 109. As shown, the first and second nicks 101a and 101b remain within the cyclic construct 109 (as part of the crosslinking region 108), and thus, in certain exemplary embodiments, provide two respective 3' ends available for polymerase binding and bidirectional extension, as described herein.
[0074] Continuing the above example, Figure 1C is a schematic diagram showing the binding of polymerases and the initiation of bidirectional extension of the cyclic construct 109 according to a particular exemplary embodiment. As shown (in step 1b, Figure 1B), once the cyclic construct 109 is formed, DNA polymerases, indicated as DNA polymerases 110a and 110b, are added to initiate the replication reaction. For example, in one embodiment, DNA polymerase 110a binds to the nick site 101a of EA100. In this regard, the nick site 101a, along with its available 3' end, causes the primer end to function for the binding and extension initiation of DNA polymerase 101a. Similarly, DNA polymerase 110b binds to the nick site 101b of EA100, and the nick site 101b at its 3' end functions as the primer end for the binding and extension of DNA polymerase 110b. As indicated by the opposing arrows, polymerases 110a and 110b are positioned to extend the annular structure 109 bidirectionally in opposite directions (Figure 1C, top panel).
[0075] In step 1c of Figure 1C, the first and second polymerases 110a and 110b extend the cyclic construct 109 bidirectionally in opposite directions (see Figure 1C, bottom panel, arrows). For example, polymerase 110a extends the 3' end of the nick site 101a while simultaneously substituting the 5' end of the nick site 101a (and its associated parent template strand 107a). That is, as polymerase 110a proceeds, it uses the parent strand 107b as a template to extend the 3' end of the nick site 101a', synthesizing a new daughter strand 107a that is complementary to the sequence of the parent strand 107b (and thus shares the same sequence as the parent strand 107a). The new daughter strand 107a' also contains a replicated ssDNA daughter strand crosslinking portion 108a as part of the DNA crosslinking region 108.
[0076] Similarly, polymerase 110b uses the parent strand 107a as a template to extend the 3' end of the nick site 101b, while also substituting the 5' end of the nick site 101b (and its associated parent template strand 107b) (Figure 1C, bottom panel). In other words, as step 1c proceeds, polymerase 110b uses the parent strand 107a as a template to extend the 3' end of the nick site 101b to synthesize a new daughter strand 107b' that is complementary to the sequence of the parent strand 107a (and therefore shares the same sequence as the parent strand 107b). The new daughter strand 107b' also contains the replicated ssDNA daughter strand crosslinking portion 108b as part of the crosslinking region 108.
[0077] Figure 1D is a schematic diagram showing the continuous polymerase extension of a cyclic construct 109 and the formation of a double-length DNA template according to a particular exemplary embodiment. As shown (top panel), polymerases 110a and 110b continue to the ends of the parent template 107. For example, polymerase 110a continues along the parent template strand 107b to the 5' end of the parent template strand 107b, completing the synthesis of a new daughter strand 107a'. Similarly, polymerase 110b continues along the parent template strand 107a to the 5' end of the parent template strand 107a, completing the synthesis of a new daughter strand 107b'. In step 1d, once polymerases 110a and 110b have completed the synthesis of daughter strands 107a' and 107b', respectively, polymerases 110a and 110b dissociate from the cyclic construct 109 to form a double-length DNA template 111 (bottom panel, as shown in Figure 1D).
[0078] As shown in Figure 1D (bottom panel), the double-length DNA template 111 contains two copies of the original target DNA template 107, namely the first and second copies 111a and 111b, respectively, each adjacent to a crosslinking region 108, which contains a spacer region 102. In particular, each template copy contains both the parent polynucleotide strand (shown in black) and the newly synthesized daughter polynucleotide strand (shown in gray). For example, template copy 111a contains the original (parent) template strand 107b and the newly synthesized daughter strand 107a'. On the opposite side of the crosslinking region 108, template copy 111b contains the original (parent) template strand 107a and the newly synthesized daughter strand 107b'. Furthermore, the double-length DNA template 111 contains the first terminal end 112a and the second terminal end 112b. The first terminal end 112a includes, for example, a portion 100b of the EA100 strand associated with the 5' end EA100 of the nick site 101b (a white-outlined black rectangle in terminal end 112a) and a copy thereof (a white-outlined gray circle in terminal end 112a). Similarly, the second terminal end 112b of the double-length DNA template 111 includes a portion 100a of the EA100 strand associated with the 5' end EA100 of the nick site 101a (a white-outlined black circle in terminal end 112b) and a copy thereof (a white-outlined gray rectangle in terminal end 112b).
[0079] In particular, both strands of the double-length DNA template also contain a parent strand (black) ligated to a newly synthesized daughter copy (gray) of the parent strand. For example, parent strand 107a is covalently adjacent to the newly synthesized daughter strand 107a' in the 5'→3' direction via the strand of bridging region 108 (i.e., the strand of bridging region 108 containing strand portions 100a and 108a). Furthermore, for polymerase-mediated elongation of the cyclic construct 109 as described herein, the nucleotide sequence of parent strand 107a matches the nucleotide sequence of the new daughter strand 107a'. That is, daughter strand 107a' is a sequence copy (i.e., daughter copy) of parent strand 107a of the target DNA template.
[0080] Similarly, on the complementary strand of the double-stranded DNA template, the parent strand 107b is covalently adjacent to the new daughter strand copy 107b' via the strand of DNA crosslinking region 108 (i.e., the crosslinking strand 108 containing strand portions 100b and 108b), also in the 5'→3' direction. Similarly, and again in this case, for polymerase-mediated elongation of the cyclic construct, as described herein, the nucleotide sequence of the parent strand 107b matches the nucleotide sequence of the new daughter strand 107b'. Thus, each strand of the double-stranded DNA contains both the parent polynucleotide sequence and the daughter polynucleotide sequence copy on each strand, in addition to the respective parent template strands and their complementary daughter strands of the two target DNA copies 111a and 111b (Figure 1D, lower panel).
[0081] Duplicate DNA template with Y-branched terminal adapter In certain exemplary embodiments, the design of the end adapter (EA) shown in Figure 1A can be modified to impart additional features to the resulting double-length DNA template. These include, for example, features that facilitate subsequent PCR amplification and / or DNA sequencing. These are shown in Figures 2A to 2E, which collectively show the modified EAs and illustrate a method of using the modified EA to synthesize a double-length DNA template containing a primer-binding sequence, according to certain exemplary embodiments.
[0082] Referring to Figure 2A, a diagram is shown illustrating a Y-branched end adapter 200 (YBEA) having hybridized strands 200a (circular) and 200b (rectangular) according to a particular exemplary embodiment. As shown, the YBEA200 has a general double-stranded polynucleotide EA structure like that in Figure 1A, except that the first Y-branching element 213a and the second Y-branching element 213b are ligated to the 5' ends of the first and second nick sites 201a and 201b, respectively. Each Y-branching element 213a and 213b may contain, for example, a predetermined oligonucleotide sequence, and its design can be adapted to achieve a specific purpose, such as PCR amplification or DNA sequencing of a double-length DNA template, but is not limited thereto. Also, as with the EA in Figure 1A, the reference to "200a" refers to the entire 5'→3' strand of the YBEA200, with the nick site 201a located within strand 200a. Similarly, the reference to "200b" refers to the entire 5'→3' chain hybridized to chain 200a, with the nick site 201b located within chain 200b.
[0083] In certain exemplary embodiments, each Y-branch element 213a and 213b sequence may include a predetermined oligonucleotide sequence that provides a complementary or hybridizable primer-binding site useful, for example, for PCR amplification. That is, each Y-branch element 213a and 213b may contain, for example, 10 to 30 nucleotides, e.g., 15 to 25 nucleotides or 18 to 22 nucleotides, and its complementary strand contains a primer-binding site sequence. In certain exemplary embodiments, Y-branch elements 213a and 213b contain the same sequence, but in other exemplary embodiments, Y-branch elements 213a and 213b contain different sequences. In certain exemplary embodiments, Y-branch elements 213a and 213b are the same length, but in other exemplary embodiments, Y-branch elements 213a and 213b may be different lengths.
[0084] YBEA200 also includes terminal ends 203 and 204 adjacent to each nick site 201a and 201b, each terminal 203 and 204 being suitable for efficient ligation to the end of the target DNA template. That is, the terminal ends are ligable to the target DNA template. Preferably, as described herein, the nick sites 201a and 201b are separated by a spacer region 202 so that EA can accept binding of two polymerases for bidirectional extension. That is, the nick sites 201a and 201b are far enough apart so that the binding of one polymerase does not sterically interfere with and / or substitute for the binding of the second polymerase. This configuration is shown, for example, in Figure 2A, where the spacer region 202 linearly offsets the first nick site 201a from the second nick site 201b.
[0085] Figure 2B is a schematic diagram showing the circularization of a target DNA template using YBEA200 according to a particular exemplary embodiment. Referring to Figure 2B, YBEA200, and its respective first and second Y-branch elements 213a and 213b associated with terminals 203 and 204, respectively, are combined with the target DNA template 207 to form a circular construct 209 (similar to the formation of the circular construct 109 in Figure 1B). That is, YBEA200 is combined with the target DNA template 207, which has first and second terminal terminals 205 and 206 and includes complementary strands 207a and 207b (Figure 2B).
[0086] In step 2a, for example, YBEA200 is ligated to either end of the target DNA template 207, which includes parent polynucleotide strands 207a and 207b. For example, the terminal end 203 of YBEA200 is ligated to the template end 206. Alternatively, in step 2a, although not shown for simplification, the other terminal end of YBEA200 (i.e., 204) is ligated to the terminal end 205 of the target DNA template 207.
[0087] In step 2b of Figure 2B, the unligated (free) ends of YBEA200 are ligated to the remaining free ends of the target DNA template 207 to form a circular construct 209. For example, if the terminal end 203 of YBEA200 is ligated to the template end 206 in step 2a, then in step 2b, the YBEA terminal end 204 of YBEA200 is ligated to the template terminal end 205, thereby forming the circular construct 209 of the original (parent) target DNA template 207. Alternatively, if the terminal end 204 of YBEA200 is ligated to the template end 205 in step 2a, then in step 2b, the YBEA terminal end 203 is ligated to the template terminal end 206, thereby forming the circular construct 209.
[0088] In either case, in step 2b of Figure 2B, the two terminal ends 203 and 204 of YBEA200 join the respective ends 205 and 206 of the template 207, thereby forming a YBEA crosslinking region 208 between the ends of the template 207. That is, the entirety of YBEA200 forms a crosslinking region 208 between the two ends 205 and 206 of the parent target DNA template 207. This forms a circular construct 209 containing YBEA200 (as the crosslinking region 208) together with the complementary parent template strands 207a and 207b of the parent target DNA template 207. In this way, YBEA200 (Figure 2A) acts as a crosslinking precursor for the crosslinking region 208. Furthermore, the respective first and second nick sites 201a and 201b remain within the annular structure 209 (as part of the crosslinking region 208), thereby providing two respective 3' free ends available for polymerase binding and bidirectional extension, as described herein. Y-branched elements 213a and 213b are also present in the annular structure 209.
[0089] Continuing the above example, Figure 2C is a schematic diagram showing the binding and initiation of bidirectional extension of polymerases to a cyclic construct containing YBEA200 according to a particular exemplary embodiment. As shown in Figure 2C (and also as in Figure 1C), once the first and second polymerases 210a and 210b combine with the cyclic construct 209, they bind to the nick sites 201a and 201b, respectively, and proceed in opposite directions (as indicated by the arrows in the upper panel of Figure 2C). Polymerases 210a and 210b also replace the 5' ends of the parent chains 207b and 207a, respectively (including their respective associated 213b and 213a Y-branching elements).
[0090] In step 2c (Figure 2C, bottom panel), polymerases 210a and 210b extend the cyclic construct 209 bidirectionally in opposite directions (see arrows). Specifically, polymerase 210a, as it proceeds, uses the parent strand 207b as a template to extend the 3' end of the nick site 201a to synthesize a new daughter strand 207a' that is complementary to the sequence of the parent strand 207b (and therefore shares the same sequence as the parent strand 207a). The new daughter strand 207a' also contains the replicated ssDNA daughter strand crosslinking portion 208a as part of the crosslinking region 208. Furthermore, the Y branch element 213a remains at the 5' end of the substituted template strand 207a.
[0091] Similarly, polymerase 210b uses the parent strand 207a as a template to extend the 3' end of the nick site 201b, while also substituting the 5' end of the nick site 201b (and its associated parent template strand 207b) (Figure 2C, bottom panel). Thus, in step 2c (bottom panel), polymerase 210b, as it proceeds, uses the parent strand 207a as a template to extend the 3' end of the nick site 201b to synthesize a new daughter strand 207b' that is complementary to the sequence of the parent strand 207a (and therefore shares the same sequence as the parent strand 207b). As shown, the Y branch element 213b also remains attached to the parent strand, which has been substituted at its 5' end. The new daughter strand 207b' also contains the replicated ssDNA daughter strand crosslinking portion 208b as part of the crosslinking region 208.
[0092] Continuing with the exemplary embodiments described above, Figure 2D is a schematic diagram showing the continuous polymerase elongation of a cyclic construct and the formation of a double-length DNA template using YBEA200 according to a particular exemplary embodiment. As shown in Figure 2D (top panel), polymerase 210a continues along the parent template strand 207b to the 5' end of the parent template strand 207b, completing the synthesis of a new daughter strand 207a'. Similarly, as shown in Figure 2D (top panel), polymerase 210b continues along the parent template strand 207a to the 5' end of the parent template strand 207a, completing the synthesis of a new daughter strand 207b'. In particular, the new daughter strand 207a' contains a Y-branched element daughter strand 213b', which is complementary to the Y-branched element 213b that binds to the parent strand 207b. Similarly, the new daughter chain 207b' contains a Y-branching element daughter chain 213a', and the Y-branching element daughter chain 213a' is complementary to the Y-branching element 213a that is attached to the parent chain 207a.
[0093] In step 2d of Figure 2D, once the first and second polymerases 210a and 210b have completed the synthesis of daughter strands 207a' and 207b', respectively, polymerases 210a and 210b dissociate from the cyclic construct 209 to form a double-length DNA template 211 (bottom panel of Figure 2D). As shown in Figure 2D (bottom panel), the double-length DNA template 211 contains two copies of the target DNA template 207, namely the first and second copies 211a and 211b, respectively, each adjacent to a crosslinking region 208 (this crosslink originates from YBEA 200 and contains crosslinking daughter strand portions 208a and 208b).
[0094] As shown, each template copy 211a and 211b contains both the parent polynucleotide strand (shown in black) and the newly synthesized daughter polynucleotide strand (shown in gray). For example, template copy 211a contains the original (parent) template strand 207b and the newly synthesized daughter strand 207a'. On the opposite side of the crosslinking region 208, template copy 211b contains the original (parent) template strand 207a and the newly synthesized daughter strand 207b'. Furthermore, the double-length DNA 211 template contains a first terminal end 212a and a second terminal end 212b. The first terminal end 212a contains, for example, a portion 200b of the YBEA200 strand related to the 5' end YBEA200 of the nick site 201b (white-outlined black rectangle in terminal end 212a) and a copy thereof (white-outlined gray circle in terminal end 212a). Similarly, the second terminal end 212b of the double-length DNA template 211 contains a portion 200a of the YBEA200 strand associated with the 5' end YBEA200 of the nick site 201a (white-out black circle in 212b) and a copy thereof (white-out gray rectangle in terminal end 212b).
[0095] Similarly, as shown, template copy 211a also includes the Y-branching element 213b and its complementary sequence in the Y-branching element daughter strand 213b' at the terminal end 212a, while template copy 211b includes the Y-branching element 213a and its complementary sequence in the Y-branching element daughter strand 213a' at the terminal end 212b (Figure 2D, bottom panel). That is, the Y-branching elements 213a and 213b of YBEA200 are present at the terminal end of the double-length DNA template 211 when derived from YBEA200 as described herein. Thus, the combination of YBEA200 and the parent target DNA template 207 via steps 2a to 2d in Figures 2B to 2D yields a double-length DNA template 211 containing predetermined oligonucleotide sequences at each end (Figure 2D, bottom panel).
[0096] Furthermore, similar to the exemplary double-length DNA template 111 shown in Figure 1D (bottom panel), both strands of the double-length DNA template 211 in Figure 2D also contain a parent polynucleotide strand (black) covalently linked to a newly synthesized daughter copy (gray) of the target DNA template. For example, the parent strand 207a is covalently adjacent to the newly synthesized daughter strand 207a' in the 5'→3' direction via the strand of the crosslinking region 208 (i.e., the strand of the crosslinking region 208 containing strand portions 200a and 208a). Moreover, as described herein, for polymerase-mediated elongation of the cyclic construct, the nucleotide sequence of the parent strand 207a matches the nucleotide sequence of the new daughter strand 207a'. That is, the daughter strand 207a' is a sequence copy of the parent strand 207a of the target DNA template.
[0097] Similarly, the parent strand 207b is covalently adjacent to the new daughter strand 207b' in the same 5'→3' direction via the strand of the crosslinking region 208 (i.e., the strand of crosslinking region 208 containing strand portions 200b and 208b). Likewise, the nucleotide sequence of the parent strand 207b matches the nucleotide sequence of the copy 207b' of the new daughter strand. That is, the daughter strand 207b' is a sequence copy (i.e., daughter copy) of the parent strand 207b of the target DNA template. In this way, each strand of the double-stranded DNA contains both the parent polynucleotide sequence and the daughter polynucleotide sequence copy on each strand, in addition to the respective parent template strands and their complementary daughter strands of the two target DNA copies 211a and 211b (Figure 2D, lower panel).
[0098] As described herein, in certain exemplary embodiments, Y-branch elements 213a and 213b can be used to facilitate amplification, such as PCR amplification. Thus, Figure 2E shows a denatured (single-stranded) form of a double-length DNA template 212 (lower panel) of Figure 2D, which provides predetermined oligonucleotide primer-binding sequences when the original Y-branch elements 213a and 213b are replicated, according to a particular exemplary embodiment. For example, if Y-branch element 213a is replicated as 213a' (Figure 2D), then the 213a' daughter Y-branch element contains the Y-branch element 213a' or a primer-binding site within it. Similarly, if Y-branch element 213b is replicated as 213b' (Figure 2D), then the 213b' daughter Y-branch element contains the Y-branch element 213b' or a primer-binding site within it. As shown in Figure 2E, after the PCR denaturation step, primer 214 binds to the primer-binding site of Y-branch element 213a', and primer 215 binds to the primer-binding site of Y-branch element 213b'. Furthermore, similar to conventional PCR protocols and procedures, primers 214 and 215 provide 3' ends for polymerase elongation. In this way, a template strand can be replicated using the YBEA200 embodiment, and then primer binding sites can be provided for conventional PCR-based template amplification.
[0099] In certain exemplary embodiments, primers 214 and 215 have the same sequence and therefore bind to the same sequence within the respective primer binding sites of the Y branch elements 213a' and 213b'. That is, primers 214 and 215 are the same. Alternatively, in certain exemplary embodiments, primers 214 and 215 have different sequences and therefore bind to different sequences within the respective primer binding sites of the Y branch elements 213a' and 213b'. Thus, the YBEA and its associated Y branch elements offer a unique ability to customize the replication of the target DNA template strand for downstream use such as PCR amplification.
[0100] Epigenetic analysis using double-length DNA templates As described herein, in certain exemplary embodiments, the target DNA template includes a native target sequence. Therefore, in certain exemplary embodiments, the target DNA template can retain epigenetic information about the target sequence, such as the methylation pattern of the target sequence. Furthermore, since the parental polynucleotide strands of the target DNA template are retained in the diploid DNA template described herein, i.e., each strand of the diploid DNA template includes a parental polynucleotide strand from the target DNA template (referred to as the “parent copy” of the target sequence in the context of the diploid DNA template), the diploid DNA template also stores epigenetic information from the target sequence.
[0101] Furthermore, daughter copies of the target sequence can be synthesized under conditions that preserve the genetic information of the target sequence, as further described herein. Therefore, the presence of both parental and daughter copies of the target sequence on the same strand of the diploid DNA template is particularly beneficial for “intra-strand” comparisons for identifying epigenetic information. Also, since each parental copy of the target DNA template in the diploid DNA template hybridizes to complementary daughter sequences, in certain exemplary embodiments, this arrangement also enables “inter-strand” comparisons for identifying epigenetic information. Dual means of comparing parental and daughter sequences advantageously enhance the accuracy (and reliability) of the epigenetic information detected in the target sequence. These and other exemplary embodiments are illustrated and described with respect to Figures 3A–3G.
[0102] To facilitate intra-strand and / or inter-strand comparisons, in certain exemplary embodiments, the terminal adapters provided herein, such as the YBEA in Figure 2A, can be further modified to provide features that enable bioinformatics grouping of sequence reads. For example, terminal adapters such as the YBEA in Figure 2A can be modified to include a unique molecular identifier (UMI). In certain exemplary embodiments, the UMI can be included within the spacer region 202 of the YBEA, in which case the UMI sequence on the double-stranded double-length DNA template has a reverse complementary sequence. Such modified terminal adapters (and their use in a method for performing epigenetic analysis of a target DNA template strand, i.e., a target sequence) are shown in Figures 3A to 3F.
[0103] Referring to Figure 3A, a diagram is shown illustrating a Y-branched terminal adapter containing a UMI ("YB-UMI-EA") according to a particular exemplary embodiment. As shown, YB-UMI-EA300 has a typical double-stranded polynucleotide YBEA structure as shown in Figure 2A. This includes, for example, hybridized strands 300a and 300b, where "300a" refers to the entire 5'→3' strand of YBEA300 (with a nick site 301a within strand 300a) and "300b" refers to the entire 5'→3' strand hybridized to strand 300a (with a nick site 301b within strand 300b). Additionally, single-stranded Y-branched elements 313a and 313b are attached to the 5' ends of the first and second nick sites 301a and 301b of YB-UMI-EA, respectively. For example, the Y branch elements 313a and 313b may contain predetermined oligonucleotide sequences, and their design can be adjusted to achieve specific purposes, such as PCR amplification, as described above with respect to Figures 2A to 2E.
[0104] For example, as shown in Figure 3A, each Y branch element 313a and 313b may contain a predetermined oligonucleotide sequence that provides a complementary or hybridizable primer-binding site useful for PCR amplification of a new daughter strand. That is, each Y branch element 313a and 313b may contain, for example, 10 to 30 nucleotides, e.g., 15 to 25 nucleotides or 18 to 22 nucleotides, and its complementary strand contains a primer-binding site sequence. In certain exemplary embodiments, the Y branch elements 313a and 313b contain the same sequence (313a and 313b) as shown in Figure 3A. Alternatively, the Y branch elements 313a and 313b contain different sequences. In certain exemplary embodiments, the Y branch elements 313a and 313b are the same length, but in other exemplary embodiments, the Y branch elements 313a and 313b may be of different lengths.
[0105] YB-UMI-EA300 also includes terminal ends 303 and 304 adjacent to each nick site, each terminal 303 and 304 being adapted for efficient ligation to the end of the target DNA template. Preferably, as described herein, the nick sites 301a and 301b are separated by a double-stranded spacer region 302 so that the EA can accept binding of two polymerases for bidirectional elongation. That is, the nick sites 301a and 301b are far enough apart so that the binding of one polymerase does not sterically interfere with and / or substitute for the binding of the second polymerase. This is shown in Figure 3A, where the spacer region 302 linearly offsets the first nick site 301a from the second nick site 301b.
[0106] As also shown, for example, the UMI sequence 316 is positioned within the spacer region 302 of YB-UMI-EA300. UMI, also known as molecular barcode or random barcode, contains a short, random and / or predetermined nucleotide sequence incorporated into an oligonucleotide sequence. Typically, UMI is 5 to 20 nucleotides long, e.g., 8 to 16 nucleotides long. Needless to say, this length can vary depending on the application. For example, UMI can have a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. More conventionally, UMIs contain a sequence of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.
[0107] As shown in Figure 3A, since YB-UMI-EA300 is double-stranded, UMI sequence 316 has complementary UMI strand sequences 316a and 316b, with strand 316a exhibiting 5'→3' polarity and strand 316b having complementary 3'→5' polarity. Furthermore, since UMI strand sequences 316a and 316b are complementary sequences, the sequences of the different strands of the resulting double-length DNA template can be bioinformatically paired, as described herein, enabling the comparison of sequencing reads of the different strands.
[0108] Referring to Figure 3B, a schematic diagram is provided showing the circularization of a target DNA template using YB-UMI-EA300 according to a particular exemplary embodiment. As shown, YB-UMI-EA300, together with its respective first and second single-stranded Y branch elements 313a and 313b and double-stranded UMI sequence 316, is combined with a double-stranded target DNA template 307 to form a circular construct 309. That is, YB-UMI-EA300 is combined with the target DNA template 307, which has first and second terminal ends 305 and 306 and includes complementary strands 307a and 307b. In step 3a, for example, YB-UMI-EA300 is ligated to either end of the target DNA template 307, which includes polynucleotide strands 307a and 307b. For example, the terminal end 303 of YB-UMI-EA300 is ligated to the template end 306. Alternatively, in step 3a, although not shown for simplification, the other terminal end (i.e., 304) of YB-UMI-EA300 is ligated to the terminal end 305 of the target DNA template 307.
[0109] In step 3b of Figure 3B, the unligated (free) end of YB-UMI-EA300 is ligated to the other free end of the target DNA template 307 to form a circular construct 309 (or circular construct). For example, if the terminal end 303 of YB-UMI-EA300 is ligated to the template end 306 in step 3a, then in step 3b, the terminal end 304 of YB-UMI-EA300 is ligated to the template terminal end 305, thereby forming a circular construct 309 of the original (parent) target DNA template 307. Alternatively, if the terminal end 304 of YB-UMI-EA300 is ligated to the template terminal end 305 in step 3a, then in step 3b, the terminal end 303 is ligated to the template end 306, thereby forming a circular construct 309.
[0110] In either case, in step 3b of Figure 3B, the two endpoints 303 and 304 of YB-UMI-EA300 connect to the respective endpoints 305 and 306 of the template 307, thereby forming a YB-UMI-EA crosslinking region 308 between the endpoints of the template 307. That is, the entire YB-UMI-EA300 forms a crosslinking region 308 between the two endpoints 305 and 306 of the target (parent) template 307. This forms an annular structure 309 (or annular structure) containing YB-UMI-EA300 (as the crosslinking region 308) together with the complementary parent template chains 307a and 307b of the parent template 307. In this way, YB-UMI-EA300 (Figure 3A) acts as a crosslinking precursor for the crosslinking region 308. Furthermore, the respective first and second nick sites 301a and 301b remain within the cyclic construct 309 (as part of the crosslinking region 308), thereby providing two respective 3' ends available for polymerase binding and bidirectional extension, as described herein. The single-stranded Y-branch elements 313a and 313b are also present in the cyclic construct 309 together with the UMI sequence 316.
[0111] Referring to Figure 3C, a magnified view of a portion of target DNA template 307 is shown, illustrating an exemplary nucleic acid sequence according to a particular exemplary embodiment. As shown, the portion of target DNA template strand 307a contains endogenous methylated cytosine residues (i.e., 5-methylcytosine or "5mc") at nucleotide positions 3, 8, and 10 of the exemplary sequence, and an unmethylated cytosine residue (arrow) at position 5 (if strand 307a is read from left to right, i.e., in the 5'→3' direction of strand 307a). However, the cytosine residue at position 5 (see asterisk in strand 307a) is not protected (i.e., unmethylated). Furthermore, target DNA template strand 307b is complementary to target DNA template strand 307a and contains an endogenous methylated cytosine residue at position 9 (if read from left to right, i.e., in the 3'→5' direction of strand 307b). The cytosine residue at position 6 (note the asterisk is on strand 307b) is unprotected (i.e., unmethylated). In this example, the 5mC residue represents the epigenetic methylation pattern of the parent target DNA template target.
[0112] Continuing with exemplary embodiments of YB-UMI-EA300, Figure 3D is a schematic diagram showing the binding and bidirectional extension initiation of polymerases to a circular construct containing YB-UMI-EA300 according to a particular exemplary embodiment. As shown in Figure 3D (and also as in Figure 2C), when the first and second polymerases 310a and 310b combine with the DNA circular construct 309, they bind to the nick sites 301a and 301b, respectively, and proceed in opposite directions (as indicated by the arrows in the upper panel of Figure 3D). Polymerases 310a and 310b also replace the 5' ends of the parent strands 307a and 307b, respectively (including their associated and respective 313a and 313b Y branch elements). The UMI strand sequences 316a and 316b of UMI316 remain unchanged.
[0113] In step 3c (Figure 3D, bottom panel), polymerases 310a and 310b extend the cyclic construct 309 bidirectionally in opposite directions (see arrows). Specifically, polymerase 310a, as it proceeds, uses the parent strand 307b as a template to extend the 3' end of the nick site 301a, synthesizing a new daughter strand 307a' that is complementary to the sequence of the parent strand 307b (and therefore shares the same sequence as the parent strand 307a). The new daughter strand 307a' also contains the replicated ssDNA daughter strand crosslinking portion 308a as part of the crosslinking region 308. Furthermore, the Y branch element 313a remains at the 5' end of the substituted template strand 307a, while the UMI sequence 316 remains unchanged along with its strand sequences 316a and 316b.
[0114] Similarly, polymerase 310b uses the parent strand 307a as a template to extend the 3' end of the nick site 301b, while also substituting the 5' end of the nick site 301b (and its associated parent template strand 307b) (Figure 3D, bottom panel). Thus, in step 3c (bottom panel), polymerase 310b, as it proceeds, uses the parent strand 307a as a template to extend the 3' end of the nick site 301b to synthesize a new daughter strand 307b' that is complementary to the sequence of the parent strand 307a (and therefore shares the same sequence as the parent strand 307b). As shown, the Y branch element 313b also remains bound to the parent strand 307b, which has been substituted at its 5' end. The new daughter strand 307b' also contains the replicated ssDNA daughter strand crosslinking portion 308b as part of the crosslinking region 308. Similarly, UMI sequence 316, along with its chain sequences 316a and 316b, remains unchanged because it is not replicated during chain elongation.
[0115] Referring to Figure 3E, a schematic diagram is shown illustrating the continuous polymerase elongation of a circularized target DNA template and the formation of a double-length DNA template using an exemplary embodiment of YB-UMI-EA300 according to a particular exemplary embodiment. As shown in Figure 3E (top panel), polymerase 310a continues along the parent template strand 307b to the 5' end of the parent template strand 307b, completing the synthesis of a new daughter strand 307a'. Similarly, as shown in Figure 3E (top panel), polymerase 310b continues along the parent template strand 307a to the 5' end of the parent template strand 307a, completing the synthesis of a new daughter strand 307b'. In particular, the new daughter strand 307a' contains a Y-branched element daughter strand 313b', which is complementary to the Y-branched element 313b that binds to the parent strand 307b. Similarly, the new daughter strand 307b' contains the Y-branching element daughter strand 313a', which is complementary to the Y-branching element 313a that binds to the parent strand 307a. The UMI sequence 316 remains unchanged, along with its strand sequences 316a and 316b.
[0116] In step 3d of Figure 3E, once the first and second polymerases 310a and 310b have completed the synthesis of the daughter strands 307a' and 307b', respectively, polymerases 310a and 310b dissociate from the circular construct 309 to form a double-length DNA template 311 (Figure 3E, bottom panel). As shown in Figure 3E (bottom panel), the double-length DNA template 311 contains two copies of the original (parental) target DNA template 307, namely the first and second copies 311a and 311b, respectively, each adjacent to a crosslinking region 308. The crosslinking region 308 arises from YB-UMI-EA300 and contains the crosslinking daughter strand portions 308a and 308b together with the invariant UMI316 and its complementary strand sequences 316a and 316b.
[0117] As shown, each template copy 311a and 311b contains both the parent polynucleotide strand and the newly synthesized daughter polynucleotide strand. For example, template copy 311a contains the original (parent) template strand 307b and the newly synthesized daughter strand 307a' of the target DNA template 307 (Figures 3B and 3G). On the opposite side of the crosslinking region 308, template copy 311b contains the original (parent) template strand 307a and the newly synthesized daughter strand 307b' of the target DNA template 307 (Figures 3B and 3G). Furthermore, the double-length DNA template 311 contains the first terminal end 312a and the second terminal end 312b. The first terminal end 312a includes, for example, a portion 300b of the YB-UMI-EA300 strand (a white black rectangle in terminal end 312a) related to the 5' end of nick site 301b YB-UMI-EA300, and a copy thereof (a white gray circle in terminal end 312a). Similarly, the second terminal end 312b of the double-length DNA template 311 includes a portion 300a of the YB-UMI-EA300 strand (a white black circle in 312b) related to the 5' end of nick site 301a YB-UMI-EA300, and a copy thereof (a white gray rectangle in terminal end 312b).
[0118] As shown similarly, template copy 311a also includes the Y-branching element 313b and its complementary sequence in the Y-branching element daughter strand 313b' at the first terminal end 312a, while template copy 311b includes the Y-branching element 313a' and its complementary sequence in the Y-branching element daughter strand 313a at the second terminal end 312b (Figure 3E, bottom panel). In this way, the combination of YB-UMI-EA300 and target DNA template strand 307 via steps 3a to 3d in Figures 3B to 3D yields a double-length DNA template 311 containing predetermined oligonucleotide sequences at each end, along with UMI 316 (and its strand sequences 316a and 316b) (Figure 3E, bottom panel).
[0119] Similar to the double-length DNA templates 111 and 211 in Figures 1D and 2D, both strands of the double-length DNA template 311 shown in Figure 3D also contain a parent copy (black) ligated to a newly synthesized daughter copy (gray) of the target DNA template. For example, parent copy 307a is covalently adjacent to the newly synthesized daughter copy 307a' in the 5'→3' direction via the strand of the crosslinking region 308 (i.e., the strand of the crosslinking region 308 containing strand portions 300a and 308a). Furthermore, as described herein, for polymerase-mediated elongation of the cyclic construct, the nucleotide sequence of parent copy 307a matches the nucleotide sequence of the new daughter copy 307a'.
[0120] Similarly, parent copy 307b is covalently adjacent to the new daughter copy 307b' in the same 5'→3' direction via the strand of crosslinking region 308 (i.e., the strand of crosslinking region 308 containing strand portions 300b and 308b). Likewise, the nucleotide sequence of parent copy 307b matches the nucleotide sequence of the new daughter copy 307b'. In this way, each strand of the double-stranded DNA contains both the parent template DNA and the daughter copy DNA on each strand, in addition to the parent template DNA that has hybridized to the complementary daughter DNAs of the two target DNA copies (Figure 3E, bottom panel).
[0121] In certain exemplary embodiments, the double-length DNA template shown in Figure 3E (bottom panel) can be advantageously used to identify epigenetic information associated with the parent target DNA template 307. For example, during the extension steps of polymerases 310a and 310b shown in Figures 3E-3F, protected nucleotides, such as methylated cytosine nucleotide residues, can be used for daughter strand extension and synthesis. This incorporates the protected nucleotides, such as methylated cytosine nucleotide residues, into the newly synthesized daughter strands 307a' and 307b'. In such embodiments, the daughter strands having the protected cytosine residues preserve the genetic information of the target DNA template, for example, during a bisulfite treatment process that converts native cytosine to uracil. Subsequently, after the bisulfite conversion reaction and DNA sequencing, bioinformatics analysis of the sequence information can be performed, as further described herein, to identify the methylated cytosine residues in the original (parent) target DNA template.
[0122] In certain exemplary embodiments, the identification of methylated cytosine residues in the original (parent) target DNA template 307 provides epigenetic information related to the original (parent) target DNA template 307. This is shown, for example, in Figure 3F, which shows a schematic diagram of exemplary bisulfite conversion of a double-length DNA template 311 and its subsequent PCR amplification product by using a Y-branched end adapter having UMI (i.e., YB-UMI-EA300) in Figure 3A, according to a particular exemplary embodiment.
[0123] Referring to Figure 3F (and prior to step 3e), an exemplary cleaved portion of the double-length DNA 311 of Figure 3E is shown, and the cleaved portion contains a double copy of the exemplary template sequence portion shown in Figure 3C. That is, Figure 3F shows only a portion of the sequence of the original target DNA template 311 (simply for the sake of simplifying the figure), and the portion shown contains two copies of the exemplary sequence (311a and 311b) according to Figure 3C. As shown, each strand contains, in its intra-strand configuration, a parent copy (black) and a daughter copy (gray), for example, these copies resulting from the formation of the double-length DNA template as described herein.
[0124] As shown in Figure 3F, each template copy portion 311a and 311b contains 10 exemplary nucleotide pairs corresponding to the nucleotide pairs in Figure 3C, and this nucleotide sequence is mirrored on both sides of the diploid DNA template as a result of forming a diploid DNA template. For example, the sequence of parent strand 307a (of template copy 311b) corresponds to the same sequence on daughter strand 307a' (of template copy 311a), and both sequences 307a and 307a' are related to UMI strand sequence 316a. For example, reading the polynucleotide sequence related to UMI strand 316a from left to right (i.e., 5'→3'), the exemplary parent copy (black) is TACACGACGC (SEQ ID NO: 1), while the daughter copy (gray) is the same polynucleotide sequence, namely TACACGACGC. Therefore, the exemplary 5'→3' sequence related to UMI 316a is TACACGACGC--UMI--TACACGACGC.
[0125] Similarly, considering complementary base pairing of the strands, the sequence of the parent strand 307b (of template copy 311a) corresponds to the same sequence on the daughter strand 307b' (of template copy 311b), but both sequences 307b and 307b' are related to UMI strand sequence 316b. That is, reading the sequence related to UMI strand 316b from left to right (i.e., 3'→5'), the exemplary daughter strand sequence (gray) is ATGTGCTGCG (SEQ ID NO: 2), and the parent sequence (black) is also ATGTGCTGCG. In other words, the exemplary 5'→3' sequence related to UMI316b is ATGTGCTGCG--UMI-ATGTGCTGCG. Thus, each UMI strand sequence 316a and 316b of UMI316 is related to the parent strand (black) and a portion of the new daughter strand (gray) (Figure 3F).
[0126] As also shown in the epigenetic evaluation of this exemplary target DNA template strand, prior to step 3e, the parent strand 307a of template copy 311b contains endogenously methylated (protected) cytosine residues at positions 3, 8, and 10, and an unmethylated cytosine residue (arrow) at position 5 (from left to right, i.e., 5'→3', as shown in Figure 3C). Furthermore, the parent strand 307b of template copy 311a contains an endogenously methylated (protected) cytosine residue at position 9, and an unmethylated residue at position 6 (when read from 3'→5', as shown in Figure 3C). However, each daughter strand 307a' and 307b' contains only methylated cytosine residues as a result of polymerase elongation, and only methylated cytosine residues are obtained from the elongation reaction. That is, neither daughter strand 307a' nor 307b' contains an unmethylated cytosine residue. Thus, the protected daughter strand 307a' or 307b' (gray) preserves the genetic information of the target DNA template, while the natural parental target DNA template strands 307a and 307b (black) preserve the epigenetic information of the target DNA template (and consequently, the target sequence).
[0127] In step 3e of Figure 3F, a double-length DNA template is subjected to bisulfite conversion using a conventional method. For example, bisulfite conversion is a method that uses bisulfite to determine the methylation pattern of DNA, such as methylation of a target DNA template. For example, DNA methylation is an endogenous biochemical process that involves the addition of methyl groups to cytosine or adenine DNA nucleotides. For example, DNA methylation stably alters gene expression in cells when cells divide and differentiate from embryonic stem cells into specific tissues. In bisulfite conversion (also known as bisulfite sequencing), the target nucleic acid is first treated with a bisulfite reagent that specifically converts unmethylated cytosine residues to uracil residues (i.e., C→U conversion), but does not affect methylated cytosine residues (i.e., methylated cytosine residues are "protected" from C→U conversion). Subsequently, the converted uracil residues are replaced with thymine residues (i.e., U→T substitution) by PCR reactions with natural adenine (A), cytosine (C), guanine (G), and thymine (T) nucleotides. In this way, unmethylated (i.e., "unprotected") cytosine residues are converted to thymine via the intermediate uracil (i.e., C→U→T).
[0128] Therefore, as shown in step 3e of Figure 3F, when the denatured double-length DNA template strand is subjected to the bisulfite conversion reaction, the unmethylated cytosine residue at position 5 of parent strand 307a is converted to a uracil residue (SEQ ID NO: 3, shown in bold and underlined), i.e., a 5C→5U conversion occurs. Similarly, the unmethylated cytosine residue at position 6 of parent strand 307b is converted to uracil (SEQ ID NO: 4, shown in bold and underlined), i.e., a 6C→6U conversion occurs. However, the bisulfite reaction does not affect either parent strand 307a or the methylated (protected) cytosine residues in 307a; that is, these cytosine residues remain cytosine residues. This includes the methylated cytosine residues of daughter strands 307a' and 307b' that were contained within daughter strands 307a' and 307b' via polymerase elongation using methylated cytosine nucleotides, as described herein.
[0129] Following the bisulfite conversion reaction in step 3e, in step 3f of Figure 3F, the bisulfite-converted strand of the product of the denatured double-length DNA template 311 is subjected to PCR amplification and sequencing using conventional methods. For example, the strand of the denatured double-length DNA template 311 can be amplified as described herein using PCR primers targeting the Y branch elements 313a' and 313b'. As shown in step 3f and as can be determined by conventional DNA sequencing, the PCR product (shown in a denatured state for illustrative purposes) yields separate strands, each associated with either UMI strand sequence 316a or 316b. In the PCR product, the uracil residues generated by the bisulfite conversion of unmethylated (unprotected) cytosine residues are substituted with thymine. For example, the uracil residue at position 5 of the parent strand 307a is substituted with a thymine (T) residue during the PCR reaction (see arrow), i.e., a 5U → 5T substitution. Furthermore, the 5U→5T substitution in parent strand 307a corresponds to UMI strand sequence 316a. Similarly, the uracil residue at position 6 of parent strand 307b is substituted with a thymine (T) residue during the PCR reaction (see arrow), i.e., a 6U→6T substitution. Furthermore, the 6U→6T substitution in parent strand 307b corresponds to UMI strand sequence 316a. Thus, each strand of UMI (316a and 316b) in this example is localized by a strand-specific nucleotide conversion (C→U→T) of the original (parent) DNA template.
[0130] In step 3g, after the PCR reaction in step 3f, the PCR product is sequenced, and the resulting sequenced reads identify the methylation pattern of the original parent copy of the DNA target sequence by intra-strand comparison of the parent and daughter sequences. That is, the daughter strand copy has protected cytosine residues, is resistant to bisulfite conversion, and therefore preserves the parent template gene sequence. Thus, at each position in the original parent strand sequence containing native (unmethylated) cytosine, the whole-strand sequence read shows a discrepancy between the parent and daughter sequences, and in contrast, at each position in the parent strand sequence containing methylated cytosine, the whole-strand sequence read shows a match between the parent and daughter sequences.
[0131] Additionally or alternatively, the parental sequence methylation pattern can also be identified and / or confirmed by comparing complementary parental and daughter strand sequences (i.e., inter-strand comparison). That is, comparing parental and daughter strand sequences from different strands of a diploid DNA template (made possible by bioinformatics grouping of UMI read sequences) reveals mismatches between pair bases at the positions of native cytosines in the parental sequence, while the positions of methylated cytosines show normal complementarity with respect to the daughter sequence. Such intra-strand and inter-strand comparisons and analyses are shown in Figure 3G, and both methods can be used independently or in combination to evaluate epigenetic information related to the original target template sequence.
[0132] (Before step 3h) Referring to Figure 3G, a schematic diagram is shown illustrating a comparison of both intra- and inter-strand portions of a double-length DNA template to confirm epigenetic information associated with the original target DNA template, according to a specific exemplary embodiment. In this exemplary schematic diagram, the same exemplary sequence is carried over from Figure 3F, and the strands are shown in an aligned double-length DNA double-template configuration for illustrative purposes only. As shown, in intra-strand analysis, for example, (reading the strand sequence related to UMI sequence 316a in the 5'→3' direction (i.e., from left to right, from the 5' end of strand fragment 307a to the 3' end of strand fragment 307a') an intra-strand TC mismatch is identified at the 5th nucleotide (see arrow related to UMI sequence 316a). That is, the sequence of strand fragment 307a (black) contains a thymine (T) residue, while the sequence of strand fragment 307a' (gray) contains a cytosine (C) residue. Importantly, since the daughter strand is synthesized using a protected cytosine analog, the TC mismatch within the strand identifies strand fragment 307a' as the daughter copy and fragment 307a as the parent copy of the target DNA template.
[0133] Similarly, in the chain associated with UMI sequence 316b, when the sequence of the chain associated with UMI sequence 316b is read from 3' to 5' (i.e., from left to right, from the 3' end of chain fragment 307b' to the 5' end of chain fragment 307b), the TC mismatch within the chain is identified at the 6th position of the nucleotide (see the arrow associated with UMI sequence 316b). That is, the sequence of chain fragment 307b' (black) contains a thymine residue, while the sequence of chain fragment 307b (gray) contains a cytosine residue. Also, similar to the TC mismatch associated with UMI sequence 316a described earlier, the presence of a cytosine residue at the 6th position of chain fragment 307b' identifies this chain fragment as a daughter chain (gray), and chain fragment 307b (black) is the parent-derived chain. Furthermore, the presence of a substituted thymine residue at the 6th position of chain 307b indicates that this thymine nucleotide was an unprotected cytosine residue in the original target sequence, as fully described below.
[0134] Additionally or alternatively, prior to step 3h, in certain exemplary embodiments, inter-chain mismatch analysis may be used to identify, evaluate, and / or confirm epigenetic information related to the original target sequence. As shown, for example, the inter-chain alignment of the sequence of exemplary parent strand 307a (black) and the sequence of daughter strand 307b' (gray) reveals a TG mismatch at position 5 of the 307a / 307b' aligned sequence. Based on the presence of this mismatch, it can also be determined that the sequence of strand 307a corresponds to the parent target sequence. This is because only unmethylated (unprotected) cytosine residues undergo C→U→T bisulfite / PCR conversion, while daughter strand elongation with methylated (protected) cytosine residues incorporates only protected cytosine residues into the daughter strand. Therefore, during bisulfite / PCR conversion, only unprotected cytosine residues in the parent strand are converted to thymine residues (i.e., cytosine residues in the daughter strand are not converted). Once strand fragment 307a is identified as a parent-derived copy, when read from left to right (i.e., 5'→3'), this parent-derived copy can be identified as being related to the 5' end of UMI sequence 316a, and daughter strand fragment 307a' is located downstream of the 3' end of UMI 316a (as shown).
[0135] Similarly, the alignment of the sequence of the exemplary parental fragment 307b (black) with the sequence of the exemplary daughter fragment 307a' (gray) reveals a TG mismatch at position 6 of the 307b / 307a' aligned sequence. Thus, the presence of a thymine residue in the TG mismatch identifies fragment 307b as the parental derived strand, and fragment 307a' as the complementary daughter strand. Therefore, once strand 307b is identified as the parental derived copy, when read from left to right (i.e., 3'→5'), this parental derived copy can be identified as relating to the 3' end of UMI fragment 316a, and as shown in the figure, daughter fragment 307b' is located upstream of the 5' end of UMI 316b. In certain exemplary embodiments, such inter-chain and intra-chain analyses can be used to identify and confirm methylation patterns by UMI-based read grouping across multiple sequence reads. This is particularly useful, for example, when a large region of the target sequence (such as one conserved in the target DNA template) contains methylated cytosine residues.
[0136] In step 3h of Figure 3G, protected (methylated) cytosine residues associated with the original (parent) target DNA template 307 can be identified based on the intra-strand and / or inter-strand analysis described herein. This provides epigenetic information about the original (parent) DNA template 307. For example, as described above, the C→U→T bisulfite / PCR conversion occurs only at unprotected (native) cytosine residues. Therefore, the presence of any cytosine residue in the strand identified as corresponding to the original (parent) target DNA template strand (i.e., strand fragments 307a and 307b in the example above, shown in black) can be identified as a previously protected (methylated) cytosine residue. This is shown, for example, after step 3h, where the arrows indicate the identification of protected cytosine residues in the original parent (target) template (e.g., template 307). As shown for strand fragment 307a, for example, previously protected cytosine residues are present at positions 3, 8 and 10 (see arrows, read the strand from left to right). Similarly, the cytosine residue at position 9 of chain fragment 307b can also be identified as previously protected (see arrow, read from left to right).
[0137] Finally, in step 3i of Figure 3G, in certain exemplary embodiments, the strand fragments identified as corresponding to the parent strands (e.g., 307a and 307b) of the original (parent) target DNA template 311 (and their associated methylation patterns) can be aligned to reveal the epigenetic pattern associated with the original (parent) target DNA template 307. That is, by using the methods described in Figures 3A to 3G, epigenetic information associated with the original (parent) target DNA template 307 can be obtained. As shown, for example, the presence of these cytosine residues in the strand corresponding to the parent template strand fragment 307a was necessarily protected (methylated) in the original (parent) template strand 307a (and therefore not converted by bisulfite conversion), so the aligned sequences of strand fragments 307a and 307b show methylation at positions 3, 8, and 10 of strand fragment 307a. Furthermore, using the TC mismatch and / or TG mismatch to identify the presence of a substituted thymine residue (for bisulfite conversion), the C residue at position 5 can be assigned in place of the thymine residue in chain fragment 307a (see the asterisk on the cytosine residue at position 5 in chain fragment 307a). Similarly, by this same or similar rationale, chain fragment 307b shows methylation at position 9 (read from left to right) and an unprotected cytosine at position 6 (asterisk) (when the sequence of chain fragment 307b is read from left to right, i.e., 3'→5'). And in particular, this identified epigenetic methylation pattern corresponds to the exemplary methylation pattern shown as an example in Figure 3C (see the inset in Figure 3C).
[0138] Therefore, by incorporating methylated cytosine nucleotides during polymerase elongation of a circularized target DNA template, and then subjecting the diploid DNA template to bisulfite / PCR conversion, epigenetic information related to the original target DNA template can be easily obtained by intra- and inter-strand parent / daughter sequence comparisons.
[0139] In consideration of the disclosure herein, epigenetic detection methods can be incorporated into the methods of the present invention. For example, enzymatic conversion of the modified base of interest, or any other biochemical or chemical reaction that specifically converts the modified nucleic acid base of interest compared to the natural base (or, alternatively, converts the unmodified nucleic acid base of interest, as discussed herein in relation to the bisulfite conversion of natural cytosine to uracil). Specific exemplary methods of enzymatic conversion of the modified base of interest are disclosed, for example, in the concurrently pending U.S. Provisional Patent Applications No. 63 / 380439 and No. 63 / 147959 of the present applicant, which are incorporated herein by reference in their entirety.
[0140] Double-length DNA template for use in PCR multiplexing In certain exemplary embodiments, the terminal adapters described herein may be modified to include, additionally or alternatively, one or more sequence indices (SIDs). That is, the terminal adapters, for example, those in Figures 1A, 2A, and / or 3A, may be modified to include, for example, one or more specific nucleotide sequences in which the original source identifies the target DNA template (and thus the target sequence) when multiple target DNA templates / target sequences are being analyzed. Such SIDs are very useful, for example, in applications such as DNA multiplexing, i.e., processing multiple different samples simultaneously via PCR. Thus, SIDs are also called sample identifiers.
[0141] In certain exemplary embodiments, for example, the same or different SIDs may be included adjacent to the Y-branched sequence element described herein within a sequence contiguous with the 3' end of the Y-branched sequence element described herein. Additionally or alternatively, one or more SIDs may be included on the same strand having the Y-branched sequence element, and intervening non-SID nucleotides or sequences of nucleotides may separate the SID from the Y-branched element. Nevertheless, each SID may be specific to a target sequence having a sequence complementary to the SID found on the opposing (complementary) strand of the terminal adapter. Subsequently, double-stranded DNA molecules having different SIDs can be processed in a single PCR reaction, for example, this SID allows for the differentiation of different DNA samples after sequencing. Furthermore, since multiple copies of the SID appear on a single replicated PCR product strand, the SID can be identified with high bioinformatics accuracy. This reduces or eliminates the need for additional error correction. In such exemplary embodiments, the SID can also be used as a landmark on a given strand, allowing for additional analysis. Furthermore, such embodiments including SIDs may also include UMIs, as shown in Figures 3A-3E.
[0142] Referring to Figure 4A, a diagram of a Y-branched end adapter containing two SID sequences and (optionally) UMI is shown according to a particular exemplary embodiment. As shown, YB-UMI / SID-EA400 has a common polynucleotide double-stranded YBEA structure as shown in Figure 3A, and includes UMI416 having UMI strands 416a and 416b. This includes Y-branching elements 413a and 413b that bind to the 5' ends of the first and second nick sites 401a and 401b of YB-UMI / SID-EA, respectively. For example, Y-branching elements 413a and 413b may contain a predetermined oligonucleotide sequence, and its design can be tailored to achieve a specific purpose, such as PCR amplification of a double-length DNA template, as described above with respect to Figures 2A-2E and 3A-3E.
[0143] As shown in Figure 4A, each Y-branch element 413a and 413b may contain a predetermined oligonucleotide sequence that provides a complementary primer-binding site useful for PCR amplification. That is, each Y-branch element 413a and 413b may contain, for example, 10 to 30 nucleotides, for example, 15 to 25 nucleotides or 18 to 22 nucleotides, and its complementary strand contains the primer-binding site sequence. In certain exemplary embodiments, the Y-branch elements 413a and 413b contain the same sequence, as shown in Figure 4A (413a and 413b). Alternatively, the Y-branch elements 413a and 413b contain different sequences. In certain exemplary embodiments, the Y-branch elements 413a and 413b are the same length, but in other exemplary embodiments, the Y-branch elements 413a and 413b may be of different lengths.
[0144] The YB-UMI / SID-EA400 also includes terminal ends 403 and 404 adjacent to each nick site, with each terminal 403 and 404 adapted for efficient ligation to the end of the target DNA template. Preferably, as described herein, the nick sites 401a and 401b are separated by a double-stranded spacer region 402 so that the EA can accept binding of two polymerases for bidirectional elongation. That is, the nick sites 401a and 401b are far enough apart so that the binding of one polymerase does not sterically interfere with and / or substitute for the binding of the second polymerase. This is shown, for example, in Figure 4A, where the spacer region 402 linearly offsets the first nick site 401a from the second nick site 401b.
[0145] As also shown, for example, the UMI sequence 416 is positioned within the spacer region 402 of YB-UMI / SID-EA400. Typically, the UMI is 5 to 20 nucleotides long, for example, 8 to 16 nucleotides long. Needless to say, this length can vary depending on the application. For example, the UMI can have a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. More conventionally, UMI contains a sequence of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides. As shown in Figure 4A, since YB-UMI / SID-EA400 is double-stranded, UMI sequence 416 has complementary UMI strand sequences 416a and 416b, with strand 416a exhibiting 5'→3' polarity and strand 416b having complementary 3'→5' polarity.
[0146] In addition to UMI416, which is shown in YB-UMI / SID-EA400 but can be optionally included, YB-UMI / SID-EA400 includes diagonally positioned SID417a (gray circle with black crossing lines) and SID418a (direction, shaded box), each shown adjacent to and connected to Y branch elements 413a and 416b (solid black circle). That is, in the example shown in Figure 4A, the sequences of each SID417a and 418a arise in series with the 5'→3' polynucleotide sequences of Y branch elements 413a and 413b, respectively. Also shown, each of SID417a and 418a has its respective complementary strand, namely SID complementary strands 417b (shaded box) and 418b (gray filled circle).
[0147] Traditionally, a SID includes short, random nucleotide sequences and / or predetermined nucleotide sequences that can be incorporated into a polynucleotide sequence. Typically, a SID is 5 to 20 nucleotides long, for example, 8 to 16 nucleotides long. Needless to say, this length can vary depending on the application. For example, a SID can have a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. More conventionally, a SID includes a sequence of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.
[0148] While the YB-UMI / SID-EA400 shows SIDs 417a and 418a located adjacent to and connected to Y-branch elements 413a and 413b, respectively, it should be understood that one or more SIDs can be placed anywhere within the YB-UMI / SID-EA400 to facilitate sample differentiation. For example, one or more SIDs may be located in continuity with the UMI chain 416a or 416b, for example, on the 5' side of UMI chain 416a or the 5' side of UMI chain 416b. In other exemplary embodiments, the SIDs may be contained within and / or included as part of UMI 416. Additionally or alternatively, one or more SIDs may be placed at any end of the YB-UMI / SID-EA400. For example, SID 417a may be located at the 3' end portion of terminal 404, while SID 418a may be located at the 3' end portion of terminal 403. Therefore, the SIDs described herein can be placed in or within multiple different locations on the YB-UMI / SID-EA400, insofar as the SIDs enable the differentiation of samples as described herein.
[0149] Referring to Figure 4B, a double-length DNA template resulting from the use of YB-UMI / SID-EA400 in Figure 4A is shown according to a particular exemplary embodiment. That is, YB-UMI / SID-EA400 is ligated to both ends of a target DNA template, thereby forming a circular end adapter / target DNA template as shown in Figures 1B, 2B, and 3B. In other words, a circular construct containing YB-UMI / SID-EA400 can be formed using the same or similar steps shown in Figures 1B, 2B, and 3B, where YB-UMI / SID-EA400 forms a crosslink that covalently connects both ends of the target DNA template. Subsequently, using the same or similar steps as described in Figures 1C-1D, 2C-2D, and 3D-3E, the 3' ends of the nick sites 401a and 401b of YB-UMI / SID-EA400 can be joined and extended, for example, in opposite directions, using the first and second polymerases to form daughter strand copies as described herein. For example, as described in step 1d (Figure 1D), step 2d (Figure 2D), and step 3d (Figure 3E), once the polymerases have completed their extension, the polymerases dissociate from the cyclic construct to obtain the double-length DNA template shown in Figure 4B.
[0150] As shown in Figure 4B, the double-length DNA template 411 contains two copies of the original (parental) DNA template 407, namely the first and second copies 411a and 411b, each adjacent to a crosslinking region 408. The crosslinking region 408, derived from YB-UMI / SID-EA400, contains the daughter strand SID sequence 417a' and the complementary sequence 417b, and the daughter strand (gray) is formed via polymerase-mediated elongation of the 3' end of the nick site 401a. The crosslinking region 408 also contains the daughter strand SID sequence 418a' and the complementary sequence 418b, and the daughter strand (gray) is formed via polymerase-mediated elongation of the 3' end of the nick site 401b. UMI416 contains the complementary UMI strand sequences 416a and 416b.
[0151] As shown, each DNA template copy 411a and 411b contains both the parent polynucleotide strand (black) and the newly synthesized daughter polynucleotide strand (gray). For example, template copy 411a contains the original (parent) template strand 407b and the newly synthesized daughter strand 407a'. On the opposite side of the crosslinking region 408, template copy 411b contains the original (parent) template strand 407a (dashed line and black) and the newly synthesized daughter strand 307b' (gray). Furthermore, the double-length DNA 411 template contains the first terminal end 412a and the second terminal end 412b. The first terminal end 412a includes, for example, a portion 400b of the YB-UMI / SID-EA400 strand (a white black rectangle in terminal end 412a) and a copy thereof (a white gray circle in terminal end 412a) related to the 5' end of the nick site 401b YB-UMI / SID-EA400. Similarly, the second terminal end 412b of the double-length DNA template 411 includes a portion 400a of the YB-UMI / SID-EA400 strand 400 (a white black circle in 412b) and a copy thereof (a white gray rectangle in terminal end 412b).
[0152] Similarly, as shown, template copy 411a also contains, at its terminal end 412a, SID 418a and its complementary daughter SID copy 418b, along with its complementary sequence in the Y-branch element 413b and the Y-branch element daughter strand 413b'. Likewise, at the other end of the double-length DNA template (i.e., terminal end 412b), SID 417a and its complementary daughter SID copy 417b are shown, along with its complementary sequence in the Y-branch element 413a and the Y-branch element daughter strand 413a'.
[0153] In this way, the combination of YB-UMI / SID-EA400 and the target DNA template strand results in a double-length DNA template 411 containing a predetermined oligonucleotide sequence at each end (i.e., Y branch elements 413a' and 413b' and their respective complementary copies 413a and 413b), a SID and its complementary copy at each end (i.e., SID 418a and 417a and their respective complementary copies 418b' and 417b'), and UMI 416 (and its strand sequences 416a and 416b). Figure 4B shows the positions of the SID sequences 418a, 418b', 417a, and 417b' consecutive to the Y branch elements 413b, 413b', 413a, and 413a', but it should be understood that the SID can be located at different positions within the double-length DNA template. For example, if there is no Y-branch element on the terminal adapter containing the SID, the resulting double-length DNA template can contain the SID at the first and second terminal ends 412a and 412b without the Y-branch element.
[0154] The strand-specific information encoded in each strand of such a double-length DNA template 411 allows for easy identification and differentiation of individual strands of the double-length DNA template 411 in a multiplex PCR reaction. In fact, since multiple copies of the SID appear on a single replicated PCR product strand, the SID of YB-UMI / SID-EA400 and the resulting double-length DNA template 411 can be determined with high accuracy bioinformatics, thereby reducing or eliminating the need for additional error correction. The SID can also be used as a landmark for a given strand, enabling further analysis.
[0155] Elongation and formation of asymmetric DNA templates In certain exemplary embodiments, methods for producing asymmetric DNA template copies and asymmetric DNA templates are provided. That is, the methods and compositions provided herein can be used to generate asymmetric DNA templates in which only one strand of a target DNA template is replicated. Thus, an asymmetric DNA template is an asymmetric DNA template in that only one strand of the parent template is replicated. Such asymmetric DNA templates are used in sequence preparation workflows that require a single-stranded DNA molecule as a target template, for example, the “Sequencing by Expansion” method developed by the inventors, which is incorporated herein in its entirety by reference (see, for example, U.S. Patent Application Publication No. 20220042075).
[0156] Referring to Figure 5A, a modified Y-branched end adapter according to Figure 2A is shown according to a particular exemplary embodiment, however, it is modified to accept only single polymerase binding and unidirectional extension. In other words, a modified Y-branched end adapter (or "modified YBEA") accepts only unidirectional extension of the target DNA template (when ligated to the target DNA template). For example, modified YBEA500 contains hybridized strands 500a (circular) and 500b (rectangular), thus forming a polynucleotide double helix. Modified YBEA500 also has a general EA structure like that in Figure 2A, in that a first Y-branching element 513a and a second Y-branching element 513b are added to the 5' ends of the first and second nick sites 501a and 501b, respectively. Furthermore, the modified YBEA500 includes terminal ends 503 and 504 adjacent to each nick site 501a and 501b, each terminal end 503 and 504 being suitable for efficient ligation to the ends of the target DNA template as described herein. That is, the terminal ends are ligable to the target DNA template.
[0157] Furthermore, unlike YBEA200 in Figure 2A, one 3' end of the nick site (providing a polymerase elongation site as described herein) is blocked or otherwise modified to prevent polymerase binding and / or elongation in the modified YBEA. In particular, any method known in the art can be used to modify the 3' end to prevent polymerase binding and / or elongation. For example, the 3' end can be phosphorylated. As shown in the exemplary modified YBEA500 in Figure 5A, the 3' end associated with nick site 501a includes a phosphorylated 3' end, which in turn prevents polymerase elongation of the 3' nick site 501a as described herein. Also, although not shown for simplicity, in certain exemplary embodiments, the modified YBEA500 may include UMI and / or one or more SIDs as described herein.
[0158] In certain exemplary embodiments, a modified YBEA (having a single extensible nick site) can be combined with a target DNA template to form a circular construct. Specifically, the modified YBEA can be ligated to both ends of the target DNA template, as shown in Figures 2B and 3B. Following ligation, a crosslink is formed between the two ends of the target DNA template, with the modified YBEA 500 acting as the crosslink. Once the modified YBEA forms a crosslink connecting the terminal ends of the target DNA template, a circular construct is formed as shown in Figures 2B (steps 2a and 2b) and 3B (steps 3a and 3b).
[0159] Referring to Figure 5B, a schematic diagram is shown illustrating the binding of polymerase and the initiation of unidirectional elongation of the circular construct according to a particular exemplary embodiment. As shown, the circular construct 509 includes the parent strands (black) 507a (dashed line) and 507b (solid line) of the target DNA template, along with Y branch elements 513a and 513b. Furthermore, the nick site 501a contains a phosphorylated 3' end, which consequently prevents polymerase elongation of the 3' end at the nick site 501a. However, when polymerase 510b is combined with the DNA circular construct 509, it binds to the nick site 501b (which in this example lacks 3' modification) and proceeds in the opposite direction to the nick site 501a (as indicated by the arrow). Polymerase 510b also replaces the 5' end of the parent template strand 507b and its associated Y branch element 513b. However, if there is no polymerase binding / extension at the nick site 501a, the parent strand 507a remains bound to its complementary strand 507b at the nick site, i.e., there is no substitution of the 5' end of strand 507a, as in the case of bidirectional targeted DNA template extension as described herein.
[0160] In step 5a of Figure 5B, polymerase 510b continues to extend the cyclic construct 509 in one direction (see arrow). That is, polymerase 510b uses the parent chain 507a as a template to extend the 3' end of the nicked site 501b while simultaneously substituting the 5' end of the nicked site 501b (and its associated parent template chain 507b) (Figure 5B, bottom panel). Thus, in step 5a (bottom panel), polymerase 510b, as it proceeds, uses the parent chain 507a as a template to extend the 3' end of the nicked site 501b to synthesize a new daughter chain 507b' (gray) that is complementary to the sequence of the parent chain 507a (and therefore shares the same sequence as the parent chain 507b). As shown, the Y-branch element 513b also remains bound to the parent chain 507b, which has been substituted at its 5' end. The new daughter strand 507b' also contains the replicated ssDNA daughter strand crosslinking portion 508b as part of the crosslink 508. However, at its phosphorylated (and therefore blocked) 3' end, the nick site 501a remains and is not extended.
[0161] Continuing with the exemplary embodiments described above, Figure 5C is a schematic diagram showing the continuous polymerase elongation and asymmetric template formation of a cyclic construct using modified YBEA500 according to a particular exemplary embodiment. As shown, polymerase 510b continues along the parent template chain 507a to the 5' end of the parent template chain 507a, completing the synthesis of a new daughter chain 507b'. In particular, the new daughter chain 507b' contains a Y-branched element daughter chain 513a', which is complementary to the Y-branched element 513a that binds to the parent chain 507a. As shown in Figure 5C, in this exemplary embodiment, the parent template chain 507b is completely replaced, but this chain is not replicated due to the blocked 3' end associated with 501a (see Figure 5B).
[0162] In step 5b of Figure 5C, once polymerase 510b completes the synthesis of the daughter strand 507b', polymerase 510b dissociates the DNA complex to form an asymmetric DNA template 511 (bottom panel). As shown in Figure 5D (bottom panel), the asymmetric DNA template 511 contains parent template strands 507a and 507b, each adjacent to a crosslink 508 (the crosslink arises from modified YBEA 500 and contains the crosslinked daughter strand portion 508b). The symmetric DNA template 511 also contains a first terminal end 512a and a second terminal end 512b. However, with unidirectional extension by polymerase 510b, the asymmetric DNA template 511 contains only a single daughter strand, i.e., strand 507b'. As shown, the asymmetric portion of the asymmetric DNA template 511 contains strand 507b', and the Y branch element 513b is located at the 5' end (terminal end 512a) of the 507b' strand. Furthermore, the bridging portion 508 has a newly synthesized daughter bridging portion 508b and is covalently connected adjacent to the parent template chain 507b which has a newly synthesized daughter chain 507b'.
[0163] As also shown, template copy 511b includes a daughter strand 507b' as a complementary strand to the parent template strand 507a. The parent strand 507a also includes a Y branch element 513a at its 5' end, while the new daughter strand 507b includes a daughter Y branch element 513a' at its 3' end. In this way, the modified YBEA 500 and the parent target DNA template are combined to form a circular construct, which is then subjected to polymerase-mediated extension as described in steps 5a and 5b of Figures 5B and 5C to form an asymmetric DNA template.
[0164] Multi-precision mold stretching In certain exemplary embodiments, the methods and compositions described herein can be repeated any number of times (starting from a first diploid DNA template) to form a polyploid DNA template. For example, after forming a diploid DNA template according to the methods and compositions described herein, both ends of the diploid DNA template can be ligated to a second end adapter (EA), for example, a second EA having the features of the EA in Figure 1A. - This forms a circular construct containing the diploid DNA template and the ligated EA. - The circular construct can then be replicated bidirectionally as described herein to form a quadruploid DNA template or a "double-double" template, i.e., a DNA molecule containing replicated copies of the original diploid DNA template. The quadruploid DNA template includes, for example, two parent target DNA template strands arising from the original diploid DNA template, along with complementary daughter strands, as described herein. However, the quadruploid DNA template also includes replication of these strands, and thus contains four copies of the target DNA template. The formation of such a quadruploid DNA template is shown, for example, in Figure 6.
[0165] As shown, the target DNA template in the example of Figure 6 is a double-stranded DNA template containing two copies of the original target DNA template as described herein, and is therefore referred to in this example as target double-stranded DNA template 607. The two copies of the original DNA template contain, for example, hybridized polynucleotide strands 607a and 607b' (first copy) and hybridized polynucleotide strands 607b and 607a' (second copy). Also, as described herein, both copies contained both the parental DNA from the original target DNA template (strands 607a' and 607b' shown in black) and their complementary copy strands (strands 607b and 607a shown in gray). These two copies are separated by a first double-stranded crosslinking region 608a (i.e., the original crosslink) derived from the first (or first) terminal adapter used to form target double-stranded DNA template 607 as described herein. - The first crosslink 608a has, for example, strands 620 and 630. The target double-length DNA template 607 also includes the first and second template terminal ends 605 and 606, respectively, both of which are ligated to the second EA 600.
[0166] A second terminal adapter (EA) 600 is also shown, which has a structure such as EA100 in Figure 1A. For example, the second EA600 includes a first nick site 601a and a second nick site 601b, both of which can accept polymerase binding and extension (e.g., bidirectional extension as described herein). The EA600 also has first and second EA terminal ends 605 and 606, respectively. For example, both EA terminal ends 605 and 606 can be ligated to a target double-length DNA template 607 as described herein.
[0167] In step 6a, for example, the second EA600 is ligated to either end of the target double-length DNA template 607. For example, the terminal end 603 of the second EA600 is ligated to the template end 606 (Figure 6, step 6a). Alternatively, although not shown for simplification in step 6a, the other terminal end of the second EA600 (i.e., end 604) is ligated to the terminal end 605 of the target double-length DNA template 607. In either case, one end of the second EA600 is ligated to the end of the target double-length DNA template.
[0168] In step 6b of Figure 6, the remaining free end of the second EA 600 is ligated to the remaining free end of the target double-length DNA template 607 to form a circular construct 609. That is, in step 6b of Figure 6B, the two terminal ends 603 and 604 of the second EA 600 ligate to the respective ends 605 and 606 of the target double-length template 607, thereby forming a second DNA crosslink 608b between the ends of the target double-length template 607. That is, the entirety of the second EA 600 forms a second DNA crosslink 608b between the two ends 605 and 606 of the target double-length DNA template 607. - This forms a circular construct 609 including the first EA (as crosslink 608a) and the second crosslink 608b (formed from the second EA 600). In this way, the second EA600 acts as a crosslinking precursor for the second crosslinking region 608b of the cyclic construct 609. The first and second nicks 601a and 601b remain within the cyclic construct 609 (as part of the second crosslinking region 608b), and thus, in certain exemplary embodiments, provide two separate 3' ends available for polymerase binding and bidirectional extension, as described herein.
[0169] In steps 6c to 6d, the circular construct 609 is replicated as described for steps 1c to 1d in Figures 1C and 1D (these steps are combined in Figure 6 for simplification). As shown, once steps 6c to 6d in Figure 6 are completed, a quadruple-length DNA template 611 is obtained. For example, in step 6c, the circular construct 609 is brought into contact with polymerases (e.g., first and second polymerases (not shown)) that bind to the nick sites 601a and 601b of the circular construct 609. These polymerases then elongate the circular construct 609 in both directions. For example, one polymerase elongates the 3' end of nick site 601a while also substituting its 5' end. Similarly, the other polymerase elongates the 3' end of nick site 601b while also substituting its 5' end. In step 6d, once the polymerase completes the extension reaction of the cyclic construct 609, they dissociate from the cyclic construct 609 to form a quadruple-length DNA template 611.
[0170] As shown, the quadruple-length DNA template 611 contains four copies of the target sequence. For example, copy 1 contains the original parental target template DNA strand 607a (carried from the original parental target DNA template to the diplic-length DNA template) and its complementary, newly synthesized non-parental strand 607c. Copy 2 contains, for example, the non-parental strand 607a' (derived from the diplic-length DNA template) together with the newly synthesized non-parental strand 607d. As shown, copies 1 and 2 are separated by the strand segment 620 (black and gray circles) of the first crosslink 608a, along with their newly synthesized complementary portion (gray rectangle).
[0171] Similarly, copy 3 contains the non-parent strand 607b' (derived from the double-length DNA template) together with the newly synthesized non-parent strand 607e. As shown, copies 2 and 3 are separated by the second crosslinking region 608b, and the second crosslinking region 608b, which contains the portion, forms EA600 (black) and its newly synthesized portion (gray). Furthermore, copy 4 contains the original parent target template DNA strand 607b (carried from the original parent target DNA template to the double-length DNA template) and its complementary newly synthesized non-parent strand 607f. As shown, copies 3 and 4 are separated by the strand segment 630 (white-outlined gray box) of the first crosslink 608a, together with its newly synthesized complementary portion (gray shaded circle).
[0172] In particular, in the example in Figure 6, in the 5'→3' direction, the parent polynucleotide 607a is adjacent to non-parent copies 607a', 607e, and 607f via a crosslinking region sequence. Parent 607a also shares the same 5'→3' polynucleotide sequence as non-parent copies 607a', 607e, and 607f. Similarly, in the 5'→3' direction, the parent polynucleotide 607b is adjacent to non-parent copies 607b', 607d, and 607c (via a crosslinking region sequence). Parent 607b also shares the same 5'→3' polynucleotide sequence as non-parent copies 607b', 607d, and 607c.
[0173] Figure 6 shows the formation of a quadruple-length DNA template using EA in Figure 1A, but it should be understood that the method in Figure 6 can be repeated multiple times, for example, by doubling the number of copies of the target DNA template each time. For example, the first double-length DNA template contains two copies of the target DNA template as described herein, but further duplication (as in the exemplary method in Figure 6) generates four copies of the target DNA template, i.e., a quadruple-length DNA template 611. Further iterations thereafter yield the initial target DNA template with 8, 16, 32, 64 copies, etc.
[0174] Furthermore, it should be understood that a quadruple-length DNA template can be formed using any of the end adapters and their associated uses described herein. Alternatively, it should be understood that a multiplicative-length DNA template can be formed using any of the end adapters and their associated methods described herein, when the replication shown in Figure 6 is repeated. This includes, for example, the use of different EAs in different iterations when forming a multiplicative-length DNA template.
[0175] For example, the first duplicate DNA template may be formed using the EA in Figure 1A, and the second repeat may form a quadruplicate DNA template 611, also using the EA in Figure 1A. Subsequently, additional repeats may form a multiplicate DNA template of eight copies using the EA in Figure 2A (EA200), Figure 3A, and / or the EA in Figure 4A (EA300). Thus, in certain exemplary embodiments, the quadruplicate or multiplicate DNA template may include a Y-branched EA to facilitate subsequent PCR amplification and / or include UMI and SID to facilitate bioinformatics (including genetic and epigenetic analysis as described herein). In fact, such quadruplicate and multiplicate DNA templates are particularly useful for sequencing reactions and the validation of their associated data. Additionally or alternatively, in certain exemplary embodiments, one or more EAs may include protected nick sites as described herein (e.g., Figures 5A and 5B) to form an asymmetric duplicate or multiplicate DNA template.
Claims
1. A linear end adapter for replicating a target DNA template, A first polynucleotide chain that hybridizes to a second polynucleotide chain, thereby forming a polynucleotide double helix, wherein the polynucleotide double helix includes a first terminal end and a second terminal end, A first nic region and a second nic region, wherein the first nic region is located within the first polynucleotide chain of the polynucleotide double helix, and the second nic region is located within the second polynucleotide chain of the polynucleotide double helix, A spacer region that separates the first and second nick portions from each other, thereby linearly offsetting the first nick portion from the second nick portion. A linear terminal adapter, including one.
2. The linear end adapter according to claim 1, wherein the first nick site of the first polynucleotide chain includes a discontinuous cleavage in the sequence of the first polynucleotide chain, and / or the second nick site of the second polynucleotide chain includes a discontinuous cleavage in the sequence of the second polynucleotide chain.
3. The linear end adapter according to claim 1 or 2, wherein each terminal end is configured to ligate to both ends of the target DNA template.
4. The linear end adapter according to claim 3, wherein the first and / or second terminal ends of the polynucleotide double helix include ligable blunt ends.
5. The linear end adapter according to claim 3, wherein the first and / or second terminal ends of the polynucleotide double helix include a ligationable nucleic acid overhang.
6. A linear end adapter according to any one of claims 1 to 5, wherein each nick site is configured for a polymerase-mediated elongation reaction.
7. A linear end adapter according to any one of claims 1 to 6, wherein the linear offset between the first nick site and the second nick site corresponds to the distance at which polymerase can be bound to the first nick site and the second nick site.
8. The linear end adapter according to any one of claims 1 to 7, wherein the first nick portion and / or the second nick portion includes a 3' end and a 5' end.
9. The linear end adapter according to claim 8, further comprising a first Y-branched element array coupled to the 5' end adjacent to the first nick site and / or a second Y-branched element array coupled to the 5' end adjacent to the second nick site.
10. The linear end adapter according to claim 9, wherein the first Y-branched element sequence and / or the second Y-branched element sequence encodes a primer-binding sequence.
11. The linear terminal adapter according to claim 9 or 10, wherein the first Y-branching element sequence and / or the second Y-branching element sequence are approximately 5 to 25 nucleotides in length.
12. The linear terminal adapter according to any one of claims 1 to 11, wherein the first polynucleotide chain and / or the second polynucleotide chain comprises a unique molecular identifier (UMI) sequence.
13. The linear end adapter according to claim 12, wherein the UMI is arranged within the spacer region.
14. A linear end adapter according to any one of claims 1 to 13, wherein the first polynucleotide chain comprises a first sequence index (SID) and / or the second polynucleotide chain comprises a second SID.
15. A linear end adapter according to any one of claims 9 to 13, wherein the first polynucleotide chain includes a first sequence index (SID), the second polynucleotide chain includes a second SID, the sequence of the first SID is contiguous with the first Y branch element sequence, and the sequence of the second SID is contiguous with the second Y branch element sequence.
16. The linear end adapter according to any one of claims 1 to 15, wherein the linear end adapter is approximately 50 to 100 nucleotides in length.
17. The linear end adapter according to any one of claims 1 to 16, wherein the spacer region is approximately 10 to 50 nucleotides long.
18. The linear end adapter according to any one of claims 1 to 17, wherein the first nick site and / or the second nick site have a length corresponding to about 0 to 10 nucleotides.
19. The linear end adapter according to any one of claims 1 to 5, wherein only one of the aforementioned nick sites is configured for a polymerase-mediated extension reaction.
20. The linear end adapter according to claim 19, wherein either the first nick site or the second nick site contains a 3'-blocking group that prevents polymerase-mediated elongation.
21. The linear terminal adapter according to claim 20, wherein the 3'-blocking group is a phosphate group.
22. A linear end adapter according to any one of claims 19 to 21, wherein the first nick portion and / or the second nick portion includes a 5' end, the 5' end of the first nick portion includes a first Y-branch element array, and / or the 5' end of the second nick portion includes a second Y-branch element array.
23. The linear terminal adapter according to claim 22, wherein the first Y-branch element and / or the second Y-branch element encode a primer-binding sequence.
24. The linear terminal adapter according to any one of claims 19 to 23, wherein the spacer region includes a UMI array.
25. A method for replicating a target DNA template, A ligation reaction is performed between a target DNA template and a linear end adapter according to any one of claims 1 to 24, thereby forming a circular construct. The target DNA template includes a first target DNA template end and a second target DNA template end, The ligation reaction involves (i) ligating the first terminal end of the terminal adapter to the terminal end of the first target DNA template, and (ii) ligating the second terminal end of the terminal adapter to the terminal end of the second target DNA template, thereby forming the circular construct. The aforementioned circular construct is subjected to a DNA polymerase-mediated extension reaction, thereby replicating the target DNA template. Methods that include...
26. The method according to claim 25, wherein the DNA polymerase-mediated extension reaction includes contacting the cyclic construct with a plurality of strand substitution polymerases.
27. The method according to claim 26, wherein the strand substitution polymerase is selected from the group consisting of KAPA HiFi DNA polymerase, Q5® High-Fidelity DNA polymerase, and Pfu DNA polymerase such as Pfu-X.
28. The method according to claim 26, wherein the chain substitution polymerase is phi29 polymerase.
29. The method according to any one of claims 25 to 28, wherein the polymerase-mediated extension reaction includes extension of the 3' end of the first nick site or the 3' end of the second nick site of the terminal adapter.
30. The method according to claim 29, wherein (i) polymerase-mediated elongation of the 3' end of the first nick site of the terminal adapter or polymerase-mediated elongation of the 3' end of the second nick site of the terminal adapter forms an asymmetric DNA template, or (ii) polymerase-mediated elongation of both the 3' end of the first nick and the 3' end of the second nick site of the terminal adapter forms a double-length DNA template.
31. The method according to claim 30, wherein the strands of the asymmetric DNA template or the double-length DNA template include a unique molecular identifier (UMI).
32. The method according to claim 31, wherein the UMI is located within the chain of the crosslinking region of the asymmetric DNA template or the double-length DNA template.
33. The method according to any one of claims 30 to 32, wherein the strands of the asymmetric DNA template or the double-length DNA template include a sequence index (SID).
34. The method according to claim 33, wherein the SID is located within the chain of the crosslinking region of the asymmetric DNA template or the double-length DNA template, and / or is located at the terminal end of the asymmetric DNA template or the double-length DNA template.
35. The method according to any one of claims 25 to 33, wherein the terminal end of the asymmetric DNA template or the double-length DNA template includes a Y-branched terminal adapter.
36. The method according to claim 35, wherein the Y-branched terminal adapter codes for a primer binding site.
37. A method for preparing a double-length DNA template a from a target DNA template, A ligation reaction is performed between a target DNA template and a terminal adapter according to any one of claims 1 to 18, thereby forming a circular construct. The target DNA template includes a first target DNA template end and a second target DNA template end, The ligation reaction forms a circular construct in which (i) the first terminal end of the terminal adapter is joined to the first target DNA template terminal end, and (ii) the second terminal end of the terminal adapter is joined to the second target DNA template terminal end. The cyclic construct is subjected to a DNA polymerase-mediated extension reaction, thereby forming a double-length DNA template containing a first copy and a second copy of the target DNA template. Methods that include...
38. The method according to claim 37, wherein the DNA polymerase-mediated elongation includes contacting the cyclic construct with a plurality of strand substitution polymerases.
39. The method according to claim 38, wherein the strand substitution polymerase is selected from the group consisting of KAPA HiFi DNA polymerase, Q5® High-Fidelity DNA polymerase, and Pfu DNA polymerase such as Pfu-X.
40. The method according to claim 38, wherein the chain substitution polymerase is phi29 polymerase.
41. The method according to any one of claims 37 to 40, wherein the polymerase-mediated extension reaction includes extension of the 3' end of the first nick site and the 3' end of the second nick site of the terminal adapter.
42. The method according to any one of claims 37 to 41, wherein the polymerase-mediated elongation is bidirectional.
43. The method according to any one of claims 37 to 41, wherein the first copy of the target DNA template and the second copy of the target DNA template are linked adjacent to each other via DNA crosslinking regions.
44. The method according to claim 43, wherein the crosslinking region originates from the terminal adapter.
45. The method according to claim 43 or 44, wherein each polynucleotide strand of the double-length DNA template includes a 5' to 3' parent strand of the target DNA template and a 5' to 3' daughter strand copy of the parent strand of the target DNA template.
46. The method according to claim 45, wherein the parent strand and the daughter strand copy of the target DNA template are adjacent to each other via the 5' to 3' strands of the DNA crosslinking region.
47. The method according to claim 46, wherein the chain in the crosslinked region includes a unique molecular identifier (UMI).
48. The method according to claim 46 or 47, wherein the chain in the crosslinked region includes an array index (SID).
49. The method according to any one of claims 45 to 48, wherein the double-length DNA template includes a first terminal end and a second terminal end, and the first terminal end and / or the second terminal end includes an SID.
50. The method according to any one of claims 37 to 49, wherein the linear end adapter includes a first Y branching element sequence and a second Y branching element sequence, and by performing the DNA polymerase-mediated extension reaction, the first Y branching element sequence and the second Y branching element sequence are positioned at the 5' ends of each parent strand of the double-length DNA template.
51. The method according to claim 50, wherein the polymerase-mediated extension reaction of the DNA circular construct synthesizes a first daughter Y branch sequence and a second daughter Y branch sequence, wherein the first daughter Y branch sequence is complementary to the first Y branch sequence and the second daughter Y branch sequence is complementary to the Y branch sequence.
52. The method according to claim 51, wherein the first daughter Y branch element sequence and the second daughter Y branch element sequence are located at the 3' end of each daughter strand copy of the double-length DNA template.
53. The method according to claim 50 or 52, wherein the Y-branched element codes for a primer binding site.
54. The method according to any one of claims 37 to 52, wherein the above method is repeated in succession to form a quadruple-length DNA template or a multiplicative-length DNA template.
55. A double-length DNA template formed by the method according to any one of claims 37 to 53.
56. A method for identifying epigenetic information associated with a target nucleic acid sequence, (a) Ligating linear target DNA templates to both ends of a linear end adapter according to any one of claims 1 to 18, thereby forming a circular DNA construct, (b) Performing a DNA polymerase-mediated bidirectional elongation reaction of the circular DNA construct in the presence of multiple protected cytosine nucleotides, thereby forming a double-length DNA template containing the protected cytosine nucleotides, (c) Denaturing the double-length DNA template, (d) The denatured double-length DNA template is subjected to a bisulfite conversion reaction, thereby forming a bisulfite-converted double-length DNA template strand of the double-length DNA template, (e) Perform polymerase chain reaction (PCR) amplification of the bisulfite-converted double-length DNA template strand, (f) Sequence determination of the PCR-amplified bisulfite-converted double-length DNA template strand, (g) Identifying epigenetic information related to the target nucleic acid based on the sequencing of the PCR-amplified bisulfite-converted double-length DNA template strand. Methods that include...
57. The method according to claim 56, wherein each polynucleotide strand of the double-length DNA template in step (b) includes a parent template strand derived from the target DNA template and a daughter copy strand of the parent template strand.
58. The method according to claim 57, wherein the parent mold chain is connected adjacent to the daughter copy chain of the parent mold chain via a single-strand bridging region.
59. The method according to claim 58, wherein the single-chain crosslinking region is derived from the terminal adapter.
60. The method according to claim 57 or 58, wherein the protected cytosine nucleotide is incorporated into the daughter copy strand of the parent template strand during the DNA polymerase-mediated bidirectional extension reaction of step (b).
61. The sequencing of the PCR-amplified bisulfite-converted double-length DNA template strand in step (f) provides the polynucleotide sequences of the parent template strand and the daughter copy strand. The method according to any one of claims 57 to 60, wherein identifying the epigenetic information related to the target nucleic acid includes an intra-chain comparison of the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the daughter copy strand.
62. The method according to claim 61, wherein the sequence mismatch locations between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the daughter copy strand identify the unprotected cytosine residue locations in the parent template strand.
63. The method according to claim 62, wherein the unprotected cytosine residue position in the parent template strand corresponds to the unprotected cytosine residue position in the target nucleic acid sequence.
64. The method according to any one of claims 61 to 63, wherein the position of a cytosine residue in the sequence of the parent template chain indicates the corresponding position of a protected cytosine in the target nucleic acid sequence.
65. The method according to claim 56, wherein the double-length DNA template in step (b) includes a first copy and a second copy of the target DNA template.
66. The method according to claim 65, wherein the first copy and the second copy of the target DNA template are linked together via a double-stranded crosslinking region.
67. The method according to claim 66, wherein the double-strand crosslinking region is derived from the terminal adapter.
68. The method according to any one of claims 65 to 67, wherein each copy of the target DNA template in the double-length DNA template comprises a parent template strand and a daughter strand that is complementary to and hybridized with the parent template strand.
69. The method according to claim 68, wherein during the DNA polymerase-mediated bidirectional extension reaction of step (b), the protected cytosine nucleotide is incorporated into the hybridized complementary daughter strand.
70. The sequencing of the PCR-amplified bisulfite-converted double-length DNA template strand in step (f) provides the polynucleotide sequences of the parent template strand and its hybridized complementary daughter strand. The method according to claim 68 or 69, wherein identifying the epigenetic information related to the target nucleic acid includes an inter-strand comparison between the polynucleotide sequence of the parent template strand and the polynucleotide sequence of the hybridized complementary daughter.
71. The method according to claim 70, wherein the nucleotide mismatch sites between the polynucleotide sequence of the parent template chain and the hybridized complementary daughters identify the unprotected cytosine residue sites in the parent template chain.
72. The method according to claim 71, wherein the unprotected cytosine residue position in the parent template strand corresponds to the unprotected cytosine residue position in the target nucleic acid sequence.
73. The method according to any one of claims 56 to 72, wherein the protected cytosine nucleotide comprises a methylated cytosine residue.
74. The method according to any one of claims 56 to 72, wherein the unprotected cytosine nucleotide is a non-methylated cytosine residue.
75. The method according to any one of claims 56 to 74, wherein the double-length DNA template in step (b) includes a unique molecular identifier (UMI).
76. The method according to claim 75, wherein the UMI is located in the single-chain crosslinking region described in claim 59 or the double-chain crosslinking region described in claim 67.
77. The method according to any one of claims 56 to 76, wherein the double-length DNA template in step (b) includes a sequencing index (SID).
78. A double-length DNA template comprising a first copy and a second copy of a target DNA template, wherein the first copy and the second copy of the target DNA template are linked adjacent to each other via double-stranded crosslinking regions.
79. The double-length DNA template according to claim 78, wherein each polynucleotide chain of the double-length DNA template includes a parent template strand derived from the target DNA template and a daughter copy strand of the parent template strand.
80. The double-length DNA template according to claim 78 or 79, wherein the parent template strand is connected adjacent to the daughter copy strand of the parent template strand via the cross-linking region strand.
81. The double-length DNA template according to claim 78, wherein each copy of the target DNA template within the double-length DNA template comprises a parent template strand and a daughter strand that is complementary to and hybridized with the parent template strand.
82. The double-length DNA template according to any one of claims 78 to 81, wherein the double-length DNA template includes a first terminal end and a second terminal end, and either terminal end includes a sequence encoding a primer binding site.
83. The double-length DNA template according to any one of claims 78 to 82, wherein the crosslinking region or its chain contains a unique molecular identifier (UMI).
84. The double-length DNA template according to any one of claims 78 to 83, wherein the crosslinking region or its strand includes a sequencing index (SID).