Methods and compositions for sequencing double-stranded nucleic acids
By developing a method and composition for forming nucleic acid clusters on a solid support, and utilizing primer hybridization and extension techniques to determine sequences on the sense and antisense strands of nucleic acids, this invention solves the problem of high time consumption and cost associated with existing double-stranded sequencing, achieving efficient and accurate double-stranded sequencing that is adaptable to nucleic acid templates of different lengths.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PACIFIC BIOSCIENCES OF CALIFORNIA INC
- Filing Date
- 2021-03-03
- Publication Date
- 2026-04-24
AI Technical Summary
Existing double-stranded or paired-end sequencing methods are time-consuming, costly, and prone to introducing sequence errors, making it difficult to efficiently sequence multiple regions of the target nucleic acid.
A method and composition are provided for determining sequences on both strands by forming nucleic acid clusters on a solid support, comprising sense and antisense strands in tandem, using primer hybridization and extension techniques, and processing the strands after sequencing to avoid errors, including capping or blocking portions to control extension.
It enables efficient and accurate extraction of sequence information from both strands of nucleic acid without relying on thermal cycling, reducing replication steps and error rates, adapting to nucleic acid templates of different lengths, and supporting parallel or serial sequencing.
Smart Images

Figure CN115516104B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 984,438, filed March 3, 2020, the contents of which are incorporated herein by reference in their entirety.
[0003] References to sequence lists
[0004] This application contains an electronic sequence list. The sequence list is provided as a file entitled Sequence_Listing_42HB-328033-WO, created on February 26, 2021, and is 3 kilobytes in size. The information in the electronic sequence list is incorporated herein by reference in its entirety.
[0005] background
[0006] This disclosure generally relates to obtaining sequence information from nucleic acids and has specific applicability to sequencing double-stranded or paired ends of nucleic acids.
[0007] Many nucleic acid sequencing methods involve the extension of primers along a target nucleic acid template. Primers can be extended by sequentially incorporating nucleotides in a template-dependent manner (usually using polymerases), and the sequence of the incorporated nucleotides can be observed to determine the target nucleic acid sequence. When performing nucleic acid sequencing, it is often necessary to sequence more than one region of the target nucleic acid. For example, sequencing the opposite strand of the target nucleic acid in opposite directions can provide information. This double-stranded sequencing offers the advantage of increasing the total amount of sequence information obtainable from a given target nucleic acid, especially when the target length exceeds the read length typically obtainable from the sequencing method employed. In cases where the read length of the sequencing method is on the same order of magnitude as the template length, bidirectional sequencing can offer advantages such as independent verification of sequence information obtained by comparing “sense” reads with corresponding “antisense” reads. Various double-stranded or paired-end sequencing methods have been proposed, but these methods increase time and cost compared to single-stranded sequencing methods. For example, the requirements for synthesizing complementary strands and / or removing strands between two sequencing reads can be cumbersome, time-consuming, and expensive, and if the strand used for sequencing the second read is derived from a strand that has already been sequenced, it may introduce errors into the sequence results.
[0008] Therefore, there is a need for efficient methods to sequence more than one region of a target nucleic acid. This invention addresses this need and provides related advantages.
[0009] Brief Overview
[0010] This disclosure provides a method for determining a sequence from the sense and antisense strands of a nucleic acid. The method may include the following steps: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a tandem sense strand and a tandem antisense strand, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand within the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand. Optionally, after step (c) and before step (d), a capping or blocking portion may be incorporated into the primer extending along the antisense strand such that the primer does not extend further during step (e) when the second primer extends along the sense strand.
[0011] This document also provides a composition comprising a solid support to which a nucleic acid cluster is attached, the nucleic acid cluster comprising a sense strand of a tandem strand and an antisense strand of the tandem strand, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site, optionally wherein the sense strand is covalently attached to the solid support, and optionally wherein the antisense strand is covalently attached to the solid support. Brief description of the attached diagram
[0013] Figure 1A A schematic diagram of a method for generating circular template nucleic acids that hybridize with immobilized nucleic acid primers (e.g., capture primers) is shown. Figure 1B A schematic diagram is shown of primers immobilized along the circular template nucleic acid extension via rolling circle amplification to generate clusters with immobilized single tandem strands. Figure 1C A schematic diagram is shown illustrating the use of more than one primer (e.g., an amplification primer) to extend a fixed sense strand along a tandem strand to produce a cluster having a fixed sense strand that hybridizes with more than one antisense strand of the tandem strand. Figure 1D The diagram illustrates the generation of a fixed sense strand by extending a primer (e.g., a capture primer) along a circular template nucleic acid via rolling circle amplification to produce a fixed sense strand in a tandem, and the generation of a cluster of fixed sense strands having a fixed sense strand that hybridizes with more than one antisense strand by extending more than one primer (e.g., a capture primer) along the fixed sense strand in the same reaction. The primers in this disclosure (such as capture primers or amplification primers) can be any nucleic acid, such as RNA, DNA, DNA / RNA chimeras, or other molecules capable of hybridizing with another nucleic acid (e.g., template nucleic acid, sense strand, or antisense strand). Figure 1E A schematic diagram illustrates the extension of more than one primer (e.g., amplification primer) along a fixed sense strand of a tandem to generate a cluster of fixed sense strands having hybridization with more than one antisense strand of the tandem, wherein one, more, or each of the more than one antisense strand contains one or more nucleotides or bases that can tag or target the antisense strand for degradation. For example, the antisense strand may contain one or more nucleotides that are uridine monophosphates. As another example, the antisense strand may contain one or more nucleotides that are deoxyribopseudouridine monophosphates. For example, the antisense strand may contain one or more bases that are uracil. As another example, the antisense strand may contain one or more modified bases or one or more modified nucleotides. As another example, the antisense strand may contain one or more bases that are atypical bases or one or more atypical nucleotides.
[0014] Figure 2A A schematic diagram is shown of a method for processing nucleic acid clusters to close or cap the 3' end of the second strand (e.g., antisense strand), and then sequencing the second strand in the cluster. Figure 2B A schematic diagram is shown of a method for processing nucleic acid clusters after sequencing the second strand (e.g., antisense strand) in the cluster to block or cap the 3' end of the primer extension product, and then sequencing the first strand (e.g., sense strand) in the presence of the second strand in the cluster. Figure 2C A schematic diagram is shown of a method for processing nucleic acid clusters after sequencing the second strand in the cluster to remove the second strand (e.g., antisense strand) and primer extension products, and then sequencing the first strand (e.g., sense strand) in the absence of the second strand and primer extension products in the cluster. Figure 2D A schematic diagram is shown of a method for processing nucleic acid clusters after sequencing the second strand in a cluster to digest the second strand (e.g., antisense strand) and primer extension products, and then sequencing the first strand (e.g., sense strand) in the absence of the second strand and primer extension products in the cluster.
[0015] Figures 3A-3F A non-limiting exemplary illustration of a method for determining a sequence from the first and second strands (e.g., sense and antisense strands) of a nucleic acid is shown. Figure 3A Clustering is shown to generate the first chain. Figure 3B The sequencing of the first strand is shown. Figure 3C The extension products of the sequencing primers are shown to be further extended using strand displacement polymerase. Figure 3D The diagram shows how further extensions of the extension products from first-strand sequencing produce second-strand scales. Figure 3E The 3' end of the closed second chain is shown. Figure 3F The initiation and sequencing of the second-strand scales are shown. Figures 3A-3FIn the diagram, the first chain is indicated by a dashed line, and the second chain is indicated by a solid line.
[0016] Figures 4A-4F A non-limiting exemplary illustration of a method for determining sequences from the sense and antisense strands of nucleic acids is shown. Figure 4A Two surface primers attached to the surface are shown. Figure 4B The diagram shows the libraries hybridizing and being ligated on a splint. Figure 4C The polymerase is shown performing RCA on the first chain. Figure 4D The second surface primer that hybridizes with the first strand is shown. Figure 4E The scaling that occurs in the same RCA reaction via polymerase is shown (not shown for simplicity). RCA and MDA can occur simultaneously or sequentially. For example, the 3' end of the primer amplification during RCA may contain a capping portion, allowing RCA and MDA to occur sequentially. Figure 4F The image shows scales that have undergone substitution initiated by sequencing primers.
[0017] Figure 5 The extracted signal intensities of the clusters are shown, obtained from sequencing 25 nucleotides from each of the two strands at more than one nucleic acid site in the array.
[0018] Figure 6 The extracted signal intensities of the clusters are shown, obtained from sequencing 100 nucleotides from each of the two strands at more than one nucleic acid site in the array.
[0019] Figure 7 The distribution of fragment lengths is shown, determined from paired end reads sequenced from 100 nucleotides from each of the two strands at more than one nucleic acid site in the array. The extracted signal intensity of the aggregates sequenced from 100 nucleotides is shown in... Figure 6 The explanation is as follows.
[0020] Figures 8A-8B The sequencing methods using RCA and MDA with capture primers and amplification primers attached to a solid support, as shown in this disclosure, do not favor target sequences of a specific size and can produce high-quality reads.
[0021] Figure 9A This demonstrates high-quality signal density for sequencing the first strand generated from a nucleic acid template using capture primers that are amplified using rolling circle and attached to a solid support. Figure 9B The diagram shows a high-quality signal density for sequencing a second strand generated from a first strand using multiple displacement amplification and amplification primers attached to a solid support, wherein the first strand is generated from a nucleic acid template using rolling loop amplification and capture primers attached to a solid support.
[0022] Detailed description
[0023] Various methods have been developed to determine the sequences of two distinct parts of a nucleic acid template. In many cases, these methods are configured to determine the sequences of opposite ends of the nucleic acid template and are therefore referred to as “paired-end” sequencing. Typically, the sequence of the first end is determined by extending a first primer along the first strand of the nucleic acid template, and the sequence of the second end is determined by extending a second primer along the second strand of the nucleic acid template. Since the nucleotides are oriented oppositely on one strand to their orientation on the other strand (these two strands are called “antiparallel”), the first and second primers extend in opposite directions and toward each other. In particular, because of this orientation, other paired-end sequencing methods require sequencing each strand in the absence of the other strand. Therefore, paired-end methods typically require a step of removing or synthesizing a strand between the two sequencing reads. This disclosure provides a method for sequencing the sense strand of a target nucleic acid in the presence of the antisense strand of the target nucleic acid. Figures 2A-2B This describes a non-limiting exemplary method for sequencing the sense strand of a target nucleic acid in the presence of an antisense strand of the target nucleic acid.
[0024] The methods and compositions described herein relate to paired-read sequencing of tandem repeat templates. Primers are annealed to the tandem repeat and extended (typically using a strand displacement polymerase) to produce polymeric, partially double-stranded molecules with continuous tandem repeat chains and a series of distinct extension products, which are annealed to the tandem repeat chains at their 3' ends. Their 5' ends are single-stranded and exposed for sequencing, for example using primers annealed to conserved regions in the tandem repeats in which they are replicated. After sequencing these templates, they are optionally removed or degraded, thereby exposing the polymeric, partially double-stranded molecules and allowing sequencing using primers annealed to, for example, conserved regions in the tandem repeats. In an alternative embodiment, the tandem repeat strand is sequenced prior to generating the series of distinct extension products. Similarly, in yet another embodiment, the tandem repeat strand is sequenced simultaneously with the generation of the series of distinct extension products, such that the extension products are a result of sequencing.
[0025] In a more detailed description of the compositions and methods herein, we begin with linear nucleic acids, such as library components having different 5' and 3' adaptors. The linear nucleic acid is annealed to a guiding oligonucleotide, such as a surface-bound oligonucleotide, having regions complementary to the 5' and 3' adaptors of the linear nucleic acid, and oriented to position the 5' and 3' ends of the linear nucleic acid in close proximity. See also Figure 1A Connect the 5' and 3' ends of the linear nucleic acid to circularize it.
[0026] Then, the guide oligonucleotide was extended via rolling circle amplification to add more than one monomeric unit of the original linear nucleic acid to its 3' end. The result was a tandem polymer of the original linear nucleic acid tethered to the surface by the guide oligonucleotide, as shown below. Figure 1B As seen in the diagram. This reaction can be terminated by thermal inactivation of the polymerase or by washing, alone or in combination with chemical inactivation.
[0027] The methods and compositions described herein involve one or more steps, such that practice of the disclosure herein may include some or all of the steps disclosed herein. Some methods effectively "begin" midway through the process described herein. Thus, for example, in Figure 1D The process observed begins with a circularized library component having adjacent first and second adaptor regions, rather than a linear nucleic acid with 5' and 3' adaptors. Subsequent steps are observed to proceed in a very similar manner.
[0028] Multiple substitutional amplification oligonucleotides contact and extend with the tandem strand, forming a series of different extension products that anneal to the tandem strand at their 3' ends, such as... Figure 1C As seen in the reference. In some implementations, these extension reactions are also sequencing reactions, as in the reference. Figures 3B-3D As described. Optionally, in some cases, these extension reactions are performed without sequencing, thereby maintaining the integrity of the template and the extended strand by avoiding nucleic acid damage that often occurs in other steps of the imaging process or sequencing reaction. The oligonucleotides are provided in solution or optionally tethered to the surface. When the oligonucleotides are tethered to the surface, the result is tethering a surface with more than one tandem nucleic acid molecule, wherein the tandem nucleic acid molecule contains more than one monomeric unit, and wherein the nucleic acid molecules locally, partially hybridize to form locally double-stranded segments adjacent to the replaced single-stranded segments. For example, a sense strand may locally, partially hybridize with multiple antisense strands. In some embodiments, the antisense strand is sequenced while the antisense strand is locally, partially hybridized with the sense strand, and / or the sense strand is sequenced while the sense strand is locally, partially hybridized with the antisense strand. Optionally or additionally, the antisense strand and the sense strand are separated by, for example, thermal denaturation, such that when the antisense strand (or sense strand) is sequenced, the antisense strand and the sense strand do not locally, partially hybridize.
[0029] The elongation reaction initiated by MDA oligonucleotides may optionally be terminated, for example, by introducing a 3'-blocking nucleotide. See also Figure 2AIn some cases, the 3' nucleotide binds to a portion of the strand that is large or shaped enough to inhibit the formation of the ternary complex, preventing the polymerase from binding to the 3' end of the terminated strand, even if it binds to the template strand. This has the benefit of reducing background signal when using sequencing by binding in subsequent sequencing reactions, as described elsewhere in this article.
[0030] Then, the exposed single-stranded segments of a series of different extensions annealed to the tandem strand at their 3' ends are used as templates for sequencing reactions, optionally initiated by primers annealed to or collinear with conserved regions of the tandem strand (e.g., regions originally used to circularize components of linear nucleic acid libraries). See also Figure 2A .
[0031] Optionally, the sequencing reaction can be terminated using the same, similar, or different terminator portions as those used above. Similarly, in some cases, it is beneficial to use partial terminators that prevent the formation of ternary complexes so that they do not generate background signals in subsequent binding sequencing reactions. Optionally, these molecules are removed or degraded prior to subsequent sequencing of the complementary strand.
[0032] Typically, these sequencing reactions are not allowed to extend further after sequencing is complete, thus preventing the formation of complete monomers of tandem nucleic acids based on sequencing. Instead, they are terminated, either by incorporation of the blocking portion or by removal of sequencing reagents, or both. Optionally, some sequencing reactions are allowed to extend, thereby forming copies of the complete monomers that they sequenced for a portion of.
[0033] The second strand of the paired read is sequenced using any of a variety of methods. For example, in some cases, the primers are annealed at the 3' end of a single-stranded tandem with a monomer unit that is not annealed with an MDA oligonucleotide or an MDA oligonucleotide extension product, such as... Figure 2B As shown in the diagram. Optionally or in combination, a series of distinct extension products annealed to the tandem strand at their 3' ends are typically removed or degraded along with the sequencing reaction products annealed with them, thereby exposing the single-stranded tandem template. Removal can be achieved by unwinding the molecules, so that the molecules are no longer held together by their base pairings and some of the double-stranded nucleic acid products detach from the tandem template. Optionally, the MDA-initiated extension products can be incorporated with cleavable or cleavage-guided moieties, such as dUTP incorporated as uracil bases, which can be used to tag the extension products for degradation. Cleavage-guided moieties promote the degradation of some double-stranded nucleic acid products, such that they degrade in some cases while still attaching to the single-stranded tandem, and the reverse complementary sequencing primers maintain hybridization.
[0034] Sequencing primers, such as the previously used MDA primers or other primers annealed in the monomeric unit of the tandem, are then used to initiate a series of sequencing reactions. These sequencing reactions run on the opposite strand and in the opposite direction relative to the previous sequencing reaction, producing reads that can be paired with the sequencing reads from the previous round to produce sequencing data for two distinct regions of the template (such as the two ends of the original linear template used to generate the tandem molecule).
[0035] These methods offer numerous advantages over methods in the art or those used in sequencing. First, sequencing templates typically undergo no more than two, three, or four replication steps from linear or circular nucleic acids such as… Figure 1A The library components described herein are removed. That is, in some cases of the techniques described herein, no library component undergoes more than 2, 3, or 4 rounds of replication, depending on colony formation and sequencing. Replication may occur before colony formation and produce many more than 2, 3, or 4 copies, but in some cases, no colony sequencing product undergoes more than 2, 3, or 4 replication reactions from the starting library material. Optionally, in some cases, the library component undergoes more than 2, 3, or 4 replication cycles, such as 5, 6, 7, 8, 9, 10, or more than 10 replication cycles. Therefore, errors introduced by amplification-heavy processes, such as bridging amplification, which rely on using polymerase products as templates for subsequent amplification, are not introduced into the paired read methods disclosed herein. Therefore, the reaction may exhibit significantly higher accuracy than reactions that rely on other amplification methods.
[0036] The starting linear nucleic acid molecules, such as library components, are relatively unrestricted by length. Although bridging amplification is less efficient for library components that are too small or too large relative to the oligonucleotides located on the cluster, or whose replication takes a long time or carries the risk of polymerase shedding and incomplete extension, the method described in this paper is neutral to the length of the starting nucleic acid and can accommodate library components of less than 50, 45, 40, 35, 30, 25, 20, or less than 20 bases, or optionally at least 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, or more than 3000 bases.
[0037] Some methods described herein do not require thermal cycling and can be performed in an environment with some or all steps that can be completed isothermally. Other methods described herein do not require thermal cycling except for the thermal inactivation of the enzymes involved in the RCA and / or MDA reactions, and can be performed in an environment with some or all steps that can be completed isothermally. Optionally, inactivation is accomplished by chemical inactivation or removal (e.g., by washing). That is, rolling circle amplification, MDA primer annealing, sequencing, and degradation (such as uracil incorporation-mediated degradation) are, in some cases, performed at a single temperature, or specific sub-steps of the process are performed at a single temperature, such that thermal cycling is independent of its extent within the bridge amplification-mediated sequencing reaction. For example, in some cases, tandem formation isothermally occurs, MDA-initiated second-strand synthesis isothermally occurs, and sequencing of the second and first strands isothermally occurs, wherein the first, second, and third portions of these parts of this disclosure may occur at the same temperature or at temperatures different from each other.
[0038] In some cases, the method described in this paper does not rely on sequencing reaction products as templates for second-strand sequencing reactions. That is, in some cases, all templates are generated before any sequencing reaction is performed. This separation of template formation from sequencing facilitates isothermal reactions and reduces thermal stress on the sequencing apparatus, because all template formation can occur at a single temperature, optionally under isothermal conditions, while all sequencing subsequently occurs at a temperature suitable for sequencing, such as an isothermal temperature. Separating template generation from sequencing allows for reduced temperature variations and allows each set of steps to proceed in its own temperature mode.
[0039] Therefore, the methods and compositions of this paper generate paired sequence data in which none of the sequences in the paired sequence data is replicated from the library template more than four times. That is, in some cases of the techniques of this paper, depending on colony formation and sequencing, no library component undergoes more than 2, 3, or 4 rounds of replication. Replication may occur before colony formation and produce many more than 2, 3, or 4 copies, but in some cases, the sequencing product of no colony is removed from the starting library material by more than 2, 3, or 4 replication reactions. Optionally, in some cases, the library component undergoes more than 2, 3, or 4 replication cycles, such as 5, 6, 7, 8, 9, 10, or more than 10 replication cycles. The number of replication cycles can be, is about, is at least about, is at most, or is at most about: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, or any number or range between any two of these values. In some cases, paired sequencing reads are generated without thermal cycling during sequencing, or without thermal cycling during amplification, or isothermally. Some methods described herein do not require thermal cycling, except for enzyme inactivation by heating, and can be performed in an environment with some or all steps that can be performed isothermally. Alternatively, enzyme inactivation is accomplished by chemical inactivation or washing. In many cases, paired sequence reads are generated from templates that are not generated from previous templates subjected to sequencing reaction conditions. That is, all templates are generated before any sequencing reaction, resulting in cleaner templates and a more efficient sequencing workflow. Optionally, in some cases, the strand is sequenced before the reverse complement of the synthesized strand, which is then sequenced (see [link to documentation]). Figures 3A-3F (Examples).
[0040] The method described herein can be readily multiplexed, enabling the parallel acquisition of paired-end sequence data from more than one target nucleic acid. For example, more than one target sequence can be distributed across sites on a nucleic acid array. Each site can contain clusters of sense strands of the specific target nucleic acid and antisense strands of the target, and both strands can be present in each cluster during sequencing. Those skilled in the art will recognize that this method can also be configured to acquire paired-end sequence data from more than one target nucleic acid by performing serial processing on the target nucleic acids from more than one target nucleic acid.
[0041] This disclosure provides a method for preparing nucleic acid clusters containing both sense and antisense strands of a target nucleic acid. In a specific configuration, clusters are formed by a combination of rolling circle amplification (RCA) of a circular template nucleic acid to form a tandem sense strand and multiple substitution amplification (MDA) of one or more antisense strands to form a tandem. The sense strand can be attached to a solid support, for example, by means of primers fixed by extending the solid support during the RCA reaction. The antisense strand can be attached to a solid support, for example, by means of primers fixed by extending the solid support during the MDA reaction. The RCA and MDA reactions can be performed simultaneously (e.g., by amplification in the presence of both sense and antisense primers). Figure 1D A non-limiting exemplary method for performing RCA and MDA reactions simultaneously is described. Alternatively, the RCA and MDA reactions can be performed sequentially (e.g., by performing RCA to generate a sense strand in the absence of an antisense primer, and then delivering an antisense primer to generate an antisense strand via MDA). Figures 1A-1C A non-limiting exemplary method for sequentially performing RCA and MDA reactions is described. Surprisingly, the antisense strand does not need to be generated by fixed primers to remain in the cluster for double-stranded sequencing. Instead, due to the non-covalent interaction of the antisense strand with the sense strand or with other fixed parts of the cluster, the antisense strand generated by the combination of fixed primer RCA and soluble primer MDA can be retained in the cluster.
[0042] For some implementations that employ sequencing of a single strand within a cluster that also contains other strands, modifying the 3' ends of the other strands prior to sequencing can be useful. The methods described herein can employ primer modification procedures suitable for the specific sequencing method used. For example, unwanted 3' ends can be modified to incorporate a 3' blocking portion. This blocking portion can be used to prevent unwanted background signals. Unwanted background signals may originate from, for example, Sequencing By Binding. TM (SBB TM (Binding sequencing) or sequencing-by-synthesis (SBS) reactions. During the SBS process, undesirable 3' end elongation can lead to unwanted background. Particularly useful blocking portions are irreversible (e.g., dideoxynucleotides) or inert to removal by the reagents used in the SBS process (e.g., blocking portions are reversible terminators orthogonal to reversible terminators used in the SBS process). Optionally or in combination, the 3' ends in other strands can be modified to incorporate a 3' capping portion. The capping portion can be used to prevent 3' end elongation in the SBS process. TM Unwanted background signals result from the formation of a stabilized ternary complex at the 3' end during the process. Typically, the capping is irreversible, or at least for those affected by SBB. TMThe reagents used in the process are inert. If necessary, the capping can be reversible, for example, removed by denaturation, linker cleavage, etc.
[0043] Unless otherwise stated, the terms used herein shall be understood to have their common meanings in the relevant fields. Several terms used herein and their meanings are explained below.
[0044] As used herein, the term “about” refers to a quantity or range that spans + / - 10% of that quantity, or + / - 10% of the previously stated range limit.
[0045] As used herein, the term "array" refers to a group of molecules attached to one or more solid supports in a manner that allows molecules to be distinguished from one another. An array may comprise different molecules, each located at a different addressable site on a solid support. An array may include individual solid supports, each serving as a site for carrying different molecules, wherein different molecules can be identified based on the position of the solid support on the surface to which it is attached, or based on the position of the solid support in a liquid such as a fluid flow. The molecules in an array may be, for example, nucleotides, nucleic acid primers, nucleic acid templates, or nucleases (such as polymerases, ligases, exonucleases, or combinations thereof).
[0046] As used herein, the phrase “at least one of” A, B, and C (in any order) refers to a set that may include A, or may include A and B, or may include A, B, and C (alone or in combination with other unlisted elements).
[0047] As used herein, the term "attachment" refers to the state in which two objects are joined, fastened, adhered, connected, or bound together. For example, a reactive component, such as an initiating template nucleic acid or polymerase, can be attached to a solid component via covalent or non-covalent bonds. Covalent bonds are characterized by the sharing of electron pairs between atoms. Non-covalent bonds are chemical bonds that do not involve shared electron pairs and can include, for example, hydrogen bonds, ionic bonds, van der Waals forces, hydrophilic interactions, and hydrophobic interactions.
[0048] As used herein, the term "closing portion," when used to refer to a nucleotide, means a nucleotide portion that inhibits or prevents the 3' oxygen of the nucleotide from forming a covalent bond with the next correct nucleotide during nucleic acid polymerization. The closing portion of a "reversibly terminated" nucleotide may be removed from a nucleotide analogue or otherwise modified to allow the 3'-oxygen of the nucleotide to covalently attach to the next correct nucleotide. Such a closing portion is referred to herein as a "reversible terminator portion." Exemplary reversible terminator portions are described in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference. A nucleotide having a closing portion or a reversible terminator portion may be located at the 3' end of a nucleic acid (such as a primer), or the nucleotide may be a monomer not covalently attached to a nucleic acid. The closing portion does not need to prevent or exclude the formation of the ternary complex at the 3' end of the nucleic acid to which the closing portion is attached. Particularly useful closing portions will be present at the 3' end of the nucleic acid involved in the formation of the ternary complex.
[0049] As used herein, the term "capped portion," when referring to nucleic acids, means, when present in nucleic acids, a portion that prevents or inhibits the binding of the 3' end of the nucleic acid to polymerase and the next correct nucleotide to form a ternary complex. Portions that generate steric blocks that prevent ternary complex formation are particularly useful and include, for example, polymeric or ligation products that extend a primer to the end of a template that hybridizes with the primer. Another example of a steric block is a mismatched nucleotide. Capped portions can have a positive or negative charge that prevents or inhibits ternary complex formation. Capped portions can contain ligands that bind to receptors to prevent or inhibit ternary complex formation, such as biotin (or its analogues) that binds to streptavidin (or a similar protein), an epitope that binds to an antibody (or a functional fragment thereof), a carbohydrate that binds to a lectin, etc. Therefore, a ternary complex inhibitor can be a ligand-receptor complex that inhibits ternary complex formation. Further examples of portions that can be used as inhibitors of ternary complexes include base modifications and nucleotide analogs described in U.S. Patent Application Publication No. 2020 / 0032322A1 or Turcatti et al. Nucl. Acids. Res. 36(4)e25(2008), each of which is incorporated herein by reference.
[0050] As used herein, the term "circular," when referring to a nucleic acid strand, means that the strand lacks ends (i.e., it lacks both 3' and 5' ends). Therefore, the 3' oxygen and 5' phosphate portions of each nucleotide monomer in a circular strand are covalently attached to adjacent nucleotide monomers in the strand. Circular DNA strands can serve as templates for generating tandem amplicones via rolling circle amplification (RCA), where each sequence unit of the tandem amplicon is the reverse complement of the circular nucleic acid strand. Circular nucleic acids can be double-stranded. One or both strands of a double-stranded nucleic acid may lack both 3' and 5' ends. One strand of a double-stranded nucleic acid may have a nick (the absence of at least one nucleotide monomer relative to the other strand) or a cleavage (the absence of a phosphodiester bond between two nucleotide monomers), provided that the other strand is circular.
[0051] As used herein, the term "cluster," when referring to nucleic acids, refers to a group of nucleic acids attached to a solid support, for example, at a site in a site array on a solid support. The term "clonal population" refers to a group of nucleic acids that are homogeneous relative to a particular nucleic acid sequence. Homogeneous sequences are typically at least 10 nucleotides long, but can be even longer, including, for example, at least 50, 100, 500, 1000, or 2500 nucleotides long. Clonal populations can originate from a single template nucleic acid. Clonal populations can include at least 2, 10, 100, 1000, or more copies of a particular nucleic acid sequence. Copies can be present in a single nucleic acid molecule, for example, as tandem copies, or copies can be present on separate nucleic acid molecules. Typically, all nucleic acids in a cluster will have the same nucleotide sequence. It should be understood that negligible amounts of contaminating nucleic acids or mutations (e.g., due to amplification artifacts) can appear in a cluster without departing from obvious clonality. Clusters can be at least 80%, 90%, 95%, or 99% cloned. Optionally, the cluster can be 100% cloned.
[0052] As used herein, the term "common sequence" refers to a nucleotide sequence that is identical to that of two or more nucleic acid molecules. The common sequence of two or more nucleic acids may include all or part of the nucleic acids being compared. The common sequence may have a length of at least 5, 10, 25, 50, 100, 250, 500, 1000, or more nucleotides. Optionally or additionally, the length may be up to 1000, 500, 250, 100, 50, 25, 10, or 5 nucleotides. A population of nucleic acid molecules may include individual molecules having common sequence regions (e.g., "universal primers" or "universal primer binding sites") and variable sequence regions (e.g., "target regions") that differ from one individual to another.
[0053] As used herein, the term "tandem" when referring to a nucleic acid molecule means a continuous nucleic acid molecule containing more than one copy of a common sequence linked in tandem. Similarly, the term "tandem" when referring to a nucleotide sequence means a continuous nucleotide sequence containing more than one copy of a common sequence linked in tandem. Each copy of the sequence can be called a "sequence unit" of the tandem. Sequence units can have a length of at least 10, 50, 100, 250, 500, or more bases. A tandem can contain at least 2, 5, 10, 50, 100, or more sequence units. Sequence units can contain subregions with any of a variety of functions, such as primer-binding regions, target sequence regions, tag regions, unique molecular identifiers (UMIs), etc.
[0054] The term “includes” is intended to be open-ended in this document, including not only the elements mentioned, but also any other elements.
[0055] As used herein, the term "cycle," when referring to a sequencing procedure, refers to a portion of the sequencing run that is repeated to indicate the presence of nucleotides. Typically, a cycle includes several steps, such as steps for delivering reagents, washing away unreacted reagents, and detecting signals indicating changes that occur in response to added reagents.
[0056] As used herein, the term "unblocking" means the removal or modification of the reversible terminator portion of a nucleotide to make the nucleotide extendable. For example, a nucleotide may be present at the 3' end of a primer, such that unblocking makes the primer extendable. Exemplary unblocking reagents and methods are set forth in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference.
[0057] As used herein, the term "exogenous," when referring to a portion of a molecule, means a chemical portion that is not present in the molecule's natural analogues. For example, an exogenous label on a nucleotide is a label that is not present on naturally occurring nucleotides. Similarly, an exogenous label present on a polymerase is not found on polymerases in their natural environment.
[0058] As used herein, the term "extension," when referring to nucleic acids, means the process of adding at least one nucleotide to the 3' end of a nucleic acid. The term "polymerase extension," when referring to nucleic acids, refers to the polymerase-catalyzed process of adding at least one nucleotide to the 3' end of a nucleic acid. The nucleotide or oligonucleotide added to a nucleic acid through extension is referred to as incorporation into the nucleic acid. Therefore, the term "incorporation" can be used to refer to the process of attaching a nucleotide or oligonucleotide to the 3' end of a nucleic acid by forming a phosphodiester bond.
[0059] As used herein, the term "extensible," when referring to a nucleotide, means that the nucleotide has an oxygen or hydroxyl moiety at the 3' position and is capable of covalently linking to the next correct nucleotide. Extensible nucleotides can be located at the 3' position of a polymeric nucleic acid or can be monomeric nucleotides. Extensible nucleotides will lack a closing moiety, such as a reversible terminator moiety.
[0060] As used herein, the term "fixed" when referring to a molecule means that the molecule is directly or indirectly, covalently or non-covalently attached to a solid support. In some configurations, covalent attachment may be preferred, but generally what is required is that the molecule (e.g., nucleic acid) remains fixed or attached to the support when the support is intended to be used, such as when sequencing nucleic acids fixed to an array site or another solid support.
[0061] As used herein, the term "label" refers to a molecule or portion thereof that provides a detectable characteristic. Detectable characteristics can be, for example, optical signals such as absorbance of radiation, fluorescence emission, luminescence emission, fluorescence lifetime, fluorescence polarization, etc.; Rayleigh and / or Mie scattering; binding affinity to a ligand or acceptor; magnetic properties; electrical properties; charge; mass; radioactivity, etc. Exemplary labels include, but are not limited to, fluorophores, chromophores, nanoparticles (e.g., gold, silver, carbon nanotubes), heavy atoms, radioactive isotopes, mass labels, charge labels, spin labels, acceptors, ligands, etc.
[0062] As used herein, the term "next correct nucleotide" refers to a nucleotide or type of nucleotide that binds to and / or is incorporated into the 3' end of a primer to be complementary to a base in the template strand that the primer hybridizes to. The base in the template strand is called the "next base" and is the base in the template immediately adjacent to the 5' end of the primer that hybridizes to the 3' end. The next correct nucleotide can be called a "cognate" of the next base, and vice versa. Cognate nucleotides that interact with each other in a ternary complex or double-stranded nucleic acid are called "paired." Nucleotides with bases that are not complementary to the next template base are called "incorrect," "mismatched," or "non-homologous" nucleotides.
[0063] As used herein, the term "non-catalytic metal ion" refers to a metal ion that, in the presence of a polymerase, does not promote the formation of the phosphodiester bonds required for the chemical incorporation of nucleotides into primers. Non-catalytic metal ions can interact with polymerases, for example, through competitive binding compared to catalytic metal ions. Therefore, non-catalytic metal ions can act as inhibitory metal ions. A "divalent non-catalytic metal ion" is a non-catalytic metal ion having a divalent oxidation state. Examples of divalent non-catalytic metal ions include, but are not limited to, Ca. 2+ Zn 2+ Co 2+ Ni 2+ and Sr 2+ Trivalent Eu 3+ and Tb 3+ The ion is a non-catalytic metal ion with a trivalent oxidation state.
[0064] As used herein, the term "nucleotide" may refer to natural nucleotides or their analogues. Examples include, but are not limited to, nucleoside monophosphates, nucleoside diphosphates, and nucleoside triphosphates (NTPs) such as ribonucleoside triphosphates (rNTPs), deoxyribonucleoside triphosphates (dNTPs), or their non-natural analogues such as dideoxynucleoside triphosphates (ddNTPs) or reversibly terminated nucleoside triphosphates (rtNTPs).
[0065] As used herein, the term "polymerase" can be used to refer to an enzyme that synthesizes nucleic acids, including but not limited to DNA polymerases, RNA polymerases, reverse transcriptases, primases, and transferases. Typically, a polymerase has one or more active sites at which nucleotide binding and / or nucleotide polymerization can be catalyzed. A polymerase catalyzes the polymerization of a nucleotide to the 3' end of the first strand of a double-stranded nucleic acid molecule. For example, a polymerase catalyzes the addition of the next correct nucleotide to the 3' oxygen group of the first strand of a double-stranded nucleic acid molecule via a phosphodiester bond, thereby covalently incorporating a nucleotide into the first strand of the double-stranded nucleic acid molecule. Optionally, a polymerase does not need to be able to perform nucleotide incorporation under one or more of the conditions used in the methods described herein. For example, a mutant polymerase may be able to form a ternary complex but cannot catalyze nucleotide incorporation. A polymerase may have strand substitution activity, such as Phi29. A polymerase may lack strand substitution activity. A polymerase may have 5'→3' exonuclease activity. A polymerase may lack 5'→3' exonuclease activity.
[0066] As used herein, the term "primer" refers to a nucleic acid having a sequence that binds to a nucleic acid at or near a template sequence. Typically, primers bind in a conformation that allows, for example, a polymerase to extend and replicate the template. A primer may be a first portion of a nucleic acid molecule that binds to a second portion of the nucleic acid molecule, the first portion being the primer sequence and the second portion being the primer-binding sequence (e.g., a hairpin primer). Alternatively, a primer may be a first nucleic acid molecule that binds to a second nucleic acid molecule having the template sequence. Primers may include or may be DNA, RNA, or analogs thereof. Primers may have an extendable 3' end, a blocked 3' end that prevents primer extension, or a capped 3' end to prevent or exclude the formation of a ternary complex. In some embodiments, primers may contain nucleotide modifications not at the 5' and 3' ends. Optionally or additionally, primers may contain modified nucleotides at the 5' end.
[0067] As used herein, the term "primer-template nucleic acid hybrid" or "primer-template hybrid" refers to a nucleic acid having a double-stranded region such that one strand is the primer and the other is the template. The two strands can be part of a continuous nucleic acid molecule (e.g., a hairpin structure), or the two strands can be separable molecules that are not covalently attached to each other.
[0068] The terms “sense” and “antense” are used in this paper to distinguish members of a pair of complementary nucleic acid molecules or sequences. These terms are intended to serve as context-specific identifiers. Based on their use in the field of molecular biology, these terms are interchangeable. Thus, a strand identified as a “sense” strand in one context can be referred to as an “antense” strand in a second context. This is unrelated to how similar or different the first and second contexts are.
[0069] As used herein, the term "site," when referring to an array, means the location within the array where a specific molecule is present. A site may contain only a single molecule or a group of molecules of the same kind (an ensemble of molecules). Optionally, a site may contain groups of molecules of different kinds (e.g., a group of ternary complexes with different template sequences). Sites in an array are typically discrete. Discrete sites may be adjacent or may have gaps between them. Arrays used herein may have sites spaced, for example, less than 100 micrometers, 50 micrometers, 10 micrometers, 5 micrometers, 1 micrometer, or 0.5 micrometers apart. Optionally or additionally, an array may have sites spaced greater than 0.5 micrometers, 1 micrometer, 5 micrometers, 10 micrometers, 50 micrometers, or 100 micrometers apart. Each of these sites may have an area of less than 1 square millimeter, 500 square micrometers, 100 square micrometers, 25 square micrometers, 1 square micrometer, or smaller. Sites may also be referred to as "features" of the array.
[0070] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquids. The substrate may be non-porous or porous. The substrate may optionally be able to absorb liquids (e.g., due to porosity), but typically has sufficient rigidity such that the substrate does not substantially swell when absorbing liquids and does not substantially shrink when the liquid is removed by drying. Non-porous solid supports are typically impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene with other materials, polypropylene, polyethylene, polybutene, polyurethane, Teflon, etc.). TM Materials include cycloolefins, polyimides, nylon, ceramics, resins, Zeonor, silica or silica-based materials (including silicon and modified silicon), carbon, metals, inorganic glasses, fiber bundles, and polymers.
[0071] As used herein, the term "template" means a nucleic acid or a portion thereof having a nucleotide base sequence that serves as a guide for producing a complementary copy of that sequence. The template can be replicated by extending a primer that hybridizes to or is adjacent to the template. Extension can be mediated by a polymerase or ligase. The template may include or may be DNA, RNA, or analogues thereof.
[0072] As used herein, the term "ternary complex" refers to the intermolecular association between a polymerase, a double-stranded nucleic acid, and a nucleotide. Typically, the polymerase promotes the interaction between the next correct nucleotide and the template strand of the initiating nucleic acid. The next correct nucleotide interacts with the template strand via Watson-Crick hydrogen bonds. The term "stabilized ternary complex" refers to a ternary complex with the presence of either a promoting or elongating effect, or a ternary complex whose disruption has been inhibited. Typically, stabilization of a ternary complex prevents the covalent incorporation of the nucleotide component of the ternary complex into the initiating nucleic acid component of the ternary complex.
[0073] The embodiments shown below and cited in the claims can be understood in accordance with the above definitions.
[0074] This disclosure provides a method for determining a sequence from the sense and antisense strands of a nucleic acid. For example, the sense and antisense strands may be a first strand and a second strand, respectively. The method may include the steps of: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a tandem sense strand and a tandem antisense strand, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand in the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand.
[0075] Nucleic acids used in the methods or compositions described herein can be DNA, such as genomic DNA, synthetic DNA, amplified DNA, complementary DNA (cDNA), etc. RNA, such as mRNA, ribosomal RNA, tRNA, etc., can also be used. Nucleic acid analogs can also be used in the methods or compositions described herein. For example, nucleic acid analogs can be used as templates for the amplification or sequencing processes described herein. The nucleic acids used herein, for example, as templates for generating nucleic acid clusters or as targets for sequencing, can be of biological origin, synthetic origin, or amplification products. Primers used herein can include or can be DNA, RNA, or analogs thereof.
[0076] The nucleic acid template containing the target sequence used in the sequencing methods described in this disclosure can be derived from or generated from a sample. The sample contains one or more organisms. The nucleic acid template can be obtained or derived from the sample without performing a polymerase chain reaction. The nucleic acid template can be obtained or derived from the sample by performing a polymerase chain reaction for several cycles, such as up to one cycle, two cycles, three cycles, four cycles, five cycles, six cycles, seven cycles, eight cycles, nine cycles, or ten cycles or more.
[0077] This paper envisions nucleic acid templates (or target sequences) of different lengths. The length of the nucleic acid template can be the following, approximately the following, at least the following, at least approximately the following, at most the following, or at most the following nucleotides: 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490. 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, or a quantity or range between any two of these values.
[0078] Exemplary organisms from which nucleic acids can be derived include, for example, mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cattle, cats, dogs, primates, humans, or non-human primates; plants such as Arabidopsis thaliana, maize, sorghum, oats, wheat, rice, rapeseed, or soybeans; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, bees, or spiders; fish such as zebrafish; reptiles; amphibians such as frogs or Xenopus laevis; dictyostelium discoideum; and fungi such as Pneumocystis carinii and Takifugu. Nucleic acids can originate from prokaryotes, such as bacteria like Escherichia coli, Staphylococcus aureus, or Mycoplasma pneumoniae; archaea; viruses such as hepatitis C virus or human immunodeficiency virus; or viroids. Nucleic acids can be derived from homogeneous cultures or populations of the above organisms, or optionally from a collection of several different organisms in a community or ecosystem. Nucleic acids can be isolated using methods known in the art, including, for example, those described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory, New York (2001) or Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference.
[0079] Nucleic acids can be obtained from preparative methods such as genomic, transcriptomic, or other nucleic acid isolation, genome fragmentation, gene cloning, and / or amplification. One or more nucleic acids can be obtained from amplification techniques such as polymerase chain reaction (PCR), emulsion PCR, random priming amplification, rolling circle amplification (RCA), multiple substitution amplification (MDA), etc. RCA and MDA are particularly suitable for producing tandem products. Exemplary methods for isolating, amplifying, and fragmenting nucleic acids to produce templates for analysis on an array are set forth in U.S. Patent Nos. 6,355,431 or 9,045,796, each of which is incorporated herein by reference. Amplification can also be performed using methods set forth in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory, New York (2001) or Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference.
[0080] Nucleic acid clusters can contain one or more tandem nucleic acid chains. For example, a cluster can contain only a single tandem nucleic acid chain. A single tandem chain can be generated through an RCA reaction, for example, as in... Figure 1B As illustrated in the diagram. Optionally, the nucleic acid cluster may comprise a first strand (e.g., sense strand) as a tandem strand and one or more second strands (e.g., antisense strands) complementary to the first strand. For example, one or more second strands can be generated by multiple substitution amplification (MDA) of the tandem template. See, for example, by [example from...] Figure 1C or Figure 1D The method illustrated in the diagram produces double-stranded clusters. For ease of reference, following convention used in molecular biology, one of the two complementary strands may be referred to as the "sense" strand, and the complement of the sense strand may be referred to as the "antense" strand.
[0081] The tandem strand in a cluster can contain more than one copy of tandemly linked sequence units. For example, the tandem strand can contain at least 2, 10, 25, 100, or more sequence units. The number of sequence units in the tandem strand can be, for example, up to 100, 25, 10, or 2 sequence units. The number of sequence units in the tandem strand produced by RCA will vary depending on the number of times the polymerase completes one revolution around the circular template during replication. The content of each sequence unit produced by RCA will be the inverse complement of the content of the replicated circular template. For example, as in... Figure 1BAs shown, the circular template comprises two connective subregions (indicated by hollow rectangles and cross-shaded rectangles) and a target region (indicated by dashed lines). The connective subregions can have any of a variety of functions, including, but not limited to, providing a binding site complementary to a capture probe (e.g., a capture probe attached to a solid support), providing a primer binding site for replicating the circular template, providing a primer binding site for replicating a complement to the circular template, providing a tag for association with the target region (e.g., a tag indicating the origin of the target region or a tag for identifying errors introduced during target region amplification, etc.). The connective subregions or portions thereof may be common to a population of circular templates or a population of tandems. Regardless of whether the connective subregions have a common sequence, the target regions in a population of circular templates or in a population of tandems can have different sequences. Therefore, when comparing sequence units between two or more tandems or between two or more circular templates, the sequence units may have common sequence regions (e.g., universal primer binding sites or universal capture probe binding sites) and / or the sequence units may have regions with different sequences (e.g., different target sequences).
[0082] The length of the sequence units in the tandem or the length of the ring template can be selected to suit the specific application of the method described herein. For example, the length can be at least about 50, 100, 250, 500, 1000, or 1 x 10^ ... 4 1 x 10 5 One or more nucleotides. Optionally or additionally, the length may not exceed 1 x 10^6 nucleotides. 5 1 x 10 4 1, 1000, 500, 250, 100, or 50 nucleotides. It should be understood that the above length ranges apply to the target region of the sequence unit or the target region of the circular template, but do not include adaptor sequences (e.g., common or universal adaptor sequences) that may also be present in the sequence unit. For example, length ranges may describe the size of genomic fragments or other nucleic acid fragments used to generate clusters or otherwise exist within clusters.
[0083] A cluster may contain one or more chains of concatenation. In some configurations, a cluster may contain no more than one chain of concatenation. Optionally, a cluster may contain more than one chain of concatenation, for example, at least 2, 4, 10, 50, 100 or more meaningful chains containing concatenation. Optionally or additionally, the number of chains of concatenation in a cluster may be, for example, at most 100, 50, 10, 4, 2 or 1 meaningful chains containing concatenation. In some implementations, the number of tandem chains can be: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 500 The number or range of 0, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 100000, or any two of these values. Chains of tandem sequences within a particular cluster can have the same target sequence, for example, sense chains that are part of the same tandem sequence. Optionally, a cluster can have more than one distinct chain of tandem sequences; such clusters are non-clonal.
[0084] A cluster containing at least one sense chain of a tandem may also contain at least one antisense chain of the tandem. For example, a cluster containing one or more sense chains of a tandem may also contain more than one antisense chain. More than one antisense chain may include at least 2, 4, 10, 50, 100 or more antisense chains for a particular tandem. Optionally or additionally, the number of antisense chains in a cluster may be, for example, at most 100, 50, 10, 4, 2 or 1 antisense chains for a particular tandem. In some implementations, the number of antisense chains can be: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 The number or range of entries, including 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 100000, or any two of these values. In a specific configuration, a cluster may contain a single meaningful concatenation chain (no more than one meaningful chain) and a single antisense concatenation chain (no more than one antisense chain). Optionally, a cluster may contain at least one meaningful chain of a concatenation and more than one antisense chain of a concatenation. The number of antisense chains in a cluster can exceed the number of sense chains in the cluster. Optionally, the number of sense chains in a cluster can exceed the number of antisense chains in the cluster. Note that the antisense chains of a tandem aggregate do not need to be the same length as the sense chains. For example, the antisense chain can have more sequence units than the sense chain, or the antisense chain can have fewer sequence units than the sense chain. The number of sequence units in the antisense chain can fall within the range described herein for the sense chains of a tandem aggregate. The antisense chain of a tandem aggregate does not need to have more than one sequence unit. In fact, the antisense chain does not need to have a complete sequence unit.
[0085] A cluster containing at least one sense strand of a tandem hybrid may also contain at least one antisense strand of the tandem hybridizing with the sense strand via Watson-Crick base pairing. For example, the sense strand of a particular tandem hybridizes with at least 2, 4, 10, 50, 100 or more antisense strands of that particular tandem hybridizes. Optionally or additionally, the sense strand of a particular tandem hybridizes with, for example, up to 100, 50, 10, 4, 2 or 1 antisense strand of that particular tandem hybridizes.
[0086] Nucleic acid clusters can be attached to a solid support. The solid support can be made from any of a variety of materials used in analytical biochemistry. Suitable materials may include, for example, glass, polymer materials, silicon, quartz (fused silica), borofloat glass, silica, silica-based materials, carbon, metals, optical fibers or fiber bundles, sapphire, or plastic materials. Materials can be selected based on the properties required for a specific application. For example, materials that are transparent to a desired wavelength of radiation are useful for analytical techniques utilizing that wavelength. Conversely, it may be desirable to select materials that do not allow a certain wavelength of radiation to pass through (e.g., opaque, absorbing, or reflecting). Wavelength regions that may or may not pass through a particular material include, for example, UV, VIS (e.g., red, yellow, green, or blue), or IR. Other properties that can be utilized in materials include inertness or reactivity to certain reagents used in downstream processes (such as those described herein), ease of handling, or low manufacturing cost.
[0087] A particularly useful solid support is particles, such as beads or microspheres. Populations of beads can be used for the attachment of nucleic acid populations. In some embodiments, it may be useful to use a configuration where each bead has a single target sequence. A single bead may have a single nucleic acid molecule with the target sequence, or alternatively, a single bead may have more than one nucleic acid molecule, each of which has the target sequence. In some configurations, beads may be attached to a nucleic acid tandem, where more than one copy of the target sequence is present in a single nucleic acid molecule. Beads in a population may have different target sequences from one another. When compared with each other, beads in a population may have a common nucleic acid sequence. For example, a population of beads may be attached to universal primers such that the same primer sequence is present on more than one bead in the population.
[0088] The composition of the beads can vary, for example, depending on the form, chemical, and / or attachment method to be used. Exemplary bead compositions include a solid support, optionally incorporating chemical functionalities used in protein and nucleic acid capture methods. These compositions include, for example, plastics, ceramics, glass, polystyrene, melamine, methyl styrene, acrylic polymers, paramagnetic materials, thorium dioxide sol, carbon graphite, titanium dioxide, latex, or cross-linked dextran such as Sepharose. TM Cellulose, nylon, cross-linked micelles and Teflon TM This is in addition to other materials described in the "Microsphere Detection Guide" from Bangs Laboratories, FishersInd. (which is incorporated herein by reference).
[0089] The geometry of the particles, such as beads or microspheres, can also correspond to a variety of different forms and shapes. For example, particles can be symmetrical shapes (e.g., spherical or cylindrical) or irregular shapes (e.g., controlled-pore glass). Furthermore, the particles can be porous, thus increasing the surface area available for trapping ternary composites or their components. Exemplary sizes of the beads used herein can range from nanometers to millimeters or from about 10 nm to about 1 mm.
[0090] In certain implementations, the beads may be arranged or otherwise spatially differentiated. Exemplary bead-based arrays that may be used include, but are not limited to, the BeadChip available from Illumina, Inc. (San Diego, CA). TM Or arrays such as those described in: U.S. Patent Nos. 6,266,459; 6,355,431; 6,770,441; 6,859,570; or 7,622,294 or PCT Publication No. WO 00 / 63437, each of which is incorporated herein by reference. The beads may be located at discrete positions on a solid support, such as holes, whereby each position accommodates a single bead. Alternatively, each discrete position where a bead resides may each include more than one bead, for example, as described in: U.S. Patent Application Publications Nos. 2004 / 0263923A1, 2004 / 0233485A1, 2004 / 0132205A1, or 2004 / 0125424A1, each of which is incorporated herein by reference.
[0091] As will be appreciated from the above bead array implementations, the methods of this disclosure can be implemented in multiple forms, thereby enabling the parallel detection of more than one different type of nucleic acid using one or more steps of the methods described herein. Other types of arrays can be used instead of bead arrays, including, for example, those described in further detail below. While it is possible to process different types of nucleic acids sequentially using one or more steps of the methods described herein, parallel processing can provide cost savings, time savings, and condition consistency. The arrays or methods of this disclosure can be configured to include at least 2, 10, 100, or 1x10⁻⁶. 3 1x10 4 1x10 5 1x10 6 1x10 9 Or more different nucleic acids. Optionally or additionally, the array or method of this disclosure can be configured to include up to 1x10 9 1x10 6 1x10 5 1x10 4 1x10 3100, 10, 2, or fewer different nucleic acids. Nucleic acids can be attached to different sites in the array. Therefore, the number of sites in the array can be within the range of different nucleic acids illustrated here. Furthermore, the various reagents or products described herein (e.g., primer-template nucleic acid hybrids or stabilized ternary complexes) can be multiplexed to have different types or kinds within these ranges.
[0092] Further examples of commercially available arrays that can be used in the methods or compositions described herein include arrays synthesized via photolithography of nucleic acids, such as the Affymetrix GeneChip. TM Arrays. According to some implementations, dotted arrays can also be used to attach pre-synthesized nucleic acids to array sites. Exemplary dotted arrays are available from CodeLink, Amersham Biosciences. TM Arrays. Another useful type of array is one manufactured using inkjet printing methods, such as SurePrint, available from Agilent Technologies. TM The techniques used to attach nucleic acid probes to these arrays can be modified to attach nucleic acid primers for amplification (e.g., via RCA) and / or sequencing of target nucleic acids that hybridize with the primers.
[0093] Other useful arrays include those used in nucleic acid sequencing applications. For example, methods and compositions for attaching amplicon to genomic fragments (often called clusters) to form arrays can be particularly useful. Examples are described in Bentley et al., Nature 456:53-59 (2008), PCT Publication WO91 / 06678; WO 04 / 018497 or WO 07 / 123744; U.S. Patent Nos. 7,057,026; 7,211,414; 7,315,019; 7,329,492 or 7,405,281; or U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.
[0094] Nucleic acids can be attached to solid supports (e.g., sites on an array) via covalent or non-covalent bonds. For example, a solid support can be covalently or non-covalently attached to or near the 5' end of a tandem nucleic acid. This configuration may occur, for example, when the tandem has been generated by an RCA (Reactive Carbon Acetate) using primers that extend to or near their 5' ends and are attached to the solid support. The attachment of nucleic acids to solid supports can be mediated by any of a variety of surface chemistry methods, such as the reaction of a carboxylic acid ester or succinimidyl ester moiety on the solid support with an amine-modified nucleic acid, the reaction of an alkylating agent (e.g., iodoacetamide or maleimide) on the solid support with a thiol-modified nucleic acid, the reaction of a siloxane or isothiocyanate-modified solid support with an amine-modified nucleic acid, the reaction of an aminophenyl or aminopropyl-modified solid support with a succinylated nucleic acid, the reaction of an aldehyde or epoxide-modified solid support with an acylhydrazine-modified nucleic acid, or the reaction of a thiol-modified solid support with a thiol-modified nucleic acid. The members of the aforementioned reaction pairs can be switched depending on whether they are present on the solid support or the nucleic acid. Click chemistry can be used to attach nucleic acids to solid supports. Exemplary reagents and methods for use in click chemistry are set forth in U.S. Patents 6,737,236, 7,375,234, 7,427,678, and 7,763,736, each of which is incorporated herein by reference.
[0095] In certain embodiments, a stabilized ternary complex, polymerase, primer, template, primer-template nucleic acid hybrid, or nucleotide is attached to the surface of the flow cell or to a solid support within the flow cell. One or more nucleic acid clusters may be attached to the surface of the flow cell or to a solid support within the flow cell. The flow cell allows for convenient fluid manipulation by delivering a solution into and out of a fluid chamber in contact with the analyte bound to the support. The flow cell also provides detection of the fluid manipulation components. Detectors can be positioned to detect signals from the solid support, such as signals from tags recruited to the solid support during sequencing, for example, due to the formation of the stabilized ternary complex. Exemplary flow cells that may be used are described, for example, in U.S. Patent Application Publication No. 2010 / 0111768A1, WO 05 / 065814, or U.S. Patent Application Publication No. 2012 / 0270305A1, each of which is incorporated herein by reference.
[0096] In certain configurations, nucleic acid clusters can be attached to a solid support by covalently attaching one or both strands to it. The nucleic acid can be single-stranded or double-stranded. In some configurations, the first strand of a double-stranded nucleic acid is covalently attached to the solid support, while the second strand is not covalently attached. For example, the second strand may be retained in the cluster due to Watson-Crick base pairing with the first strand. Clusters formed by covalently attaching the sense strand of a tandem nucleic acid to a solid support can retain one or more antisense strands by effectively functioning a region of more than one base pairing of strand entanglement within the cluster.
[0097] Nucleic acid clusters may be prepared prior to or as part of the methods described herein. In a particular configuration, a nucleic acid cluster may comprise a nucleic acid tandem strand. A cluster may consist of a single strand of the tandem strand. A cluster does not need to contain an antisense strand of the tandem strand, nor does it need to contain an antisense strand of any region of the tandem strand. Optionally, a cluster may comprise a sense strand of the tandem strand and at least one antisense strand complementary to all or part of the sense strand.
[0098] A useful method for generating tandem nucleic acids on a solid support is rolling circle amplification (RCA). Typically, the method involves polymerase extending and annealing primers to a circular template, causing the polymerase to rotate more than one turn around the template to produce a tandem single-stranded DNA containing more than one tandem repeat, each repeat complementary to the circular template. In one configuration, RCA can be initially performed in the presence of a low concentration of a polymer such as a dendritic macromolecule (e.g., polyamidoamine (PAMAM)) and subsequently in the presence of the polymer. In one configuration, the RCA reaction is stopped by denaturing the polymerase, for example, by heating the sample at 60°C, 65°C, 70°C, 75°C, 80°C, or higher. In one configuration, the RCA reaction is stopped by removing one or more components of the RCA, such as the polymerase and dNTPs. Components of RCA can be removed, for example, by washing. Optionally, one or more antisense strands can be produced by replicating the sense strand of the tandem, for example, using multiple substitution amplification (MDA). Both RCA and MDA methods can be performed isothermally. Typically, the polymerase used for RCA or MDA is a chain displacement polymerase. Methods and reagents that can be used for RCA, MDA, or combinations thereof are described, for example, in: Lizardi et al., Nat. Genet. 19:225-232 (1998), U.S. Patent Nos. 6,830,884; 6,797,474; 6,670,126; 6,576,448; 6,323,009; 6,280,949 or US 2007 / 0099208 A1, each of which is incorporated herein by reference.
[0099] Figures 1A-1DA schematic diagram is provided for a method of generating immobilized tandem nucleic acid clusters on a solid support. (As shown in...) Figure 1A As shown, nucleic acid primers (indicated by hollow and lined rectangles) are attached to a solid support (indicated by dashed rectangles) via adapters (indicated by gray lines). Primers can be used to capture target nucleic acids through primer binding sites complementary to the primers. Figure 1A In one configuration shown, the fixed primers can hybridize partially with primer-binding sites located at opposite ends of the target sequence (the target sequence is indicated by dashed lines, and the flanking primer-binding site regions are indicated by hollow and lined rectangles, respectively). Therefore, the fixed primers act as a clamp to hold the two ends of the target nucleic acid together. The two ends can ligate while hybridizing with the clamp nucleic acid to form a circular form of the target nucleic acid. Figure 1A In another configuration shown, the target nucleic acid is circularized before hybridization with primers immobilized on a solid support. For example, the linear target nucleic acid can be ligated simultaneously with hybridization with splice oligonucleotides in solution, or simultaneously with hybridization with splice oligonucleotides on a solid support other than the solid support used for RCA. Alternatively, when using, for example, Circligase... TM When using enzymes such as Epicenter, Madison WI, or others capable of splinting nucleic acid ends, splinting is not required. Similarly, immobilization can occur due to the complementarity between the immobilization primer and the primer binding site in the circular template.
[0100] Figure 1B A schematic diagram is provided showing the production of single-stranded tandem strands via rolling circle amplification of a circular template induced by hybridization with immobilized primers. Primers are immobilized with their 3' ends available for polymerase extension (e.g., primers may be attached at or near the 5' end). The product of the first sub-step is shown as having progressed to the point where two copies of the circular template (two sequence units) have been produced, and the circular template has hybridized with a portion of a replicating third copy (the third sequence unit). Each sequence unit contains a region complementary to the target sequence (indicated by a solid black line) and a region complementary to the primer (indicated by a hollow and lined rectangle). The product of the second sub-step has progressed to the point where nearly six copies of the circular template have been produced. Figure 1B The final product of the RCA reaction is shown after the absence (e.g., removal) of the cyclic template in the third sub-step. For illustrative purposes, two regions of the final product are shown: a region depicting the sequence units (indicating the tandem primary structure of the sense strand), and a region where the number and conformation of the sequence units are not specified (indicating the dynamic and variable secondary structure of the entire cluster). Figures 1A-1D The diagram illustrates that the steps described above can be performed in multiple ways, so that the steps in the diagram occur at more than one individual site in the array.
[0101] In a specific configuration, nucleic acid clusters can be prepared by: (i) providing a sense strand in a tandem on a solid support, and (ii) synthesizing an antisense strand using amplification primers that bind to primer binding sites in the sense strand.
[0102] Therefore, this disclosure provides a method for determining a sequence from the sense and antisense strands of a nucleic acid, comprising the steps of: (a) (i) providing a tandem sense strand on a solid support; and (ii) synthesizing an antisense strand of the tandem strand using amplification primers that bind to primer binding sites in sequence units of the sense strand, thereby providing a nucleic acid cluster attached to the solid support, wherein the nucleic acid cluster comprises a tandem sense strand and a tandem antisense strand, wherein the tandem strand comprises more than one copy of tandemly linked sequence units, wherein the sequence units comprise a target sequence and a primer binding site; (b) hybridizing primers to primer binding sites in sequence units of the antisense strand within the cluster; (c) extending primers along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to primer binding sites in sequence units of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand.
[0103] It can generate one or more antisense chains, such as in Figure 1C As described in the text. Amplification can begin with a tandem sense strand (indicated by black lines) fixed to a solid support (indicated by dotted rectangles) via a adapter (indicated by gray lines). The sense strand of the tandem strand contains sequence units, each containing a primer binding site. Amplification primers complementary to all or part of the primer binding sites can hybridize with the sense strand of the tandem strand and extend in the MDA reaction to produce an antisense strand (shown as gray lines at least partially annealed to the sense strand). For illustrative purposes, a portion of the MDA product is shown to indicate the direction of extension (see arrows) and to indicate the primary structure of the tandem antisense strand with various lengths and annealing modes. Another portion of the MDA product is shown with less structure specificity to indicate the dynamic and variable secondary structure of the cluster as a whole.
[0104] Nucleic acid clusters may be prepared prior to or as part of the methods described herein, for example by (i) providing a solid support with capture primers, (ii) hybridizing a circular nucleic acid template with primers, (iii) extending the capture primers along the circular nucleic acid template by rolling circle amplification to synthesize the sense strand of the tandem, and (iv) synthesizing the antisense strand of the tandem by extending amplification primers that bind to primer binding sites in the sequence units of the sense strand.
[0105] Therefore, this disclosure provides a method for determining sequences from the sense and antisense strands of nucleic acids, comprising the steps of: (a)(i) providing a solid support having capture primers, (ii) hybridizing a circular nucleic acid template with primers, (iii) synthesizing a sense strand of a tandem by extending the capture primers along the circular nucleic acid template via rolling circle amplification, and (iv) synthesizing an antisense strand of the tandem by extending amplification primers that bind to primer binding sites in the sequence units of the sense strand, thereby providing a nucleic acid cluster attached to the solid support, wherein the nucleic acid cluster contains The tandem strand comprises a sense strand and an antisense strand, wherein the tandem strand contains more than one copy of tandemly linked sequence units, wherein each sequence unit contains a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand in the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand.
[0106] The sense and antisense chains of a tandem chain can be like... Figure 1D The process described is as follows. Optionally, the annular template can be trapped on a solid support, such as in... Figure 1A The diagram is provided. Regardless of whether a diagrammatic capture method is used, primers can be immobilized in a manner where the 3' end is available for polymerase extension (e.g., primers can be attached at or near the 5' end). Figure 1D The combined RCA / MDA reaction is shown, where the product of the first sub-step is shown as having progressed to the point where the sense strand contains two copies (two sequence units) of the circular template, and the circular template hybridizes with a portion of the replicating third copy (the third sequence unit). Each sequence unit in the sense strand contains a region complementary to the target sequence (indicated by a solid black line) and a region complementary to the primer (indicated by a hollow and lined rectangle).
[0107] because Figure 1D The reaction involves amplification primers complementary to the primer binding sites in the sense strand, and the product of the first sub-step also includes three antisense strands shown at various stages of extension. Each sequence unit in the antisense strand contains a region that serves as a copy of the target sequence (indicated by a solid gray line) and a region complementary to the primer (indicated by hollow and lined rectangles). Figure 1D The product of the second sub-step has progressed to the point of producing nearly six copies of the cyclic template and five antisense strands, which are shown in various stages of the extension. Figure 1DThe final product of the RCA / MDA reaction after the cyclic template is absent (e.g., has been removed) in the third sub-step is shown. Two regions of the final product are shown, including a region depicting the sequence units to illustrate the primary structure of the tandem and a region where the number and conformation of the sequence units are not specified (indicating the dynamic and variable secondary structure of the entire cluster). Figure 1D The diagram and the steps described above can be multi-pathed, so that the diagram steps occur at more than one individual site in the array.
[0108] MDA methods, such as in Figure 1C or Figure 1D The methods illustrated in the context can be used to generate one or more antisense strands complementary to at least a portion of the tandem strand. Amplification primers can be in solution, as shown in the figure. Alternatively, amplification primers can be attached to a solid support, for example, using the covalent or non-covalent attachment chemistry for nucleic acids described herein. Figures 4A-4F This describes a non-limiting exemplary MDA method using amplification primers attached to a solid support. The antisense strand of a tandem extension resulting from covalently fixed primers is covalently attached to the solid support. Amplification primers do not require covalent attachment to the solid support; instead, they attach to the cluster via hybridization with the sense strand. The antisense strand of a tandem extension resulting from primers not covalently attached to the solid support can attach to the cluster via hybridization with the sense strand or through non-covalent bonds with other parts of the cluster.
[0109] like Figure 1C As exemplified, the sense and antisense strands of the tandem can be generated in separate reactions. For example, the sense tandem strand can be generated via an RCA reaction, and then this sense tandem strand can be used as a template for a subsequent MDA reaction. Thus, RCA can be performed to generate a sense tandem strand attached to a solid support, the reagents used for RCA can then be removed from contact with the solid support, and the reagents used for MDA (e.g., primers complementary to the primer binding sites in the sense strand) can then be contacted with the solid support. Alternatively, the sense and antisense strands of the tandem can be generated in a “single-pot” reaction, such that the RCA reagents are not separated from the MDA reagents. For example, a first primer for amplifying a circular template via RCA can be present together with a second primer for amplifying the tandem via MDA. See, for example, Figure 1D The first and second primers can be extended in the presence of each other. If desired, the sense and antisense strands of the tandem can be generated simultaneously, for example, by simultaneously performing RCA and MDA reactions.
[0110] One or more antisense chains can be as follows: Figure 1EThe generation process is described below. Amplification can begin with a tandem sense strand (indicated by black lines) anchored to a solid support (indicated by dot-filled rectangles) via a adapter (indicated by gray lines). The sense strand of the tandem strand contains sequence units, each containing a primer binding site. Amplification primers complementary to all or part of the primer binding sites can hybridize with the sense strand of the tandem strand and extend in the MDA reaction to produce an antisense strand (shown as gray lines at least partially annealed to the sense strand). For illustrative purposes, a portion of the MDA product is shown to indicate the direction of extension (see arrows) and to indicate the primary structure of the tandem antisense strand with various lengths and annealing modes. Another portion of the MDA product is shown with less structure specificity to indicate the dynamic and variable secondary structure of the cluster as a whole.
[0111] The MDA reaction can be carried out in the presence of deoxyribonucleoside triphosphate (dATP), deoxyribonucleoside triphosphate (dTTP), deoxyribonucleoside triphosphate (dGTP), and deoxyribonucleoside triphosphate (dCTP) (or their analogues). The resulting antisense strand can contain adenine, guanine, cytosine, and thymine bases. The MDA reaction can also be carried out in the presence of deoxyribonucleoside triphosphate (dUTP) (or its analogues). When the MDA reaction is carried out in the presence of dUTP in addition to dATP, dTTP, dGTP, and dCTP, the resulting antisense strand can contain uracil bases (indicated by an asterisk in the antisense strand) in addition to adenine, guanine, cytosine, and thymine bases. The position of the uracil base in the antisense strand is not predetermined. See below for reference. Figure 2D As described in further detail, the antisense strand containing uridine bases can be digested after the antisense strand is sequenced. Although this application describes that the MDA reaction can be carried out in the presence of dUTP (or its analogues) such that the resulting antisense strand contains uracil bases, this is for illustrative purposes only, and alternative methods to achieve similar results are also considered. For example, the MDA reaction can be carried out in the presence of deoxyribonucleotide triphosphates with modified or atypical bases such that the resulting antisense strand contains one or more modified or atypical bases. Such modified or atypical bases can target the antisense strand for degradation, such as reference... Figure 2D Enzymatic digestion as described below. As another example, the MDA reaction can be carried out in the presence of modified or atypical deoxyribonucleotide triphosphates (e.g., deoxypseudouridine triphosphate), such that the resulting antisense strand contains one or more modified or atypical nucleotides (e.g., deoxypseudouridine monophosphate). Such modified or atypical nucleotides can target the antisense strand for degradation, as referenced below. Figure 2D The enzyme digestion mentioned above.
[0112] Whether the base in the antisense strand is thymine or uracil (when the corresponding base in the sense strand is adenine) depends on the relative concentrations of dTTP and dUTP in the MDA reaction. The concentration of dUTP (or a deoxyribonucleotide triphosphate with modified or atypical bases, or a modified or atypical deoxyribonucleotide triphosphate) in the MDA reaction can be lower than the concentration of the other deoxyribonucleotide triphosphate in the MDA reaction, resulting in a low percentage of uracil bases (or modified or atypical bases, or modified or atypical nucleotides) present in the antisense strand. Uracil bases (or modified or atypical bases, or modified or atypical nucleotides) can be randomly distributed and present in a low percentage, such that the two antisense strands (any two antisense strands) contain uracil bases (or modified or atypical bases, or modified or atypical nucleotides) at different positions.
[0113] The concentration of deoxyribonucleic acid triphosphates (e.g., dATP, dTTP, dGTP, or dCTP) in the MDA reaction, or the concentration of all deoxyribonucleic acid triphosphates, may be about, at least, at least about, at most, or at most about: 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, 0.5 mM, 0.6 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 ...6 mM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 6 mM, 6 mM, 6 mM, mM, 7mM, 8mM, 9mM, 10mM, 11mM, 12mM, 13mM, 14mM, 15mM, 16mM, 17mM, 18mM, 19mM, 20mM, 25mM, 30mM, 35mM, 40mM, 45mM, 50mM, 55mM, 60mM, 65mM, 70mM, 75mM, 80mM, 85mM, 90mM, 95mM, 100mM, or any number or range between any two of these values. The concentration of dUTP in the MDA reaction can be, approximately, at least, at least approximately, at most, or at most approximately: 0.001 mM, 0.002 mM, 0.003 mM, 0.004 mM, 0.005 mM, 0.006 mM, 0.007 mM, 0.008 mM, 0.009 mM, 0.01 mM, 0.02 mM, 0.03 mM, 0.04 mM, 0.05 mM, 0.06 mM, 0.07 mM, 0.0 8mM, 0.09mM, 0.1mM, 0.2mM, 0.3mM, 0.4mM, 0.5mM, 0.6mM, 0.7mM, 0.8mM, 0.9mM, 1mM, 2mM, 3mM, 4mM, 5mM, 6mM, 7mM, 8mM, 9mM, 10mM, 11mM, 12mM, 13mM, 14mM, 15mM, 16mM, 17mM, 18mM, 19mM, 20mM, or any quantity or range between any two of these values.The ratio of the concentration of dUTP (or deoxyribonucleotide triphosphates with modified or atypical bases, or modified or atypical deoxyribonucleotide triphosphates) to the concentration of dTTP (or the concentration of another deoxyribonucleotide triphosphate or the total concentration of deoxyribonucleotide triphosphates other than dUTP, or deoxyribonucleotide triphosphates with modified or atypical bases, or modified or atypical deoxyribonucleotide triphosphates) can be as follows: approximately below, at least below, or up to. The following are considered minimum, maximum, or maximum approximations: 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70. 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, or a number or range between any two of these values.The percentage of deoxyribonucleotide triphosphates (dUTP, or deoxyribonucleotide triphosphates with modified or atypical bases, or modified or atypical deoxyribonucleotide triphosphates) in the MDA reaction can be, is about, is at least about, is at most, or is at most about: 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.00 8%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any amount or range between any two of these values.
[0114] In the MDA reaction, the concentration of dUTP (or a deoxyribonucleotide triphosphate with modified or atypical bases, or a modified or atypical deoxyribonucleotide triphosphate) relative to another deoxyribonucleotide triphosphate such as dTTP can be low, resulting in a low percentage of uracil bases (or modified or atypical bases, or modified or atypical deoxyribonucleotides) present in the antisense strand. The ratio of a nucleotide with a uracil base (or a modified or atypical base, or a modified or atypical nucleotide) to a thymine base (or another base, or all bases that are not uracil) can be one of the following, about one of the following, at least one of the following, at least one of the following, at most one of the following, or at most one of the following: 1:10000, 1:9000, 1:8000, 1:7000, 1:6000, 1:5000, 1:6000, 1:5000, 1:4000, 1:3000, 1:2000, 1 :1000, 1:900, 1:800, 1:700, 1:600, 1:500, 1:400, 1:300, 1:200, 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1 :74, 1:73, 1:72, 1:71, 1:70, 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1: 37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, or a number or range between any two of these values.The percentage of uracil bases in the antisense strand can be, is about, is at least about, is at most or is at most about: 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any number or range between any two of these values.
[0115] Nucleic acid clusters may be prepared prior to or as part of the methods described herein, for example by (i) providing a solid support with capture primers, (ii) hybridizing a circular nucleic acid template with primers, (iii) synthesizing the sense strand of the tandem by extending the capture primers along the circular nucleic acid template via rolling circle amplification, and (iv) synthesizing the antisense strand of the tandem by extending amplification primers that bind to primer binding sites in the sequence units of the sense strand, wherein the antisense strand contains one or more bases that are uracil.
[0116] Therefore, this disclosure provides a method for determining sequences from the sense and antisense strands of nucleic acids, comprising the steps of: (a)(i) providing a solid support having capture primers, (ii) hybridizing a circular nucleic acid template with primers, (iii) synthesizing the sense strand of the tandem by extending the capture primers along the circular nucleic acid template via rolling circle amplification, and (iv) synthesizing the antisense strand of the tandem by extending amplification primers that bind to primer binding sites in the sequence units of the sense strand, wherein the antisense strand contains one or more uracil bases, thereby providing a nucleus attached to the solid support. A nucleic acid cluster, wherein the nucleic acid cluster comprises a sense strand and an antisense strand of a tandem strand, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand in the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand.
[0117] The method disclosed herein for determining sequences from the sense and antisense strands of nucleic acids can be used to determine sequences from the sense and antisense strands of nucleic acids in a sample. Advantageously, this method can have little or no bias or preference for nucleic acids of certain sizes, such that the size distribution of nucleic acids in the sequenced sample is the same as, similar to, or comparable to the size distribution of the target sequence being sequenced. The size distribution of nucleic acids in the sequenced sample is the same as, similar to, or comparable to the size distribution of reads generated from sequencing the first and / or second strands. This distribution can be, for example, a normal distribution.
[0118] A particularly useful sequencing procedure that can be performed in the methods described herein is a cyclic procedure employing repeated reagent delivery cycles. Each cycle may include one or more steps. For example, each cycle may include all the steps required to detect a single nucleotide position in a template nucleic acid. Some sequencing procedures employ cyclic reversible terminator (CRT) chemistry, where each cycle includes the following steps: (i) adding a single reversibly terminating nucleotide to augment a nascent primer to the nucleotide position to be detected; (ii) detecting the nucleotide at the single nucleotide position; and (iii) unblocking the nascent primer to allow a return to step (i) to begin subsequent cycles.
[0119] A method for determining a sequence from the sense and antisense strands of a nucleic acid may include the following steps: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand and an antisense strand in a tandem, wherein the tandem comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand in the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand. Optionally, extending the second primer in step (e) includes (i) adding a reversibly terminated nucleotide to the second primer and (ii) a repetitive cycle of unblocking the reversibly terminated nucleotide on the second primer. As an alternative or alternative, the extension primer in step (c) may include (i) adding a reversibly terminating nucleotide to the primer and (ii) a repeating cycle of unblocking the reversibly terminating nucleotide on the primer.
[0120] Regardless of whether one or both of steps (c) and (e) include repeated steps (i) and (ii) as described above, the repeated cycles in steps (c) and / or (e) may optionally include (iii) detection of a stabilized ternary complex comprising a polymerase, the next correct nucleotide, and a second primer that hybridizes to either the antisense strand or the sense strand, respectively. Alternatively, the next correct nucleotide or polymerase may have a marker detected in steps (c) and / or (e) to generate a signal used to determine a sequence from at least a portion of the target sequence in either the antisense strand or the sense strand, respectively.
[0121] A concrete example of a useful CRT nucleic acid sequencing process is Sequencing By Binding. TM (SBB TM The reaction, for example, is described in the following: jointly owned U.S. Patent Application Publication Nos. 2017 / 0022553A1; 2018 / 0044727A1; 2019 / 0169688A1; 2019 / 0345544A1; and 2019 / 0367974A1, each of which is incorporated herein by reference. Typically, SBBs are used to determine the sequence of a template nucleic acid molecule. TM The method can be based on the formation of a stable ternary complex (between polymerase, initiating nucleic acid, and homologous nucleotides) under specific conditions. The method may include an inspection phase and a nucleotide incorporation phase.
[0122] SBB can be performed on at least one template nucleic acid molecule initiated by primers, polymerase, and at least one type of nucleotide. TM The process includes a check phase. Under conditions where nucleotides are not covalently added to the primers, the interaction between the polymerase and the nucleotides with the initiated template nucleic acid molecule can be observed; and the next base in each template nucleic acid can be identified based on the observed interactions. For example, the nucleotides may contain detectable labels. In some embodiments, the polymerase may be labeled. During the check phase, various conditions and reagents can be used to stabilize the ternary complex. For example, the primers may contain a reversible blocking portion that prevents covalent attachment of nucleotides; and / or the polymerase cofactors required for extension, such as divalent metal ions, may be absent; and / or an inhibitory divalent cation that inhibits polymerase-based primer extension may be present; and / or the polymerase present in the check phase may have chemical modifications and / or mutations that inhibit primer extension; and / or the nucleotides may have chemical modifications that inhibit incorporation. In a particular embodiment, the ternary complex is transcribed by Li + It is stable in the presence of betaine or both. For example, the reagents and techniques described in U.S. Patent No. 10,400,272 (which is incorporated herein by reference) can be used.
[0123] It should be understood that the options described herein for stabilizing ternary complexes are not mutually exclusive, but can be used in various combinations. For example, ternary complexes can be stabilized by a combination of one or more methods, including but not limited to cross-linking of polymerase domains, cross-linking of polymerase with nucleic acids, polymerase mutation for stabilizing ternary complexes, allosteric inhibition of small molecules, anti-competitive inhibitors, competitive inhibitors, non-competitive inhibitors, absence of catalytic metal ions, presence of blocking portions on primers, and other methods described herein.
[0124] Typically, detection can be achieved during the inspection step by sensing the inherent properties of the ternary complex or the labeled portion thereto. Exemplary properties upon which detection can be based include, but are not limited to, mass, conductivity, energy absorption, luminescence, etc. Luminescence detection can be performed using methods known in the art related to nucleic acid arrays. The luminescent group can be detected based on any of a variety of luminescent properties, including, for example, emission wavelength, excitation wavelength, fluorescence resonance energy transfer (FRET), quenching, anisotropy, or lifetime. Other detection techniques that can be used with the methods described herein include, for example, mass spectrometry for sensing mass; surface plasmon resonance for sensing the binding of polymerases or other analytes to a surface; absorbance for sensing the wavelength of energy absorbed by the label; calorimetry for sensing temperature changes due to the presence of the label; conductivity or impedance for sensing the electrical properties of the label; or other known analytical techniques.
[0125] The extension phase can be performed by creating conditions in which nucleotides can be added to primers that hybridize with template nucleic acid molecules. In some implementations, this involves removing the reagents used in the inspection phase and replacing them with reagents that promote extension. For example, the inspection reagents can be replaced with polymerases and reversibly terminating nucleotides.
[0126] Nucleotide analogs involved in stabilizing ternary complexes, or nucleotide analogs added to primers via polymerase catalysis, may include a terminator moiety that reversibly prevents subsequent nucleotide incorporation into the 3' end of the primer after the analog has been incorporated. For example, U.S. Patents 7,544,794 and 8,034,923 (the disclosures of which are incorporated herein by reference) describe reversible terminator moieties in which the 3'-OH group is replaced by a 3'-ONH2 moiety. Another type of reversible terminator moieties are linked to a nitrogenous base of the nucleotide, such as those described in U.S. Patent 8,808,989 (the disclosure of which is incorporated herein by reference). Other reversible terminator moieties that can be similarly used in conjunction with the methods described herein include those described in references cited elsewhere herein or in U.S. Patents 7,956,171, 8,071,755, and 9,399,798 (the disclosures of which are incorporated herein by reference). In some implementations, the reversible terminator portion can be modified or removed from the primer in a process known as “unblocking”, thereby allowing subsequent nucleotide incorporation.
[0127] When included in the methods described herein, the unblocking process can facilitate the sequencing of primer-template nucleic acid hybrids. The unblocking process can be used to convert a reversibly terminated primer into an extendable primer. Primer extension can then be used to move the site of the ternary complex formation along the template nucleic acid to different locations. Repeated cycles of extension, checking, and unblocking can be used to reveal the sequence of the template nucleic acid. Each cycle reveals subsequent bases in the template nucleic acid. Exemplary reversible terminator portions, methods for incorporating them into primers, and methods for modifying primers for further extension (commonly referred to as "unblocking") are described in U.S. Patent Nos. 7,427,673; 7,414,116; 7,544,794; 7,956,171; 8,034,923; 8,071,755; 8,808,989; or 9,399,798. Further examples are illustrated in the following: Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. Patent No. 7,057,026; WO 91 / 06678; WO 07 / 123744; U.S. Patent No. 7,329,492; U.S. Patent No. 7,211,414; U.S. Patent No. 7,315,019; U.S. Patent No. 7,405,281 and US 2008 / 0108082, each of which is incorporated herein by reference.
[0128] In certain embodiments, prior to the step of forming a stable ternary complex with the primer-template hybrid, reagents used during the primer modification step (e.g., extending the primer by adding nucleotides or capping the primer by adding a ternary complex inhibitor moiety) are removed to avoid contact with the primer-template hybrid. For example, it may be desirable to remove the nucleotide mixture used in the extension step when one or more types of nucleotides in the mixture would interfere with the formation or detection of the ternary complex in a subsequent assay step. Similarly, it may be desirable to remove the polymerase or cofactor used in the primer modification step to prevent unwanted catalytic activity in the subsequent assay step. However, if desired, nucleotides from the extension step can be removed, while the polymerase from the extension step is retained and proceeds to the step of forming the ternary complex. See, for example, U.S. Patent Application Publication No. 2020 / 0032317A1, which is incorporated herein by reference. Thus, the ternary complex detected in the assay step may contain a polymerase molecule used in the preceding primer extension step. After removing the reaction components, a washing step can be performed, in which an inert fluid is used to remove primer-template hybrids from the residual components of the reagent mixture used for primer modification.
[0129] Another useful CRT sequencing process is sequencing by synthesis (SBS). SBS typically involves the enzymatic extension of nascent primers by iteratively adding nucleotides to the template strand that hybridizes with the primers. In short, SBS is initiated by contacting the initiated target nucleic acid with one or more labeled nucleotides, DNA polymerase, etc. The primers incorporate the detectable labeled nucleotides. The label in the labeled nucleotides can be removed. After the label in the labeled nucleotides is removed, molecular scarring may remain. In contrast, the nucleotides incorporated into the initiated template during the extension phase can be unlabeled nucleotides, such as native nucleotides. Therefore, it is not necessary to remove the label from the incorporated nucleotides. Optionally, the labeled nucleotides may also contain a reversible terminator portion. Thus, a single reversibly terminating nucleotide is added to the primer so that subsequent extension does not occur until a deblocking agent is delivered to remove the reversible terminator portion from the primer. An SBS cycle can be performed n times to extend the primer by n nucleotides, thereby detecting sequences of length n. Exemplary SBS procedures, reagents, and detection components that can be readily applied to the systems or devices described herein are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Patent Nos. 7,057,026; 7,329,492; 7,211,414; 7,315,019 or 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082A1, each of which is incorporated herein by reference. Also useful are SBS methods commercially available from Illumina, Inc. (San Diego, CA).
[0130] Therefore, a method for determining a sequence from the sense and antisense strands of a nucleic acid may include the following steps: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand and an antisense strand in a tandem, wherein the tandem comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand in the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand. Optionally, extending the second primer in step (e) includes (i) adding a reversibly terminated nucleotide to the second primer and (ii) a repetitive cycle of unblocking the reversibly terminated nucleotide on the second primer. As an alternative or alternative, the extension primer in step (c) may include (i) adding a reversibly terminating nucleotide to the primer and (ii) a repeating cycle of unblocking the reversibly terminating nucleotide on the primer.
[0131] Regardless of whether one or both of steps (c) and (e) include repeated steps (i) and (ii) as described above, the reversibly terminated nucleotide added to the primer in step (c) or (e) may optionally have a marker that is detected to generate a signal used to determine at least a portion of the target sequence in the antisense or sense strand, respectively. Alternatively, the repeated cycle in one or both of steps (c) and (e) may also include (iii) removal of the marker after it has been detected.
[0132] Some SBS implementations are cyclic but do not require the use of reversible terminator nucleotides. A particularly useful method involves detecting protons released when nucleotides are incorporated into the extension product. For example, sequencing based on proton release detection can use reagents and electrodetectors commercially available from Thermo Fisher (Waltham, MA) or described in: U.S. Patent Application Publication No. 2009 / 0026082 A1; No. 2009 / 0127589 A1; No. 2010 / 0137143 A1; or No. 2010 / 0282617 A1, each of which is incorporated herein by reference. Another cyclic sequencing process that does not require the use of reversible terminator nucleotides is pyrosequencing. Pyrosequencing detects the release of inorganic pyrosequencing (PPi) when nucleotides are incorporated into nascent primers that hybridize with template nucleic acid chains (Ronaghi et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al., Science 281(5375), 363 (1998); U.S. Patent Nos. 6,210,891, 6,258,568 and 6,274,320, each of which is incorporated herein by reference).
[0133] Ligation sequencing reactions are also useful, including, for example, those described in: Shendure et al., Science 309:1728-1732 (2005); U.S. Patent No. 5,599,675; or U.S. Patent No. 5,750,341, each of which is incorporated herein by reference. Hybridization sequencing procedures can be used, for example, as described in: Bains et al., Journal of Theoretical Biology 135(3),303-7 (1988); Drmanac et al., Nature Biotechnology 16,54-58 (1998); Fodor et al., Science 251(4995),767-773 (1995); or WO1989 / 10977, each of which is incorporated herein by reference. In both ligation sequencing and hybridization sequencing procedures, primers hybridizing with a nucleic acid template are ligated by oligonucleotides, for example, using fluorescently labeled oligonucleotides for repeated extension cycles.
[0134] Reagent removal or washing procedures can be performed between any of the various steps described herein. These procedures can be used to remove one or more reagents present in the reaction vessel or on the solid support. For example, a reagent removal or washing step can be used to separate a primer-template hybrid from other reagents that have been in contact with the primer-template hybrid under ternary complex stability conditions. In a particular embodiment, reagent separation is facilitated by attaching the reagent of interest, such as the primer-template hybrid, to the solid support and removing fluids in contact with the solid support. One or more reagents described herein may be attached to the solid support or provided in solution, as needed, to suit the specific application of the method or apparatus described herein.
[0135] Sequencing methods may include more than one repetition of the cycles or steps within cycles described herein. For example, the checking and primer modification steps may be repeated more than once, as may optional steps such as unblocking primers or washing away unwanted reactants or products between various steps. Therefore, nucleic acids may undergo at least 2, 5, 10, 25, 50, 100, 150, 200, or more repetitions of the sequencing methods described herein. Fewer cycles may be performed when shorter read lengths are required. Therefore, nucleic acids may undergo up to 200, 150, 100, 50, 25, 10, 5, or 2 cycles of the sequencing methods described herein. The above range of cycle numbers is applicable to any reversible termination subprocess of the cycles described herein. In some embodiments, the sequencing method may perform a predetermined number of repetitions of the cycles. Optionally, the cycles may be repeated until a specific empirical observation state is reached. For example, the loop can be repeated as long as the signal is above the observable threshold, the noise is below the observable threshold, or the signal-to-noise ratio is above the observable threshold.
[0136] Other examples of sequencing methods that can be used in the methods described in this article include those developed by Illumina. TM ,Inc. (e.g., HiSeq) TM MiSeq TM NextSeq TM or NovaSeq TM Systems), Life Technologies TM (e.g., ABI PRISM) TM or SOLiD TM (systems), Pacific Biosciences (e.g., using SMRT) TM Technology systems such as Sequel TM or RS II TM (System), MGI Tech (DNBSEQ-T7, MGISEQ-2000, MGISEQ-200 or BGISEQ-500), or Qiagen (e.g. Genereader) TM The methods used on commercialized platforms (systems).
[0137] The method described in this article for determining the sequences of the sense and antisense strands of nucleic acids can be performed in the presence of the antisense strand while the sense strand has been sequenced (see [link to article]). Figure 2B (Examples). In an optional configuration of this method, the antisense strand may be absent during sequencing of the sense strand (see [example]). Figure 2C and Figure 2D (Examples).
[0138] The method described in this article for determining the sequences of the sense and antisense strands of nucleic acids can be performed in the presence of the sense strand and with the antisense strand sequenced (see [link to article]). Figure 2A (Example of the method). In an optional configuration of this method, the sense strand may be absent during sequencing of the antisense strand.
[0139] The methods described in this article for determining the sequences of the sense and antisense strands of nucleic acids can be performed in the presence of the sense strand and the antisense strand being sequenced, as well as in the presence of the antisense strand and the sense strand being sequenced (see [link to article]). Figure 2A and Figure 2B (Examples). In an optional configuration of this method, the sense strand may be absent during sequencing of the antisense strand, and the antisense strand may be absent during sequencing of the sense strand.
[0140] The method described in this article for determining the sequences of the sense and antisense strands of nucleic acids can be performed under conditions in which the antisense strand hybridizes with primers (e.g., non-extending or extending primers generated by sequencing the antisense strand) when the sense strand is sequenced (see [link to article]). Figure 2B (Examples). In an optional configuration of this method, the sequencing primers extended during antisense sequencing may not be present during sense sequencing (see [example]). Figure 2C and Figure 2D (Examples).
[0141] The method described in this paper for determining the sequences of the sense and antisense strands of nucleic acids can be performed under conditions in which the sense strand hybridizes with primers (e.g., non-extending or extending primers generated by sequencing the sense strand) when the antisense strand is sequenced. In an optional configuration of the method, sequencing primers extended during sense strand sequencing may be absent during antisense strand sequencing.
[0142] The method described in this paper for determining the sequences of the sense and antisense strands of nucleic acids can be performed under the following conditions: when the antisense strand is sequenced, the sense strand hybridizes with primers (e.g., non-extending or extending primers generated by sequencing the sense strand), and wherein when the sense strand is sequenced, the antisense strand hybridizes with primers (e.g., non-extending or extending primers generated by sequencing the antisense strand). In an optional configuration of the method, sequencing primers extended during sense strand sequencing may be absent during antisense strand sequencing, and sequencing primers extended during antisense strand sequencing may be absent during sense strand sequencing.
[0143] When the steps of determining the sense strand sequence and synthesizing the antisense strand occur simultaneously, the method described herein for determining the sense and antisense strand sequences of nucleic acids can be performed. For example, during the step of determining the sense strand sequence, a portion of the antisense strand is synthesized. Figures 3A-3FThis paper describes a non-limiting exemplary method for determining the sequences of the sense and antisense strands of a nucleic acid, wherein the synthesized antisense strand includes an extension product derived from sequencing the sense strand.
[0144] Therefore, this disclosure provides a method for determining a sequence from the sense and antisense strands of a nucleic acid. The method may include the following steps: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand and an antisense strand in a tandem, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit contains a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand within the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) removing the antisense strand from the cluster; (e) after the antisense strand removal, hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (f) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand.
[0145] Optionally, this disclosure provides a method for determining a sequence from the sense and antisense strands of a nucleic acid. The method may include the following steps: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand and an antisense strand in a tandem, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit contains a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand within the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (e) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand, wherein steps (d) and (e) are performed when the antisense strand is present in the cluster. It should be understood that steps (d) and (e) may be performed before or after steps (b) and (c) as needed.
[0146] This disclosure provides a method for sequencing one strand of a double-stranded nucleic acid in the presence of the other strand. In some configurations, sequencing will involve extending a first primer along one strand of the double-stranded nucleic acid in the presence of a second primer that hybridizes to the other strand. A blocking or capping portion may be added to the second primer, thereby allowing the first primer to extend selectively in the presence of the second primer. Figure 2A and Figure 2B The diagram provides an example. The starting point of the method, such as... Figure 2AAs shown, a cluster has a tandem sense strand (indicated by black lines, with primer binding sites indicated by hollow and lined rectangles) and a tandem antisense strand (indicated by gray lines, with primer binding sites indicated by hollow and lined rectangles). The presence of more than one antisense strand on the sense strand is referred to herein as an antisense scale. The extendable 3' ends within the cluster, such as the 3' ends of the antisense strand (indicated by arrows on the gray lines), can optionally be modified with either closed or capped portions. The antisense strand can then be sequenced using a first set of sequencing primers (indicated by hollow arrows) and nucleotides (indicated by gray diamonds). Continuing to... Figure 2B The extended primers can then be modified with either a blocking or a capping portion. The sense strand can then be sequenced using a second set of sequencing primers (indicated by the ribbon arrows) and nucleotides (indicated by the black diamonds). Sequencing of the sense strand can be performed in the presence of the extension product from the antisense strand sequencing, because the extension product has been blocked or capped during sense strand sequencing to prevent it from generating background signals.
[0147] It should be understood that nucleic acid clusters can have any of a variety of 3' ends, which may be undesirably extended under conditions intended to extend the primer of interest. These 3' ends may be present not only on the second primer but also at the ends of the sense or antisense strands within the cluster. The presence of unwanted 3' ends in the cluster generates background signals caused by the extension of these 3' ends, and these background signals can hinder the ability to resolve the desired signal caused by sequencing-based extensions of the first primer. Figure 2A As shown in the first step, in the method described herein, this artifact can be avoided by modifying these 3' ends to incorporate a blocking portion that prevents polymerase extension of the modified 3' ends or by incorporating a capping portion that inhibits the formation of a ternary complex at the modified 3' ends.
[0148] In another example, the array site can contain sense and antisense strands, and one strand can be capped, allowing selective sequencing of the other strand without background interference from the capped strand. Selective primer capping can be used to sequence paired regions of larger nucleic acid molecules to determine the structural relationship between two regions in the genome. Selective primer capping can also allow for the separate sequencing of the target sequence and associated tag sequences. Exemplary tags include, but are not limited to, those attached to the target sequence to identify the nucleic acid sequence from which the target sequence originates (specific tag sequences are uniquely attached to target nucleic acids harvested from specific cells, tissues, or other samples). Tags can also be used to distinguish mutations or polymorphisms of biological (e.g., clinical) interest from errors introduced during sample extraction and preparation procedures. Such tags are often referred to as unique molecular identifiers (UMIs).
[0149] The blocking portion can be added to the primer or other nucleic acid in any of the various ways described herein. For example, a nucleotide having the blocking portion can be added to the 3' end of the primer or other nucleic acid via polymerase catalysis. Alternatively, the primer can be chemically modified to incorporate the blocking portion prior to hybridization with the nucleic acid molecule or cluster. Regarding the blocking and unblocking steps in the CRT sequencing process, exemplary methods for blocking the 3' end are described herein. In some embodiments, the primer or other nucleic acid can be modified with a reversible terminator portion. In such a configuration, the reversible terminator can be removed or modified to unblock the primer or other nucleic acid for subsequent extension.
[0150] The capping portion can be added to a primer or other nucleic acid in any of the various ways set forth herein. For example, the capping portion can be added to the primer before hybridization with a template. In another instance, the capping portion can be added using synthetic techniques. Optionally, the primer can be hybridized with a template, and the primer can subsequently be modified to include the capping portion. Optionally, the primer can be extended to incorporate a nucleotide containing a ligand, and the ligand can then bind to a receptor that acts as an inhibitor of ternary complex formation. Useful reagents and methods for capping nucleic acids are set forth in U.S. Patent Application Publication No. 2019 / 0367974 A1, which is incorporated herein by reference.
[0151] In another example, the primer may have a chemically modifiable moiety that reacts with another reagent to form a capping portion. The primer may have the chemically modifiable moiety before hybridization with the template, or alternatively, after hybridization, the primer may be extended to incorporate a nucleotide with the chemically modifiable moiety. Subsequently, the chemically modifiable moiety may covalently react to attach a ternary complex inhibitor to the primer. Suitable chemically modifiable moieties may include functional groups such as amino groups, carboxyl groups, maleimide groups, oxo groups, or thiol groups. Functional groups illustrated herein in the context of attaching nucleic acids to a solid support, such as click chemistry, may also be useful.
[0152] The methods disclosed herein may include primer modification processes that add nucleotides or other portions to primers. Primer modification processes can be used to prepare primer-template nucleic acid hybrids for use in sequencing processes or as part of a screening process used during nucleic acid analysis (such as sequencing analysis). For example, primer modification may add a reversible terminator portion to the 3' end of the nucleic acid to prevent elongation of the nucleic acid during amplification or sequencing. Optionally or additionally, primer modification may add a cap portion to the nucleic acid to prevent the formation of a ternary complex at the 3' end of the nucleic acid.
[0153] Therefore, this disclosure provides a method for determining a sequence from the sense and antisense strands of a nucleic acid. The method may include the steps of: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand and an antisense strand in a tandem, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit contains a target sequence and a primer binding site; (b) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand within the cluster; (c) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) adding a closing or capping portion to the extended primer; (e) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (f) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand, wherein the closing portion prevents further extension of the extended primer, or wherein the capping portion prevents polymerase from binding to the 3' end of the extended primer.
[0154] This disclosure also provides a method for determining a sequence from the sense and antisense strands of a nucleic acid, comprising the steps of: (a) providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand and an antisense strand in a tandem, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) adding a closing or capping portion to the 3' end of the nucleic acid cluster attached to the solid support; (c) hybridizing a primer to a primer binding site in a sequence unit of the antisense strand in the cluster; (d) extending the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (e) hybridizing a second primer to a primer binding site in a sequence unit of the sense strand; and (f) extending the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand. Optionally, the method may further include adding a closing portion to the extended primer prior to step (d) to prevent further extension of the extended primer during step (f). As an optional or alternative approach, the method may also include adding a capping portion to the extended primer prior to step (d) to prevent polymerase from binding to the 3' end of the extended primer during step (f).
[0155] A particularly useful capping chemistry would be selective for the 3' end of a primer or other nucleic acid. Capping chemistry can be selective for the oxygen or hydroxyl portion of the 3' end of the nucleic acid. For example, capping chemistry can be reactive to the 3' end of a non-blocking primer but inert to a modified blocking primer. Chemistry using enzymes specific to the natural 3' end of the primer is particularly useful, including, for example, ligases and polymerases. A more efficient capping procedure when modifying the ends of natural 3' primers compared to blocking primers is beneficial for many applications.
[0156] In a particular embodiment, a ternary complex inhibitor is used to cap the primer. Any of a variety of moieties can be added to the primer to inhibit or prevent the subsequent formation of a ternary complex at the 3' end of the primer. Optionally, the primer can be attached to the oligonucleotide moiety due to the attachment of the 5' end of the oligonucleotide to the 3' end of the primer by a ligase-catalyzed process, or due to the extension of the primer by a series of nucleotide pairs catalyzed by a polymerase. Ligases, polymerases, and other primer-modifying enzymes can be useful. However, chemical techniques can also be used to modify the primer in the methods described herein. The oligonucleotide moiety attached to the primer can contain one or more non-natural nucleotide analogs. These analogs can be selected based on their ability to form base pairs with the template, but analogs that do not pair with the template can also be used. Similarly, natural nucleotides can be present in the oligonucleotide moiety at one or more sites that form mismatches that disrupt base pairing between the oligonucleotide and the template. Sites particularly useful for mismatches or non-natural nucleotide analogs are at or near the 3' end of the oligonucleotide moiety, where they can act as an inhibitor of subsequent ternary complex formation.
[0157] Another example of a useful portion that can be added to a primer to inhibit or prevent the subsequent formation of a ternary complex at the 3' end of the primer is a single nucleotide (e.g., a natural nucleotide or a non-natural nucleotide analog). For example, a mismatched nucleotide can be attached to the 3' end of a primer to prevent ternary complex formation. Particularly useful mismatches and polymerases whose ability to recognize mismatches is affected are described in Kwok et al., Nucleic Acids Res. 18(4):999–1005 (1990), which is incorporated herein by reference. The mismatched nucleotide can be present at the 3' end of the primer, introduced by an oligonucleotide motif added to the primer. In some embodiments, a series of two or more mismatched nucleotides may be present at or near the 3' end of the primer to achieve inhibition of ternary complex formation.
[0158] Regardless of whether the nucleotide added to the 3' end of the primer matches the template, this nucleotide can contain an exogenous portion that acts as a ternary complex inhibitor. The exogenous portion can have a steric hindrance effect, thereby preventing the polymerase, nucleotide, or both from forming the ternary complex. This portion can have other effects on ternary complex formation, including but not limited to charge repulsion of the polymerase or homologous nucleotide, perturbation of the primer-template nucleic acid hybrid structure, and repulsion of the polarity of the polymerase or homologous nucleotide. The exogenous portion can be attached to a nucleotide via a linker, such as to the nucleotide attached to the 3' end of the primer. Particularly useful linkers have a relatively short length, such that ternary complex inhibition occurs proximal to the 3' end of the primer. For example, linkers can have a length from a single covalent bond to up to 2, 5, 10, or 15 covalent bonds. Optionally or additionally, linkers can have a length shorter than 20, 15, 10, or 5 covalent bonds. Thus, the linker can maintain a distance of at most, for example, between the attachment site on the nucleotide and the attachment site on the ternary complex inhibitor. Or even smaller. A relative lack of flexibility can be a useful feature of the linker, similarly allowing the inhibitor portion to be located close to the 3' end of the primer. Therefore, linkers with amide, ester, carbon-carbon double, carbon-carbon triple, ring, and other rotationally bound bond structures can be useful.
[0159] Particularly useful exogenous moieties that can attach to nucleotides include, but are not limited to, biotin or other ligands that can inhibit or prevent the formation of ternary complexes due to their presence at the 3' end or due to their interaction with avidin, streptavidin, or other receptors, thereby preventing the formation of ternary complexes. Exemplary biotin analogs that can be used include, but are not limited to, biotin carbonate 5 or biotin carbamate 6 (see Yamamoto Chem Asian J. 10:1071-1078 (2015), which is incorporated herein by reference), 2-iminobiotin, diaminobiotin, or desulfobiotin. Biotin analogs that reversibly bind to avidin or streptavidin can be particularly useful when the cap is removed under relatively mild conditions that do not denature avidin or streptavidin. Peptides with affinity for streptavidin or avidin can also be useful, such as peptides for SBP-tag systems (see Keefe et al. Protein Expression and Purification 23:440-446 (2001), which is incorporated herein by reference).
[0160] Other ligand-receptor pairs that can be used to form ternary complex inhibitors include, but are not limited to, antibodies (or functional fragments thereof, such as Fab or ScFv) and epitopes; carbohydrates and lectins; or nucleic acids (or their analogues) and their complementary nucleic acids (or their analogues). Any of a variety of functional antibody fragments can be used, including, for example, monovalent types such as Fab or scFv, or other monovalent types or their engineered variants such as F(ab')2, double-stranded antibodies, triple-stranded antibodies, microantibodies, or single-domain antibodies (see Holliger and Hudson, Nat. Biotechnol. 23:1126-1136 (2005), which is incorporated herein by reference). Another class of particularly useful ligand-receptor pairs are peptides and their binding conjugates, which are commonly used as purification tags when fused with recombinant proteins. Exemplary peptides include, but are not limited to, those with divalent cations such as Ni 2+ This includes peptides with multiple histidine residues (a polypeptide sequence of 6 or more histidine residues), glutathione S-transferase (GST) bound to glutathione, myc-tags (e.g., peptides with the sequence: EQKLISEEDL (SEQ ID NO:1)) bound to antimyc antibodies (available from the Developmental Studies Hybridoma Bank at the University of Iowa); calmodulin-binding peptides (CBP, e.g., peptides with the sequence KRRWKKNFIAVSAANRFKKISSSGAL (SEQ ID NO:2)) bound to calmodulin; FLAG tags (e.g., peptides with the sequence: DYKDDDD (SEQ ID NO:3) or DYKDDDDK (SEQ ID NO:4) or DYKDDDK (SEQ ID NO:5)) bound to antiFLAG antibodies; or maltose-binding proteins bound to amylose or maltose. It should be understood that when using the above-described purified tagging systems, the peptide or the molecule bound to it may be covalently linked to a nucleotide or primer. Typically, it is preferred to attach the smaller partner body to the nucleotide via a linker and then bind it to the larger partner body to form a ternary complex inhibitor.
[0161] The advantage of using a ligand-receptor combination as the capping part is that the ligand does not need to inhibit the formation of the ternary complex until it binds to the receptor. For example, the nucleotide attached to the ligand can bind to the polymerase and the primer-template nucleic acid to form a ternary complex, and the nucleotide can be incorporated into the primer to position the ligand at the 3' end of the primer. The receptor can then bind to the ligand at the 3' end of the primer to inhibit the formation of the ternary complex at the 3' end of the primer.
[0162] The blocking nucleotides added to the primers in the methods described herein do not need to have exogenous labels. This is because the methods described herein do not require detection of extended primers. However, if desired, one or more types of reversible termination nucleotides used in the methods described herein can be detected, for example, by attaching exogenous labels to the nucleotides. Exemplary reversible terminator portions, methods for incorporating them into primers, and methods for modifying primers for further extension (often referred to as “unblocking”) are described elsewhere in this document and in the references cited herein.
[0163] Similarly, the ternary complex inhibitors or other primer caps added to the primers in the methods described herein do not need to have exogenous labels. This is because the methods described herein do not require the detection of primer-template nucleic acid hybrids containing the capped portion. However, if desired, one or more types of capped portions used in the methods described herein can be detected, for example, by attaching an exogenous label to the cap or ternary complex inhibitor.
[0164] This disclosure provides a method for sequencing one nucleic acid strand in the presence of another nucleic acid strand, then removing the sequenced strand, and then sequencing the other strand. Examples are provided in... Figure 2A , Figure 2C and Figure 2D The diagram provides the starting point for the method, such as... Figure 2A As shown, a cluster has a tandem sense strand (indicated by black lines, with primer binding sites indicated by hollow and lined rectangles) and a tandem antisense strand (indicated by gray lines, with primer binding sites indicated by hollow and lined rectangles). The extendable 3' ends within the cluster, such as the 3' end of the antisense strand (indicated by arrows on the gray lines), can optionally be modified with closed or capped portions. The antisense strand can then be sequenced using a first set of sequencing primers (indicated by hollow arrows) and nucleotides (indicated by gray diamonds). Continuing to... Figure 2C Then, the antisense strand and extended primers can be removed from the cluster, for example, through denaturation, degradation, or other methods. (Reference) Figure 2D Then, the antisense strand and the extended primer can be removed from the cluster by enzymatic digestion. For example, the antisense strand may contain a base that is uracil (indicated by an asterisk in the antisense strand), as shown in the reference. Figure 1EAs described above, uracil DNA glycosylation enzyme (UDG) can be used to catalyze the excision of uracil bases in the antisense strand, forming a base-free (pyrimidine-free) site while maintaining the integrity of the phosphodiester backbone of the antisense strand. The lysinic activity of endonuclease VIII can then be used to disrupt the phosphodiester backbone on the 3' and 5' sides of the base-free site, thereby releasing the remaining base-free deoxyribose portion of the base-free site, resulting in a nick on the antisense strand. The remaining antisense strand and extended primers can then be digested using a nuclease that digests double-stranded DNA. For example, a double-strand-specific exonuclease with 5' to 3' exonuclease activity (such as T7 exonuclease) can be used to digest the remaining antisense strand and extended primers from the 5' ends (the 5' ends of the antisense strand and extended primers, and the 5' end at the antisense strand nick). Although this application references... Figure 1E The description of MDA reactions in the presence of dUTP (or its analogues) such that the resulting antisense strand contains uracil bases is for illustrative purposes only and is not intended to be limiting. For example, MDA reactions can be carried out in the presence of deoxyribonucleotide triphosphates with modified or atypical bases, such that the resulting antisense strand contains one or more modified or atypical bases. Such modified or atypical bases can target the antisense strand for degradation, such as reference... Figure 2D The described enzymatic digestion. As another example, the MDA reaction can be carried out in the presence of modified or atypical deoxyribonucleotide triphosphates (e.g., deoxyribopseuuridine triphosphate), such that the resulting antisense strand contains one or more modified or atypical nucleotides (e.g., deoxyribopseuuridine monophosphate). Such modified or atypical nucleotides can target the antisense strand for degradation, such as reference... Figure 2D The enzyme digestion described.
[0165] refer to Figure 2C and Figure 2D Then, the sense strand can be sequenced using a second set of sequencing primers (indicated by ribbon arrows) and nucleotides (indicated by black diamonds). Sequencing of the sense strand can be performed in the absence of the antisense strand and extensions from antisense strand sequencing.
[0166] Figures 3A-3F A schematic diagram is provided for a method of determining sequences from the first and second strands of nucleic acids. Figures 3A-3F In the diagram, the first chain is indicated by a dashed line, and the second chain by a solid line. For example, in... Figure 3A As shown, nucleic acid capture primer 302 is attached to solid support 304 via a adapter. The primer can be used to capture target nucleic acid 306 through primer binding sites 306A and 306P, which are complementary to the primer. Figure 3AIn one configuration shown, the immobilized primers can partially hybridize with primer binding sites 306A and 306P located at the opposite ends of the target sequence 306T. Therefore, the immobilized primers act as a clamp to hold the two ends of the target nucleic acid together. The two ends can ligate simultaneously with the clamped nucleic acid to form a circular form 306c of the target nucleic acid. The single strand 308 is produced by a polymerase (such as Phi29) through isothermal amplification (such as rolling circle amplification) of a circular template initiated by hybridization with the immobilized primers. The synthesized single strand can be a tandem strand. The single-stranded tandem strand synthesized from nucleic acid primers may be referred to herein as the first strand (or the sense strand) of the tandem strand. This single strand may contain more than one copy of the circular template and is a tandem strand of sequence units. The first strand can be generated from the nucleic acid template without polymerase chain reaction. Primers and synthesized single strands in Figure 3A It is shown in dashed lines. Figure 3A In the process, two copies (two sequence units) of the circular template have been generated, and the circular template has hybridized with a portion of the third copy (third sequence unit) that is being replicated. Figure 3B The final product of the RCA reaction in the absence of a cyclic template is shown. During subsequent steps of the method following the RCA reaction, the final product of the RCA reaction may not hybridize with a cyclic template. In some embodiments, the final product of the RCA reaction hybridizes with a cyclic template during one or more subsequent steps of the method following the RCA reaction.
[0167] refer to Figures 3B-3D The steps of determining the first-strand sequence using sequencing primer 310 and synthesizing the second-strand (e.g., antisense strand) using amplification primers can occur simultaneously. For example, the sequencing primers can initiate the first strand. The sequencing primers can be extended to produce extension product 312 to determine the first-strand sequence using, for example, a polymerase (not shown). Any sequencing method of this disclosure, such as SBB... TM Alternatively, SBS can be used to determine the sequence of the first strand and generate the second strand. The extension products of the sequencing primers used to determine the first strand sequence can be further extended with MDA to generate the tandem second strand 314 (see [link to MDA]). Figure 3D (Examples). Polymerases, such as those with strand displacement activity, can be used to further extend sequencing extension products with MDA to produce a second strand. The synthesized second strand comprises the extension product from sequencing the first strand using sequencing primers. Sequencing primers, the extension products of the sequencing primers, and the resulting second strand are shown as solid lines in the figure. The number of sequencing primers binding to the first strand, the number of sequencing primer extension products, and the number of second strands produced shown in the figure are for illustrative purposes only and are not intended to be limiting. More than one second strand produced by further extending the extension products of sequencing primers is considered... Figures 3D-3F It appears as scales, and in this article it is referred to as second-chain scales (or antisense chain scales).
[0168] refer to Figure 3E The 3' end of the second strand can be capped by incorporating a 3' blocking or capping portion 316 to prevent further extension and / or hinder or exclude the formation of the ternary complex during sequencing of the second strand. Many of the 3' blocking or capping portions described elsewhere in this disclosure can be used. The second strand can be initiated with sequencing primer 318 and sequenced by extending sequencing primers to produce a ternary complex. Figure 3F The extended sequencing primer 320 is described in the instructions.
[0169] Figures 4A-4F A schematic diagram is provided for a method of determining sequences from a first and second strand of nucleic acid, wherein the first and second strands are attached to a solid support. Figures 4A-4F In the diagram, the capture primer and the first strand are indicated by dashed lines, while the target nucleic acid, amplification primer, and second strand are indicated by solid lines. Corresponding regions are shown by double lines (e.g., the binding site 406AC of the capture primer and the binding site 406AA of the amplification primer, and the corresponding sequences on the capture and amplification primers) or single lines (e.g., the target sequence 406T and the corresponding sequences on the first and second strands, and the binding site 406P' of the sequencing primer and the corresponding sequence on the sequencing primer).
[0170] As in Figure 4A As shown, nucleic acid capture primer 402 is attached to solid support 404 via a adapter, and nucleic acid amplification primer 422 is attached to solid support via a adapter. The capture primers can be used to capture target nucleic acid 406 via primer binding sites 406A and 406P', which are complementary to the capture primers. Figure 4B In one configuration shown, the immobilized capture primers can partially hybridize with primer-binding sites 406A and 406P' located at the opposite ends of the target sequence 406T. Therefore, the immobilized capture primers act as a clamp to hold the two ends of the target nucleic acid together. The two ends can be ligated using, for example, a T4 ligase to form a circular form 406c of the target nucleic acid while hybridizing with the clamped nucleic acid. Kinases, such as T4 polynucleotide kinase, can phosphorylate the 5' end of the target nucleic acid prior to ligation to form a circular target nucleic acid. Reference Figures 4C-4D The single-stranded 408 is produced by a polymerase (such as Phi29) through isothermal amplification (such as rolling circle amplification) of a circular template initiated by hybridization with immobilized primers. The synthesized single strand can be a tandem strand. The single strand synthesized from nucleic acid primers may be referred to herein as the first strand (or the sense strand of the tandem strand). This single strand may contain more than one copy of the circular template and is a tandem strand of sequence units. The first strand can be generated from the nucleic acid template without polymerase chain reaction. Primers and synthesized single strands are shown in the figure as dashed lines. Figure 4DThree copies (three sequence units) of the circular template are shown for illustrative purposes only and are not intended to be limiting. Figure 4D The final product of the RCA reaction after removal of the cyclic template is shown.
[0171] Figures 4A-4F The method described may include the step of synthesizing a second strand (e.g., as an antisense strand) using amplification primers attached to a solid support. (See reference) Figure 4D Primer 422 can initiate the first strand 408. The amplification primers can be extended with MDA to produce... Figure 4E The second chain 414 of the tandem structure is described in the diagram. The second chain of the tandem structure may contain one or more sequence units. Figure 4E This illustrates a second chain, where one sequence unit has already been generated and the second sequence unit has begun to be generated. Figure 4E This illustrates the near completion of the generation of another second strand, representing the second sequence unit. Polymerases, such as those with chain displacement activity, can be used to generate the second strand, such as via MDA. The generated second strand is shown as a solid line in the figure. The generation of more than one second strand is shown in... Figure 4E The first strand appears as scales and is referred to herein as the second-strand scale (or antisense scale). In one configuration, the steps of synthesizing the second strand using amplification primers and determining the sequence of the first strand using sequencing primers can occur simultaneously, as described in the reference. Figures 3A-3F The description states that in this configuration, the amplification primers are sequencing primers.
[0172] refer to Figure 4A The illustration shows that the capture primer 402, which binds to nucleic acid 406, can contain a complementary sequence to a portion of the primer binding site 406A. Using RCA, the resulting first strand 408 can contain the complementary sequence of the entire primer site 406A. The amplification primer 422, used to generate more than one second strand from the first strand 408, can contain a portion of the primer binding site sequence. The partial sequences of the primer binding sites on the capture primer 402 and the amplification primer 422 may not overlap, as shown in... Figure 4A As illustrated in the illustration, primer binding site 406A is shown as having two regions: one region 406AC (shown as a double line with two lines of equal thickness) that binds to capture primer 402, and another region 406AA (shown as a double line with two lines of unequal thickness) that binds to amplification primer 422.
[0173] Surface-bound oligonucleotides can anneal to the adaptor site of the first-strand tandem. Alternatively, in some cases, surface-bound oligonucleotides anneal to adjacent or other distinct regions of the first-strand tandem. These regions are typically derived from the 5' or 3' end of the original library adaptor, although it is also conceivable that oligonucleotides target internal regions of the original library or derived from it.
[0174] refer to Figure 4F The second strand can be initiated and sequenced using sequencing primer 418. In one configuration, the 3' end of the second strand can be capped by incorporating a 3' blocking or capping portion to prevent further extension of the 3' end of the second strand and / or to prevent or exclude the formation of a ternary complex at the 3' end of the second strand. Many of the 3' blocking or capping portions described elsewhere in this disclosure can be used. The second strand can be initiated and sequenced before the first strand is initiated or sequenced. Alternatively, the second strand can be initiated and sequenced after the first strand is initiated or sequenced.
[0175] The solid support may contain two (or more, such as three, four, five, six, seven, eight, nine, ten or more) types or groups of primers. Two or more types of primers may, for example, serve as more than one capture primer and more than one amplification primer. Optionally or additionally, primer types may include more than one capture primer and more than one sequencing primer for sequencing (e.g., the first strand produced by extending the capture primer). The density of one type or group of primers (e.g., capture primers) on the solid support may be higher than the density of another type or group of primers (e.g., amplification primers) on the solid support. The density of one type or group of primers on the solid support may be the same as the density of another type or group of primers on the solid support. The density of primer (or all primers) types or groups may vary. The density of primer (or all primers) types or groups on the solid support may be: 1 x 10-1, approximately 1 x 10-1, at least 1 x 10-1, at least approximately 1 x 10-1, at most approximately 1 x 10-1, or at most approximately 1 x 10-1. 10 2x10 10 3x10 10 4x10 10 5x10 10 6x10 10 7x10 10 8x10 10 9x10 10 1x10 11 2x10 11 3x10 11 4x10 11 5x10 11 6x10 11 7x10 11 8x10 11 9x10 11 1x10 12 2x10 12 3x10 12 4x10 12 5x10 126x10 12 7x10 12 8x10 12 9x10 12 1x10 13 2x10 13 3x10 13 4x10 13 5x10 13 6x10 13 7x10 13 8x10 13 9x10 13 1x10 14 2x10 14 3x10 14 4x10 14 5x10 14 6x10 14 7x10 14 8x10 14 9x10 14 1x10 15 2x10 15 3x10 15 4x10 15 5x10 15 6x10 15 7x10 15 8x10 15 9x10 15 1x10 16 2x10 16 3x10 16 4x10 16 5x10 16 6x10 16 7x10 16 8x10 16 9x10 16 primers / m 2 Or the quantity or range between any two of these values.
[0176] This paper envisions various separation distances or average separation distances between two adjacent primers of the same type or population (or two different types or populations). The separation distance or average separation distance between two adjacent primers of the same type or population (or two different types or populations) can be, approximately, at least, at least approximately, at most, or at most approximately: 10nm, 11nm, 12nm, 13nm, 14nm, 15nm, 16nm, 17nm, 18nm, 19nm, 20nm, 21nm, 22nm, 23nm, 24nm, 25nm, 26nm, 27nm, 28nm, 29nm, 30nm, 31nm, 32nm, 33nm, 34nm, 35nm, 36nm, 37nm, 38nm, 39nm, 40nm, 41nm, 4 2nm, 43nm, 44nm, 45nm, 46nm, 47nm, 48nm, 49nm, 50nm, 51nm, 52nm, 53nm, 54nm, 55nm, 56nm, 57nm, 58nm, 59nm, 60nm, 61nm, 62nm, 63nm, 64nm, 65nm, 66nm, 67nm, 68nm, 69nm, 70nm, 71nm, 72nm, 73nm, 74nm, 75nm, 76nm, 77nm ,78nm,79nm,80nm,81nm,82nm,83nm,84nm,85nm,86nm,87nm,88nm,89nm , 90nm, 91nm, 92nm, 93nm, 94nm, 95nm, 96nm, 97nm, 98nm, 99nm, 100nm, 11 0nm, 120nm, 130nm, 140nm, 150nm, 160nm, 170nm, 180nm, 190nm, 200nm, 21 0nm, 220nm, 230nm, 240nm, 250nm, 260nm, 270nm, 280nm, 290nm, 300nm, 3 10nm, 320nm, 330nm, 340nm, 350nm, 360nm, 370nm, 380nm, 390nm, 400nm, 4 10nm, 420nm, 430nm, 440nm, 450nm, 460nm, 470nm, 480nm, 490nm, 500nm, 510nm, 520nm, 530nm, 540nm, 550nm, 560nm, 570nm, 580nm, 590nm, 600nm, 610nm, 620nm, 630nm, 640nm, 650nm, 660nm, 670nm, 680nm, 690nm, 700nm, 710nm, 720nm, 730nm, 740nm, 750nm, 760nm, 770nm, 780nm, 790nm, 800nm,810nm, 820nm, 830nm, 840nm, 850nm, 860nm, 870nm, 880nm, 890nm, 900nm, 910nm, 920nm, 930nm, 940nm, 950nm, 960nm, 970nm, 980nm, 990nm, 1000nm, or any number or range of these values between any two of these values.
[0177] This disclosure envisions various ratios of the quantity of one type or population of primers to the quantity of another type or population of primers. The ratio of the quantity of one type or population of primers to the quantity of another type or population of primers can be, approximately, at least, at least approximately, at most, or at most approximately: 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70, 1:69, 1:68, 1:67. 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:183:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, or the quantity or range between any two of these values.
[0178] Two adjacent primers (or any two adjacent primers) of the same type or population can form contact with each other. Two adjacent primers (or any two adjacent primers) of the same type or population cannot form contact with each other. Two adjacent primers (or any two adjacent primers) of different types or populations can form contact with each other. Two adjacent primers (or any two adjacent primers) of different types or populations cannot form contact with each other. The average distance between the positions of two adjacent or nearest capture primers attached to the solid support can be greater than (or less than or equal to) the length of one of the two capture primers, the length of the two capture primers, the average length of the two capture primers, or the total length of the two capture primers (or 0.1x, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x of the length). The average distance between the positions of two adjacent or nearest amplification primers attached to the solid support in more than one amplification primer may be greater than (or less than or equal to) the length of one of the two amplification primers, the length of the two amplification primers, the average length of the two amplification primers, or the total length of the two amplification primers (or any length of 0.1x, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x).
[0179] Two primers (or each primer) of a type or population (e.g., capture primers) can have the same length. Two primers (or each primer) of a type or population (e.g., capture primers) can have different lengths. The length of primers (or two or more primers of a type or population, or each primer of a type or population) can be: 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57. 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 9 1, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000 nucleotides. The primer (or two or more primers of a type or group, or each primer of a type or group) can have the following lengths: approximately, at least, at least approximately, at most, or at most approximately. 0.2μm, 0.3μm, 0.4μm, 0.5μm, 0.6μm, 0.7μm, 0.8μm, 0.9μm, 1μm, 2μm, 3μm, 4μm, 5μm, 6μm, 7μm, 8μm, 9μm, 10μm, or a value or range between any two of these values.
[0180] The ratio of the length of a primer of a certain type or population (e.g., capture primer) to the length of primers of the same type or population (e.g., capture primer), or the ratio of the length of a primer of a certain type or population (e.g., capture primer) to the length of primers of another type or population (e.g., amplification primer), can vary. The ratio of the lengths of two primers of one type or population, or the ratio of the lengths of two primers of different types or populations, can be as follows: approximately, at least, at least approximately, at most, or at most approximately: 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:7 1, 1:70, 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:2 6, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:170:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, or the quantity or range between any two of these values.
[0181] Any of the various polymerases can be used in the methods or apparatus described herein, for example, to replicate nucleic acid templates, form stable ternary complexes, or to modify primers. Usable polymerases include naturally occurring polymerases and their modified variants, including but not limited to mutants, recombinants, fusions, genetically modified forms, chemically modified forms, synthetics, and analogs. Naturally occurring polymerases and their modified variants are not limited to polymerases capable of catalyzing polymerization reactions. Optionally, naturally occurring and / or modified variants thereof have the ability to catalyze polymerization reactions under at least one condition not used in the formation or inspection of the stable ternary complex. Optionally, naturally occurring and / or modified variants of the stabilizing ternary complex have modified properties, such as enhanced binding affinity for nucleic acids, reduced binding affinity for nucleic acids, enhanced binding affinity for nucleotides, reduced binding affinity for nucleotides, enhanced specificity for the next correct nucleotide, reduced specificity for the next correct nucleotide, reduced catalytic rate, no catalytic activity, etc. Mutant polymerases include, for example, polymerases in which one or more amino acids are substituted by other amino acids, or polymerases in which one or more amino acids are inserted or deleted. Exemplary polymerase mutants that can be used to form stable ternary complexes include, for example, those set forth in: U.S. Patent Application Serial No. 15 / 866,353, published as U.S. Patent Application Publication No. 2018 / 0155698 A1; U.S. Patent Application Publication No. 2017 / 0314072; or U.S. Patent Application Serial No. 16 / 567,476, each of which is incorporated herein by reference.
[0182] Modified polymerases include polymerases containing an exogenous label moiety (e.g., an exogenous fluorophore) that can be used to detect the polymerase. Optionally, the label moiety can be attached after at least partial purification of the polymerase using protein separation techniques. For example, the exogenous label moiety can be covalently linked to the polymerase using a free thiol or free amine moiety of the polymerase. This can include covalent linkage to the polymerase via a side chain of a cysteine residue or via a free amino group at the N-terminus. The exogenous label moiety can also be attached to the polymerase via protein fusion. Exemplary label moiety that can be attached via protein fusion include, for example, green fluorescent protein (GFP), phycobiliproteins (e.g., phycocyanin and phycoerythrin), or wavelength-shifted variants of GFP or phycobiliproteins. In some embodiments, the exogenous label on the polymerase can function as a member of a FRET pair. The other member of the FRET pair can be an exogenous label of a nucleotide that binds to the polymerase and is attached to a stabilized ternary complex. Thus, the stabilized ternary complex can be detected or recognized by FRET.
[0183] Optionally, the polymerase involved in stabilizing the ternary complex, or the polymerase used to extend or modify primers, does not need to be attached to a foreign label. For example, the polymerase does not need to be covalently attached to a foreign label. Instead, the polymerase may lack any label until it associates with a labeled nucleotide and / or a labeled nucleic acid (e.g., a labeled primer and / or a labeled template).
[0184] Different activities of polymerases can be utilized in the methods described herein. Polymerases can be used, for example, in template amplification processes, primer modification processes (such as primer extension or primer capping steps), inspection steps, or combinations thereof. Different activities can be the result of structural differences (e.g., through natural activity, mutation, or chemical modification). However, polymerases can be obtained from a variety of known sources and applied according to the teachings set forth herein and the accepted activities of polymerases. Useful DNA polymerases include, but are not limited to, bacterial DNA polymerases, eukaryotic DNA polymerases, archaea DNA polymerases, viral DNA polymerases, and bacteriophage DNA polymerases. Bacterial DNA polymerases include E. coli DNA polymerases I, II and III, IV and V, the Klenow fragment of E. coli DNA polymerase, Clostridium stercorarium (Cst) DNA polymerase, Clostridium thermocellum (Cth) DNA polymerase, and Sulfolobus sofataricus (Sso) DNA polymerase. Eukaryotic DNA polymerases include DNA polymerases α, β, γ, δ, ε, η, ζ, λ, σ, μ, and k, as well as Rev1 polymerase (terminal deoxycytidine transferase) and terminal deoxynucleotide transferase (TdT). Viral DNA polymerases include T4 DNA polymerase, phi-29 DNA polymerase, GA-1, phi-29-like DNA polymerase, PZA DNA polymerase, phi-15 DNA polymerase, Cp1 DNA polymerase, Cp7 DNA polymerase, T7 DNA polymerase, and T4 polymerase. Other useful DNA polymerases include thermostable DNA polymerases and / or thermophilic DNA polymerases, such as *Thermus aquaticus* (Taq) DNA polymerase, *Thermus filiformis* (Tfi) DNA polymerase, *Thermococcus zilligi* (Tzi) DNA polymerase, *Thermus thermophilus* (Tth) DNA polymerase, *Thermus flavusu* (Tfl) DNA polymerase, *Pyrococcus woesei* (Pwo) DNA polymerase, *Pyrococcus furiosus* (Pfu) DNA polymerase and TurboPfu DNA polymerase, *Thermococcus litoralis* (Tli) DNA polymerase, and species of *Pyrococcus* sp.GB-D polymerase, *Thermotoga maritima* (Tma) DNA polymerase, *Bacillus stearothermophilus* (Bst) DNA polymerase, *Pyrococcus Kodakaraensis* (KOD) DNA polymerase, Pfx DNA polymerase, *Thermococcus sp.* JDF-3 (JDF-3) DNA polymerase, *Thermococcus gorgonarius* (Tgo) DNA polymerase, *Thermococcus acidophilium* DNA polymerase, *Sulfolobus acidocaldarius* DNA polymerase, *Thermococcus sp. 9°* N-7 DNA polymerase, *Pyrodictium* DNA polymerases from *Methanococcus voltae*, *Methanococcus thermoautotrophicum*, *Methanococcus jannaschii*, *D. Tok Pol* (a strain of *Desulfurococcus*), *Pyrococcus abyssi*, *Pyrococcus horikoshii*, *Pyrococcus islandicum*, *Thermococcus fumicolans*, *Aeropyrum pernix*, and heterodimeric DNA polymerases DP1 / DP2. Engineered and modified polymerases can also be used in the disclosed techniques. For example, a modified version of the extremely thermophilic marine archaea *Thermococcus* species 9°N (e.g., Therminator DNA polymerase from New England BioLabs Inc.; Ipswich, MA) can be used. Other other useful DNA polymerases (including 3PDX polymerase) are disclosed in U.S. Patent 8,703,461, the disclosure of which is incorporated herein by reference.
[0185] Useful RNA polymerases include, but are not limited to, viral RNA polymerases such as T7 RNA polymerase, T3 polymerase, SP6 polymerase, and Kll polymerase; eukaryotic RNA polymerases such as RNA polymerase I, RNA polymerase II, RNA polymerase III, RNA polymerase IV, and RNA polymerase V; and archaeal RNA polymerases.
[0186] Another useful polymerase is reverse transcriptase. Exemplary reverse transcriptases include, but are not limited to, HIV-1 reverse transcriptase from human immunodeficiency virus type 1 (PDB1HMV), HIV-2 reverse transcriptase from human immunodeficiency virus type 2, M-MLV reverse transcriptase from Moloney murine leukemia virus, AMV reverse transcriptase from avian myeloblastemia virus, and telomerase reverse transcriptase for maintaining eukaryotic chromosome telomeres.
[0187] Polymerases possessing inherent 3'-5' proofreading exonuclease activity can be useful in some implementations. Polymerases substantially lacking 3'-5' proofreading exonuclease activity are also useful in some implementations, such as in most genotyping and sequencing implementations. The absence of exonuclease activity can be a wild-type trait or a feature conferred by variant or engineered polymerase structures. For example, the exo minus Klenow fragment is a mutant form of the Klenow fragment that lacks 3'-5' proofreading exonuclease activity. The Klenow fragment and its exo minus variants can be useful in the methods or compositions described herein.
[0188] Methods can be used to determine sequences from the sense and antisense strands of a target nucleic acid to determine all or part of the target nucleic acid sequence. In some configurations, most or all of the full length of the sense and antisense strands are sequenced. Therefore, the sequences determined for the two strands will completely or almost completely overlap. In some cases, the sequence determined from the antisense strand will be complementary to the full length of the sequence determined from the sense strand, and vice versa. The more complete the overlap between the sequences determined from the two strands, the more accurate the sequencing results will be, because the two sequences can be compared to identify errors. In some cases, the sequence of one strand can be used to correct the sequence of the other strand. For example, if a difference is found between the two strands, the difference can be resolved by discarding (or statistically downgrading) the determination made in the strand with sequence motifs known to be error-prone in sequencing methods.
[0189] In other configurations of the method for determining sequences from the sense and antisense strands of a target nucleic acid, a first portion of the target sequence is determined from the antisense strand, and a second portion of the target sequence is determined from the sense strand. Depending on the length of the target sequence and the lengths of the two reads, the first and second portions may partially overlap. Therefore, the first portion is partially complementary to the second portion of the target sequence. In some cases, the sequence determined from the antisense strand may be complementary to the full length of the sequence determined from the sense strand.
[0190] In some applications of the methods described in this paper, gaps may appear between two parts when they are considered to be aligned with the complete target sequence. The length of the gap can be at least 1, 10, 100, 1000 or more bases. Optionally or additionally, the length of the gap can be at most 1000, 100, 10 or 1 base. Knowledge of the opposite orientation of the two sequences and the size of the gap between the two sequences can be advantageous for alignment with a reference genome. Information related to the relative orientation of the two sequences not only improves genome reconstruction by alignment with a reference genome, but also allows for the identification of structural variations in the genome. The expected genome alignment deviation between the two ends of a paired end read can indicate the structural variations of the sequenced sample compared to the reference sequence aligned with the read.
[0191] This disclosure provides systems configured to perform the methods described herein. For example, a system may be configured to generate and detect a ternary complex formed between a polymerase and a primer-template nucleic acid hybrid in the presence of nucleotides to recognize one or more bases in a template nucleic acid sequence. Systems of this disclosure may include containers, solid supports, or other devices for performing nucleic acid amplification and / or nucleic acid detection methods. For example, the system may include arrays, flow cells, multiwell plates, or other convenient devices. Devices may be removable, allowing them to be placed into or removed from the system. Therefore, the system may be configured to process more than one device (e.g., container or solid support) sequentially or in parallel. The system may include a fluid component having a reservoir for containing one or more reagents described herein (e.g., polymerase, primers, template nucleic acid, nucleotides for ternary complex formation, nucleotides for primer extension, unblocking reagents, ternary complex inhibitors, or mixtures of these components). The fluid system may be configured to deliver reagents to the container or solid support, for example, via channels or droplet transfer devices (e.g., electrowetting devices). Any of a variety of detection devices can be configured to detect containers or solid supports in which reagents interact. Examples include luminescent detectors, surface plasmon resonance detectors, and other detectors known in the art. Exemplary systems having fluid and detection components that can be readily modified for use in the systems described herein include, but are not limited to, those set forth below: U.S. Patent Application Publication No. 2018 / 0280975A1, which claims priority to U.S. Patent Application Serial No. 62 / 481,289; U.S. Patent Nos. 8,241,573, 7,329,860, or 8,039,817; or U.S. Patent Application Publication Nos. 2009 / 0272914A1 or 2012 / 0270305A1, each of which is incorporated herein by reference.
[0192] Optionally, the system of this disclosure also includes a computing system or components thereof. The computing system may include a computer processing unit (CPU) configured as an operating system component. The CPU may include one or more processors or processing units. The same or different CPUs may interact with the system to acquire, store, and process signals (e.g., signals detected in the methods described herein). In certain embodiments, the CPU can be used to determine the identity of a nucleotide present at a specific location in the template nucleic acid from the signal. In some cases, the CPU will identify the nucleotide sequence of the template from the detected signal.
[0193] The computing system can be a personal computer system, server computer system, thin client, fat client, handheld or laptop device, multiprocessor system, microprocessor-based system, set-top box, programmable consumer electronics, network PC, minicomputer system, mainframe computer system, smartphone, and a distributed cloud computing environment that includes any of the above systems or devices. The computing system may include a memory architecture that includes RAM and non-volatile memory. The memory architecture may also include removable / non-removable, volatile / non-volatile computer system storage media. Furthermore, the memory architecture may include one or more readers for reading and writing from non-removable, non-volatile magnetic media, such as hard disk drives; disk drives for reading and writing from removable, non-volatile disks; and / or optical disk drives for reading or writing from removable, non-volatile optical disks, such as CD-ROMs or DVD-ROMs. The computing system may also include a variety of computer system readable media. Such media can be any available media accessible in a cloud computing environment, such as volatile and non-volatile media, as well as removable and non-removable media.
[0194] The memory architecture may include at least one program product having at least one program module, which is implemented to execute instructions configured to perform one or more steps of the methods described herein. For example, executable instructions may include an operating system, one or more application programs, other program modules, and program data. Typically, a program module may include routines, programs, objects, components, logic, data structures, etc., that perform the specific task described herein.
[0195] CPU components can be coupled via an internal bus, which can be implemented as one or more of several types of bus architectures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using any of a variety of bus architectures. Such architectures include, by way of example and not limitation, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, the Enhanced ISA (EISA) bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0196] The CPU can optionally communicate with one or more external devices, such as a keyboard, pointing device (e.g., mouse), display (e.g., graphical user interface (GUI)), or other devices that facilitate user interaction with the nucleic acid detection system. Similarly, the CPU can communicate with other devices (e.g., via a network interface card, modem, etc.). This communication can occur through the I / O interface. Furthermore, the CPU of the system described herein can communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a suitable network adapter.
[0197] Example I
[0198] Paired-end sequencing via two 22-cycle runs (PE2x22)
[0199] This embodiment describes a method for generating clusters with sense chains and antisense chains using simultaneous RCA and MDA, as referenced herein. Figure 1D As described. This embodiment demonstrates the sequencing of two strands in a paired-end sequencing method. Furthermore, this embodiment shows that both strands can be sequenced without removing one strand to sequence the other.
[0200] Materials and methods
[0201] The flow cell containing the template nucleic acid was prepared as follows. A layer of SHARK248 primer (5'-CGCCGTATCATTCAAGCAGAAGAC*G*G-3', where the asterisk represents a phosphate thioester bond; SEQ ID NO:6) was attached to the inner surface of the flow cell by click chemistry. The template loop hybridized with the SHARK248 primers via a universal adaptor complementary to the SHARK248 primers, and was constructed using a scaling primer containing 33 mM Tris-HCl pH 8.0 (Sigma-Aldrich, St. Louis, MO), 30 mM MgCl2 (Invitrogen, Carlsbad, CA), 10 mM DTT (Sigma-Aldrich, St. Louis, MO), 90 mM KCl (Teknova, Hollister, CA), 2% sucrose (Sigma-Aldrich, St. Louis, MO), 0.5 M betaine (Sigma-Aldrich, St. Louis, MO), 0.2% Tween-80 (Sigma-Aldrich, St. Louis, MO), 0.8 mM dNTPs (New England Biolabs, Ipswich, MA), and 0.1 μM scale formation primers. Primer) SHARK291 (5'-ATCTCGTATGCCGTCTTCTGCTT*G-3', where the asterisk indicates a phosphate thioester bond; SEQ ID NO: 7) was used as the Phi29 DNA polymerase (Thermo Fisher Waltham, MA) Scientific mixture extension primer. The RCA / MDA reaction was performed at 37°C for 4.5 h, with Phi29 DNA polymerase replenished every 15 minutes. After extension, Phi29 was removed by heat denaturation, followed by washing in a buffer containing 40 mM Tris-HCl pH 8.0, 110 mM KCl, 0.02 mM EDTA (Sigma-Aldrich, St. Louis, MO), and 0.1% Tween-80.
[0202] The clusters generated by the above method are then processed to add a capping portion to any extendable 3' end of the cluster. The capping reagent is prepared as a mixture containing 50 mM trimethylglycine (Sigma-Aldrich, St. Louis, MO), U.S. Patent Application Publication No. 2020 / 0032322A1 (which is incorporated herein by reference). Figure 5The four biotinylated dideoxynucleotides shown are each 2 μM, 0.1 mM biotin (Sigma-Aldrich, St. Louis, MO), 50 mM KCl (Teknova, Hollister, CA), 0.1% Tween-80 (Sigma-Aldrich, St. Louis, MO), 5 mM MgCl2 (Invitrogen, Carlsbad, CA), 40 U / ml M15 DNA polymerase (see U.S. Patent Application Serial No. 16 / 567,476, which is incorporated herein by reference), and 0.1 mM EDTA (Invitrogen, Carlsbad, CA). The capping reagent is introduced into the flow cell, and the capping reaction is allowed to proceed at 55°C for 2 minutes. The capping reagent is removed, and the flow cell is washed with high salt to remove any bound DNA polymerase. The next step in the capping process is incubation of the biotinylated primer extension product with streptavidin. The streptavidin mixture contained 50 mM trimethylglycine (Sigma-Aldrich, St. Louis, MO), 0.076 mg / ml streptavidin (New England Biolabs, Ipswich, MA), 50 mM KCl (Teknova, Hollister, CA), 0.1% Tween-80 (Sigma-Aldrich, St. Louis, MO), 5 mM MgCl2 (Invitrogen, Carlsbad, CA), and 0.1 mM EDTA (Invitrogen, Carlsbad, CA). The streptavidin mixture was introduced into the flow cell and allowed to bind at 55°C for 2 minutes. Prior to the subsequent sequencing step, the clusters were treated with formamide to denature any double-stranded regions.
[0203] More than one SHARK231 primer (5'-GTGACTGGAGTTCAGACGTGTGCTCTTC-3'; SEQ ID NO:8) was hybridized to the complementary primer binding site in the antisense strand connective region of the cluster. Sequencing ByBinding was performed using cycles. TM (SBB TM The method sequenced the antisense strand of the cluster, wherein each cycle included the following steps: (i) extension: adding a reversibly terminated nucleotide to the primer of a fixed primer-template hybrid; (ii) detection: forming and detecting a stable ternary complex on the fixed primer-template hybrid with reversible termination; and (iii) activation: cleaving the reversible terminator from the extended primer. Each cycle resulted in the addition of a single nucleotide and subsequent detection of the nucleotide position.
[0204] Sequencing cycles are initiated by incorporating a reversible terminator nucleotide at the 3' end of a primer of a fixed primer template heterozygote. This is accomplished via an extension step in which the flow cell is contacted with an unlabeled reversible terminator nucleotide (an analogue of dATP, dGTP, dCTP, and dTTP) in the presence of M15 polymerase (described in U.S. Patent Application Sequence No. 16 / 567,476, which is incorporated herein by reference). The reversible terminator nucleotide used in this illustrative procedure comprises a 3'-ONH2 reversible terminator moiety. A description of such a reversible terminator nucleotide can be found in U.S. Patent No. 7,544,794 (which is incorporated herein by reference).
[0205] Next, a washing step is performed to remove dNTPs from the flow-through cell. The washing solution contains isopropanol, Tween-80, hydroxylamine, and EDTA. The washing step retains the M15 polymerase (described in U.S. Patent Application Serial No. 16 / 518,321, which is incorporated herein by reference).
[0206] The cycle then proceeds to a check subroutine in which each of four different nucleotides is individually delivered to the flow cell (Cy5-labeled dTTP, Cy5-labeled dATP, Cy5-labeled dCTP, and Cy5-labeled dGTP). The system pauses fluid flow to allow ternary complex formation, free nucleotides are removed from the flow cell by delivery of imaging fluid, and the flow cell is then checked for ternary complex formation at the immobilized primer-template hybrid. The imaging fluid contains LiCl, betaine, Tween-80, KCl, ammonium sulfate, hydroxylamine, and EDTA, which stabilize the ternary complex after removal of free nucleotides (see U.S. Patent No. 10,400,272, which is incorporated herein by reference). The flow cell is imaged by fluorescence microscopy to detect the ternary complex containing labeled nucleotides, which are homologs of the next correct nucleotide in each template nucleic acid. During the ternary complex formation and detection steps, a reversible terminator portion on the 3' nucleotide of the primer strand excludes nucleotide incorporation.
[0207] After the check subroutine, the flow cell is rinsed with washing buffer to remove nucleotides from the check subroutine. The sequencing cycle then proceeds to a lysis step, in which reversible terminator portions are removed from the primers using sodium acetate and sodium nitrite, as described in U.S. Patent No. 7,544,794 (which is incorporated herein by reference). The lysis reagent is then removed, and the flow cell is washed with imaging fluid to remove residual polymerase from the check step. The sequencing process then proceeds to the next nucleotide position by returning to the first step of the next sequencing cycle.
[0208] In performing the above SBB TMAfter 27 cycles of the method (the first five cycles are control cycles), the extended primers generated by the sequencing reaction are capped using the capping reagents and methods described earlier in this embodiment.
[0209] After adding the cap, perform the above SBB procedure. TM Five loops of the method are used as a control to demonstrate the effectiveness of the capping process.
[0210] After 5 control cycles, more than one SHARK276 primer (5'-CGGCGACCACCGAGATCGGCGACCACCGAGATCTACACTCTTTCC CTACACGACGCTCTTCCGATCT-3'; SEQ ID NO:9) was hybridized to the complementary primer binding site of the sense strand connective region in the cluster. SBBs were used for the antisense strand. TM The method sequenced the sense strand of the cluster.
[0211] result
[0212] The results from the sequencing process described above are analyzed below. The "open" intensity (the brightest nucleotide signal intensity obtained from a given cluster in a given inspection step) is tracked to monitor the formation of ternary complexes. Figure 5 A plot showing the 50th percentile of the normalized "on" intensity on the y-axis and the cycle number on the x-axis is presented. Data from each of the four nucleotide types are plotted as separate lines. The data points from the first five cycles (control cycles) appear noisy in both the "on" and "off" plots because the first five nucleotides are common across all evaluated clusters. Data points from subsequent cycles are more informative of the signal-to-noise ratio levels achieved during sequencing, as each data point is an average of all four nucleotide types spanning more than one different template.
[0213] Figure 5 The results demonstrate that antisense sequencing (cycles 6-27 in the figure) achieved excellent separation between "on" and "off" signals. These results indicate that the presence of a second strand in the cluster does not significantly adversely affect the sequencing of the first strand. Furthermore, the signal attenuation during first-strand sequencing is comparable to that of previous SBBs performed on arrays with only one strand per site. TM The observed signal attenuation was comparable.
[0214] Figure 5 The results show that only a negligible "on" signal was detected during 5 control cycles (cycles 28–32). This indicates that effective 3' capping was performed on the extended primers and other strands present after antisense sequencing.
[0215] Figure 5The results demonstrate that sense strand sequencing (cycles 33-54 in the figure) achieved excellent separation between "on" and "off" signals. In fact, the second read had a higher, on average, "on" signal compared to the first read, while the "off" signal was comparable between the two reads. These results indicate that the presence of antisense strands within the cluster does not significantly adversely affect sense strand sequencing. Furthermore, the signal attenuation during the first-strand sequence reads was comparable to the signal attenuation observed from the second-strand sequence reads.
[0216] Paired-end alignment analysis was performed on a subset of sequenced clusters using Bowtie 2 (Langmead B, Salzberg S. Fast gapped-read alignment with Bowtie 2. Nature Methods, 9:357-359 (2012), which is incorporated herein by reference). For this analysis, 574 reads were evaluated. Of these, 574 (100.00%) were paired; 7 (1.22%) were co-aligned 0 times; 567 (98.78%) were co-aligned exactly once, and 0 (0.00%) were co-aligned more than once. A total of 7 read pairs were co-aligned 0 times; of these, 4 (57.14%) were not co-aligned once. Three read pairs showed 0 instances of consistent or inconsistent alignment; six pairs were formed; two (33.33%) showed 0 instances of alignment; three (50.00%) showed exactly 1 instance of alignment; and one (16.67%) showed more than 1 instance of alignment. Overall, 98.7% of paired end reads were consistently aligned, demonstrating accurate sense and antisense sequencing from a single cluster.
[0217] Example II
[0218] Paired-end sequencing via two 100-cycle runs (PE2x100)
[0219] This embodiment confirms and extends the results of Embodiment I, demonstrating that each cluster can be sequenced from two strands, with at least 100 cycles per strand. The results confirm that both strands can be sequenced without removing one strand to sequence the other.
[0220] Materials and methods
[0221] Clusters were prepared using the simultaneous RCA and MDA methods described in Example I. The antisense and sense strands were sequenced as described in Example I, except that each strand was sequenced for 100 cycles.
[0222] result
[0223] The sequencing results were analyzed as described in Example I. Figure 6The results demonstrate that antisense sequencing (cycles 1-100 in the figure) achieved excellent separation between "on" and "off" signals, and further confirm the signal attenuation during first-strand sequencing (antisense sequencing) compared to previous SBBs performed on arrays with only a single strand at each site. TM The observed signal attenuation was comparable. Results from five control cycles (cycles 101-105 in the figure) indicate effective capping of the 3' ends in the clusters after the first sequencing read. Results from sense strand sequencing (cycles 106-205 in the figure) demonstrate excellent separation between the "on" and "off" signals. These results indicate that the presence of antisense strands in the clusters does not significantly adversely affect sense strand sequencing. Furthermore, the signal attenuation during the first strand sequence read was comparable to that observed from the second strand sequence read.
[0224] Table 1: Independent analysis of the chain
[0225] antisense chain There is a chain of righteousness Cluster count 28437 30835 Percentage of aligned clusters 90.03 91.8 Q score 32.9 33.4
[0226] Bowtie 2 was used in conjunction with Omniome sequencing software to evaluate 100 cyclic runs of paired ends. When each strand was analyzed independently, over 90% of reads were aligned with the reference genome (Table 1). Epigenetic Q scores of 32.9 and 33.4 were observed for the antisense and sense strands, respectively (Table 1). Furthermore, paired end reads as buddy pairs were analyzed using Bowtie 2. In this case, 81.9% of read pairs were consistently mapped to the reference genome (Table 2). Buddy pair distances calculated by Bowtie 2 were plotted as a distribution to further demonstrate, along with the number of consistent read pairs, the specificity of priming and sequencing the two strands of DNA within each cluster. See also Figure 7 This demonstrates that these clusters were sequenced from the relative ends of the library fragments. Partner pair distances were mapped and aligned with the expected original fragment sizes of the sequenced library.
[0227] Table 2: Bowtie 2 Analysis
[0228] % of the total Consistently aligned read pairs 81.92% Unaligned reading segments 0.65% One partner in a pair aligns 3.79% Overall alignment 84.46%
[0229] Paired-end alignment analysis was performed on a subset of the sequenced clusters using Bowtie 2. For this analysis, 21,418 reads were evaluated. Of these, 21,418 (100.00%) were paired; 3,873 (18.08%) were co-aligned 0 times; 17,545 (81.92%) were co-aligned exactly once, and 0 (0.00%) were co-aligned more than once. Of the 3,873 pairs with 0 co-alignments, 138 pairs (3.56%) were not co-aligned once. Within the reads, 3,735 pairs were co-aligned or not co-aligned 0 times; 7,470 pairs were paired; 6,657 (89.12%) were aligned 0 times; 813 (10.88%) were aligned exactly once; and 0 (0.00%) were aligned more than once.
[0230] Example III
[0231] Enzymatic digestion of the second strand and primer extension products was performed before sequencing the first strand.
[0232] The first and second chains containing the target sequence (insertion) are generated using RCA and MDA, as referenced herein. Figure 1E The description describes a second strand containing a uracil base. The 3' end of the antisense strand contains a biotin-conjugated nucleotide. The 3' end of the antisense strand is sealed with streptoacidin, which binds to the biotin-conjugated nucleotide. After initiating the second strand with sequencing primers and sequencing the second strand using extension sequencing primers, the second strand and the extension sequencing primers are digested with enzymes, as described in the references herein. Figure 2D Described. Uracil DNA glycosylase (UDG) and endonuclease VIII are used to generate a nick in the second strand. The remaining antisense strand and extended primer are digested from the 5' ends (the 5' ends of the antisense strand and the extended primer, as well as the 5' end at the antisense strand nick) using T7 exonuclease. The size of the target sequence is determined using paired end reads. Figure 8A This is a histogram showing that the method does not have a preference for target sequences of a specific size. The sequencing fluorescence intensity of paired reads in... Figure 8B The results were presented in [the diagram]. Each reaction was run for 50 cycles. The observation that the signal intensity of the second read was comparable to that of the first read demonstrates the effectiveness of our method in generating a population suitable for paired read sequencing.
[0233] Example IV
[0234] Rolling circle amplification and multiple displacement amplification using capture primers and amplification primers attached to solid surfaces
[0235] Figure 9A The diagram shows high-quality signal density from sequencing the first strand generated from a nucleic acid template using rolling circle amplification with capture primers attached to a solid support. Figure 9BThe diagram shows a high-quality signal density for sequencing a second strand generated from a first strand, which was produced using multiple displacement amplification with amplification primers attached to a solid support, while the first strand was generated from a nucleic acid template using rolling circle amplification with capture primers attached to a solid support.
[0236] Numerous publications, patents, and / or patent applications have been referenced throughout this application. The disclosures of these documents are hereby incorporated in their entirety by reference.
[0237] In at least some of the previously described embodiments, one or more elements used in one embodiment may be used interchangeably in another embodiment, unless such substitution is technically impractical. Those skilled in the art will understand that various other omissions, additions, and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter defined by the appended claims.
[0238] Regarding the use of substantially any plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural where appropriate for the context and / or application. For clarity, various singular / plural arrangements may be explicitly set forth herein. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context explicitly indicates otherwise. Unless otherwise stated, any reference to “or” herein is intended to cover “and / or.”
[0239] Those skilled in the art will understand that, generally, the terminology used herein, and particularly in the appended claims (e.g., the body of the appended claims), is generally intended to be “open-ended” terms (e.g., the term “including” should be interpreted as “including but not limited to”, the term “having” should be interpreted as “having at least”, the term “includes” should be interpreted as “includes but is not limited to”, etc.). Those skilled in the art will further understand that if a particular number is intended to be introduced in the claim statement, such an intention will be explicitly stated in the claim, and if such a statement is absent, such an intention does not exist. For example, to aid understanding, the appended claims may include the introductory phrases “at least one” and “one or more” to introduce the claim statement. However, the use of such wording should not be construed as meaning that introducing a claim statement with the indefinite article “a(a)” or “an” would limit any specific claim in a claim statement containing such an introduction to embodiments containing only one such statement, even when the same claim includes the introductory wording “one or more” or “at least one” and indefinite articles such as “a(a)” or “an” (e.g., “a(a)” and / or “an” should be interpreted as meaning “at least one” or “one or more”); the same applies to the use of definite articles to introduce a claim statement. Furthermore, even if a specific number in an introductory claim statement is explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as meaning at least the number stated (e.g., simply stating “two statements” without other modifiers means at least two statements or two or more statements). Furthermore, in cases where conventions like "at least one of A, B, and C" are used, this syntactic structure is generally intended to be understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" will include, but is not limited to, systems having a single A, having a single B, having a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In cases where conventions like "at least one of A, B, or C" are used, this syntactic structure is generally intended to be understood by a person skilled in the art (e.g., "a system having at least one of A, B, or C" will include, but is not limited to, systems having a single A, having a single B, having a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.).Those skilled in the art will further understand that, in practice, any separate words and / or wording presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to account for the possibility of including one, any, or both terms. For example, the phrase “A or B” should be understood to include the possibility of “A” or “B” or “A and B”.
[0240] Furthermore, when features or aspects of this disclosure are described in terms of the Markush group, those skilled in the art will recognize that this disclosure is also described in terms of any individual member or subgroup of the Markush group.
[0241] As those skilled in the art will understand, for any and all purposes, such as providing a written description, all scopes disclosed herein also include any and all possible subscopes and combinations of subscopes. Any enumerated scope can be readily considered sufficiently described and such that the same scope can be divided into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each scope discussed herein can be readily divided into lower thirds, middle thirds, and upper thirds, etc. As those skilled in the art will also understand, all linguistic terms such as “up to,” “at least,” “greater than,” “less than,” etc., include the cited numbers and refer to a scope that can subsequently be divided into subscopes as discussed above. Finally, as those skilled in the art will understand, a scope includes members of each individual. Thus, for example, a group having 1-3 items means a group having 1, 2, or 3 items. Similarly, a group having 1-5 items means a group having 1, 2, 3, 4, or 5 items, and so on.
[0242] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and are not intended to be limiting, and the actual scope and spirit are indicated by the following claims.
[0243] Many embodiments have been described. However, it should be understood that various modifications can be made. Therefore, other embodiments are within the scope of the appended claims. sequence list <110> California Pacific Biosciences Co., Ltd. Harry K.K. Subramaniam Aaron W. Feldman Zhi-Yuan Chen D. Malyshev Greg Richmond <120> Methods and compositions for sequencing double-stranded nucleic acids <130> 42HB-328033-WO <150> 62 / 984,438 <151> 2020-03-03 <160> 9 <170> PatentIn version 3.5 <210> 1 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Synthetic peptides <400> 1 Glu Gln Lys Leu Ile Ser Glu Glu Asp Leu 1 5 10 <210> 2 <211> 26 <212> PRT <213> Artificial Sequence <220> <223> Synthetic peptides <400> 2 Lys Arg Arg Trp Lys Lys Asn Phe Ile Ala Val Ser Ala Ala Asn Arg 1 5 10 15 Phe Lys Lys Ile Ser Ser Ser Gly Ala Leu 20 25 <210> 3 <211> 7 <212> PRT <213> Artificial Sequence <220> <223> Synthetic peptides <400> 3 Asp Tyr Lys Asp Asp Asp Asp 1 5 <210> 4 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Synthetic peptides <400> 4 Asp Tyr Lys Asp Asp Asp Asp Lys 1 5 <210> 5 <211> 7 <212> PRT <213> Artificial Sequence <220> <223> Synthetic peptides <400> 5 Asp Tyr Lys Asp Asp Asp Lys 1 5 <210> 6 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> Synthetic oligonucleotides <220> <221> misc_feature <222> (24) (25) <223> Thiophosphate bond <220> <221> misc_feature <222> (25)..(26) <223> Thiophosphate bond <400> 6 cgccgtatca ttcaagcaga agacgg 26 <210> 7 <211> twenty four <212> DNA <213> Artificial Sequence <220> <223> Synthetic oligonucleotides <220> <221> misc_feature <222> (23)..(24) <223> Thiophosphate bond <400> 7 atctcgtatg ccgtcttctg cttg 24 <210> 8 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> Synthetic oligonucleotides <400> 8 gtgactggag ttcagacgtg tgctcttc 28 <210> 9 <211> 67 <212> DNA <213> Artificial Sequence <220> <223> Synthetic oligonucleotides <400> 9 cggcgaccac cgagatcggc gaccaccgag atctacactc tttccctaca cgacgctctt 60 ccgatct 67
Claims
1. A composition comprising A solid support to which a nucleic acid cluster is attached, the nucleic acid cluster comprising a sense strand and an antisense strand of a tandem strand, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site, wherein the sense strand is covalently attached to the solid support, and wherein the antisense strand is not covalently bonded to the solid support but is attached to the nucleic acid cluster by pairing with the Watson-Crick bases of the sense strand.
2. The composition according to claim 1, further comprising a first primer that hybridizes to a primer binding site in the sequence unit of the antisense strand.
3. The composition of claim 2, comprising an extended first primer that hybridizes to a primer binding site in the sequence unit of the antisense strand.
4. The composition of claim 3, wherein the extended first primer comprises a labeled reversibly terminated nucleotide.
5. The composition of claim 4, wherein the reversibly terminated nucleotide is covalently attached to the extended first primer.
6. The composition of claim 4, wherein the reversibly terminated nucleotide is unblocked.
7. The composition according to claim 4, wherein the marking is removed.
8. The composition of claim 3, wherein the extended first primer comprises a closing portion.
9. The composition of claim 3, wherein the extended first primer comprises a capped portion.
10. The composition according to claim 3, wherein the primer binding site, together with the first primer, polymerase and nucleotide hybridized thereto, forms a ternary complex.
11. The composition of claim 3, wherein the extended first primer is at least 100 nucleotides longer than the first primer.
12. The composition of claim 1, further comprising a second primer that hybridizes to a primer binding site in the sequence unit of the sense strand.
13. The composition of claim 12, comprising an extended second primer that hybridizes to a primer binding site in the sequence unit of the sense strand.
14. The composition of claim 13, wherein the extended second primer comprises a labeled reversibly terminated nucleotide.
15. The composition of claim 14, wherein the reversibly terminated nucleotide is covalently attached to the extended second primer.
16. The composition of claim 14, wherein the reversibly terminated nucleotide is unblocked.
17. The composition of claim 14, wherein the marking is removed.
18. The composition of claim 13, wherein the extended second primer comprises a closing portion.
19. The composition of claim 13, wherein the extended second primer comprises a capped portion.
20. The composition of claim 13, wherein the primer binding site, together with the second primer, polymerase and nucleotide hybridized thereto, forms a ternary complex.
21. The composition of claim 13, wherein the extended second primer is at least 100 nucleotides longer than the second primer.
22. The composition of claim 1, wherein the antisense strand is synthesized from the sense strand using amplification primers that bind to primer binding sites in the sequence units of the sense strand.
23. The composition of claim 22, wherein the amplification primer is non-covalently attached to a primer binding site in a sequence unit of the sense strand.
24. The composition of claim 1, wherein the solid support comprises a capture primer, wherein the sense strand is synthesized by a capture primer that extends along the nucleic acid template via rolling circle amplification and hybridizes with the nucleic acid template, and wherein the antisense strand is synthesized by an amplification primer that extends to bind to a primer binding site in a sequence unit of the sense strand.
25. The composition according to claim 24, wherein the nucleic acid template is a circular nucleic acid template.
26. The composition of claim 24, wherein the nucleic acid template is circularized from a linear nucleic acid template.
27. The composition of claim 24, wherein the capture primer and the amplification primer are extended using a strand displacement polymerase.
28. The composition of claim 24, wherein the amplification primer is non-covalently attached to a primer binding site in a sequence unit of the sense strand.
29. The composition of claim 1, wherein the cluster comprises more than one antisense chain of the tandem strand.
30. The composition of claim 29, wherein the more than one antisense strand exceeds the number of sense strands in the cluster.
31. The composition of claim 29, wherein the cluster comprises more than one meaningful chain of the tandem.
32. The composition of claim 31, wherein the number of the more than one sense strand exceeds the number of the more than one antisense strand in the cluster.
33. The composition of claim 1, wherein the antisense strand of the tandem strand has fewer copies of sequence units than the sense strand.
34. The composition of claim 1, wherein the antisense strand comprises a first portion of the target sequence or its inverse complement, and wherein the sense strand comprises a second portion of the target sequence or its inverse complement.
35. The composition of claim 34, wherein the target sequence includes a gap between the first portion of the target sequence and the second portion of the target sequence.
36. The composition of claim 34, wherein the first portion of the target sequence is partially complementary to the second portion of the target sequence.
37. The composition of claim 34, wherein the first portion of the target sequence is complementary to the full length of the second portion of the target sequence.
38. The composition of claim 34, wherein the second portion of the target sequence is complementary to the full length of the first portion of the target sequence.
39. The composition of claim 1, wherein the target sequence comprises at least 100 base pairs.
40. The composition of claim 1, wherein the sense strand is generated from a nucleic acid template by isothermal amplification.
41. The composition of claim 40, wherein the isothermal amplification is rolling circle amplification.
42. The composition according to claim 40 or 41, wherein the nucleic acid template is derived from or generated from the sample without polymerase chain reaction.
43. The composition according to claim 40 or 41, wherein the nucleic acid template is derived or generated from the sample by performing a polymerase chain reaction for up to five cycles.
44. The composition of claim 1, wherein the antisense strand comprises one or more modified or atypical nucleotides.
45. The composition of claim 1, wherein the antisense strand comprises one or more nucleotides having modified or atypical bases, wherein the bases are uracil.
46. The composition according to claim 44 or 45, wherein the ratio of the number of uracil bases to the number of thymine or non-uracil bases in the antisense chain is from 1:1000 to 1:
10.
47. The composition according to claim 44 or 45, wherein the percentage of uracil bases in the antisense chain is 0.001% to 1%.
48. The composition of claim 1, wherein the antisense strand is synthesized by extending a primer in the presence of a deoxyribonucleotide triphosphate comprising dATP, dTTP, dGTP, dCTP and dUTP.
49. The composition according to claim 48, wherein the ratio of dUTP to another deoxyribonucleic acid triphosphate or all other deoxyribonucleic acid triphosphate is 1:1000 to 1:
10.
50. The composition of claim 48, wherein the concentration of dUTP is from 0.01 mM to 1 mM.
51. A method for determining sequences from the sense and antisense strands of nucleic acids, comprising: (a) Providing a nucleic acid cluster attached to a solid support, wherein the nucleic acid cluster comprises a sense strand of a tandem strand and an antisense strand of the tandem strand, wherein the tandem strand comprises more than one copy of a tandemly linked sequence unit, wherein the sequence unit comprises a target sequence and a primer binding site; (b) Hybridize the primers to primer binding sites in the sequence units of the antisense strand in the cluster; (c) Extend the primer along the antisense strand to determine a sequence from at least a portion of the target sequence in the antisense strand; (d) Remove the antisense chain from the cluster; (e) Hybridize the second primer to the primer binding site in the sequence unit of the sense strand; and (f) Extend the second primer along the sense strand to determine a sequence from at least a portion of the target sequence in the sense strand.
52. The method of claim 51, further comprising, prior to step (e), adding a closing portion to the extended primer to prevent the extended primer from extending further during step (f).
53. The method of claim 51, further comprising, prior to step (e), adding a capping portion to the extended primer to prevent polymerase from binding to the 3' end of the extended primer during step (f).
54. The method of claim 51, wherein extending the primer in step (c) comprises (i) adding a reversibly terminating nucleotide to the primer and (ii) a repeated cycle of unblocking the reversibly terminating nucleotide on the primer.
55. The method of claim 54, wherein the reversibly terminated nucleotide comprises a marker detected to generate a signal for determining a sequence from at least a portion of the target sequence in the antisense strand.
56. The method of claim 55, wherein the repeated loop further comprises (iii) removing the mark after the mark has been detected.
57. The method of claim 54, wherein the repeated cycle further comprises (iii) detecting a stabilized ternary complex comprising a polymerase, the next correct nucleotide, and the primer hybridizing with the antisense strand.
58. The method of claim 57, wherein the next correct nucleotide or the polymerase comprises a marker detected to generate a signal for determining a sequence from at least said portion of the target sequence in the antisense strand.
59. The method of claim 57, wherein during step (c)(iii), the reversibly terminated nucleotide is covalently attached to the primer.
60. The method of claim 54, wherein the extension of the primer in step (c) comprises at least 100 repeating cycles to determine a sequence from at least 100 bases of the target sequence in the antisense strand.
61. The method of claim 51, wherein the extension of the second primer in step (f) comprises (i) adding a reversibly terminating nucleotide to the second primer and (ii) unblocking the reversibly terminating nucleotide on the second primer in a repeated cycle.
62. The method of claim 61, wherein the reversibly terminated nucleotide comprises a marker detected to generate a signal for determining a sequence from at least said portion of the target sequence in the sense strand.
63. The method of claim 62, wherein the repeated loop further comprises (iii) removing the mark after the mark has been detected.
64. The method of claim 61, wherein the repeated cycle comprises (iii) detecting a stabilized ternary complex comprising a polymerase, the next correct nucleotide, and the second primer hybridizing with the sense strand.
65. The method of claim 64, wherein the next correct nucleotide or the polymerase comprises a marker detected to generate a signal for determining a sequence from at least said portion of the target sequence in the sense strand.
66. The method of claim 64, wherein during steps (f) and (iii), the reversibly terminated nucleotide is covalently attached to the second primer.
67. The method of claim 61, wherein the extension of the second primer in step (f) comprises at least 100 repeating cycles to determine a sequence from at least 100 bases of the target sequence in the sense strand.
68. The method of claim 51, wherein step (a) comprises (i) Providing the meaningful chain of the tandem body on the solid support, and (ii) The antisense strand is synthesized using amplification primers, which bind to the primer binding site in the sequence unit of the sense strand.
69. The method of claim 68, wherein the amplification primers are covalently attached to the solid support.
70. The method of claim 68, wherein the amplification primers are not covalently attached to the solid support.
71. The method according to any one of claims 51-67, wherein step (a) comprises (i) Provide a solid support containing the capture primers. (ii) Hybridize the nucleic acid template with the capture primers. (iii) The sense strand is synthesized by extending the capture primer along the nucleic acid template via rolling circle amplification, and (iv) The antisense strand is synthesized by extending amplification primers, the amplification primers binding to the primer binding site in the sequence unit of the sense strand.
72. The method according to claim 71, wherein the nucleic acid template is a circular nucleic acid template.
73. The method of claim 71 further comprises circularizing the nucleic acid template prior to steps (a) and (ii).
74. The method of claim 71 further comprises circularizing the nucleic acid template after step (a)(ii) and before step (a)(iii).
75. The method of claim 71, wherein the extension of the capture primer and the extension of the amplification primer occur simultaneously.
76. The method of claim 71, wherein during steps (a) and (iii), the capture primer hybridizes with the primer binding site in the sequence unit of the sense strand.
77. The method of claim 71, wherein the extension of the capture primer and the extension of the amplification primer are performed using a strand displacement polymerase.
78. The method of claim 71, wherein the amplification primers are covalently attached to the solid support.
79. The method of claim 71, wherein the amplification primers are not covalently attached to the solid support.
80. The method of claim 51, further comprising, prior to step (b), adding a blocking portion to the 3' end of the nucleic acid cluster attached to the solid support.
81. The method of claim 51, further comprising, prior to step (b), adding a capping portion to the 3' end of the nucleic acid cluster attached to the solid support.
82. The method of claim 51, wherein the cluster comprises more than one antisense chain of the concatenation.
83. The method of claim 82, wherein the number of more than one antisense chain exceeds the number of sense chains in the cluster.
84. The method of claim 82, wherein steps (e) and (f) are performed prior to steps (b) and (c).
85. The method of claim 82, wherein the cluster comprises more than one meaningful chain of the tandem.
86. The method of claim 85, wherein the number of the more than one sense chain exceeds the number of the more than one antisense chain in the cluster.
87. The method of claim 86, wherein steps (e) and (f) are performed prior to steps (b) and (c).
88. The method of claim 51, wherein the sequence is determined under isothermal conditions.
89. The method of claim 51, wherein the series formation is carried out under isothermal conditions.
90. The method of claim 51, wherein all sequence templates are generated before any sequence is determined.
91. The method of claim 51, wherein a first sequencing reaction template and a second sequencing reaction template are generated prior to sequence determination in step (c).
92. The method of claim 51, wherein a first sequencing reaction template and a second sequencing reaction template are generated prior to the sequence determination in step (c) and the sequence determination in step (f).
93. The method of claim 51, wherein the antisense chain of the tandem has fewer copies of sequence units than the sense chain.
94. The method of claim 93, wherein steps (e) and (f) are performed prior to steps (b) and (c).
95. The method of claim 51, wherein the nucleic acid cluster is attached to the solid support by covalent attachment of the sense strand to the solid support.
96. The method of claim 95, wherein the antisense strand is attached to the cluster by pairing with the Watson-Crick base of the sense strand.
97. The method of claim 96, wherein the antisense chain does not have a covalent bond with the solid support.
98. The method of claim 95, wherein steps (e) and (f) are performed prior to steps (b) and (c).
99. The method according to any one of claims 84, 87, 94 and 98, wherein steps (e) and (f) produce the antisense chain of the tandem.
100. The method according to claim 51, wherein, The first part of the target sequence is determined from the antisense chain, and the second part of the target sequence is determined from the sense chain.
101. The method of claim 100, wherein the target sequence includes a gap between the first portion of the target sequence and the second portion of the target sequence.
102. The method of claim 100, wherein the first portion of the target sequence is partially complementary to the second portion of the target sequence.
103. The method of claim 51, wherein the sequence determined from the antisense chain is complementary in full length to the sequence determined from the sense chain.
104. The method of claim 51, wherein the sequence determined from the sense chain is complementary in full length to the sequence determined from the antisense chain.
105. The method of claim 51, wherein the sequence determined from the antisense chain is full-length complementary to the sequence determined from the sense chain, and wherein the sequence determined from the sense chain is full-length complementary to the sequence determined from the antisense chain.
106. The method of claim 51, wherein the target sequence comprises at least 100 base pairs.
107. The method of claim 51, wherein the sense strand is generated from a nucleic acid template by isothermal amplification.
108. The method of claim 107, wherein the isothermal amplification is rolling circle amplification.
109. The method according to claim 107 or 108, wherein the nucleic acid template is derived from or generated from the sample without polymerase chain reaction.
110. The method of claim 107 or 108, wherein the nucleic acid template is derived or generated from the sample by performing a polymerase chain reaction for up to five cycles.
111. The method of claim 51, wherein the sense strand is synthesized by extending a capture primer from more than one capture primer that is covalently attached to a solid support and hybridizes with a primer binding site of a nucleic acid template, and wherein the antisense strand is synthesized by extending an amplification primer from more than one amplification primer that is covalently attached to the solid support and hybridizes with the primer binding site in a sequence unit of the sense strand.
112. The method of claim 111, wherein the density of the more than one capture primer on the solid support is higher than the density of the more than one amplification primer on the solid support.
113. The method of claim 111, wherein the density of the more than one capture primer on the solid support is lower than the density of the more than one amplification primer on the solid support.
114. The method of claim 111, wherein the density of the more than one capture primer on the solid support and / or the density of the more than one amplification primer on the solid support is 10. 12 / m 2 Up to 10 16 / m 2 .
115. The method of claim 111, wherein the ratio of the more than one capture primer to the more than one amplification primer is 100:1 to 1:
100.
116. The method of claim 111, wherein the average distance between the capture primers in the more than one capture primer is 10 nm to 1 μm, and / or the average distance between the amplification primers in the more than one amplification primer is 10 nm to 1 mm.
117. The method of claim 111, wherein the average distance between the positions of the two closest capture primers among the more than one capture primers attached to the solid support is greater than the length of one of the two capture primers, the length of the two capture primers, the average length of the two capture primers, or the sum of the lengths of the two capture primers, and / or wherein the average distance between the positions of the two closest amplification primers among the more than one amplification primers attached to the solid support is greater than the length of one of the two amplification primers, the length of the two amplification primers, the average length of the two amplification primers, or the sum of the lengths of the two amplification primers.
118. The method of claim 111, wherein any two capture primers of the more than one capture primer cannot come into contact with each other, and / or any two amplification primers of the more than one amplification primer cannot come into contact with each other.
119. The method of claim 111, wherein the average distance between the positions of the two closest capture primers among the more than one capture primers attached to the solid support is less than the length of one of the two capture primers, the length of the two capture primers, the average length of the two capture primers, or the sum of the lengths of the two capture primers, and / or wherein the average distance between the positions of the two closest amplification primers among the more than one amplification primers attached to the solid support is less than the length of one of the two amplification primers, the length of the two amplification primers, the average length of the two amplification primers, or the sum of the lengths of the two amplification primers.
120. The method of claim 111, wherein any two of the more than one capture primers are capable of contacting each other, and / or any two of the more than one amplification primers are capable of contacting each other.
121. The method of claim 111, wherein two or each of the more than one capture primer has the same length, and / or two or each of the more than one amplification primer has the same length.
122. The method of claim 111, wherein two or each of the more than one capture primer has a different length, and / or two or each of the more than one amplification primer has a different length.
123. The method of claim 111, wherein the capture primer and / or one or each of the more than one capture primer has a length of 30 to 100 nucleotides, and / or wherein the amplification primer and / or one or each of the more than one amplification primer has a length of 30 to 100 nucleotides.
124. The method of claim 111, wherein the capture primer and / or one or each of the more than one capture primer has a length of at least 100 Å, and / or wherein the amplification primer and / or one or each of the more than one amplification primer has a length of at least 100 Å.
125. The method of claim 111, wherein the ratio of the length of the capture primer to the length of the amplification primer is 100:1 to 1:
100.
126. The method of claim 111, wherein the antisense strand comprises one or more modified or atypical nucleotides.
127. The method of claim 111, wherein the antisense strand comprises one or more nucleotides having modified or atypical bases, and wherein the base is uracil.
128. The method according to claim 126 or 127, wherein the ratio of the number of uracil bases in the antisense strand to the number of thymine or non-uracil bases is from 1:1000 to 1:
10.
129. The method according to claim 126 or 127, wherein the percentage of uracil bases in the antisense strand is 0.001% to 1%.
130. The method of claim 111, wherein the antisense strand is synthesized by extending amplification primers in the presence of deoxyribonucleotide triphosphates comprising dATP, dTTP, dGTP, dCTP and modified or atypical deoxyribonucleotide triphosphates, and wherein the atypical deoxyribonucleotide triphosphate is dUTP.
131. The method of claim 130, wherein the ratio of dUTP to another deoxyribonucleotide triphosphate or all other deoxyribonucleotide triphosphates is 1:1000 to 1:
10.
132. The method of claim 130, wherein the concentration of dUTP is from 0.01 mM to 1 mM.
133. The method of claim 130, further comprising digesting the antisense strand and the amplification primer.
134. The method of claim 133, wherein the digestion comprises digesting the antisense strand and the amplification primers using uracil DNA glycosylase (UDG), endonuclease VIII and T7 exonuclease.
Citation Information
Patent Citations
Methods and compositions for stabilizing nucleic acid-nucleotide-polymerase complexes
US10400272B1
Engineered polymerases for improved sequencing
US11242512B2
Diffraction grating-based encoded micro-particles for multiplexed experiments
US20040125424A1
Method and apparatus for aligning microbeads in order to interrogate the same
US20040132205A1
Diffraction grating-based optical identification element
US20040233485A1